Discussion about this post

User's avatar
Foreman's avatar

validate-evaluator handles LLM judges calibrating against human labels: TPR/TNR, bias correction, all of it. But eval-audit is itself an evaluator, and the new-user case it targets has no human-label ground truth to calibrate the audit against. Two pipelines can both score “high severity” on judge validation.

The user who needs eval-audit most is the user least equipped to notice the gap. Is the implicit answer “trust the priors until you’ve labelled enough of your own data to recalibrate”, or is there something tighter in the meta-skill I missed?

No posts

Ready for more?