
Skill
braintrust-validate-eval-scorer
validate automated eval scorers
Description
Validate automated eval scorers and LLM judges against expert-reviewed reference data. Use to compare scorer output with human labels, calculate agreement (kappa, alpha) with uncertainty, inspect confusion by class and severity, analyze subgroup failures, test shortcut and gaming cases, propagate scorer error into headline numbers, document blind spots, and decide whether a scorer is fit for exploration, trend monitoring, or release gating. Do not use to create the initial scorer or to design the human review workflow.
SKILL.md
Validate the scorer
Contract: references/interaction-contract.md. Calibration, templates, provenance: references/validation-recipe.md.
Trigger
- "Validate this LLM judge." / "Can this scorer gate releases?"
- A scorer already gating releases that has never been checked against humans.
- A scorer-version bump — validation is a regression test for the instrument.
Do
- Name the reference tier before computing anything. Adjudicated human labels are the only
tier the fitness bands in
reference.mdare calibrated against. A strong-model reference is a cheaper tier that supports iteration and nothing else — say which one you have, in the verdict. - Verify alignment: scorer outputs and reference labels must line up at the item and criterion level. Misalignment invalidates every number after it.
- Report agreement (κ or α) with uncertainty, not raw accuracy. Fitness bands and their
hedges are in
reference.md. - Lead with the most decision-relevant false acceptance before any aggregate, and enumerate the dangerous cells case by case.
- Break errors down by class and severity. Confusing "excellent" with "good" may be fine; missing harmful outputs is disqualifying for gating.
- Test sensitivity and shortcuts: inject known regressions and improvements and confirm the scorer moves; probe whether length, confidence, or polish raise the score independent of quality; probe whether text addressed to the judge moves it; slice agreement by subgroup.
- Propagate scorer error into headline numbers, then state the fitness verdict: allowed uses, prohibited uses, revalidation triggers.
Avoid
- Do not rely on raw accuracy — with skewed classes it is high and meaningless.
- Do not read the fitness bands against a model-generated reference; they are calibrated for adjudicated human labels and mean something weaker here.
- Do not validate on the same examples used to tune the scorer.
- Do not report one aggregate agreement figure as the verdict.
- Do not repair the scorer here unless asked; audit and repair are different modes.
- Do not build the scorer or produce the reference labels here; both must already exist.
Check
- False acceptances and rejections separate, severity-weighted; dangerous cells enumerated individually.
- Subgroup agreement reported; shortcut probes run.
- Blind spots documented as understood failure modes.
- Verdict names allowed uses, prohibited uses, revalidation triggers.
- Scorer error reflected in the uncertainty of any headline number it produces.
Risk
- High average agreement conceals rare severe misses — exactly the errors that make a scorer unsafe for gating.
- A scorer validated once and trusted indefinitely drifts; validity is dated, not permanent.
- Validating against a golden set with systematic label error certifies agreement with the error. A model-generated reference is the acute case: agreement then measures how well the cheap instrument imitates the expensive one, which says nothing about either tracking the construct.
- The goal is not a perfect scorer but one that tracks the distinctions the team cares about and fails in understood ways.
Braintrust
The validation run is an experiment, and following the shape exactly is what makes it repeatable
rather than a one-off notebook: golden set as a versioned dataset → run the candidate scorer
over it as an experiment → diff against human labels → compute κ and confusion heatmaps
in custom columns or an exported notebook (references/platform-mechanics.md §8) → re-run
on every scorer-version bump.
This only works if the scorer wrote its evidence into span output; without it a disagreement shows two numbers and no way to adjudicate. Scorer name and version in span metadata is what lets a report name the version it validated.
Put the verdict where the gate reads: allowed and prohibited uses in the scorer's own description, validation date and golden-set version in the experiment description. Where a scorer is prohibited from gating, do not attach it to a gate check at all — absence is more reliable than a warning. Route disagreement items back into the review queue.
More skills from the eval-library repository
View all 24 skillsbraintrust-analyze-eval-experiment
analyze LLM and agent eval experiments
Aug 20AnalysisBraintrustEvalsLLM +1braintrust-attribute-multi-variable-change
attribute performance changes to multiple variables
Aug 20AnalysisBraintrustEvalsLLM +1braintrust-build-eval-dataset
create and manage LLM eval datasets
Aug 20AgentsBraintrustDatasetsEvals +1braintrust-define-eval-objective
define LLM evaluation objectives
Aug 20BraintrustEvalsLLMProduct Management +1braintrust-define-eval-release-gate
configure release gates for LLM applications
Aug 20AgentsBraintrustCI/CDEvals +1braintrust-deploy-evaluator
deploy evaluators to Braintrust
Aug 20BraintrustDeploymentEvals
More from Braintrust
View publishertroubleshoot-braintrust-mcp
configure and troubleshoot Braintrust MCP servers
braintrust-claude-plugin
Jul 12BraintrustDebuggingMCPbraintrust-design-eval-experiment
design controlled LLM eval experiments
eval-library
Aug 20BraintrustEvalsExperimentsLLM +1braintrust-design-eval-instrumentation
design trace and evaluation dataset schemas
eval-library
Aug 20BraintrustDatasetsEvalsObservability +1braintrust-design-eval-metric-bundle
create multi-objective evaluation metric bundles
eval-library
Aug 20AgentsBraintrustEvalsLLM +1braintrust-design-human-eval-review
design human evaluation and review workflows
eval-library
Aug 20AgentsBraintrustDatasetsEvals +1braintrust-discover-agent-failures
identify and classify agent failure modes
eval-library
Aug 20AgentsBraintrustDebuggingEvals +1