Problem
LLM-as-judge scores are only comparable when the evaluator is the same. The
judge currently returns { scores, rationale, weightedScore, passed } with no
record of who judged or how. If the judge model, rubric, prompt, parser, or
decoding settings change between a baseline and a candidate, a score delta
(e.g. 0.71 → 0.79) can be read as an improvement when it is really a change of
evaluator. The recorded-transcript replay makes the transport deterministic; it
does not make two differently-judged scores comparable.
Proposal
Persist a small judge-identity envelope with every verdict and fail closed on
comparative claims when a material field differs.
JudgeVerdict gains an identity field: judge provider/model, rubric digest,
parser version, decoding settings (temperature), and the scenario id.
- The
judge check / report surfaces the judge model and rubric digest.
- A comparison helper classifies a baseline→candidate delta as
comparable or
insufficient-evidence when any material identity field differs.
- Add one small, frozen, human-adjudicated calibration slice so judge drift can
be detected rather than treating the judge as ground truth.
Scope
Follow-up to the judge scorer. Keep it small: identity + digest + a comparison
guard + one calibration fixture. Do not build a general evaluation platform.
Acceptance criteria
Problem
LLM-as-judge scores are only comparable when the evaluator is the same. The
judge currently returns
{ scores, rationale, weightedScore, passed }with norecord of who judged or how. If the judge model, rubric, prompt, parser, or
decoding settings change between a baseline and a candidate, a score delta
(e.g. 0.71 → 0.79) can be read as an improvement when it is really a change of
evaluator. The recorded-transcript replay makes the transport deterministic; it
does not make two differently-judged scores comparable.
Proposal
Persist a small judge-identity envelope with every verdict and fail closed on
comparative claims when a material field differs.
JudgeVerdictgains anidentityfield: judge provider/model, rubric digest,parser version, decoding settings (temperature), and the scenario id.
judgecheck / report surfaces the judge model and rubric digest.comparableorinsufficient-evidencewhen any material identity field differs.be detected rather than treating the judge as ground truth.
Scope
Follow-up to the judge scorer. Keep it small: identity + digest + a comparison
guard + one calibration fixture. Do not build a general evaluation platform.
Acceptance criteria
evidence, not as an improvement