Skip to content

feat(evals): record judge identity and refuse cross-identity comparisons #8573

Description

@sudoKrishna

Problem

LLM-as-judge scores are only comparable when the evaluator is the same. The
judge currently returns { scores, rationale, weightedScore, passed } with no
record of who judged or how. If the judge model, rubric, prompt, parser, or
decoding settings change between a baseline and a candidate, a score delta
(e.g. 0.71 → 0.79) can be read as an improvement when it is really a change of
evaluator. The recorded-transcript replay makes the transport deterministic; it
does not make two differently-judged scores comparable.

Proposal

Persist a small judge-identity envelope with every verdict and fail closed on
comparative claims when a material field differs.

  • JudgeVerdict gains an identity field: judge provider/model, rubric digest,
    parser version, decoding settings (temperature), and the scenario id.
  • The judge check / report surfaces the judge model and rubric digest.
  • A comparison helper classifies a baseline→candidate delta as comparable or
    insufficient-evidence when any material identity field differs.
  • Add one small, frozen, human-adjudicated calibration slice so judge drift can
    be detected rather than treating the judge as ground truth.

Scope

Follow-up to the judge scorer. Keep it small: identity + digest + a comparison
guard + one calibration fixture. Do not build a general evaluation platform.

Acceptance criteria

  • Every verdict carries a stable judge-identity envelope
  • The report shows the judge model and rubric digest
  • A delta across different judge identity is reported as insufficient
    evidence, not as an improvement
  • A frozen calibration slice with known expected scores exists
  • Documented how to refresh the calibration slice

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions