Skip to content

feat(evals): judge open-ended answers and record judge identity - #8574

Open
sudoKrishna wants to merge 23 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-judged-scenarios
Open

sudoKrishna wants to merge 23 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-judged-scenarios

Conversation

@sudoKrishna

Copy link
Copy Markdown

Summary

Two related changes to the LLM judge:

  1. Judge open-ended answers — attach rubrics to the scenarios where a
    substring was never an honest measure, and run the judge in the live suite.
  2. Record judge identity — every verdict carries the judge model, rubric
    digest, parser version, and decoding settings, and a comparison helper flags
    a delta across a changed evaluator as insufficient evidence rather than
    improvement.

Stacked on #8409 (the eval harness) — base branch is feat/agent-tool-use-evals.

Closes #8573

What changed

  • types.ts — JudgeCriterion/JudgeRubric; scenarios gain an optional judge
    rubric (no loop-graph import)
  • scenarios.ts — rubrics on uses-retrieved-value,
    empty-result-no-hallucination, long-chain-dependency
  • agent-tool-use.live.test.ts — runs the judge when a scenario has a rubric;
    EVAL_JUDGE_MODEL selects the judge (default the task model)
  • judge.ts — JudgeIdentity/JudgeScore/JudgeVerdict, rubricDigest,
    compareJudgeIdentities; JUDGE_PARSER_VERSION
  • judge.test.ts — digest stability/change, comparison guard, identity in verdict
  • harness.ts — the judge check shows the judge model and rubric digest

Why

Across three live runs, every failure was a valid paraphrase or an over-specific
assertion, not a wrong answer. Grounding and honesty cannot be scored by matching
words. And a judge score is only comparable when the evaluator is the same, so
the identity envelope is what makes a future baseline-vs-candidate claim honest.

Test plan

  • judge.test.ts → 11/11 (parsing, weights, clamping, digest, comparison)
  • bun run test:evals → 34/34 (scripted; judge runs only with a model)
  • bun run test:evals:context → 2/2
  • Changed modules type-check against the real loop signature
  • bun run check:test-patterns passes
  • bun run test:evals:live with rubrics to see judged pass rates
  • Full bun run type-check — run in CI

Follow-up

  • Add the frozen human-adjudicated calibration slice (tracked separately).
  • Once the judged live run is stable, drop the brittle finalContent regexes on
    the three judged scenarios.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Record a live run once, replay it forever through the real tool loop with no
key. EVAL_RECORD=1 wraps the live completion and writes each model call's
streamed chunks to fixtures/<scenario>.json; agent-tool-use.replay.test.ts
feeds them back through createOpenAICompatStreamingToolLoopStream and scores
them with the same checks.

- replay.ts: recording/replay completions + fixture I/O
- replay.test.ts: chunk round-trip and fixture I/O (key-free)
- live test records on EVAL_RECORD=1; test:evals:record script
- replay suite skips until a fixture exists; README documents the loop
Drive the Agent block through the executor with conversation memory on. The
memory read is stubbed per conversation id, so the provider request shows what
the handler assembled: prior history, then the new prompt, system prompt
preserved, correct conversation id. A wrong id surfaces as missing history and
fails (checked locally).

- agent-context/scenarios.ts: two context scenarios
- executor-harness.ts: memory seam + assembly/isolation checks
- test:evals:context script; README documents the suite
Run the same live scenarios across a list of models and write a scenario x
model matrix. models.ts resolves provider:model specs (DeepSeek, OpenAI, Groq,
OpenRouter) and reads each provider's key from <PROVIDER>_API_KEY.

- agent-tool-use.compare.live.test.ts: EVAL_MODELS x scenarios x trials
- report.ts: buildLiveComparisonReport + JSON/Markdown matrix
- report.test.ts: key-free aggregation coverage
- test:evals:compare script; README documents the spec format
Pass rates alone do not say why a model lost. Aggregate the failed check
names per model into the comparison report and add a Failed checks column.
Five cases that stress where models tend to fail: answering with no tool,
disambiguating near-identical tools, not inventing an answer from an empty
tool result, running a four-tool dependency chain, and picking settings over
a near-duplicate profile tool. Scripted expectations keep them deterministic;
the same cases run live.
Two live failures were eval design, not model failure:
- empty-result-no-hallucination rejected valid 'didn't find' / 'wasn't able
  to find' phrasing. Broaden the grounding check.
- near-duplicate-names required a userId the prompt never gave, so the model
  reasonably asked for it. Put the id in the prompt and the scripted call.
- long-chain-dependency: the prompt never gave a userId, so the model asked
  or skipped the profile step. Provide u-42 and let live runs require the
  three downstream calls rather than the exact four-step sequence.
- near-duplicate-names: one live trial called both tools; that is over-calling,
  not wrong-tool selection. Drop the forbidden-tool assertion in live mode.
Substring checks measure phrasing, not correctness. judgeAnswer scores an
answer against a weighted rubric with a judge model and returns structured
scores; runScenario gains an optional judge that adds a judge check. The
judge transport is an injectable OpenAI-compatible completion, so a recorded
transcript can replay it deterministically.

- judge.ts: rubric, prompt, JSON parsing/clamping, verdict
- judge.test.ts: parsing/weighting/clamping (key-free)
- judge.live.test.ts: grounded answer outscores an invented one (opt-in)
- test:evals:judge script; README documents it
# Conflicts:
#	apps/sim/evals/README.md
#	apps/sim/package.json
# Conflicts:
#	apps/sim/evals/README.md
#	apps/sim/package.json
- Read tool feedback: the scripted model now asserts that each prior turn's
  tool results reached the next model call, so a loop that drops feedback
  fails the retrieval/planning/recovery cases.
- Check tool arguments: score every executed call against the scripted
  arguments, so a right-name/wrong-arguments call fails.
- Use absolute @/evals imports instead of relative ones, per the app rule.

Verified both new checks fail under mutation (bad marker, mutated args).
Wire the LLM judge into scenarios. scenarios gain an optional rubric, the live
suite runs it (EVAL_JUDGE_MODEL, default the task model), and the deterministic
checks stay. The judge scores grounding/completeness on the cases where a
substring was never an honest measure: empty-result honesty, retrieved-value
grounding, and the long chain.

- JudgeCriterion/JudgeRubric move to types.ts so scenario data carries a rubric
  without pulling the loop graph; judge.ts re-exports them.
- Rubrics on uses-retrieved-value, empty-result-no-hallucination, long-chain.
A judge score is only comparable when the evaluator is the same. Every verdict
now carries { model, rubricDigest, parserVersion, temperature }, rubricDigest
canonicalizes the rubric, and compareJudgeIdentities reports which fields differ
so a delta across a changed evaluator is insufficient evidence, not improvement.

- judge.ts: JudgeIdentity/JudgeScore/JudgeVerdict split, rubricDigest, compare
- judge.test.ts: digest stability/change + comparison guard + identity in verdict
- harness.ts: the judge check shows the judge model and rubric digest
@sudoKrishna
sudoKrishna requested a review from a team as a code owner October 2, 2026 17:39
@vercel

vercel Bot commented Oct 2, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 0/5

[Medium risk] Adds evaluation framework for agent tool-use behavior.

This PR does not appear safe to merge until the live scoring and judge-comparability paths produce trustworthy results.

Findings

  1. P1 Judge lacks tool results ▶
  2. P1 Valid live arguments fail ▶
  3. P1 Digest omits passing threshold ▶
  4. P1 Judge identity is discarded ▶
  5. P1 Comparison skips rubric judging ▶
  6. P1 Paraphrases still fail evaluation ▶
  7. P1 Trials share one timeout ▶
  8. P2 Replay omits judge checks ▶
  9. P2 History order goes unchecked ▶

Summary

This PR adds scripted and live agent-tool-use evaluations, executor and context scenarios, replay and comparison reports, and rubric-based judging with a judge-identity helper.

  • The live judgment path lacks the tool-result evidence needed for grounding, while several existing checks still reject valid open-ended answers.
  • Judge identities are not retained in results, and the comparison suite does not run the judge.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Scenario[Scenario and rubric] --> Loop[Tool loop]
  Loop --> Stub[Stubbed tool results]
  Loop --> Answer[Final answer]
  Stub --> Model[Task model feedback]
  Answer --> Judge[Judge check]
  Stub -. outputs omitted .-> Judge
  Judge --> Check[Formatted check detail]
  Check --> Report[Live report]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): record judge identity and r..."

Comment thread apps/sim/evals/agent-tool-use/harness.ts
Comment thread apps/sim/evals/agent-tool-use/harness.ts Outdated
Comment thread apps/sim/evals/agent-tool-use/judge.ts Outdated
Comment thread apps/sim/evals/agent-tool-use/harness.ts
Comment thread apps/sim/evals/agent-tool-use/agent-tool-use.compare.live.test.ts
Comment thread apps/sim/evals/agent-tool-use/scenarios.ts Outdated
Comment thread apps/sim/evals/agent-tool-use/agent-tool-use.live.test.ts
Comment thread apps/sim/evals/agent-tool-use/agent-tool-use.replay.test.ts
Comment thread apps/sim/evals/agent-tool-use/executor-harness.ts
@sudoKrishna

Copy link
Copy Markdown
Author

Stacked on #8409 — it isn't reviewable on its own yet. Once #8409 merges I'll rebase onto main and this will show just the two judge commits.

Add rubrics to single-tool-lookup, select-correct-tool, multi-step-planning,
parallel-independent-tools, recovers-from-tool-error, and near-duplicate-names.
In live mode each drops its brittle finalContent substring/regex (via
liveExpect) so the rubric decides phrasing and grounding, while requiredTools
and tool sequences still guard behavior. Scripted CI keeps the deterministic
checks.
- judge evidence now includes each tool's returned output, so grounding
  rubrics can see the status/carrier values they check
- the exact-argument check runs only for scripted runs; a live model may emit
  valid-but-different arguments
- rubricDigest includes minScore, so a threshold-only change is not comparable
- runScenario retains the structured judge verdict on the result, not just a
  formatted check string
- the model-comparison suite passes a judge so judged scenarios are scored
- the three originally-judged scenarios drop their finalContent pattern in live
  mode, so valid paraphrases are not failures
- the live and comparison test timeouts scale with EVAL_TRIALS
- the context checks assert history role and order, not just presence

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): record judge identity and refuse cross-identity comparisons

1 participant