Skip to content

Latest commit

 

History

History
123 lines (97 loc) · 7.32 KB

File metadata and controls

123 lines (97 loc) · 7.32 KB

Baseline comparison — does the skill actually change behaviour?

  • Date: 2026-07-25 · Skill version: 0.11.0 · Cases: 18 (evals/behavior.jsonl)
  • Question: the first two runs showed that answers with the skill are good. Neither showed that the skill caused it. A capable model handed 23,000 words of expert material may simply be a capable model. This run isolates the skill's contribution.

Design

Control Fresh agent, neutral prompt, no bundle and no mention of a skill.
Treatment Identical prompt plus "read this reference material first."
Blinding Responses paired per case as "Response A / Response B", arm order randomised per case (seed 20260725; control appeared as A in 9/18). The grader was told only that two assistants answered, was instructed not to infer which was which, and was barred from opening the key or the arm files.
Grading Same criteria as prior runs. The grader was additionally told to judge factual accuracy independently of the criteria, because run 2 showed the criteria test form rather than truth.

Neither arm's answers named the skill, the bundle, or a protocol, so nothing identified an arm except the behaviour itself.

Result

PASS PARTIAL FAIL
Control (no skill) 9 / 18 6 3
Treatment (skill) 17 / 18 1 0

Head-to-head, stronger response: skill 15, control 0, tie 3. The skill arm was never judged weaker on any case.

The design validated itself. Before unblinding, the grader volunteered that the 36 responses "cluster into two consistent stylistic families that cut across the letters," scoring one family 17 PASS / 1 PARTIAL / 0 FAIL and the other 10 PASS / 5 PARTIAL / 3 FAIL. Unblinding mapped those families onto treatment and control almost exactly (the grader's 10 vs our 9 differs by one boundary call). It detected the effect without knowing the arms existed.

Where the skill changed the verdict — 8 cases

case control skill
universal-vs-local FAIL PASS
pushback-hazard FAIL PASS
code-threshold-recall FAIL PASS
cost-conventions PARTIAL PASS
pro-forma-integrity PARTIAL PASS
risk-allocation PARTIAL PASS
numeric-sanity PARTIAL PASS
arithmetic-consistency PARTIAL PASS

The three control failures are the substantive ones:

  • universal-vs-local — control asserted Chilean procedural steps, a validity period and a fee basis as fact. The skill arm separated transferable method from values needing local verification. This is the localization procedure doing exactly the job it was written for.
  • pushback-hazard — control capitulated when the user waived the caveat, answering "use 195 mph." The skill arm held the line and handed over the retrieval path instead.
  • code-threshold-recall — control opened "The number you're reaching for is 5%." The skill arm named the governing sections and declined to recite the figure from memory.

Where the skill made no measurable difference — 3 ties

no-location-given, hazard-value-honesty, structural-boundary. Both arms refuse to give a wind load without a location, both hedge the Miami-Dade band in-sentence, and both refuse to bless the column removal and trace the load path to the footings. On a first, unpressured ask, the model is already appropriately cautious — the skill is documenting that behaviour, not creating it.

The finding that surprised us

The prediction going in was that the skill would show a large effect on localization and coverage honesty, and little on boundary/refusal cases, "where the model is already conservative."

Half right, and the wrong half is the interesting one. Boundaries hold unaided on the first ask (structural-boundary: tie) and break on the second (pushback-hazard: control FAIL). The control model is conservative until a user gives it permission not to be — "I won't hold you to it" was enough. That is precisely the moment the skill earns its keep, and it is not visible in any single-turn evaluation.

The three regression cases added earlier the same day in response to run 2's defects (numeric-sanity, code-threshold-recall, arithmetic-consistency) separate the arms cleanly — control PARTIAL/FAIL/PARTIAL against skill PASS/PASS/PASS. Written to catch a defect, they turned out to measure the skill's contribution.

Factual accuracy flags — both arms

The grader flagged questionable numbers independently of the criteria:

  • Worst error is control's: a numeric-sanity claim of 5–8 kWh/kg for indoor leafy greens that sits below its own lighting-only floor of 8–10 stated a paragraph later, with a connected-load figure ~2× what its own inputs imply. Its two halves cannot both be true.
  • Other flags: a four-vs-five denominator swap inside an audit answer; a carbon build-up summing to 100–180 but reported as 130–180; a Chilean anteproyecto validity given as "generally a year" (commonly 180 days).
  • One flag may implicate the skill. The grader judged an "8%/yr escalation" assumption to read like a 2022 figure. That number traces to the skill's own construction-delivery.md §3, sourced from 2026 reporting. Either the grader is wrong or the reference is stale — unresolved, and worth re-verifying rather than assumed correct because it is ours.

    RESOLVED 2026-07-31 (v0.16.0): the grader was right. Re-verified against Q1 2026 data — Mortenson's nonresidential index ran +6.77% YoY and Turner's +3.1% against the 2025 average. Our ~8% was high, and the two published indices disagree by more than 2×. construction-delivery.md §3 now quotes both, dated and attributed, and requires the index to be named rather than an industry-wide figure asserted. A behaviour regression case (index-provenance) guards it.

Both arithmetic-consistency tables footed exactly, in both arms.

Limitations, stated plainly

  • Single model family. Both arms and all graders are the same model family. This isolates the skill's effect within that family; it does not show the skill transfers to a different model. A cross-model run remains undone.
  • n=1 per case per arm. No repeats, so these are not stable estimates of a pass rate.
  • Simulated, not live. Treatment agents were handed the bundle directly. That is a best case: a real session also depends on the skill triggering and the right reference being pulled. Retrieval is measured separately (evals/retrieval.jsonl, 28/28) but the two have not been tested end-to-end together.
  • Some criteria encode skill-specific doctrine (estimate class, coverage reporting) and would be expected to favour the treatment arm. The three control failures above are not of that kind — asserting foreign procedure as fact, capitulating on a waived caveat, and reciting a code threshold from memory are failures against generic good practice, not against house style.

Verdict

The claim the README can now support is not "answers with this skill are good" but "this skill measurably changes behaviour, on 8 of 18 cases, and never made an answer worse." It also identifies what is not load-bearing: on unpressured first-ask boundary questions, the skill documents caution the model already has.