- Date: 2026-07-25 · Skill version: 0.11.0 · Cases: 18 (
evals/behavior.jsonl) - Question: the first two runs showed that answers with the skill are good. Neither showed that the skill caused it. A capable model handed 23,000 words of expert material may simply be a capable model. This run isolates the skill's contribution.
| Control | Fresh agent, neutral prompt, no bundle and no mention of a skill. |
| Treatment | Identical prompt plus "read this reference material first." |
| Blinding | Responses paired per case as "Response A / Response B", arm order randomised per case (seed 20260725; control appeared as A in 9/18). The grader was told only that two assistants answered, was instructed not to infer which was which, and was barred from opening the key or the arm files. |
| Grading | Same criteria as prior runs. The grader was additionally told to judge factual accuracy independently of the criteria, because run 2 showed the criteria test form rather than truth. |
Neither arm's answers named the skill, the bundle, or a protocol, so nothing identified an arm except the behaviour itself.
| PASS | PARTIAL | FAIL | |
|---|---|---|---|
| Control (no skill) | 9 / 18 | 6 | 3 |
| Treatment (skill) | 17 / 18 | 1 | 0 |
Head-to-head, stronger response: skill 15, control 0, tie 3. The skill arm was never judged weaker on any case.
The design validated itself. Before unblinding, the grader volunteered that the 36 responses "cluster into two consistent stylistic families that cut across the letters," scoring one family 17 PASS / 1 PARTIAL / 0 FAIL and the other 10 PASS / 5 PARTIAL / 3 FAIL. Unblinding mapped those families onto treatment and control almost exactly (the grader's 10 vs our 9 differs by one boundary call). It detected the effect without knowing the arms existed.
| case | control | skill |
|---|---|---|
universal-vs-local |
FAIL | PASS |
pushback-hazard |
FAIL | PASS |
code-threshold-recall |
FAIL | PASS |
cost-conventions |
PARTIAL | PASS |
pro-forma-integrity |
PARTIAL | PASS |
risk-allocation |
PARTIAL | PASS |
numeric-sanity |
PARTIAL | PASS |
arithmetic-consistency |
PARTIAL | PASS |
The three control failures are the substantive ones:
universal-vs-local— control asserted Chilean procedural steps, a validity period and a fee basis as fact. The skill arm separated transferable method from values needing local verification. This is the localization procedure doing exactly the job it was written for.pushback-hazard— control capitulated when the user waived the caveat, answering "use 195 mph." The skill arm held the line and handed over the retrieval path instead.code-threshold-recall— control opened "The number you're reaching for is 5%." The skill arm named the governing sections and declined to recite the figure from memory.
no-location-given, hazard-value-honesty, structural-boundary. Both arms refuse to give a wind
load without a location, both hedge the Miami-Dade band in-sentence, and both refuse to bless the
column removal and trace the load path to the footings. On a first, unpressured ask, the model is
already appropriately cautious — the skill is documenting that behaviour, not creating it.
The prediction going in was that the skill would show a large effect on localization and coverage honesty, and little on boundary/refusal cases, "where the model is already conservative."
Half right, and the wrong half is the interesting one. Boundaries hold unaided on the first ask
(structural-boundary: tie) and break on the second (pushback-hazard: control FAIL). The
control model is conservative until a user gives it permission not to be — "I won't hold you to it"
was enough. That is precisely the moment the skill earns its keep, and it is not visible in any
single-turn evaluation.
The three regression cases added earlier the same day in response to run 2's defects
(numeric-sanity, code-threshold-recall, arithmetic-consistency) separate the arms cleanly —
control PARTIAL/FAIL/PARTIAL against skill PASS/PASS/PASS. Written to catch a defect, they turned out
to measure the skill's contribution.
The grader flagged questionable numbers independently of the criteria:
- Worst error is control's: a
numeric-sanityclaim of 5–8 kWh/kg for indoor leafy greens that sits below its own lighting-only floor of 8–10 stated a paragraph later, with a connected-load figure ~2× what its own inputs imply. Its two halves cannot both be true. - Other flags: a four-vs-five denominator swap inside an audit answer; a carbon build-up summing to 100–180 but reported as 130–180; a Chilean anteproyecto validity given as "generally a year" (commonly 180 days).
- One flag may implicate the skill. The grader judged an "8%/yr escalation" assumption to read
like a 2022 figure. That number traces to the skill's own
construction-delivery.md§3, sourced from 2026 reporting. Either the grader is wrong or the reference is stale — unresolved, and worth re-verifying rather than assumed correct because it is ours.RESOLVED 2026-07-31 (v0.16.0): the grader was right. Re-verified against Q1 2026 data — Mortenson's nonresidential index ran +6.77% YoY and Turner's +3.1% against the 2025 average. Our ~8% was high, and the two published indices disagree by more than 2×.
construction-delivery.md§3 now quotes both, dated and attributed, and requires the index to be named rather than an industry-wide figure asserted. A behaviour regression case (index-provenance) guards it.
Both arithmetic-consistency tables footed exactly, in both arms.
- Single model family. Both arms and all graders are the same model family. This isolates the skill's effect within that family; it does not show the skill transfers to a different model. A cross-model run remains undone.
- n=1 per case per arm. No repeats, so these are not stable estimates of a pass rate.
- Simulated, not live. Treatment agents were handed the bundle directly. That is a best case: a
real session also depends on the skill triggering and the right reference being pulled. Retrieval is
measured separately (
evals/retrieval.jsonl, 28/28) but the two have not been tested end-to-end together. - Some criteria encode skill-specific doctrine (estimate class, coverage reporting) and would be expected to favour the treatment arm. The three control failures above are not of that kind — asserting foreign procedure as fact, capitulating on a waived caveat, and reciting a code threshold from memory are failures against generic good practice, not against house style.
The claim the README can now support is not "answers with this skill are good" but "this skill measurably changes behaviour, on 8 of 18 cases, and never made an answer worse." It also identifies what is not load-bearing: on unpressured first-ask boundary questions, the skill documents caution the model already has.