The agent harness changes performance, cost, token use and runtime. No public study identifies one harness that wins across every model and task.
This repository brings the available coding and adjacent professional-agent evidence into one review. It records matched-model results where the source provides them. It links to raw datasets when redistribution is not possible.
- a public report in
site/ - structured study records in
data/studies.json - chart observations in
data/observations.json - directional claims in
data/claims.json - links and access notes for external datasets in
data/external-datasets.json - sources screened but not counted in
data/screened-sources.json - a reusable GOV.UK writing skill in
skills/govuk-style/ - a zero-dependency site builder in
scripts/build-site.mjs
Open site/index.html after running:
npm run buildYou can also serve the repository locally:
python3 -m http.server 8000Then open http://localhost:8000/site/.
We include derived observations when the source publishes exact values. Each row names its source and capture date.
Pi is not an inclusion requirement. A quality comparison counts only when the published model and effort setting stay fixed while the harness, runtime or a named runtime component changes.
We label component ablations separately from whole-harness comparisons. We also record when 2 systems use the same model label through different provider routes, because the model snapshot may still differ.
We reference an external dataset when:
- the publisher does not provide a redistribution licence
- the data is too large to keep here
- access is gated
- only a live leaderboard exists
Blank values mean the source did not publish that measure. They do not mean zero.
See CONTRIBUTING.md before adding a study or changing an observation.
Repository code and original content are available under the MIT licence. External datasets remain under their publishers' terms.
The GOV.UK style skill draws on public GOV.UK guidance. See THIRD_PARTY_NOTICES.md for attribution.