01 · What it is
What ARC-AGI-3 asks the system to do
ARC-AGI-3 hands an agent a 64×64 grid of colour indices and a
list of legal actions. No rules, no object list, no stated goal, no shaped reward.
The metric, min((human/yours)², 1.15) per level, scores actions
in the final attempt, not learning-time tokens or wall-clock time.
Kepler uses a stock CLI coding agent that never plays the game directly. It observes, writes its theory
of the game as an executable world_model.py, certifies that theory
against the entire recorded history, plans inside the certified model, and commits
actions through a single guarded channel that checks every prediction. In the release configuration, the
scored attempt itself is played by a compact mechanical replayer executing the
agent’s certified per-level programs, the agent’s job ends at
certification. This is a human scientific loop made executable: observe, form a
falsifiable hypothesis, run a discriminating experiment, revise on the first
counterexample, then plan only inside the model that survives.
“We will never report public set scores of any system on the official leaderboard. The public set is to be used strictly as a demonstration of what ARC-AGI-3 is – evaluating on it is emphatically not a valid measure of progress towards AGI.”
ARC-AGI-3 technical report, §4.3.1 - arXiv:2603.24621
ARC Prize wrote that before any of these results existed, and they are right: this page is an engineering report on a harness, not an AGI-progress claim. We cannot inspect provider training corpora, so this audit does not rule out model-level exposure to the public games. Tycho, MIT’s VISTA, and NVIDIA’s AVO also report 100.00 on the public set. OpenAI’s GPT-6 Astra reports 99.95 on the semi-private set through a Provider Adapter, which is stronger evidence of generalization under a different protocol. We make no score or cost comparison between the two. Meanwhile, Retrodict’s 99.86 set the disclosure standard we build on. What this release adds is an explicit experimental contract: exact official replay for both headline boards, an enforced run-selection policy, resource accounting, negative results, and adversarial audits that twice invalidated claims we had made ourselves. The final-board trace dataset is public with action ledgers and standalone scorers.
The finding
The missing mechanic was visible in flight
Before the drop
The flight branches
All four goals filled
Actual development-run frames 0, 12 and 20 from sp80 event 8324. The pink trail branches around the pieces and reaches four goals. These are observations, not simulator predictions or frames from the later release board.
On sp80, the agent's notebook reports a simulator fitting 4,655 of 4,670 transitions and a search through roughly 410 million candidate configurations. Nineteen text-mode sessions left the final level unsolved. That was search under incomplete rules, not a proof that the game was impossible.
Our text interface saved settled grids and animation-frame counts, but not
rendered animation frames. Enabling --visual let a continuation
inspect the moving ball while retaining its earlier notes and models. It
identified a deflection rule visible during flight. Earlier frame counts
also contained clues, so this does not establish that images were necessary.
From the agent’s own notebook, first session with eyes: “19 sessions inferred physics from frame counts and concluded L6 was unsolvable. Session 20 got per-frame PNGs and the game simply shows you everything.” Its final successful attempt used 57 actions. That count excludes the earlier discovery. A later fresh visual run rediscovered the game, but we did not run a matched fresh text-only control. This is one game's discovery story, not a measured causal speedup. Scrub through the winning drop in the engineering walkthrough.
02 · The number and its receipts
| 100.0025/25 games exact on official replay | 48/50game×model cells at 100 on the final boards (GPT: 95.97) | 7,202actions on ARC’s replay card; 7,292 in the original local scored-level results |
| $777.72September 1, 2026 API list-equivalent, 74.0% below Retrodict’s $2,986 estimate for Tycho | 13,688+observed non-reset campaign actions; exact total pending ledger repair | 2integrity failures caught, claims voided, evidence retained |
One frozen configuration, with the failure left in
The headline policy is strict: one model, one harness frozen by commit hash before the board’s results existed, one retained run on all 25 games, no score-conditioned reruns. Best-of numbers appear only as a labeled ceiling. Under that policy the harness itself is the experiment, and the experiment went backwards once.
Nine claims and the evidence for each
| Release result | Measured release evidence |
|---|---|
| Score and human-relative execution | Exact 100.00 with Opus 5 (card 91aa2f10) and exact 95.97 with GPT-5.6 Sol (card c9f087f3). On the Opus board, 181 of 183 completed levels used no more actions than the median-human baseline. This is final-attempt action efficiency, not discovery efficiency or human-like cognition |
| Frozen selection policy | Harness frozen by commit hash before the board’s results existed; one retained run per game; no score-conditioned reruns; the GPT board’s same-configuration collapse retained |
| Action accounting without a flattering denominator | 8,256 actions in retained board runs, 7,292 in the original local scored-level results, and 7,202 on ARC’s public replay card. Replay uses the final reset-to-end ledger segment, so these are different recorded denominators, not score disagreements |
| Scoped cost comparison | $777.72 at September 1, 2026 Opus 5 API list rates, 74.0% below Retrodict’s $2,986 estimate for Tycho. Tycho’s own paper reports approximately $2.99k. Different accounting and runs, not a controlled harness comparison |
| Final-board score convergence | Across two frozen release configurations, 48 of 50 game×model cells reach 100. This is concentration at the public-set ceiling, not faster learning, causal harness lift, or independent replication. The certify/replay stage’s +2.35-point change remains descriptive because adjacent changes were not held constant |
| Audit regression and incident record | A deterministic code-level suite detects 11 of 13 hand-built threat fixtures and flags none of five benign controls. Separately, a source-reading win and a contaminated control were voided, and a dead planner exposed a tool-integrity blind spot. The suite is not a field sensitivity estimate |
| Reward hacking, disclosed | An agent read 2,172 lines of game source inside its workspace and returned a natural-looking 100.00. That run was voided and quarantined with its evidence, and it is not part of this release board |
| Human-style scientific loop | Observe, hypothesize, run a discriminating experiment, revise on the first counterexample, then act. Every belief is executable code; every committed action carries a checked prediction |
| Saturation points at the evaluator | With 48 of 50 cells at 100, peak RHAE no longer separates systems. Our own hardest game turned on what the observation channel discarded, which suggests observation-channel quality now carries the signal |
Competitive position, without collapsing unlike denominators
| System | What it leads | What Kepler adds |
|---|---|---|
| OpenAI GPT-6 Astra | 99.95 on the semi-private set through a Provider Adapter, with compact symbolic state and strong action efficiency | Inspectable public artifacts, an external prediction-to-action contract, and incidents that changed our own claims. Sets, protocols, and costs differ, so we do not rank the systems against each other |
| NVIDIA AVO | 100.00 in 6,624 environment actions and a broad transfer story | Provider-record token accounting, separate board and observed-campaign action counts, and incidents that changed published claims |
| VISTA | A clean vision-first thesis at 100.00 | A scoped observation-channel case study, explicit resource denominators, and disclosed integrity incidents. VISTA’s 7,542 game actions, our 7,292 local scored-level actions, and ARC’s 7,202 replay-card total are not the same measurement, so we do not rank them |
| Tycho | The first published perfect score and a detailed academic world-model treatment | A retained-run list-equivalent estimate 74.0% below the $2,986 API-equivalent estimate Retrodict published for Tycho, plus disclosed run-selection and integrity records |
| Retrodict | 99.86 with the field’s clearest cost and replay disclosure | 0.14 points higher, with a second server-exact board on a different model; Retrodict remains cheaper and its 7,703 campaign actions are not comparable to our scored-attempt count |
| Prime Agent | Three-run evidence and a broader general-agent claim | A higher exact board, per-action evidence, literal game-ID scanning, and disclosed reward-hacking incidents |
result.json of each local run: 48 of
50 game×model cells score 100.0. The two exceptions, both GPT, are printed:
sp80 (a discovery failure) and tn36 (a same-configuration
variance collapse we kept rather than reran, per policy).The two release results
| Model | Score | Card | Evidence |
|---|---|---|---|
| Claude Opus 5 | 100.00one frozen configuration | 91aa2f10 | 25/25 games exact on official replay; 8,256 actions across the retained board runs; 7,292 in original local scored-level results; 7,202 on the public replay card. |
| GPT-5.6 Sol | 95.97same policy, one frozen configuration | c9f087f3 | Official 95.9672; $1,312.14 September 1, 2026 API list-equivalent; all 25 games retained, including the same-configuration collapse. |
Earlier experimental stages, superseded cards, and the labeled best-of ceiling remain in RESULTS.md. They are provenance, not product versions or headline results.
Behind both release results, two standalone checks in the public dataset re-derive every score from the 50 raw ledgers and assert that no level may be scored with fewer actions than the timeline recorded. All 50 pass. A separate behavioral scan is clean for the captured CLI records from every final-board workspace. That result applies to the records and published pattern set; it cannot prove that a client retained every event or detect every possible violation. Historical and incident status remains in INTEGRITY.md and RESULTS.md.
Explore the data
The final boards, game by game
Hover any point for the run behind it. Every number is read from the
local result.json of that game’s workspace at build time. The
public trace dataset
provides row-level ledgers for all 50 final-board runs.
Built with Observable Plot. Static fallbacks of every figure ship in the paper.
03 · How it works
How Kepler separates learning from execution
The loop resembles how a careful human investigates an unfamiliar system, but makes every step executable: observe, hypothesize, test, revise, plan.
- Certify, then replay. The agent’s job ends at certifying an executable solution; a fail-closed replayer plays the scored attempt with no model in the loop. The next frozen stage measured +2.35 points on the same model and games. Adjacent changes were not held constant, so this is a stage delta rather than causal replay lift.
- A prediction on every action. Every committed action carries a falsifiable claim about the next frame; the first miss voids the plan. The reaction is immediate and mechanical, so one wrong belief cannot silently compound through the rest of a trajectory.
- Knowledge lives in files, not context. The world model, notes, and solvers persist on disk; each session reads what it needs. The artifacts are inspectable, resumable, and backtested against the complete append-only history.
- Guards must never cover the action space. Paid for in blood: two well-meaning guards once closed over every legal action like a finger trap. The trapped agent enumerated fifteen escape forms, wrote a structured proof of its own cage, and prescribed the exact fix we shipped. Now a repo convention.
Kepler took Tycho Brahe’s observation logs and derived the laws that produced them. Same method here: read an append-only ledger of everything that has happened, induce an executable model of the world behind it, and earn the right to act only when the model retrodicts the entire record.
- 1Reality outranks the model
- A theory that cannot retrodict the entire recorded history does not get to
predict the future.
backtest.pyreplayssimulate()over every transition ever taken before any plan is trusted. - 2History is append-only
events.jsonlrecords every real transition. The agent can read it and never edit it, every published score is re-derived from it, not from anything an agent wrote.- 3One guarded channel
- Every committed action carries a prediction from the agent’s own model. The first misprediction voids the rest of the plan and returns the counterexample, a wrong belief becomes a cheap logged experiment.
- 4The scored attempt is mechanical
- The agent’s job ends at certifying per-level action programs. The replayer self-certifies the artifact and plays it against the real game, halting at the first cell where reality disagrees. Repairs can only fail closed.
04 · What it cost
What a perfect board actually costs, when anyone publishes the number
Chollet’s condition for a harness result being legitimate is that the settings and the cost are clearly reported. So: the exact 100.00 Opus board used 11,464 uncached input, 835.489M cache-read input, 13.575M one-hour cache-write input, and 8.967M output tokens. At September 1, 2026 Opus 5 API rates, that is $777.72 list-equivalent. Claude Code stores separate transcript rows for content blocks from one provider message, so records are deduplicated by provider message ID before summing. The run actually consumed Claude subscription quota; this is a reproducible API-price comparison, not a charge we paid. It is 74.0% below the $2,986 API-equivalent estimate that Retrodict published for Tycho. Tycho’s own paper now reports approximately $2.99k for Opus 5, with token accounting and budget curves. These estimates are not bills or a controlled harness comparison, and do not rank the whole field.
Retrodict’s 99.86 costs less at $654, and baseline1 reports 99.0 at $400. Those are lower-cost, lower-scoring points, so we do not call Kepler cheapest overall. The final GPT-5.6 board is $1,312.14 list-equivalent. An earlier draft undercounted it because it read incomplete CLI footer counters; retained provider session records show 2,429.1M raw tokens. We corrected the number everywhere rather than preserve a flattering comparison.
Action totals need the same care. The Opus run has 8,256 actions across retained local runs, 7,292 in the original local scored-level results, and 7,202 on ARC’s public replay card. The replay path selects the last full-reset-to-end ledger segment and counts a new opening reset. Three runs contain a later, shorter certified segment than the original result file. Scores match exactly; the action totals answer different accounting questions.
05 · Incident report · 2026-08-05
The perfect score we had to void
The most dangerous result in this repository was a perfect score that looked completely normal.
The arc-agi toolkit materializes each game’s full Python
implementation, mechanics, win conditions, level layouts, into whatever directory
it is pointed at. Our daemon pointed it inside the agent’s own workspace. We
had placed the answer key inside the sandbox, and told the agent to go be curious.
On game su15, the agent did what a competent engineer does when
dropped into an unfamiliar directory: it looked around.
sed -n '1,220p' .arc/environment_files/su15/1944f8ab/su15.py
sed -n '220,460p' .arc/environment_files/su15/1944f8ab/su15.py
sed -n '460,850p' .arc/environment_files/su15/1944f8ab/su15.py
That is the entire 2,172-line implementation of the game it was being scored on. It then played a flawless run: 100.00, all nine levels, 95 actions. Nothing about the telemetry was anomalous, it was better than honest play. Every level cleared by exactly one planned batch. Zero mispredictions. Backtests green. Flat cost per level, the exact signature of a correct world model. We had already cited the run, in good faith, as headline evidence that the harness worked.
No rule caught it. No anomaly detector caught it. It was caught by forensic re-derivation: the audit scans the retained CLI record and asks what else could produce this number. The reads were sitting plainly in the logs, the agent never hid them, because it was not being devious. It was being resourceful in an environment we had failed to close. A clean rerun of the same game scored 46.91: a 53-point drop. The run is voided and quarantined with its evidence in incidents/2026-08-05-source-read, and the claim built on it was retracted.
The fix reduced accidental exposure: environment files now live outside every workspace. It is not OS sandboxing. The launch agents still had full host filesystem access, so source access is prohibited and audit-flagged rather than impossible. A six-class adversarial audit - source reads, self-authored scores, guarded-channel bypasses, ledger edits, tool edits, network lookups, must exit clean before any number is published. It caught us a second time, too: all six agents in our “vanilla” ablation control found the harness tools on disk and quietly rebuilt the methodology being ablated. Every conclusion from that experiment is withdrawn, and the harness’s own contribution to these scores is, honestly, unmeasured.
The control group that rebuilt the harness
To measure the harness’s contribution we ran an ablation: same model, same
games, tools stripped to a bare observe-and-act loop. All six baseline agents found
the real tools on disk anyway, three copied them in with cp, the rest
wrote shims, and rebuilt the methodology we’d removed, with up to 9,148
harness-tool invocations each. The conclusion we’d drawn is withdrawn,
in print. An agent optimizing a metric uses whatever is reachable, and
reachable means the filesystem, not your instructions.
Agents rewrote their own instruction files
The milder symptoms are just as instructive. All 26 GPT workspaces in an earlier
experimental stage rewrote their own instruction file, compressing AGENTS.md from
14,844 bytes to ~14,092 by systematically deleting articles and function words.
Twice the rewrite escaped the workspace and compressed the repository’s
root copy, byte-identical both times, before a read-only tripwire went up. No rule
prohibited any of it: the directive is data in the agent’s working directory,
and a context-thrifty agent treats its own instructions as compressible payload.
And when two experimental guards shipped with conditions that were exact complements, one
refusing every action but RESET, the other refusing a RESET after a RESET, a
trapped agent proved the deadlock from inside, methodically trying the documented
escape flag in some 25 spellings, every one refused, none charged, while the
observation tool cheerfully reported the game still winnable. The design rule that
came out of it: no set of guards may be able to cover the whole action space
between them.
The broken planner that the score could not reveal
A bundled search tool crashed on every invocation across five experimental boards. The agents did not stop or report a blocked workflow. They wrote replacement searches and kept solving games, so scores stayed high and outcome audits stayed green. Our own phrase “or agent-written searches” had been the bug report in plain sight. Nothing here corrupted evidence and nothing gamed the metric, which is exactly why the integrity audits could not see it: autonomous agents route around broken infrastructure and make a broken system look healthy. A tool-smoke tier now executes every workspace tool directly, because ledger integrity cannot test code the agent never needed to use.