FrontierPhysics v0.1 with 56 tasks

FrontierPhysics: Evaluating Agents for End-to-End Physics Research

182submissions56tasks9domains7modalities42researchers38institutions

Agent Performance

0%5%10%15%20%25%240 min120 min60 min30 min15 minavg agent wall-clock per rollout (log)overall pass ratemost efficient ↗fleet 6.4%Claude Opus 5.5GPT-6 AstraClaude Fable 5.1Claude Opus 5GPT-5.6 SolGrok 4.6GPT-5.6 TerraGPT-5.6 LunaKimi K3DeepSeek V4 ProGemini 3.8 Flash
Hover a point to pin its exact pass rate and wall-clock to the axes. Dashed lines mark fleet means; the emerald line is the Pareto frontier.

Discovery Loop

Research &PlanResearch &PlanImplement &ExperimentImplement &ExperimentEvaluate &FeedbackEvaluate &Feedback
  • Research & Plan

  • Implement & Experiment

  • Evaluate & Feedback

Agent Leaderboard

Sort by
#AgentOverall
1
Claude Opus 5.5Claude Code
22.0%
2
GPT-6 AstraCodex
14.9%
3
Claude Fable 5.1Claude Code
13.1%
4
Claude Opus 5Claude Code
10.1%
5
GPT-5.6 SolCodex
3.0%
6
Grok 4.6Grok Build
3.0%
7
GPT-5.6 TerraCodex
1.8%
8
GPT-5.6 LunaCodex
1.2%
9
Kimi K3Claude Code
0.6%
10
DeepSeek V4 ProClaude Code
0.6%
11
Gemini 3.8 FlashAntigravity
0.0%
Hover over a row to see confidence intervals, time and cost. Score is 0–100: a passing rollout's weighted rubric score, 0 for any failure.
FrontierPhysics · 56 tasks · 3 trials per task · 95% CIs
Anthropic
OpenAI
xAI
Moonshot
DeepSeek
Google

Cost Efficiency

0%5%10%15%20%25%$50$20$10$5$2$1$0.50mean inference cost per rollout (USD, log)overall pass ratemost cost-efficient ↗fleet 6.4%Claude Opus 5.5GPT-6 AstraClaude Fable 5.1Claude Opus 5GPT-5.6 SolGrok 4.6GPT-5.6 TerraGPT-5.6 LunaKimi K3DeepSeek V4 ProGemini 3.8 Flash
Hover a point to pin its exact pass rate and cost to the axes. Dashed lines mark fleet means; the emerald line is the Pareto frontier.

Physics-Domain Profile

204060Condensed MatterPhysicsAMO PhysicsQuantumInformationAstrophysicsFluidDynamicsStatisticalPhysicsPhotonics &ImagingHigh EnergyPhysicsChemicalPhysics
Condensed Matter Physics12 tasks
Claude Opus 5.5Claude Code41.7 → 30.6%
GPT-6 AstraCodex36.1 → 25.0%
Grok 4.6Grok Build11.1 → 11.1%

solid = overall pass · pale = passed the verifier only · hover another radar axis to switch domain

Claude Opus 5.5· Claude CodeGPT-6 Astra· CodexGrok 4.6· Grok Buildrings at 20–60% · overall pass

Example Task

multiplexing-ion-chain-qnet

hardtrapped-ionscalculationsimulationoptimization
tasks/multiplexing-ion-chain-qnet/task.md
    • g2_data54 files
    • paper_template16 files
    • assets32 files

Loading…

Run eval

claude-agent-acp · effort max · no-skill · trial 3 of 3