Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2%...

Claude Opus 5 reaches 30.2%, materially outperforming Fable
Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
Of these, it was able to beat 4 of them matching or surpassing human level efficiency
Newly beaten environments: ar25, ft09, lp85, r11l, s5i5
6 of the 25 public demo environments have now been solved
See Opus 5 play s5i5: arcprize.org/replay/cb8fbc9…
Opus 5 used advanced logical reasoning to turn ARC-AGI-3 layouts into algebraic notation. On action 23 it described the scene as "4_center = 2×axis − 5_center"
This is the first explicit reflection equation by a model we've analyzed
By action 248 the model went further and generalizes this to two dimensions: “ghost = reflection about axis strips: br' = 2·br_axis − br, bc' = 2·bc_axis − bc”
ar25 was previously unbeaten. View replay: arcprize.org/replay/b71b579…
- Reproduce the results: github.com/arcprize/arc-a…, github.com/arcprize/arc-a…
- Testing policy: arcprize.org/policy
- Full Claude Opus 5 Results: arcprize.org/results/anthro…
arcprize.org/scorecards/mod…


