Is AIME the right benchmark for reasoning in LLMs? OpenAI and...

@DimitrisPapail
Dimitris Papailiopoulos@DimitrisPapail
7 views Feb 20, 2025 ~2 min read
1
Is AIME the right benchmark for reasoning in LLMs?

OpenAI and DeepSeek have effectively used AIME to demonstrate a key point about test-time compute: accuracy improves with increased inference compute (more output tokens), and both trend upward with more rounds of reinforcement learning.

AIME has been an incredible vehicle for story telling, and clearly played an important role conveying the narrative that "LLMs can solve competition-level math problems, and their accuracy increases as a function of inference time compute, with no saturation in sight".

It honestly sounds amazing!

However, as a benchmark, AIME has several major drawbacks:
1) It's difficult to scale fresh AIME-like problem generation, as we're limited to waiting for each year's mid-February release of just 30 questions.
2) There are concerns about train-test leakage, as many of these questions look near identical to stuff on the internet.
3) It is unclear if it correlates with real world use cases.

If AIME's primary utility is in storytelling, and we’re less concerned about practical relevance, it’s worth asking whether it’s still the best choice. Other tasks could serve the same purpose: N-digit multiplication, maze solving, or some ARC like puzzle could equally demonstrate that "accuracy scales with test-time compute and RL rounds." These might not sound as amazing as competition-level math solving, but might better serve the scientific study of reasoning.

Where I am going with this is that alternative benchmarks could be designed that are 1) easier to generate at any scale, 2) come with automatically verifiable problems, and 3) a knob for controlled difficulty—logic puzzles, algorithmic/coding questions, auto-generated mathematical proofs, etc, etc. Perhaps if we're lucky these could be designed to have higher alignment with real-world reasoning tasks.

AIME has served its purpose in storytelling and demonstrating that LLMs can "reason their way" to better performance. But given its significant limitations, perhaps it’s time to retire it in favor of more purposeful, scalable, and reliable benchmarks.

I have my doubts that this will happen as the community still uses MMLU and GSM8K that share similar issues (and then some more). We seem content with these benchmarks perhaps because they were created by humans, and also appear challenging to humans. Perhaps that is good enough for storytelling.
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial