BREAKING: Apple just proved AI "reasoning" models like Claude,...

They tested Claude Thinking, DeepSeek-R1, and o3-mini on problems these models had never seen before.
The result ↓
They used fewer tokens and gave up faster, despite having unlimited budget.
Like handing someone step-by-step instructions to bake a cake.
The models still failed at the same complexity points.
They can't even follow directions consistently.
Then they fall apart like a house of cards.
Instead, they hit hard walls and start giving up.
Is that intelligence or memorization hitting its limits?
Current "reasoning" breakthroughs may be hitting fundamental walls that can't be solved by just adding more data or compute.
The industry is chasing metrics that don't measure actual intelligence.
• They avoid data contamination
• They require pure logical reasoning
• They can scale complexity precisely
• They reveal where models actually break
Smart experimental design if you ask me.
Is Apple just "coping" because they've been outpaced in AI developments over the past two years?
Or is Apple correct?
Comment below and I'll respond to all.
1. Follow me @rubenhassid for more threads around what's happening around AI and it's implications.
2. RT the first tweet

They just memorize patterns really well.
Here's what Apple discovered:
(hint: we're not as close to AGI as the hype suggests)



