In partnership with

A year ago, Apple tried to test the world’s best AI reasoning models. Models like OpenAI’s o-series, Anthropic’s thinking Claude, and DeepSeek-R1. These were the same models many people saw as a big step toward AGI.

And Apple managed to break them.

When the puzzles became hard enough, all of the models dropped to zero accuracy. Then Apple gave them the full step-by-step solution and simply asked them to follow it. But they still failed at the exact same point.

Apple published this in a paper called “The Illusion of Thinking.” Their conclusion: these models are not truly reasoning. They are mostly pattern-matching, and once a problem becomes too long or complex, that pattern starts to fall apart.

I covered that paper last year. It went a little viral, a lot of you found me through it. And I agreed with most of it.

Subscribe to keep reading

This content is free, but you must be subscribed to ninzaverse to continue reading.

Already a subscriber?Sign in.Not now

Keep Reading