Ilya Sutskever’s interview with @dwarkesh_sp is a great showcase of...

In this clip, Ilya shows that he’s inching towards the truth with this interesting analogy: “Suppose you have two students. One of them decided they wanted to be the best competitive programmer, so they will practice 10,000 hours for that domain. They will solve all the problems, memorize all the proof techniques, and be very skilled at quickly and correctly implementing all the algorithms. By doing so, they became one of the best.
Student number two thought, “Oh, competitive programming is cool.” Maybe they practiced for 100 hours, much less, and they also did really well. Which one do you think is going to do better in their career later on?”
Dwarkesh, like most other sane humans, picked the second student. They don’t get closer to why they’d do so than that the second student has the ‘It.’
This is the closest he gets to the truth. But the analogy holds the key. The reason you pick the second person is that they seem to be following their genuine interests. And solving real world problems. The first person, on the other hand, is driven by a desire to ace an artificial test or benchmark. While that could be their genuine interest (and if so, more power to them), we infer that it may be driven by something else, the notion that acing the test is merely a means to an end: a highly paid job, social acclaim, etc. So the test is a nuisance to overcome. The second developer just does it for fun and the instant it’s no longer fun, does something else.
Even more important than their differences, though, are the similarities between these two programmers. They both have problems and they both follow interests (and as an aside, we should refrain from making value judgements about their interests and problems). They follow their interests and solve their problems. And as they do, their knowledge can grow. AI can do none of this.
To start with, AIs are static. They are ‘trained’ and released. There’s no learning after training. But actually, there’s no learning in their ‘training’ either. The models ‘absorb and compress’ a vast corpus of knowledge, and when prompted can respond with next token prediction. But this doesn’t imply that the models themselves have knowledge. They don’t. They hallucinate 100% of the time. They are increasingly correct – especially on close-ended tasks, but they have no idea of knowing when they are (correct or not). They, in fact, know nothing. That is why they can blow you away at times and at other times seem dumb beyond belief.
When ‘superintelligence-believers’ see a human, they see limitations: compared to AIs we are limited in the breadth of facts we can store. Our learning process is slow and error-prone. And worse (in their view) it is often obstructed by limited attention. “Attention is all you need”, right? This is false. Rather than being limitations, it’s exactly these traits that are required for AGI. The underlying model of learning that dominates our discourse is “the bucket theory of mind”, that learning is the transfer of facts from the teacher’s mind to that of the students. But, as Karl Popper has shown (and I contend we all know from experience), this is impossible. We can’t simply absorb facts. Nothing can. We can only learn what we are interested in (which unfortunately can also be driven by threats/coercion), and only via guessing the meaning, exposing that guess to criticism, and updating the guess: ‘Conjecture and refutation’ as Popper called it.
Continual learning is the only way to create knowledge. Ilya is right to call for research instead of scaling. And a great starting point is to understand how we became AGI. It’s not about our genes (our ‘hardware’), it is something about the ‘software’ that we run. As David Deutsch said brilliantly: “The thing that provides the G in AGI is bound to be simple. We know it’s simple in humans because we only differ slightly from chimpanzees (…) The thing that is causing the generality must be a relatively simple program, only a few k or a few 10s of k (…) When doing the mental arithmetic, we are not using the G part of our brain, we are using the dumb computation part. We are harnessing a part of the brain that’s not human at all. We can offload that kind of task to computers and we already do. That is how I’d like to think of a future AGI. It’s a program similar to our program with a lot of compute power that we (ourselves) also have access to (via computers).”
We should not prophesy, but I will suggest that the path to AGI will be faster if the AI community becomes interested in the work of David Deutsch and Karl Popper, then guesses what it means, and finally, attempts to program it. For as David Deutsch says: “If you can’t program it, you haven’t understood it.” And only when we understand it, we will have it.