The Einstein Test and Why LLMs Can't Jump

Google DeepMind has proposed an interesting problem called The Einstein Test.

Imagine an LLM that has access only to all scientific knowledge and published papers up to 1905, that is, exactly the information Einstein had before presenting the theory of relativity. Now the question is:

Could this model have discovered the theory of relativity on its own?

It’s a clean thought experiment because it removes the usual excuse. You can’t say the model lacked data. It has everything Einstein had, every paper, every result, every open problem sitting in the physics of the day. The Michelson-Morley experiment and its awkward null result are in there. Maxwell’s equations are in there. The tension between them and Newtonian mechanics is in there. Einstein didn’t have a secret extra dataset. He had the same pile everyone else had. So the test isolates the one thing that’s actually in question. Not knowledge. The leap.

The answer, and the name

The answer from DeepMind researchers is: No. That is why they named their presentation “LLMs can’t jump”, meaning language models cannot perform a “leap of thought.”

The reason is that LLMs are extremely strong at finding patterns, combining information and reasoning based on what they have previously learned. However, when it comes to creating a completely new idea or presenting a fundamental principle that no one has ever proposed before, they run into trouble.

That distinction is the whole thing. Pattern-finding and recombination are exactly what these models are built to do, and they do it better than any human. Hand an LLM the 1905 corpus and it will summarise every position, spot every inconsistency, connect papers nobody thought to connect. What it won’t do is stand up and declare that time itself is relative, that simultaneity depends on the observer, that the problem isn’t in the measurements but in an assumption everyone treated as bedrock. That move isn’t in the data. It’s a decision to reject part of the data’s own framing. And a system trained to predict what comes next based on what came before has nothing to push off from when the right answer is something that never came before.

Induction, deduction, and the one they’re missing

In fact, models work mostly based on Induction and Deduction, but what Einstein did was more like Abduction.

Worth being concrete about the three, because the labels carry the argument. Induction is generalising from examples. See enough cases, infer the rule. Deduction is the reverse, take the rule and derive what must follow. LLMs live in both of these and they’re formidable at them. Both stay inside the existing frame. Induction reads patterns out of what’s there. Deduction works out consequences of what’s already assumed. Neither invents the frame.

Abduction is different. It’s the leap to the best explanation, inventing a new hypothesis that would make the confusing facts suddenly make sense. It’s not read off the data and it’s not derived from the rules. It’s a guess at a cause nobody had proposed, judged by how much it explains once you assume it. That’s what Einstein did. He didn’t extend the existing physics. He proposed a new starting principle and the mess resolved.

A mental leap that cannot be achieved merely by putting previous data together.

What LLMs can actually do

Of course, this does not mean that LLMs are completely incapable. Modern models, when equipped with tools such as web search, code execution, simulation, iterative experimentation and autonomous agents, perform much better at generating new hypotheses and assisting scientific research. In the last couple of years, AI systems have already helped researchers in areas such as drug design, new materials discovery, and solving certain mathematical problems.

So this isn’t a dismissal. Give a model tools and a loop, let it search, run code, simulate, test its own guesses and iterate, and it becomes a genuinely powerful research partner. The results are real. Drug design, new materials, hard maths problems, all of it moved with AI in the loop. However, there is still a huge gap between “assisting in discovery” and “performing a scientific revolution on the scale of the theory of relativity.”

Perhaps the most important point is this: at least with the current architecture, LLMs play more the role of an incredibly fast and accurate researcher, not a “new Einstein.”

Paper: https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20(17).pdf

Logo

Naomi Nour - building AI that's genuinely useful.

Twitter Github YouTube ADPList