HN
Today

Position: LLMs Can't Jump

A DeepMind researcher's position paper argues that LLMs are structurally incapable of making true 'leaps of intuition' to create foundational axioms, using Einstein's General Relativity as a case study. The paper sparked significant debate on Hacker News, leading the author to clarify that his intent was to explore the conditions for such AI leaps, not to dismiss AI's scientific potential entirely. This recurring discussion taps into fundamental questions about AI's intelligence, creativity, and the nature of scientific discovery itself.

58
Score
25
Comments
#1
Highest Rank
9h
on Front Page
First Seen
Aug 5, 12:00 PM
Last Seen
Aug 5, 8:00 PM
Rank Over Time
11382024232528

The Lowdown

The paper "Position: LLMs Can't Jump" posits that large language models are inherently limited in their ability to generate foundational axioms or make significant intuitive leaps, similar to those seen in groundbreaking scientific discoveries. It uses Albert Einstein's formulation of General Relativity as a prime example, suggesting that the kind of conceptual leap required for such a breakthrough is beyond the current structural capabilities of LLMs, especially when data is sparse.

  • The core argument is that LLMs struggle to create entirely new, foundational axioms, contrasting with the human capacity for intuitive, theory-building thought experiments.
  • The author, Tom Zahavy of DeepMind, clarified in a public statement that the paper represents his personal views and is not a company position. He stressed that the paper is not a dismissal of AI for science, but rather an exploration of what it would take for AI systems to achieve a specific kind of profound, axiom-generating jump.
  • Zahavy, himself a contributor to AI systems that have made scientific discoveries (like AlphaProof), noted that improving current AI recipes is still highly valuable and that his paper focuses on a particularly challenging, long-term goal for AI.

Ultimately, the paper and subsequent discussion highlight the ongoing debate within the AI community about the limits of current LLM architectures and what future advancements might be necessary to enable more profound, human-like scientific creativity.

The Gossip

Authorial Aims and AI Alchemy

The discussion immediately highlighted that the paper's provocative title might be misleading. The author, Tom Zahavy, clarified in a follow-up that "LLMs Can't Jump" is a personal position paper, not a company stance, and doesn't assert LLMs can *never* make scientific discoveries. Instead, it's an exploration of what it would take for AI to achieve profound, axiom-creating leaps like Einstein's, a problem he admits might not be the most urgent. Some commenters pointed out potential weaknesses in the paper's core arguments or suggested alternative theoretical frameworks the author could have used.

Leaping Limitations & AI's Environment

A significant portion of the debate centered on the inherent capabilities and limitations of LLMs regarding true intuition and creative leaps. Some argue that an LLM's "intuition" is fixed after training, making genuine novel axiom creation impossible without retraining or specific architectural changes. Others contend that the limitations might be more about the "harness and environment" an LLM operates within, suggesting that with richer interaction and sensory grounding, LLMs could indeed achieve more profound insights, or that introducing elements like noise or randomness could simulate human-like creative jumps.

Historical Hurdles in Human Intuition

Commenters also drew parallels to human intelligence and scientific discovery, questioning the simplistic narrative of human "leaps of intuition." One commenter argued that the popular retelling of Einstein's work, for instance, is often reductive, suggesting that even human breakthroughs are built on extensive groundwork rather than pure spontaneous insight. Others contemplated whether humans are truly "good" at abstract leaps, or if our perceived ability is simply due to a lack of a better comparative metric. This perspective adds a layer of complexity to what we expect from AI.