George Hotz said humans are “100T models that train on only 10B tokens.”

Too busy? Skip to TLDR

The argument

Yann LeCun has been the most vocal critic of autoregressive LLMs, arguing that animals and humans get very smart with vastly smaller amounts of data and that a child learns to recognize cats from a few examples, whereas GPT needed the entire internet.

His alternative to reach human-level intelligence is JEPA: Joint Embedding Predictive Architecture, which is a promising research directions. We need a richer way to represent the world than language-based tokens, even though LLMs can already build latent representations of the board state in games like Othello and chess (cool video about it) out of it.

But the central claim that humans are far more sample-efficient learners oversimplifies how we got there.

3.5 billion years of evolutionary pretraining

The AI vs Human comparison only counts postnatal sensory inputs of humans against the training corpus of an LLM, and ignores the immense “pretraining” encoded in the child’s genome. The human brain starts from ~750 megabytes of DNA which is the result of ~3.5 billion years of evolutionary optimization across different organisms, which cannot be compared to a random weight initialization of AI. Anthony Zador calls this the genomic bottleneck.

Every generation of organisms that ever lived was a “training run” where the organisms that failed to meet survival-relevant requirements were eliminated from the gene pool. Over millions of generations, hominid brains went from ~400 cc in Australopithecus afarensis (3.5 million years ago) to ~1,350 cc in modern Homo sapiens.

The evolutionary records suggest that the first stone tools appeared ~3.3 million years ago with ~400 cc brains. Homo erectus made handaxes with ~900 cc. Compound tools and the habitual use of fire came around 500,000 years ago. Symbolic behavior emerged 100,000–70,000 years ago. And then, around 50,000 years ago, came what some archaeologists call the “Great Leap Forward” with cave paintings, complex tools, and behavioral modernity.

Each jump in capability likely followed both increases in brain size (scaling laws?) and changes in brain architecture, though we can only guess it from indirect evidence like tools, endocasts, and genetics. We will probably never be able to reverse engineer each step to find out what capability appeared at what point.

What babies know before they “learn” anything

We can be biased by our own case, since human babies seem to go from not being able to do anything to learning things from scratch during childhood. But other species’ capabilities show a different picture: calves can walk within hours of birth, sea turtles find the ocean right out of the egg, migratory birds can fly thousands of miles with zero parental guidance. Human babies take about 12 months to walk, partly because growing our oversized brain eats up so much energy during pregnancy that we’re born with less mature circuits than most other animals. However, there is an interesting phenomenon called the stepping reflex, present at birth, which disappears and later returns. That could mean some motor program is genetically encoded, and that the baby is recovering a capability instead of learning it from scratch.

Vision shows similar untrained capabilities through a subcortical pathway that handles fast, unconscious threat detection. In monkeys, pulvinar neurons fire faster and stronger for snake images than for faces or shapes. Babies between 7 and 10 months show stronger neural responses to snakes than to frogs or caterpillars, with no prior exposure. This shows that we are not necessarily few-shot learning these capabilities: some object detection and threat classification already comes wired in.

Sleep

Roughly a third of human life is spent sleeping. During the REM phase of the sleep, the brain paralyzes the body and replays the day’s learning in bursts, moves memories from short-term to long-term storage, and prunes weak connections while strengthening useful ones. REM sleep kinda looks like reinforcement learning where difficult (unfamiliar) tasks are replayed and the brain optimizes itself to successfuly complete them in the future.

Synapses vs Parameters

The human brain contains ~86 billion neurons forming an estimated 100–150 trillion synapses. The current frontier LLMs are estimated to be between 2 to 5 trillion total parameters, approaching the synaptic count of a cat brain. But comparing synapses to parameters massively understates brain complexity, which makes the geohot’s estimation of “100T models” and “10B tokens” pretty off.

Parameters are fixed numbers that do the same thing to whatever reaches them. A synapse fires probabilistically: depending on the synapse, it responds anywhere from 95% of the signals to 2% of them. How it responds depends on the timing of the signals carries information by itself. The thousand synapses of one neuron are not copies of each other, each can be rewritten from the outside by other chemicals, and the brain creates and destroys them continuously. A synapse behaves like a small self-modifying circuit (cool video about that). And that’s for one synapse. It takes a neural network 5 to 8 layers deep just to approximate one cortical neuron.

TLDR

LeCun is right that current LLMs are limited by their architecture: they lack world models, persistent memory, self-training, etc.

But comparing human sample-efficiency to artificial intelligence models trained from scratch is misleading. The brain is not randomly initialized: it’s a ~750 megabyte compressed checkpoint that already includes pretrained visual processing, threat detection, physics intuitions, social cognition, and learning algorithms. New borns arrive with capabilities nobody taught them, the brain keeps retraining itself every night and synapses are not comparable to parameters.

Life is the product of the largest and longest training process on Earth, one that consumed countless organisms across billions of years.

Written language is the one thing we actually do learn from scratch. It is too recent for evolution to have prepared us for it, and we need years of explicit training just to read and write. It is also humanity’s most abstract representation of its environment, and LLMs being able to reason out of this abstraction is really impressive.