In 1960 the physicist Eugene Wigner published an essay whose title became a genre: “The Unreasonable Effectiveness of Mathematics in the Natural Sciences.” His puzzle was simple to state and has never been resolved. Mathematics is a game played with symbols, developed largely for its own internal beauty. Why should it describe the orbit of a planet or the spectrum of hydrogen to ten decimal places? Wigner called it “a wonderful gift which we neither understand nor deserve.”
For sixty-odd years this stayed a philosopher’s puzzle, because we had only one instance of it. Mathematics was the only instrument we had that worked this well, so nobody could tell whether the mystery belonged to the instrument or to the world.
We now have a second instance, and I think it changes the question.
The second miracle
Consider what a large language model is. The objective is almost embarrassingly simple: given some text, predict what comes next. The optimizer is crude: nudge billions of numbers slightly downhill, over and over. The architecture is generic and contains no theory of grammar, physics, law, or mind. Nobody programmed in the ability to write a proof, debug a kernel, or explain a joke.
Yet those abilities appear. By the theory we had when this started, they shouldn’t. Classical learning theory says a model with vastly more parameters than it needs should memorize its data and generalize terribly. These networks can in fact memorize pure noise, which was shown experimentally years ago. When given real data, they don’t. They find the pattern instead of the lookup table, and nobody has a complete account of why.
Rich Sutton called the broader pattern the bitter lesson: across decades of AI research, general methods plus scale have beaten human cleverness encoded by hand, every time. It was bitter because we had assumed intelligence required bespoke design. We assumed that because we flattered ourselves that our own intelligence was bespoke.
So we have Wigner’s situation again: a method that works far better than our theory of it says it should. The obvious question is whether the two miracles are related. I think they are one fact, observed twice.
Prediction is compression
Start with an old idea from information theory: prediction and compression are the same task. If you can predict the next symbol well, you can encode the sequence in fewer bits, and the reverse also holds. A perfect predictor of a text is a perfect compressor of it.
Now consider the shortest possible description of a large body of human text. Surface statistics like word frequencies and common phrases only get you so far. Past that point, the cheapest way to encode a physics textbook is to know some physics. The cheapest way to encode a million arguments is to model how reasoning goes. The cheapest way to encode dialogue is to model the people speaking. The objective looks shallow, but its minimum is deep. A system pushed hard enough toward better prediction is pushed toward modeling whatever generated the data.
This makes the success of AI less mysterious, but only if the generating process is itself compressible. If the world were a lookup table of unrelated facts, no amount of scale would help, because there would be nothing to find.
The world is not like that. Physical reality is dominated by low-order interactions, locality, symmetry, and hierarchical composition, with big things made of smaller things by reusable rules. Max Tegmark and colleagues pointed out some years ago that this is precisely the class of functions deep networks represent cheaply. Out of the unimaginably vast space of possible functions, nature uses a vanishingly small, highly structured corner. Neural networks happen to be good at that corner.
Mathematics is good at the same corner. A differential equation is a compression: a line of symbols that unpacks into every trajectory a system could ever take. Newton’s laws are a few hundred bits that encode every orbit. Wigner’s miracle is that such short descriptions exist and that we can find them.
So here is the claim. The unreasonable effectiveness of mathematics and the unreasonable effectiveness of AI are the same phenomenon: the world is radically compressible, and anything that compresses well will look unreasonably effective at describing it. For three centuries mathematics was the only such compressor we had, so we mistook a fact about the world for a fact about mathematics. Now a second compressor, built independently, with no axioms, elegance, or understanding, has also worked. When two very different instruments succeed at the same job, the explanation is probably in the job.
What the claim predicts
A philosophical thesis earns its keep by sticking its neck out. This one does so in three places.
Effectiveness should track compressibility, domain by domain. AI should be strong wherever the generating process is simple relative to the data it produces: language, code, protein structure, short-horizon weather. It should be weak where the process is incompressible or adversarially self-modifying. Examples are the fine detail of chaotic systems past their predictability horizon, markets that adapt to any predictor, and regimes with no precedent in the data. People call AI’s uneven performance the “jagged frontier” and treat it as an engineering quirk. On this view it is a map. The jaggedness shows where the world compresses and where it doesn’t.
Math and AI should succeed in correlated places, and the exceptions should be informative. Where both work, the structure is simple enough for humans to write down. Where AI works and mathematics has not, as in protein folding and natural language itself, we have found something new. There is structure that is real and learnable but too high-dimensional to fit in any formalism a human can hold in mind. This kind of regularity can be known but not stated. I don’t think we’ve absorbed how strange this category is, or how large it might be.
Scaling laws are measurements of the world as well as of models. Model performance improves as a smooth power law across many orders of magnitude of scale. That smoothness says something about how structure is distributed in the data. There is a great deal of it, at every scale, in diminishing but never-exhausted increments. The exponents depend on architecture and data as well as on nature, so this needs care. Still, some of what we are reading off those curves is a property of reality, seen through the network.
Three objections
It’s a selection effect: we celebrate the domains where AI works and ignore the rest. This is true, and Wigner faced the same objection. But the claim here is conditional and testable. It doesn’t say AI works everywhere. It says where AI works is predicted by compressibility. If AI turned out to master incompressible domains, or failed consistently in highly structured ones, that would count against the thesis.
AI is trained on human output, so it’s compressing us, not the world. This is partly right, and important. Language is pre-distilled cognition, the residue of billions of people who already did the hard work of learning from raw reality. Much of a language model’s effectiveness is inherited. But the objection fails as a general account, because the same methods work when trained directly on nature. Protein structure predictors learn from crystallography data, and weather models learn from the atmosphere’s own record. No human understanding was in the loop, yet the compressor worked.
“The world is compressible” is a tautology. It is not, and this may be the most important point. By a simple counting argument, almost all possible sequences are incompressible, and almost all possible functions have no description shorter than themselves. A compressible universe is an extraordinarily special one. We don’t know why we live in one, though it’s fair to note that no one could live in the other kind, since an incompressible world has no stable structure in which to evolve an observer. That is a version of Wigner’s mystery, pushed down a level but not dissolved.
The separation
There is one way the two miracles differ, and I think it is the most consequential part of the story.
Mathematics gives effectiveness with understanding. The equation is the explanation. When you have Maxwell’s equations, you can predict light and you also know what light is. For the entire history of science, predictive power and comprehension arrived together, so we assumed they were the same thing, or at least inseparable.
AI gives effectiveness without understanding, at least without ours. A network predicts how a protein folds and hands us no theory of folding. We are reduced to doing natural science on our own artifacts, probing trained networks the way a biologist probes a cell, trying to recover the theory the network found but cannot state. We have built a second thing in the universe that knows without explaining. Nature was the first.
This forces a question we never had to ask. If we can have the full predictive content of a theory without the theory, what was understanding for?
The deflationary answer is that understanding was a convenience: a compression format sized for human working memory, precious to us and irrelevant to knowledge. I think that answer is wrong, and the compression thesis itself shows why.
A compressor works on what exists. It needs data, and data is a record of what has already happened, already been said, already been measured. Everything AI does so well happens in this positive space, the territory that has been mapped, however thinly. Inside it, a good enough compressor will find any regularity there is to find, and will increasingly find it faster than we do. I see no reason to expect a lasting human advantage there, and I think people searching for one are searching in the wrong place.
The frontier of knowledge is somewhere else, in the negative space: the question nobody has asked, the concept that does not yet exist, the regime with no precedent. There is nothing there to compress, because nothing has been written down yet. What operates there is judgment, not prediction: a sense of which of the infinitely many unexplored directions is valid and worth the walk. That is what understanding is for. It tells you where a model will break before it breaks. It tells you which question is worth asking. It lets you step off the edge of the data and land somewhere real.
So the separation of prediction from understanding does not make understanding obsolete. It isolates it. For all of history, understanding came bundled with the labor of calculation and derivation, and we could not see it clearly because it was always doing two jobs at once. AI is taking the first job. What remains is the pure form of the second, the capacity to originate. As compression becomes abundant and nearly free, that capacity becomes the scarce resource, and everything else gets measured against it.
Could a machine originate too? I won’t claim it is impossible in principle. I will claim that it will not come from compressing harder. Scale on a predictive objective buys more of the positive space, at finer resolution. It does not buy the negative space, because prediction has no purchase where there is nothing yet to predict. Whatever crosses that line, in a human or a machine, is a different kind of act, and we should stop confusing it with the one we have just learned to automate.
Wigner ended his essay with gratitude for a gift he couldn’t explain. We have now been given the same gift a second time, in a form that works just as well and explains nothing. That tells us the gift was always a property of the world, and that mathematics and AI are two ways of making use of it. It also tells us where the gift runs out. The compressible world is now open to anyone with enough compute. The unmapped world still needs someone to decide where to go.