You wake up, scroll through LinkedIn, and see a high school student posting that they used ChatGPT to find a missing reduction and, together with classmates, have cracked P vs. NP, the most famous open problem in computer science.
A week prior, you remember seeing a Reddit post claiming Claude solved the same problem, complete with a downloadable 142-page PDF. A few days later, someone shared a GitHub repository filled with diagrams, lemmas, and a Lean formalization created with Gemini. That same day on Facebook, a completely different person built an animated website demonstrating how she solved it, verified by ChatGPT, Kimi, and DeepSeek. Over the next few weeks, several YouTube creators published hour-long videos claiming independent solutions using various open-weight frontier models.
No major news outlet has broken the story, but social media has gone wild. Everyone is competing to claim they are the one who truly solved it with a rigorous proof. Ironically, the most crucial group remains silent: computational complexity theorists. They have stayed quiet for weeks with no interviews scheduled, offering only brief statements that they are reviewing the claims. The public demands answers. P vs. NP is famous in sci-fi and pop culture; it feels like everyone knows the question, but everyone is waiting for a true proof. The world is waiting for the mathematical community to break its silence and confirm if the problem is actually solved, which would revolutionize technology: breaking cybersecurity foundations, optimizing global energy grids, and transforming materials science.
Let’s step back for a moment: what if there was no external chaos, but rather a few small teams of mathematicians working alongside frontier AI for a year, followed by two months of rigorous peer review, culminating in a formal announcement that P vs. NP has finally been proved or disproved?
Which scenario is more likely?
When proof-shaped claims become cheap to generate, who gets to make the call?
Originality isn't the bottleneck
Here is the misconception our culture keeps smuggling into every conversation about AI and genius:
A regular person mostly needs access to one original idea. Once the idea appears, the world can take care of the rest.
You can hear it in founder mythology. You can hear it in the way breakthroughs get retold as lightning strikes rather than long slogs. You can hear it in the advice handed to young people, which increasingly amounts to a warning not to get trapped in expertise, because expertise makes you conventional, so you should just think differently.
The belief is emotionally convenient, because it makes greatness look like a door that might open for anyone at any moment. The historical record is less convenient. Shakespeare had a trade. Darwin had barnacles. Picasso could draw like an academic master before he shattered the figure. Bach copied scores by hand as a boy. Katalin Karikó stayed with mRNA for decades before the world caught up with her. None of these stories begins with a blank mind receiving a perfect bolt from the sky.
The pattern is not that deep work guarantees greatness; it plainly does not. The pattern is that field-changing work almost never appears before long contact with a domain's materials, constraints, failures, and standards. 4 5 6 56 58
Two things are worth fixing in place before we go further, because the argument is narrower and sturdier than it first looks.
Depth is necessary, not sufficient. The world is full of people with thirty years in a field who never had an original thought. Depth is the price of admission to the room where originality becomes possible; it buys nothing once you are inside. Necessity is all the argument needs.
And depth is not the same thing as a credential. The usual rebuttal to everything that follows is a list of famous dropouts, so it is worth separating the two ideas early. Ramanujan had no degree. Faraday was a bookbinder's apprentice. Karikó was demoted and defunded and worked for decades without the standing her work deserved. None of them was shallow; what they lacked was a certificate. Even the modern examples people reach for, Gates and Zuckerberg, left the credential rather than the domain: Gates had thousands of hours at a terminal and a shipped product behind him before he walked away from Harvard. The variable that matters is hours in contact with a domain's real constraints and real feedback. School is one way to accumulate them, and in some fields a poor one.
The myth survives for a second reason, though, one that outlasts any list of examples. It quietly confuses production with recognition. It asks whether a new thing can be made, and that is no longer the hard question. The hard question is that if a new thing is made, who can tell whether it matters. 1 3 7 8 24
Who would recognize Hamlet?
There is a thought experiment nearly everyone has heard. Put enough monkeys in front of enough typewriters, give them enough time, and eventually one of them types out the complete works of Shakespeare. It is almost always told as a story about probability, a way of saying that with enough random tries even order will surface from noise.
Sit with it a moment longer and a second problem appears, one the probability version quietly skips. Suppose it actually happened. Suppose that tomorrow, somewhere in a room full of monkeys and paper, one of them types out Hamlet, clean, from the first line to the last.
Those pages fall into the pile with all the rest. The room now holds an ocean of gibberish and a single masterpiece, and nothing in the room can tell them apart. The monkey cannot. The typewriter cannot. The paper does not glow, and the masterpiece carries no alarm to announce itself.
So the masterpiece is real, and for the moment it is also worthless. Its value does not switch on until a reader walks in who knows what they are holding, and that reader would have to bring a great deal with them: the English language, the conventions of Elizabethan tragedy, the shape of the blank-verse line, some sense of what had already been tried and what would count as new. Lacking that, they set Hamlet back down on the pile with everything else.
Who in that room would know?
That is the missing half of creativity. The standard definition asks for originality and effectiveness together, and Csikszentmihalyi adds the social mechanism that ties them: an act becomes creative when a field recognizes a valuable change to a domain. 1 3
Originality is the monkey's part. Recognition is the reader's part, and recognition is not free.
For most of history the room stayed small, because producing a plausible page was itself hard work. That is the thing that just changed. AI has driven the cost of the monkey's half toward zero, so the pile now grows by orders of magnitude. What it has not touched is the reader.
The flood of plausible work
The door opens and the printer starts.
Every minute, the machine can produce an artifact with the surface of expertise. A grant proposal, a legal memo, a proof sketch, a market analysis, a diagnosis, a molecule, a codebase. A PhD-shaped object complete with tables, citations, caveats, and the calm tone of someone who knows exactly what they are talking about.
This is a real change in the cost of production, not fake progress. The floor rises. More people can make competent-looking work, more first drafts become usable, more bounded tasks get solved, and for many purposes that is genuinely good.
The measured pattern already shows the catch, though. AI can improve individual outputs while pulling them toward one another. Doshi and Hauser found that AI assistance raised story ratings, with the largest gains going to the least experienced writers, while the stories grew measurably more similar to each other. The floor rose and the room narrowed. 38
Research ideation shows the same split. Expert reviewers can rate AI-generated ideas as highly novel while judging them weaker on feasibility. That is the whole machine in miniature: novelty has become cheap, and appropriateness remains expensive. 44
The danger is not that the output always fails. It is that the output sometimes fails beautifully. The labs have documented the mechanism themselves. Post-training can make a model more helpful while damaging its calibration, and benchmark incentives can reward confident guessing over admitting ignorance, so the object arrives polished while the internal alarm that should say check this has been muffled. 47 48
The most expensive thing in the world is a plausible falsehood.
In late 2024 a paper moved through the AI-and-science conversation carrying exactly the result everyone wanted. It was titled Artificial Intelligence, Scientific Discovery, and Product Innovation, and its author was Aidan Toner-Rodgers, then a PhD student in economics at MIT. The claim was clean enough to travel: an AI tool rolled out inside a large materials-science lab, and the researchers who got access supposedly discovered 44% more materials, filed 39% more patents, and produced 17% more prototypes. 51
The story was almost engineered to be irresistible. It was not a chatbot writing a poem; it looked like field evidence from a real industrial lab, with more than a thousand scientists, a staggered rollout, and innovation measured in materials and patents. It seemed to answer the exact question everyone was asking in 2024, whether AI merely makes knowledge work feel faster or actually accelerates discovery, and the answer appeared to be yes. It spread before peer review had finished. Major outlets covered it, and prominent economists including David Autor and Daron Acemoglu, a Nobel laureate, discussed and praised it even though it had not appeared in a refereed journal.
That is what made it dangerous. It did not read like spam or like a hallucinated citation at the foot of a student essay. It read like a rare piece of evidence from inside the machinery of scientific production, with numbers, institutional proximity, and the right topic at the right moment, the sort of thing people build forecasts and funding arguments and public stories around.
Then the floor gave way. In May 2025 MIT stated that, after concerns were raised and a confidential internal review conducted, it had no confidence in the provenance, reliability, or validity of the data, and no confidence in the veracity of the research. It asked arXiv and the Quarterly Journal of Economics to withdraw or flag the paper, and said the author was no longer at the institution. Acemoglu and Autor publicly urged that the findings not be relied on in academic or public discussion. 51
The shape of the failure is the point. The paper did not look sloppy to everyone; it looked good enough to pass through elite filters before the right kind of scrutiny caught up. According to reporting, the person who helped bring the problem to MIT was a computer scientist with materials-science experience, who read it and knew from domain knowledge that the tool and the lab it described did not add up. The single most celebrated empirical claim that AI accelerates discovery was itself a fabrication, and catching it took exactly the deep domain contact the myth says we no longer need. A coda makes it almost too neat: in October 2025 an EU research commissioner appeared to cite the retracted study in a speech promoting AI in science. The error had already reached policy, and no one in the chain had the depth to stop it.
A plausible falsehood is expensive because it borrows the trust reserved for truth.
AI does not remove the verification problem. It multiplies the number of things that need verifying. Which returns us to the reader in the room, and to a question we skipped past: what does verification actually take?
What the expert actually sees
Start inside a human head, with an experiment that isolates the thing.
Put two people in front of a chessboard for five seconds, one a beginner and one a master, then cover the board and ask each to rebuild the position from memory. If the position came from a real game, the master reconstructs it with startling accuracy and the beginner cannot. Now scatter the same pieces at random, show the board again for the same five seconds, and the master's advantage very nearly disappears. Same eyes, same thirty-two pieces, same five seconds, and suddenly the expert looks ordinary. 11
What did the master see in the real position that vanished in the random one?
An expert is not simply someone with more facts filed away. An expert is someone whose training has changed the unit of perception. Where the beginner sees individual pieces, the master sees a threat, a pawn structure, a familiar weakness, a shape with a history behind it.
Chase and Simon gave the mechanism its name. Masters do not remember thirty-two isolated pieces; they see a handful of meaningful structures, or chunks. Ericsson and Kintsch later described how this works as long-term working memory: experts encode a problem into stored structures, so a tiny working-memory window behaves as though it were far larger. 11 12
Notice what this does to combination, which is where novelty actually happens. Working memory holds only about four things at once, for the expert and the novice alike. The difference is what each of those four things contains. The novice is combining four small items. The expert is combining four compressed worlds, each one an entire structure. The expert is therefore searching a combinatorial space orders of magnitude larger, not because they are quicker, but because their units are bigger. Depth is the aperture, not the cage. 12
The same thing happens in the body. When Calvo-Merino and colleagues scanned expert ballet dancers, expert capoeira dancers, and non-dancers all watching the same clips, the brain regions involved in action observation responded far more strongly when a dancer watched movements from their own trained repertoire. Same footage, different perception. A trained dancer sees quotations, errors, departures, possibilities; an untrained viewer sees someone moving. 14
This generalizes, and it is why the expert's advantage is nearly invisible from the outside. A mathematician sees a proof step straining under too much weight. A chemist sees a molecule that looks elegant and will fail at synthesis. A software architect sees a beautiful demo that will be unmaintainable in six months. They are not reasoning their way to these judgments so much as seeing them, and the seeing was built by years of contact that cannot be acquired by reading about it.
Economists Cohen and Levinthal gave this capacity a name that will run through the rest of this essay: absorptive capacity, the ability to recognize the value of new information, take it in, and use it. The hard part is that absorptive capacity is built out of prior related knowledge. You cannot hand the job of recognizing value to someone who has not yet learned what value looks like. 18
Information has no value to an agent with no capacity to absorb it.
That sentence is the key to the whole room. When there is no cheap way to check an answer from the outside, the trained reader is the only verifier there is.
Where AI wins, and where it stalls
Sometimes, though, you do not need the reader at all, because the world supplies a fast, mechanical judge that says yes or no on its own. Those are the domains where the machine looks like magic, and it is worth seeing exactly why.
Start on a Go board in Seoul. In March 2016, AlphaGo beat Lee Sedol, one of the greatest Go players alive, four games to one. The famous moment came in game two, move 37, a stone placed so strangely that professional commentators first read it as a mistake. Much later in the game the board revealed what it had been for. The move had been waiting for the future. 71
That result was genuinely astonishing, and it was not a demo or a marketing page. It was a public contest inside a world with rules sharp enough to cut glass. A Go move is legal or it is not; a game is won or lost; a self-play system can play again and again and receive a hard answer every time.
Why could AI become superhuman there?
Because the board was a machine for judgment. The model could generate moves, and the game could punish them. Search had somewhere to go, reinforcement learning had a gradient, and millions of trials were not millions of opinions but millions of encounters with a rule-bound world that returned a verdict.
The tighter the verifier, the more generation turns into progress.
Leave the board and the magic begins to thin.
Protein structure prediction is a real achievement, and AlphaFold mattered because the field had benchmarks, physical constraints, and decades of experimentally determined structures to compare against. A predicted structure is not a medicine, though. A binding hypothesis is not efficacy, and a promising molecule is not a safe drug. Between a protein model and a treatment on the market sit wet labs, synthesis, assays, toxicity, dosing, patients, regulators, failures, and time. The next section on medicine follows that chain in detail. 55 30 52
Code splits the same way, and it is worth being exact, because the phrase AI can code now hides most of what matters. Writing a function that passes its tests is a bounded task with a cheap verifier: the compiler, the type checker, and the test suite say yes or no in milliseconds, so the machine can generate many candidates and let the oracle sort them. A great deal of real software is exactly this, the plumbing and glue and well-worn patterns that a thousand prior examples already solved, and here the machine is genuinely, remarkably good.
Complex software is not a pile of solved coding problems, though. Figma, Ubuntu, or Stripe are architectures: sustained structures of abstraction held together by human intention, in which a thousand decisions about layering, data flow, and interface are made so the whole stays fast, extensible, and comprehensible as it grows, while it also carries invisible obligations like latency, security, permissions, drivers, package dependencies, money movement, and backwards compatibility. An architecture has no cheap verifier. The tests pass, the demo works, the surface looks identical to the real product, and the thing can still be quietly, structurally wrong in a way that only surfaces later as sluggish performance, or a feature that will not integrate, or an extension that forces a rewrite. The user sees the same surface; the behavior under load is not the same. That difference lives precisely where the verifier is slow and expensive, which is precisely where the machine is weakest and the expert is hardest to replace.
This is why productivity is hard to measure even among experts. In METR's early-2025 study, experienced open-source developers working in their own mature repositories expected AI to speed them up, and reported afterward that it had, while the measured time showed a slowdown. The developers predicted a large speed-up, felt a large speed-up, and were in fact somewhat slower, a wide gap between perception and reality on home turf. METR later noted that newer tools and adoption patterns complicate the picture, and cited honestly it still lands: when the verifier is not a simple pass or fail signal, even experts misread whether AI helped. If they cannot calibrate on their own code, a person with no domain knowledge has no calibration signal at all. 40 41
Writing carries the same trap in a softer form. A paragraph can be locally fluent while a long essay has to hold a thesis across distance, and a serious novel has to sustain voice, character, plot, memory, surprise, and reader trust for hundreds of pages. There is no compiler for Anna Karenina, and no unit test for whether a chapter earns the next one.
Mathematics has to be handled with the same care. Formal systems can check fully formalized arguments, but ordinary mathematical prose is not automatically formal. A proof assistant helps once the work has been translated into its language, and even then the choice of theorem, the formalization, the link back to the intended claim, and the community's judgment remain human work. The recent AI-assisted results on Erdős problems are useful case studies rather than proof that mathematics has become easy to automate. They show a division of labor in which the machine searches and mathematicians frame, recognize, repair, and certify. 49
So AI can help the process, and help it a great deal. It can generate candidates, compress first drafts, search a design space, propose tests, write scaffolding, and find analogies, making competent work cheaper across the board. Creating an artifact that matters in the world is a different thing from creating one that looks complete on a screen, because the world does not give a single exam. It gives many, and most of them are slow. Pointing two million agents at a problem does not change this unless the agents have a hard way to know which of their outputs are real. More agents can mean more search, and they can just as easily mean more polished candidates for humans to evaluate. Without a verifier, scale is volume rather than judgment.
The frontier labs reveal the same thing through their hiring. OpenAI advertises roles asking for seven or more years of software engineering experience on research-platform and frontier-systems work. Google DeepMind lists senior machine-learning engineering roles requiring multi-year engineering and AI experience. Anthropic advertises senior and staff-level infrastructure and research-tooling roles. If frontier models had made deep engineering judgment cheap, the labs closest to the technology would be hiring armies of clever beginners and letting the models improve around them. They are paying instead for experienced people who can build, evaluate, debug, and hold complex systems together. 72
This is the jagged frontier. Inside it, AI assistance improves performance; outside it, people accept plausible output and get worse, as Dell'Acqua and colleagues found with a large group of consultants. The catch is that seeing where the frontier lies requires domain knowledge, and the people most likely to wander past it are the least equipped to notice they have. 39
| Domain | Verifier | Prediction |
|---|---|---|
| Go | Legal moves, self-play, win/loss | Superhuman play becomes possible |
| Small coding tasks | Compiler, tests, runtime | Fast local gains |
| Product-grade software | Architecture, security, scale, users, time | Expert guidance remains load-bearing |
| Protein structure | Benchmarks and experimental comparison | Real progress, bounded artifact |
| Drug discovery | Wet lab, clinical trials, regulators | Many costly reality checks |
| Long fiction or serious essays | Reader memory, taste, culture, time | No cheap pass/fail signal |
| Mathematical claims | Expert proof review; formal checks when formalized | Search can help, certification is not automatic |
The pattern is simple enough to be uncomfortable. The easier it is to grade the work, the more powerful generation becomes. The harder and more layered the reality checks, the more valuable trained human judgment becomes.
The apprenticeship montage
That trained judgment is not issued with a diploma. It is grown, slowly, in contact with a domain, and the growing is the part the myth skips. Watch it happen once.
Go back to molecular biology in the late 1970s, before PCR, and the room feels slower. DNA is no longer mysterious. The double helix has been on blackboards for decades, and sequencing, restriction enzymes, gels, probes, plasmids, bacterial colonies, and radioactive film are all standard furniture. The trouble is quantity. What a researcher usually needs is not the whole genome but one small piece of it, a single gene, or the few hundred letters around a suspected mutation, and in a tube of human DNA that piece sits in only a handful of copies, lost among billions of others. The sequence is there, spelled out correctly inside the cell. There is simply not enough of that one stretch to see it, test it, or diagnose from it, and almost everything a researcher wants to do next begins with getting more copies of exactly that segment.
The established way to get them ran through living things. You splice your fragment into a plasmid, coax bacteria to take it up, let the colonies grow overnight while the cells copy your DNA along with their own, then screen the plates and hope the piece you need is somewhere in the dish. It works. It is also slow, roundabout, and full of ways to fail. So the field pushed in the sensible directions, toward better probes, better cloning, better ways to make bacteria carry a fragment long enough to be caught, each of them treating the shortage as a matter of labor. If the target is too faint, build a longer route around it.
Kary Mullis was inside that world at Cetus Corporation in Emeryville, California, though his path there had been crooked rather than ceremonial. He earned his PhD at Berkeley in 1972 after wandering through astrophysics, a spell of writing fiction, and medical research, and came to DNA synthesis late, out of curiosity. By the early 1980s his job was to make oligonucleotides, the short custom strands that other scientists used as primers. A new machine had just automated the work, so primers were suddenly abundant in his lab, things he could design, order, and hold by the fistful, while the group next door wrestled with a stubborn problem: spotting a single changed letter, a point mutation, at one exact position in the genome. He was soaked at once in cheap primers and in the difficulty of pinning down one faint target. 70
Then came the drive. One night in the spring of 1983, Mullis was steering his Honda up Highway 128 toward his cabin in Mendocino County, his girlfriend asleep beside him, turning the mutation problem over in his head. He pictured a primer settling next to his target and a polymerase extending it into a fresh copy, and then a question surfaced: what if a second primer sat on the opposite strand, facing the first, so each one seeded a copy of the other's stretch? That was when he saw the loop. Heat the DNA to pull its two strands apart, cool it so the primers bind, let the polymerase build new strands, and the products of one round become the templates for the next. Run it again and two copies become four, then eight, then sixteen. He knew programming, and he recognized the shape of a reiterative procedure, a result fed back into the same operation again and again until it runs away with itself. He pulled onto a turnout and worked the arithmetic in the dark. Ten cycles would make about a thousand copies, twenty a million, thirty a billion. One faint segment could become as much clean DNA as anyone could ever want, and by his own account he understood, sitting there on the shoulder of the road, that everyone who cared about DNA would want it.
The idea was the easy part. Turning it into a method that worked was a long grind, and that grind is the rest of the story. The first versions of PCR used an ordinary DNA polymerase from E. coli, which carried a fatal inconvenience: the same heat that separated the strands each cycle also killed the enzyme, so a technician had to stand at the bench and pipette in fresh polymerase after every single round, for hours, tube by tube. It worked, and it was tedious enough to stay a niche trick. The fix came from an improbable place. In the near-boiling water of Yellowstone's hot springs lives a bacterium, Thermus aquaticus, whose DNA polymerase keeps working at temperatures that destroy most enzymes. The organism had been described in the late 1960s and its heat-proof polymerase purified in 1976, years before anyone had a use for it. Researchers at Cetus saw what it could do and swapped it in, and now the enzyme survived the heat instead of dying in it. You could mix the reaction once and let it run untended, and automated machines were soon cycling the temperatures on their own. A finicky hand method became an instrument, and a faint segment of DNA became routine to read. 70
Why did that solution become visible to Mullis in particular?
Because it was never, to him, a vague wish for more DNA. It was a sequence of concrete moves he already had in his hands: separate the strands, mark the boundaries with two primers, extend, repeat. Primers were his daily material. The pain of a single-copy target was the problem in the next room. The instinct that a doubling process runs away with itself came from code. Depth did not hand him one extra fact; it handed him a decomposition of the problem that the people treating it as a labor shortage never reached. The pieces had been sitting in the same building for years. He was the one positioned to see how they fit, and even then the insight was only the beginning of the work.
The same hidden apprenticeship is easy to miss in more famous lives, because we remember the flash and forget the decade under it. Shakespeare drilled Latin rhetoric in grammar school until the moves became reflex, then spent roughly twenty years in the London theatre learning what keeps a crowd breathing, and he built Hamlet and Lear out of stories other people had already shaped. Picasso completed a decade of rigorous academic training before Cubism, and Guernica, the archetype of the sudden bolt, survives in some forty-five preparatory sketches that show an incremental, effortful process rather than a single stroke of lightning. Darwin had the essential idea in 1838 and published in 1859, and spent eight of those intervening years on barnacles, earning the taxonomic authority without which Origin would have been dismissed as speculation. Karikó spent close to forty years on modified mRNA, through demotion and defunding, until the 2005 result with Drew Weissman made the later vaccines possible. Ramanujan worked for years in isolation through Carr's Synopsis, and when his results reached England they still required Hardy to validate, contextualize, and in places repair them.
What runs through all of them is contact: books, tools, materials, feedback, rivals, failed attempts, and the humbling intimacy of knowing exactly how a thing breaks. 4 5 60 62 65
This is also where the famous ten-thousand-hours idea becomes useful, once it stops being magical. The number came from Ericsson's work, and Ericsson spent the rest of his life objecting to how it was used. Ten thousand hours is not a threshold, and it is not any ten thousand hours. His construct was deliberate practice: effortful work at the edge of current ability, with feedback that exposes specific errors. Twenty comfortable years in a job is not ten thousand hours of that; it is closer to one year repeated twenty times. And there is an honest correction to keep alongside it. Macnamara, Hambrick, and Oswald ran the meta-analysis and found that deliberate practice explains a real but partial share of the variance, roughly a quarter in games, a fifth in music, less in professions. Practice is not destiny. What that finding measures, though, is variance among people who have already done the work; it is the spread at the top of the distribution, not the gap between the top and the street. Depth remains the entry ticket even where it does not, by itself, decide the winner. 7 8
The more defensible version of the pattern comes from John Hayes, who studied roughly five hundred notable compositions by seventy-six eminent composers and found, with almost no exceptions, that no masterwork appeared before about ten years of immersion. Even Mozart's early pieces read as competent juvenilia; the works that entered the repertoire came after the decade mark. It is now known, unglamorously, as the ten-year rule. 4
A better image than a ladder with a prize on top is a set of eyes slowly learning to focus. At first the beginner sees content. Later they see structure. Later still they see pressure points, the places where an assumption is doing quiet work, where a source is suspiciously clean, where a result is promising but underpowered, where a design will buckle once the second product line arrives. That last stage is judgment, and it is the stage everyone is most eager to skip.
The cancer drug
The most seductive place to imagine skipping it is medicine. Take the story people most want to be true. A teenager sits in a bedroom, prompts a model, and finds a cure for cancer.
The appeal is obvious. It lets intelligence outrun institutions, turns medicine into a locked door that only needed the right sentence, and hands the hero a laptop instead of a billion-dollar lab. The reason it does not work is not that the machine cannot generate a plausible molecule, because it can. The reason is that generating the molecule was never the hard part.
The first generated molecule is the title card, not the end of the movie.
Target identification has to survive the literature, and here the trouble runs deeper than most people outside the lab realize, because written knowledge has two ceilings that both bite at exactly this step.
The first ceiling is tacit knowledge. In the 1970s the sociologist Harry Collins studied laboratories trying to build TEA lasers, and found that none of them succeeded from the published papers alone. Every successful build involved personal contact with someone who already had a working one. The record was necessary and radically insufficient, and the missing knowledge could not simply be written down, because the people who held it did not fully know they had it. Biology is full of that kind of knowledge, along with bad maps, cell lines that lie, assays that drift, mouse models that flatter, and human bodies that refuse to behave like diagrams. 31 32
The second ceiling is that the written record is substantially wrong in places. When Amgen scientists tried to reproduce fifty-three landmark preclinical cancer papers, they succeeded with six; a similar exercise at Bayer landed in the same range. Much of the map has landmarks in the wrong place, with no annotation saying which, and the knowledge of which results to distrust lives socially, in labs and corridors and the shared judgment of people who have watched things fail to replicate. 33 34
Both ceilings fall on the machine at once, because the model's picture of cancer biology is distilled from that same literature. It has no reliable way to know which parts of its own map are fiction, and it will describe the fiction as fluently as the fact. The expert's most valuable asset is often knowing which parts of the record are wrong, and that asset is, almost by definition, not in the record.
After the molecule comes everything else. Synthesis, assays, pharmacokinetics, toxicology, an IND filing, then Phase I, II, and III, then regulators, then doctors, then patients whose lives cannot be debugged after deployment. Oncology is the harshest version of this verifier. Across large clinical-trial datasets, the odds of a compound moving from Phase I to approval are extremely low and the median time in the clinic runs well over a decade. Reality returns its verdict slowly and at enormous cost. 30
As of mid-2026 the strongest AI-originated drug-discovery story is still not a story about a pharmacy shelf. Insilico's rentosertib has reached Phase III for idiopathic pulmonary fibrosis, which is a major milestone, and the company's own release notes that the drug remains investigational and has not been approved by any regulator. Getting there took roughly a decade, hundreds of scientists, robotic wet laboratories, an IPO, and industry partnerships. The single most successful AI drug-discovery effort on Earth still required a company full of experts, physical laboratories, and ten years. It was not a prompt. 52
In medicine, the verifier is not a vibes check. It is a decade of biology asking whether you were serious.
The economics of the room
Oncology is the extreme case, where the verifier is not just slow but catastrophic. Most of the world sits between that and the Go board, and the general pattern shows best not in a single hard trial but in an ordinary room full of ideas.
Imagine a grant committee at 11:47 at night. The lights are too bright and the coffee has gone cold beside a stack of applications. Every folder holds a plausible problem, a polished method, and a confident paragraph about impact. The reviewers are not sorting nonsense from genius. They are deciding which of a hundred reasonable ideas deserves a decade of work, a laboratory, a team, and millions of dollars.
At first the room feels rich with possibility. Then the arithmetic turns. If every existing idea can be combined with every other, the number of possible projects grows faster than any committee can read, test, or fund. The scarce thing stops being a candidate and becomes the capacity to tell which candidate is worth carrying into reality.
How many good ideas can a room evaluate before the room itself becomes the bottleneck?
Martin Weitzman called this recombinant growth. Ideas become raw material for further ideas, and the possible combinations multiply far faster than the people who must evaluate them. The implication is easy to miss: an explosion of possible combinations does not produce an explosion of progress. It produces an explosion of triage. 19
The binding constraint on innovation is the human and institutional capacity to evaluate, test, and develop an astronomical supply of possible combinations.
A small case makes the abstract point physical. Scott Shane gave different entrepreneurs access to the same MIT technology, and they did not see the same opportunity inside it. Each noticed a different use, and which use each one saw was predicted by what they already knew. The technology was identical; the opportunities were not, because prior knowledge had reshaped the world each person could perceive. Opportunity recognition is domain knowledge in action, not a general talent, and it is absorptive capacity doing its work one more time. 25
Scientific novelty follows the same shape. In an analysis of 17.9 million papers, the highest-impact work tended to pair a deep conventional core with a small tail of unusual connections. Novelty paid when it was anchored in enough established knowledge to be understood, tested, and built upon; unfamiliar combinations on their own carried more variance and lower average usefulness. Fleming found the complement in patents, where recombining unfamiliar components raised the variance of the outcome and lowered its mean usefulness. Originality pays when it is grounded in mastery, and on its own it is mostly noise. 20 21
Now widen the view from the single committee. Benjamin Jones found that the age at first invention has risen, teams have grown, and specialization has deepened as the stock of knowledge has piled up. Bloom and colleagues found the macro consequence: ideas are getting harder to find, and research productivity is falling. Progress increasingly asks more people to carry more of the map before anyone can see the next opening. 22 23
The cultural story runs the other way, forever searching for the twenty-two-year-old who wins by refusing to learn the old rules. The administrative data is less cinematic. In a Census-based study of 2.7 million founders, the mean age among the fastest-growing top 0.1% was about forty-five, and prior experience in the specific industry was the strongest single predictor of success; founders in their early twenties had the lowest likelihood of a successful exit. Youth brings energy and some freedom from convention. It does not repeal the burden of knowledge. 24
The young genius is a memorable exception. Experience is the quiet base rate.
Depth is not the same as infallibility, and it is worth being exact about where a single expert's judgment can be trusted. Kahneman and Klein found that expert intuition earns its keep mainly where feedback is fast and unambiguous, as in chess or firefighting. The domains this essay cares about most, science, engineering, and the building of a company, are the opposite: feedback there is slow, noisy, and rarely repeated, and a lone expert's gut is not an oracle. Their intuition may frame the question better than anyone else's, and it still has to be treated as a hypothesis to be tested. 27
That is exactly why the reliable unit is a field rather than a person. A hard problem is not handed to one deep expert who then pronounces on it. It goes to a room of experts from different subdomains, each blind where the others can see, who argue over the proposal, decide what gets funded, and try to break the result, and it becomes knowledge only once independent teams can reproduce it. A field also changes its mind in a way a machine does not, and the difference is worth naming. A community of experts can be overturned by a single decisive, reproducible result, so that one clean experiment retires a theory a thousand papers had assumed. A language model's picture of a field is a weighted average of everything it was trained on, and in that average a lone critical result carries almost no weight, so the very evidence that would flip a human field tends to wash out. 73 54 The field updates hard on the exception; the model regresses to the mean. None of this makes evaluation cheaper. Needing a whole community of deep, independent people rather than one appointed oracle makes the capacity to judge scarcer still.
So the pile is not valuable because one folder happens to contain a good idea. It is valuable only if the room can find that folder before its attention, budget, and patience run out. When a machine begins printing ten thousand plausible folders a day, the economics do not shrink the room's importance. They make the room the entire story.
The false shortcut
If the room is the whole story, the obvious objection walks in wearing a lab coat. Can the machine not simply be the room? Run three models, make them debate, build a panel, let one generate, one criticize, one judge. It sounds like peer review, and it works in demos because the transcript looks like deliberation.
Peer review is not several voices, though. It is a machine for disagreement: independent people, trained differently, with incentives to catch one another's mistakes, checking claims against a reality outside the conversation. Stacking similar models reproduces the voices and none of that machinery, and it is worth being concrete about why.
Start with the math behind the wisdom of crowds. Condorcet's jury theorem says that a majority vote grows more reliable as the crowd grows, but only under two conditions: each voter has to be better than a coin flip, and the voters' mistakes have to be independent. The second condition is the one that breaks here. 37 When everyone shares the same blind spot, extra voters add no new information, and you have multiplied one opinion rather than assembled many. Frontier models share training corpora, architectures, post-training recipes, and a great deal of distillation from one another, so three of them agreeing is closer to one reasoning process echoed three times than to three independent checks. 82
This is not a hypothetical, and it has a name. Kleinberg and Raghavan studied what happens when many decision-makers lean on the same algorithm, using hiring as their example, and found a result that surprises people: even when the shared algorithm is more accurate than any single firm's own judgment, everyone adopting it can make the system as a whole worse. The same applicants get screened out everywhere at once, and the independent second look that would have rescued a good candidate never happens. They called it algorithmic monoculture, after farming, where a single high-yield crop is efficient and uniquely exposed, so one blight can take the whole field. 46 Bommasani and colleagues carried the idea into modern AI, where systems increasingly share components like training data and a common base model, and measured what they named outcome homogenization: the same individuals handed the same result across unrelated deployments. In a later analysis of millions of real job applications screened by tools from a single vendor, some applicants who applied to many positions were flagged for rejection at all of them, more often than chance would allow. That is what correlated failure looks like when it reaches people. Not scattered misses that average out, but the same miss, everywhere, at the same time. 83
A small, comic failure shows the deeper problem. Ask a model whether you should walk or drive to a car wash that is a hundred meters away, and many models are pulled toward walk by the distance cue, missing the constraint that the car has to be at the car wash. When researchers built this into a benchmark, they found the distance cue outweighed the goal by a wide margin, that the failure looked like keyword association rather than compositional inference, and that no model cleared a modest bar under strict evaluation, with the presence constraint that the object must physically be there sitting especially low. A structured reasoning scaffold lifted performance dramatically, which is to say a human still has to supply the frame. The model can possess the fact and fail to apply the constraint. 50
Scale that up and it stops being funny. A model can hold sentences about clinical-trial design and still miss the constraint that makes a proposed trial useless. It can hold sentences about distributed systems and still miss the abstraction that fails under load. It can hold sentences about proof technique and still let one hidden assumption carry the whole bridge. Banerjee offers an architectural version of the same limit: a model can interpolate within the captured horizon of its training experience, and recombining that record does not by itself supply an unobserved, long-horizon constraint, so when a proposal turns on such a constraint, verification still runs through an expert community with access to the relevant reality. 73
Three models sharing a blind spot do not become three independent judges. They become one blind spot with better formatting. And the loop closes badly, because models can prefer their own style of answer, and systems trained recursively on generated output degrade. A verification system cannot live forever on its own exhaust; it needs contact with independent human and physical ground truth. 45 54
The vanishing apprentice
So the machine needs the trained human after all. The danger runs deeper than that, because the machine may quietly reduce how many trained humans get made.
A junior analyst used to read the ugly source documents. A junior scientist used to fight the assay. A junior engineer used to trace the bug through the codebase. A student used to sit with the blank page long enough to discover which part they did not understand. Those tasks are slow, annoying, and often beneath the glamour of a field, and they are also how judgment gets installed.
Lisanne Bainbridge named the pattern in 1983, in a paper on the ironies of automation. The more reliable the automation, the less the human practices, so the more their skill decays, and the human is kept on precisely to handle the cases the automation cannot. Automation therefore tends to leave the operator least capable at the exact moment they are most needed. Air France 447 is the textbook case, an autopilot disengaging in conditions it could not handle while the crew had lost the manual recovery skill they were there to supply. 43
Modern knowledge work shows the same shape. In survey work from Microsoft Research and Carnegie Mellon covering hundreds of real workplace uses of generative AI, workers reported less cognitive effort across most tasks, and two findings sit at the center: higher confidence in the AI went with less critical thinking, and higher confidence in one's own ability went with more. The work shifts from doing to supervising, and supervision is not an entry-level skill. The people least equipped to check are the ones who check least. 42
This is the loop that should make everyone uneasy:
We need experts to verify machine output.
Experts are made by doing the hard parts.
The machine is increasingly used to remove the hard parts.
The system consumes its own precondition.
The labs price this with their own budgets. Expert-evaluation markets pay far more for physicians, lawyers, PhDs, executives, and domain specialists than for general annotation, because the frontier model still needs scarce human judgment as ground truth. The rate card spans two or three orders of magnitude, from a few dollars an hour for bulk labeling to hundreds or thousands an hour for credentialed specialists and senior operators, and the top rungs are paid to judge rather than to generate. Providers maintain tens of thousands of doctorate-holding contractors, expert rates are rising, and the whole arrangement exists partly because models trained on their own outputs degrade, so the loop must be replenished with fresh human ground truth. Generation got cheap; verification did not. 53
There is a caution worth stating, so this is not mistaken for reassurance. That experts are paid more does not mean experts are respected more. The people paying the highest rates work closest to the technology and have been burned by its failures often enough to see what verification is worth. The general public sees the polished surface and never the caught error, and tends to draw the opposite conclusion. A rising price in a thin expert market can sit alongside a falling public regard for expertise, and it is the second, not the first, that decides whether the next generation of experts gets trained. The market is bidding up the very thing the culture keeps telling young people they no longer need to become.
The outsider who wasn't
There is an objection this whole argument has to meet head on, because it is the one people reach for first. The story goes that real breakthroughs come from outside. Insiders are captured by their field's assumptions, and it takes a newcomer, unburdened by the orthodoxy and reasoning from first principles, to see what everyone else missed. If that is how progress works, then depth is a cage, and the amateur with a good model is exactly who we should bet on.
It is a beautiful story, and it survives mostly by cropping the photograph.
Start with the case people cite most. SpaceX is the modern parable of the outsider: a software entrepreneur with no aerospace training walks in and rebuilds an industry that Boeing and Lockheed had run for decades. Elon Musk did study physics, and he did bring a first-principles habit of pricing a rocket by the raw cost of its materials rather than the going rate. But look at who he hired before the first engine existed. His first employee was Tom Mueller, a propulsion engineer who had spent roughly fourteen years at TRW and had built the largest amateur liquid-fuel engine in the world in a friend's warehouse. Structures went to Chris Thompson, out of McDonnell Douglas. Avionics went to Hans Koenigsmann, out of Microcosm. The person who would come to run the company, Gwynne Shotwell, had a decade at the Aerospace Corporation behind her. Musk supplied the capital, the mission, and the cost logic, then assembled a room of the deepest aerospace people he could find and immersed himself until he could argue engine design with them. SpaceX was not depth defeated by an outsider. It was one kind of depth, in software and business and hard-nosed cost engineering, aimed at a problem and fused with a great deal of borrowed aerospace depth. 75
Stripe tells the same story in a different key. The Collison brothers are cited as college kids who reinvented payments, an industry owned by banks and processors they had never worked inside. But Patrick Collison won Ireland's national young-scientist prize at sixteen for work in Lisp, left for MIT, and had already built and sold a software company before he was twenty. John was cut from the same cloth. They were not outsiders to anything that mattered here. They were unusually deep software engineers who had felt, in their own hands, how painful it was to accept a payment on the web, and they rebuilt that experience as a few lines of developer code. The banking, licensing, and compliance depth they did not have, they hired, because a five-person startup cannot personally become a licensed money transmitter in fifty states. What looked like an outsider storming a fortress was a domain expert attacking the one seam where his expertise gave him an edge, and buying the rest of the expertise he needed. 76
Notice what the founder's own expertise usually is in these stories, because it is the part most often misread. It is frequently not deep knowledge of the industry being disrupted. It is deep knowledge of how to turn a technology into something a market will actually take, together with enough technical depth to tell which experts are worth listening to and which are bluffing. That is a real and rare expertise, built over years, and it is not the same thing as having no expertise at all. Beneath the visible founder, the load-bearing work is very often done by people with profound depth in the hard problem, whose names never reach the magazine cover.
Now run the film the other way, because the failures are the clean test. Theranos is the outsider myth taken literally. Elizabeth Holmes left Stanford after two semesters of chemical engineering and set out to reinvent blood diagnostics, a field with a brutal physical verifier: the assay either returns the right number or it does not. She had the vision, the narrative, and the turtleneck, and she did not have the depth, so when the depth mattered she overrode it. A Pfizer scientist who reviewed the technology found the claims implausible and advised against working with the company, and the response was to add pharmaceutical logos to a report to manufacture an endorsement that did not exist. When the devices produced unreliable results, the company hid the failures behind staged demonstrations and quietly ran samples on other manufacturers' machines. None of it changed what the blood actually did. Regulators eventually found immediate jeopardy to patient safety, patients received false results, the company collapsed, and Holmes was convicted of fraud. The tire hit the road, and reality, which does not read pitch decks, returned its verdict. 77
That is what the pure outsider story looks like when it is true: not a genius unburdened by convention, but a person without the depth to know which of her own results to believe, in a domain where being wrong is not a rounding error. The distance between Holmes and the Collisons is not audacity or vision, since both had plenty. It is that one side had the depth to tell a real result from a plausible one, and the discipline to defer to the experts who held the depth they lacked.
The same lesson arrives from a sharper angle, because the most instructive failure here was authored by the hero of this section. In 2013 Musk published the Hyperloop Alpha white paper, a first-principles design for a "fifth mode" of transport: passenger pods fired through near-vacuum tubes at close to the speed of sound. It was reasoning from fundamentals at its most seductive, and Musk did not build it. He released the concept and let others chase it. Over the next decade the best-funded of them, Hyperloop One, later Virgin Hyperloop, raised more than $450 million, built a test track in the Nevada desert, ran a single short passenger test that topped out near 100 miles per hour rather than 700, never signed a contract to build a working line, pivoted to freight, and shut down at the end of 2023. By early 2026 the main European efforts had followed it into bankruptcy, and the sector was effectively gone. The physics of holding a hard vacuum over hundreds of miles, the thermal expansion of the tube, the cost of the right-of-way, and the economics of moving enough people to pay for any of it never closed. 78
Set that beside SpaceX and the real variable stands out. Same founder, same instinct to reason a design up from physical fundamentals rather than from what the industry already does, opposite outcomes. The difference is not the quality of the reasoning on the page. It is that SpaceX paired the reasoning with a room of aerospace veterans and a decade of building, launching, failing, and rebuilding against a verifier that answered every attempt, while Hyperloop stayed a derivation on paper handed to a crowd. Reasoning from fundamentals is a method for generating a candidate, not for confirming one, which puts it on the cheap side of this essay's divide. Whether the candidate holds is settled elsewhere, by the physical world or the market, and reaching that verdict takes people deep enough to build the thing and run the test. An elegant derivation buys no exemption from that step. It only marks the step as worth taking.
The research on who actually wins from the outside says the same thing, and it is worth being precise, because the myth leans on a real finding and then misreads it. Jeppesen and Lakhani studied 166 broadcast science problems that R&D-heavy firms posted to more than twelve thousand outside solvers, and found that the winners really did tend to come from the margins: the further a solver's own field sat from the field of the problem, the more likely they were to crack it. But read the finding closely. The winning solvers were scientists with deep expertise, applying it from an unexpected angle. Marginality helped because it supplied a different trained toolkit, not because it supplied a blank slate. The Census work on founders points the same way, with the highest-growth founders averaging around forty-five and prior experience in the relevant industry the strongest predictor of success, and Shane found that which opportunity a person even sees inside a new technology is set by the knowledge they already carry. The outsider who succeeds is almost always an insider of somewhere else. 74 24 25 28 29
So the objection, examined, changes shape. It is true that the person who reinvents a field is often not a veteran of that field. It is not true that they arrive empty. They arrive with deep expertise built elsewhere, a nose for turning it into something the world will use, and the judgment to assemble and trust the experts who hold the depth they lack. AI can hand the aspiring outsider more raw material than ever, and it can help a deep person move faster across the seam where their expertise applies. What it cannot supply is the depth itself, or the trained judgment that tells a founder which expert is right when the experts disagree. Strip those away and you do not get the next SpaceX. You get the confident reinvention that meets reality and loses.
Back to P = NP
Return to the opening scene, a week later.
The LinkedIn post is still there. The Reddit PDF is still being argued over. The GitHub repository has more stars, and a new issue reports that the claimed Lean formalization does not prove the theorem stated in the README. The website has been translated, and the YouTube series has spawned response videos. None of this is yet the conventional announcement anyone imagined. It is a public queue of answer-shaped objects.
What would it take for one of them to become knowledge?
Not attention, which is easy now. Not confidence, which can be generated. Not length, diagrams, citations, model names, or a calm explanatory voice. The claim has to move from the cheap side of the verifier map to the expensive side, where people can say exactly what has been proved and why the hard step has not merely been renamed.
The field would first have to identify the claim. Is it a proof that P = NP, a proof that P is not NP, a conditional result, a misunderstanding of the problem, or a clever reduction that buries the hard step inside a definition. Then experts would have to read it, not skim it, with the full weight of professional attention: test the definitions, isolate the load-bearing lemmas, check the reductions, hunt for smuggled assumptions, try to break it, try to formalize it, try to explain it to one another. And the proof would need a community around it, independent readers with different priors and real incentives to find the flaw, formal tools where they genuinely apply, public standards, and time.
A formal checker can be part of that story, but there is no button that compiles ordinary mathematical prose into truth. Someone has to choose what to formalize, understand what the formalization means, connect it back to the domain, and decide whether the proof has answered the question people actually care about. 49
This is why P = NP is not like the Go board. The board punishes a bad move at once, and a compiler rejects a broken program at once, and a clinical trial rejects a drug slowly and brutally and at great expense. A major proof claim lives in a stranger space, where there is a right answer and the path to recognizing it runs through trained human judgment, through formalization where it applies, and through a community strong enough to withstand a flood of plausible falsehoods. 30 71
So the question in the title is not a joke. We might know, because a field still exists that can know. Or we might not know quickly, because the answer arrives inside a pile so large that the field's reading capacity becomes the bottleneck.
Where there is no cheap external verifier, the verifier is a trained human community.
Ask who, in practice, is likely to write the proof the field eventually accepts, and the romantic answer starts to look unlikely. The world holds very few people who can verify a proof of P = NP, and their attention is finite. Reading a million claims to the end is not something that scarce set of readers can afford. So the recognized proof, when it comes, will most plausibly come from inside the field: a young mathematician, quite possibly working with AI to close the gaps, whose work reaches people who will actually read it because the author is already legible to them. The most economical way a field has to survive a flood of originality is to route the work back through its experts, and to let its experts be the ones who carry a result into the world.
That leaves an uncomfortable possibility. A correct proof may already be sitting somewhere in the outsider pile, posted before any credentialed paper appears, and the field may never find it. Not because anyone stole it, and not because the gatekeepers are corrupt, but because checking every claim is not affordable, and recognition is rationed to the candidates a scarce set of readers can reach. The proof that becomes knowledge is not necessarily the first one that was true. It is the first one that reaches a reader who can confirm it.
This is the strange thing the machine does. It floods the experts with plausible proofs, and in the same motion it makes those experts the only instrument that can still pull signal from the noise. The more convincingly the machine can counterfeit a result, the more the scarce human reader matters, and the more the reward gathers around whoever the field already trusts to read.
So it is worth being clear about who the era actually rewards. The machine may well let a regular person generate something extraordinary. Generation, though, is not recognition, and recognition is where the reward lives. The value still settles on the people who can verify a result, vouch for it, and move it into the world, which means the circle of who benefits does not widen. It narrows. The tool hands almost everyone the power to produce and concentrates the power to be believed. It shrinks the share rather than expanding it.
None of this argues against the tool. The Erdős results are real, the protein structures are real, and the everyday usefulness is large. The argument is against a misunderstanding of it, because the misunderstanding is what does the damage. Taken for what it is, a stupendous generator whose output is only as valuable as the judgment that selects and verifies it, the machine pairs with human expertise and the pair is stronger than either alone. Taken for a replacement of that judgment, it invites us to dismantle the very thing that lets us tell its good output from its bad. Knowing the limits precisely is not anti-AI. It is the precondition for using AI well.
The people who make the things that matter in this era will not be the ones with the best prompts. They will be the ones who can read the pile, because they did the years in contact with a domain's real constraints and real feedback, long enough to have built the perception, the chunks, the critic, and the taste. That community is not decoration around a breakthrough. It is the instrument by which a breakthrough becomes knowledge.
Which leaves the obvious, comfortable reply: we do not all need to become experts, we only need to keep a few readers on hand to check what matters. It sounds like a solution, and it is where this argument hands off to the next essay, When AI Makes the Ads, Who Becomes the Art Director?, which asks what happens to the readers themselves. Verification turns out not to be a switch that a few experts can flip on demand. It is a scarce resource that can be exhausted, and the people who can perform it are made slowly, through markets, classrooms, and laboratories that we are now, with the best intentions and impeccable local logic, quietly starving. We are producing pages faster than we ever have, and readers more slowly than we ever have. Whether we keep making them is the question the pile leaves behind.
Predictions
An argument this strong should make bets. Here are mine.
-
The first approved AI-originated drug will come from an organization dense with domain experts, physical laboratories, clinical infrastructure, and regulatory skill.
-
AI progress will stay fastest where outputs meet usable external feedback: tests, compilers, simulations, clear scoring functions, and formal systems once the relevant work has actually been formalized.
-
If P = NP is ever proved or disproved, the proof will emerge from people who understand complexity theory well enough to verify it, and they will be the ones to recognize it, not someone with no such background simply prompting a chatbot.
-
In domains without usable feedback loops, AI will produce more plausible candidates than institutions can responsibly evaluate.
-
Mean output quality will rise while collective diversity falls, most visibly in writing, design, pitch decks, code scaffolds, and research ideation.
-
The market price of verification skill will rise even as public respect for expertise grows more fragile.
-
Model ensembles will help on bounded tasks and show diminishing returns wherever the failure mode is shared constraint blindness.
-
If AI truly substitutes for accumulated expertise, the age and experience burden of top innovators should fall. If it does not fall, the expertise bottleneck is still binding. Watch it; if it falls, I am wrong.
The prediction I care about most is the one underneath all of these: societies that protect apprenticeship will outperform societies that only buy tools. The pile is coming either way. The future belongs to whoever can read it.
Sourcing Note
This version keeps the original evidence base and changes the pacing. Load-bearing empirical claims, including clinical-trial success rates, preclinical reproducibility, AI-assisted creativity, the METR productivity findings and their follow-up caveats, the Toner-Rodgers research-integrity event, and rentosertib's Phase III status, should be checked against primary or primary-adjacent sources before any publication update. The Erdős-problem material is treated as an illustrative case study rather than statistical evidence that mathematics has become easy to automate or verify. Quotations are kept short, and the argument does not lean on any single anecdote carrying more weight than the cited evidence can bear. 30 33 38 40 41 49 51 52