October 1, 2026
    Sajjad Khazipura
    Verification, Hallucinations, Karl Popper

    Why Verification Is AI's Next Category

    Third in a series, following “Computational Knowledge and Reasoning: From Plato to Production Systems” and “Knowledge in Motion: From Plato to Heraclitus.” This essay draws on Karl Popper.

    A large language model is not a database, much less a knowledge system. At its core it is a pattern matcher with imagination, good grammar and polite manners — trained to predict, then tuned to please. Call it a conjecture machine: it proposes plausible claims — conjectures, in Karl Popper’s sense — and can refute them only where someone has built it an external checker. The industry has now spent well over a trillion dollars on the infrastructure behind the most powerful conjecture machines in history. This essay reviews what that money bought, where it fell short, and the new category of products that will unlock the rest: verification — external checkers, domain-knowledge verifiers, reasoners and calculators that formalize the ability to refute, so conjecture machines can be trusted with consequential work.

    AI infrastructure spend

    This accounting excludes every forward projection and counts only what companies have reported. The four largest hyperscalers — Microsoft, Alphabet, Amazon and Meta — spent about $508 billion on capital expenditure from 2022 through 2024, and about $376 billion more in 2025 alone; Oracle adds $20–55 billion a year on top (Microsoft, Alphabet, Amazon, Meta, Oracle earnings releases). Not all of this spending is for AI, but Amazon, for one, attributes the surge to it: its increase “primarily reflects investments in artificial intelligence.” And 2026 is running at nearly twice the 2025 pace: first-half spending by the four reached $295 billion, up from $160 billion a year earlier (Microsoft Q3 and Q4, Alphabet, Amazon, Meta releases). Altogether, these four companies alone have spent nearly $1.2 trillion since 2022 — most of it in the last 24 months.

    On top of that rides the capital raised by the model builders themselves. Third-party trackers estimate $114 billion flowed into AI companies in 2024, $202 billion in 2025, and an extraordinary $399 billion in the first half of 2026 alone (Crunchbase and Dealroom estimates), headlined by single rounds larger than the GDP of small nations. Much of this cycles back into the compute layer, so the total actually spent on the large language model buildout to date is well over a trillion dollars.

    Against this spending, enterprise AI revenue reached an estimated $37 billion in 2025 (Menlo Ventures survey estimate). No general-purpose technology in history has seen such a wide gap between capital employed and returns measured at this stage of investment. That gap is either the greatest deferred payoff ever, or a signal that the employed capital is not delivering. The evidence below suggests a specific answer: while we are attempting to optimize generation, what’s conspicuous by its absence is verification.

    Capability scaled; reliability didn’t

    On raw capability, scaling delivered, and continues to deliver. On reliability — the property enterprises actually price — the curve has visibly flattened. Reliability failures come in several kinds: fabrication — hallucination proper, inventing things out of thin air — factual errors about real things, unfaithfulness to a source the model was given, flawed reasoning, and broken rules. The evidence below spans all of them.

    One clean signal comes from a 2026 study of frontier models writing code (Churilov, May 2026): the five 2026 frontier models tested invent non-existent software packages 4.6 to 6.1 percent of the time, and none has surpassed the best model of 2024 (3.6 percent) on that metric. Two years of unprecedented investment compressed the spread between models — the worst got much better — but did not lower the floor. That is the signature of an asymptote.

    The asymptote is task-shaped. Even on grounded summarization, where the source sits in front of the model, most frontier models still introduce unsupported claims into more than 10 percent of summaries; only the best reach low single digits (Vectara, Nov 2025). Multi-step agentic workflows compound errors: once an agent accepts a false premise, later steps inherit it, and agents repair such errors only 36 percent of the time (Tracing the Cascade, Aug 2026). Even GPT-4o succeeds on all eight repeated attempts at the same retail task less than a quarter of the time (τ-bench, 2024). Meanwhile the externality grows in the wild: an audit of 111 million references estimates nearly 150,000 fabricated citations entered the scientific literature in 2025 alone, spread across many papers rather than concentrated among a few bad actors (Zhao et al., May 2026). Deployment is scaling faster than reliability.

    Most telling of all: the frontier labs have stopped promising zero. OpenAI’s own research argues that hallucination is structural — next-token prediction, trained against benchmarks that penalize “I don’t know,” systematically rewards the confident bluff over the calibrated abstention (Kalai et al., Sep 2025). Some reasoning-optimized models hallucinate more than their predecessors: on OpenAI’s PersonQA test, o3 hallucinated 33 percent of the time and o4-mini 48 percent, against 16 percent for the older o1 (OpenAI system card, Apr 2025). The field’s own theory now predicts what its own measurements show: you cannot scale your way out of a problem that scaling’s objective function is rewarding.

    Independent measurement confirms the bluff. Artificial Analysis’s AA-Omniscience benchmark measures how often a model answers wrongly when it should have admitted it did not know. On its current leaderboard, today’s leading frontier models from OpenAI, Anthropic and Google do so in roughly 45 to 75 percent of such cases. Spending more does not buy restraint: the most expensive models per task score no better than cheap ones, and the lowest rates belong to inexpensive models (Artificial Analysis, accessed Sep 29, 2026).

    Jagged intelligence: the productivity evidence

    Productivity tells a subtler story — not diminishing returns, but jagged ones, and jagged in a revealing pattern.

    The most rigorous longitudinal evidence is shifting: a controlled trial found experienced open-source developers took 19 percent longer with AI tools in early 2025 (METR, Jul 2025); seven months later, METR tentatively estimated a speedup, while calling its data “very weak evidence” (METR, Feb 2026). Randomized studies show real gains — 14 percent more issues resolved per hour in customer support (Brynjolfsson et al., QJE 2025), 40 percent less time on professional writing (Noy & Zhang, Science 2023), and 55.8 percent faster on a scoped coding task (Peng et al., 2023). The frontier is genuinely moving.

    But three findings frame those gains. First, perception runs far ahead of measurement: in METR’s trial, developers expected to be 24 percent faster, believed afterwards they had been 20 percent faster, and were actually 19 percent slower — an overestimate of roughly 40 percentage points (METR, May 2026). Second, the gains are concentrated precisely where errors are cheap: large savings on routine tasks collapse to single digits on complex work, and reverse entirely outside the model’s capability frontier — the “jagged frontier” documented since 2023 (Dell’Acqua et al., HBS 2023) and still visible in every serious study since. Third, the aggregate effect remains modest: AI users report saving 5.4 percent of their work hours, but across all workers the figure is about 2 percent (St. Louis Fed, Feb 2025; FRED Blog, Aug 2026).

    Read the reliability data and the productivity data together and they turn out to be one finding, not two. Productivity gains flourish exactly where errors are tolerable, and stall exactly where trustworthiness binds — law, medicine, engineering, finance, operations. The productivity ceiling is the reliability floor. A trillion dollars of compute moved the first curve and left the second nearly where it was.

    A new “engineering” every season

    The industry has responded to this gap with a proliferating vocabulary. Prompt engineering. Context engineering. Loop engineering. Graph engineering. Harness engineering. Each term arrives with the confidence of a discipline and the half-life of a fashion.

    It is easy to mock the churn. It is more useful to read it as evidence. Each of these terms is the industry discovering, one layer at a time, the same underlying truth: the model is not the system. Prompt engineering discovered that the conjecture machine is steerable. Context engineering discovered that what you put in front of it matters more than how you ask. Graph engineering discovered that structured knowledge beats prose recall. Harness and loop engineering discovered that models must be wrapped in scaffolding — retries, tool calls, checks, escalations — before they can be trusted with consequential work.

    Every one of these disciplines is real, and every one of them shares a limitation: they optimize the inputs to generation. They make the conjectures better. None of them, by itself, refutes anything. A perfectly engineered prompt, fed perfectly engineered context, through a perfectly engineered harness, still terminates in an unverified claim — a bolder, better-dressed conjecture. This is why errors — hallucinations included — have not gone to zero under any of these regimes and will not: the missing half of Popper’s loop cannot be supplied by improving the first half.

    Generation is not verification

    Here is the structural diagnosis. A language model is an extraordinary generator — arguably the most powerful abductive engine ever built, proposing hypotheses, drafts, plans, and completions at negligible marginal cost. But generation and verification are different computational faculties. Verification requires an authority outside the generator: a knowledge base to check claims against, a logic to test inferences with, a constraint system to enforce, a provenance chain to audit. The generator cannot grade its own homework, and every architecture that asks it to — self-critique, self-consistency, “think step by step” — inherits the generator’s failure modes at a higher price point. Even a judge from a different model family is another conjecture machine: its errors are less correlated, but it has nothing to check against.

    Subbarao Kambhampati and his colleagues gave this diagnosis its sharpest formulation in the LLM-Modulo framework (Kambhampati et al., ICML 2024): treat the language model as a powerful idea generator inside a loop with sound, external verifiers and critics — symbolic reasoners, formal checkers, domain models — that test each conjecture and feed refutations back for revision. The model proposes; the verifiers dispose; the loop iterates until the output survives scrutiny. It is Popper’s epistemology rendered as a system architecture: conjecture from the neural component, refutation from the symbolic one.

    This is where our first two essays converge with the economics. Verification is only as strong as the explicit knowledge it verifies against. A verifier needs facts with provenance — where did this claim come from? It needs proof structure — what does this conclusion depend on? It needs time — was this true when it mattered, and what did we know when? That is the knowledge unit we formalized as κ = (S, P, O, τ, μ, J), and it is why explicit knowledge representation and reasoning — ontologies, knowledge graphs, temporal validity, justification chains — is not a detour from deep learning but a prerequisite for verification: one that augments deep learning rather than replacing it.

    And the evidence says verification is where the hard work — and the value — now lies. On reasoning and planning tasks, when language models critique their own work, performance collapses; when a sound external verifier checks the same outputs, performance rises significantly (Stechly, Valmeekam & Kambhampati, ICLR 2025). More compute has not bought trust — the most expensive models bluff no less often than cheap ones. Scaling offers a seductive recipe — more data, more compute, more money, a better model — and for capability it has worked. Trust follows a different recipe: grunt work and first-principles science. There is no substitute for the effort: modelling domain knowledge, building verifiers, optimizing the pipeline, and discovering and fixing the corner cases one by one.

    A new category

    Markets eventually organize around binding constraints. When compute was the constraint, the market built a compute industry, and it is now worth trillions. Trust is the constraint today, and a market category is forming around it: verifiers, reasoning layers, grounding infrastructure, claim-checking pipelines, knowledge substrates built for justification rather than mere retrieval. The category does not yet have a settled name — which, as the stream of new “engineering” disciplines shows, has never stopped an industry before. But its function is settled: it is the refutation half of the loop, sold as infrastructure.

    The inventiveness now flowing into this layer — neuro-symbolic architectures, formal verification of model outputs, graph-constrained generation, multi-stage claim verification — is doing for trust what the last four years of capex did for capability. It will not make errors impossible; Popper would remind us that no finite process certifies truth. What it makes possible is something more valuable and more honest: systems that can show their justifications, expose their evidence, time-stamp their beliefs, and catch their own fabrications before a customer, a court, or a clinician does.

    Over a trillion dollars bought the conjecture machine. The next era belongs to whoever builds the refutation machine to sit beside it. That is not a diminishing return. It is where the returns went.

    In homage: Karl Popper (1902–1994)

    This essay borrows its vocabulary from Karl Popper. Born in Vienna and later a professor at the London School of Economics, Popper argued in The Logic of Scientific Discovery (1934; English 1959) that what makes a theory scientific is not that it can be proven, but that it can be refuted. In Conjectures and Refutations (1963) he described knowledge as growing through bold guesses exposed to severe tests. He never saw a language model, yet he described its proper place exactly: a generator of conjectures, trustworthy only in the company of a critic. Every verifier we build is a small tribute to his idea that we learn by finding out where we are wrong.

    We use cookies for analytics and personalization. Privacy Policy