Epoch AI tracked the price curves of five AI benchmarks over the past three years: the per-query cost for equivalent capability drops roughly 47% each quarter, an annual rate approaching 13x. In January 2025, getting o3 to answer a PhD-level multiple-choice question correctly cost 30 cents; by GPT-5.6 Luna, it was down to $0.04 — a 725x drop in 18 months. But the quarterly decline for newly achieved SOTA capabilities has been cut in half over two years, from 66% to 32%. The steepest price drops are behind us.
30 cents to get a PhD physics question right; 18 months later, just 4 tenths of a cent
In late January 2025, OpenAI released a model called o3, which scored 75% accuracy on PhD-level physics, chemistry, and biology multiple-choice questions, at a cost of 30 cents per question. Less than 18 months later, GPT-5.6 Luna hit the same score for just $0.04 per question. 725x. Epoch AI calls this "the price of thought."
GPQA Diamond is 448 multiple-choice questions covering PhD-level physics, chemistry, and biology — questions a layperson can barely even parse. When o3 was released on January 31, 2025, it burned through an average of 30 cents per question.
Epoch AI pulled OpenAI's published API pricing and the model's benchmark scores onto the same coordinate plane to plot this "cost per correct thought" curve. In under 18 months, the cost of crossing that same 75% threshold crashed to $0.04 per question — four-tenths of a cent.
Translated into everyday terms, it's even more jarring: a new car sticker-priced at $50,000, 18 months later tagged at $69, title transfer included. This is the real price of AI inference(having the model "think" step by step before answering, rather than spitting out an answer in one shot). Epoch AI formally named this curve in a report published in September 2026 — the price of thought — to distinguish it from input-side costs like "compute."
The 725x figure is just the distance between the starting and ending points. Using three years of data, Epoch AI calculated that the cost of equivalent AI capability drops an average of 47% per quarter, or 13x per year — faster than the steam engine, electricity, and the internet. How that average was derived gets unpacked below.
But 725x is the endpoint, and 47% is the average speed. Epoch AI's data shows the price drops are decelerating: when a capability first becomes SOTA (state-of-the-art, the current strongest level), costs fall 66% per quarter; two years later, only 32%. The price-reduction dividend on new capabilities is thinning out.
Five benchmarks, three years of data — how the 47% quarterly drop was assembled
The researchers weren't tracking absolute prices. They plotted score against price over time: how much it costs to reach a given score on a given question, and how that cost moves. The curve has fallen 47% every quarter for three years.
Epoch AI researchers Luke Emberson and David Roodman compiled every piece of data they could get over the past three years: five AI capability benchmarks spanning math, hard sciences (chemistry, physics, biology), and game-like tasks. For each, they recorded the model's score and what it cost to hit that score at the time.
The x-axis is time; the y-axis is the lowest model price needed to clear a given score threshold. Cutting the data this way isolates the rate of decline in marginal cost for equivalent capability (calculated quarterly), shielded from new model price hikes or shifts in benchmark difficulty.
The spread across categories is the sharpest finding. Math problems fall fastest, at 50% to 52% per quarter — 16 to 19x per year. Game-like puzzles are slowest, at 39% to 43% per quarter, or 7 to 10x per year. The researchers' explanation: math has strong right/wrong signals, so every inch of training improvement cashes out as an inch of benchmark gain; game-like problems lean more on strategy search and compute stacking, leaving less room for algorithmic optimization.
Across all five lines together, costs slide down 47% per quarter — 13x per year. No general-purpose technology has kept that pace: not the steam engine, not electricity, not the internet.
The data doesn't go back before GPT-3 opened up commercial inference, but the researchers believe the 47% quarterly pace most likely kicked off as early as November 2021. Continuous model pricing data just wasn't available in the early days, so that stretch of the curve couldn't be precisely reconstructed in the paper.
The cost data came from publicly listed model API (cloud-based interfaces billed per call) pricing tables and official technical reports — vendor-reported figures, not independently verified by third parties. Treat these numbers as a report card the vendors graded themselves.