If this is the cheapest intelligence has ever been, why does everyone using it feel like they're running out of time?

This report began as a discussion between two builders. Both agreed the moment is extraordinary. Both were anxious. But they were anxious about different clocks. One believes the window is closing: today's AI is subsidized, underpriced, and temporary — exploit it before the bill arrives. The other believes the floor is falling: every generation of capability gets radically cheaper, and the real edge was never the price. The data, it turns out, has something to say to both of them.

$20.00
Price per million tokens at GPT-3.5-class performance — Nov 2022 → Oct 20241
cheaper in under
two years
280×+
Collapse in cost of GPT-3.5-class intelligence1
200×/yr
Median price decline to fixed capability since Jan 20243
49%
Of U.S. employees never use AI at work12
1 in 9
U.S. adults report elevated AI-specific FOMO15

Builder Anxiety Is Now a Measurable Thing

The feeling has a name in the literature now. Researchers call it FOMO-AI — the worry that your AI skills, access, or output lag behind others — and it has a validated psychometric scale. In a 557-person U.S. study, elevated FOMO-AI predicted greater anxiety and depressive symptoms, with younger adults and women disproportionately affected. One finding stands out: AI literacy buffers the effect. The people who understand the tools best worry the most productively — and the least destructively.1516

More than 1 in 9 U.S. adults report elevated FOMO-AI

Not generalized tech anxiety — a specific, measurable fear of being left behind by AI. It correlates with anxiety and depressive symptoms, which in turn reduce well-being. This is the first peer-reviewed anchor for what builders have been describing anecdotally since 2023: staying up until 3 a.m. because the tools feel too good to put down, and too temporary to trust.15

Two Builders. Same Tools. Different Clocks.

Strip away the specifics and almost every AI-literate builder is running one of two mental countdowns. Both are rational. Both are backed by real evidence. They cannot both be the clock that matters most.

Clock No. 1

The Closing Window

The Arbitrageur's Thesis

AI capability is massively ahead of its price. For a few hundred dollars a month, one person can produce the output of an extremely capable team. That gap exists because labs are burning investor capital to win users — the Uber playbook: subsidize until habits form, then charge what it actually costs.

If that's true, this is a brief, historic arbitrage. The people who build aggressively inside the window will look like the people who registered domains in 1996. Everyone else will look back with regret.

"The tools are too cheap for how powerful they are. If I don't exploit this window, I may miss the moment."
Clock No. 2

The Falling Floor

The Compounder's Thesis

The moment is real — but the pricing premise is backwards. Every model generation gets cheaper per unit of capability, at rates with no precedent in any prior technology. Today's frontier model becomes tomorrow's cheap baseline. The premium tier may stay expensive; the floor keeps dropping.

If that's true, the window doesn't slam shut — access keeps broadening. The durable edge isn't getting in before the prices rise. It's taste, judgment, and knowing where premium intelligence is actually needed.

"The arbitrage is real, but the edge is knowing how to use the tools well — not accessing them during a temporary pricing moment."

Four Years of Contradictory Signals

Both clocks were wound by the same events. Read the timeline once through the Arbitrageur's eyes, then again through the Compounder's. It supports both stories — which is exactly why the anxiety won't resolve on its own.

November 2022

The $20 Baseline

ChatGPT launches. Querying a model at GPT-3.5-class performance costs about $20.00 per million tokens. Nobody calls this expensive — there is nothing to compare it to. The clock starts.1

2023

The Frontier Premium Era

GPT-4-class models arrive at premium prices, and the first builder migration begins. The pattern that will define the era emerges immediately: the newest capability commands a steep premium, while last year's miracle quietly gets discounted.

January 2024 Onward

The Floor Starts Falling Faster

Price declines accelerate. Across benchmarks, the cost to hit a fixed capability level falls between 9× and 900× per year depending on the task — and after January 2024 the median rate accelerates from roughly 50× to roughly 200× per year. No prior technology curve looks like this.34

2024–2025

Open Weights Commoditize the Middle

Open-weight models close the quality gap with closed models from 8% to 1.7% in a single year on human-preference leaderboards. By late 2025 they serve roughly a third of all tokens on OpenRouter, with Chinese open-weight models alone surging from a 1.2% weekly share to nearly 30% in some weeks. Capability is leaking out of the labs' pricing power.15

June–July 2025

The Repricing Wave

The other clock chimes. Replit replaces flat $0.25-per-checkpoint Agent pricing with effort-based pricing that scales with AI consumed — bills spike, users revolt.6 Weeks later, Anthropic introduces rate limits to curb Claude Code power users.9 The flat-rate, all-you-can-eat phase of AI products is visibly ending.

Late 2025

The Jevons Turn

Reasoning models — which burn far more tokens per task — grow from negligible to more than half of all tokens routed on OpenRouter. Unit prices keep falling; total consumption rises faster. The bill goes up anyway.5

March 2026

The Margins Come Into View

Reporting on Anthropic's internal figures shows a gross margin of −94% in 2024 swinging to a projected ~+40% for 2025 — undercut by inference costs running 23% over plan. The labs really were losing money serving you. And the loss really is shrinking. Both builders claim vindication.78

The Case for the Closing Window

The Arbitrageur's argument rests on a simple suspicion: nobody sells dollars for ninety cents forever. If the current price of intelligence is a customer-acquisition expense, then today's builders aren't customers — they're the habit being formed.

The Subsidy Thesis

Today's AI pricing reflects what labs are willing to lose, not what intelligence costs. Venture and investor capital absorbs the gap while usage habits form. Once dependence is established and capital tightens, prices converge upward toward true cost plus margin — and the builders who structured their economics around subsidized intelligence get repriced overnight. It is the Uber playbook applied to cognition: cheap rides until you've sold your car.

01

Subsidize

Serve intelligence below cost. Anthropic's gross margin in 2024: −94%. Every token you bought, someone else partly paid for.

02

Habituate

Builders restructure entire workflows, products, and businesses around the subsidized price. The dependency is the product.

03

Reprice

Flat rates become usage rates. Checkpoints become "effort." Unlimited becomes rate-limited. 2025 provided the first documented wave.

04

Extract

With habits locked in and alternatives behind, prices settle at what the market will bear — which, for the dependent, is a lot.

The Labs Really Were Losing Money on You

Anthropic gross margin, reported & projected78
2024
−94%
2025E
~+40%

Unaudited internal projections reported by The Information via Axios (March 2026). The 2025 projection was itself cut roughly 10 points because inference costs on Google and Amazon clouds ran ~23% higher than planned, with 2025 inference spend growing more than threefold to roughly $2.7B. Note what the Arbitrageur must concede: this chart shows serving economics improving dramatically — the strongest version of the window thesis is about product pricing, not raw model economics.

Case Study · June 2025
Replit

Before: Replit Agent charged a flat $0.25 per "checkpoint" — a predictable, arguably subsidized unit price for agentic work.

After: "Effort-based pricing." Complex tasks bundle into a single checkpoint that can cost substantially more than $0.25, scaling with the AI effort consumed. Users documented bill spikes; the backlash was loud enough that Replit acknowledged rollout problems and issued credits.6

Weeks later, Anthropic imposed rate limits on Claude Code's heaviest users.910 Cursor's 2025 pricing changes triggered the same cycle of confusion, anger, and refunds.11 One quarter, three repricings — each moving the same direction: toward usage-aligned billing.

What Would Falsify This Clock

Sustained gross-margin improvement at the labs without corresponding end-user price increases; flat-rate products that survive their power users; per-task costs (not just per-token prices) continuing to fall through 2026–2027. The margin chart above is already uncomfortable evidence: the subsidy is shrinking because serving got cheaper, not because prices went up.

The honest version of the window thesis survived verification in a narrower form than its believers tell it: product-level subsidies are demonstrably ending — flat rates, unlimited tiers, loss-leader checkpoints. What did not survive is the claim that intelligence itself is about to get more expensive.

The Case for the Falling Floor

The Compounder's argument doesn't need a story. It needs a chart. The price of any fixed level of AI capability has collapsed at rates with no precedent in the history of technology pricing — faster than Moore's Law, faster than solar, faster than genome sequencing.

The Floor Thesis

Whatever the frontier costs, last year's frontier becomes this year's commodity. Smaller models hit the same benchmarks, hardware gets 30% cheaper per unit of performance every year, and open-weight models put a hard ceiling on what anyone can charge for mid-tier capability. The premium tier may stay premium — but the capability most execution work actually requires keeps falling toward free. The window doesn't close. The floor drops out from under it.

The Same Intelligence, Two Years Apart

Price per million tokens at GPT-3.5-class performance (64.8% MMLU)12
Nov 2022 · GPT-3.5
$20.00
Oct 2024 · Gemini-1.5-Flash-8B
$0.07

A more than 280-fold price reduction for equivalent benchmark performance, driven by increasingly capable small models. Bar lengths are illustrative — at true scale, the second bar would be less than half a pixel wide. Source: Stanford HAI 2025 AI Index, data from Epoch AI. Figures are API list prices at MMLU-equivalence, not raw compute cost.

How Fast the Floor Falls, by Measure

Annual decline in price to reach fixed benchmark performance3
Slowest task tracked
9×/yr
GPT-4-level, GPQA Diamond
~40×/yr
Median, all tasks
~50×/yr
Median, after Jan 2024
~200×/yr
Fastest task tracked
900×/yr

Epoch AI, March 2025: log-linear regression of the lowest-priced model meeting fixed performance thresholds across six benchmarks, ~2022–early 2025. Bar lengths are proportional to the logarithm of the decline rate. Epoch cautions the fastest recent rates may not persist, and the analysis predates reasoning-model economics. Underneath it all: ML hardware costs fall ~30% per year performance-adjusted, and energy efficiency improves ~40% per year.1

Open Weights Close the Gap

Performance gap between best closed and best open-weight models, Chatbot Arena1
8.0%
Jan 2024
1.7%
Feb 2025

By late 2025, open-weight models served roughly one-third of all tokens on OpenRouter — with Chinese open-weight models alone growing from a 1.2% weekly share in late 2024 to nearly 30% in some weeks.5 Whatever closed labs decide to charge, this is the ceiling on mid-tier pricing power. Caveat: the gap figure is a dated snapshot on a human-preference leaderboard, and OpenRouter telemetry skews toward cost-sensitive routing traffic.

What Would Falsify This Clock

Price-per-fixed-capability flattening or reversing across 2026–2027; open-weight models failing to track the frontier as reasoning workloads dominate; or per-task costs rising persistently because capability gains demand disproportionately more tokens. Epoch's own caveat applies: the historic rates were measured before reasoning models changed what a "task" consumes.

The floor thesis won the verification battle decisively on unit economics. But notice what it quietly assumes: that the capability you need tomorrow is the capability that's getting cheap today. That assumption is where the third clock starts ticking.

The Floor Falls. The Bill Rises Anyway.

In 1865, William Stanley Jevons noticed that more efficient steam engines didn't reduce England's coal consumption — they exploded it. Cheaper intelligence is doing the same thing. As unit costs collapse, builders don't pocket the savings. They upgrade to more compute-hungry capability.

Reasoning Models Eat the Token Supply

Share of all tokens routed through OpenRouter going to reasoning-optimized models5
Negligible
Early 2025
>50%
Late 2025

Reasoning models consume far more tokens per task — they think out loud, at length, before answering. As they became the default for serious work, total token consumption surged even as per-token prices fell. OpenRouter data represents one API marketplace skewed toward developer workloads, not the whole market.

The Frontier Premium Refuses to Die

Blended list price per million tokens, late 20255
GPT-5 Pro (frontier)
$34.97
Gemini 2.0 Flash (budget)
$0.147

A roughly 240× spread between frontier and budget pricing persisted through 2025 — even as both tiers got cheaper in absolute terms. Bar lengths are illustrative, not to scale. This spread is precisely what underwrites the working pattern both builders already converged on: frontier intelligence for planning, architecture, and first-pass thinking; cheap models for execution and edits.

This is the synthesis hiding inside the argument: both builders are describing the same curve from different ends. The Arbitrageur watches total spend rise and calls it the subsidy ending. The Compounder watches unit prices fall and calls it the floor dropping. Jevons says: yes. Both. That's what the curve does.

Total spend can rise while unit cost collapses — when falling prices unlock demand faster than they save money.

You Are Arguing Inside a Bubble of Power Users

Both clocks assume the race is crowded. The adoption data says otherwise. The feeling that "everyone is already ahead" is an artifact of who builders follow, not of who actually uses these tools. Most of the labor market hasn't seriously started.

55%
U.S. adults 18–64 who have ever used generative AI (Aug 2025)14
45%
U.S. employees using AI at work even a few times a year (Q3 2025)12
23%
Using AI at work weekly or more12
10%
Using AI at work daily12

Where Adoption Is Actually Growing

Generative AI use among U.S. adults 18–64, Aug 2024 → Aug 202514
Personal use, Aug 2024
36.0%
Personal use, Aug 2025
48.7%
Work use, Aug 2024
33.3%
Work use, Aug 2025
37.4%

Federal Reserve Bank of St. Louis / Real-Time Population Survey (Bick, Blandin, Deming). Personal use is sprinting; work use is crawling — roughly 63% of U.S. workers still don't use generative AI at work at all. Gallup's Q4 2025 wave shows workplace adoption roughly plateauing at 46%, with 49% answering "never."13 Habitual use is rising (28% weekly, 13% daily by Q1 2026) but remains a minority behavior.

Sit with the asymmetry: builders are having a sophisticated argument about which decimal of the token price matters, while half the working-age population has never opened the tools at all. Whatever the right clock is, the stadium is mostly empty. The race has barely started — and the anxiety of the people in it is calibrated to the wrong crowd.

What Actually Expires, and When

Put both clocks on the table and the disagreement gets surprisingly small. They aren't predicting different futures. They're watching different layers of the same stack — and each layer has its own expiration behavior.

The Closing Window The Falling Floor
What it watches Product pricing — subscriptions, flat rates, agent checkpoints, usage tiers. Capability pricing — the cost of hitting a fixed benchmark via API.
What the data shows Real and ending: Replit's effort-based repricing, Claude Code rate limits, Cursor's billing revolt. Flat-rate AI was a loss leader, and 2025 was the year it died.6911 Real and accelerating: 280×+ collapse in two years, ~200×/yr median decline post-2024, open weights one model-generation behind.13
What actually expires The deal — unlimited tiers and below-cost products. Not access to intelligence itself. The premium — on any fixed capability level. The frontier stays expensive; everything behind it commoditizes.
The blind spot Confuses the end of product subsidies with the end of cheap intelligence. The margin data shows serving costs falling, not prices rising.7 Assumes per-token price is the bill. Jevons says the bill is tokens × price — and tokens are exploding faster than prices fall.5
Rational response Build aggressively now — but structure unit economics to survive usage-aligned pricing, because that's what's coming. Invest in taste, judgment, and model-routing skill — knowing where frontier intelligence is needed and where the cheap baseline suffices.

The verified record, compressed: the window thesis is right about products and wrong about intelligence. The floor thesis is right about intelligence and silent about bills. The anxiety is real, validated, and aimed at the wrong target — the scarce resource was never this year's pricing. It's the judgment to know what to build while everyone else is still deciding whether to show up.

Which clock are you building against?

The window — exploit the subsidy before it ends 24%
The floor — capability keeps getting cheaper, compound on it 31%
Both — products reprice while intelligence commoditizes 38%
Neither — the bottleneck was never price, it's taste 7%

Two builders argued about which clock was real. The data's answer is less dramatic and more useful: the subsidized deal expires, the price of intelligence keeps falling, and your total bill rises anyway — while half the market hasn't entered the building. The anxiety is measurable, the window is mislabeled, and the moat was never the discount. It was knowing what to do with the machine.

Citations & Sources

  1. Stanford Institute for Human-Centered AI (2025). The 2025 AI Index Report. Stanford University. Inference cost, hardware cost, and open-weight gap data from Epoch AI. hai.stanford.edu/ai-index/2025-ai-index-report
  2. Stanford HAI (2025). The 2025 AI Index Report: The State of AI in 10 Charts. hai.stanford.edu/news/ai-index-2025-state-of-ai-in-10-charts
  3. Epoch AI (2025). LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks. March 12, 2025. epoch.ai/data-insights/llm-inference-price-trends
  4. Appenzeller, G. (2024). Welcome to LLMflation: LLM Inference Cost Is Going Down Fast. Andreessen Horowitz. a16z.com/llmflation-llm-inference-cost
  5. OpenRouter & Andreessen Horowitz (2025). State of AI: An Empirical 100 Trillion Token Study. December 2025. openrouter.ai/state-of-ai
  6. Replit (2025). Introducing Effort-Based Pricing for Replit Agent. June 18, 2025 (updated July 2, 2025). blog.replit.com/effort-based-pricing
  7. Axios (2026). Reporting on Anthropic Gross Margin Projections and AI Model Serving Costs. March 12, 2026, citing The Information. axios.com/2026/03/12/ai-models-costs-ipo-pricing
  8. The Information (2026). Anthropic Lowers Gross Margin Projection as Revenue Skyrockets. Unaudited internal projections; figures relayed via secondary reporting.
  9. TechCrunch (2025). Anthropic Unveils New Rate Limits to Curb Claude Code Power Users. July 28, 2025. techcrunch.com
  10. VentureBeat (2025). Anthropic Throttles Claude Rate Limits; Devs Call Foul. venturebeat.com
  11. FinTech Weekly (2025). Cursor Pricing Change: User Backlash and Refunds. fintechweekly.com
  12. Gallup (2025). AI Use at Work Rises. Gallup Panel, Q3 2025, n=23,068, ±1.0pp. gallup.com/workplace/699689
  13. Gallup (2025–2026). Workplace AI Adoption Tracking, Q4 2025 and Q1 2026 Waves. gallup.com/workplace/701195
  14. Bick, A., Blandin, A. & Deming, D. (2025). The State of Generative AI Adoption in 2025. Federal Reserve Bank of St. Louis / Real-Time Population Survey; NBER Working Paper 32966. stlouisfed.org
  15. Telematics and Informatics Reports (2025). FOMO-AI Prevalence and Mental Health Correlates in a U.S. Adult Sample (N=557). Elsevier. Cross-sectional; associations are correlational. sciencedirect.com
  16. Telematics and Informatics, Vol. 100 (2025). Development and Validation of the FOMO-AI Scale. doi:10.1016/j.tele.2025.102283.

Methodology note: every quantitative claim in this dossier passed a three-vote adversarial verification process against primary sources. Figures that failed verification — including a widely circulated OpenAI burn projection and several token-share statistics — were excluded. Anthropic margin figures are unaudited internal projections relayed through secondary reporting and are flagged as such. Epoch AI's decline rates measure price-to-fixed-capability (API list prices), not frontier-model pricing or raw compute cost.