The AI demand bubble will be decided by who pays

Existing users can consume dramatically more AI by running agents, delegating longer jobs, and replacing labor with compute. The buildout survives only when outside customer cash replaces capital cycling among labs, hyperscalers, and investors.

Written By
Grant Harvey
Grant Harvey
Aug 5, 2026
27 minute read

So Ed Zitron is one of AI's most famous haters, and generally speaking, I like his work. His analysis (and occasional first party reporting) is about as good as it can be from outside the labs, and given his disdain for big tech. His new piece, The AI Demand Bubble, makes the strongest version of the AI bear case, which we haven't talked about in awhile: essentially, analyst estimates suggest OpenAI and Anthropic generate most hyperscaler AI revenue these days, and that's despite those labs depending on investors and the very same cloud companies to help fund their compute bills. It's a Cloud Compute Ouroboros all the way down, I guess?

Naturally, since reading this, I've been thinking about the assumption underneath Ed's argument. He seems pretty convinced about AI supply outrunning AI demand. And yet, the AGI-pilled crowd in Silicon Valley, convinced human-level AI (called AGI) is close, see demand swallowing every chip, power plant, and data center we can build.

So why does Ed keep claiming there will be a demand-based bubble? Shouldn't it be a supply-constraint bubble?

For example, Dwarkesh Patel and Gavin Baker both reached the same bullish conclusion about demand outpacing supply, but from different directions. Dwarkesh starts with the future: if models become capable enough to replace expensive human work, every scarce GPU becomes dramatically more valuable. Gavin's recent comments start squarely with the present: rising rental prices, pricier contract renewals, growing token usage, and private AI companies already suggest compute demand is stronger than public markets can see.

Ed audits today’s invoices. Dwarkesh and Gavin price in tomorrow’s machine labor. As a neutral third party observer, it feels like, because of their particular information bubbles, both are only seeing half the picture.

In my opinion, the biggest argument against Ed's demand take is that even if AI is a product no one wants, which seems increasingly untrue, AI as a product can actually multiply demand without adding net new users, because with each subsequent increase in intelligence and reduction in price, each existing user will want to increase their own usage. BUT... that demand still has to (eventually) become profitable cash for said users, or their ability to absorb the costs of AI will dry up. So whether we're building toward a demand glut, or a supply backlog, either way we're in for a bumpy ride.

Advertisement

Let's break that down, shall we?

The two conversations that inspired this rebuttal

These are the two videos we're going to unpack, because they had a lot of data and numbers and projections in them that are worth understanding.

Dwarkesh Patel: Why smarter AI models could drive up compute prices 10x

  • (0:00) Patel frames the essay around what the compute situation for frontier labs could look like over the next few years.
  • (0:05) Anthropic revenue has reportedly grown 10x year over year for three straight years; Patel says it ended last year at $9B and could end this year at $100B to $150B.
  • (0:17) If that 10x trend continued, Anthropic would need roughly $1T in revenue by the end of next year, a conclusion Patel calls "very wild" and dependent on capabilities.
  • (0:39) Lab compute is increasing only about 3x per year, creating a gap if revenue keeps increasing 10x.
  • (0:46) Patel says the gap can close through some combination of three mechanisms: higher lab margins, higher compute prices, or a larger share of compute going to inference instead of training.
  • (1:01) His understanding is that all three mechanisms are already happening.
  • (1:05) Anthropic inference margins reportedly rose from about 40% in the middle of last year to upwards of 80% for Fable.
  • (1:14) Compute spot prices were more than 40% above their February 2026 trough.
  • (1:24) Epoch estimated OpenAI spent only about one-quarter of its compute on inference in 2024; Patel says the share is likely closer to 50% or higher now.
  • (1:34) Labs would prefer not to keep shifting compute toward inference because they view inference revenue mainly as proof that helps them raise money for the next, larger training run.
  • (1:49) If most compute goes to inference, Patel says a lab is effectively declaring that progress has stalled and it has become a cloud provider rather than an AGI company.
  • (2:03) The labs believe that within a year they will build models that make current systems look "extremely shitty," so they need the majority of compute for training and experiments.
  • (2:22) If labs preserve most compute for training, the remaining escape valves are higher lab margins or higher compute prices.
  • (2:39) Margins above 90% would require a leading model to stay so far ahead that competition could not substitute it away, which Patel finds difficult to imagine for intelligence.
  • (3:09) That leaves rising compute prices as the most plausible balancing mechanism in his thought experiment.
  • (3:22) Frontier labs cannot rely on ordinary spot instances; they need large, efficient, flexible, and secure clusters that protect model weights and customer data.
  • (3:44) Patel cites Google renting 110,000 mixed GB200 and GB300 GPUs from SpaceX for $900M per month, about 2x the spot price per GPU-hour, while spot itself was already more than 40% above February.
  • (4:06) The central conclusion is that smarter models can monetize the same amount of compute much more effectively.
  • (4:11) A human-level software engineer running on an H100 equivalent could justify more than $250K per year at current engineer prices, over 15x the current H100 spot price, before counting nights and weekends.
  • (4:34) Patel considers the objection that millions of new AI engineers would reduce the marginal value of engineering labor.
  • (4:50) He compares that objection to the lump-of-labor fallacy.
  • (4:55) The analogy is high-skill immigration: standard economics often expects innovation and specialization to preserve or increase long-run labor value rather than depress it.
  • (5:06) He allows that an AI labor shock could be so large and fast that the usual economic heuristic fails.
  • (5:20) As the best labs monetize compute better and compute gets more expensive, competitors will struggle to bid against firms that can earn more from each unit.
  • (5:38) Patel says the most interesting implication is that the best, most compute-efficient model could charge a much larger premium.
  • (5:46) He invokes the Alchian-Allen effect: when the shared input becomes expensive, buyers rationally pay more for the higher-quality option that economizes that input.
  • (5:50) At $20 per H100-hour, using a weaker model that burns more tokens for the same result becomes economically irrational.
  • (6:11) A model that achieves the same result with less compute has, in effect, "created more compute," which makes scarce compute more valuable.
  • (6:21) Many current low-value AI applications could be priced out once AI can perform work comparable to top humans.
  • (6:35) Labs may be willing to pay far more for tokens that automate AI research than consumers will pay for "more AI slop talk."
  • (6:46) Patel worries the argument resembles past scarcity forecasts that underestimated substitution and innovation, especially the Simon-Ehrlich bet.
  • (7:24) He still thinks the analogy may fail because compute supply is less elastic, less substitutable, and less able to absorb demand shocks than commodity extraction.
  • (7:51) The roughly 3x annual compute increase decomposes into about 1.4x from Moore's Law, 1.2x from new fabs, and 1.8x from reallocating leading-edge wafers from smartphones and PCs to AI.
  • (8:01) Rather than accelerating Moore's Law, Patel says maintaining the current pace for a few more years would itself be a "miracle."
  • (8:12) New fab growth may remain bottlenecked through 2030 or beyond by the rate at which ASML can build EUV lithography machines.
  • (8:23) The wafer-reallocation boost could hit a wall by the end of next year, when AI's share of TSMC N3 capacity may rise from 60% to 86%.
  • (8:46) Patel questions whether the industry can even sustain 3x annual compute growth for several more years, much less exceed it.
  • (10:05) He expects compute to become cheap again eventually, in a post-scarcity world where robots can turn raw materials into chips; the essay is about the current pre-singularity regime.
  • (10:34) Anthropic growing revenue 10x while compute grows 3x illustrates unusually strong economies of scale in the model business.
  • (10:43) A model pays a one-time training cost to learn skills that can be shared across all users, unlike human labor, where every new worker must be trained separately.
  • (10:58) Patel dislikes the resulting concentration of power, but concludes that intelligence appears to have these strong scale economies.

Gavin Baker: Why the markets are pricing AI wrong

  • (0:00) Baker says he came to Silicon Valley trying to be scared and to find a negative quantitative metric, but instead found improving fundamentals and falling stock prices.
  • (0:25) As recorded, Nvidia was at its lowest forward P/E in a decade, which Baker reads as the market pricing in significant overearning.
  • (1:18) He describes July 2026 as "2022 in a month."
  • (1:30) Many AI stocks fell roughly 40% to 60% from their highs in a straight-line monthly selloff.
  • (1:55) Baker says he found no clear quantitative deceleration; GPU, memory, and token metrics were accelerating.
  • (2:09) The accelerating indicators he names include GPU availability, GPU rental pricing, DRAM spot prices, and token growth.
  • (2:28) A major market blind spot is poor visibility into Anthropic, OpenAI, and U.S. open-source inference clouds such as Fireworks, Baseten, Modal, and Together.
  • (2:42) Open-source usage accelerated sharply after GLM 5.2 and Kimi K3, while Nvidia's Nemotron continued advancing.
  • (3:07) Baker says Anthropic was still growing strongly and was almost certainly producing significant free cash flow.
  • (3:14) Public charts comparing semiconductor cash flow with hyperscaler cash flow miss the private labs and inference providers.
  • (3:32) In 2024 and 2025, even bulls expected GPU rental prices to decline slowly; almost nobody expected old GPU prices to rise vertically in 2026.
  • (4:03) Long-term GPU contracts that helped neoclouds finance hardware now price installed compute at a large discount to the spot market.
  • (4:25) As those contracts expire, installed compute can reprice higher even if spot prices fall from current levels.
  • (4:48) Baker says operating cash flow from Microsoft, Meta, and Amazon accelerated from 28 to 32, or to 35 after adjusting for unusual one-time items.
  • (5:19) That cash-flow acceleration came before Rubin systems and before many below-market compute contracts repriced.
  • (5:41) The market interpreted Meta renting compute as evidence of excess capacity and future capex cuts; Baker argues the actual capex telemetry never weakened.
  • (5:58) His alternative explanation is that Meta saw SpaceX sell specialized clusters at huge premiums and recognized an opportunity.
  • (6:18) He speculates Meta may have considered proving high returns on a small capacity slice, then raising equity and increasing capex.
  • (7:01) Meta's Muse 1.1 was, in Baker's view, its best model in a long time, which further argued against slowing investment.
  • (7:20) The Silicon Data token index flattened partly because model mix shifted after GLM 5.2 and Kimi K3, not necessarily because total compute demand weakened.
  • (7:44) The mix moved from expensive frontier tokens with very high inference margins toward cheaper open-source tokens.
  • (8:00) Baker's core claim is that, broadly, a token still requires flops, memory, and watts regardless of whether the model is frontier or open source.
  • (8:24) Open source therefore shifts margin dollars out of the frontier-model layer and toward infrastructure rather than eliminating compute demand.
  • (8:33) Lower token prices create elasticity, which can increase token consumption and total infrastructure demand.
  • (9:04) Baker treats Jensen Huang's strong support for open source as evidence that open source is not bad for Nvidia's business.
  • (9:36) He also argues society benefits from many capable models rather than one or two frontier providers charging roughly 90% margins.
  • (9:55) Reports of a Chinese DUV lithography machine triggered another semicap equipment selloff.
  • (10:04) The serious market concern, in Baker's view, is tighter credit: real yields rose, spreads widened, and CDS prices increased.
  • (10:23) He says the overwhelming majority of the AI buildout is still funded from operating cash flow, although credit is becoming a larger component.
  • (11:09) Rising credit costs become truly dangerous only if the buildout must depend heavily on debt.
  • (12:33) The key near-term question is how much credit the next six months of construction will require.
  • (12:45) Debt-financed buildouts need prompt repayment and can unwind rapidly when supply and demand diverge, as happened during the internet boom.
  • (13:28) Consensus models the coming Blackwell and Rubin gigawatts as monetizing at roughly Ampere-era rates, two generations behind Hopper.
  • (13:56) On that conservative assumption, Baker cites about $1.3T to $1.4T of hyperscaler operating cash flow.
  • (14:28) If the capacity monetizes below current Blackwell rates but above Ampere rates, his estimate rises toward $2T, removing roughly $700B of credit demand.
  • (14:45) Higher monetization and contract repricing would improve credit ratios, making any remaining debt easier to finance.
  • (15:14) He revisits the "Blackwell air pocket" risk: hundreds of billions spent on chips used first for training, which does not immediately produce a return.
  • (15:45) Strong Anthropic results helped the market look through that risk in April through June; July's narrative pileup changed the mood just as operating cash flow improved.
  • (16:16) Microsoft brought a large amount of new capacity online in June that had not yet appeared in second-quarter results.
  • (16:24) The bull thesis depends on private-company demand staying strong enough that installed compute reprices upward as contracts expire.
  • (16:46) At current repricing levels, Baker believes operating cash flow could fund most, and perhaps all, of the buildout for several years.
  • (17:16) The continuing selloff itself is unsettling because, aside from credit, the market's concrete fundamental objection is unclear.
  • (18:25) One startup rented several thousand B200s in the mid-$2 per GPU-hour range, then seven months later expected to pay just under $4 for an equivalent cluster.
  • (19:06) That example implies a 50% to 60% increase in GPU rental prices over six or seven months.
  • (19:24) An inference cloud reportedly expected to pay 100% more for Blackwells when its contract expired.
  • (19:31) Those repricing examples imply that hyperscalers are currently under-earning on contracted installed compute.
  • (19:59) The only potentially negative metric Baker found was third-party data suggesting Anthropic's growth curve may have softened slightly.
  • (20:07) Even if Anthropic slowed, he says OpenAI and open source accelerated enough that aggregate demand still appeared to accelerate.
  • (20:24) Baker calls open-source demand "dark matter" for public markets because it is difficult to observe directly.
  • (21:21) Nvidia's lowest forward multiple in ten years tells Baker the market "100%" believes the company is overearning.
  • (22:47) He still sees Anthropic in pole position, but says OpenAI, Grok, and Cursor also had a transformational July.
  • (23:21) A Fidelity investor summarized the last three years as doing "the dumbest most superficial thing as quickly as possible" and rapidly cycling risk; Baker thinks July followed that pattern.
  • (24:02) Even if credit becomes unavailable, Baker argues that scarcity would make the compute already online more valuable.
  • (24:28) He proposes that Claude has become a kind of "Walter Cronkite for the stock market," because many investors feed the same news into the same model and receive similar interpretations.
  • (25:18) Claude may be smart, but its interpretation is not always correct, especially when investors are making probabilistic judgments about the future.
  • (25:53) A chart of Japanese capacitor stocks suggested an entire multi-year capacitor cycle had been compressed into about six weeks.
  • (26:55) The technical development Baker finds most capable of disrupting demand is progress toward continual learning and sample-efficient learning.
  • (27:25) He contrasts models trained on roughly 300T tokens with a human estimate around 20B, then imagines training on 10T tokens and learning efficiently in deployment.
  • (27:46) If that happens, training could fall toward a small share of semiconductor compute demand, though not literally zero.
  • (28:10) Even so, Baker finds it hard to believe better learning efficiency would be net negative for total AI infrastructure demand.
  • (28:44) The clearest thesis breaker would be hyperscaler operating cash flow failing to accelerate.
  • (29:16) A sustained, dramatic fall in GPU prices, or GPUs becoming easy to obtain, would also be a major warning signal.
  • (29:48) A plateau or decline across Anthropic, OpenAI, Grok, Cursor, and other labs would be negative unless open-source growth was expanding the overall pie.
  • (30:05) Baker expects a multimodel future in which companies fine-tune an open-source model on proprietary data and place it behind a router with frontier models.
  • (30:30) That stack can sometimes produce slightly better outcomes at roughly half the user-facing cost.
  • (30:45) He argues the lower price comes mainly from shifting tokens from models with roughly 90% gross margins to models with perhaps 30% margins, not from using less compute per token.
  • (31:24) Companies that let AI spending rise 20x can use routers to stabilize budgets while still increasing total token volume.
  • (32:17) Cheaper tokens can therefore raise the number of GPU-hours consumed even when the customer's dollar spend stabilizes or falls.
  • (32:24) AI-native companies are leaning into tokens instead of hiring humans, while other regions and industries remain at much earlier adoption stages.
  • (33:11) Baker estimates only roughly 250,000 to 500,000 people use agentic AI today despite an acute compute shortage, then asks what happens at 1%, 100M, or 500M users.
  • (33:56) The revenue supporting the buildout must ultimately come from faster productivity-driven economic growth or labor substitution.
  • (34:23) At AI-native companies, labor substitution often appears as hiring fewer people rather than firing existing workers.
  • (34:56) Token spending at aggressive adopters can reach 20% to 25% of compensation spend; examples mentioned in the conversation reached 30% and even 50%.
  • (35:22) Applying a 20% token share to roughly $25T of knowledge work implies a potential $5T pool.
  • (35:47) A technology CEO told Baker that founder-controlled companies were not conducting broad layoffs after adjusting for COVID-era overhiring, suggesting they still see abundant opportunities for people plus AI.
  • (36:16) In the bull case, AI spending drives growth rather than pure labor replacement; Baker points to Cognition, Ramp, and Stripe data linking higher AI spend with faster growth.
  • (36:55) Current shortages span chips, memory, power, and infrastructure; where demand looks weak, the constraint may simply be an inability to energize gigawatts quickly enough.
  • (37:05) Supply responses have become extreme, including reconditioning turbines from old aircraft for data-center power.
  • (37:29) Memory suppliers and customers are shifting toward long-term agreements with prepayments, price floors, and ceilings, sacrificing short-term upside for durability.
  • (38:48) Baker says adding more memory per unit of compute is the single most important way to increase token output.
  • (39:19) He frames the competition among Amazon Trainium, Google TPUs, AMD, and Nvidia as "Game of Thrones" because supply-chain allocations can determine market share.
  • (39:39) For the next several years, share may depend heavily on what capacity each company has pre-purchased.
  • (40:02) Breaking a memory LTA to chase lower prices could be fatal if leverage later returns to suppliers and future allocations move to a competitor.
  • (40:57) That game theory can hold even during temporary oversupply because memory remains cyclical and oversupply eventually reverses.
  • (42:08) Baker thinks the current environment structurally favors Nvidia and makes its low valuation difficult to understand.
  • (42:16) When projects require financing, Nvidia GPUs are the easiest accelerators to finance.
  • (42:33) He describes Nvidia's newer model as a "credit wrapper" combined with a revenue share when GPU prices remain above a floor.
  • (42:45) That arrangement could quickly create a very large cloud-like royalty business for Nvidia.
  • (43:03) Baker distinguishes it from vendor financing because third parties lend to GPU buyers while Nvidia may invest equity and share in upside.
  • (45:05) The structure extends the logic of LTAs: Nvidia trades some immediate upside for durable recurring revenue.
  • (45:24) Baker says Nvidia should explain the structure more clearly and notes that avoiding equity stakes in major AI winners has usually been a mistake.
  • (45:49) Jensen Huang sees the plans and progress of many labs, including continual-learning efforts, and Baker interprets Huang's behavior as informed bullishness.
  • (46:04) Equity upside plus revenue sharing can bridge the period before customers generate enough operating cash flow to self-fund expansion.
  • (46:37) The model also strengthens Nvidia against startups that pay more for TSMC capacity and HBM and cannot finance chips on equally favorable terms.
  • (47:30) Baker argues Anthropic could have run away with the market if it had bought compute as aggressively as OpenAI; OpenAI and Grok returned to the frontier because they had the compute.
  • (47:54) After seeing what under-buying compute can cost, he doubts any frontier lab will voluntarily back off soon.
  • (50:29) Baker says nearly everyone he met in Silicon Valley was more bullish than he was.
  • (50:39) Dwarkesh Patel's scenario of an H100 renting for about $250K per year, roughly 15x spot, sat outside even Baker's previously considered tail outcomes.
  • (51:20) He summarizes the acceleration as multiplying more compute, higher compute economics, and higher inference margins.
  • (52:09) China's reported DUV progress can be both decades behind and a real phase transition because China previously lacked the capability entirely.
  • (53:12) The market likely overreacted to the immediate ASML impact, which might take five years to appear, but Baker says the development should not be dismissed.
  • (54:22) Semiconductor progress depends on learning by doing, so China cannot simply teleport from a 2001 capability level to the 2026 frontier.
  • (55:31) Open-source models approaching the frontier, plus easy customization from inference clouds, are a "godsend" for software and AI-native companies.
  • (56:08) He cites companies reaching roughly $50M in annual revenue and positive cash flow within about nine months, though durability was previously uncertain.
  • (56:34) Proprietary domain data combined with customized open-source models can turn a former "wrapper" into a more defensible business.
  • (56:48) Fireworks Nexus can ingest a company's data, reinforcement-learn a model, and route queries with only a few lines of code.
  • (57:32) An AI-native company might send 30% to 60% of tokens to frontier models and the rest to its own tuned model, improving defensibility and cost.
  • (57:49) The conversation cites a system where a frontier model plans and delegates tasks to weaker models at roughly 15x greater efficiency.
  • (58:20) A frontier-model-maximalist view says a model that reaches recursive self-improvement could distill every lower intelligence level cheaply and leave little room for open source; Baker considers that possible but unlikely.
  • (59:06) AI-native companies with proprietary data may prefer open source for independence, durability, safety, and control over terms.
  • (59:43) Cheaper open-source models may increase the economic value of the most capable frontier model that orchestrates them.
  • (1:00:22) Inference clouds are growing nearly as fast as early frontier labs while burning far less cash.
  • (1:01:05) Frontier models may continue capturing most economic value even if open source produces most tokens.
  • (1:01:40) Baker identifies regulation as the largest risk to the AI infrastructure thesis.
  • (1:02:26) He points to a New York data-center moratorium and a public narrative that data centers raise electricity prices, consume water, and eliminate jobs.
  • (1:03:12) Baker argues that newer behind-the-meter deals can lower surrounding electricity prices rather than raise them.
  • (1:03:29) He says developers increasingly offer communities hospitals, schools, police and fire infrastructure, and lower power bills as part of data-center deals.
  • (1:03:52) Data centers create continuing demand for plumbers, electricians, HVAC contractors, and maintenance workers, not only temporary construction jobs.
  • (1:04:35) He cites a data-center water-use claim that was allegedly overstated by 10,000x and persisted even after correction, comparing it with the spinach-and-iron myth.
  • (1:05:28) Baker wants the industry to mount a broad public-information campaign explaining what data centers do and how communities benefit.
  • (1:06:31) He argues that well-structured projects can bring cheaper power, wealthier communities, more jobs, and minimal water or environmental impact.
  • (1:06:48) The industry should also tell stories about AI aiding medical breakthroughs, rare-disease treatment, and life-saving research.
  • (1:07:38) People inside Silicon Valley often assume AI's benefits are obvious, but Baker says that view diverges sharply from much of the public.
  • (1:08:33) A missing piece in many compute forecasts is SRAM-based acceleration, which can use older nodes and avoid some HBM constraints.
  • (1:08:53) He describes disaggregating inference into prefill, attention, and feed-forward stages that can run on different specialized chips.
  • (1:09:25) Adding SRAM accelerators to existing and new clusters could materially improve AI return on investment.
  • (1:10:30) Potential dark-horse players include Fireworks and Cognition, whose leaders Baker sees as unusually capable.
  • (1:11:11) Baker calls SpaceX the most important newly public company and thinks markets do not yet understand it as a compute business.
  • (1:11:37) He says SpaceX fundamentals improved after its IPO through Grok 4.5, the Cursor acquisition, and demonstrated ability to deploy compute quickly and cheaply.
  • (1:12:05) SpaceX added a vast amount of compute and the market absorbed it without a visible pricing blip, which Baker sees as bullish for demand.
  • (1:12:29) A public report suggested SpaceX could try to bring 8GW of compute online, a target Baker finds extraordinary and probably too high.
  • (1:12:52) He cites monetization near $50B per gigawatt and next-year consensus revenue around $73B, before fully counting Starlink upgrades, Grok 4.5, Cursor, and the core business.
  • (1:13:39) Only hyperscalers, CoreWeave, Crusoe, and SpaceX had reportedly energized more than 500MW in a year; SpaceX did so fastest and at the lowest cost.
  • (1:14:23) The bearish SpaceX case assumes compute spot prices fall roughly 90%, leaving new capacity far less valuable than expected.
  • (1:15:18) Baker thinks very little of SpaceX's possible compute expansion is reflected in market estimates, even though 8GW is unlikely.
  • (1:15:51) After spending time at Starbase, he says orbital compute feels more real every day.
  • (1:16:05) Benchmark backing StarCloud, an orbital-compute company working with SpaceX and Starlink laser technology, serves as an outside sanity check for the concept.
  • (1:17:26) Baker closes with humility: the future is probabilistic, and the market will reveal over the next year which thesis was right.

Okay, so let's now analyze these datapoints all together.

User count =/= demand, as it misses the main AI demand engine

Software usually grows by adding customers, subscriptions, or seats. AI adds a second growth engine: each customer can buy more work.

The same company can keep its workforce flat while increasing AI usage through:

  • Longer reasoning jobs.
  • More agents running at the same time.
  • Automated workflows that operate all day.
  • The most capable models supervising cheaper models.
  • Repeated testing, checking, and revision.
  • Compute replacing tasks that previously required another employee.

A company with 10,000 employees could still have 10,000 employees next year. Its AI workload might rise tenfold because each worker supervises several agents.

Microsoft's April FY26 Q3 call showed the first signs of this shift. Microsoft 365 Copilot had more than 20M paid seats, first-party agent usage had risen 6x year to date, and queries per user were up nearly 20% in one quarter. By July 29, Microsoft said paid seats had passed 30M.

Advertisement

Google told a similar story in February. Gemini Enterprise had more than 8M paid seats, nearly 350 customers each processed more than 100B tokens in December, and its enterprise agents handled more than 5B customer interactions during the quarter.

Google also said it reduced Gemini's serving cost per unit by 78% during 2025. That number belongs in both the bull and bear cases. Lower costs can unlock much more usage, while competition passes much of the savings to customers.

A simple stress test

The following numbers are The Neuron's illustrative scenario, separate from any company guidance, analyst consensus, and the figures supplied by the videos.

Assume the number of serious AI users doubles between 2026 and 2030.

Then assume each user consumes 25 times more AI as occasional chat gives way to persistent agents.

That produces:

  • 2× as many serious users
  • 25× more work per user
  • 50× more total AI workload

Now assume better chips, smaller models, routing, and software improvements cut the average price of that workload by 85%.

Revenue still grows 7.5 times:

50× workload × 15% of today's price = 7.5× revenue

The simple scenario shows how AI revenue can rise through heavier use even while unit prices collapse. A 25× increase is an assumption, not a prediction. If Anthropic's engineers' internal AI use is any indication, it could be much more.

Gartner estimates agent-based jobs may already consume five to 30 times more tokens than ordinary chatbot requests. The unresolved race is between that workload growth and the combined effect of new power, better chips, cheaper models, and smarter routing.

Baker's examples point toward scarcity at the high end. One startup expected an equivalent B200 cluster to cost 50% to 60% more after seven months. An inference provider reportedly expected its Blackwell renewal price to double.

Those two private-market anecdotes do not form a market-wide price index. They suggest top-end compute can remain scarce while older chips or badly located data centers become oversupplied.

Economics says both camps are right on different clocks

Economics has names for nearly every part of this disagreement.

AI demand can grow through the intensive margin, meaning the same customer buys more rather than the company finding more customers.

A February 2026 NBER working paper models production as a sequence of steps that humans and AI can divide between them. The authors found that better AI can create nonlinear productivity gains once it becomes capable enough to handle several connected steps as a chain.

Advertisement

That is the economic version of the subagent argument. A small capability improvement can unlock eight hours of autonomous work rather than making one chatbot answer slightly better.

Cheaper AI can also increase total compute consumption. Economists call this the rebound effect, with the strongest version known as the Jevons paradox. The basic idea is simple: when a resource becomes more efficient and cheaper to use, people find many more reasons to use it.

A 90% decline in the price of an agent job could produce more than ten times the usage. But a review of the research found mixed evidence on whether efficiency gains cause consumption to rise enough to erase all the savings.

So Jevons is a possibility, not a cheat code proving infinite AI demand. The bull case works when usage grows faster than the cost per task falls.

More AI activity also does not guarantee proportionally more economic output.

A May 2026 NBER study followed more than 100,000 software developers. Autonomous coding agents increased commits, or saved code changes, by 180%. The gain fell to 50% at the project level and 30% for actual software releases.

The agents produced much more code. Human review, coordination, testing, and deployment became the bottlenecks. The researchers estimated that AI and human effort remained strong complements rather than clean substitutes.

That gives Ed one of his strongest points. Token consumption and GPU utilization can explode while finished, sellable output rises much more slowly.

Businesses may also experience value before their financial statements show it.

A March 2026 NBER survey of nearly 750 executives found positive AI productivity gains, especially in finance and high-skill services. Executives' perceived gains were larger than measured gains, which the researchers linked to delays before productivity turns into revenue.

That is this entire debate in miniature. The bulls may correctly see operational value that accounting has not captured yet. Ed may correctly point out that data centers are financed with cash, not perceived productivity.

The older economic literature makes the timing problem even worse.

The Productivity J-Curve says general-purpose technologies often require years of workflow redesign, training, software, and organizational change before their productivity gains appear in the data.

Advertisement

Classic time-to-build research warns that long construction periods and uncertain future revenue can produce gross overinvestment. Companies commit capital years before they know what the finished asset will earn.

Those theories fit together almost too well. AI demand could arrive later than Ed expects and still arrive too late for some of the companies financing today's data centers.

Supply constraints are already visible. A March 2026 Dallas Fed paper estimated that existing data centers had raised wholesale electricity prices by 3% to 5% nationwide. Depending on construction and utilization, the researchers projected increases of roughly 20% to 50% by 2028.

Local scarcity still does not guarantee attractive returns. The industry can have too little power, too many badly financed projects, and too much obsolete hardware at the same time.

Ed can be right about the financing bubble while underestimating eventual workload. Silicon Valley can be right about workload while badly mispricing the return.

Putting that together, a technology can transform the economy and ALSO destroy an extraordinary amount of capital along the way...

Ed Zitron follows the money

Zitron asks where the money paying for all that compute begins.

An AI lab can consume $20B of cloud infrastructure because:

  • Customers paid it $20B.
  • Investors supplied the cash.
  • A cloud company invested in the lab.
  • The lab signed a commitment payable later.
  • A lender financed the data center against expected demand.

Every path creates workload. Only the first proves that outside customers value the output enough to fund it.

Zitron cites analyst estimates placing OpenAI and Anthropic at 59% to 75% of AWS AI revenue. Barclays analyst Ross Sandler estimated 73% in 2026 and 2027, rising to 75% in 2028. A separate estimate from Stephen Ju put the concentration at 59% in 2026 and 55% in 2027.

He also cites a UBS model attributing 28% of total Google Cloud revenue in 2026, and 49% in 2027, to OpenAI and Anthropic. A Wells Fargo estimate puts the two labs at 70% or more of Microsoft AI revenue.

Amazon, Google, and Microsoft have not confirmed those figures. The estimates still expose the risk: the companies driving infrastructure demand may depend on money supplied by the companies building the infrastructure.

Advertisement

Anthropic illustrates the tension. In May, it reported annualized revenue above $47B and raised another $65B, partly to expand compute capacity.

That is extraordinary growth paired with extraordinary capital intensity.

"Demand" is doing too much work

The industry uses one word for four different things:

  • Technical workload: An agent runs, generates tokens, or occupies an AI chip.
  • Paid usage: A customer pays for ChatGPT, Claude, Copilot, Gemini, an API, or an agent platform.
  • External customer cash: The money comes from a hospital, retailer, software company, or other buyer outside the AI financing loop.
  • Profitable demand: That payment covers the model, the cloud, hardware replacement, debt, and a reasonable return on the capital invested.

A GPU can stay busy while the infrastructure owner loses money.

The internet boom proved that point. Traffic grew, fiber carried real demand, and many companies still collapsed because they borrowed too much, built too early, or watched prices fall faster than usage grew.

The official numbers give each side ammunition

In April, Microsoft reported an AI annual revenue run rate above $37B, up 123% year over year. Azure grew 40%, and management said demand continued to exceed capacity.

The same quarter showed the cost. Microsoft Cloud's gross margin, the share left after direct service costs, fell to 66%, partly because of AI infrastructure and growing usage. Intelligent Cloud cost of revenue rose 47%, driven by AI infrastructure and GitHub Copilot.

Google Cloud passed a $70B annual run rate, while signed future business reached $240B. Revenue from products built on Google's generative models grew nearly 400% year over year.

Those are real businesses serving real customers. They still leave the questions investors need answered:

  • How much revenue comes from frontier labs?
  • How much comes from ordinary businesses?
  • How much cash is collected rather than promised?
  • What does the infrastructure earn after replacement costs?
  • How concentrated is the customer base?

The hyperscalers report impressive growth without enough detail to separate those buckets.

Eventually, the finance team asks the rude question: who is paying for all this?

A range, not a prophecy

Our first model produced a base case near $1.2T in annual model and agent revenue by 2030.

Zitron's criticism showed where that forecast was too confident. It treated workload, customer revenue, cloud pass-through revenue, and capital-funded lab spending as equally durable.

The following ranges are our current scenario map:

  • Bear case, $250B to $400B: Usage grows, but prices fall faster than customer spending. Major compute commitments get renegotiated, and older or poorly located infrastructure loses value.
  • Base case, $550B to $850B: AI workload grows dramatically. Top-end clusters remain scarce, some commodity capacity becomes oversupplied, and outside customer revenue supports much of the buildout.
  • Bull case, $1.2T to $2T: Agents deliver clear labor-substitution economics. Businesses move meaningful payroll and operating budgets into machine work, allowing labs to pay for compute from operations.

These ranges map possible outcomes rather than precise probabilities.

The customer is the economy

The deeper issue is not how much AI the labs can sell. It is whether the companies buying AI can earn more than they spend, then keep spending from the returns.

Every AI dollar travels through a chain. A lab sells tokens to a software company. The software company sells an agent to a retailer, hospital, bank, or manufacturer. That business sells something to an end customer. Several companies can record revenue along the way, but the chain only closes when somebody outside the AI financing loop pays because AI made a product cheaper, better, faster, or newly possible.

If a company spends $1 on AI and creates more than $1 in additional profit, demand can compound. The company has both a reason and the cash to buy more. If it creates 70 cents, usage may still look impressive for a while, but the budget eventually runs into a wall.

This is where scarce compute could turn AI into a pay-to-win economy. The largest companies can reserve the best chips, pay for the strongest models, automate more work, and reinvest the gains. Smaller companies may pay higher spot prices, accept weaker systems, or wait for capacity. That creates a feedback loop in which the firms that can afford more intelligence become the firms best positioned to afford even more.

But bigness cannot eliminate final demand. If the largest technology companies mostly sell chips, clouds, models, and software to one another, one company's revenue remains another company's cost. The technology sector can capture a much larger share of the economy, but it cannot permanently float above the rest of it. Households, governments, and non-AI businesses still have to buy the products created with all that compute.

There are three broad ways this resolves:

  • A broad productivity boom: AI creates more output than it costs, lowers prices, expands purchasing power, and gives companies across the economy enough return to keep increasing their usage.
  • A pay-to-win economy: A few cash-rich companies capture most of the gains, smaller competitors get priced out, and AI demand stays enormous but increasingly concentrated.
  • A closed-loop boom: Labs, clouds, chipmakers, and investors keep funding one another while outside customer profits lag, making the system look healthier than its final demand really is.

A truly detached AI economy would require machines to produce, transact, and reinvest without relying on human purchasing power. That is a much stronger claim than saying AI demand is high, and nothing in current cloud revenue proves it yet.

Until then, chip scarcity determines who gets access. Customer returns determine how long they can keep paying.

The number that settles the argument

AI can generate more workload without generating more users. It cannot generate purchasing power by itself.

One human can summon 100 agents. Those agents still need to produce something another person or business can afford to buy.

Existing customers can probably absorb a huge amount of compute through longer jobs, parallel agents, automated checking, and lower prices.

The harder test is whether those customers create enough profit to pay for the supply.

OpenAI, Anthropic, and the application layer have to turn usage into customer-funded gross profit before their compute commitments, equipment write-downs, and financing costs arrive. The cloud companies eventually have to disclose enough detail for investors to calculate the return.

The number to watch is external free cash flow per energized gigawatt, or how much real customer cash AI businesses generate for each gigawatt of data-center power.

The final test is whether AI's productivity J-curve rises before the data-center industry's loomung (and hopefully not too funded with too much credit) bills come due.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.