If AI companies have already built too many data centers, somebody forgot to tell the data centers.
Microsoft's latest quarter gave us the contradiction in one line: Azure grew 43%, yet management still said customer demand exceeded available capacity. Memory remains constrained. Nvidia says its supply chain is stretched. And investor Gavin Baker says the price of some identical B200 GPU clusters has risen roughly 50%-60% in six or seven months instead of falling.
That makes the simplest version of Ed Zitron's bubble argument hard to square with today's physical market. Oversupply is a strange description for a scarce asset that keeps getting more expensive.
But Zitron does not actually need today's GPUs to sit empty for his argument to work.
A GPU can run at 100% utilization and still be a terrible investment if the customer eventually refuses to pay enough to cover the GPU, power, data center, financing, model training, and everything wrapped around it.
That gives us a better question than “Is AI a bubble?”
Is the industry building too much compute, or are we still building too little at a price nobody yet knows how to set?
- First up, the TL;DR
- Jensen is building a pressure valve, not ending concentration
- Why $100B cannot buy you a gigawatt tomorrow
- Ed's best argument survives the shortage
- The whole debate comes down to what a token is worth
- Undersupply is how you build an oversupply crash
- NVIDIA may be the pressure valve and the accelerator
- The six signals I'd actually watch now
- What you can actually do with this
- Our take: shortage first, glut second
- Full video insights
First up, the TL;DR
So the AI industry is preparing to spend trillions on compute (chips and server racks to serve AI workloads). AI bears and bulls are aligned on this being a very big deal, but misaligned in terms of whether there is currently a shortage of compute, or soon to be a glut of compute. Hence, the AI bubble debate. Are we spending too much on AI infrastructure, or too little?
The physical evidence currently looks more bullish than bearish. Microsoft says AI demand still exceeds available capacity. Server memory remains tight. NVIDIA says demand exceeds supply. Investor Gavin Baker's newest checks show the same thing in GPU rental markets.
Meanwhile, SemiAnalysis founder Dylan Patel thinks OpenAI and Anthropic could consume roughly half of the world's new AI compute surprisingly soon because the labs can make more money from each megawatt than most competitors.
In his conversation with Dwarkesh Patel, Dylan Patel estimates roughly $11T of AI infrastructure investment from 2024 through 2029, with more than $5T potentially financed by debt. He thinks OpenAI and Anthropic alone could consume roughly half of new compute by the end of 2027. And the numbers get pretty astronomical.
Gavin Baker sees the spending very differently from skeptics. His argument is that demand is still accelerating, frontier labs, open models, clouds, apps, and chipmakers can all win, and transformative technologies often produce bubbles because investors fund infrastructure before demand fully catches up.
In particular, Baker sees a pressure valve: NVIDIA is helping outside investors finance the data centers, chips, and power needed to expand supply instead of forcing OpenAI and Anthropic to own every server rack themselves.
AI skeptic Ed Zitron sees something else. His counterargument goes straight at the word “demand.” He argues AI usage can look enormous while subscription subsidies, investor capital, and cloud financing obscure whether outside customers will ultimately pay the full cost of what they consume.
Here's what happened:
- Dylan estimates AI infrastructure spending could reach roughly $11T through 2029.
- Baker sees compute, memory, and token demand accelerating into shortages.
- NVIDIA is helping third parties finance more AI infrastructure.
- Zitron argues utilization can still hide terrible underlying economics.
Our take: Current evidence favors undersupply, but that shortage may be how the eventual bubble gets built.
Scarcity raises prices. Higher prices make massive construction look rational. Those projects arrive years later.
If AI creates enough value to absorb them, Dylan and Gavin win.
If customers refuse the full price once supply catches up, Ed gets his crash.
Shortage first. Glut second.
Jensen is building a pressure valve, not ending concentration
Dylan and Dwarkesh's centralization thesis begins with a simple economic rule: the company that can make the most money from a scarce megawatt can afford to pay the most for it.
Dylan estimates base AI compute has recently cost roughly $10M-$15M per megawatt, while Anthropic has at times generated as much as $50M in revenue per megawatt. That creates a flywheel where the best model earns more from scarce compute, reinvests the proceeds, improves, then bids even harder for the next unit.
Follow that far enough and OpenAI and Anthropic can consume an extraordinary share of new compute.
Except Jensen Huang appears to be building an enormous financial detour around the ownership version of that outcome.
Nvidia announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR designed to mobilize more than $500B of third-party infrastructure capital. Investors can finance AI factories instead of requiring frontier labs to fund every gigawatt themselves.
Baker's newest conversation describes three related mechanisms:
- Offtake agreements: A large customer commits to using future compute, which helps the data-center owner finance the GPUs.
- LTAs: Long-term agreements lock in future supply, especially for scarce memory, often with volume commitments, prepayments, and price floors or ceilings.
- Credit support: Nvidia helps make infrastructure financeable so third-party lenders can fund it, while Nvidia can participate in upside.
Baker calls Nvidia's newer structure a “credit wrapper.” Nvidia's own earnings commentary describes the same basic mechanism: outside capital underwrites the project, Nvidia can provide a minimum revenue commitment on part of the capacity, and Nvidia can share in revenue above a floor.
Nvidia is already doing a version of this in Ohio. SB Energy will own and operate the PORTS-Pike campus, while Nvidia provides credit support for land, power, and the data-center shell.
And the customer for the planned 8 GW?
OpenAI.
That gives us a hard distinction: Nvidia can decentralize ownership of compute without decentralizing who consumes it.
Jensen can stop two labs from owning every server rack without stopping them from buying most of what those racks produce.
In fact, making independent compute easier to finance could help OpenAI and Anthropic consume more of it.
Why $100B cannot buy you a gigawatt tomorrow
Dwarkesh looks at the economics and asks an obvious question: if a few billion dollars of semiconductor equipment can ultimately enable hundreds of billions in AI revenue, why doesn't capitalism simply build vastly more equipment?
Dylan's answer is basically: because factories take time.
The chain includes:
- ASML lithography machines and the precision optics inside them.
- Advanced semiconductor fabs.
- HBM and server DRAM.
- Networking gear and optical components.
- Transformers and turbines.
- Data-center buildings.
- Power generation and grid connections.
Money can make every company in that chain expand.
Money cannot make all of them finish expanding Tuesday.
Dylan calls the constraint industrial latency. Baker reaches the same conclusion from a different direction. In his latest conversation, he says “everything is at a shortage right now” and that the immediate weakness is the industry's inability to energize gigawatts quickly enough.
The independent data points the same way. Microsoft says demand exceeds capacity even after aggressive expansion. Memory suppliers are locking customers into long-term agreements. Nvidia says its entire supply chain remains challenged.
Conclusion #1: Today's AI infrastructure problem is primarily a time-to-build problem, not a lack-of-money problem.
The money is trying very hard to solve it.
The calendar refuses to cooperate.
Ed's best argument survives the shortage
Zitron's strongest argument gets weaker when it becomes “there is no AI demand.”
There clearly is.
Azure is growing while capacity constrained. Nvidia's data-center business is expanding rapidly. Google has reported broad paid enterprise adoption of Gemini and enormous token consumption from large customers.
I would therefore reject the strongest version of the “demand does not exist” claim.
There is substantial paying outside demand.
His financial critique is harder to dismiss.
Zitron's objection is the gap between utility and the amount of capital being committed to produce it. His cleanest market test is simple: what happens when customers pay something close to the full marginal cost of the AI they use?
A $20 or $200 consumer subscription does not tell you whether that user would pay the actual cost of every token they consumed. Flat subscriptions hide the meter.
Agent loops make the distinction sharper. A million tokens generated while solving a valuable engineering problem and a million tokens generated while an agent repeatedly screws up both count as demand for compute.
One created much more economic value.
Conclusion #2: Ed's strongest argument is about price, not utilization.
Busy GPUs prove scarcity.
They do not prove somebody paid the right price to build them.
The whole debate comes down to what a token is worth
This is where Baker and Zitron are talking past each other.
There are at least four different prices hiding inside one AI token:
- What the compute costs to produce it.
- What the model provider charges for it.
- What the customer is willing to pay for it.
- What economic value the customer creates with it.
Baker and Dylan are focused heavily on #4.
Dylan's Jane Street example is intentionally extreme. If Anthropic earns $100M from a megawatt while a customer creates $300M-$500M of value from that same intelligence, the token provider is leaving a huge amount of economic surplus with the buyer.
If AI saves a lawyer two hours, finds a trading opportunity, ships a feature, or replaces part of an employee's workload, the token can be worth multiples of its production cost.
That supports higher compute prices.
Zitron is much more skeptical of #3.
He sees subsidized subscriptions, bundled products, venture-backed application companies, and expensive agent loops and asks whether customers will still consume the same amount when the invoice reflects the whole cost.
OpenRouter's GPT-5.6 discount experiment gives us a useful test. When prices fell, Terra token consumption rose 5.6x and Luna jumped 13.8x. After discounts ended, roughly 32% of customers retained some usage.
The bullish interpretation: cheaper intelligence unlocks enormous latent demand.
The bearish interpretation: much of that demand appears only when intelligence is cheap enough.
Baker makes the infrastructure version of the same point. Cheaper open models can reduce model-layer margins while still increasing total tokens and GPU hours because customers use more of them.
Conclusion #3: We know people want vastly more intelligence as its price falls. We do not know the long-run clearing price of frontier intelligence.
That is the $11T question.
Undersupply is how you build an oversupply crash
Now the arguments finally snap together.
Imagine the next few years going exactly as today's physical evidence suggests.
- Compute stays scarce.
- GPU rental prices rise.
- Memory stays tight.
- OpenAI and Anthropic bid aggressively.
- Cloud companies panic about falling behind.
- Everybody signs LTAs.
- Customers sign offtake agreements.
- Nvidia brings hundreds of billions of institutional capital into AI infrastructure.
- Hyperscalers issue more debt.
- Suppliers order more factories.
- Power developers build more generation.
Congratulations, you have now given capitalism every possible signal to build a comical amount of compute.
But semiconductor fabs, data centers, transmission infrastructure, and power plants have long construction times.
The capacity ordered during the shortage arrives later.
And Baker himself accepts that is the dangerous part of that cycle. In the newest interview, he says debt-fueled buildouts can unwind very quickly if supply and demand move out of balance.
Zitron explains what makes that overbuilding dangerous: debt, fixed leases, depreciating hardware, and customer commitments remain after the demand forecast changes.
So the plausible cycle looks like this:
undersupply → rising compute prices → LTAs / offtake / credit support → debt-funded construction → delayed supply wave → price/value mismatch → falling compute prices → financing unwind
That is much more convincing than saying we are already drowning in unused GPUs.
Different pieces of the supply chain can also cross from scarcity to abundance at different times. Memory markets already show how one component can stay tight while another starts moving toward looser supply.
The eventual oversupply can be real without the current shortage being fake... or continuous, if demand remains elevated.
NVIDIA may be the pressure valve and the accelerator
NVIDIA's financing strategy solves a real problem today.
There are customers who want compute. Labs and hyperscalers cannot finance every gigawatt themselves. NVIDIA can help turn AI factories into an investable asset class and bring huge pools of outside capital into the market.
That weakens the most literal version of Dylan and Dwarkesh's centralization scenario.
OpenAI does not need to own the Ohio data center.
SB Energy can own it.
A pension fund, infrastructure fund, or private-credit vehicle can finance the asset.
But if OpenAI is still the highest-value customer, OpenAI gets the intelligence.
And the more successful NVIDIA becomes at solving today's shortage, the faster supply can eventually catch up.
NVIDIA is therefore both a pressure valve for today's scarcity and a potential accelerator of tomorrow's overbuild.
The six signals I'd actually watch now
Forget total token counts by themselves. Forget another giant data-center announcement.
A genuine transition from scarcity toward an AI infrastructure bust would look more like several of these turning together:
- Same-generation GPU rental prices fall for a sustained period. Baker himself says a dramatic, persistent drop would scare him.
- GPUs become easy to obtain. Short lead times would tell us physical scarcity has broken.
- HBM and server-memory prices loosen. LTAs matter much less when nobody fears future allocation.
- Full-price AI retention weakens. Users love cheaper tokens. The important question is how much consumption survives normalized pricing.
- Operating cash flow stops catching up with CapEx and debt. Then Zitron's financing problem becomes much more dangerous.
- Outside enterprise demand stops broadening. If growth increasingly depends on a small circle of mutually financed counterparties, Zitron's concentration argument gets stronger.
Falling compute prices plus weakening full-price customer demand would get my attention very quickly.
What you can actually do with this
The macro predictions are fun. The useful part is turning them into a short list of signals that help with real decisions.
1. Track value per AI dollar, not AI activity
For your own company, measure completed work, revenue, cost savings, cycle-time reduction, or quality improvement against the full cost of the AI workflow.
Token volume is a usage metric. It is not a return-on-investment metric.
2. Build model portability into important workflows
Baker and Dylan's scarcity scenario implies frontier intelligence can become more expensive or less available. Zitron's scenario implies providers can retrench when subsidies become uneconomic.
Both point to the same operational advice: avoid building a critical process that only works with one model, one API, or one pricing tier when a fallback is practical.
3. Watch the inference-to-training split
Inference is compute used to serve customers. Training and R&D are compute used to make future systems better.
If frontier labs keep redirecting capacity toward research, temporary shortages or slower product growth may coexist with accelerating internal progress. That distinction matters for anyone interpreting revenue slowdowns or model-release gaps.
4. Separate outside demand from financed demand
When evaluating an AI vendor or infrastructure provider, ask where the money ultimately begins.
Customer budgets funded by measurable productivity are stronger evidence than demand funded primarily through equity raises, cloud credits, or debt that assumes future adoption.
5. Treat compute, power, and credit as linked inputs
The cost of AI may increasingly be determined by electricity, data-center capacity, memory, networking, and financing rather than model quality alone.
A team budgeting a large agent deployment should stress-test costs instead of assuming today's token prices remain stable forever.
6. Keep human review where AI errors have asymmetric costs
Zitron is right about one thing even if his macro thesis proves too bearish: failed AI output still consumes compute, and agent errors can propagate into downstream actions.
For high-stakes workflows, the relevant metric is successful outcomes per dollar, not merely autonomous runs per dollar.
Our take: shortage first, glut second
We have covered the AI bubble's bull and bear case before. These newer conversations make the present condition easier to call.
AI compute is undersupplied today.
Microsoft says demand exceeds capacity. NVIDIA says the same thing through its supply-chain commentary. GPU rental anecdotes and memory agreements point in the same direction. Paid enterprise AI demand is broad enough that the entire story cannot reasonably be dismissed as two labs passing money around.
That does not make $11T of investment safe.
It tells us where the danger probably comes from.
The shortage creates extraordinary prices.
Extraordinary prices justify extraordinary financing.
Extraordinary financing creates extraordinary supply.
And by the time that supply arrives, the critical variable is no longer how badly OpenAI wanted another gigawatt in 2026.
It is how much the eventual customer is willing to pay for the intelligence that gigawatt produces.
That is where Ed's critique becomes useful. He may be too bearish on how much economic value AI already creates, and the outside-demand data pushes against some of his strongest claims. But he is asking the correct question about the other side of the ledger.
Gavin thinks productivity and labor substitution can ultimately fund enormous token consumption.
Dylan thinks the best labs may create so much value per megawatt that they can simply outbid everyone else.
Jensen is building a financing system capable of turning those expectations into physical infrastructure.
Ed is asking what happens if the customer at the very end of that chain says: actually, not at that price.
Our conclusion is that the next AI infrastructure problem is more likely to be too little compute than too much. The eventual crash, if it comes, is more likely to be caused by the industry's response to that shortage than by an oversupply that already exists today.
The signal to watch is not CapEx by itself.
Watch when compute becomes easy to buy, prices start falling, and customers stop following them down with more spending.
Until then, oversupply is a forecast.
The most plausible AI bubble path is shortage first, glut second.
And the question sitting underneath $11T of investment remains unanswered:
What is the next useful unit of intelligence worth when somebody finally has to pay the whole bill?
Full video insights
The sections below preserve the detailed, timestamped insight extractions used to build this article. Each timestamp jumps directly to the relevant moment in the source video.
Gavin Baker: The Positive-Sum Future of AI
- (0:00) Gavin Baker predicts the 21st century may be remembered as “the age of Elon and Jensen” because of how deeply they are changing infrastructure and industry.
- (0:13) He argues every profound technology creates a bubble because markets get ahead of reality, and that overvaluation then finances an overbuild.
- (0:31) Baker argues data centers may be a major positive for working-class Americans because they are helping reindustrialize parts of the US.
- (0:45) He predicts an increasing fraction of global compute will move into orbit and says asteroid mining will become a real industry.
- (1:22) His favorite reality check when talking to companies is: “Can you tell me one quantitative data point in your business that’s getting worse?” He says he struggled to find one in July and August.
- (1:47) Baker says OpenAI, open source, and Grok were all accelerating even while public AI stocks had experienced large drawdowns.
- (2:41) He embraces a positive-sum thesis: frontier labs, open source, neoclouds, inference clouds, applications, and chip companies can all win simultaneously.
- (3:31) He hypothesizes Anthropic may have cleaned up accounting and rebased metrics ahead of an IPO, setting up future reported reacceleration.
- (4:04) Baker describes release timing among frontier labs as strategic gamesmanship because companies often hold more capable checkpoints than the ones currently public.
- (4:41) He thinks having OpenAI and Anthropic as public companies would improve transparency for investors trying to model the AI economy.
- (5:10) Baker is uneasy with Anthropic asking candidates how they would feel if equity went to zero; his point is that even mission-driven labs need valuable equity to finance compute.
- (5:34) He calls Anthropic an “accidental enterprise company,” suggesting commercialization emerged from its research mission rather than leading it.
- (5:50) A key lab-economics thought experiment: if a lab has 10 GW and shifts most of it from monetized inference to training, reported revenue can fall sharply even as long-run capability improves.
- (7:32) Baker says public markets will have to learn that model labs can control near-term revenue through checkpoint releases, pricing, and how they allocate compute between training and inference.
- (8:08) His core mental model is “and, not or”: frontier models, N-minus-one models, open source, apps, clouds, and labs do not have to be mutually exclusive winners.
- (8:48) He places Nvidia at the center of that positive-sum stack because nearly every winning scenario still consumes more accelerated compute.
- (9:11) Baker expects labs to reinvest incremental operating cash flow into training for a long time rather than optimize for free cash flow.
- (11:18) He frames the compute race as asymmetric: underinvesting can permanently lose share, while overinvesting can threaten solvency.
- (11:52) Baker cites neocloud examples with roughly 9- to 10-month payback periods to argue some infrastructure projects can be unusually fast-returning.
- (14:49) On demand, he argues the industry may still be serving only a small fraction of the world’s knowledge workers, leaving enormous room for diffusion.
- (16:35) He reports his own firm’s token usage rose roughly 100x from March to August and says agent adoption could multiply that again.
- (17:40) Baker says younger employees who grew up with AI workflows are beginning to look structurally more productive than older workers who are adopting the tools later.
- (18:04) He describes Grokbot rebuilding automations in seconds that previously took hours with coding agents and calls the experience a “ChatGPT moment.”
- (19:16) His distinction is that coding assistants are mostly reactive, while emerging agents can recommend and initiate actions, making the workflow more proactive.
- (20:28) If AI diffuses across 1.5B knowledge workers, Baker expects token demand to be effectively “endless” relative to today’s usage.
- (20:48) He sees bubbles as a normal part of transformative technologies, but says debt makes the timing risk more dangerous because liabilities survive the hype cycle.
- (21:48) Physical constraints such as copper, transformers, mines, and grid buildout may slow AI deployment, which he says could be socially useful even if economically frustrating.
- (23:12) Baker thinks the AI industry has over-indexed on doom narratives and needs to communicate concrete positive outcomes more effectively.
- (24:07) His data-center case is grounded in local economics: construction and trades jobs, tax revenue, and industrial investment in towns that have not seen major capital flows.
- (25:12) He argues “we need to beat China” is too abstract to persuade the public; the industry should explain tangible everyday benefits.
- (27:03) Low US natural-gas prices relative to Europe and Asia are presented as a structural advantage for electricity-intensive manufacturing and compute.
- (28:52) Baker’s communications advice is to copy early Meta: tell specific stories about how normal people and businesses benefit, not abstract stories about technological destiny.
- (29:27) His base case is underbuild through 2028, not oversupply; he worries more about compute scarcity than a near-term glut.
- (30:02) He says severe scarcity could drive compute prices or effective token costs up dramatically, even by an order of magnitude in a stress case.
- (30:21) Compute scarcity can become inequality: wealthy users and large enterprises buy the best intelligence while ordinary users are priced out.
- (31:52) Baker stresses that open-weight models do not eliminate compute demand; cheap or free weights still require expensive inference.
- (33:10) He uses Starlink as a model for positive-sum infrastructure: build where terrestrial economics fail, then unlock demand that was previously unreachable.
- (34:25) The orbital-compute concept involves airplane-sized data centers in sun-synchronous orbit, with thermal design using radiators on the dark side.
- (35:36) Baker recounts a skeptical physics PhD who changed his mind after discussing orbital compute with SpaceX engineers, which Baker uses as evidence the engineering case is stronger than it sounds.
- (36:18) His threshold claim is not that orbital compute is easy, but that there is no obvious physics reason it cannot work.
- (37:06) The economics hinge on full Starship reusability: if launch cost falls enough, terrestrial power, cooling, and land savings can offset orbital deployment cost.
- (38:39) He expects training to remain mostly terrestrial because of latency and networking, while an increasing share of inference may be able to move to orbit.
- (39:51) Baker sees Starlink broadband plus mobile connectivity as a huge addressable market before even counting AI services.
- (40:33) He predicts a future SpaceX/xAI bundle that combines Starlink connectivity, Grokbot, and advertising.
- (40:51) The SpaceX/xAI infrastructure bet is presented as “heads we win, tails we win”: if the first-party app works, compute is useful; if not, scarce compute can still be sold.
- (42:17) Baker expects a global network of “starbases” and ultimately thousands of Starship launches per year if reusable launch economics work.
- (44:10) He predicts asteroid mining, with robots capturing resource-rich asteroids and processing them in controlled orbits.
- (47:50) His long-run industrial vision is heavy industry migrating off Earth while the planet becomes more residential and service-oriented.
- (48:07) Baker argues Microsoft’s failure to own the best frontier model is less damaging in a future where enterprises use ensembles of specialized models.
- (48:28) He does not expect one model to dominate every task; large companies will combine frontier systems, open source, and their own fine-tuned models.
- (48:52) Chip companies have an incentive to sponsor powerful open models because cheaper model margins can increase total token demand.
- (51:29) Baker’s likely enterprise stack is an internal model for proprietary data plus one or two frontier models, with the frontier model planning and cheaper systems executing.
- (52:31) He calls the competition over the enterprise intelligence stack “corporate chess,” with routing and orchestration as strategically important layers.
- (54:09) The prize is the “abstraction layer of intelligence”: the system that decides which model, data, and tool should handle each job.
- (55:44) Baker praises Cursor for staying product-focused and shipping practical autonomy rather than centering grand AGI rhetoric.
- (57:04) He thinks coding is unusually suitable for agents because outputs are verifiable and documented; enterprise workflows are messier and may favor incumbent platforms.
- (57:39) He expects the abstraction layer to be contested by frontier labs, Databricks, Palantir, inference providers, Microsoft, Snowflake, Salesforce, Workday, and vertical applications.
- (59:40) Over time, Baker thinks the lowest-cost intelligence provider will probably own or tightly control its compute stack.
- (1:00:06) He suggests evaluating hyperscalers using enterprise value versus net property, plant, and equipment as a kind of “AI price-to-book” ratio.
- (1:00:35) Nvidia’s strategy is described as vertically integrated but horizontally open: own many system components while supporting a broad ecosystem.
- (1:00:58) Baker’s advice to semiconductor startups is “thank you Jensen”: do not attack Nvidia head-on; find edges of the stack where a small share can still create enormous value.
- (1:01:37) He notes Nvidia spans accelerators, CPUs, networking, and other components, making competition a system problem rather than a single-chip problem.
- (1:05:27) Open source is strategically good for Nvidia because lower model margins can push more value into compute and increase token consumption.
- (1:06:00) Baker emphasizes Nvidia’s supply-chain advantage: it locks up fab capacity, memory, networking, optics, capacitors, racks, and financing, not merely GPU designs.
- (1:11:45) He calls Nvidia the “central bank” or “Federal Reserve” of AI because financing and allocation decisions can influence the entire ecosystem.
- (1:12:16) In a supply-constrained market, Baker says deal structure reveals preference: strategic investments, warrants, and financing terms show which customers chip companies most want to support.
Dylan Patel and Dwarkesh Patel: Two labs will soon control most of the world's workforce
- (0:37) Dylan says that by the end of last year, most U.S. GDP growth was coming from AI infrastructure, and roughly one-third of compute coming online this year is ultimately destined for OpenAI and Anthropic.
- (1:03) Dylan estimates AI CapEx at a little over $1T this year and more than $2T in 2028, with the labs taking an increasing share.
- (1:41) Dylan says Anthropic turned profitable in Q2 and believes OpenAI could turn profitable in Q3 as Codex, GPT-5.6, and related products grow.
- (2:26) He says lab margins have “really skyrocketed” over the last year and a half, while base compute costs remain roughly $10M-$15M per megawatt.
- (3:08) Dylan says Anthropic has reached as much as $50M of revenue per megawatt, creating a flywheel where $10 of inference capacity can generate $50 of revenue that can then be reinvested into training.
- (4:20) Dylan expects OpenAI and Anthropic together to take 40%-50% of new compute next year.
- (5:00) He predicts that by the end of next year, roughly half of the world's incremental compute will already be serving OpenAI and Anthropic.
- (5:26) Dwarkesh summarizes the trajectory as world compute roughly doubling while frontier-lab compute triples, potentially taking a lab from 2 GW to 6 GW, then 18 GW, then 54 GW.
- (6:20) Dylan notes that new accelerators are also getting 3-5x more performance per watt than prior generations, so counting gigawatts alone understates the concentration of usable compute.
- (6:45) His base-case implication is that by late 2028, if current trends continue, OpenAI and Anthropic could control most of the usable FLOPs in the world.
- (7:18) Dwarkesh reframes the compute bottleneck through semiconductor manufacturing: how much fab tooling is actually needed to produce another gigawatt of leading-edge AI chips every year?
- (8:19) Dwarkesh estimates roughly $6B of fab CapEx could create the manufacturing capacity for one additional gigawatt of compute per year.
- (8:54) He argues that over five years, that $6B of fab investment could ultimately enable more than $1T of downstream AI revenue.
- (9:39) Their point is that when a supply-chain bottleneck sits between $1 of investment and something like $100 of eventual economic value, capitalism will throw extraordinary amounts of money at expanding that bottleneck.
- (10:19) Dylan jokes that if someone could buy an ASML EUV machine for $400M, they might be able to wait and resell it for north of $1B because of scarcity.
- (10:40) Dylan's constraint is latency in the industrial supply chain: the economic signal can be overwhelming, but the supply chain does not react immediately.
- (11:41) Dylan says roughly 100 ASML tools in 2030 still looks like the current production trajectory, though throwing $10B at bottleneck suppliers such as Carl Zeiss could change that.
- (12:06) He does not expect the labs to directly force that kind of all-out manufacturing expansion within the next few years because the world is capital constrained.
- (12:29) Dylan estimates next year's AI buildout at well north of $2T of CapEx, far more than labs' near-term cash flows can fund themselves.
- (13:41) Dylan says getting OpenAI and Anthropic to a combined 100 GW in 2028 would require them to absorb perhaps 70%-80% of incremental compute, which would severely disrupt compute markets.
- (14:37) Dylan argues even an ordinary operator can currently acquire modern GPUs, run open models with software like vLLM or SGLang, sell the inference, and generate more revenue than the compute costs.
- (15:00) That basic profitability is already pushing the historical $10M-$15M per megawatt price of compute upward.
- (15:17) Dwarkesh asks whether the equilibrium price becomes $25M or $40M per megawatt if frontier labs continue generating much more value from each megawatt than everybody else.
- (15:32) He adds that recursive self-improvement could widen the monetization gap further if labs possess internal models that help build better successors before those models are released publicly.
- (16:03) Dylan believes the labs might need to pay $25M, $30M, or even $50M per megawatt to capture 70% of 2028 compute.
- (16:15) Dylan introduces the major counterforce: regulations and safety constraints may slow frontier labs more than open-source Chinese model developers.
- (17:06) If labs cannot release their strongest models, Dylan says their revenue per megawatt may stop climbing fast enough for them to outbid everybody for compute.
- (17:38) Dwarkesh's intuition pump: if models were literally equivalent to fully automated white-collar workers, the economically rational price of AI compute would become enormous.
- (18:25) Dylan stresses that frontier labs currently capture only a fraction of the total economic value their models create; most surplus still accrues to users.
- (18:43) His example is Jane Street: it can potentially create far more value with GPT-5.6 or Anthropic tokens than the lab itself earns as profit from selling those tokens.
- (19:25) Dwarkesh asks whether markets eventually equilibrate so that compute prices rise close to whatever OpenAI and Anthropic can earn from that compute.
- (20:05) He highlights the strange current arbitrage: Anthropic can effectively take compute that costs $10 and turn it into something worth $100.
- (20:16) Dylan lays out the AI value stack: end users currently capture the most value, apps very little, models have recently shifted from negative to massive positive gross margins, and hardware historically captured most of the surplus.
- (20:48) A year earlier, Dylan says model companies were effectively destroying value at the model layer by selling tokens for less than the underlying infrastructure cost while chips and fabs captured the margin.
- (21:56) Dylan says OpenAI and Anthropic are now rapidly increasing their share of total value capture.
- (22:03) SpaceX showed another possible equilibrium: compute owners can simply raise prices to $25M-$40M per megawatt when labs desperately need capacity.
- (22:31) Despite those high spot deals, Dylan expects most compute to still transact below $20B per gigawatt by the end of next year because financing constraints matter.
- (22:59) SpaceX's advantage was having compute already built and uncommitted, giving it leverage to sell scarce capacity to whoever could monetize it best.
- (23:35) Dylan argues Meta and SpaceX are the only plausible number-three players because their balance sheets let them build compute before they have an end customer.
- (24:05) That gives Meta and SpaceX valuable optionality: use the compute internally or rent it to OpenAI and Anthropic at dramatically higher prices.
- (25:40) Dylan says OpenAI or Anthropic could plausibly reach $70M-$80M of blended revenue per megawatt by the end of 2027 if they can keep releasing their best models.
- (26:16) The higher lab economics get, the more every upstream supplier has an incentive to raise prices too, from Nvidia through memory vendors and foundries.
- (26:40) Dylan describes a bullwhip effect in value capture: price increases ripple backward through the supply chain instead of rebalancing everything instantly.
- (27:19) Dwarkesh challenges the relatively conservative revenue-per-gigawatt forecast by asking what happens if current model progress continues or recursive self-improvement begins.
- (27:32) Dylan's answer is that the best model in existence may not be commercially deployable, making regulation rather than technical capability the binding constraint.
- (28:14) Dylan argues that unreleased or “neutered” models cannot be fully used for things like inference optimization, reducing the revenue growth labs could otherwise get from improved capability.
- (28:49) They point out that data-center restrictions themselves decrease supply and raise the cost of compute, citing emerging restrictions or proposed restrictions in New York, Texas, and Ohio.
- (29:16) Dwarkesh says in a takeoff scenario, a frontier lab could rationally keep its strongest internal model six months ahead of what outsiders can access.
- (29:44) Dwarkesh poses a public-markets conflict: if inference generates $100B per gigawatt, investors may resist moving another 10% of a lab's compute from serving customers into training.
- (30:31) Dylan gives what he calls a very non-consensus prediction: frontier labs will allocate less and less compute to inference over time.
- (30:51) He argues the consensus has it backward: most future compute will go toward training-related forward passes rather than revenue-generating inference.
- (31:10) Dylan's reasoning is simple: if the choice is dividends and buybacks versus building AGI, OpenAI and Anthropic's executives and boards will choose to build AGI.
- (31:46) Dylan says the labs may find their best and fastest models more valuable internally than as external products.
- (32:01) His core tradeoff is that $100M per megawatt of immediate inference revenue may be worth less than the future progress created by applying that same compute to AI research.
- (32:28) Dylan says inference primarily serves one strategic purpose for the labs: generating enough cash to keep expanding the training fleet.
- (32:43) He believes this shift is already happening, with the share of lab compute going to R&D increasing during the last few months.
- (33:06) Dylan's evidence is that Anthropic keeps adding compute while monthly revenue additions have plateaued, implying more marginal compute must be flowing into R&D rather than inference.
- (33:29) Dylan projects global AI compute additions of roughly 30 GW this year, 50 GW next year, 70 GW in 2028, and 90-100 GW in 2029.
- (34:02) He says once the world is adding 100+ GW per year, you're implicitly talking about extraordinary GDP growth, potentially entering an RSI-driven economic regime.
- (34:37) Dylan says the U.S. share of newly deployed compute has climbed dramatically since 2022 while China's has fallen.
- (34:58) He estimates roughly 70% of AI data-center watts are now being deployed in America, while China accounts for less than 10% of incremental deployment.
- (35:31) Dylan expects China to remain below 10% of new global compute for now and estimates China could have 30 GW or less of AI compute in 2028.
- (35:43) He expects China's domestic compute buildout to begin accelerating sharply around 2028 as new fabs come online.
- (36:10) By 2028, Dylan expects SMIC, CXMT, and related domestic manufacturing to be producing millions of units per year and adding perhaps 5-10 GW of domestic AI compute that year alone.
- (36:39) Dylan stresses that Chinese gigawatts will likely represent much less actual performance than American gigawatts because the chips will be worse.
- (37:01) The speed of China's catch-up depends heavily on future U.S. export controls, domestic Chinese semiconductor equipment progress, and legislation such as the MATCH Act.
- (37:19) Dylan nevertheless expects China to hockey stick because scaling manufacturing rapidly is a major Chinese strength.
- (37:42) Dylan says 50 incremental GW in China in 2029 is completely reasonable, though some may still come from foreign chips.
- (38:10) Dwarkesh notes this could imply a single leading Western lab in 2028 has more quality-adjusted compute than all of China has in 2029 or 2030.
- (38:37) Dwarkesh says the magnitude of the current compute gap makes him view U.S. export controls as a potentially major strategic success if automated coding and automated research arrive before China catches up.
- (39:03) He argues the compute gap could leave China far behind precisely when automated coders begin turning into automated AI researchers.
- (39:30) Dylan adds that U.S. financial markets' willingness to aggressively fund startups is another source of the gap, not only export controls.
- (40:11) A major puzzle is that Chinese labs remain surprisingly competitive in public model quality despite possessing dramatically less compute.
- (40:17) Dylan estimates leading Chinese labs commonly have only 100-200 MW of total compute, aside from ByteDance Seed, while Anthropic is heading above 5 GW.
- (40:50) Dylan breaks a frontier lab's current compute allocation into roughly 50% research, 10% development, and 40% inference.
- (41:03) Research means running experiments on architectures, datasets, hyperparameters, attention methods, and other ideas before committing to a large production training run.
- (41:22) Dylan says Anthropic's Mythos pretraining run used under 200 MW for roughly two months, even though the company had multiple gigawatts available overall.
- (41:55) The reason labs cannot simply pour every gigawatt into one run is that clusters are difficult to co-locate and coordinate, multi-site training is hard, and adding more RL rollouts does not necessarily improve results.
- (42:28) As automated coding and automated research improve, Dylan expects the distinction between research compute and training compute to become increasingly fuzzy.
- (42:42) Continual learning is another force that could push a much larger fraction of total compute directly into ongoing model training.
- (42:51) Dylan estimates that building 100 GW per year at current prices would mean roughly $5T of annual IT CapEx before including supporting infrastructure.
- (43:24) He says headline AI CapEx numbers often count only servers, networking, fiber, transceivers, and related IT, while excluding the data centers and power plants that must be financed in advance.
- (44:10) Once those future data-center and power requirements are included, Dylan says annual incremental AI CapEx could plausibly approach $10T by the end of 2030.
- (44:26) Dwarkesh notes that if most of that buildout happened in the U.S., data centers could consume something like a quarter to a third of today's U.S. GDP, making politics a likely constraint.
- (44:49) Dylan's broader view is that pure capitalist reallocation points toward enormous AI investment, but politics, credit markets, and capital markets prevent the theoretical maximum from being built.
- (45:24) He asks the central financing question: if 2028 requires $3T-$4T of total AI CapEx and nobody is yet generating that much cash, where does the money come from?
- (45:55) Dylan says hyperscalers are already increasingly issuing debt to pay for CapEx, citing Meta, Amazon, Google, and eventually Microsoft.
- (46:17) By 2028, he expects hyperscalers and their suppliers to be raising hundreds of billions of dollars of debt, creating huge new competition for global capital.
- (46:25) Potential financiers include semiconductor companies themselves, traditional infrastructure funds, and ordinary investors reallocating away from mortgages or government debt.
- (46:51) Dylan says capital may shift from financing homes and governments toward hyperscaler, data-center, or even Anthropic debt because AI projects can offer much higher returns.
- (47:00) He suggests Anthropic could rationally pay something like 20% interest for incremental capacity if owning that capacity is still much cheaper than renting scarce compute at extreme prices.
- (48:51) The brothers introduce their most macroeconomic scenario: AI investment could trigger a sovereign debt crisis by radically increasing global demand for capital.
- (49:18) Dwarkesh's mechanism is that exceptionally high returns on AI infrastructure bid up interest rates for everybody else in the economy.
- (49:42) If AI borrowers can pay more, governments, conventional companies, consumers, and mortgage borrowers all face higher financing costs.
- (50:10) Dwarkesh thinks the U.S. can ultimately survive this because much of the AI infrastructure would be physically located in America and could become a new tax base.
- (51:34) He calculates that a 1 percentage-point increase in rates could push U.S. debt service from around 20% to 25% of tax revenue over time, while a 5-point increase could push it above 40% before accounting for new borrowing.
- (52:08) Dwarkesh thinks heavily indebted countries with weak tax bases, naming Pakistan and Nigeria as examples, could be hit especially hard under such a high-interest-rate regime.
- (52:34) Dylan says debt-heavy sectors such as consumer goods, telecom, and banking would also face severe pressure as AI infrastructure crowds them out of credit markets.
- (52:57) The relevant change may show up not only in central-bank policy rates but in credit spreads as massive AI borrowers compete for the same pool of lenders.
- (53:19) Dylan says labs rationally want to invest more than their cash flows because the future returns on additional compute are so attractive.
- (53:34) Regulations, voter backlash, safety restrictions, and rising interest rates all act as forces bending the compute-buildout curve below what simple profit-maximizing economics would suggest.
- (54:16) Asked how much debt the AI ecosystem might need, Dylan turns to SemiAnalysis's modeling.
- (54:30) Their model has roughly $11T of AI CapEx from 2024 through 2029.
- (54:42) Dylan estimates around $6T could be funded with cash and more than $5T with credit.
- (55:19) The $11T estimate includes the fact that labs increasingly generate their own cash, but Dylan still thinks they cannot cash-flow the entire desired buildout.
- (55:34) Even $5T of new credit would still leave supply below what unconstrained model-driven compute demand would want, pushing revenue per megawatt even higher.
- (56:06) Dylan says using substantial debt is economically optimal even if labs become highly profitable because you want to build more capacity than current earnings alone allow.
- (56:45) Asked to vibe an interest-rate number for the end of the decade, Dylan says growth itself is a reason rates should rise.
- (57:06) Meta has recently borrowed around 5%-6%; Dylan says he sees no reason a hyperscaler would not happily pay 8% if compute returns remain enormous.
- (57:31) The problem is that if hyperscaler borrowing costs rise by 250 basis points, Dylan expects borrowing costs for the rest of the economy to rise similarly.
- (58:03) Dwarkesh notes a second-order effect: higher interest rates mean higher discount rates, which can cause the present value of long-duration equities to collapse.
- (58:45) They compare the possible shock for developing countries to the Volcker shock, when sharply higher U.S. rates contributed to sovereign defaults across dozens of countries.
- (59:11) Dwarkesh predicts a similar wave of sovereign defaults could happen again as AI investment drives another global repricing of capital.
- (59:27) Moving further out, Dwarkesh says he thinks it is very likely the world economy will eventually reach a regime where it can double every year.
- (59:48) The key difference in a fully automated economy is labor: if you can double the effective labor force every year, labor stops being the bottleneck that makes today's economic doubling times so long.
- (1:00:03) Dwarkesh thinks annual growth in such a world could be tens of percent at minimum and potentially around 100%.
- (1:00:16) He expects interest rates to roughly track that much higher growth regime and speculates that rates could reach tens of percent in the 2030s.
- (1:00:30) Dwarkesh's extreme implication is that countries outside AI production default, non-AI stocks become worth very little under huge discount rates, and governments face radical fiscal stress.
- (1:01:00) Dylan translates the nerd speak: society has entered a regime where the opportunity cost of using capital for anything other than compounding automated production becomes extraordinarily high.
- (1:01:25) Rising opportunity cost of capital is, in Dylan's view, the underlying cause connecting AI investment, rising interest rates, falling equity multiples, and fiscal stress.
- (1:01:33) Dylan argues that if you are extremely bullish on AI, then even highly profitable semiconductor companies may deserve only 2x-3x earnings multiples because the discount rate for everything should be much higher.
- (1:02:05) His paradox is that if memory demand becomes as extraordinary as bulls expect, that itself implies an economic regime where equity valuations broadly compress and the stock market crashes.
- (1:02:20) Dylan thinks Meta may be logically worth far more than its current valuation because of its cash flows plus the enormous strategic value of the compute it is hoarding.
- (1:02:45) He describes the macro process as reallocating all capital toward AGI by pricing everybody else out.
- (1:03:00) The true limiter on AGI may therefore be how much disruption the rest of society tolerates, not how quickly individual research engineers can make technical progress.
- (1:03:32) Dylan says economic, regulatory, and political forces create substantial downward pressure against a perfectly vertical straight takeoff, even if models technically become capable of one.
- (1:03:52) His hope for a slower takeoff depends on governments restricting releases, internal usage, data centers, and other channels through which labs could rapidly compound capability.
- (1:04:05) Dwarkesh worries that blocking external deployment could perversely increase centralization because labs would keep more of their strongest capability private.
- (1:04:24) His nightmare scenario is a six-month public-release delay during RSI: internally, a lab could improve by orders of magnitude while everybody outside remains stuck on models years behind.
- (1:04:53) Dwarkesh argues that even slowing AI deployment by one calendar year may not mean much if compute is growing 2-3x annually and RSI compresses several years of algorithmic progress into one.
- (1:05:00) He estimates an RSI year could contain something like 3-6 years of today's AI progress.
- (1:05:16) Dylan counters that governments can also restrict internal use of powerful models, not only public release, which could genuinely slow the recursive loop.
- (1:05:39) Dylan expects future U.S. governments to tell labs to slow down internally as well because elected officials and constituents will increasingly dislike the effects of AI.
- (1:06:20) His bottom line is that real-world economic and political constraints will limit AI development and deployment even if eventual progress remains inevitable.
- (1:07:54) Dwarkesh introduces what he finds most startling: how much of the future world's labor supply could end up controlled by only a few companies.
- (1:08:12) His math combines frontier FLOPs growing 4-5x annually with the compute needed for a fixed capability level falling about 3x annually, implying the frontier's effective AI population could grow roughly 10x per year.
- (1:08:46) Dwarkesh sketches OpenAI going from an effective 10 million AI laborers to 100 million, then one billion, if those trends continue.
- (1:09:04) He says it is very plausible that before the end of this decade, one lab could contain more effective AI labor than there are humans on Earth.
- (1:09:17) This reframes centralization: most economically relevant people, measured by work output, could reside inside two labs that also consume an increasing share of world compute.
- (1:10:06) Dylan says if you believe frontier labs are the best users of scarce compute, plus believe in AI researchers, AGI, or RSI, then centralization of compute follows almost mechanically.
- (1:10:26) Even without RSI, Dylan says the effective population available at a given capability frontier is already increasing roughly 10x per year.
- (1:10:45) Once RSI starts, they speculate effective population or intelligence could grow 100x or even 1,000x per year, or some mixture of the two.
- (1:10:59) Dylan directly asks what world avoids centralization because every force they have discussed seems to point the same way.
- (1:11:31) Dwarkesh identifies the first structural force: AI training has enormous economies of scale because one training improvement can be amortized across billions of sessions or users.
- (1:11:49) The second force is scarce compute: the slightly better model can monetize each scarce GPU more effectively, earn a higher markup, and therefore outbid rivals for more compute.
- (1:12:00) A possible third flywheel is continual learning: the most widely deployed model gets the most real-world interaction data, potentially making it even stronger.
- (1:12:24) Dwarkesh calls it an important intellectual project to find a credible decentralized, broadly empowered post-AGI future that still takes these economies of scale seriously.
- (1:12:41) The obvious alternative is government control, but Dylan says plainly that he does not trust the government, Dario, or Sam.
- (1:12:58) They struggle to see a path that does not require choosing between different kinds of centralization unless some currently unforeseen force changes the structure of AI progress.
- (1:13:18) Dylan argues AI may flip a traditional advantage of decentralized capitalism: a highly centralized AI economy could theoretically grow faster because intelligence and decision-making themselves scale centrally.
- (1:13:52) Even today, he notes, a remarkably large portion of the AI economy is concentrated among Nvidia, OpenAI, Anthropic, and a handful of hyperscalers.
- (1:14:04) Dylan sees only a few obvious brakes on extreme concentration: slower AI progress, heavy government regulation, or political resistance that slows resource accumulation by the leaders.
- (1:14:42) Dylan offers one potential saving grace: Anthropic still does not capture most of the value created by Anthropic's models.
- (1:14:59) His example is that if Anthropic earns $100M per megawatt, a customer like Jane Street might be creating $300M-$500M of value from that same megawatt.
- (1:15:33) Dwarkesh points out why even that saving grace may disappear: shifting inference compute into R&D assumes returns to labor inside the AI lab eventually exceed returns outside it.
- (1:15:52) Dylan agrees that if tokens are worth more inside the lab than outside, the rational move is to stop exporting as much intelligence and use the compute internally.
- (1:16:20) The conversation ends on the centralization thesis: if Anthropic can create hundreds of millions of dollars per megawatt internally, why allocate that scarce compute to outside customers at all?
Ed Zitron: The Man Who Calls BS on AI
- (2:50) Ed Zitron’s central thesis is that generative AI is “at its heart” a con because he believes it is marketed as far more capable, transformative, and financially sound than the underlying products justify.
- (3:19) He argues AI has been sold as “magic” that will replace jobs and cure major problems, while he sees today’s products as expensive, unreliable cloud software.
- (3:52) Zitron says he is a technology enthusiast, not an anti-tech critic, but believes current AI fails basic expectations that its marketing creates.
- (4:03) He mocks the complexity of modern AI setups, arguing that a supposedly intelligent autonomous system should not require elaborate combinations of models, prompts, and harnesses to function reliably.
- (4:44) Zitron claims a large share of hyperscaler AI revenue is tied to spending by OpenAI and Anthropic, which he describes as structurally unprofitable customers.
- (5:30) He criticizes public companies for not cleanly disclosing AI revenue and for using “annualized run rate” figures he says are too loosely defined.
- (6:22) Zitron accepts that AI has some value but argues the value must be judged against more than $1T of capital expenditure.
- (6:58) He uses large data-center power requirements to illustrate the physical scale of the investment, citing Stargate Abilene as his example.
- (7:46) His financial concern is that companies are increasingly using debt and enormous capex for an industry whose direct revenue he believes remains comparatively small.
- (8:36) Zitron disputes the idea that headline adoption proves organic demand, calling AI integration across Google, Microsoft, and Amazon “the largest non-consensual push of technology in history.”
- (9:11) He says many people use AI as a search substitute because it is heavily promoted and sometimes better at broad query ingestion, not because every use case has durable standalone value.
- (9:34) He argues the current investment wave has no obvious post-bubble reuse story because AI GPUs are specialized assets.
- (11:28) He links generative AI to the older SEO-slop problem, arguing models make it easier to manufacture generic content designed to rank rather than help readers.
- (12:03) Zitron explains token billing as the hidden meter underneath AI services: users pay for input tokens and output/reasoning tokens even when consumer subscriptions hide the meter behind rate limits.
- (12:58) He cites estimates that a $200 monthly ChatGPT subscription can theoretically consume far more than $200 worth of API-equivalent tokens, using that as evidence of subsidy.
- (13:18) His key point is that consumers do not see the true marginal cost of heavy AI usage because subscription pricing obscures it.
- (13:54) He says enterprise enthusiasm changes quickly when companies move from flat subscriptions to paying something closer to actual token consumption.
- (15:03) Zitron emphasizes the capacity-planning problem: inference providers pay to stand up GPU capacity whether or not demand arrives, but underprovisioning creates outages and churn.
- (15:53) He argues AI is unusual because users pay for compute even when the model hallucinates or damages the task result.
- (16:44) Zitron rejects the standard “spending ahead of value” defense because he believes inference economics have not improved enough and may be getting more expensive.
- (17:07) He says he does not think the industry began as a deliberate conspiracy; his view is that companies initially expected chips, efficiency, and willingness-to-pay to catch up.
- (17:20) His criticism is that investors should have reassessed once losses remained enormous, rather than continuing because AI narratives supported stock valuations.
- (18:23) He cites Microsoft-related figures to argue AI-derived revenue is too small relative to capex plans, concluding “the math does not make sense.”
- (20:13) The interviewer offers the Innovator’s Dilemma counterargument: disruptive technologies often begin worse and uneconomic before their higher ceiling becomes obvious.
- (20:36) The interviewer compares hallucinating AI with early cars that were unreliable, expensive, and constrained by strange rules before eventually overtaking horses.
- (24:31) The interviewer says his personal experience is that outright hallucinations have become much rarer, while Zitron disagrees and offers counterexamples.
- (24:56) Zitron uses Bloomberg’s natural-language query interface as an example of AI that can be useful because it generates queries against a known data source, but says it can still quietly return bad results.
- (27:00) The interviewer cites a hallucination leaderboard showing a steep drop on simple summarization tasks and argues the trajectory matters even if errors remain.
- (27:28) The interviewer compares AI’s early technical limitations with dial-up internet, arguing new platforms can be frustrating before infrastructure matures.
- (31:37) The conversation turns to whether process matters when a human and AI produce the same output; the disagreement centers on trust, accountability, and error modes rather than aesthetics alone.
- (36:00) Zitron argues that human professionals have reputational and organizational accountability for mistakes in a way a model does not, making the same-looking error different in practice.
- (39:56) The interviewer pushes on whether users already treat AI probabilistically, similar to asking another person for advice, while Zitron keeps returning to the scale and opacity of model error.
- (43:26) Zitron argues the clearest market test would be whether people keep using frontier AI at prices that reflect its full cost instead of heavily subsidized flat-rate plans.
- (48:09) The dot-com comparison becomes a core debate: the interviewer sees bubbles as capable of financing durable infrastructure, while Zitron argues AI’s assets and economics are different enough that the analogy is weak.
- (52:10) Zitron argues data-center commitments create unusually large fixed risks because hardware depreciates quickly and power/buildout commitments persist even if model demand falls.
- (56:51) He says Google has already made search worse through ad and ranking incentives and sees generative AI as another layer that can worsen the information environment.
- (1:04:23) On jobs, Zitron argues there is not convincing macroeconomic evidence that generative AI is already replacing work at the scale implied by public narratives.
- (1:08:47) He distinguishes genuine task automation from executive rhetoric about replacing workers, arguing the latter is often used to pressure labor even before systems can fully substitute for staff.
- (1:12:57) The interviewer challenges whether Zitron’s constant focus on fraud and hype risks becoming its own narrative that ignores real utility; Zitron responds that scrutiny is necessary precisely because the bullish story is dominant.
- (1:18:18) On cybersecurity, Zitron acknowledges offensive AI capability can be useful or dangerous but resists framing current models as autonomous cyber superintelligence.
- (1:21:39) He argues claims of AI-driven economic growth often conflate growth in existing cloud, advertising, and platform businesses with revenue directly caused by AI.
- (1:23:45) On the US-China “AI race,” Zitron rejects spending trillions simply because leaders invoke China, asking what concrete public benefit the race is meant to produce.
- (1:25:19) He is similarly skeptical that robotics means imminent mass unemployment, separating impressive demos from widespread economically viable deployment.
- (1:30:02) On agentic AI, Zitron argues adding tools and loops does not remove core reliability and cost problems; it can compound them by letting errors propagate through actions.
- (1:32:44) The interviewer compares AI adoption with the internet’s rise, while Zitron argues bundling, defaults, and fear-based messaging make today’s adoption less cleanly comparable.
- (1:37:20) Zitron says the phrase “AI” is doing too much work, allowing companies to package many different software improvements into a single transformative narrative.
- (1:41:16) When asked what he personally uses generative AI for, Zitron says there are narrow utilities he finds useful but does not see those uses as evidence for the scale of current investment.
- (1:45:14) He rejects the idea that larger models necessarily imply a smooth path to general intelligence, arguing intelligence claims often move faster than measurable product reliability.
- (1:47:30) On future jobs, the disagreement is less about whether AI can automate tasks and more about whether that automation will be dependable and cheap enough to replace whole roles.
- (1:50:21) Zitron’s future scenario is not “AI destroys humanity” but “AI spending breaks financially,” followed by a painful revaluation of the companies and infrastructure built around it.
- (1:52:51) The interviewer argues workflows have already changed for many power users; Zitron’s reply is that anecdotes of transformed work do not yet justify trillion-dollar system-wide economics.
- (1:55:27) The interviewer reads bullish CEO statements from Google, Amazon, and Meta that explicitly say underinvesting is a greater risk than overinvesting.
- (1:56:08) Zitron jokes that Meta’s capex posture resembles Lord Farquaad’s “some of you may die” line, because he sees executives socializing the downside of a gamble they personally control.
- (1:58:08) Asked what would change his mind, Zitron says he would need a dramatic hardware breakthrough that reduces costs by roughly three orders of magnitude.
- (2:00:50) He challenges stories about models “blackmailing” people, arguing some widely reported examples involved researchers prompting models inside contrived scenarios rather than autonomous real-world coercion.
- (2:08:34) Zitron says yes, he believes there is an AI bubble, but describes it as broader than startups because public mega-cap valuations and infrastructure commitments are intertwined.
- (2:09:01) He sees OpenAI’s ability to finance commitments and eventually access public markets as a critical weak link in the bubble scenario.
- (2:13:33) His “tech depression” thesis is that a sharp revaluation of AI-exposed mega-caps would hit ordinary retirement accounts because those companies dominate major indices.
- (2:14:11) He predicts Nvidia revenue could fall 50% to 70% in a bust, explicitly tying that forecast to his belief that today’s demand contains circular financing and overbuild.
- (2:14:53) He argues venture capital is also exposed because more than half of recent VC dollars went into AI and many application companies depend on subsidized model economics.
- (2:16:20) Zitron says the scale of OpenAI and Anthropic infrastructure commitments makes a simple bailout or bridge financing story difficult in his view.
- (2:17:44) He criticizes “paper gains,” noting that strategic investments can lift reported investment values even before those holdings produce realized cash returns.
- (2:19:13) Asked what the public should do, Zitron says he personally holds cash, does not trust the market, and is uncomfortable giving financial advice.
- (2:19:40) His practical advice is to act as if volatility is coming and to be skeptical of tech-company promises rather than treating investor-relations language as neutral fact.
- (2:21:53) Zitron says his anger toward AI CEOs comes from feeling that ordinary people would never receive the same tolerance for misleading claims or repeated failure.
- (2:22:38) He says he writes at length because he wants readers to see the evidence chain behind his conclusions rather than accept a contrarian slogan.
- (2:25:20) The closing answer pivots from criticism to relationships: Zitron says showing appreciation, love, and support for people around you is a central antidote to getting consumed by negative work.
- (2:26:50) His final actionable point is simple: tell people you love their work or that you care about them; he thinks people do this far less often than they should.