AI Economics 101: How AI Companies Make Money, Why Chips Are the Bottleneck, and What It All Means for Your Job

We watched three of the most important conversations about AI's future, ran every number through a calculator, and came out with one question answered: how fast can this actually go?

Written By
Grant Harvey
Grant Harvey
Mar 17, 2026
34 minute read

Three conversations happened in the span of a few weeks this year that, taken together, give you the complete picture of where AI is going, how fast it can get there, and what's physically stopping it from going faster.

Dario Amodei (CEO of Anthropic) explained how AI companies plan to make money. Dylan Patel (CEO of SemiAnalysis, the firm that tracks every chip order and data center on Earth) explained what physically limits how much AI we can build. And Jensen Huang (CEO of NVIDIA) showed what the factories of the future look like and how many tokens they'll produce.

Each one has a piece of the puzzle. None of them have all of it. So we put them together, ran every number through a calculator, and came out with a unified picture of AI's economic future.

Think of this as your cheat sheet for understanding the business of AI, from the trillion-dollar deals to the $2.5B optics company in Germany that's quietly holding up the entire industry.

First Up, The TL;DR

If you only have three minutes, read this:

  • How AI companies make money: They split their computing power roughly 50/50 between training new models and serving customers. The serving side (called "inference") has fat profit margins. So each model is technically profitable. The reason companies aren't profitable yet is because they keep spending 10x more to train the next model. It's like a restaurant that makes money on every plate but keeps building a bigger kitchen. Dario Amodei predicts this eventually stabilizes into a 3-4 player oligopoly (like cloud computing), with healthy margins, within a few years (you could argue we're basically there now, plus healthy open source competitors).
  • What limits growth: Not power. Not data center construction. Chips. Specifically, one Dutch company called ASML makes the only machine on Earth (an EUV lithography tool) that can print the circuits small enough (3 nanometers or smaller, to be precise) on advanced AI chips. They produce ~70-100 machines per year. Each gigawatt of AI data center capacity requires about 3.5 of these tools. Dylan Patel did the math: by 2030, the world can produce roughly 115-320 gigawatts of AI compute, depending on how fast the supply chain wakes up. That's the physical ceiling.
  • What happens inside those data centers: Jensen Huang showed that NVIDIA's new Vera Rubin architecture will produce 350x more tokens from the same amount of power in just two years. So even though the gigawatts grow "only" 50-60% per year (constrained by chips), the actual AI output grows far faster, though the rate is front-loaded: ~2,200% from 2026 to 2027 as Vera Rubin arrives, decelerating to ~124% by 2029-2030 as each new generation adds less relative improvement. FWIW, Jensen sees $1 trillion+ in confirmed chip demand through 2027.
  • What this means for jobs: We ran both revenue-based and token-based calculations. The physical compute infrastructure can support the displacement equivalent (as measured by tokens per equivalent salary) of just ~1,300 fully autonomous agent-equivalents in 2026, rising to ~1.2 million by 2028, and ~10-90 million by 2030 depending on the scenario (we estimate a theoretical maximum of ~3.6 quintillion tokens produced per year). So while we only have enough compute to power ~1.3K fully autonomous agents today, it rises to 1-9% of global knowledge workers by decade's end.
  • What does this all mean? The technology arrives faster than the supply chain can build it. The supply chain builds it faster than the economy can absorb it. So the actual pace of change is set by the slowest of these three gears.
Advertisement

Chapter 1: How AI Companies Make Money (The Dario Amodei Playbook)

Anthropic's revenue trajectory tells you everything you need to know about why every tech company on Earth is scrambling right now. Zero in 2022. $100M in 2023. $1B in 2024. $9-10B in 2025. And then, in January 2026 alone, they added another few billion.

That's a 10x growth rate, year after year.

But how does an AI company actually turn that revenue into profit? Dario broke it down on the Dwarkesh Patel podcast with a toy model that's surprisingly simple.

The 50/50 Split

Take a company spending $100B/year on compute. Roughly half goes to training new models. The other half supports inference (serving customers). If the inference side generates $150B in revenue with 50%+ gross margins, the company makes ~$50B in profit.

So why aren't AI companies profitable?

Because they're in an exponential scale-up phase. Last year's model cost $1B to train and generated $4B in revenue. Great. But this year, you're spending $10B to train the next model. Each model makes money; the company doesn't. Yet.

The Bankruptcy Math

Here's where it gets scary. Data centers take 1-2 years to build. You have to guess your 2027 revenue in 2025 and buy accordingly.

Dario walked through the nightmare scenario: if you assume $1 trillion in 2027 revenue and buy $1 trillion of compute, but revenue comes in at $800 billion instead? "There's no force on earth, there's no hedge on earth that could stop me from going bankrupt."

Being off by one year, or having growth be 5x instead of 10x, is catastrophic at these scales.

Why It'll Be 3-4 Players, Not a Monopoly

Dario predicts the AI industry will look like cloud computing at its maturation stage: 3-4 major players with healthy margins. Not one winner. Why? Extremely high costs of entry keep the field small. But models are more differentiated than cloud services ("everyone knows Claude is good at different things than GPT"), which supports even healthier margins.

The one wild card: if AI models can eventually build other AI models, the moat around everything could dissolve. But Dario says that's far past the "country of geniuses" milestone. It's a post-AGI question.

Advertisement

The Not-All-Tokens-Are-Equal Insight

One of Dario's sharpest observations: the value of tokens varies by orders of magnitude. Someone asks Claude to restart their Mac? That response is worth a few cents. But if a model tells a pharmaceutical company to move an aromatic ring (a structural element of a molecule) from one end of a drug to the other and it works? Those tokens could be worth tens of millions of dollars.

This is why "pay for results" pricing models are coming. And it's why the API business model is more durable than people think, because there will always be a thousand new startups experimenting with the latest capabilities in ways nobody predicted.

On Timelines

Dario was refreshingly specific:

  • 90% confident: "Country of geniuses in a data center" within 10 years. "It's crazy to say this won't happen by 2035."
  • 50/50 hunch: More like 1-3 years.
  • ~99% on coding: End-to-end software engineering in 1-2 years.
  • Revenue: "Hard to see" that there won't be trillions in AI revenue before 2030.

But even with those timelines, he emphasized "fast but not infinitely fast" diffusion. Legal review. Security compliance. Change management. Enterprise procurement. Each step is reasonable. Together, they add months or years between when the technology is ready and when it actually reshapes the economy.

Advertisement

Chapter 2: What Physically Limits How Fast We Can Build? (The Dylan Patel Reality Check)

Dario gives you the demand side. Dylan Patel gives you the supply side. And the supply side is where revenue dreams meet cold hard supply chains (also physics).

The Bottleneck Shifts Over Time

Most people fixate on one constraint at a time. What is it: Power? Data centers? Chips? But the real picture is a cascade that shifts year by year. According to Dylan's research:

  • 2026-2027: Memory fabs and cleanroom space are the binding constraint. Memory vendors (Micron, SK Hynix, Samsung) didn't build new fabs for 3-4 years because they were losing money in 2023. Fabs take 2 years to build. New capacity won't arrive until late 2027.
  • 2028-2030: The constraint shifts to EUV tools. ASML, a Dutch company, makes the only machine in the world that can print advanced chip circuits. They produce ~70-100 per year. Each gigawatt of NVIDIA's Rubin-generation data center needs about 3.5 of them.

Power is never the binding constraint. Dylan is explicit. Doubling energy cost adds $0.10/hour to a GPU's total cost. There are 16+ types of gas-powered generation alone, plus solar, wind, batteries, fuel cells. The US grid is terawatt-scale. This is a solved problem (so long as we deploy and scale more power at the right time, which Dylan is confident we'll do).

The $2.5B Company Holding Up the Entire Industry

The EUV tool supply chain is mind-boggling. Each machine costs $300-400M. It works by dropping tin droplets, hitting them three times with a laser to generate extreme ultraviolet light, then bouncing that light off 18 multilayer mirrors made by Carl Zeiss (sadly a private company atm), through a system where the reticle and wafer move in opposite directions at 9 Gs of acceleration while maintaining sub-nanometer alignment. If that makes your head just reading that, imagine trying to build & run it.

Carl Zeiss, the company making those mirrors? Market cap: $2.5B. Fewer than 1,000 specialists. All artisanal production; meaning you can't train replacements quickly.

Meanwhile, a single gigawatt data center built on top of those chips generates $50B in economic value, supported by ~$1.2B worth of EUV tooling. Dylan says that's a 42:1 leverage ratio. The entire AI revolution is balanced on a supply chain with 10,000+ suppliers that nobody outside the semiconductor industry has heard of.

Advertisement

The "AGI-Pilled Cascade" Problem

Dylan describes why the demand signal gets distorted on its way down the supply chain. The AI labs want X compute. NVIDIA, "not quite as AGI-pilled," builds X - 1. TSMC sees NVIDIA's order and builds X - 1. ASML sees TSMC's order and builds X ÷ 2. By the time the signal reaches the Zeiss lens-polisher in Germany, it's been attenuated by an order of magnitude. To them, what's the rush?

Then when everyone finally "gets it," the lead times are years.

Your iPhone Is Going to Cost $250 More

The memory crunch has a nasty downstream effect. DRAM prices have tripled. An iPhone's memory cost went from ~$50 to ~$150. Add NAND increases, and the total BOM increase is ~$150. Apple won't eat all that margin.

And it's worse at the low end. Global smartphone volumes are projected to fall from 1.4 billion to 500-600 million by 2027, driven almost entirely by mid-range and budget phones getting priced out. The AI industry is literally cannibalizing the consumer electronics industry for memory chips.

Dylan's assessment: "People are going to hate AI even more." This too could slow things down.

Chapter 3: The Token Factory (Jensen Huang's GTC 2026 Vision)

If Dario told us who's buying and Dylan told us who's building, Jensen just told us what they're building and how productive it'll be.

The Inference Inflection

Jensen's core claim: computing demand has increased by roughly 1 million times in the last two years. He breaks that into two factors: the work per task went up ~10,000x (from simple chat to reasoning to multi-step agentic workflows), and usage went up ~100x.

Three inflections drove this: generative AI (ChatGPT, late 2022), reasoning AI (o1/o3, 2024), and agentic AI (Claude Code, 2025). Each one made AI capable of doing more, which required more compute per task, which drove demand through the roof.

The critical insight: AI went from answering questions to doing work. As Jensen put it: "You don't ask AI what, where, when, how. You ask it create, do, build."

Advertisement

The Most Important Chart for Every CEO

Jensen presented what he called "the single most important chart for the future of AI factories", and for once, the hype is justified.

  • The Y-axis: Tokens per watt. This is throughput. How many tokens your factory produces per unit of power. Since every data center is physically power-constrained ("a one gigawatt factory will never become two"), this is your production capacity.
  • The X-axis: Token speed. How fast each individual response generates tokens. Faster speed means smarter AI (larger models, more context, more reasoning steps). Jensen says explicitly: "This axis is the same as smartness of the AI."

The trade-off: Throughput and speed are enemies. High throughput comes from batching many requests efficiently. High speed requires dedicating resources to each request. You can't maximize both.

This creates a tiered market, just like hotels or airlines. Jensen laid out five tiers from free ($0/M tokens, small models, high throughput) all the way to premium ($150/M tokens, frontier models, maximum speed). Each hardware generation doesn't just make tokens cheaper; it opens up entirely new tiers that weren't previously possible at any price.

350x in Two Years

The headline performance number Jensen cited: in a one-gigawatt factory, token generation will go from 2 million to 700 million tokens per second in two years. A 350x increase.

For context: 700 million tokens per second means ~60 trillion tokens per day from a single gigawatt factory. That's enough to support tens of millions of agentic sessions running simultaneously.

And switching from Blackwell to Vera Rubin produces 5x more revenue per gigawatt under Jensen's tiered pricing model. So the same power envelope, same data center, same land, same permits: 5x more money. That's why every company is racing to upgrade; at least, so goes the sales pitch.

$1 Trillion Through 2027

Last year, Jensen saw $500B in confirmed demand. This year, standing at GTC, he claimed, "at least $1 trillion through 2027." And he said he's "certain computing demand will be much higher than that."

The Anthropic vs OpenAI Compute Gamble

Jumping back to Dylan's interview for a second, this is where the compute gamble comes into play. Dylan framed it bluntly: Dario "screwed the pooch" compared to OpenAI on compute procurement.

OpenAI signed aggressive deals with everyone, including companies that had never built a data center in their lives. For months, credit markets panicked. Then OpenAI raised $110B* and suddenly everyone believed them. (*it's more like $15B right now, and $95B of commitments for more money in the future).

Anthropic, meanwhile, was principled. Conservative, even. And now they're paying 50% markups for last-minute compute through revenue shares and spot pricing. The most responsible kid in the class is now paying the highest prices.

Why an H100 NVIDIA Chip Is Worth MORE Today Than 3 Years Ago

This was maybe the most counterintuitive finding. The "bears" (the Michael Burrys of the GPU world) said GPUs should depreciate fast because new hardware keeps getting better. They predicted H100 prices would crash as Blackwell shipped.

What actually happened? H100 prices went UP. Labs are signing 2-3 year deals at $2.40/hr on hardware that costs $1.40/hr to deploy. Why?

Dylan says this is because the value of a GPU isn't determined by what you can buy today. It's determined by the value you can extract from it today. GPT-5.4 running on an H100 produces more tokens of a better model than GPT-4 ever did. The TAM for GPT-4 tokens was a few billion dollars. The TAM for GPT-5.4 tokens is probably north of $100B. Same GPU, much higher value.

Taking Annual Token Budgets Into Account

Jensen dropped one of the most quotable predictions of the year: every engineer at NVIDIA will soon get an annual token budget worth roughly half their base salary. Engineers making a few hundred thousand a year will get another $100-150K in tokens so they can be "amplified 10x."

"It is now one of the recruiting tools in Silicon Valley," he said. "How many tokens comes along with my job."

This is going to be very important for our analysis. You'll see what I mean in a minute.

Chapter 4: How Much AI Compute Will We Actually Have by 2030?

Now let's do an experiment. I want to put all three of these perspectives together and try to do the math and see what the maximum threshold of compute we can produce per year is, given these constraints (note: these are back of the napkin projections, i.e estimates, to at least attempt to try to understand the limits of supposed "artificial general intelligence" on the job market. I'm sure some of this will be wrong; if you work at SemiAnalysis or another firm tracking this, please email me with my email at the bottom with any corrections you see).

The Supply Ceiling (Dylan's Constraints)

We built a scenario model stacking every constraint Dylan described across four different scenarios and how much gigawatts of compute could be used each year. These are cumulative deployed compute numbers, meaning the total installed base of AI data center capacity running at any given point in time, accumulated across all years.

We'll break down how we got here in a minute. But here are how the scenarios could change as the binding bottleneck shifts from memory/fabs (2026-2027) to EUV tools (2028-2030):

  • Conservative (~52%/yr growth): 30 GW (2026) → 52 GW (2027) → 82 GW (2028) → 162 GW (2030)
  • Median (~61%/yr growth): 35 GW (2026) → 67 GW (2027) → 115 GW (2028) → 238 GW (2030)
  • Optimistic (~69%/yr growth): 39 GW (2026) → 79 GW (2027) → 144 GW (2028) → 319 GW (2030)
  • Catastrophic, i.e. Taiwan Crisis (~40%/yr growth, max, and potentially regionally limited): 30 GW (2026) → 38 GW (2027) → 50 GW (2028) → 115 GW (2030).

Each year's increment (from Dylan's scenarios) adds new hardware on top of what's already deployed; old hardware doesn't disappear, since GPUs typically have a 5+ year useful life (and per Dylan's assessment, older servers could become more valuable as compute remains constrained).

So when we say "238 GW in 2030," that means 238 GW of AI compute infrastructure operating simultaneously worldwide: 15 GW of pre-2026 Hopper-era hardware still running, plus 20 GW of Blackwell deployed in 2026, plus 32 GW of Vera Rubin from 2027, and so on.

This is different from annual production (how many new GW of chips ship in a given year, which is constrained by ASML to ~23-65 GW/year depending on the scenario) or annual capacity (how many GW the EUV tool fleet could theoretically support). The cumulative number is what matters for the job ceiling because all deployed hardware is producing tokens simultaneously. On that front...

Total tokens per year (quadrillions), by scenario:

  • Conservative: 1.4Q (2026) → 28Q (2027) → 250Q (2028) → 1,041Q (2029) → 2,325Q (2030)
  • Median: 1.7Q (2026) → 40Q (2027) → 391Q (2028) → 1,604Q (2029) → 3,591Q (2030)
  • Optimistic: 2.0Q (2026) → 50Q (2027) → 519Q (2028) → 2,178Q (2029) → 5,024Q (2030)
  • Taiwan Crisis: 1.4Q (2026) → 13Q (2027) → 102Q (2028) → 571Q (2029) → 1,642Q (2030)

(Q = quadrillion = 10^15 tokens)

The median scenario goes from 1.7 max quadrillion tokens in 2026 to 3,591 quadrillion in 2030. That's a ~2,100x increase in annual token production while the GW base only grows ~7x. The entire difference is the vintage effect: newer hardware producing orders of magnitude more tokens per watt. Now, let's break down how we got here before we break down what this means for jobs.

Step 1: How Much Compute Exists Today

Entering 2026, the world has roughly ~15 GW of AI data center capacity already deployed, mostly Hopper-era hardware. This is our starting base.

For reference on the major players: Anthropic and OpenAI each sit at roughly 2-2.5 GW today. Google, Meta, Microsoft, and Amazon hold the rest across a mix of internal AI workloads and cloud services. Dylan says roughly 20 GW of incremental capacity is coming online in the US alone this year.

Step 2: How Much New Compute Can We Produce Each Year

This is where ASML enters the picture. The global EUV tool fleet (currently ~275 tools, growing by 70-100/year) determines how many advanced AI chips can be manufactured. At 3.5 EUV tools per gigawatt and ~60% of the fleet allocated to AI, the annual production capacity is:

  • 2026: EUV fleet of 345 tools → ~61 GW/yr producible for AI
  • 2027: 420 tools → ~75 GW/yr
  • 2028: 500 tools → ~89 GW/yr
  • 2029: 590 tools → ~105 GW/yr
  • 2030: 690 tools → ~122 GW/yr

Now here's the catch: EUV capacity is NOT the binding constraint in every year. In 2026, the fleet could produce 61 GW of AI chips, but only ~20 GW actually gets deployed because there aren't enough memory fabs or cleanrooms to build the rest (memory vendors didn't build new fabs for 3-4 years while they were losing money). By 2029-2030, those bottlenecks have eased, but now the EUV fleet itself becomes the ceiling.

The actual new GW deployed each year, after ALL bottlenecks (memory, fabs, EUV, power, construction), gives us four scenarios:

  • Conservative: 15 GW (2026) → 22 (2027) → 30 (2028) → 38 (2029) → 42 (2030)
  • Median: 20 → 32 → 48 → 58 → 65
  • Optimistic: 24 → 40 → 65 → 80 → 95
  • Crisis: 15 → 8 → 12 → 25 → 40

In the median scenario, only 33% of EUV capacity is used in 2026 (memory is the bottleneck). By 2030, that rises to 53% (EUV starting to bite). In the optimistic scenario, 2030 deployment hits 78% of EUV capacity; it's clearly becoming the hard ceiling, exactly as Dylan predicts.

Step 3: Total Installed Capacity Per Year

Jensen sees $1T+ in committed demand through 2027. At ~$50B per GW of total CapEx, we can assume that's ~20 GW of committed orders, or at least forecasted ones (not financial advice; check NVIDIA's actual public reporting for how much is actually committed), not including unmet demand. And demand exceeds supply at every price point.

Jensen claims Nvidia can manufacture "multi-gigawatts of AI factories per month" worth of rack systems. But the total EUV fleet can produce roughly 61-122 GW/year of AI chips (at 60% AI allocation, depending on the year). The reason actual deployment is far lower; 20-65 GW/year in the median scenario; is that other bottlenecks prevent using all that chip capacity. In 2026-2027, memory fabs and cleanrooms are the limiting factor (memory vendors didn't build for 3-4 years). By 2029-2030, EUV tools themselves become the ceiling as demand catches up to the fleet's theoretical capacity. The assembly of racks is never the constraint; it's always something further down the supply chain.

And as we said above, each year's new hardware stacks on top of what's already running (GPUs have 5+ year useful life, and per Dylan, older servers may actually appreciate as compute stays constrained). So cumulative deployed AI compute:

  • Conservative (~52%/yr growth): 30 GW (2026) → 52 GW (2027) → 82 GW (2028) → 120 GW (2029) → 162 GW (2030)
  • Median (~61%/yr growth): 35 → 67 → 115 → 173 → 238 GW
  • Optimistic (~69%/yr growth): 39 → 79 → 144 → 224 → 319 GW
  • Taiwan Crisis (~40%/yr): 30 → 38 → 50 → 75 → 115 GW

The EUV hard ceiling: by 2030, the cumulative global EUV fleet of ~690 tools, at ~60% AI allocation and ~65% practical utilization, supports a theoretical ~270 GW of maximum compute. The optimistic scenario exceeds this because it assumes supplemental compute from older 7nm nodes, packaging innovations, and 3D DRAM (as Dylan suggested).

Step 4: Total Token Capacity Per Year

Here's where Jensen's numbers transform the picture. Dylan forecasts GW (the power envelope), while Jensen forecasts tokens-per-GW (the efficiency). These multiply:

Even though GW grows "only" 50-65% per year, total token output grows far faster because each new hardware generation produces dramatically more tokens per watt. But the growth rate isn't steady; it's violently front-loaded. In the median scenario, token output grows ~2,200% from 2026 to 2027 (as the first Vera Rubin deployments come online and produce 17x more tokens per GW than Blackwell), then ~870% from 2027 to 2028, then ~310% from 2028 to 2029, then "only" ~124% from 2029 to 2030 as the architectural leaps get smaller.

Jensen's "350x in two years" claim s a one-time jump from the Blackwell era to the Vera Rubin era. After that, each new generation adds 1.4-2.5x, not 10-17x. But of course, this is a roadmap, and the specs are likely to change. Even still, the biggest efficiency leap in AI history is happening right now, as we move towards 2027-2028. By 2030, the installed base has largely caught up to the new architectures, and growth starts tracking closer to GW growth again.

And this brings us to where the vintage mix transforms the picture. Not all gigawatts are created equal. A GW of 2026 Blackwell hardware produces ~2 million tokens/sec. A GW of 2030 Feynman hardware produces ~700 million. So even though total GW grows ~7x from 2026 to 2030, total token throughput grows ~2,000x.

So let's say we track each year's deployed hardware by its "vintage" (what generation chip it runs) and apply Jensen's performance data from his GTC slides: Hopper at ~1M tokens/sec/GW, Blackwell at ~2-12M (improving with software over its life), Vera Rubin at ~35-500M depending on maturity and Groq integration, and Rubin Ultra/Feynman at ~500-700M. Each vintage also improves over time as Nvidia pushes software updates (the Fireworks 7x example).

Total token production per year, in quadrillions (Q = 10^15 tokens):

  • Conservative: 1.4Q (2026) → 28Q (2027) → 250Q (2028) → 1,041Q (2029) → 2,325Q (2030)
  • Median: 1.7Q (2026) → 40Q (2027) → 391Q (2028) → 1,604Q (2029) → 3,591Q (2030)
  • Optimistic: 2.0Q (2026) → 50Q (2027) → 519Q (2028) → 2,178Q (2029) → 5,024Q (2030)
  • Taiwan Crisis: 1.4Q (2026) → 13Q (2027) → 102Q (2028) → 571Q (2029) → 1,642Q (2030)

The median scenario goes from 1.7 quadrillion tokens per year in 2026 to 3,591 quadrillion in 2030. That's a ~2,100x increase in annual token production, driven almost entirely by the vintage effect: by 2030, over 90% of all tokens come from hardware deployed in 2028-2030, even though older Hopper and Blackwell systems still account for a meaningful share of installed gigawatts.

This means the industry can serve exponentially more users, more agentic sessions, and more valuable work, even while the physical infrastructure grows at a more modest pace. The software and architecture eat what the hardware can't provide (and theoretically, more efficient AI model architectures could have an even bigger impact here).

Note: we could also break this out to how many tokens could be produced per second, like so:

  • Conservative: 0.05B tok/s (2026) → 0.9B (2027) → 7.9B (2028) → 73.7B tok/s (2030)
  • Median: 0.06B tok/s (2026) → 1.3B (2027) → 12.4B (2028) → 113.9B tok/s (2030)
  • Optimistic: 0.06B tok/s (2026) → 1.6B (2027) → 16.5B (2028) → 159.3B tok/s (2030)
  • Taiwan Crisis: 0.05B tok/s (2026) → 0.4B (2027) → 3.2B (2028) → 52.1B tok/s (2030)

In more intuitive units (tokens per day, median scenario): ~5 trillion tokens/day in 2026, ~111 trillion/day in 2027, ~1,070 trillion/day in 2028, and ~9,838 trillion tokens/day by 2030.

Chapter 5: How many jobs could be impacted by 2030?

Here's the question we really want to know though: how many jobs can AI actually replace, given physical constraints?

We ran the math two ways; once from pure token throughput (bottom-up), once from Jensen's revenue-per-GW slides (top-down). Both methods converge on the same range. Then we checked the results against Anthropic's own labor market research, which tracks how Claude is actually being used in the economy today, to see if the numbers make sense.

The Calculation Chain

There are four variables that multiply together to determine the maximum number of jobs AI can affect in a given year:

  1. How many gigawatts of compute exist (from Dylan's supply chain forecast).
  2. How many tokens per second each gigawatt produces (from Jensen's architecture specs, tracked by hardware vintage since old Blackwell-era GW produces far fewer tokens than new Vera Rubin GW).
  3. What fraction of those tokens go to autonomous work (vs. training, free-tier chatbots, search, recommendations, research).
  4. How many tokens it takes to fully replace one knowledge worker for a year.

Here is Our Assessment of the Maximum Job Impact Per Year

Now we convert tokens to agent-equivalents. The simplest version is to take annual tokens produced as a proxy for salary, where each agent gets an annual token budget per year (as Jensen said, we can roughly expect each engineer getting $100K-$150K in annual token budget to manager their agents). But we need to make three cuts to narrow the actual total token pool down to "tokens doing autonomous work that replaces a human":

  1. Cut 1: Training. Roughly 50% of all compute goes to training new models, not serving customers (Dario's 50/50 framework). Cut the token numbers in half.
  2. Cut 2: Work-replacement share. Of the remaining inference tokens, most go to consumer chatbots, search, recommendations, free-tier usage, and human augmentation (making people faster, not replacing them). Anthropic's labor market research shows that even the most AI-penetrated occupations are mostly seeing augmentative, not automated, use (so far). Let's assume the share going to fully autonomous work replacement grows over time: ~3% in 2026 (agentic AI still early), ~7% in 2027, ~12% in 2028, ~18% in 2029, ~25% in 2030.
  3. Cut 3: Tokens per replaced worker. Jensen says $100-150K/year in tokens makes an engineer 10x more productive. Full autonomous replacement (no human in the loop) costs more because agents need extra reasoning, retries, and verification. Our mid estimate: 20 billion tokens per fully replaced knowledge worker per year (~$200K at $10/M tokens; equivalent to a mid software salary).

After all three cuts, here are the MAXIMUM agent-equivalent tokens produced per year that the physical infrastructure can support:

  • Conservative: ~1,100 (2026) → ~50K (2027) → ~750K (2028) → ~4.7M (2029) → ~14.5M agents (2030)
  • Median: ~1,300 (2026) → ~71K (2027) → ~1.2M (2028) → ~7.2M (2029) → ~22.4M agents (2030)
  • Optimistic: ~1,500 (2026) → ~88K (2027) → ~1.6M (2028) → ~9.8M (2029) → ~31.4M agents (2030)
  • Crisis (Max): ~1,100 (2026) → ~22K (2027) → ~310K (2028) → ~2.6M (2029) → ~10.3M agents (2030)

As a share of the world's ~1 billion knowledge workers: the median scenario for agent adoption reaches only 2.2% by 2030; with the current architecture, that is. Even the optimistic scenario tops out at 3.1%. At the most aggressive token-efficiency assumption (5B tokens/worker instead of 20B), the median 2030 ceiling rises to ~90M (9.0%).

According to this model, that's the absolute maximum impact to the job market, assuming everything goes right for agentic adoption. Keep in mind, a ~9% impact to the job market would still be MASSIVE. But it wouldn't be every job, and it doesn't account for new roles opening up to fill those old ones. So no need to cancel all your plans for a future where AI does all the jobs for us, or burn yourself out to avoid joining the permanent underclass. We've got plenty of time to keep working.

Cross-Checking Against Revenue

We can sanity-check these numbers by converting them into dollars. Specifically: how much total revenue would the entire AI industry (Anthropic, OpenAI, Google, Meta, and every other company selling access to AI models) earn from customers paying for tokens, API calls, and AI-powered services?

Jensen's GTC slides show Blackwell generates $30B/GW/year in revenue, Vera Rubin $150B, and VR+Groq achieving $300B in revenue. But Dylan's data shows Anthropic currently realizes ~$8-10B/GW; roughly 30% of Jensen's ceiling. Applying that realization rate across the vintage-weighted model...

  • The median 2030 scenario produces ~$1.8T in realistic AI inference revenue.
  • If 40% of that displaces human wages at ~$65K average, that implies ~11M displaced workers; consistent with the token model's maximum 22.4M agents supported per year within the expected uncertainty range.
  • The truth is likely 10-22M by 2030, or roughly 1-2% of global knowledge workers could be completely replaced with agent equivalents.

Here's the revenue math stepped out a bit: Take total GW, cut it in half (Dario's 50% goes to training new models, 50% to serving customers). Multiply the inference half by Dylan's $10-13B/GW in annual compute rental costs. Then divide by the cost ratio (Anthropic's margins are sub-50%, meaning roughly $0.55 of every dollar customers pay goes to compute; the rest is gross profit). That gives us total AI inference revenue; the combined money that all AI companies worldwide earn from selling token access, whether through APIs, chat subscriptions like Claude Pro or ChatGPT Plus, enterprise contracts, or embedded AI services:

  • Conservative: 162 GW → ~$1.8 trillion in combined AI industry token/service revenue by 2030.
  • Median: 238 GW → ~$2.6 trillion.
  • Optimistic: 319 GW → ~$3.5 trillion.

These are industry-wide totals (all AI companies combined, globally), not per-company. For context, global cloud computing revenue today is around $600B. So the median scenario has AI token revenue alone reaching 4x the size of today's entire cloud industry by decade's end.

Dario said it's "hard to see" that there won't be trillions before 2030. Even our conservative scenario produces $1.8T. The numbers check out.

Now, if we want to step out how we got here even more thoroughly, we can...

Input 1: Tokens per GW (The Vintage Problem)

Not all gigawatts are equal. A GW deployed in 2026 on Blackwell doesn't magically become Vera Rubin in 2028. So we track each year's hardware "vintage" separately, using Jensen's actual performance data from his GTC slides:

  • Pre-2026 vintage (Hopper era): ~1-2M tokens/sec/GW. Jensen's throughput chart shows Hopper at ~0.15M TPS/MW (tokens per second per megawatt), which translates to ~150K TPS/GW at the free tier, dropping to near zero at the premium tier. With software updates, it climbs modestly over time.
  • 2026 vintage (Blackwell NVL72): Jensen's chart shows ~0.8M TPS/MW at the free/medium tiers. That's ~800K TPS/GW at those tiers. But Jensen's stated starting point for a blended factory is ~2M tokens/sec/GW. With software optimization (the Fireworks 7x example), this matures to ~10-12M by 2028-2029.
  • 2027 vintage: This is a transition year. Not all 2027 deployments will be Vera Rubin; some will be late Blackwell. We model it as 40% Blackwell / 60% early Vera Rubin. Jensen's chart shows Vera Rubin at ~1.6M TPS/MW at the free tier (2x Blackwell), but this is for a mature deployment. Year-one Vera Rubin starts lower as software catches up to the hardware. Blended estimate for 2027 vintage: starts at ~35M tokens/sec/GW, matures to ~200M by 2030.
  • 2028 vintage (Vera Rubin + Groq): Jensen's slide shows VR+LPX maintaining ~0.5M TPS/MW all the way out to 400+ TPS/User, which is 35x Hopper at the premium tier. The Groq integration primarily adds value at the high-speed tiers. Blended: starts at ~200M, matures to ~500M.
  • 2029-2030 vintage (Rubin Ultra / Feynman): These architectures aren't on Jensen's slides yet, but the roadmap shows NVLink 144 (Kyber rack), LP35 with NVFP4, and eventually Feynman with LP40. Extrapolating Jensen's generational improvements: 500-700M tokens/sec/GW.

Jensen's throughput chart reveals something critical: by 2030, over 90% of all token throughput will come from hardware deployed in 2028-2030. The pre-2026 Hopper fleet and 2026 Blackwell fleet still running will contribute less than 1% of total tokens, even though they represent a meaningful chunk of deployed gigawatts. When the new chips arrive matters far more than how many total gigawatts exist.

Input 2: What Share Goes to Autonomous Work?

This is the hardest variable to estimate, and Anthropic's labor market research gives us the best available ground truth.

Dario says roughly 50% of compute goes to training. That's our first cut: half of all tokens never reach a customer.

Of the remaining inference tokens, most still go to things that aren't "replacing a worker": consumer chatbots on the free tier, search augmentation, recommendation engines, internal hyperscaler workloads, research and experimentation, and human augmentation (making people faster without eliminating their role).

Anthropic's research paper, "Labor Market Impacts of AI" (March 2026), helps calibrate this. They built a measure called "observed exposure" that tracks what tasks Claude is actually doing in professional settings, weighted toward automated (rather than augmentative) use cases. Their key findings:

  • Even in Computer & Math occupations (the most theoretically exposed category at 94%), Claude currently covers just 33% of tasks in real-world professional usage
  • The most exposed single occupation is Computer Programmers at 75% task coverage
  • 30% of all workers have literally zero AI task coverage (cooks, mechanics, bartenders, lifeguards, etc.)
  • 97% of observed Claude usage falls on tasks rated as theoretically feasible; the models can do these things. The gap is adoption, not capability.
  • Critically: they find no systematic increase in unemployment for highly exposed workers since late 2022. The only detectable signal is a ~14% slowdown in hiring of workers aged 22-25 into exposed occupations.

What does this tell us about our "work replacement share" assumption? Even in the most AI-penetrated occupation (programming), a quarter of tasks aren't covered. And Anthropic's measure weights automated use more heavily than augmentation; most of that 33-75% coverage is people using AI as a tool, not AI doing the work autonomously.

Therefore, our estimates for the share of inference tokens going to fully autonomous work replacement are as follows:

  • 2026: ~3% (agentic AI still early; Anthropic's data shows mostly augmentative use)
  • 2027: ~7% (Claude Code and coding agents scaling; but most usage still augmentative)
  • 2028: ~12% (broader enterprise agent adoption; more API-driven automated workflows)
  • 2029: ~18% (agents handling full workflows in the most exposed occupations)
  • 2030: ~25% (significant share is autonomous work, but still well below theoretical maximum)

This reflects what Anthropic's own data shows about how slowly automated use is growing relative to augmentative use.

Input 3: Tokens Per Replaced Worker

Jensen says giving an engineer ~$100-150K/year in tokens makes them 10x more productive. That's the augmentation budget. Full autonomous replacement requires more tokens because agents need extra reasoning steps, retries, verification loops, and can't take the human shortcuts that make augmentation efficient.

We model three estimates:

  • Low (5B tokens/worker/year): Efficient models on routine, highly structured tasks like data entry, basic customer service, form processing. At ~$10/M tokens, that's ~$50K; cheaper than hiring a human for these roles.
  • Mid (20B tokens/worker/year): Blended across task types including some complex reasoning. At $10/M tokens, that's ~$200K. Cross-checks against Jensen's augmentation budget at ~2-3x the cost (you need more tokens for full autonomy than for augmentation).
  • High (50B tokens/worker/year): Complex engineering and research work with heavy reasoning chains, long context, multi-step agentic workflows. At premium pricing ($45/M), that's ~$2.25M per agent per year; expensive, but potentially worth it if the output is high-value enough.

The Results: Median Scenario

Here's the full calculation for Dylan's median GW forecast. Each year, we sum the token throughput across all hardware vintages (each producing at its generation-specific rate), apply the 50% training cut, apply the work-replacement share, then divide by 20B tokens per worker (mid estimate):

2026 (35 GW deployed, ~55M total tokens/sec):

  • 2026 Blackwell vintage (20 GW at 2M tok/s/GW): 40M tok/s (73%)
  • Pre-2026 Hopper vintage (15 GW at 1M): 15M tok/s (27%)
  • After 50% training cut → 27.5M tok/s inference
  • After 3% work share → 825K tok/s for autonomous work
  • ~1,300 agent-equivalents (near zero; barely a rounding error on the labor force)

2027 (67 GW, ~1.3B total tokens/sec):

  • 2027 vintage (32 GW, 60/40 VR/Blackwell blend at ~35M tok/s/GW): 1.12B tok/s (87%)
  • 2026 Blackwell (20 GW at 7M, software-improved): 140M tok/s (11%)
  • Pre-2026 Hopper (15 GW at 1.5M): 22.5M tok/s (2%)
  • After training cut + 7% work share → ~71,000 agent-equivalents

2028 (115 GW, ~12.4B total tokens/sec):

  • 2028 VR+Groq vintage (48 GW at 200M): 9.6B tok/s (78%)
  • 2027 vintage (32 GW at 80M, maturing): 2.56B tok/s (21%)
  • 2026 Blackwell (20 GW at 10M): 200M tok/s (2%)
  • After training cut + 12% work share → ~1.2 million agent-equivalents

2029 (173 GW, ~50.9B total tokens/sec):

  • 2029 Rubin Ultra vintage (58 GW at 500M): 29.0B tok/s (57%)
  • 2028 VR+Groq (48 GW at 350M, maturing): 16.8B tok/s (33%)
  • 2027 vintage (32 GW at 150M): 4.8B tok/s (9%)
  • After training cut + 18% work share → ~7.2 million agent-equivalents

2030 (238 GW, ~113.9B total tokens/sec):

  • 2030 Feynman vintage (65 GW at 700M): 45.5B tok/s (40%)
  • 2029 Rubin Ultra (58 GW at 650M): 37.7B tok/s (33%)
  • 2028 VR+Groq (48 GW at 500M): 24.0B tok/s (21%)
  • 2027 vintage (32 GW at 200M): 6.4B tok/s (6%)
  • After training cut + 25% work share → ~22.4 million agent-equivalents

All Four Scenarios (Mid Estimate: 20B tokens/worker/year)

  • Conservative: ~1,100 (2026) → ~50K (2027) → ~750K (2028) → ~4.7M (2029) → ~14.5M (2030, 1.5% of knowledge workers).
  • Median: ~1,300 (2026) → ~71K (2027) → ~1.2M (2028) → ~7.2M (2029) → ~22.4M (2030, 2.2%).
  • Optimistic: ~1,500 (2026) → ~88K (2027) → ~1.6M (2028) → ~9.8M (2029) → ~31.4M (2030, 3.1%).
  • Crisis: ~1,100 (2026) → ~22K (2027) → ~305K (2028) → ~2.6M (2029) → ~10.3M (2030, 1.0%).

Sensitivity: What If Agents Are More or Less Token-Hungry?

The tokens-per-worker assumption matters a lot. Here's the median scenario at all three estimates for 2030:

  • At 5B tokens/worker/year (efficient agents, routine tasks): ~90M agents (9.0% of knowledge workers)
  • At 20B tokens/worker/year (mid, blended tasks): ~22.4M agents (2.2%)
  • At 50B tokens/worker/year (complex engineering/research): ~9.0M agents (0.9%)

The range spans from ~9M to ~90M depending on how token-efficient autonomous agents become. If models get dramatically more efficient (fewer tokens per task through better reasoning), the upper end comes into play. If agentic workflows remain token-hungry (lots of retries, long chains of thought, verification loops), we stay near the lower end.

Cross-Check: Jensen's Revenue-Per-GW

Jensen's GTC slides show the revenue ceiling per GW by architecture generation (reading directly from his bar charts):

  • Blackwell: $30B/GW/year (assumes 25% power per tier, full utilization)
  • Vera Rubin: $150B/GW/year
  • VR + LPX (Groq): $300B/GW/year

But Dylan's data gives us a reality check on Jensen's ceilings. Anthropic currently runs ~2-2.5 GW at ~$20B ARR. That's roughly $8-10B of actual revenue per GW; about 30% of Jensen's $30B Blackwell ceiling. The gap is competition, utilization, pricing pressure, and the fact that not every GW runs at full capacity on the optimal tier mix.

We can use these ceilings as a second model by tracking revenue by hardware vintage the same way we track tokens. In the median 2030 scenario: 15 GW of Hopper at $10B/GW, 20 GW of Blackwell at $30B, 32 GW of Vera Rubin at $150B, 48 GW of VR+Groq at $300B, 58 GW of Rubin Ultra at $350B, and 65 GW of Feynman at $400B. After the 50% training cut, that gives a total revenue ceiling of ~$33T.

Obviously nobody thinks AI revenue will actually hit $33T by 2030 (due to adoption and supply constraints). That ceiling assumes Jensen's projections for architectures that don't exist yet (Rubin Ultra, Feynman) are correct, and that every GW runs at full utilization on the optimal tier mix. Reality will be much lower. Applying the same ~30% realization we observe today on Blackwell gives ~$9.9T ceiling → ~$3.0T realistic. But the 30% rate was calibrated on Blackwell; if future architectures can't fill the premium/ultra tiers ($45-150/M tokens), realization could be lower, maybe 10-20%. At 10% realization, it's ~$3.3T ceiling → ~$1.0T realistic.

So the realistic revenue range is roughly $1-3T depending on how literally you take Jensen's future pricing tiers. If 40% of that revenue displaces human wages and the average displaced knowledge worker earns ~$65K globally, that implies roughly ~6-18M displaced-equivalent workers.

That brackets our token model's mid estimate of 22.4M, with the revenue model tending lower because it's anchored to what companies actually earn today rather than projecting from hardware specs. The two models converge on a likely range of 10-22M fully autonomous agent-equivalents by 2030, or a maximum threshold (pending no new architecture changes, which is unlikely) at roughly 1-2% of the global knowledge workforce.

What Anthropic's Labor Data Tells Us Is Actually Happening

The token math gives us the ceiling. Anthropic's labor market research gives us the floor; what's actually happening right now.

Their March 2026 paper studied unemployment trends using their "observed exposure" measure (which combines theoretical AI capability with real-world Claude usage data, weighted toward automated professional use). The findings are striking:

  • No detectable unemployment increase. Workers in the top quartile of AI exposure have shown no statistically significant increase in unemployment compared to unexposed workers since ChatGPT launched. The confidence interval is tight enough that they could detect a 1 percentage point differential increase. It hasn't happened.
  • The only signal: young worker hiring has slowed. Workers aged 22-25 entering exposed occupations saw a ~14% drop in job-finding rates compared to 2022. This echoes findings from other researchers (Brynjolfsson et al.). But it's subtle, barely statistically significant, and could reflect young workers staying in school longer, taking different jobs, or other factors unrelated to AI.
  • The exposure gap is enormous. The most theoretically exposed category (Computer & Math) is 94% theoretically feasible for AI but only 33% actually covered in real professional usage. AI can do far more than it currently is. The red area (what AI is doing) is a fraction of the blue area (what AI could do).
  • Who's most exposed isn't who you'd think. The top quartile of exposed workers are 16 percentage points more likely to be female, earn 47% more on average, and are almost 4x more likely to have a graduate degree than the unexposed group. AI displacement risk is concentrated in well-paid, well-educated, white-collar professions; not the hourly workers people typically worry about.

This is exactly the "diffusion lag" Dario described: the technology is ready for many more tasks than the economy has adopted it for. The theoretical ceiling is 5-10x higher than actual usage. And even in occupations where AI covers 75% of tasks (computer programming), nobody's getting fired at statistical scale. Yet.

Why This Is a Hockey Stick, Not a Line

The progression from ~1,300 full agent-equivalents in 2026 to ~21.7 million in 2030 is a hockey stick because three multipliers compound simultaneously:

  • GW grows ~60%/year (constrained by ASML)
  • Tokens per GW grow ~5-10x per hardware generation (Jensen's architecture improvements; each new vintage produces orders of magnitude more tokens than the last)
  • Work-replacement share of inference grows from 3% to 25% (as agentic adoption matures and automated use cases expand beyond early adopters)

Multiply those three together and you get a ~17,000x increase in agent-equivalents from 2026 to 2030. A curve that barely registers in 2026-2027, then explodes in 2028-2030 as Vera Rubin deployments come online and the agentic share of inference ramps.

Anthropic's coverage data validates this shape. The "observed exposure" measure shows AI is still early in its penetration of even the most theoretically exposed occupations. But the blue area (what's possible) is vast, and the red area (what's happening) is growing every quarter as adoption deepens and models improve. The gap between blue and red is where the next few years of disruption lives.

What This Means

  • In 2026-2027, the physical compute infrastructure can support somewhere between 1,100 and 88,000 fully autonomous agent-equivalents. That's a rounding error on the global labor force. The constraint is absolute: there simply aren't enough high-throughput tokens being produced, because Blackwell-era hardware makes relatively few tokens per watt. This is consistent with Anthropic's finding of no detectable unemployment signal. We can assume that achievements in model architecture could increase this at least somewhat.
  • The real inflection comes in 2028-2029, when Vera Rubin and Groq deployments hit scale. That's when token throughput jumps by an order of magnitude, the ceiling on autonomous work rises to millions, and the agentic share of inference breaks past 10%.
  • By 2030, the compute ceiling allows for 10-22 million agent-equivalents at mid-range estimates (converging the token model and revenue cross-check), roughly 1-2% of global knowledge workers. Even at the most optimistic token-efficiency assumptions, it tops out around 90-126 million (9-13% of the global workforce).
  • But Anthropic's 30% zero-exposure workforce is untouchable. About 30% of workers have tasks that AI simply doesn't cover; physical labor, manual trades, in-person services. No amount of additional compute changes this. The ceiling applies only to the ~70% of the workforce whose tasks have any AI overlap, and within that group, the most exposed occupations (programming, customer service, data entry, financial analysis) will feel it first; arguably, we're not even yet at full automation and they already are.
  • Most importantly: Mass displacement from AI is physically impossible through at least 2028. There aren't enough tokens, and Anthropic's real-world data confirms it: no unemployment signal detected yet, even in the most exposed occupations. After 2028, the ceiling lifts fast, but Dario's diffusion lag becomes the new governor. The supply chain builds it faster than the economy can absorb it.

The Three Gears

The pace of AI's economic impact is set by three gears, each turning at a different speed:

  • Technology (model capability) — Fastest gear. Algorithmic gains of 10x/year; multiple breakthroughs per year.
  • Infrastructure (compute supply) — Medium gear. 50-65% GW growth/yr; capped by ASML, TSMC, memory fabs.
  • Adoption (economic diffusion) — Slowest gear. Legal, compliance, change management, trust, organizational inertia.

The actual pace of change is always set by the slowest gear. Today (2026), that's infrastructure. By 2028-2029, it shifts to adoption. Jensen gives us the technology. Dylan gives us the infrastructure ceiling. Dario gives us the adoption lag. Together, they tell us: the technology arrives faster than the factories can build it, and the factories build it faster than the economy can absorb it.

The real question is how fast a company in the Netherlands can teach artisans to polish mirrors, and how fast a Fortune 500 can get through its vendor procurement process. Which, when you think about it, is the most human constraint of all.

Sources: Dario Amodei on Dwarkesh Patel (Feb 2026). Dylan Patel on Dwarkesh Patel (Mar 2026). Jensen Huang GTC 2026 Keynote (Mar 2026). Anthropic: Labor Market Impacts of AI (Mar 2026). Full timecoded insights from all three conversations, supply chain analysis, and Python forecast model available in our companion pieces.

As mentioned, if you catch any errors here, please email me at grant@theneurondaily.com and I'll gladly update the piece. This is an important metric to track and I want to get it as close to right as possible, even though this is just an experiment just for educational purposes.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.