Is Gemini 3 Flash the "intelligence too cheap to meter" moment for AI?

Google just shipped a model that’s priced like a commodity, scores like a frontier system, and is being pushed into products with billions of touchpoints.

Written By
Grant Harvey
Grant Harvey
Dec 18, 2025
5 minute read

Google launched Gemini 3 Flash yesterday and is rolling it into the Gemini app and AI Mode in Search as the default "fast" brain. The thing costs $0.50 per 1M input tokens and $3 per 1M output tokens—and it performs better than Gemini 2.5 Pro.

This is the trend with Gemini Flash models: they typically surpass the last generation's Pro model. But this time, the gap feels different.

The specs: Gemini 3 Flash posted 78% on SWE-bench Verified (real-world software tasks), 90.4% on GPQA Diamond (hard science questions), and 33.7% on Humanity's Last Exam (a really hard exam, lol). It also uses 30% fewer tokens on average than 2.5 Pro on typical traffic. Google is bundling context caching that can cut costs by 90% on repeated-token workloads, plus a Batch API for additional savings.

The positioning: Google is explicitly framing this as "agentic / vibe-coding" infrastructure for iterative work. JetBrains, Bridgewater, and Figma are among the early adopters Google name-checks, and those are some pretty big names.

The bigger play: But the Flash story isn't just "a faster model." Google is also simultaneously building TorchTPU to make its TPU chips run the AI framework PyTorch more seamlessly (working with Meta, where PyTorch was founded) to reduce the switching costs that keep developers locked into rival chipmaker NVIDIA's CUDA software ecosystem. In other words, Google is trying to line up the entire stack (model, tooling, chips) so cheap intelligence makes it cheaply into production.This

This is essentially Google targeting NVIDIA’s moat by making TPUs friendlier to PyTorch, which is basically a shot at NVIDIA’s default-dev software, CUDA. Although, as we covered recently, some like investor Gavin Baker think Google's current TPU advantage could be a temporary situation.

This is a popular idea. X-famous dev Pieter Levels pretty publicly said he bought over $1M of Google stock, calling the move a personal reversal from "Google hater" to "Google is dominating in real usage." His thesis is simple: the endgame advantage in AI looks like chips + distribution + multimodal data, and that points to a small set of players (to him, they are Google, Elon, and China).

So how's it playing in the wild? We dug through X and Reddit to find out. The verdict: mostly positive, with users calling it "fantastic" and a "big step up."

On X, early testers describe it as "dangerously close to Pro levels" for coding, with one user noting it handles complex tasks without the "AI slop" you usually get from flash models. Jeff Dean and other Google DeepMind leads have been promoting it, and one team deployed it in production within hours for faster navigation and JSON handling. Benchmarks show 2x faster and 4x cheaper performance than predecessors on browser agents.

Advertisement

Reddit's reception is more mixed but leans optimistic. The GitHub Copilot integration has sparked debates about cost efficiency—total tokens generated matters more than per-token pricing. And another thread we were following had these concerns: 

  • The quantization problem: One user noted that models perform well at launch but get quantized "to like 4 bits" weeks later, becoming "worse than the previous version." Their example: "Gemini 2.5 Pro back in June was way better than Gemini 3 Pro for the past few weeks." The concern is that "these benchmarks are meaningless when you can't rely on the models actually delivering these results consistently."
  • Rate limit design flaw: Multiple users discovered that Flash Thinking shares usage limits with Gemini Pro, which creates a bizarre incentive structure. As one user put it: "There's no reason to use the thinking model if you have the AI pro plan since it's shared with the pro model for some reason." Another added: "You're not incentivized to use this model when it counts against your limit for 3 pro as well."
  • The pricing trade-off: One user made an interesting economic argument: "I'd rather have a model 50% of the price of Pro that I have confidence in, even for coding, than a model 30% of the cost that is too dumb for general use." In other words, Flash 3 being more expensive than old Flash might actually be "the right move" if it's genuinely capable.
  • The reliability problem: A user described Gemini 2.5 Pro as "Russian roulette—you never knew what you'd get. Sometimes I got amazing answers and sometimes I got extremely dumb answers." They claim 3.0 is "much more reliable" but still has this problem, just less frequently.
  • Benchmark skepticism: One highly upvoted comment captured the general mood: "I'm getting tired of these benchmarks" and "Lots of posts about benchmarks like anyone knows what TF it really means." The suggestion was that instead of charts, people should show "genuine output that's post worthy" with context about who scored it and who paid for the test. We noticed this same feeling around the Image edit benchmarks. Nobody trusts them.

What people are building: The demos are wild. Someone emulated macOS in HTML, while another created an animated 3D procedural rooms from prompts. We'll gather more as time goes on; this is all very new.

Why this matters: Flash at $0.50 per million tokens makes "AI everywhere” economically viable for the first time. That's the context for this new Vanguard data: AI-exposed wages rose 3.8% over the past two years (vs 0.7% elsewhere) because companies are paying more for workers who can work alongside cheap AI, not paying less because AI replaced them. Meanwhile, open source is closing the gap and getting cheaper to run (6x cheaper, according to some reports), which means Google's bet on "speed-per-dollar" is also a bet that the developers who learn to work with Flash-priced 

Counterpoint: The labor snapshot is early, and impacts can lag. Also, cheaper intelligence doesn't only multiply productivity: it multiplies everything. The New York just wrote that A.I.-generated content passed a kind of "audiovisual Turing test" this year, accelerating the slop problem: spam, deepfakes, and synthetic sludge that looks real enough to travel fast.

Advertisement

Bottom line: Gemini 3 Flash is rolling out globally now, which means "fast frontier" intelligence is about to become ambient: cheap enough to be always-on, everywhere. Google is betting on speed to make AI feel like a utility. 

Now here's the tension: if AI becomes too cheap, companies will be tempted to replace entry-level workers instead of training them. Recently, AWS CEO Matt Garman heard business leaders say they could "replace all of our junior people" with AI, and called it "the dumbest thing I've ever heard." His reasoning? "How's that going to work when ten years in the future you have no one that has learned anything?" In other words: intelligence too cheap to meter only works if you still have humans who, y’know, know how to use the meter.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.