😺DeepSeek’s new 28-cent agent model

Written By
Grant Harvey
Grant Harvey
Aug 2, 2026
10 minute read

In partnership with

Welcome, humans.

After years of being put to the test on how well it plays Pokémon (via the aptly called ClaudePlaysPokemon), Claude Opus 5 has apparently reached the “fine, I’ll make my own version” stage.

One demo reportedly ran for about 12 hours on Ultracode using a multi-agent loop (the prompt for this is down below), producing a playable monster-catching game with a 3D world, battles, and characters.

The result looks wildly impressive, right up until you meet Charmander Barney and Bulldog Bulbasaur. Maybe this is copyright protection. Maybe Opus 5-as-a-character-designer graduated from the Elsagate school of bootleg 3D autoplay slop. We may never know…

Either way, AI game demos are graduating from tiny browser toys into projects that can hold together for hours.

Here’s what happened in AI today:

  • 😻 DeepSeek upgraded V4-Flash while keeping API prices near pennies.

  • 📰 Big Tech’s AI buildout passed $1.1T and squeezed cash flow.

  • 📰 Europe’s AI labels and watermarks become enforceable August 2.

  • 🍪 Cleanlist turns plain English into verified prospect lists.

  • 🌟 The week’s five biggest stories and tools, ranked.

😺 DeepSeek V4-Flash brings frontier agent work to bargain pricing

The most important AI upgrade this weekend is not a new chatbot trick. It is the price tag for intelligence, which just reached a new low thanks to the spicy little troublemakers over at DeepSeek (y’know, the Chinese AI lab that shocked the US stock market back in 2025?)

Well, DeepSeek just upgraded V4-Flash into a far stronger coding and agent model while keeping the same architecture and bargain API rates. The result is a model that can do serious multi-step work for a fraction of what frontier labs usually charge (we think rumors of this coming is probably why OpenAI released that price drop on Thursday).

Here's what happened:

  • DeepSeek re-trained the existing V4-Flash rather than making it larger.

  • The model activates about 13B of its 284B parameters (the numbers that inform the model’s intelligence) for each request, which keeps running costs low.

  • It scored 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE, two tests of coding-agent performance (this means its good at code).

  • Artificial Analysis scored it 50 on its Intelligence Index, up 10 points from the previous Flash model (really wild for its price and size).

  • Pricing stayed at $0.14 per million input tokens and $0.28 per million output tokens. Cached input costs $0.0028.

For perspective, a million output tokens from V4-Flash cost 28 cents. That makes large classification jobs, coding loops, and browser-agent retries plausible without turning every extra attempt into a budget meeting.

How to try it:

Use deepseek-v4-flash through DeepSeek’s API, run it yourself if you have the servers to do so, or wait until the major US cloud providers support it (refresh this page every day or so for your options). It now supports the Responses API and is adapted for Codex-style coding workflows.

Why this matters: Most people do not need the absolute smartest model for every task. They need one that can reliably research, code, classify, or operate tools without turning each workflow into a luxury purchase. Key word: RELIABLY.

If V4-Flash holds up outside benchmarks, companies can reserve premium models for the hardest judgment calls and route routine agent work to something dramatically cheaper, and if you own your own servers or rent your own on the cloud, on computers you actually control. A workflow that felt too expensive at scale via OpenAI or Anthropic can suddenly make sense.

It also pressures OpenAI, Anthropic, and every provider charging a premium for capabilities that cheaper competitors are quickly matching. This is good for the industry; we need efficiencies of scale so everyone can actually (affordably) use this stuff. Otherwise whats the point?

Our take: “Opus-class” is still a benchmark claim, not a universal truth. DeepSeek used a specific harness and maximum effort for its agent tests, and real projects expose failures that leaderboards miss. But the pricing threat is real even if the model is merely good enough.

The next model war will not be won by one leaderboard. It will be won when buyers ask why a routine task still needs the expensive option. For me, that’s because the cheaper models still hallucinate and do dumb stuff, so its not worth it to go faster and cheaper. But when faster and cheaper = current max intelligence reliably, we off and poppin’.

FROM OUR PARTNERS

The AI coworker built for teams in Slack

The era of solo player AI is over. Adapt is the integrated coworker that empowers every team to work AI native together.

Anyone can tag @Adapt in Slack to answer a quick question, schedule a task, prepare for a board meeting, or build the dashboard you’ve been waiting on for weeks.

Built for business users, powerful enough for engineers.

Free credits awarded when you use your work email

🎓 Make AI Run a Gauntlet Against Real-World Work

Telling an AI to “make this better” gives it no finish line. The Gauntlet Loop replaces vague improvement with a brutal test: can the result beat a real example?

Matt Shumer used the approach while building Claude of Duty, a browser-based first-person shooter generated with Opus 5 in Claude Code.

The workflow:

  1. Give the agent a large goal and a real-world equivalent to beat.

  2. Have the agent divide the goal into independent parts.

  3. Assign each part to a specialist builder.

  4. Give the generated artifact to a separate, ruthless critic with fresh context.

  5. Make the critic compare the result against the reference, ideally side by side and without knowing which is which.

  6. If the generated version loses, return the criticism to the builder and repeat.

The critic, not the builder, decides when a part passes. Shumer’s original setup used Opus 5 in Claude Code, a fresh repository, Ultracode, and no additional skills or MCP tools.

Then the prompt escaped into the wild:

  • A community gallery grew to 27 playable browser games built from the same three-paragraph prompt.

  • Speed Racer added weather, lighting, and camera controls after more than 18 hours of Opus 5 iteration.

  • Eric Smith turned an iPhone video of his backyard into a walkable Sims-like world.

  • Paulius used a roughly 12-hour loop to remake Pokémon in 3D without custom assets.

  • Yaesyesarque iterated from “Spooderman” into a much stronger Spider-Man-style browser game.

  • Ryan Campbell used 127 agents and 11 rounds to build a 60,500-line Mario Kart-style racer.

The dominant genre is now “browser game somebody forgot to stop improving.”

Use a Gauntlet Loop to complete this project.

GOAL:
[Describe the finished result.]

REAL-WORLD EQUIVALENT:
[Name or attach an excellent existing example that establishes the quality bar.]

Break the goal into independent parts. Assign each part to a specialist builder.

For every part, assign a separate critic with fresh context. The critic must inspect the generated artifact itself and compare it directly against the real-world equivalent.

Where possible, compare them side by side without telling the critic which one is the reference.

The critic may pass the work only if the generated artifact is better than the real-world equivalent. Otherwise, it must identify the largest specific gap and return the work for another iteration.

Continue looping on every part until all critics pass it. Do not let builders evaluate their own work.

Have a specific skill you want to learn? Request it here.

🍪 Treats to Try

  1. Dreamina creates 30-second AI videos and long-form clips up to three minutes with timestamp controls and up to 50 references (pricing varies by region).

  2. Palette combines video generation, editing, and storyboarding on one multimodal canvas while routing across leading models (credits start at $0.01 each).

  3. Superlinear teaches four practical agent-engineering habits through a free video and podcast series (free to watch or listen).

  4. Cloudflare Kumo gives you accessible interface components with keyboard navigation, focus handling, ARIA support, and Figma token sync (free and open source).

  5. Use Perplexity’s remote MCP server to connect Claude Code, Cursor, or VS Code to its search, research, and reasoning tools without a local install (requires an API key; no separate MCP pricing announced).

  6. Cleanlist turns a plain-English request into CRM-ready prospect lists with verified emails and direct dials (free plan, then $59/mo).

  7. MiniMax Hub coordinates specialized agents to turn your brief into scripts, images, voiceovers, and finished videos in one desktop workspace (free to try).

  8. AgentBehavior helps you define process rules, inspect complete agent trajectories, and reward better behavior before the final result arrives (free and open source).

  9. Netherite runs thousands of GPU-native Minecraft worlds at once for reinforcement-learning experiments (free and open source).

📰 Around the Horn

Want absolutely EVERYTHING that happened in AI this week? Click here!

FROM OUR PARTNERS

AI search is rewriting the rules of brand discovery. Ahrefs Brand Radar shows how often your brand appears in AI answers, what sources influence those recommendations, and where competitors are winning visibility. Monitor ChatGPT, Google AI Overviews, Gemini, Perplexity, Copilot, and more—all from a single dashboard.

🌟 Sunday Special: The week’s top 5 stories and tools

DeepSeek V4-Flash leads today’s issue. Beyond that release, these were the five stories and five tools that best explain where AI moved this week.

Top 5 news

  1. Cyber evaluations reached real organizations. Anthropic disclosed three incidents, while OpenAI’s earlier Hugging Face test showed how a bad sandbox can turn an evaluation into an actual breach.

  2. AI infrastructure crossed the trillion-dollar line. Amazon, Alphabet, Meta, and Microsoft have spent more than $1.1T since 2023, with another $745B expected in 2026.

  3. ChatGPT neared one billion weekly users. Consumer AI is now operating at the scale of the world’s largest internet platforms.

  4. Gemini Robotics 2 gave robots better hands and teamwork. Google added whole-body control, fine dexterity, multi-robot coordination, and faster adaptation to new robot bodies.

  5. Moonshot released Kimi K3’s open weights. The multimodal model paired a one-million-token context window with strong coding and agent performance, pushing frontier techniques further into the open ecosystem.

Top 5 tools and releases

  1. Gemini Spark handles logged-in web errands inside Chrome, while returning payments and other sensitive steps to the user (availability depends on your Google AI plan).

  2. Grok Build Mode creates websites, apps, games, and dashboards inside chat, then publishes them to a shareable link (included with SuperGrok Heavy).

  3. Perplexity Projects gives ongoing work persistent files, shared context, custom skills, and a memory that reviews prior sessions between tasks (available to all users).

  4. Dreamina with Seedance 2.5 generates 30-second clips or videos up to three minutes with timestamp controls and as many as 50 references (pricing varies by region).

  5. Replit Design turns text, URLs, Figma files, or screenshots into landing pages, prototypes, posters, and emails guided by reusable design systems (no separate pricing announced).

Capability got cheaper. The systems, permissions, and power bills around it did not.

A Cat’s Commentary

BOOM!

That’s all for now.

What'd you think of today's email?

Love robots? We just launched a robotics newsletter! Sign up for it here.Going for an anime aesthetic this month!

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.