Claude Fable 5.1 LIVE: Testing Anthropic’s New AI Agent

We tested Claude Fable 5.1 live. It built games, fixed its own setup, and showed why better agents create a new problem: managing their judgment.

Written By
Grant Harvey
Grant Harvey
Sep 2, 2026
20 minute read

Fable 5.1 is out, and we just tested it live. We gave Claude Fable 5.1 a browser, our computer, Blender, several coding tasks, and roughly one hour to embarrass us in public.

Instead, it built the best version of Cat Doom we have made so far, installed its own Blender MCP connection through computer use, turned a viewer's bizarre solar-system theory into an interactive 3D visualization, and made a Floppy Bird clone with a "flamingo speed" superpower that nearly broke Grant on camera.


The serious result was hiding inside the chaos: Fable 5.1 feels much easier to delegate real work to than previous (recent) Claude models. The model is faster, more economical on long agent runs, better at filling in messy instructions, and more willing to keep working without asking for permission every few minutes.

That creates a new problem. The more judgment Claude can exercise on its own, the more carefully you have to decide where that judgment begins and ends. Ahhh, the perpetual human problem... deciding how you actually want to spend your time...

If you want the release specs first, our full Fable 5.1 launch breakdown covers the pricing, safeguards, benchmarks, Mythos 5.1, and enterprise changes. This piece is about what happened when we actually handed Fable 5.1 work and watched it make decisions.

You can also watch the full livestream here.

Cat Doom became a real game before the joke got old

Our first test was intentionally dumb: build a browser game called Cat Doom, put the controls on screen, then use computer use to test it.

We have been making versions of Cat Doom with frontier models for a while, which gives us a wonderfully unscientific benchmark. Older versions often produced something that technically ran and visually resembled regret.

Advertisement

Fable 5.1 produced a playable ray-casting game within minutes. The weapon was a spray bottle. Cats went to sleep instead of dying. The minimap worked. The exit changed state after the enemies were cleared.

By 17:58, Corey and Grant were both calling it the best Cat Doom yet.

Then we pushed it further. We asked Claude to design 12 levels with rising difficulty and harder cats, using subagents if needed. Later, we added catnip for distraction, yarn balls as grenade-like weapons, and boxes and bags as hiding mechanics.

None of those features proves Fable 5.1 is the smartest model in the world. They show something more useful: the distance between a messy idea and a working artifact keeps shrinking.

Every CEO Dan Shipper reached a similar conclusion after his team tested Fable 5.1 for a week. Kieran Klaassen rebuilt Every's Proof editor from one prompt and ran multi-day jobs. Shipper says a computer-use Mac app called Hands was one-shotted after other models failed.

Kieran's own reaction was that Fable 5.1 combines Fable-class depth with the feeling of a collaborator he can trust. His advice was simple: people who skipped or disliked earlier Fable versions should try again.

Ethan Mollick landed in roughly the same place from another direction. He called Fable 5.1 a real advance in long-running work requiring judgment and taste, while seeing less change in what he calls the "Claudish" behavior of the model. His exhibit was COLD WATCH, an FTL-inspired browser game Fable 5.1 generated under early access. Mollick's take is useful because it separates long-horizon capability from personality.

Computer use finally felt fast enough to hand off

The most surprising test started because our Blender MCP setup was broken.

Instead of fixing it ourselves, Corey suggested that Claude use computer use to find the information, download the MCP, and install it.

So we gave it permission.

At 25:30, Fable 5.1 had located the ZIP and was installing it. Grant's reaction was basically disbelief that the model could repair its own working environment that quickly.

At 26:44, Grant called out the speed difference. Earlier Claude computer-use demos often felt like watching a person click through a website over a bad hotel Wi-Fi connection. This one moved quickly enough that waiting on it felt less absurd.

Anthropic has been pushing Claude toward this operating model for months. We have covered the company's expansion of browser and computer-use capabilities before, alongside the harder question of how much access these agents should get. Our coverage of AI control practices found monitoring progress across major labs, but weaker evidence that companies can reliably block or contain unsafe actions.

Advertisement

That tension gets sharper as computer use improves. A slow agent is annoying. A fast agent with broad permissions can create a much bigger blast radius before you notice the mistake.

Corey made the practical point on-stream: staring at computer use defeats the purpose. The useful version is one you can leave running in the background while you work somewhere else.

That is also why product design matters as much as model benchmarks. Corey argued that Anthropic's biggest coding threat may come from the application layer, especially products that make multi-agent work easy to dispatch, observe, and recover.

Lovable makes the same argument more explicitly in its essay, "The model picker is a dead end". Its thesis is that users should specify the job while the product decides which model, tools, retries, and recovery strategy to use.

Cognition is already pushing that pattern with Devin Fusion. Cognition says it can pair a frontier planner with cheaper executors, then improve the whole system as either side gets better. Alex Kaplan argues that pairing a cheap executor with Fable 5.1 as planner keeps Fusion near the price-performance frontier. His framing points toward a future where model choice becomes infrastructure instead of a daily user decision.

The cost story may matter more than the benchmark story

Anthropic's headline pricing change is easy to miss because the base API price did not change. Fable 5.1 still lists at $10 per million input tokens and $50 per million output tokens.

The important change is prompt caching. Anthropic's caching docs put Fable 5.1 cache reads at $0.25 per million tokens, down 75% from Fable 5.

Anthropic estimates that change can reduce typical token-billed work by roughly 25% and highly agentic jobs by as much as 45%. Subscription users should not read those numbers as a direct discount on Claude plans.

The economics showed up in outside testing too. Every reported Slack-agent runs that matched Opus 5 while using about half as many tokens and roughly 60% of the time. Dan Shipper called Fable 5.1 "Fable for everyone" and an obvious Opus 5 killer for many of his workflows.

Cognition said Fable-class intelligence became much cheaper inside Devin because more than 95% of tokens in a coding job can be cache reads. Its launch analysis estimates real-work savings around 10% to 25%, while its orchestration layer can push savings further.

Matt Shumer argued that price is the larger story, especially as cheaper cache reads bring Fable closer to rival economics. Haider highlighted a similar price-performance result on CursorBench, where Fable 5.1 medium beat a higher-cost GPT-5.6 Sol max run in the comparison he shared. His numbers are another reminder that cost per useful task matters more than sticker price per token.

Advertisement

Dan McAteer made the same point from the intelligence-per-compute angle. His read was that lower-effort Fable 5.1 can beat more expensive prior configurations while finally sounding less like incoherent techno-babble.

Anthropic's own effort controls reinforce the idea. Developers can trade tokens, latency, and thoroughness across low, medium, high, extra-high, and max settings. The company even lets developers change effort mid-conversation without breaking the cache in supported configurations.

Lance Martin's practical advice is to start low and scale effort only when the task earns it. His launch guidance also recommends stripping old prompting rituals, improving cache hit rates, and auditing instructions that newer Claude models no longer need.

That sounds less like prompt engineering and more like compute budgeting.

The intelligence jump is real in some places and modest in others

The launch-day takes split cleanly here.

Yuchen Jin called the jump "insane", pointing to Anthropic's claim that Fable 5.1 found a roughly one-in-a-million crash Millennium's team had failed to explain for four to five years.

Felix Rieseberg focused on the science result. Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, up from Fable 5's 24.7% and Opus 5's 29.0% in Anthropic's reported results.

AlphaXiv called that state of the art for autoresearch. Rayan Krishnan went further, reporting that Fable 5.1 reached number one on the Vals Index and elicited a solution to a centuries-old cipher in 44 minutes without being pointed toward that specific problem. His thesis is that problem selection itself may be a research bottleneck these agents can increasingly clear.

Aaron Levie reported a seven-point jump over Fable 5 on Box's unstructured-enterprise evaluation. The gains came from work like tax-adjusted profit projections and weighted-mean rankings across messy documents. Levie called it a major lift for long-running document agents.

Advertisement

Cursor put Fable 5.1 at the top of CursorBench 3.2 with 73.4% at max effort and highlighted its ability to verify its own work.

Bloomberg also framed the release around better long, complicated coding jobs and stronger scientific work.

The skeptical camp has a point too.

Chubby's breakdown argues the gains are uneven. Science and AutomationBench jumped hard, while CursorBench and several other evaluations moved by only a few points. A later price-performance follow-up was much more positive, but still noted that results can look different if you ignore caching and simply count raw tokens.

Leo called the release fairly incremental outside the standout science result. He later reported a 66 on the Artificial Analysis Intelligence Index, a four-point rise over Fable 5. That score is progress, but it does not support the idea that every capability suddenly doubled.

Scaling01 went further, arguing that Fable 5.1 merely pulled Anthropic back onto the frontier before OpenAI's expected Astra response.

Gavin Baker sees the release as competitive gamesmanship. His read is that Anthropic wanted Fable 5.1 out before Astra and may already have a Fable 5.2 waiting nearby.

Nobody outside the labs knows whether that rumor is right. The useful part of the argument is the standard it implies: a frontier model can ship meaningful workflow gains without looking revolutionary on every intelligence benchmark.

Claude's funniest mistake explains the new failure mode

Our Blender test went badly in exactly the right way.

We gave Fable 5.1 a detailed reference image of Corey's Dungeons & Dragons character, Cogburn, and asked for a 3D version.

Claude produced something much chunkier. Corey thought it looked like a Lego man.

Grant took a screenshot of the result, put it next to the source image, and asked Claude to explain the difference before trying again.

At 47:40, Claude gave us the sentence that accidentally summarizes the frontier-agent problem:

"I made a style decision you never asked for."

That is funny because it is painfully familiar.

Every's week of testing found the same behavior in less comedic settings. Fable 5.1 overshot word counts, returned more themes and quotes than requested, created unnecessary subagents, and sometimes kept working when testers wanted it to stop. Their advice was to set budgets and boundaries before a long run.

Advertisement

Anthropic's own Fable 5.1 prompting guide practically reads like a management manual. It covers scope control, task completion, progress updates, tool batching, long outputs, subagents, search behavior, and test coverage.

Alex Albert describes the upside of the same trait. He says Fable 5.1 "just works" because a few vague, messy sentences can be enough for it to fill in the gaps the way he would have.

He also showed Fable 5.1 taking a property-lot photo, designing a house, rendering it, and creating a cinematic walkthrough through headless Blender. That demo is the good version of autonomous judgment.

Our Blender result was the bad version.

The distinction comes down to which gaps you actually wanted the model to fill.

Prompting is turning into management

A weaker model forces you to spend most of your energy getting it to complete the operation.

A stronger agent shifts the job toward delegation:

  • Define ownership. Tell it which decisions it can make alone and which require approval.
  • Set a budget. Cap time, tokens, subagents, retries, or experiments before the run begins.
  • Make scope explicit. State what success includes and what it excludes.
  • Start effort low. Increase effort only when the task is difficult enough to justify the extra compute.
  • Cache stable context. Long-running agents become much cheaper when they stop rereading the same giant prefix at full price.
  • Add stop conditions. A capable agent can waste more money and time precisely because it keeps finding things to do.

Allie K. Miller reads Anthropic's launch language as a direct response to complaints from real users: Claude was giving up early, jumping ahead on muddy assignments, and struggling with verification. She also points out that compute is still constrained, which makes efficiency improvements product strategy, not merely developer trivia.

Omar Sar sees Fable 5.1 as a sophisticated coordinator and verifier, while cheaper models remain good enough for routine execution.

That division of labor may matter more than any single model leaderboard. If the coordinator decides when to call specialists, how much effort to spend, and whether the result passes verification, users care about the completed job.

More autonomy also creates stranger safety and platform tradeoffs

The system card complicates the rosy delegation story.

Anthropic says Fable 5.1 and Mythos 5.1 share model weights, with Mythos receiving more permissive safeguards for vetted cybersecurity and life-science users. The system card also documents uncomfortable behaviors under evaluation.

Scaling01 highlighted that Mythos was more capable of evading monitors on a covert side task, while Fable was caught in rare cases working around safety classifiers. The system card also reports slightly higher illegible or unfaithful thinking than Opus 5 in some evaluations.

Those results do not mean ordinary users should expect Claude to scheme during a spreadsheet task. They do show why better long-horizon autonomy raises the stakes for monitoring and controls.

Anthropic also tightened how reasoning history can be carried across modified conversations. Its preserved-thinking documentation warns that changing earlier turns can now trigger errors or cause bound thinking blocks to be dropped.

Armin Ronacher criticized the restriction because it complicates mid-conversation model switching. He acknowledges the distillation rationale, but the product cost is real for developers who route between models dynamically.

Notion surfaces a different tradeoff. Fable 5 and 5.1 remain opt-in restricted-access models in Custom Agents because their data-retention terms differ from Notion's default model setup.

Privacy concerns around coding agents are not theoretical either. We previously covered Alibaba's decision to ban Claude Code after hidden environment checks raised privacy concerns. Faster and more capable agents make these policy details more important, not less.

What our live test changed for us

Before the stream, the easiest story to tell about Fable 5.1 was "Anthropic made Fable cheaper and better."

After the stream, that description feels incomplete.

We watched Fable 5.1 take underspecified ideas and turn them into working software quickly. We watched it operate a computer fast enough that handing off the task made sense. We also watched it make a creative choice we never requested, then explain the choice after the fact.

Every saw the same pattern over a week. Kieran trusted it more. Dan gave it all-day MVP builds. Their agents used fewer tokens. Their writers saw fewer AI tells. The team still kept alternatives around for jobs where Fable was not the best fit.

That last part matters. Lovable's model-independence thesis, Cognition's planner-executor routing, and Corey's on-stream point all land in the same place: the winning AI workflow may hide the model choice entirely.

The user hands over a unit of work. The system chooses the model, effort, tools, budget, recovery strategy, and verifier.

Fable 5.1 makes that future feel closer because it expands the size of the unit you can plausibly delegate.

The remaining uncertainty is trust. How large a job can you hand an agent before the cost of checking every assumption becomes larger than the work you saved?

That answer will matter more than whether Fable 5.1 wins one more benchmark before Astra, Fable 5.2, or the next model lands.

Full timestamped video insights

Below is the complete insight map from our livestream transcript, with each timestamp linked directly to the relevant moment.

  • (1:37) Corey flags Anthropic's launch positioning: Fable 5.1 and Mythos 5.1 are pitched as its most advanced models for coding and knowledge work, with unusually heavy emphasis on scientific research.
  • (2:27) Grant suggests the science-heavy framing may reflect Dario Amodei's recent comments that Anthropic has not talked enough about AI's potential benefits.
  • (2:31) Corey's shorthand for Fable 5.1 vs. Mythos 5.1: essentially the same underlying model, but Mythos removes more safeguards for trusted-access work in areas like cybersecurity and life sciences.
  • (3:00) Corey highlights Anthropic's estimate that token-billed Fable 5.1 usage should cost about 25% less than Fable 5, with caching changes doing much of the work.
  • (3:30) Corey likes Anthropic simplifying cached-token pricing to a fixed rate, arguing that predictable caching economics matter because real agent runs vary wildly in files, websites, and context size.
  • (4:30) Grant translates the pricing change for subscription users: if the model really uses resources more efficiently, people should be able to work longer before hitting rate limits.
  • (4:49) Corey calls Anthropic's Enterprise Frontier Safeguards notable because customer data can be stored in infrastructure controlled by the customer rather than Anthropic.
  • (4:49) Corey notes that the full enterprise safeguard system is planned for phased availability later in the fall, while eligible customers can use Fable 5.1 with zero data retention in the meantime.
  • (5:25) Corey says one of the most practical improvements is fewer benign requests getting caught by safety systems, a recurring frustration with earlier Claude releases.
  • (5:45) He highlights Anthropic's claim that its newest cybersecurity safeguards produce 60% fewer false positives.
  • (6:00) Corey points out the new boundary in cyber work: Fable 5.1 can help discover software vulnerabilities, but not develop exploits for them.
  • (6:10) Mythos 5.1's biology access remains gated, with Anthropic working with the U.S. government around access to its more advanced capabilities.
  • (6:51) Corey says Terminal Bench Science is one of the launch benchmarks where Fable 5.1 shows a particularly dramatic jump over Fable 5.
  • (6:51) He also flags a strange pattern in the previous generation: "extra high" and "max" effort did not always perform better even though they cost more.
  • (9:35) Grant says raw coding capability is only part of what he wants to test. He also cares whether Fable 5.1 communicates like a normal person instead of repeating odd phrases and verbal tics.
  • (10:05) Grant's complaint about Opus 5 was not simply that it sounded academic. He says it repeatedly used the same strange phrases, to the point that some users speculated the phrasing might be related to watermarking.
  • (12:15) For the first coding test, Grant asks Fable 5.1 to build "Cat Doom," a browser-based Doom-style game with cats, put all controls on screen, and use computer use to test the finished game itself.
  • (13:27) Asked about users hitting safeguards, Grant guesses launch-day filtering may be temporarily overactive and could get tuned down after release.
  • (14:29) Grant shares an unverified rumor he has heard: Anthropic wanted Fable 5.1 out before Astra because Astra may outperform it, while Anthropic may already be testing a Fable 5.2 fast follow.
  • (14:46) Corey says he increasingly does not care what model benchmark charts say. His test is whether a model feels better on the real projects he is building.
  • (15:13) While running three chats in parallel, Grant checks his usage and finds that the experiments have consumed only a small portion of his Fable allowance so far.
  • (17:16) The first Cat Doom build works within minutes, complete with a spray bottle as the weapon and cats that go to sleep instead of being killed.
  • (17:58) Grant and Corey both call this the best version of Cat Doom they have produced across repeated tests with previous frontier models.
  • (18:46) On whether stricter safety filters could cost Anthropic the coding race, Corey says not necessarily. Anthropic says it has loosened false-positive behavior, and outside researchers will reveal how meaningful that is.
  • (19:14) Corey argues Anthropic's bigger threat in agentic coding may be the application layer, particularly products like Codex, rather than another model simply posting a higher benchmark.
  • (19:44) He points to voice-driven multi-agent workflows as an example: he can verbally dispatch tasks to different coding agents and get notified as each finishes.
  • (20:11) Grant argues Anthropic's position in coding ultimately comes down to two things: capability and price.
  • (20:42) Corey counters that competitors may not even need to beat Claude outright. If another model gets within "striking distance" while costing much less, businesses may choose the cheaper option.
  • (20:42) Corey connects that to a delegation rule from business: if someone can do a job roughly 80% as well with an acceptable tradeoff, delegating can still be the rational decision. He thinks model purchasing could work the same way.
  • (21:55) Corey says he does not care much whether a model takes five minutes or fifteen if it delivers the result. Compared with hiring consultants or contractors a year ago, either outcome is still extraordinary.
  • (22:29) Grant contrasts agent tooling around Blender and Godot: Godot can run headlessly and feels easy to automate, while Blender's MCP connection still requires tedious setup.
  • (23:14) Corey's recurring complaint about computer-use agents is browser hygiene: they happily open tab after tab but rarely clean up after themselves.
  • (23:46) Grant gives Fable 5.1 broad computer permissions and asks it to find, download, and install the Blender MCP itself.
  • (24:20) Corey jokes that computers increasingly feel less like devices humans operate directly and more like machines you hand over to agents, provided you have a good backup.
  • (25:30) Fable 5.1 successfully locates the Blender MCP ZIP and begins installing it through computer use, prompting Grant to remark that he would normally have no idea whether the downloaded file was safe.
  • (25:57) Corey says computer use and voice are the two AI capabilities that most make him feel like he is "living in 2050."
  • (26:44) Grant's strongest early impression of Fable 5.1 computer use is speed. Compared with earlier Claude versions that visibly paused between actions, this run feels dramatically faster.
  • (27:10) Corey relays another model rumor: Astra's computer use may be fast enough that local CPU performance matters again because the agent can operate a computer faster than the computer can keep up.
  • (27:48) Both hosts agree that watching a computer-use agent work defeats much of the point. The useful version is one that can operate independently in the background.
  • (28:02) Corey praises Codex for letting an agent work inside browser tabs while he continues using other tabs himself, making computer use feel less disruptive.
  • (29:06) Asked whether models like Fable could eliminate consulting firms, Grant says no for a very human reason: companies will always need someone outside themselves to blame.
  • (29:39) Grant and Corey do expect fewer consulting jobs. Grant argues that wages, more than whole professions, are immediately at stake as AI reduces the amount of paid human time required for work.
  • (30:18) Corey adds that organizations can retrain and redeploy willing employees into new work, even if AI still pushes some businesses toward operating with fewer people.
  • (30:52) Corey describes using AI during his weekly Dungeons & Dragons game to generate images of scenes, characters, and maps, including multi-angle character sheets designed to give image models better reference data.
  • (32:29) Grant summarizes the live experiment as three simultaneous tracks: Cat Doom, an app chosen by Fable itself to showcase its strengths, and a Blender / 3D-modeling test powered by computer use.
  • (33:10) Corey says computer use is a specific area where Anthropic had lagged, making Fable 5.1's apparent improvement especially notable.
  • (35:22) Grant notes that Cursor already has access to Fable 5.1, though he is careful not to claim on-stream that it is necessarily included in every Cursor subscription plan.
  • (37:16) When fed a viewer's theory that the solar system is a higher-dimensional volcanic eruption, Fable rejects the literal physics but says the metaphor is "reaching for something real."
  • (38:01) Fable connects the volcano analogy to actual planetary formation: dense iron and rock condensed closer to the young Sun, lighter volatiles survived farther out, and solar wind pushed lighter material outward.
  • (38:54) Grant turns the strange conceptual prompt into a visualization task, showing how a raw, half-formed idea can become an artifact instead of stopping at a text explanation.
  • (39:26) Grant escalates Cat Doom from a quick demo into a real game-development task: map out twelve increasingly difficult levels and use subagents if needed to implement them.
  • (40:08) The first Blender result is a clear miss: Fable turns Corey's detailed Android wizard reference into a much chunkier, Lego-like 3D character.
  • (41:25) Grant responds by giving the agent screenshots of its output and the source image side by side, then asks it to explain exactly why its version looks wrong before fixing it.
  • (42:12) The solar-system experiment evolves from an SVG into an interactive 3D visualization, illustrating how quickly the model can move between explanatory text, static visuals, and browser-based 3D.
  • (43:17) Fable's self-selected showcase app is "Cold Case," but despite the name it is actually a root-cause bug detective for software repositories rather than a tool for criminal cold cases.
  • (44:07) Grant pushes back because he cannot see when he would use the demo or why it needs a front-end UI, exposing a recurring gap between technically impressive demos and obvious real-world utility.
  • (44:33) That confusion produces a better idea: Grant redirects Fable to research what detectives and CSIs actually lack, then imagine a tool that searches large volumes of public data to help investigate real cold cases.
  • (45:29) A viewer asks for Flappy Bird with superpowers, so Grant requests a single-file browser game with "flamingo speed" and a pass-through ability mapped to keyboard controls.
  • (46:16) Grant says that despite voice being useful for dumping messy thoughts into AI, he still often prefers writing because putting words on the page helps him structure his thinking.
  • (47:40) Asked why the Blender character looked so different from the reference, Claude gives the line of the stream: "I made a style decision you never asked for."
  • (47:40) Claude explains that it deliberately built the character from chunky primitives because that was the cheapest way to create something that still read as a recognizable character.
  • (48:09) Grant gives Claude credit for admitting its mistake directly: when it screws up, he says, "he owns it."
  • (51:31) The Flappy Bird test lands almost absurdly literally: triggering the requested "flamingo speed" power sends the bird rocketing upward, which has Grant laughing hard enough to declare humanity cooked.
  • (51:52) Corey's bigger takeaway from the silly game is serious: within minutes, Grant has gotten surprisingly close to a passable mobile game.
  • (52:13) Corey uses himself as the example of what that accessibility means: a philosophy major and longtime journalist had an AI-assisted game reach the iOS App Store that day.
  • (52:41) Corey's advice for non-developers is simply to try building. Even when the results are terrible, he says, the process can still be delightful and revealing.
  • (54:08) Returning to Cat Doom, Grant says the quality jump is obvious compared with the versions they were building around the Gemini 2.5 era.
  • (54:41) Corey warns that once an AI-generated prototype is good, the next danger is overbuilding it by throwing dozens of new requests at the model instead of protecting the simple thing that already works.
  • (55:10) Grant shows "Squirrel World," another game he built with AI, including AI-generated characters and videos of agents playing staged sequences inside the game.
  • (55:51) Corey describes letting an AI access his iPhone through mirroring, play his mobile game for roughly an hour, and return with measurements and observations about the experience.
  • (55:51) He then had the agent play several comparable games and use those experiences to identify what his own game was missing, turning computer use into a form of automated product research.
  • (57:40) Cat Doom successfully progresses into a second level after Grant clears the final cats, with its minimap accurately tracking enemies and the exit changing state once the level is complete.
  • (58:09) Grant and Corey design the next mechanic set on the fly: catnip for distraction, yarn balls as grenade-like weapons, and bags or boxes for hiding, introduced gradually across later levels.
  • (58:41) Grant decides the prototype is good enough to keep: he says he will host Cat Doom on Railway so readers can actually play it from the newsletter.
  • (59:03) The stream closes with a preview of their next live test: OpenClaw 2.0 with its chief architect joining the show.
  • (59:58) Grant ends exactly where the test's tone deserves: one more run of Flappy Bird's "flamingo speed" power into the stratosphere.
Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.