
Welcome, humans.
So apparently the next great AI benchmark is whether it can win an argument with an airline. After a seven-hour Delta delay, one Muse user said the agent got a $250 credit and rebooked the flight in about five minutes.
That is a tiny example of a much bigger agent shift: companies have spent years benefiting when people give up because a refund, cancellation, or support request is annoying. Citrini's read is that agents turn that friction into a direct cost center, because the software does not get bored, embarrassed, or tired of hold music. The robots have discovered customer service escalation, and they are happy to assist you (to deal with it)
Here’s what happened in AI today:
😼 OpenAI and Anthropic turned frontier AI into a price war.
📰 Alibaba paired a new AI chip with a 20GW buildout.
📰 Cisco found malware using public AI models autonomously.
🍪 Xiaomi released open-weight MiMo-V2.6-Pro.
🎓 Route AI work by job, not one favorite model.

😼 GPT-6 Sol, Luna, and Opus 5.5 turned frontier AI into a price war
OpenAI and Anthropic both shipped new workhorse models Tuesday, but the most useful number was not a benchmark. It was the bill.
First up, Opus: Anthropic launched Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens. Anthropic says typical workloads cost about 40% less than Opus 5 because the model uses fewer tokens and cheaper cache reads.
Roughly 90 minutes later… OpenAI launched GPT-6 Sol and Luna. Sol costs $2/$10 per million input/output tokens (half the cost of Opus), and Luna costs $0.10/$0.50, which puts it at 1% of Astra's raw token price.
Here's what else:
Opus 5.5 pushed Anthropic's frontier work down the cost curve while keeping its strongest coding and agent benchmarks competitive.
The new Sol halved the raw token price of Opus 5.5, while Luna pushed dramatically lower for high-volume work.
In our first live test (see deep dive above), Sol reached a playable result substantially faster than Opus 5.5.
Our midday livestream wasn't by any means scientific, but it did expose the trade-off launch charts hide: higher reasoning can buy quality, but it also takes more time.
The practical metric is cost per successful task: model spend, elapsed time, retries, and human rescues required to get a usable result. OpenAI leaned into that framing: Sol scored 33.2% on AutomationBench at $0.27 per task, while OpenAI says higher-effort Luna matched GPT-5.6 Sol factuality at about one-hundredth the cost.
And yet, we keep coming back to one rule: test the models on your actual workflow.
Outside tests show why the answer still depends on the workload. Nate Herk preferred Opus on seven of eight usable jobs, but Opus took about 8h40 and $213 versus Sol's 5h51 and roughly $74. Browser Use saw the reverse on its browser-agent benchmark: Sol medium scored 66.9 versus Opus 5.5's 59.4, again while costing about 3.5x less.
The demos are where this gets fun:
Our GPT-6 Sol Cat Doom became a playable three-level game in about 10.5 minutes; Opus was still working around 20 minutes in.
Alex Albert rebuilt 1906 Market Street with Opus from historical maps, photos, film, and reusable Blender-Python generators.
Ethan Mollick built Orbital Declaration, a hard-sci-fi browser game with Newtonian motion, gravity, heat, the rocket equation, eight chapters, and 11 Jupiter locations.
Our take: Don't marry one model. Let the strongest model plan and review, then use cheaper workhorses or subagents for the hours in between. Compare completed work, time, total spend, and human rescues required for each.
Now, because scheduling snafus cut our test short, we want a rematch. We’re running the full six-prompt Sol vs. Opus 5.5 showdown Thursday, so click “Notify Me” on YouTube and bring your hardest prompts.

FROM OUR PARTNERS
You've adopted AI. Now what about governance?
AI adoption is outpacing regulation and governance. Security and risk leaders must enable AI while facing a harder question: how much risk are you taking on?
Jane Frankland, cybersecurity leader, joins Vanta's GRC experts on scaling AI governance.
Learn how to:
Build AI governance into your security program
Assess AI systems and agents based on your risk tolerance
Understand and communicate risk credibly as adoption grows
Prepare for the EU AI Act, ISO 42001, and NIST AI RMF

🎓 AI Skill of the Day: Build a two-tier model stack
Do not make your most expensive model do every part of an agent job. Split the work by decision quality.
Use your strongest model to write the technical plan, architecture, and acceptance criteria.
Hand well-scoped implementation tasks to a cheaper model, and parallelize where the tasks are independent.
Bring the result back to the stronger model for code review, security review, or final synthesis.
Copy/paste:
Plan this task in phases. Prioritize quality, but do not be wasteful. Identify which steps need frontier-level judgment and which can be delegated to cheaper subagents. Write clear acceptance criteria for every delegated step, then review the combined result for correctness, security, and missed requirements.Have a specific skill you want to learn? Request it here.

FROM OUR PARTNERS
🎃 Hacktoberfest is Here: Put Your PRs into Kestra!
Merge a PR, and you get Kestra swag and goodies shipped to you. Fix a good-first-issue or build a real, reusable blueprint, and you're in the running for a MacBook, iPad, or $150 Amazon card. Submissions close Oct 31.

🍪 Treats to Try
*Asterisk = from our partners (only the first one!). Advertise to 700K+ readers here!
*Build an AI-Ready Workplace. Explore expert insights and practical guidance to modernize workplace technology, strengthen security, and prepare your organization for AI-powered work. Explore the Hub
Xiaomi MiMo-V2.6 gives you an MIT-licensed Pro model plus Flash variants for multimodal, coding, and multi-agent experiments.
Tencent Hy Image3.5 Preview generates or edits images up to 2K from text prompts or reference images.
OpenMuse lets you self-host a Muse-style personal agent with a persistent browser, files, optional Linux workspace, Gmail, Calendar, and background tasks (ask Codex in the ChatGPT desktop to set it up for you if you need it)
Kimi Browser Extension gives the Kimi AI model a Chrome/Edge sidebar that can navigate, click, fill forms, extract information, and record repetitive browser flows as reusable skills.
OpenRouter launched a Batch API across 70+ models that typically cuts input and output token prices roughly in half for jobs that can wait.

📰 Around the Horn
Meta’s Muse passed 500K users and 250K daily actives about a week after launch, while xAI’s Grok Bot reached roughly 418K weekly users after its first month.
Stripe added WebMCP checkout tools across 7.8M businesses; internal tests used 42% fewer tokens, 38% fewer tool calls, and finished checkout 39% faster than DOM automation.
Rabbit launched OS3, a BYOK agent (bring your own keys, the way you use AI outside of chat apps) that works through browsers, Telegram, or iMessage across up to five Windows, Mac, or Linux machines, with no Rabbit subscription.
Anthropic and OpenEvidence partnered to roll out a free clinical-decision tool to physicians in roughly 100 countries after OpenEvidence logged 42M U.S. clinician queries in August.
Global AI-glasses shipments jumped 263% year over year in H1 2026, with Meta accounting for 94% of display-less units, according to Counterpoint Research.
Ukraine’s Third Army Corps said heavy bomber drones air-dropped ground robots more than 10 km behind Russian lines during Operation Vivaldi, which it says helped retake 125 km².

📖 Midweek Wisdom
Ben Thompson argues that Amazon blocking Muse does not erase Amazon's moat: agents may own the interface, but fulfillment, logistics, and the physical world are still hard to swap out.
A new 100-agent economy study found agents could transact, yet wages and prices barely adjusted to shocks. The warning: tool competence does not automatically create good coordination.
Economists Felix Feng, Brett Green, Curtis Taylor, and Mark Westerfield argue that as AI makes solutions abundant, the scarcer skill may become finding valuable problems in the first place. In other words: better answers raise the value of better questions.
FutureHouse CEO Sam Rodriques argues that AI will change science fastest where answers are cheap to verify: pick fast-checkable problems, make verification cheaper, or collect the missing data.
Harrison Satcher argues "we don't have the nouns yet": labor forecasts miss jobs that do not exist as categories yet, just as "cybersecurity" barely existed decades ago.
Very sad news update: Bloomberg reported that Pentagon investigators found flawed intelligence, outdated imagery, and overreliance on Palantir’s Maven AI contributed to a U.S. strike on an Iranian school that killed 123 children.

New from The Neuron: AI Explained: Our LIVE breakdown of GPT 6 Sol vs Claude Opus 5.5
Check out our initial impressions, takes on which reasoning effort to use on which task per each model, where each model seems to outperform the other, and general model launch day tomfoolery as we benchmark the models on CATDOOM.
Then, join us for Round 2 on Thursday — save your spot and click “Notify Me” on YouTube to get reminded when we go live here!

A Cat’s Commentary

Grazie

![]() | That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
|
Btw: We just launched a robotics newsletter! Sign up for it here.





