
Welcome, humans.
So apparently Alex Ziskind spent about $60K on four Mac Studios, which he then wired together, and gave the whole rig the same coding job as a cloud agent.
Here’s how that ended up:
His local Kimi K3 cluster (Kimi K3 is one of the biggest, most powerful open weight AI models that anyone with, y’know, a cluster of Mac studios can run) took roughly four hours to finish the job.
The cloud agent finished in about 15 minutes.
The takeaway? Local AI (open models you can run on your own computer / servers / nuclear powered data center) still wins when private data cannot leave your building, or when you want complete control over the machine. It just might not be fast… yet.
So in this test, the cloud hare smoked the local tortoise. He basically built the Voltron of Mac Studios and then got totally wrecked by a 12x faster Godzilla.
Here’s what happened in AI today:
😺 OpenAI launched GPT-6 Astra for longer agent jobs.
📰 xAI launched persistent Grok Bots for enterprise work.
📰 Google mapped a complete male fruit-fly nervous system.
🍪 Gemini Spark can now organize and edit Google Photos.
🎓 Write down project intent before handing work to agents.

😺 GPT-6 Astra can stay on the job and use your computer
Good news and bad news, folks.
Good news: Yesterday, OpenAI launched GPT-6 Astra, its new flagship model built to operate software and stay on longer jobs.
Bad news: you and I don’t get access… yet.
Good news: In OpenAI's demo, one Astra session turns a yellow circle into a rocket, opens Blender, and makes a printable 3D file. At the same time, it builds a game, edits a contract, drafts an eBay listing, orders lunch, and books tennis.
Bad news: Only limited partners get access to the model at first, while the rollout for other paid-plan holders starts over the coming days. Thankfully, Tibo (Head of Codex) said we’ll get a banked reset on rate limits for each day we don’t get it.
Here's what’s new and notable:
On OSWorld 2.0, a test of real computer tasks, Astra scored 72.6% and averaged about 40 minutes per task. Sol scored 65.7% and took about 75 minutes.
ARC Prize scored Astra on one of AI’s toughest tests at 62.7% with a neutral setup, then ~99.9% when it kept OpenAI's built-in reasoning state and memory-management system.
Artificial Analysis found Astra roughly tied Sol on its broad benchmark index, while using about one-third as many tokens on its coding-agent test. This has some folks saying benchmarks need a refresh in the age of long running agents.,
OpenAI says Astra reached its Critical cybersecurity threshold, a capability level that triggered extra safeguards, and found two previously unknown software flaws during testing. And then there’s the whole neuralease thing… (read the deep dive).
So why can the same model score 62.7% in one setup and 99.9% in another?
An agent (like Astra) is a system. Its harness is the software that manages tools, memory, retries, and what survives a long job. Change that surrounding software and the same model can behave very differently.
That gives you a better way to test AI in your own work, which is basically the only way to benchmark AI right now: Give two agents the same messy job, and track all the completed tasks, human rescues, elapsed time, and total cost.
So we recommend once this thing is actually available that you try it against Fable 5.1 on the same job and see which you prefer!
So... did OpenAI just ship AGI? OpenAI president Greg Brockman is now willing to say the A-word. AGI, or artificial general intelligence, is the disputed label for AI that can perform broadly across many domains at human-level or better.
Brockman said he expected AGI to arrive as one dramatic moment, but now thinks it showed up in pieces. He said calling Astra the first AGI is reasonable, and ended the briefing with “Welcome to the AGI era.”
That is an OpenAI executive’s claim, but there is no agreed AGI test, so hands-on reports from early testers on how GPT-6 performs are more useful than the label:
Latent Space burned 20B+ Astra tokens and called Astra an AI engineer you can hire for <$6/hour that can run pipelines, inspect logs, deploy systems, and manage subagents.
Every found a big jump in writing, software use, and visuals, but still preferred Fable’s product judgment when the job required simplifying instead of adding more.
Claire Vo said Astra made her far more ambitious after it cracked work prior models could not, especially when computer use let it drive real production tools.
Ethan Mollick noticed less drift on long jobs: Astra is better at remembering which ideas are still live instead of dragging discarded drafts back into the final answer.
The demos are where this gets easier to understand:
Pietro Schirano of MagicPath turned an image into a coded 3D animation, built a one-prompt underwater game, and had Astra compose a track directly inside Ableton.
Ethan Mollick’s Alexandria was built while Astra worked autonomously for days. You can walk through the playable reconstruction yourself.
Mollick also gave Astra an open-source stormy-ocean demo and asked for the world underneath it. Astra added reefs, deep-sea habitats, and procedural animals; play it here or inspect the code.
Arena AI ran a zero-cherry-pick 3D gauntlet across castles, underwater scenes, Van Gogh’s house, and open-world games, then published the prompts so you can rerun it.
Playco gives the flashy demos a production check: Astra turned one gray-box game into three themed prototypes with 50% fewer manual fixes than the previous model.
And finally, for more bad news: OpenAI says Astra's written reasoning became harder to monitor, which does partially explain the limited rollout. More autonomous systems need tighter permissions, especially if we can’t read all their thoughts, so how this model behaves in the wild may very well reveal how that tradeoff holds up.

FROM OUR PARTNERS
Want to get more out of Claude? Build personalized skills (step-by-step guide)
You're paying for Claude but not hitting your full potential. This free guide shows you how to build high-end Claude Skills with no prior tech knowledge required. Your style, your voice, your data, supercharged.

🎓 AI Skill of the Day: Give agents the “why” first
Agents can follow every instruction you give them and still build the wrong thing because they never understood why the project existed.
Anthropic's AI-native software playbook, highlighted in Rob Shocks' walkthrough, starts with a tiny file called intent.md. Think of it as the note you would leave for a smart coworker taking over the project tomorrow.
Describe the outcome and who it helps. Add the non-negotiable constraints and a concrete definition of done.
Let the agent interview you until the fuzzy parts are gone.
Save the answers as
intent.md, then use that file to create the specification and plan.
The same habit works outside coding. Before a long research, writing, or analysis job, give the agent a one-page brief it can keep returning to.
Have a specific skill you want to learn? Request it here.

FROM OUR PARTNERS
New AI audio tools are now widely available
Generate Music, Generate Speech, and Generate Sound Effects bring commercially safe audio to your creative projects. Try Firefly AI Assistant free, plus access to leading models like Gemini Omni Flash, Runway Aleph 2.0, Kling 3.0, and Adobe's own.

🍪 Treats to Try
*Asterisk = from our partners (only the first one!). Advertise to 700K+ readers here!
*Guidde turns your software workflow into a step-by-step how-to video you can share as an SOP; free plan, then $19/creator/mo billed annually.
Gemini Spark + Photos can find, enhance, organize, and prepare photos to share from one prompt while keeping originals untouched and requiring confirmation before sharing.
Warp Factory Benchmarks replays your company's real coding tasks across models, then shows which setup gives you the best cost-versus-quality tradeoff.
Hermes Desktop handles the fiddly parts of local AI by installing the software needed to run the model, matching models to your hardware, and managing memory.
Zite gives ChatGPT, Claude, Cursor, and other agents one shared database and permissions layer they can build against directly.
Perplexity for Stripe creates the project and secret key an app needs to call Perplexity, cutting out the usual credential copy-paste.
Town adds your personal assistant to group texts so it can compare calendars, research options, and book or buy things after you approve; free plan, then $15/mo.

📰 Around the Horn

Shockingly, this did NOT go off the rails…
xAI launched Grok Bot for Enterprise: persistent agents get cloud computers, learn routines by watching once, and can pass context to other Bots.
Google Research and HHMI Janelia published the largest brain wiring map yet: 166,000+ neurons and 125M connections across a male fruit fly's nervous system.
IFM released K2 Horizon, six open models from 0.9B to 375B parameters (that’s how big they are), plus training data, code, checkpoints, and logs.
WorldAgents coordinated existing image-making and image-reading models to build explorable 3D scenes without training a new world model.
MIT CSAIL's Software World puts persistent coding agents in a fake GitHub to maintain packages, file issues, review pull requests, and face hidden tests.
ComfyUI launched Forward Deployed Creatives, sending production artists into companies to build custom generative-media workflows and train the internal team to own them.
Bernie Sanders and Greg Casar proposed banning superintelligent AI and pausing advanced development; if enacted, individual violators could face up to 20 years in prison.

💡 Intelligent Insights
Terence Tao warns AI could generate proofs and experiments faster than humans build intuition around them, making scientific judgment and research taste more valuable than raw output.
Ethan Mollick says “multiplayer AI” is still underbuilt: teams need shared context where several humans and agents can work toward one goal instead of isolated chats.
Amelia Michael argues robot benchmarks often confuse weak software with weak hardware, so we should test what today’s robot bodies can already do before replacing them.
Chris Paxton thinks long-context video could become a natural robot prompt: show a machine one demonstration, then let it imitate the behavior without a custom retraining cycle.
Just Another Pod Guy argues cheaper digital intelligence shifts the bottleneck toward robotics, sensors, wet labs, and other physical systems that can turn model output into real production. Go make / fund those!

A Cat’s Commentary

Aww shucks i’m blushing

![]() | That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
|
Btw: We just launched a robotics newsletter! Sign up for it here.





