
Welcome, humans.
Okay, so GPT-6 Astra just learned the most human gaming lesson imaginable: never put all your valuables in one chest in Minecraft.
In Vals AI's long-horizon Minecraft eval, Astra apparently built a semi-automatic blaze farm, collected six blaze rods and three pearls… then watched a creeper blow up the chest holding all of Astra’s items. D-D-DEVASTED…
After it all went down, poor Astra spent hours farming potatoes and started explicitly checking whether tall green objects were sugarcane or creepers.
This is why we need better alignment, y’all. We can’t have AI out here getting so tilted gaming it becomes paranoid and traumatized like us humans! You think recursive self improving superintelligence is scary now, wait until you meet the traumatized supervillain origin story RSI…
Here's what happened in AI today:
😺 ChatGPT co-creator launched a model built for structured decisions.
📰 Dario Amodei proposed embedded third-party AI evaluators.
📰 Meta argued alignment will become a competitive advantage.
🍪 Muse gives non-technical users a consumer-friendly AI agent.
📰 Agility Robotics unveiled a 284-pound humanoid built to work beside people.

😺 JEV is a new “System One Model” that skips the whole chatbot part of AI
AI models like ChatGPT based on large language models, or LLMs, have become excellent at talking. Well, ChatGPT’s co-creator and TypeSafe's Founder Diogo Almeida thinks that may be exactly why they are still awkward at automating software.
After helping develop the research behind ChatGPT, Almeida spent two years building a different approach: System One models. The first public model, JEV, is designed to make fast, structured decisions instead of generating prose one token at a time.
Here's what happened:
JEV takes structured questions and returns typed answers plus calibrated probabilities.
TypeSafe says it answers in roughly 70 to 500 milliseconds and can be 20 to 200x faster and 40 to 400x cheaper than comparable LLM workflows.
Its training method, Reinforcement Learning for Calibrated Decisions (RLCD), is designed to make the model's confidence useful to software.
TypeSafe says JEV can work out to roughly $42 per billion tokens of equivalent workload, which is why the cost difference starts getting wild once you use AI for millions of tiny decisions.
The easiest way to think about it: a normal LLM is great when you need an explanation. JEV is aiming at the moment when you just need a decision, like when software needs to decide whether a customer request is fraud, whether an alert should escalate, or which of 1,000 records need action.
Instead of forcing a chatbot to produce JSON and then writing code to check whether the JSON is sane, the decision itself is the product.
Why this matters: Most agent workflows still use a language model as both thinker and talker. That is expensive, slow, and often overkill. JEV suggests the AI stack may split into specialists: language models for communication, coding models for implementation, math models for proofs, and decision models for high-volume judgment.
We talked about this on a surprise live yesterday: JEV looks kinda like a "true intelligence layer" that another AI could call when it needs to make multiple parallel decisions at scale. IF this interpretation is totally wrong, bare with us. This thing is brand-spankin’ new and highly technical, so we’re trying to wrap our minds around it!
Want to try it? TypeSafe is currently taking access requests, so the next test is whether those speed and cost claims survive real production workloads. If they do, the interesting part may be what developers stop using giant chat models for (if you use cheap and dumb language models for an intelligence layer in an automated workflow… this is probably for you).

FROM OUR PARTNERS
The SOC 2 Checklist that Closes Enterprise Deals
SOC 2 is often the key that unlocks enterprise customers, new markets, and investor confidence. This free checklist from Vanta maps every step — from scoping your program to continuously maintaining compliance. Whether starting from scratch or filling gaps before your audit, you'll know exactly what's required.

🎓 AI Skill of the Day: Add confidence thresholds to agent decisions
One underrated idea in JEV is not the speed. It's the probability attached to the answer.
You can use the same pattern in your own AI workflows today: make the model return a decision plus a confidence score, then route low-confidence cases to a human or a stronger model.
Try this structure:
Ask for a fixed decision, not an essay.
Require a confidence score from 0 to 1.
Set an escalation rule, e.g. anything below 0.85 gets reviewed.
Prompt:
Classify this request as APPROVE, REVIEW, or REJECT. Return only the decision, a confidence score from 0 to 1, and one sentence explaining the most important uncertainty. If confidence is below 0.85, choose REVIEW.It will not magically calibrate a general-purpose model the way Jev’s RLCD process is designed to, but it gives your workflow a useful escape hatch instead of pretending every answer deserves the same level of trust.
Have a specific skill you want to learn? Request it here.

🍪 Treats to Try
*Asterisk = from our partners (only the first one!). Advertise to 700K+ readers here!
*Why We Love It: It turns "we have an incident" into "it's already fixed." Watch how Cursor and PagerDuty deploy AI agents in Slack.
Meta launched Meta One, a subscription bundle across Instagram, Facebook, WhatsApp, and Meta AI with higher AI usage and 50+ features; Meta says its plans have already reached 15M subscriptions and trials.
Gemini 3.8 Live gives you faster multilingual voice conversations, while Live Extended Thinking adds a heavier reasoning mode that thinks longer before speaking.
Aside runs agent tasks across your logged-in websites and local Windows files while keeping credentials scoped and asking before high-risk actions like posting or payments.
Sketchpad Live turns GPT-Live or Astra into a whiteboard teacher that can draw, move shapes, and narrate a lesson while you follow along.
Concat gives you an open-source CapCut-style desktop editor with multi-track editing, offline Whisper captions, local TTS, and no account or watermark (btw, these “GitHub” links are code your agent can run for you! Learn more about how that works here)
OpenArtifacts gives Codex, Claude Code, Hermes, Pi, and OpenCode a simple place to publish reviewable HTML or Markdown artifacts you can inspect in a browser.

📰 Around the Horn

lol ruthless
Elon Musk proposed labs test one another's models, including across U.S. and Chinese companies, as a practical alternative to a universal pause, while Mark Zuckerberg argued labs should slow themselves when safety requires it, saying trust and alignment will become product advantages rather than separate compliance work.
OpenAI is reportedly in early talks for another funding round at roughly a $1.2T valuation, after annualized revenue topped $40B.
Agility Robotics unveiled Digit 5, a 5'11", 284-pound humanoid built to work beside people without safety barriers, carrying up to 50 pounds and recharging in nine minutes.
Periodic Labs connected its 1T-parameter Neon model directly to physical materials experiments, creating a loop where the AI model proposes work, the lab runs it, and the results train the next round.
Nous Research used 1,393 Fable subagents to refactor a million-line codebase (refactor = improve with the same functionality) in about 19 active hours for roughly $25K, while human review still caught regressions.
Google's AI-in-science analysis found 74% of surveyed scientists saved time with AI, averaging about 6.9 hours a week, while verification and untested hypotheses emerged as new bottlenecks.
Gensyn released open-1b with a verifiable training record, so outsiders can rerun parts of its training on different hardware and check that the model was actually trained the way Gensyn says it was.
Profound raised $180M at a $1.8B valuation after reporting 3x revenue growth in six months and more than 1,000 enterprise customers for AI-search optimization.

FROM OUR PARTNERS
The people shaping what comes next in AI are gathering in San Francisco
The people shaping what comes next in AI are gathering in San Francisco.
At The AI Conference, hear from 130+ speakers, including Chris Lattner, Emmanuel Ameisen, Peter Norvig, Illia Polosukhin, co-author of “Attention Is All You Need,” plus builders from OpenAI, NVIDIA, Google, Near AI, and more.
Hear what leading AI teams are actually building and what’s changing across agents, LLMs, infrastructure, and applied AI before it becomes common knowledge.
NEURON members save 30% with code NEURON30.

📖 Midweek Wisdom
We Must Pace the Frontier: Dario Amodei's concrete proposal for embedded evaluators, industry coordination, and eventually international limits.
Box CEO Aaron Levie argued the agent economy gets much bigger when AI starts doing background work no one explicitly prompts, like recruiting, contract review, software tests, and transcript mining.
Shopify CEO Tobi Lütke warned against “slop grenades,” where you dump cheap AI output on coworkers and make them spend the time you saved reviewing it; as generation gets cheaper, judgment matters more.
Gergely Orosz looked inside OpenAI's increasingly Codex-driven engineering process, showing how agentic teams are changing who writes code, who reviews it, and where humans stay in the loop.
François Chollet argues intelligence is better measured by how efficiently experience turns into competence, not just benchmark scores, and estimates today's AI remains roughly six orders of magnitude behind humans by that measure.
Scott Aaronson and Daniel Litt ask what happens when theorem generation gets cheap: human mathematics may shift toward understanding, exposition, verification, and defending why a proof actually matters.

Join us LIVE to learn all things OpenClaw!
OpenClaw is basically the O.G. personal agent, and this Thursday @ 4pm PT | 7pm ET, we’re going LIVE with OpenClaw’s Chief Architect, Vincent Koc, to answer all your questions and learn about the new OpenClaw 2.0. Click here to save your spot.

A Cat’s Commentary

Are you therefore asking for MORE killer robots in our reporting?

![]() | That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
|
Btw: We just launched a robotics newsletter! Sign up for it here.







