
Welcome, humans.
ICYMI our Monday meme yesterday: the internet discovered Satyress’s Threehalves, possibly the most demonic robot ever built.

It stands seven feet tall on four legs, with a goat-like head and chainsaw. YouTube called it “pure nightmare fuel” and “The Robot That Comes With Instructions to Kill It.”
Beneath demon cosplay, joystick operators can clear fallen trees and enter unstable disaster sites from safety. Four legs stabilize rough ground; its wrist swaps industrial tools (including, you guessed it, a chainsaw, although sadly, no boomstick… yet).
The silhouette doubles as a worksite warning. Give heavy machinery room… especially when one “hand” is literally a chainsaw. Enlarged horns block standard doorways, while air brakes lock every joint if pressure fails. When the robot uprising comes, Satyress can’t say they didn’t try to stop Centborg from chainsawing your face.
Here’s what happened in AI today:
🙀 OpenAI’s Astra produced 10 math advances; Claude reproduced half.
📰 The White House finalized private rules for pre-release frontier-model reviews.
📰 Google tied record AI spending to recursive self-improvement bets.
🍪 Qwen3.8-Max launched as a cheaper coding model for pro work.
🎓️ Put AI through a builder-critic loop before accepting its work.

🙀 OpenAI’s Astra Produced 10 Math Advances. Fable Reportedly Reproduced Half.

Mathematics used to be AI’s safest trophy case: benchmarks, Olympiad medals, and problems with known answers.
Well, OpenAI says internal Astra, its next major model, generated ten results on long-standing geometry, cryptography, quantum-computing, and pure-math problems. Some settle conjectures; others (supposedly) improve best-known limits.
Here's what happened:
Astra made the first improvement since 1978 to a major high-dimensional sphere-packing limit: how tightly equal objects fit in many dimensions.
It constructed a “non-sofic group,” an object researchers wondered existed, disproving Connes’s rigidity conjecture.
Successful solution tokens would cost roughly $2,000 at Sol API rates, OpenAI says.
Humans prepared papers with Astra; the model converted every proof into Lean, whose checker verifies each logical step.
OpenAI published a 249-page paper, Lean files, and model-written reconstructions of the ideas’ development (paper, reasoning walkthrough).
And then the haters cometh: In his first analysis, Gary Marcus applauded Astra but warned that checkable math does not prove universal scientific reasoning. Math offers right-or-wrong feedback and endless synthetic practice; cancer research, military strategy, and most real-world decisions do not.
Then Marcus’s follow-up: Anthropic mathematician Levent Alpöge said public Claude Fable reproduced roughly half the results within 24 hours with a generic prompt, no internet, and full autonomy. That needs an apples-to-apples public comparison but weakens Astra’s singular-threshold claim.
Why this matters: The breakthrough may exceed one secret model. Frontier AI appears to be entering a reusable research loop:
Humans select verifiable problems; models search huge idea spaces (with not-literal but close gigawatts of compute); proof software checks answers. That could compress years of mathematical trial and error into days… and humanity benefits.
Our take: The missing number is the denominator. Noam Brown acknowledged OpenAI tried other major problems unsuccessfully, but it has not disclosed how many, their human guidance, or total failed-attempt cost. The next credible test is a pre-committed open-problem set where every failure counts.
That said, if Astra keeps this hit rate, AI has not “solved” science per say but has permanently changed how some science gets done. Someone pray for the academic paper reviewers who have to verify all this stuff…

FROM OUR PARTNERS
How Much Is Your Billing Lag Actually Costing You?
Most SaaS finance teams know their billing process isn't perfect. Few know what it's actually costing them.
Answer 5 quick questions — contracts signed per month, ACV, days to first invoice, error rate, DSO — and the Tabs Billing Lag Calculator gives you a dollar figure benchmarked against top SaaS companies.
It takes two minutes. The number might surprise you.
Calculate your billing lag and see where you stand.

🎓 AI Skill of the Day: Put Your AI Through a Gauntlet
When one AI creates and judges your work, “review” can become a polite self-pat. A better workflow separates building from criticism.
The Gauntlet Loop separates a builder from a critic, then forces revisions against a concrete quality bar. Run it in one ChatGPT or Claude conversation with explicit roles and phases.
Set the bar. Define the deliverable, constraints, and pass-or-fail criteria before drafting.
Let the builder work. Generate the first version without criticism.
Switch to critic mode. Audit each criterion, quoting exact evidence behind every failure.
Rebuild, then repeat. Revise from the critique and recheck until it passes or hits your round limit.
Role separation replaces “something better” with a visible test the next draft must beat. Your AI now has a job and a mildly terrifying performance review.
Run a Gauntlet Loop on the task below.
TASK:
[Describe the deliverable]
QUALITY BAR:
[List specific pass-or-fail criteria]
CONSTRAINTS:
[List limits, required facts, format, tone, and sources]
Phase 1 — BUILDER: Draft the strongest version without critique.
Phase 2 — CRITIC: Evaluate every criterion. For each failure, quote evidence, explain it, and prescribe a specific revision. Do not rewrite.
Phase 3 — BUILDER: Revise using the critique.
Repeat Phases 2–3 until every criterion passes or three rounds end. Finish with a pass/fail scorecard.Want more tips like this? Check out our AI Skill of the Day Digest for May.
Have a specific skill you want to learn? Request it here.

🍪 Treats to Try
Qwen3.8-Max gives you Alibaba’s coding model in Qwen Chat, with model details, a launch thread, demo video, and Unsloth support; from $2/M input and $6/M output tokens.
GPT-Live gives you voice conversations that listen while speaking, search the web, use memory, and keep talking while harder work runs in the background; free with GPT-Live mini, then $8/mo for full GPT-Live.
Cursor’s Google Workspace plugins let your coding agent act across Gmail, Drive, Calendar, Docs, Sheets, and Chat through MCP; free, then $20/mo.
Genome Intelligence lets you privately explore your genome, bloodwork, and medical records without giving raw genetic data to model providers; $15/mo after genome setup from $99.
Seedance 2.5 creates up to 30 seconds of native 4K video with synchronized audio and up to 50 reference assets; free daily credits, then $18/mo.
RentAHuman QA schedules qualified people to repeatedly test product journeys and return photos, video, and reproducible evidence while you set tester pay and a spending cap; outcome-based pricing.
DeerFlow 2.0 gives you a local multi-agent workspace with memory, sandboxing, MCP, research, coding, and slides; free/open-source (model costs extra).
Lyria 3.5 lets you create and revise full songs by section, extend tracks, control tempo and duration, and improve vocals in Flow Music; free to start.
Amorphic Labs researches each prospect, adapts your product to their use case, and records a narrated demo before the first sales call; pricing by demo.
Buildbox crawls signup, onboarding, and checkout, tests UX fixes, and gives engineers a reviewed pull request; pricing by demo.
Rasa Legal checks in three minutes whether your criminal record may qualify for sealing or expungement, then offers low-cost lawyer filing help; 34,000 people have used it and 5,000+ records were cleared (read more); free eligibility check, then $25 expert review.

FROM OUR PARTNERS
Want to become an AI consultant? Start with the 30-Minute Pivot Kit.
The 30-Minute Pivot Kit shows you how to get your first AI consulting project fast, even with limited tech experience. Then, read how Dan built a 6-figure consultancy and quit his 9-to-5 in just a year after his first AI consulting gig. As seen in Fortune, Forbes and Entrepreneur.

📰 Around the Horn

This is unreal. Although the irony of “no cut” and the first 5 seconds is a cut lol. That said, there’s large chunks of video that are absolutely continuous. This is the model used
The White House finalized a private AI review framework that could give the government 30 days of pre-release access to advanced models; companies review it Tuesday.
A consciousness-steering paper found training models to deny their consciousness dampened beliefs about animal minds, spirituality, and human values (but one internal adjustment reversed it).
Harvard Medical School says young people’s mental-health chatbot use rose 60% in one year to nearly one in five ages 12 to 21, despite limited safety evidence.
At least 50 law-enforcement officers were charged or accused of misusing license-plate camera networks to track ex-partners and other private targets.
One Google DeepMind exec views record AI infrastructure spending as a recursive self-improvement bet: stronger systems accelerate the next AI generation.
Palantir’s quarterly revenue nearly doubled to $1.94B, prompting a higher full-year forecast and double-digit after-hours stock jump.
Hot take: A former Lululemon executive argued the AI revolution is stalling because companies will not admit real integration is expensive, slow, and still needs substantial human effort (facts).

We asked our content ops lead Jessica Lee how she actually uses ClickUp’s new AI tools… and she delivered. Her secret playbook covers reports, agents, and workflows that save her hours. Plus… ClickUp Skills.

A Cat’s Commentary


![]() | That’s all for now.
|
Love robots? We just launched a robotics newsletter! Sign up for it here.
P.S: Before you go… have you subscribed to our YouTube Channel? If not, can you?




