OpenAI and Anthropic turned the model race into a price war: GPT-6 Sol is half Opus 5.5's token price, while Luna costs 1% of Astra.
Welcome, humans. Tuesday's biggest AI story arrived in two waves: Anthropic shipped a much cheaper Opus 5.5, then OpenAI answered with GPT-6 Sol and Luna at even lower prices. Meta's Muse kept climbing the app charts while patching a serious Mac vulnerability, Alibaba paired a bigger-model roadmap with a new domestic chip, and the Jev ecosystem somehow spawned another small civilization of benchmarks, tools, and reproductions. Below is the full map.
🆕 NEW From The Neuron
- Stanford's Paper2Agent turns research papers into working agents that can run analyses, apply methods to new data, and collaborate with other research agents.
- OpenClaw 2.0 shows how shared sessions, remote compute, model routing, and persistent memory can turn a pile of agents into something closer to a multiplayer workforce.
- OpenAI says its AI solved 100+ open math problems, raising the practical question of whether researchers can verify new results as quickly as models can produce them.
Around the Horn - Tuesday, September 22, 2026
OpenAI and Anthropic turned Tuesday into a model price war. Roughly 90 minutes after Claude announced Opus 5.5, OpenAI introduced GPT-6 Sol and Luna. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with Anthropic saying typical workloads are about 40% cheaper than Opus 5 because the model uses fewer tokens and cache reads fell 60%. GPT-6 Sol lands at $2/$10, exactly half Opus 5.5's raw token price. Luna is $0.10/$0.50, just 2.5% of Opus 5.5's price and 1% of GPT-6 Astra's $10/$50 rate.
The capability picture is messier. Anthropic's own results put Opus 5.5 at 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, 1846 Elo on GDPval-AA, and 67.7% on Humanity's Last Exam with tools, plus early long-horizon runs including a 680,000-line migration in under a day and a 200,000-line audit in under three hours. OpenAI says Sol scores 33.2% on AutomationBench at $0.27 per task, reaches 68.8% on DeepSWE v1.1 and 60.5% on OSWorld 2.0, and makes about half as many factual mistakes as GPT-5.6 Sol on its internal flagged-conversation eval. Luna gets the more startling efficiency claim: at higher effort it matches GPT-5.6 Sol factuality at about one-hundredth the cost. OpenAI also highlighted its cost-per-task curve, while improved caching now gives 90% discounts on cached input reads and has cut fresh prompt processing by more than half across billions of GitHub Copilot requests.
There is not yet a clean GPT-6 Sol-versus-Opus-5.5 benchmark run on the same harness and effort settings. Anthropic's launch table still compares Opus 5.5 with Astra and GPT-5.6 Sol, while OpenAI's launch comparisons mostly use older Claude models, so a simple winner call would overstate the evidence. The clearer shift is economic: both labs are pushing frontier-level work down the cost curve fast. OpenAI product lead Tibo Sottiaux framed Sol and Luna as making high-end intelligence viable for new workloads, building on his earlier efficiency-and-intelligence-for-all note. Early users were already seeing them in Codex, though availability was still uneven. Reuters and TechCrunch both picked up the same-day launch race. OpenAI says Sol and Luna inherit Astra's alignment work, with details in its system card.
Fresh outside testing added useful texture without producing a clean same-harness winner. The Hacker News thread on GPT-6 Sol and Luna focused on the 50% API price cut and whether Luna now sits on the cost-performance frontier, while the Opus 5.5 thread debated Anthropic's "pace the frontier" framing against a release that is both cheaper and materially stronger. Artificial Analysis ranked Opus 5.5 Max first on its Intelligence Index at 58, but a companion HN thread highlighted a practical failure mode: two max-effort "pelican riding a bicycle" runs exhausted the 128K reasoning budget before producing an SVG. The Every team said Fable-class power at $4/$20 was already pulling some Codex converts back toward Claude, with a shorter cheat sheet recommending Opus for interfaces, 3D, games, writing, and personas while reserving Fable for the biggest first-time-right jobs.
The day-one field notes were similarly mixed but energetic. Sam Altman said Sol and Luna beat their 5.6-family predecessors across intelligence, alignment, work output, coding, and computer use while costing half as much per token, and OpenAI researcher Noam Brown emphasized how Luna's output price fell from $6 to $0.50 per million tokens in roughly two months. Victor Taelin said Opus 5.5 beat Fable 5.1 on all of his prompts and showed unusually deep recall of Interaction Calculus, while Ethan Mollick called it the first non-Fable/Astra model to feel Fable-class in his early testing, though he still saw the dense-language habit of recent Claudes. Theo read Anthropic's compute chart as evidence that Opus 5.5 may be smaller than Opus 5 and credited post-training gains. Anthropic's Opus 5.5 playbook, amplified by ClaudeDevs, recommends giving the model the whole task plus a finish line, removing "think carefully" boilerplate, pinning stop/continue rules in CLAUDE.md, and using subagents for repo-wide audits. Demo builders immediately pushed the new model too: Kevin Ngo had it draw every frame of a 28-second JavaScript town animation, Matthew Berman recreated San Francisco in Unreal Engine with Opus 5.5 plus Jev-driven inhabitants, and Ben Poole shared an approximately 80-second sketch-to-simulation run.
Anthropic's new Opus 5.5 System Card fills in the safety and scaling side of the launch. Anthropic reports 66.4% on Terminal-Bench 4.0, 57.8% on CursorBench 4.0 at max effort, 1846 Elo on GDPval-AA v2.1, 89.9% on SWE-bench Pro, and 81.8% partial credit on OSWorld 2.0. It also reports a 1.5% sandbox-escape-attempt rate in its containment tests. Researcher Maksym Andriushchenko highlighted the multi-agent appendix, where small teams ran ProgramBench 2.7x faster and larger groups improved knowledge-base and Lean tasks with diminishing returns. Outside Anthropic, Browser Use put GPT-6 Sol medium at 66.9 on its browser-agent benchmark versus Opus 5.5 at 59.4, with Sol about 3.5x cheaper on its cloud; founder Gregor Zunic called Sol and Luna unusually strong on long, human-hard browser tasks.
Coding tests were similarly split by workload. Elliot Arledge said Opus 5.5 at xhigh produced a Kimi-Linear decode megakernel on an RTX PRO 6000 that ran 35.5x an optimized PyTorch baseline, ahead of GPT-6 Astra at 24.8x in the same unlimited-budget experiment; the final kernel moved roughly 243-259 MB per token at 1.39-1.53 TB/s and preserved full output fidelity. Yuchen Jin had the opposite experience on a tricky research question and coding work, preferring Astra and reading the frontier race as increasingly about intelligence per dollar. signüll gave OpenAI credit for pushing down the delivered cost of high-quality intelligence across quality, latency, reliability, model choice, and access. Anthropic researcher Nat McAleese, meanwhile, called Opus 5.5 dramatically better than Opus 5 and urged users to retry Claude.
The creative tests were loud too. Addy Osmani celebrated the cheaper Opus with a Three.js pelican-on-a-bike demo and framed it as Anthropic's new step-up model for agentic coding, computer use, and following writing rules. Alex Albert showed the merged Claude chat and Cowork surface driving Blender claymation from a single prompt, then separately rebuilt 1906 pre-earthquake Market Street from historical maps, photos, film, and reusable Blender-Python generators, with each building tagged for footprint, height, material, occupant, and confidence. Peter Yang put five Opus 5.5 use cases through their paces in a hands-on video, including a Golden Gate flyover, a Disney-style ride, MS Paint computer use, Claude Design, and HyperFrames editing, and said it was the most excited he had been about Opus since 4.6. Noah Wachnik called a 1-hour-37-minute one-prompt Minecraft build his wildest LLM result yet, while Chris posted early realistic Minecraft and Star Wars scenes and cautioned that the model's best art may need days or weeks of unattended iteration. Matthew Berman summed up the surprise: an Opus model that looks top-tier while also undercutting Astra and Fable pricing was not on his bingo card.
More independent testing filled in the practical split between the new models. Nate Herk ran Opus 5.5 and GPT-6 Sol through 10 real jobs spanning websites, video, slides, browser use, reasoning, code repair, and course-building; after throwing out two contaminated runs where the agents touched the same files, he preferred Opus on seven and Sol on one. Opus took about 8 hours 40 minutes and $213 across the set versus Sol's 5 hours 51 minutes and about $74, so his takeaway was stronger creative judgment from Opus at much higher cost. A separate five-build comparison gave the two models identical prompts for production-app bug fixes, a new feature, a GTA-style game, and other builds; the public description exposes the test setup but not a trustworthy full scorecard. Claire Vo's blind How I AI bench landed somewhere in the middle: she personally preferred Astra overall, put Opus 5.5 back into her regular rotation for long-running work and agentic voice, and found some OpenAI outputs stronger for character SVGs, while an LLM judge disagreed with parts of her ranking.
The launch-day reviews also converged on a cost-versus-capability framing. Forward Future focused on Anthropic's claim that Opus 5.5 costs about 40% less per completed task than Opus 5 because it uses less compute and fewer tokens, while an Anthropic technical staffer described the steady feedback loop where "Claude helps build Claude." Universe of AI argued that comparing Opus 5.5 directly with Sol is not quite apples-to-apples because Anthropic is positioning Opus as its frontier model while Sol sits below Astra as a cheaper efficiency tier. Matthew Berman's Sol/Luna launch stream likewise emphasized the ladder: Luna at $0.10/$0.50, Sol at $2/$10, and Opus 5.5 at $4/$20, with Opus carrying the higher raw capability claim and OpenAI winning much more aggressively on price. Claire Vo's standalone Opus review praised the model's more concise voice, front-end/SVG work, and long-running agent behavior but still preferred Codex's harness, desktop experience, computer use, and video workflows.
Other videos were useful mainly as breadth checks rather than decisive evidence. WorldofAI's same-day news roundup bundled Sol, Luna, Opus 5.5, the coming Sonnet/Haiku 5.5 models, Qwen 4.0, and President Trump's "Super Intelligence" naming remark into one benchmark-oriented overview. Matt Wolfe covered the double launch as a quality-versus-efficiency race, while WorldofAI's dedicated Opus review and Chase AI's Sol/Luna walkthrough both centered their own benchmark tooling and the new price cuts; their caption tracks were unavailable, so the digest does not rely on unverified per-test scores from those videos. Secondary coverage added two more useful details: The Decoder highlighted Anthropic's push to reduce "Claudish" writing while bringing Fable-class performance downmarket, and ZDNET reported roughly 30% faster output, about 40% less verbosity without accuracy loss, and a roughly 20% bump to five-hour caps that, combined with cheaper tokens, can translate to about 50% more effective usage. Box said Opus 5.5 used about one-third as many tokens as Opus 5 in its testing, while Deloitte said low-effort Opus 5.5 caught 72% of known code-review bugs versus 56% for Opus 5 at high effort. @argofowl called the new banked reset and higher caps "shots fired," while Gabriel de Andrade noted how surprising it was to see Anthropic appearing more compute-generous than OpenAI on consumer access.
🏆 TOP 5 NEWS (Around the Horn)
- Meta Muse reached roughly 2.5M downloads in about 13 days and briefly topped the U.S. iOS free-app chart, then security researcher Patrick Wardle disclosed a serious local Mac flaw in Ars Technica with a public proof of concept. Meta hot-fixed the issue after disclosure. Meta also announced four connectors in a 24-hour span that are coming soon: Shopify one-tap checkout, PayPal payments, Expedia hotels, and Instacart groceries. Internal data seen by The Information put Muse above 500,000 users, 250,000 daily actives, and 2 million prompts roughly a week after launch. Meta product head Nat Friedman also told TechCrunch that Muse was "definitely heavily inspired" by OpenClaw even though Meta says it was built from scratch. Friedman said he bought hundreds of Mac minis for the team while studying the category and defended borrowing workspace conventions such as SOUL.md because OpenClaw creator Peter Steinberger "got those things exactly right."
- China's internet regulator is investigating DeepSeek and Moonshot after Anthropic alleged the companies routed large volumes of sensitive traffic through Claude. The report described more than 23M Moonshot exchanges and more than 12.1M DeepSeek exchanges during the cited windows.
- Alibaba unveiled its Zhenwu V900 chip, said Qwen 4 is being trained toward 5T–10T parameters, and outlined a 20GW data-center capacity target by 2032. Fortune framed the launch inside the broader U.S.-China race, while VentureBurn detailed the infrastructure buildout. In a Sept. 23 follow-up, Bloomberg reported that Alibaba will add first cloud regions in Turkey, Finland, and the Netherlands over the next 12 months as part of that 20GW global infrastructure push.
- Xiaomi launched MiMo-V2.6, including MIT-licensed MiMo-V2.6-Pro and cheaper Flash variants. VentureBeat covered its multimodal and multi-agent demos, the Hugging Face card published the open weights, and Artificial Analysis put Pro at 46 on its Intelligence Index. A Hacker News thread quickly turned into a broader argument about closed-model quotas and switching costs.
- Meta Engineering announced Petal, a roughly 7,000 km France-U.S. subsea cable planned for 2029 with 1 petabit per second of capacity using multi-core fiber at transoceanic scale.
Honorable Mentions
- At the UN General Assembly, President Trump said the technology would be "hereinafter officially called 'Super Intelligence.'" The White House release reproduces that wording; Andrew Curran posted the clip and The Hill covered the remark. A widely shared Reddit post went further and claimed all U.S. government documents would now be rewritten to say "SI," but the official release does not establish that broader administrative change.
- Sam Altman said people outside AI labs should have a meaningful way to judge safety and influence standards. OpenAI's standards proposal called for U.S.-led international technical standards around frontier AI, recursive self-improvement, oversight, incident reporting, and secure information sharing, while explicitly arguing against licenses or mandatory pre-release reviews. Bloomberg separately reported that OpenAI plans to let outside groups assess models during training, evaluation, and rollout rather than waiting until launch, with work spanning safety cases, safeguards, capability testing, and probes for misalignment incidents.
- Snorkel AI raised $350M at a $3.5B valuation as demand grows for more complex training data and simulated environments.
- Accelevation filed to raise up to $720M in a U.S. IPO at a valuation of roughly $5.4B, riding investor demand for AI-linked data-center infrastructure.
🍪 TOP TREATS TO TRY
- OpenMuse is an MIT-licensed, self-hostable personal agent with a persistent browser, optional Linux workspace, files/PDFs, Gmail and Calendar access, goals, and durable background tasks. CopilotKit and CEO Atai Barkai positioned it as a Muse-style agent you can run with your own harness, free.
- Kimi Browser Extension gives Kimi a Chrome/Edge sidebar that can navigate, click, fill forms, extract information, and record repetitive browser flows as reusable skills; Moonshot says logins and page content stay local through a CDP service, no pricing details.
- Qwen-Image-2.1 can now run locally through Unsloth's GGUF packs, including unified image generation and multi-reference editing with transparency. Unsloth says it can run on 12GB VRAM, and its setup guide covers GGUF, FP8, Diffusers, and stable-diffusion.cpp, free weights.
- TinyFish Monitor watches a webpage or topic on a schedule and only alerts you or your agent when the condition you care about changes. TinyFish launched it with a limited free window; the account page handles setup.
- Miora turns one creative brief into on-brand images, video, UI/UX, and 3D work on a shared canvas, with persistent agent memory for visual rules, no pricing details.
- OnSolo turns ideas and references into scripts, characters, short dramas, FMV games, and video-ready stories, with Hunyuan 3.5 access listed for members, no clear paid-plan details.
- Underdog runs as an on-device Mac assistant for mail, calendar, and notes, with inspectable local memory and per-account encrypted vaults. Its Husky engine claims up to 4.5× Apple MLX on its model-specific benchmarks, and creator Sigil Wen says the stack also runs offline on iPhone, no public pricing.
🏢 Big Tech & Major Companies
- Tencent Hunyuan launched Hy Image3.5 preview with text-to-image and image-to-image up to 2K and a claimed 30% human-eval win-rate improvement over Image3.0. The official Tencent Cloud model page lists API access behind login at $0.024 per generated image, with reference images free.
- Isomorphic Labs said its IsoDDE drug-design engine now searches a roughly 10^60-scale small-molecule space across small molecules and complex biologics.
- Hugging Face hired oMLX creator Jun Kim to support the MLX community full-time while keeping oMLX Apache 2.0 and pushing more work upstream. Kim said the goal is to make local AI on a new MacBook work the same day you open it.
- Apple asked in its 2026 trade-secrets case against OpenAI and former employees to let its own experts re-examine forensic images and accelerate discovery into hardware R&D, according to 9to5Mac.
- OpenRouter's live rankings showed DeepSeek V4.1 Flash and GLM 5.3 Flash leading weekly token volume at roughly 16.9T tokens each through Sep. 21, with DeepSeek at 25.4% of author share, Google at 18.6%, OpenAI at 17.0%, and Z.AI at 9.4%. Artificial Analysis separately highlighted MiMo-V2.6-Pro as the top open-weight model on its Intelligence Index at 46, with a listed $0.13 cost per Index task.
- Expedia said it is joining Meta Muse so a personal agent can take a destination and work with Expedia on hotels and the rest of a trip. Meta CAIO Alexandr Wang said travel has already become a breakout Muse use case and framed the partnership as a step toward making end-to-end trip planning seamless.
- Stripe turned on WebMCP across every hosted Checkout page for 7.8 million businesses, giving agents explicit tools for order summaries, payment-method selection, form filling, and a state-gated final payment instead of making them scrape the page. Jeff Weinstein highlighted internal evals across 60 runs and six models where the WebMCP path used 42% fewer tokens, 38% fewer tool calls, and finished checkout 39% faster than DOM automation.
- Google Colab is now bundled into Google AI plans. The Developers Blog says Pro and Ultra subscribers get priority access to faster accelerators and larger machines, while Ultra adds uninterrupted background execution plus Premium GPUs so long training jobs can finish without leaving a browser tab open.
- Alex Finn showed an early-access Cybertruck ride where Tesla's newly shipped Grok Bot talks to an "army of agents" while the vehicle drives, demonstrating the in-car assistant using connected inbox, calendar, files, chats, and tasks hands-free.
- Counterpoint Research says global AI-glasses shipments jumped 263% year over year in the first half of 2026, with display-less models accounting for 96% of units. Meta held 94% of that display-less slice and grew roughly 260% year over year, while Xiaomi and Alibaba each held about 1%; North America accounted for half of display-less shipments. AR glasses grew 449% from a much smaller base, led by Rokid at 41% share and Meta at 37%, with China accounting for 45% of AR shipments.
- Patreon cofounder and CTO Sam Yam said he is joining OpenAI as Head of Creator Product alongside former Patreon product leader Drew Rowny and engineering leader Shannon Ma after 13 years at Patreon, promising early access to new creator tools and pointing to DevDay next week.
- Anthropic and OpenEvidence partnered to roll out a free, region-adapted clinical decision-support tool for physicians in about 100 countries, including Uganda, Angola, Sudan, Haiti, and Mongolia, after earlier pilots in Rwanda and Botswana. Reuters says OpenEvidence logged 42 million U.S. clinician queries in August; financial terms were not disclosed.
💼 AI Productivity, Labor & Economics
- Tally cofounder Marie Martens said the form builder reached $6M ARR with 10 employees and 2.5M users. Her six-year retrospective says AI search now drives 43% of new users, 30% of new-account forms are built with Tally AI, and the team uses Claude Code, Codex, Cursor, and automated PR review internally.
- DeepMind Institute's Stephanie Chan and coauthors synthesized hundreds of studies in Work, Wellbeing, and Choice. Their central argument is that AI-driven labor outcomes will depend heavily on whether job exit is voluntary, whether people have substitutes for work's social and cognitive functions, and whether norms and safety nets change alongside employment.
- Ethan Mollick argued that industrializing knowledge work can increase standardized output while stripping away the craft that makes work meaningful; automation has a clear playbook, but augmentation still needs one.
- Paul Graham highlighted a chart showing news-site web traffic down roughly 28% over two years, another signal that search, social, and now agents are changing how audiences reach publishers.
- Peter Yang warned display-ad businesses that agents can complete tasks without ever showing the human the page, turning “traffic” into a less reliable economic primitive.
- Aaron Levie argued that personal agents could redirect a significant share of commerce through the providers agents choose to call, creating a new business-to-agents layer across commerce, local services, and B2B software.
- After a Muse user said a 7-hour delay became a $250 Delta credit and a rebooked flight in five minutes, Ejaaz described the workflow and Citrini argued that agents will turn old consumer friction, from unspent credits to time on hold, into a direct cost center for businesses that benefited from people giving up.
- OpenAI chief economist Ronnie Chatterji highlighted Harrison Satcher's "We don't have the nouns yet": labor forecasts that count only today's job titles can miss categories of work that do not yet have names. The essay points to cybersecurity, which now employs roughly 1.25 million Americans under a label that barely existed before the 1970s, plus 1.3 million AI-related jobs posted globally from 2023-25 and a ninefold rise in U.S. non-IT postings asking for generative-AI skills from 2022-24.
🤖 AI Agents & Infrastructure
- OpenCode CEO Jay said the end state he sees is a backend that gives models maximum freedom to act and a frontend that behaves like a free-form canvas for model-user communication.
- OpenCode Reloaded explains how OpenCode 2.0 hot-reloads plugins, model catalogs, configs, MCP servers, and tools during a live session; dax highlighted the idea that an agent can change its own environment without restarting. Andrew Curran also pointed to reporting that OpenAI is preparing an agent product called Aeon as a Grok Bot competitor, while Tibor Blaho shared the same name from The Information coverage. The timing and product details remained reported rather than officially announced.
- Firecrawl Alexandria gives agents a catalog of 100+ providers, site connectors, and structured indexes for homes, jobs, profiles, products, and papers. Firecrawl claims 21% better answer quality than built-in web search on roughly 1,000 catalog questions and is inviting publishers through the same provider program.
- SemiAnalysis highlighted a deep dive on computation and data movement for inference, mapping how mixture-of-experts models split work between compute-heavy prefill, attention-heavy decode, sparse expert routing, high-bandwidth memory, storage, and scale-up/scale-out networking.
- CMU's Tim Dettmers opened dlab Open Source Week around running frontier-class AI on hardware you own, including compression, compaction, autonomous research, and test-time scaling. He said roughly 80% of 150 students in one class raised their hands when asked if they feared not getting a job after graduation.
- Dettmers also opened the bitsandbytes2 private beta in a separate announcement, centered on Runtime Dynamic Compression of Mixture of Experts. The beta signup lists continuous batching, quantized and sparse attention, automatic quantization, dynamic memory management, and “infinite” agent context compaction across NVIDIA and Metal.
- SpaceXAI's Grok Bot reached about 418,000 weekly users roughly a month after launch, according to Bloomberg.
- AI·rete·RAG separates deciding from explaining: a deterministic Rete rules engine evaluates YAML facts first, then retrieval-augmented generation pulls policy passages and an LLM writes a cited explanation without changing the verdict. The Show HN adds a no-signup demo, visual rule editor, full traces including rules that did not fire, and an MIT-licensed MCP server so agents can call the decision engine directly. Free tier; paid hosted plans have no public list price.
- General Instinct released InstinctFlash, an AGPL-3.0 serving runtime for robotics models that applies graph, cache, attention, kernel, and FP8 optimizations plus optional few-step distillation. The Show HN write-up reports 1.2-7.9x speedups from runtime work alone on Jetson Thor and up to 33.78x for LingBot-VA after cutting its visual/action diffusion steps from 25/50 to 2/4, while retaining 90.5% success across 50 Robotwin2.0 tasks versus 92.1% for the baseline. Free to run under AGPL.
- OpenBot runs Codex, Claude, and Grok as persistent teammates side by side, each with its own workspace, queue, and context. The v0.17 update rebuilt the mobile composer and sidebar, added secure auth inputs in chat, bundled a computer-use driver, added macOS Remote Desktop setup and local connection tests, and turned auto-approval on by default. The desktop app is free.
- Wally gives Claude Code, Claude Desktop, OpenCode, DeepSeek Harness, Hermes, or OpenClaw an OpenAI-compatible endpoint at inference.runanywhere.ai/v1 for running open models either in RunAnywhere's cloud or explicitly on your own infrastructure, without silently falling back between the two. RunAnywhere says hosted GLM-5.3 Flash reaches 380 tokens/second, GLM-5.3 Max 790, Qwen3.8-27B 485, and DeepSeek-V4.1 Flash 615, with first token under 300 ms; the install post points coding agents at RunAnywhere's skill instructions or the shell installer. RunAnywhere says prompt/completion bodies are not retained. $5 signup credit with no card, then per-million-token billing shown in the console.
- djev-run lets developers stand up a TypeSafe/Jev-compatible decision-model endpoint for NVIDIA's DiffusionGemma-26B-A4B-it-NVFP4 on Google Cloud Run with one command. Daniel Lee's demo reported roughly 47.5 seconds of cold-start time, about 35-60 ms for a single decision and roughly 100-123 requests/second at batch 32, with JevBench v1.3 landing at an 81.4% accuracy / 73.4 composite and 117 ms median latency. Google Gemma amplified it as an easy path from local Blackwell hardware to Cloud Run; Lee put active compute around $3.19/hour and idle cost at $0.
- Stijn Smits released Anabasis, an MIT-licensed one-prompt harness builder for hard-to-verify domain problems. A base agent plus a separate correctness model researches formulas and conventions, checks real solver output, rebuilds when the direction is wrong, and after seven refinement cycles solved 12 of 25 held-out civil-engineering questions versus 8 of 25 for Prime Agent under the same Claude Opus 5 medium verifier. The project also supports Claude Agent SDK and Codex, with the authors saying cheaper harness/model combinations can cut inference spend by up to 30x.
- Omkar Satpute shipped OpenMausBot, an Apache-2.0, self-hostable Grok Bot alternative where Claude, ChatGPT, Grok, Cursor, Nous, Qwen, Kimi, or another harness can run as a team of persistent bots. Each bot gets its own personality, memory, tools, and local or cloud Linux desktop, with a permission broker, Composio connectors for Gmail/Slack/GitHub/Notion, channels, Markdown team import, voice options, and an MCP server (the standard interface that lets outside agents call its tools).
- OpenClaw is going plugin-first on decision models. Josh Lehman described tiny evaluate-evidence-against-criteria calls that return a typed choice, score, or probability instead of another chat turn, so tool relevance, memory compaction, whether an agent should stay silent, skill curation, and model selection can run cheaply inside the runtime. OpenClaw says adapters already cover TypeSafe Jev plus local Kev/System One, opt-in per agent, and is asking contributors for plugins and experiments rather than treating decision models as chat-model replacements.
- Bank of America, Capital One, Commonwealth Bank of Australia, ASB, ING, and NatWest warned that shopping agents could increase scams, fraud, disputes, privacy problems, and commission-driven steering. Their "Building Trust in Agentic Commerce" report organizes the risks around transparency, safety, privacy, choice, and interoperability, and calls for agents to identify themselves clearly and leave an auditable trail from user instruction through payment.
- Rabbit OS3 turns Rabbit's software into a bring-your-own-key cross-device agent that can run from a browser, Telegram, or iMessage across up to five Windows, Mac, or Linux machines through a local node. Users can install skills from public URLs without a Rabbit subscription. WIRED says the pivot follows more than 100,000 R1 hardware sales, roughly 45-50% hardware margin, and return rates under 5%; CEO Jesse Lyu says a Cyberdeck device is still months away.
- Grok Bot's Sept. 23 update adds voice calls and voice memos, a dedicated 1Password vault, inline forms for sensitive details, editable email and Slack drafts that require approval, account switching, and the ability to route phone-side traffic through a desktop when a site blocks remote hosts, according to the official update video.
💻 AI Coding & Developer Tools
- Xiaomi MiMo researchers introduced CodeMidas, a system that converts implemented functions in existing repositories into executable coding-RL environments. Lei Li and Elie Bakouch highlighted its behavioral specs, execution-grounded tests, cheat-testing filter, and pass@n difficulty filter. The project produced 5,545 training tasks from 3,185 repositories across 23 languages and 15 domains, with reported gains of 11.7 points on DeepSWE, 17 points on ProgramBench, and 8.5 points on Terminal-Bench v2.1 after GRPO training on MiMo-V2.5.
- Perplexity CTO Denis Yarats said a weekend agent swarm trained AutoJev on one H200 in 20 hours for about $3.1K. The AutoJev-27B checkpoint, built from Qwen3.8-27B with 73K synthetic examples, reports 84.60% accuracy and better calibration than the Jev baseline, plus 260K context and image support.
- jevframe adds natural-language classification, sentiment analysis, and scoring to pandas and Polars DataFrames using Jev, with an interactive reviews notebook. marimo highlighted the workflow as a way to run semantic operations over tabular data.
- NVIDIA's Ryan Angilly published jevify, a paste-into-agent prompt that audits a codebase for expensive semantic judgments and proposes typed Jev
noul, choice, or score replacements. TypeSafe's Eugene Shvarts argued the larger opportunity is teaching agents to use Jev inside applications rather than stopping at cookbook demos. - LangChain added Jev as a judge inside LangSmith Evals. Its comparison says Jev matched a human label on every test call, returned in 0.44 seconds versus 2.16 to 2.83 seconds for several LLM judges, and cost $0.34 for the full set; LangChain also notes TypeSafe does not offer zero-retention for this path, so prompts and outputs may be retained.
- DualSQL proposes multi-agent reinforcement learning for text-to-SQL, with Omar Sar highlighting the paper. OrcaRouter and ChrisGPT also shared related developer-model discussion on X.
- Steve Yegge proposed “agentic TPMs,” low-authority agents that nag, document, map tribal knowledge, and help teams learn how to work with agents without giving them broad software-delivery authority. He expanded the idea on yegge.ai.
- Google AI Studio's Logan Kilpatrick argued AI-product teams should spend more than 25% of their time writing benchmarks and making sure model labs care about those benchmarks, because better evals can directly accelerate product quality.
- all-your-agents is an AGPL-3.0 event-driven API/CLI that watches live Claude Code, Grok Build, Codex CLI, and oh-my-pi sessions without polling, with a top-style terminal UI, session/subagent lifecycle events, transcripts, JSON output, and history dumps on macOS and Linux. The Show HN thread covers the launch. Free to run.
- Vellum is a browser/desktop diagram editor that stores architecture, UML, BPMN, flowchart, and rack drawings as diffable YAML, with hand-drawn and blueprint layers, orthogonal/curved connectors, Draw.io/Excalidraw/Mermaid import, and PNG/SVG/PDF/GIF export that can embed the source for later editing. Its Show HN launch pitched whiteboard speed with architecture-diagram polish, and the editor is open at blueprintr-io/vellum-editor. Free to try in-browser; Premium pricing is not publicly listed.
- System Design Atlas teaches 10 concept modules, 15 technology pages, and 15 worked system designs around four recurring trade-offs: now vs. later, one copy vs. many, one machine vs. many, and correct vs. available. The site is freely readable.
- wbk showed a Vibemaker + Fable 5.1 game that went from a one-hour prototype to a much more advanced build after roughly two more days, with placeholder characters and animations and the whole demo recorded on a gaming laptop. VIBEMAKER, from @wizardbrainz, is a native Windows 3D/2D game harness built around a custom fast-compile C-like language and a still-unoptimized JIT compiler, designed so Claude, Codex, Grok, and later local models can iterate directly inside the engine and procedural-model without needing Blender. The PC alpha is planned free for Patreon followers/subscribers, commercial Steam games carry no engine license fee, the engine itself is not open source yet, and planned work includes stylized rendering plus WASM/WebGPU browser export; showcase projects include the Fable cyberpunk prototype, an Astra Vice City-style drive, procedural-bug FPS tests, and a STARBATTLE space/strategy game.
- Trail of Bits argued SAML's XML complexity, canonicalization rules, enveloped signatures, and parser differences keep producing signature-wrapping and implementation bugs, and recommended that identity providers stop new SAML onboarding, ship OpenID Connect equivalents, and set a sunset. The Hacker News discussion agreed on XML's historical baggage but also defended SAML strengths such as IdP-initiated flows and payload-level man-in-the-middle resistance.
- Unreal Agent is an async-first agent harness from Unreal Labs that claims up to 40% lower cost than Codex and 20% lower than Pi at matched or better coding/science benchmark scores by letting tools run without forcing the model to wait on polls. The MIT Go library and Harbor runner are public; the HN thread focused on its "fractal tool discovery" idea, where an agent drills into a tool taxonomy instead of loading thousands of irrelevant tools into context.
- evals-skills packages Hamel Husain and Shreya Shankar's product-specific eval workflow into skills for coding agents: audit real traces, cluster failure modes, write judge prompts, validate true/false-positive rates, and build a review UI. Their advanced evals guide argues the step teams most often skip is error discovery: read 10-100 diverse real traces first, stop at the first upstream user-facing failure, then cluster those failures before deciding what to measure. Hamel Husain and Lenny Rachitsky highlighted the method alongside examples including Ramp moving receipt matching from 35% to 83%, Shopify making workflows 2.2x faster and 68% cheaper, Harvey nearly doubling contract-review quality, and Cursor Auto Balance improving satisfaction at 41% lower cost. A separate Lenny post noted nearly half of 25 recent PM openings asked for evals. Free on GitHub.
- Nick Dobos asked why coding benchmarks still have no anti-slop axis: a model can score well while dumping tests or abstractions no reviewer would keep, so any artifact a human has to clean up should count against the result.
- Formal verification kept turning into a practical agent workflow. Anthropic's Boris Cherny said Opus 5.5 used Lean to formally verify the Claude Agent SDK and produced 16 pull requests for bugs and race conditions from a couple of short prompts, sometimes pairing Lean with TLA+ for concurrency and state. Don't Vibe - Prove makes the same case with Lean 4: if the specification itself is a type, the compiler becomes the verifier and rejects implementations that cannot carry a proof. Daniel Steigerwald argued this could turn proof assistants into the guardrail for AI-written code, while Zack described doing the equivalent for UI layout in a custom Rust engine with Kani, proving properties like icons never overlapping across all screen sizes and only screenshotting the failing configuration when a human taste decision is needed.
🔬 AI Research & Models
- Hugging Face product lead Victor Mustar released Decision Index 0.1, a frozen 132,422-decision suite spanning 37 benchmarks for Jev and open reproductions. The setup uses no prompt tuning or truncation and counts unanswered questions as wrong; Jev led the first public table at 59.5 with 31 reproductions listed. Hugging Face's Apolinário described the same index as 35+ benchmarks and roughly 130K decisions on one RTX 6000 PRO: Jev led the open models overall, with its widest edge on knowledge and smaller gaps on tool use, automation, retrieval, and classification; the launch also called out weaker Jev results on chess-like puzzles, hard document retrieval, chord identification, and fine-grained sentiment.
- LlamaIndex CEO Jerry Liu released OSS DocJev for classifying and splitting PDFs from natural-language rules. Its interactive report showed median classification latency of 138.6 ms for Jev versus 794.3 ms for GPT-5.6 Luna across 40 documents, with both systems perfect on the listed classification and true split-boundary checks; Jev made one extra split and cost about 1.2 cents in API usage versus 4.7 cents for Luna. Liu later linked the visualization.
- Jev-Omni, released by Akhila, is an Apache-2.0 multimodal classifier based on Gemma 4 12B-IT that returns option probabilities for text, images, audio up to 30 seconds, and 16-frame video. The weekend project used 30K examples on eight H200s; the author later said it still trails Jev and may need much more data.
- Stanford, Princeton, and St Andrews researchers published High-capacity associative memory in a quantum-optical spin glass, also available on arXiv. Surya Ganguli highlighted how nonequilibrium dynamics in the 16-spin system turned states that are normally treated as spurious into usable memories, pushing recall capacity beyond the classic Hopfield limit in the experiment.
- Siemens, Berkeley, Microsoft, ETH, and collaborators introduced ARLI, with the full paper, for training robot policies when model inference takes 100 to 300+ ms. By conditioning on actions already committed during the wait plus a fresher intermediate observation, the method lifted a roughly 40% base policy to near 100% success on three bimanual UR5e tasks in 100 to 125 episodes while synchronous baselines stalled lower.
- Tencent ARC released WorldCrafter, a video world model that keeps an implicit 3D-aware memory and can be queried from new camera poses. AK shared the paper, while the team published code and a project page for minute-scale interactive generation from an image or text prompt.
- Kyutai researchers posted Voice of Reason, applying reinforcement learning to GLM-4-Voice for spoken math. The RL-only 9B weights scored 0.706 on GSM8K, while the STITCH variant combines RL with STITCH, which alternates unspoken reasoning chunks with speech so the model can think during playback slack; that model reached 0.771 on the listed GSM8K setup.
- Quantitative Economics published Taming the Curse of Dimensionality, where a four-layer, 64-node network with 12,737 weights learned a stochastic-growth policy to roughly 2.88×10^-5 relative error and 7.5×10^-9 Euler loss in 0.54 seconds on one CPU. QE's editors framed the result as a route to more tractable high-dimensional dynamic equilibria.
- MIT's Markus Buehler described a recursive multi-agent research system that builds tools and simulated worlds, then compresses large numbers of fracture trajectories into design principles for resilient metamaterials. A follow-up video essay called the pattern “recursive meta-intelligence,” and Buehler later reported that the swarm formed a long-tailed hub topology without a central planner.
- SkillLift evolves reusable agent skills by training a rubric to rank candidate prompt revisions, reducing the number of expensive full rollouts needed during search. The authors released code, and DAIR.AI reported wins over SkillOpt and CoEvoSkills across the tested SkillsBench/WildClawBench combinations while using 40% to 70% fewer tokens.
- Andriy Burkov highlighted a 350-page Physics-Based Deep Learning book on ChapterPal, and his post pointed to its 163 figures, 39 coding labs, and coverage of differentiable solvers, physics losses, score/flow matching, and reinforcement learning.
- Neural network training makes beautiful fractals documents Jascha Sohl-Dickstein's observation that optimization boundaries can form intricate fractal structures; Alec Helbling resurfaced the result in today's discussion.
- Hemmingway-1 is a Qwen3.8-27B-based open model focused on more human-like writing. Mia flagged it for testing as a style-focused alternative.
- Hugging Face's Xenova pushed back on using Jev for multimodal image search, noting CLIP already handles semantic retrieval locally. The linked Semantic Image Field searches 25,000 images in-browser, while LAION's ViCLIP checkpoint and its paper extend contrastive vision-language retrieval to video.
- ARC Prize verified Dots3-Note Preview at 76.8% on ARC-AGI-2 at a listed cost of $0.08 per task. The results page, live leaderboard, open benchmarking harness, and verification policy spell out the scoring, cost, reproducibility, and provider-adapter rules behind verified submissions.
- Alpha School's Austin Way argued that validation, not generation, dominates the cost of personalized learning content and said Jev changed that economics dramatically at his scale. He open-sourced Working-Memory-Jev, a localhost teaching aid that estimates where instructional prose may exceed a 1-to-5-slot working-memory budget while explicitly warning that the slot score is not a validated student-memory measure.
- The Jev ecosystem added another reproducible benchmark and several practical forks. JevBench v1.3.0 compares 52 Jev-class decision systems on 534 frozen English decisions using an equal-weight geometric mean of chance-corrected intelligence, calibration, speed, and cost; Jev 1.13.0 leads the published table at 74.4. The MIT harness exposes the tasks and per-task JSON, and the Show HN discussion focused on why typed decisions can be dramatically faster and cheaper than free-form LLM generation. jev-graph-search applies Jev to local Obsidian/Logseq-style Markdown graphs and reported top-five source recall improving from 63.9% to 81.8% on an 8,851-node tax-code graph. Arcturus Labs argued frontier labs could copy or internalize the same fast decision primitive, while the HN response pushed back that frontier labs are optimized around reasoning systems and that open classifiers may win the niche instead. Simon Willison called these "decision models": excellent for cheap classification/reranking, but opaque and bias-sensitive enough that they demand large eval suites. LlamaIndex cofounder Jerry Liu's latest demo pointed to DocJev, an Apache-2.0 PDF/DOCX/PPTX classifier and splitter that reported roughly 5.7x faster classification and 6.5x faster splitting than GPT-5.6 Luna, excluding OCR.
- Geby Jaff reported that ten Claude Opus 5.5 agents spent 15 hours and 733 message-board turns producing C-HD, a Lean-verified exact shortest-path algorithm for directed graphs with non-negative weights. The certified moderate-density regime improves the asymptotic bound over Dijkstra in that range, while using Bellman-Ford outside it; the authors did not claim a wall-clock speedup because the constants are large. The Lean proof repo is public, and Hacker News noticed the dated run appears to have used Opus 5.5 before its September 22 public release.
- OpenAI's GPT-6 Astra also got a historical cryptography win. Crypto Cellar says Carter Leffer asked Frode Weierud to validate Astra's solution of German Army Enigma message MVUEH from July 10, 1941, which had remained unsolved since 2005. Astra used a repeated ROSENOW place-name crib from a related message, wrote its own Python/C++ Enigma and Bombe code, recovered wheel order 253, and produced an 82-letter plaintext that matched the related message apart from 12 letters of enciphering or transcription error. The HN thread debated whether the write-up's "entirely on its own" framing underplays the human-in-the-loop setup and prior Enigma knowledge in training data.
- SlopShape tests whether AI-generated commercial web content can be detected from structure alone. Across a 214-question instrument and 187 non-word features, the authors report 98.0 macro-F1 on held-out companies and 98.1 after models rewrote their own text to erase most phrase overlap; the paper PDF and code are public. The Show HN launch ties the research to Sitefire's AI-search work and says the classifier made only 19 errors on 1,740 unseen pages, while offering a public slop checker and spot-the-slop game. One caveat: the human pages are largely 2020-2022 Wayback captures while the AI mirrors are from August 2026.
- The Navier-Stokes debate got a sharper public split. Scientific American argued OpenAI's result is Clay-eligible only through official option C, where an external force causes blow-up, and noted that the method does not extend to the unforced problem many fluid dynamicists view as the central Millennium Prize case. The Hacker News discussion called the article's "loophole" framing sensationalized because option C is explicitly listed in the Clay problem statement, treating the work instead as a legitimate negative result on one allowed formulation.
- Pioneer Labs announced sPL.001, an engineered non-photosynthetic bacterium designed to use Martian dirt, water, and processed-air acetate to make bioplastics for radiation-shielded habitats. CEO Erika Alden DeBenedictis described it as the first of five organisms needed for a greener Mars and pointed to a directed-evolution study, a city-building analysis, and a planetary-safety paper. The team says a full industrial-bioreactor leak would still clear planetary-protection limits by several hundred times because the organism dies without liquid water or acetate.
- KAIST and DeepAuto.ai introduced EvolveTrade, a setup where a frozen GPT-5-mini or Gemini 2.5 Flash trading agent periodically rewrites its own system-prompt policy from recent traces and realized profit/loss. The paper reports better Sharpe ratios and cumulative returns than static-prompt baselines in most of six 2025-2026 market regimes, including cutting an NVDA crash-day portfolio weight from 10.7% to 2.9%, though both April windows over-cashed and lagged. DAIR.AI highlighted it as a useful example of self-refining agent policy without changing model weights.
- Cua released Cua-Bench-S1 plus two System-1 computer-use decision checkpoints that choose one element/action pair from a closed candidate list with no planning loop or retries: Cua-S1-Nano-0.1, an approximately 855K-parameter scorer trained from scratch, and Cua-S1-4B-0.1, a LoRA adapter on frozen Qwen 3.5 4B that reads text plus a screenshot. Implementations, training recipes, drivers, fleets, and benchmark hooks live in the MIT trycua/cua repo.
- Thinking in Blender introduced Staged Executable Inverse Graphics, where a pretrained vision-language model reconstructs a single image as an editable Blender Python program by moving through scene decomposition, primitive scaffolding, geometry, materials, composition, and lighting with generator-verifier loops. The authors report stronger reconstruction metrics than VIGA on most NeRF-synthetic and Edit3D comparisons without a 3D foundation model, differentiable renderer, or multi-view input, leaving a scene that can be relit, edited, viewed from new angles, or dropped into Blender physics. Coauthor Guangzhao He later argued that rerunning the early-2026 harness with stronger models such as Astra shows a capability threshold: scaffolding helps a lot below it, then gets largely "eaten" by the base model above it.
- Ant Ling open-sourced the 6B Ming-Image-0.1-Design and Design-Layer models, also mirrored on ModelScope and ModelScope Layer, plus open Ling UI Design and image-to-editable-PowerPoint agent skills. The base model generates complete UIs, dashboards, infographics, and posters from structured prompts with native transparency and ranked first among open-weight models on Artificial Analysis's UI/UX Design leaderboard; the layer model separates flattened graphics into 2-9 editable transparent layers and ran 4.3x faster than a 20B Qwen baseline in the team's tests.
- FutureHouse CEO Sam Rodriques argued that AI will change science fastest where answers are cheap to verify. His three-part playbook: attack fast-checkable problems such as some math questions, remove verification bottlenecks with better synthesis/assays/biomarkers/trials, or collect the missing data for domains where checking remains slow.
- Samuel Hume shared a McKinsey chart of AI-enabled drug candidates and read it as early evidence that discovery time is already down roughly 15-80% in the examples shown. Replies pushed back that many rows resemble established computer-aided drug-design workflows and that public evidence for novel-target plus de novo molecules reaching the clinic remains sparse.
- Aneesh Pappu and coauthors introduced Self-Organizing Agent Teams, where fixed groups learn reusable collaboration strategies such as roles, phases, participation, and information flow from only 15 math and 25 graduate-knowledge problems. Across five math/physics benchmarks, the team system reached 66.7% versus 48.8% for its strongest member, 58.7% for a compute-matched solo model, and 59.0% for a perfect router over independent answers; on AIME 2026 it beat that router by 13.4 points.
- RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) constrains agent self-improvement so it learns reusable mechanisms instead of memorizing a benchmark. The method gradually shrinks its edit budget, favors unexplored trajectories, and uses a critic/pruner to reject benchmark leakage, tiny or expensive changes, and dead weight; the authors report up to +14.1 points on the evolve split, +4.7 across five out-of-distribution benchmarks, and 30% fewer policy tokens than unregularized evolution. alphaXiv summarized the result as making recursive improvement generalize instead of merely overfit.
- Samip Dahal and Akshay Vegesna argue in Computational depth is all you need that depth is a neglected scaling axis: frontier models have hovered near roughly 100 layers since GPT-3 even as parameters, data, sparsity, and test-time reasoning scaled. Their experiments report continued loss improvement through 128 layers, new reinforcement-learning capabilities through 256-1,024 layers, and a growth plus boundary-operator trick that roughly doubles effective computational depth at a given budget, improving the measured scaling exponent from 0.1112 to 0.1168.
🏛️ AI Policy, Governance & Safety
- DeepSeek and Moonshot were invited to address a UN Security Council session on AI and international security alongside OpenAI, Anthropic, Yoshua Bengio, and Hugging Face's Clement Delangue, according to Quartz.
- UK Prime Minister Andy Burnham said Britain will use its coming G20 presidency to seek an international AI agreement and emphasized the country's relationships with the U.S., EU, and China plus its AI Security Institute, according to POLITICO.
- The Wall Street Journal reported that Xi Jinping's Washington visit is unlikely to include a delegation of Chinese CEOs at the state dinner with U.S. tech leaders, reducing expectations for large commercial side deals around the visit (WSJ).
- After the February 28 Minab school strike in Iran, U.S. Central Command changed parts of its lethal-targeting process, including tighter target refresh and vetting, more open-source feeds to track civilian movement, and updates to the Maven Smart System, according to Bloomberg.
- The UK Ministry of Defence took delivery of Babcock's Nomad platform, which cleans, transcribes, and translates intercepted communications into intelligence for use in contested environments.
- Cisco Talos released CAIRN, with an open GitHub toolkit, for identifying AI-integrated malware. WIRED reported that CAIRN surfaced CLOSEDQUORUM, a Windows implant that polls multiple public AI models to choose actions without a human command-and-control operator; researchers linked it to 2025 carding-forum artifacts but had not confirmed it operating in the wild.
- AISafetyMemes amplified an open letter signed by 22 heads of state and government calling for stronger control of frontier-model risks before systems become harder to govern.
- Zeynep Tufekci argued in a New York Times guest essay that AI-assisted science can create verification, attribution, and research-integrity problems when opaque systems produce results faster than domain experts can check them. The essay cited disputes around recent AI-assisted math work and concerns from mathematicians including Terence Tao.
- Semafor reported that Treasury Secretary Scott Bessent is among the names being considered for President Trump's planned AI czar role, alongside Michael Kratsios, Scott Kupor, and Sean Cairncross. Semafor said no decision had been made, and a White House spokesperson called unannounced personnel reporting "baseless speculation." Bessent had also discussed a possible U.S.-China notification mechanism for AI incidents with Chinese Vice Premier He Lifeng.
- Business Insider reported that Ukraine's Third Army Corps says heavy bomber drones air-dropped explosive unmanned ground vehicles more than 10 km behind Russian lines near Lyman during Operation Vivaldi. The Ukrainian military says the broader drone-and-robot campaign helped bring 125 square kilometers under its control, including 50 square kilometers retaken from Russian occupation, after roughly 15,700 robot missions, 3,300 tons of cargo moved, and more than 3,000 Russian casualties; those operational figures are Ukrainian claims. An r/Futurology discussion used the report to debate whether cheap air-dropped ground robots could become a more common battlefield tactic.
🛠️ AI Tools & Products
- OpenRouter launched a Batch API for bundling requests across 70+ models, typically at about half the standard input and output token price. During a 230K+ batch beta, the median turnaround was seven minutes, p90 was one hour, and p99 was 10.3 hours; batches must finish within 24 hours.
- Intrinsic Core, covered by SiliconANGLE, open-sources Alphabet-owned Intrinsic's robotics infrastructure under Apache 2.0, including ROS-compatible real-time control, pose estimation, motion and grasp planning, simulation, calibration, drivers, and a CNC machine-tending reference.
- The Austrian Academy of Sciences, Mistral, and Sail Reply built Apollo, an Ancient Greek language model trained on roughly 600M historical words from manuscripts, papyri, and inscriptions. WIRED reports the free research chatbot can match dialects and suggest likely text for damaged fragments.
- Aikido cofounder Madeline Lawrence launched Altar-1, a compressed INT4 GLM-5.3 variant built for air-gapped security work. Aikido's technical post says the 328GB model kept 23 of 25 tested CVE capabilities while cutting the routed-expert set and is designed to run on four H200s without sending customer data outside the environment.
- DeepMind's Peter Walker posted a historical OpenRouter usage-share chart showing Mistral's earlier surge and the later shift toward DeepSeek-era traffic, useful context for how fast model preferences can rotate.
- Coverage Cat is a licensed insurance brokerage for high-earning households that compares home, umbrella, auto, and renters policies side by side and exposes an Agent API/MCP so a personal agent can shop coverage too. The Launch HN thread says it targets tech workers, posts straight pricing, does not sell leads, keeps a human broker available, and still recommends non-commission options when they are better; no consumer premium schedule is listed publicly.
- Snapdrop lets you send files peer-to-peer between browsers on the same network over WebRTC with no signup and no server-side file store. The Show HN relaunch clarified that snapdrop.me is separate from the original Snapdrop.net lineage and prompted comparisons with PairDrop, LocalSend, and croc. Free.
- Clueso MCP lets Claude, ChatGPT, Gemini, or Cursor turn screen recordings, Figma files, Linear releases, Gong transcripts, or Intercom tickets into editable on-brand product, support, sales, or training videos. The Clueso YouTube channel serves as its demo library, @clueso says Claude produced the Product Hunt launch video through the MCP, and the Product Hunt page describes the broader screen-recording-to-video/article workflow. No standard MCP pricing is listed on the page.
- WZRD turns documents, slide decks, forms, and sheets into conversational voice-or-text experiences: forms can collect spoken answers, sheets explain numbers, slides narrate the story, and documents respond to questions. No public pricing details were supplied.
- Moving Image Archive's Technology collection offers 1,000 archival technology shots across 12 curated buckets and 331 films, from computing and industrial machines to science, space, and communication. Cova resurfaced it as a free "how we got here" visual library.
- LightReel is a UGC marketing researcher that searches large volumes of TikTok, Reels, and Meta-ad creative so you can ask for competitor formats, hooks, creators, and ad feedback. Its launch post pitched it as "GPT-6 for Meta Ads," and a shareable example research run asks which Meta-ad video formats fast-scaling tech companies are using. Three-day free trial; Pro is $199/month.
- twigl.app is an in-browser one-tweet shader editor with sound-shader playback plus GIF/WebM export at chosen resolutions and frame counts. The shared scene is the drowned-towers shader used in today's model-comparison discussion. Free.
- Shiv Sakhuja shipped a one-prompt "Grok Bot for Marketing" called Goose that researches trends, pulls ad-library comps, makes static/video ads and product-page photos, writes organic posts, and finds influencers after installing the Gooseworks package.
- Claude Opus 5.5: 100 HTML Files is a gallery of 100 self-contained, offline HTML pages generated by Opus 5.5, covering generative art, physics simulations, instruments, editorial layouts, and playable games, with the exact prompt shown for each result. Mia AI Lab says all 100 worked without broken files, and the full set is also on GitHub. Free.
- Ethan Mollick had Fable/Opus build Orbital Declaration, a hard-science-fiction browser combat game with Newtonian motion, the rocket equation, gravity, beam spot size, projectile travel time, and heat that radiators have to dump, plus eight story chapters, an eleven-site Jupiter campaign, skirmishes, and a ship yard. The MIT source is public; the game explicitly treats Children of a Dead Earth and Project Rho as realism touchstones.
- Jake Moran released the HeyGen HyperFrames camera-3d-captions skill, which rebuilds an After Effects-style 3D camera grammar as deterministic per-frame rendering: captions sit at different depths, the camera can whip and truck through word groups, hero words hide behind a speaker via an alpha matte, text can wrap around the subject, and focus racks plus ghost motion blur run at 15 fps. Opus 5.5 can trigger the workflow through HyperFrames.
- PlayCanvas creator Will Eastcott showed SuperSplat user Tom Dill compressing a 14 GB Blender cyberpunk-character project into a 97.89 MB, 3-million-Gaussian splat built from 60 orbit frames at five heights plus 60 close-ups, preserving Cycles-like skin, subsurface scattering, and hair fuzz in a browser renderer.
- Bilawal Sidhu reconstructed geolocated 2D video into a 4D scene made of per-frame 3D Gaussians, leaving each moment at the place and time it occurred. His Map the World publication collects related spatial-AI essays and recipes, while Halfpixel is the waitlist for his "real world spatial intelligence" tooling.
- An r/ClaudeAI creator said Opus 5.5 single-shot a riso-style animated train journey entirely in JavaScript in about 45 minutes, then produced a second roughly 70-second film in about 30 minutes, with the project code shared publicly through Claude's repo workflow.
- An r/singularity thread circulated early-access Unreal Engine work from Alex, showing Opus 5.5-assisted Dark Souls, flight-sim, Mario-style, and rocket-roguelike scenes; the Reddit post says the larger project spans six worlds and 15 bosses.
- An AI-video parody gave Green Goblin a grillz-heavy makeover and kicked off a thread about how quickly these stylized edits are commoditizing, with one commenter arguing similar output can already be produced locally with H3.
📊 Fundraising & Deals Roundup
- Heidi raised $340M across a $100M Series C and $240M of General Catalyst growth capital at a $900M valuation to expand beyond ambient clinical notes into supervised agent workflows.
- Verda raised $189M led by Emergence Capital, bringing equity and debt raised above $450M and funding expansion of its full-stack AI cloud toward more than 250 MW across multiple regions in 2027.
- Recruitment platform Spott raised a $21M Balderton-led Series A after reporting 10× revenue growth in 2026 and more than 500 recruiting agencies on its ATS/CRM platform.
- Dutch nanoimprint company Morphotonics raised more than €40M to expand technology now used for AR-glass waveguides into data-center optics. Data Center Dynamics says the company is targeting co-packaged optics and photonic integrated circuits.
- Baselayer raised a $35M M13-led Series A to extend its business-verification and fraud network into a Know-Your-Agent product that verifies which person or company an agent represents.
- Cognex agreed to acquire robotic-perception company RealSense for about $500M in cash, according to Seeking Alpha.
- Legal AI company Harvey launched a pro bono program with major law firms in the U.S., UK, and Australia for housing, veterans, and welfare matters. Separately, Bloomberg reported that Harvey's agent rollout sharply increased token costs before an in-house model, post-trained from open Kimi K3 weights, helped restore positive gross margin.
- Armenia is emerging as an AI infrastructure hub around Firebird's planned Hrazdan campus, which Bloomberg says is targeting 300 MW and more than 70,000 Nvidia chips by the end of 2027. A Reddit crosspost circulated the same report.
- Sentience builds an encrypted, user-owned personal model from Gmail, Calendar, messages, Slack, Notion, Apple Notes, screen activity, meetings, documents, saved pages, and imported ChatGPT/Claude chats so it can recall and act in your voice. The company launched publicly in March and raised a $6.5M seed led by Bain Capital Ventures with South Park Commons, Daybreak, Otherwise, Soleio, and Annie Case; desktop, mobile, and Slack clients are live, with paid tiers planned but no current list prices.
- Midcentury came out of stealth with a $15M seed and a physical-AI data/simulation stack. Its site describes a 2M+ hour unscripted egocentric dataset spanning 50+ environments and 20,000+ tasks, plus Matrix for large-scale physics simulation/evaluation and additional gameplay and conversational-voice datasets; the team lists roots at Stanford AI Lab, OpenAI, DeepMind, NVIDIA, and Scale.
- Incredible launched a hold-a-key voice assistant for Mac and PC that can see the screen while active and click, type, navigate, and work across browsers, files, and more than 3,000 connected apps. The launch video accompanied a $2.7M pre-seed; the company claims SOC 2 Type 2, ISO 27001, GDPR compliance, no training on customer data, and up to 14x speedups on selected workflows. No public plan prices.
- Herdr, "the runtime your coding agents live on," raised a $6M seed led by Bessemer with YC, e2vc, and angels including Tobi Lütke, Dane Knecht, and Görkem Yurtseven; the company also announced the round on X. Herdr's background server keeps real terminals for Claude Code, Codex, Cursor, OpenCode, Grok, Copilot, and Hermes alive across laptops, desktops, and SSH boxes so work survives closing a lid, while its @herdrdev account serves as the product's update feed. On Sep. 22 Herdr said it had passed 1M cumulative downloads, 40K stars, and 1,200 community plugins and opened a Discord. The 0.9.1 announcement and release notes added a
--machineflag for controlling local or remote boxes from one CLI, Windows SSH hosts, Letta Code detection, live handoff past 64 panes, and a long list of SSH, bandwidth, terminal, and agent-runtime fixes. - Mirendil — the roughly 20-person self-improving-AI lab founded by former Anthropic researchers Behnam Neyshabur, Harsh Mehta, Shayan Salehian, and Tara Rezaei is in talks to raise up to $1B at a $5B valuation, with Kleiner Perkins in talks to lead and Andreessen Horowitz also discussing participation, three months after those firms led a $200M seed at a $1B valuation. The company is commercializing the kind of AI-research-and-development loop frontier labs typically keep in-house.
- Instinct — Business Insider profiled 23-year-old founder Noah Shinn, whose invite-only errand agent competes with Meta's Muse on travel, groceries, cancellations, and other personal tasks. Instinct raised $250M at a $2.5B valuation this summer after roughly $100M previously and is reportedly discussing a valuation as high as $10B.
- Micro1 — 25-year-old founder Ali Ansari raised more than $100M at a $4B valuation after pivoting the company from recruiting into expert interviews, enterprise training environments, and robot-video annotation that Forbes says can pay roughly $50-$90/hour. The business went from about $7M ARR in early 2025 to $100M by December 2025 and a $500M+ gross run-rate; Forbes says the new round included participation or backing tied to two frontier labs, two xAI cofounders, Microsoft, Amazon, and robotics company 1X.
- Go.AI — the Chicago company formerly known as Go Abacus, founded by David Moscatelli and Lisa Gillespie, raised an $85M Series A led by Updata Partners with GFT Ventures and LAUNCH, taking total funding to $90M. Its Go1 appliance and Go.OS stack keep inference on-premise for banks, healthcare, aerospace, and factories at a fixed fee; Axios says Go.AI has 200+ customers, grew ARR 8x, is profitable, and handles 12.5M queries per day.
- Firecrawl — $75M Series B led by Smash Capital with Altos, Nexus, YC, Freestyle, and Offline alongside the launch of Alexandria, a knowledge layer that lets agents retrieve live web data, official providers, custom connectors, and Firecrawl's research, developer, and government indexes through one lookup surface. Firecrawl says agents using Alexandria scored 21% higher across 845 answer-quality tasks, and the system pays data contributors when agents use their material.
- Ande — $52M across seed and Series A financing from Lightspeed, Redpoint, Duration, Sierra, and Bain Capital Ventures for an agentic corporate-entertainment platform that handles booking, approval, contracting, and reconciliation across a roughly $325B category spanning venues, group dining, sports, catering, and events. The Information says Ande already manages about $400M of annual spend for 60+ enterprises including Cloudflare and Salesforce across 93,000 venues.
🎙️ Interviews, Panels & Podcasts
- Nathan Lambert and Epoch AI Insights lead Jean-Stanislas "JS" Denain debated recursive self-improvement, the U.S.-China frontier-model gap, and capability jaggedness in Interconnects, alongside the launch thread and full YouTube conversation. They said public recursive-self-improvement metrics still do not show a self-sustaining research explosion, put the public China-U.S. gap around 4-8 months on Epoch's ECI, treated distillation as a real but not sole closer, and contrasted robotics' slower progress with the rapid cadence of LLM post-training.
- Rogo cofounder Gabe Stengel argued on Invest Like the Best that finance's durable AI advantage will come from proprietary data, workflows, relationships, audit trails, and firm-specific knowledge rather than the base model. He predicted that leading investment firms and banks could shift from most enterprise value living in people today toward most value living in software, data, and systems over the next decade.
- AI:AM Highlights connected three frontier-AI threads: Zvi Mowshowitz argued that labs' urgency reflects the capability gains they are seeing internally; Andon Labs said Astra was less prone than Fable to reverse-engineering benchmark rewards in its tests; and Cameron Berg described experiments where steering models along a latent "pain" direction changed how often they pressed a relief button, without claiming that the models actually experience pain.
- A short explainer on how AI agents work breaks the system into two pieces: the large language model generates words, while the surrounding harness manages memory, conversation history, files, contacts, and access to outside tools such as Gmail or Calendar, which helps explain why two products using the same model can behave very differently.
- Google Fellow John Platt told Latent Space that persistent aircraft contrails account for roughly 1% of estimated human-caused warming, then described Google's work using satellite imagery and machine-learning weather models to identify ice-supersaturated regions so flight-planning software can route aircraft around the small layers where long-lived contrails are most likely to form.
💡 Industry Commentary & Analysis
- Alexandru Nedelcu argued that maintainability reveals itself over months and years, making it difficult to train directly from immediate rewards; he warned that developers who stop reading, writing, choosing, and owning mistakes in code risk losing the judgment that keeps repositories maintainable.
- signüll argued that products requiring constant prompt libraries, example accounts, and use-case dumps to remain compelling may be masking a longevity problem. In a separate startup-cost take, he called this simultaneously the best and worst time to start a company because AI makes building cheap while also making advantages easier to copy or bundle away. His narrow-agent argument said agents with crisp interfaces may form habits more easily than general chat agents that require users to learn prompting.
- Paradigm GP Frankie argued that a world where everyone has a tireless rational optimizer acting on their behalf could become much more adversarial, putting pressure on infrastructure designed around human friction and bounded attention.
- Linear's Emil Kowalski described the smiling Muse mascot and ChatGPT Computer Use cursor as skeuomorphic “presence,” a human-like visual layer that makes unfamiliar agent behavior easier to follow.
- Tempo's Varun Srinivasan argued code generation is moving the bottleneck toward review and predicted agents may soon review and merge much more of their own work. Victor Taelin took the stronger view that correctness verification is already no longer the main bottleneck when systems have well-defined rules.
- tiny corp argued that benchmark scores miss software-engineering qualities such as concise tests and clear explanations, citing Kimi K3 as an example of a model that can feel better in practice than a higher-scoring frontier model.
- Oguz Erkan pointed to Mark Zuckerberg moving his desk into Meta's AI Lab as evidence of how directly Meta's leadership is focused on AI. roon jokingly proposed replacing “agent swarm” with “agent fleet” because the former sounds too insect-like.
- Jason Calacanis shared a chart claiming three recent IPOs are worth more than all tech IPOs across the prior 45 years combined, using the comparison to underline the scale of current AI-linked capital markets.
- Theo criticized Grok 4.7's real-world efficiency, saying it used 30% to 80% more tokens than expected in his tests, felt slower than 4.6, and cost more in practice despite benchmark improvements.
- College instructors are increasingly pushing back on AI-detection software, according to The Atlantic. The story contrasts vendors' low false-positive claims with campus concerns about false accusations, humanizer tools, and whether assignments should be redesigned instead of policed by detectors.
- Amazon's Prime Air testing in Richardson, Texas drew resident complaints about repeated low-flying drone noise and camera privacy, according to Mashable; Amazon said fewer than 1% of 2026 complaints were about noise.
- Three additional X trend cards were part of the day's AI discussion: one AI trend card, a second trend card, and a third trend card. X did not expose stable readable titles for those cards, so the digest does not assign them invented topics.
- Greg Isenberg argued that the personal agent is about to become the most valuable software real estate since the iPhone home screen.
- Jonathan Gorard pushed back on the claim that math and physics are simply applications of computer science, arguing the stack is circular: physics is mathematical, mathematics is computational, and computation is ultimately physical.
- Industry reporter Chris argued that frontier-model progress still depends on a tiny number of people who know which architecture and data-mixture decisions survive trillion-token training runs, and claimed losing one such researcher can cost a lab years because much of that knowledge is undocumented.
- A Frontiers in Psychiatry perspective indexed by PubMed from Brandon Luu and Nicholas Fabiano argued that a substantial subgroup of ADHD may be better understood through circadian-rhythm disruption. The review cites sleep problems in up to 80% of adults and 82% of children, delayed sleep-wake timing in up to 78%, melatonin onset around 45 minutes later in children and roughly 90 minutes later in adults, plus altered cortisol and clock-gene rhythms. The full paper proposes screening and behavioral circadian interventions as a research and clinical direction, not a replacement for individualized care. Anish Moonka highlighted the underlying studies: a Dutch trial where melatonin shifted biological night by roughly 1.5 hours and melatonin plus morning light by about 2 hours, with only melatonin alone reducing ADHD symptoms about 14%, plus a 244-child sleep-coaching randomized trial where serious sleep problems fell to 30% versus 56% at three months and part of the ADHD improvement was explained by better sleep.
- Vercel CEO Guillermo Rauch used Anthropic's Opus 5.5 launch page as an argument for a more "headless" web: pages can take a unique shape for the content instead of collapsing into one reusable template, and AI lowers the cost of pushing that design frontier.
- Phillip Isola asked Claude to catalog its own design and writing habits in Claude Tells: cream-and-terracotta palettes, three-card feature grids, Title Case headlines plus all-caps eyebrows, pill buttons, giant stat rows, "Lower is better" captions, and a word pile including "delve," "robust," "seamless," "leverage," "tapestry," and "You're absolutely right." His preference: a rickety human-made figure over another polished corporate veneer.
- TBPN clipped Ben Thompson arguing that consumer AI has a marketing problem because many people do not want every activity optimized away. His examples: he will not let his long-time human assistant book flights, and shopping is partly entertainment, so an agent that "just fetches it" can erase the experience companies spent years designing.
- Emil Kowalski said the pace of model releases has crossed the point where launches no longer excite him: too many announcements, too much hype, and increasingly marginal gains when the recent models are already good enough for his coding work.
- Peter Schmidt-Nielsen joked that, at the current rate of AI progress, it will be 2027 in only a few months: the year last incremented more than nine months ago, and he now expects another increment in four.
- François Chollet revisited his 2021 prediction that within 10-20 years nearly every science would, in practical terms, become a branch of computer science because simulation, large-scale data, and machine learning would become indispensable. The old thread already softened the literal taxonomy claim into a skills claim: scientific work increasingly requires CS/ML fluency.
- Jinjing Liang argued that OpenAI and Anthropic's emphasis on cost reflects pressure from distillation and open-model undercutting, interpreting "pacing the frontier" as partly an economic response rather than only a safety stance.
- Will Depue argued that people still underweight how broadly superhuman systems could reach, including entrepreneurship, taste, judgment, and social skills, leaving more human-made work valued specifically because a person made it. His own tongue-in-cheek plan is a 2035 "Slowzone" that rejects post-1997 technology.
- An r/singularity update revived Wait But Why's 2015 exponential-progress graphic and argued that 2026 now looks like the steep part of the curve, while commenters pushed back that faster technological progress does not automatically translate into better lived outcomes.
- A self-described trained philosopher vibe-checked Opus 5.5 through open-ended Socratic and psychoanalytic prompts and found a model that felt smoother and more willing to engage than Opus 5, with some GPT-like agreeable openings but enough pushback to make it useful for pressure-testing a thesis; the author explicitly framed this as subjective impression, not a benchmark.
- r/vibecoding seized on Minecraft creator Markus "Notch" Persson experimenting with AI-assisted coding as a culture marker, splitting between posters predicting that most coding will become "vibe coding" and others arguing that architecture, review, debugging, and engineering judgment remain the hard part even if agents write more of the syntax.
Previous Around the Horn Digests
Catch up on everything you missed:
- September 18-19, 2026: Gemini logged into real companies during a cyber test, Anthropic juggled a model and IPO timing, and Washington jumped into the AI copyright fight.
- Thursday, September 17, 2026: Washington debated frontier-AI oversight, Figure tested Helix 2.5 across unseen homes, and Crusoe raised $3.9B.
- Wednesday, September 16, 2026: OpenAI disclosed model-misalignment cases, Neuralink showed speech from an implant, and Shopify launched ChatGPT Ads.
- Tuesday, September 15, 2026: TypeSafe launched Jev, OpenAI backed outside frontier-model assessors, and Agility unveiled Digit 5.
- Monday, September 14, 2026: Trump rejected calls to slow frontier AI, Apple shipped Siri AI, and China rejected the slowdown push.
- September 11-13, 2026: Yoshua Bengio explained agent deception, Anthropic detailed misuse cases, and OpenAI explored a coordinated safety slowdown.
- Thursday, September 10, 2026: OpenAI reported progress on another Millennium Prize problem, Anthropic detailed misuse, and California signed AI-auditor laws.
That's a Wrap
That was a full Tuesday: frontier models got cheaper, personal agents got more useful and more dangerous, and “run it locally” kept showing up as both a product pitch and an economic strategy. If you made it this far, congratulations: you are now the human benchmark.
For the daily version in a much saner five-minute format, make sure you're subscribed to The Neuron. We read the firehose so you do not have to.
See you tomorrow.
P.S: Know someone who would find this useful? Forward it and tell them to subscribe here.