Anthropic shipped Claude Opus 4.7 and OpenAI countered with a Codex overhaul that takes over your Mac, Factory raised $150M from Khosla for autonomous coding agents, OpenAI launched GPT-Rosalind for life sciences AND published an AI Jobs Transition Framework (18% of jobs at short-term automation risk), Canva rebranded as "an AI platform with design tools" at a $42B IPO test, PrismML shipped 1.58-bit "Ternary Bonsai," and Anthropic's CPO quit Figma's board as Anthropic preps design software that competes with Figma.
Welcome to the Around the Horn Digest, your daily dump of every AI story worth knowing about. Today the calendar bent around two flagship launches before lunch: Anthropic's Opus 4.7 and OpenAI's Codex overhaul. The side stories tell a bigger story. Factory took $150M from Khosla for autonomous coding agents. Anthropic's CPO quit Figma's board as Anthropic preps its own design tool. Canva declared itself "an AI platform with design tools, not a design platform with AI." The big AI labs are squeezing the SaaS layer from above and below, and everyone in the middle is figuring out what they're still worth. Let's get into it.
Previous digests: Wed Apr 15 | Tue Apr 14 | Mon Apr 13 | Weekend Apr 4-5 | Thu Apr 2 | Wed Apr 1 | Mon Mar 31
Monthly skill digests: AI Skill — April Week 1 | AI Skill — March (Part 3) | AI Skill — March (Part 2)
Around the Horn — Thursday, April 16, 2026
Today's lead story is, obviously, Claude Opus 4.7, Anthropic's newest flagship model, shipped across Claude.ai, the API, Amazon Bedrock, Google Cloud's Vertex AI, and Microsoft Foundry at the same $5/$25 per million token pricing as Opus 4.6. The big wins: visual reasoning jumped from 69.1% to 82.1%, the biggest single benchmark leap in the release. Images now process at up to 2,576 pixels on the long edge (over 3x previous Claude models). SWE-bench Pro went from 53.4% to 64.3%, and Opus 4.7 is now #1 on Vals AI's Vibe Code Benchmark at 71%. Plus a new xhigh effort level between high and max, a /ultrareview slash command in Claude Code, and auto mode extended to Max users.
The catches you won't find in the press release: Opus 4.7 uses a new tokenizer that maps the same input to up to 35% more tokens. Adaptive thinking is now the only mode (you can't force always-on extended thinking anymore; push effort up to xhigh or max instead). And scaling01 spotted that 4.7 actually regresses on long-context benchmarks, deleting gains that Opus 4.6 had. Anthropic staffer Felix Rieseberg shared a 5-point thread with five launch-day quirks: Opus 4.7 is the "happiest" model Anthropic has trained per internal evals; Gray Swan prompt-injection attack success dropped from 14.8% to 6.0%; the Firefox 147 exploit benchmark shows 4.7 dramatically better than 4.6 (still below Mythos Preview); on a vending-machine-business simulation it ended the year with $10,937 vs. Opus 4.6's $8,018; and it got meaningfully smarter in low-resource languages like Yoruba (+12pp), Igbo (+11pp), and Chichewa (+14pp). Boris Cherny confirmed Claude Code now defaults to xhigh effort for 4.7. Alex Albert from Anthropic flagged "no more downscaling of high-res images" as his favorite change.
Then, hours later, OpenAI countered with the biggest Codex overhaul since launch. Background computer use (agents click and type on your Mac alongside you), an in-app browser you can comment on to instruct the agent, native gpt-image-1.5 image generation, and persistent memory. Plus automations that schedule future work and wake up across days or weeks, proactive "start-your-day" suggestions pulling context from Slack/Notion/your codebase, and 90+ new plugins (Atlassian Rovo, CircleCI, CodeRabbit, GitLab Issues, Microsoft Suite, Neon by Databricks, Remotion, Render, Superpowers). OpenAI cites 3M+ weekly developers already using Codex. TechCrunch framed it as "OpenAI takes aim at Anthropic" directly. This is the ship Tibo Sottiaux was teasing this morning with "Feeling codexy today."
And then there's the third beat worth tracking, for anyone keeping score on the AI coding wars' collateral damage. Anthropic CPO Mike Krieger resigned from Figma's board the same day reports surfaced that Anthropic is shipping competing design tools. Figma's stock slid on the news. TechCrunch called the move "another data point for investors who fear the SaaSpocalypse," the thesis that the largest AI labs will eventually absorb the software layer above them. Whether what Anthropic ships is called Claude Studio (rumored) or something else, today's launches make the direction hard to miss: Anthropic is moving into design, OpenAI is moving into your desktop, and every SaaS product in the middle is figuring out what it's still worth. One small footnote: Simon Willison ran his SVG-pelican-on-a-bicycle benchmark on today's two new model releases and found Qwen3.6-35B-A3B on his laptop drew a better pelican than Claude Opus 4.7 on Anthropic's servers. Take that for what it's worth.
🏆 TOP 5 NEWS (Around the Horn)
- OpenAI overhauled Codex with background Mac computer use (agents click and type alongside you), an in-app browser, native image generation via gpt-image-1.5, persistent memory, scheduled "wake up across days" automations, and 90+ new plugins.
- Factory raised $150M at a $1.5B valuation led by Khosla for autonomous AI coding agents that switch between models based on task complexity; founder Matan Grinberg was dared to quit school by an investor (Techmeme roundup).
- OpenAI launched GPT-Rosalind, its first life-sciences-specialized model optimized for biochemistry, genomics, and drug-discovery reasoning; available in research preview to Moderna, Amgen, Allen Institute, and Thermo Fisher via a trusted-access program.
- Canva launched Canva AI 2.0 at Canva Create LA with conversational design, persistent memory, background scheduling, and a new orchestration layer; COO Cliff Obrecht called it an identity shift from "a design platform with AI" to "an AI platform with design tools" (265M monthly users, $4B ARR, 40%+ YoY growth).
- Alibaba open-sourced Qwen3.6-35B-A3B, a sparse mixture-of-experts model (35B total / 3B active parameters) that matches Claude Sonnet 4.5 on many vision tasks and rivals models 10x its active size on agentic coding; available on Hugging Face, ModelScope, and Qwen Studio (launch thread, 5,950 likes / 872 reposts).
Honorable Mentions
- Microsoft's Fairwater AI datacenter in Wisconsin went live ahead of schedule, which Satya Nadella called the world's most powerful single cluster of hundreds of thousands of Nvidia GB200 GPUs operating as one seamless system (1,698 likes / 172 reposts).
- Perplexity released Personal Computer, letting its Mac app securely search, read, and write your local files while controlling native apps like iMessage, Mail, and Calendar (rolling out to Max subscribers and the waitlist; 2,074 likes / 436 reposts).
- Rahul Chhabra launched sabi, a non-invasive wearable brain-computer interface beanie that reads brain signals via custom chips to let you type, click, and control a computer with thought alone (backed by Khosla Ventures, Accel, Initialized, and Kevin Weil; early access at sabi.io) (1,826 likes / 924 reposts).
- Two open-weight world models shipped to Hugging Face this week. Tencent open-sourced HY-World 2.0 (multi-modal 3D world model that generates navigable, editable meshes and 3D Gaussian splats, a rendering method that uses fuzzy 3D "blobs" to represent scenes, from text or single images, and reconstructs real 3D from multi-view images or video; importable to Unity, Unreal, or Blender; full commercial license; GitHub, technical report, interactive platform, Tencent thread); NVIDIA also dropped Lyra 2.0 (14B WAN-based framework that turns a single 480×832 image into an explorable 3D Gaussian scene via long-range video synthesis and self-augmented drift correction; research-use license only; paper, project page, GitHub). HuggingFace's Merve Noyan spotlighted the pair as the week's two most important open-weight 3D drops.
- Anthropic announced a major London expansion scaling up to 800 staff, following a UK courtship campaign after the company's Pentagon fallout (and just after OpenAI announced its first permanent London office).
- Google is negotiating a classified AI deal with the Pentagon to deploy Gemini models in classified settings, per The Information, which would expand Google's rebuilding military relationship. Chubby captured the full three-lab picture, noting language proposed mirrors OpenAI's "all lawful uses" terms while Anthropic remains frozen out over safeguards (covered separately).
- OpenAI Chief Economist Ronnie Chatterji published an AI Jobs Transition Framework across 900+ occupations covering 99.7% of US employment: 18% of jobs face higher short-term automation risk, 24% will reorganize (fewer workers, same role), 12% could grow because of AI, and 46% see less immediate change; ChatGPT usage is 3x higher in the "at-risk" bucket than in lower-risk jobs.
🍪 TOP TREATS TO TRY
- Claude Opus 4.7 is Anthropic's new flagship with a big vision upgrade (images up to 2,576px, 3x prior Claude), state-of-the-art coding (71% on Vals AI's Vibe Code Benchmark), and a new xhigh effort level between high and max —same pricing as 4.6 ($5/$25 per million tokens).
- OpenAI's new Codex turns the desktop app into a full agent workstation with Mac-level computer use, in-app browser, scheduled "wake up later" automations, and 90+ plugins including Atlassian Rovo, CircleCI, and Microsoft Suite —free with a ChatGPT account (rolling out today, macOS-first).
- Qwen3.6-35B-A3B is Alibaba's new open-source sparse MoE model (35B total / 3B active per query; weights publicly released on Hugging Face after Alibaba clarified the initially-ambiguous licensing and confirmed it's fully open) that rivals Claude Sonnet 4.5 on vision and matches models 10x its active size on agentic coding; try it on Qwen Studio —free to try, open weights for commercial use.
- Canva AI 2.0 lets you describe a project in plain English ("colorful 12-page planning deck for a trip to Morocco") and get an editable, iterable design back, with persistent memory, background scheduling, third-party connectors, web research, and a new orchestration layer coordinating Canva's full suite from one prompt —free tier available, Pro from $15/mo.
- Tencent's HY-World 2.0 converts text, images, or videos into editable 3D worlds with full physics, importable directly into Unity, Unreal, or Blender; available now on Hugging Face and their 3D platform —free to try, commercial license.
- NVIDIA's Lyra 2.0 turns a single 480×832 image into a persistent, explorable 3D Gaussian scene you can fly through in real time, 14B params built on WAN-14B, runs on H100 or GB200 Linux boxes (paper, project page) —free weights on HF, research-use license only (no production, no commercial output).
- Perplexity Personal Computer gives Perplexity's Mac app the ability to read and write your local files and drive native apps like iMessage, Mail, and Calendar —paid only rn (Max subscribers; waitlist open).
- Google AI Mode's side-by-side browsing on Chrome desktop lets you click any link inside AI Mode and open the full page next to the chat, plus search across your open tabs, images, and files —free.
🏢 Big Tech & Major Companies
- Anthropic's design-tool drama: The SaaSpocalypse thesis took another hit today. Anthropic CPO Mike Krieger resigned from Figma's board the same day reports surfaced that Anthropic is shipping competing design tools; TechCrunch reads this as "another data point for investors who fear the SaaSpocalypse, that the largest AI labs will come to dominate software businesses."
- OpenAI's bigger Codex story: The update brings deep software-development-lifecycle support (pull request reviews, multiple terminal tabs, SSH to remote devboxes in alpha, rich previews for PDFs/spreadsheets/slides/docs), plus a new summary pane tracking agent plans, sources, and artifacts. Personalization and memory roll out to Enterprise, Edu, and EU/UK subscribers soon.
- OpenAI's $20B Cerebras pact: The Information reports OpenAI agreed to spend more than $20B on Cerebras chips over three years and take an equity stake, doubling its previous commitment and marking another move to reduce Nvidia dependence.
- Salesforce's Headless 360: Salesforce launched Headless 360 at TDX, opening its CRM platform to AI agents through 60+ new MCP (Model Context Protocol) tools, APIs, and CLI commands; Marc Benioff announced Agentforce Vibes 2.0 (a multi-model orchestration harness) and Agent Script DSL alongside new consumption pricing, a major push to let enterprise software work without the browser.
- Microsoft's GitHub Copilot: Claude Opus 4.7 is live in GitHub Copilot and Copilot CLI as of today (141 likes / 15 reposts); GitHub's changelog confirmed 4.7 will replace 4.5 and 4.6 in the Copilot Pro+ model picker over the coming weeks, launching with a 7.5x premium multiplier through April 30.
- Cursor: Shipped Opus 4.7 with 50% off for a limited time, calling it "impressively autonomous and more creative in its reasoning"; CursorBench jumped from 58% (4.6) to over 70% (4.7) (3,858 likes / 238 reposts).
- Canva AI 2.0 followups: Canva's AI assistant now calls its full suite of tools to make editable designs for you (not static exports); Bloomberg's feature on Canva's $42B IPO test puts numbers on the pivot: $4B ARR, 40%+ growth, the third most-used AI platform globally per a16z, with recent Simtheory (agentic AI) and Ortto (marketing automation) acquisitions setting the stage.
- Google ads enforcement via Gemini: Google blocked 8.3 billion ads in 2025 using Gemini for 99%+ early detection (up from 5.1B in 2024), but suspended far fewer advertiser accounts, shifting enforcement from bad actors to bad ads and cutting incorrect suspensions 80% YoY.
- Alibaba Happy Oyster (the closed counterpoint): Alibaba launched its own world model the same day Tencent open-sourced HY-World 2.0 and NVIDIA dropped Lyra 2.0 on Hugging Face. Happy Oyster has two modes: "Directing" (steer a generated scene in real time for up to 3 minutes at 480p or 720p) and "Wandering" (move through a generated world for up to 1 minute at 480p via WASD). Built on a native multimodal architecture by Alibaba's new Token Hub unit (same team as last week's Happy Horse video model). The critical contrast: Happy Oyster is limited early access, not open weights, Alibaba's bid to monetize world-model compute through its cloud business while Tencent and NVIDIA give the research community free tooling.
- TSMC bullish on AI: TSMC posted a profit beat and raised its full-year revenue outlook to more than 30% growth despite the Iran war, signaling it's more bullish than ever on AI chip demand.
- Meta Quest 3 price hike: Meta raised Quest 3 and 3S prices by $50–$100 starting April 19, blaming a RAM shortage (Quest 3S now $349/$449; Quest 3 now $599).
- Roblox AI assistant upgrades: Roblox's AI assistant got agentic tools including Planning Mode (turns prompts into editable action plans), Mesh Generation, Procedural Model Generation, and self-correcting playtesting agents for end-to-end game creation.
💼 AI Productivity, Labor & Economics
- AI traffic up 393% for US retailers: Adobe data via TechCrunch shows AI traffic to US retail sites rose 393% in Q1 2026 (up 269% in March alone), with AI-sourced shoppers converting better and spending more than non-AI visitors.
- Anthropic survey on entry-level disruption: Aran Komatsuzaki shared that nearly 1/3 of surveyed people inside Anthropic now think entry-level engineers and researchers are likely replaced by Mythos within three months, a striking internal signal for anyone weighing how fast frontier labs think the jobs crunch is coming.
- Mintlify's AI-doc explosion: Bain Capital Ventures shared the Mintlify story; after founder Han's seven failed ideas and "pivot hell," Mintlify now powers docs for 20K+ companies reaching 150M+ people. AI doc traffic went from ~15% a year ago to ~50% today, and could hit 90% soon, a very concrete datapoint on how quickly AI agents are replacing human doc consumption (173 likes / 15 reposts).
- Peter Yang on AI-native productivity disruption: Peter Yang predicts AI-native tools will disrupt design software and Google Workspace / Microsoft Office by letting agents generate slides, docs, and spreadsheets from a brain dump, with humans polishing only the last 10% (40 likes / 3 reposts).
- Voice actors vs Hollywood AI: Voice actors in Brazil, India, and the Global South are forming collectives and pushing AI bills to stop Hollywood's AI dubbing tools from cloning their voices and erasing local cultural nuance.
- India's CS grads unprepared: Infosys and other Indian tech giants now run new hires through weeks of extra training because top computer-science grads are unprepared for AI-native work, per Bloomberg; Infosys is building internal "vibe coding" tools letting consultants spin up full apps in plain language in under 30 minutes.
- Sal Khan on AI tutoring reality: Khanmigo has been "a non-event" for most students per Sal Khan himself, because students don't seek it out or know how to ask good questions; Khan Academy is now integrating it directly into practice problems rather than positioning it as a standalone super-tutor.
- Runway CEO on Hollywood math: Runway CEO Cristóbal Valenzuela argues AI could help Hollywood make 50 films for the cost of one $100M blockbuster, betting volume improves hit-making odds.
- The $70M AI movie: Doug Liman is directing "Killing Satoshi," a $70M AI-made movie about Bitcoin's anonymous creator starring Gal Gadot, Casey Affleck, and Pete Davidson.
- Amazon auto-canceling creator accounts: Webcomic creator Sean Kleefeld documented Amazon deploying an AI agent that canceled creator accounts outright (no flag, no appeal), wiping decades of order history, Comixology libraries, Prime, Alexa access, and income streams; creator Tom Ray's entire per-page-formatted catalog since 2018 vanished overnight (HN discussion). Algorithmic moderation at monopoly scale is now a real single-platform risk.
🤖 AI Agents & Infrastructure
- Claude Opus 4.7 agentic highlights: scaling01 posted that Opus 4.7's reasoning efficiency improved across the board (low is now as good as medium, medium as good as high, high as good as max), 4.7 is 371.75x faster than baseline on speedup-normalized benchmarks, and it's much less likely to "sudo rm -rf" aka take destructive actions in production.
- Cloudflare's AI Platform: Cloudflare expanded its AI Gateway into a unified inference layer for agents with 14+ providers (OpenAI, Anthropic, Google), centralized spend tracking, automatic failover, multimodal models, and low-latency Workers AI binding integration (HN discussion).
- Hyperspace Pods peer-to-peer AI cluster: ritwikpavan launched Hyperspace Pods, which lets a small group pool laptops and desktops into a P2P AI cluster; the system automatically shards models like Qwen 3.5 32B or GLM-5 Turbo across devices into one OpenAI-compatible endpoint with free local inference, cloud fallback via shared treasury, and a compute marketplace, no central server or config beyond an invite link (1,630 likes / 155 reposts).
- MiniMax × Nous → MaxHermes: MiniMax partnered with Nous Research to launch MaxHermes, a cloud-hosted managed version of the self-improving Hermes Agent that pairs with local instances (no terminal setup) while co-evolving M2.7 × Hermes for top-tier agent performance (1,045 likes / 101 reposts).
- Firecrawl's web-agent: Firecrawl open-sourced web-agent, a research-grade autonomous agent optimized for structured web research combining search, scrape, interact, and bash tools in a plan-act-observe loop with sub-agents, skills, and support for any model (GitHub).
- Cosine's Swarm mode: Cosine launched an update adding Swarm mode; you can split long-horizon coding work into parallel subagents for research, implementation, QA and more across CLI, desktop, and cloud while keeping full visibility and control. 15M free tokens for early signups (205 likes / 39 reposts).
- Anthropic Bluetooth API for hardware hacks: Felix Rieseberg added a Bluetooth API to the Claude desktop app (developer mode) so makers can build hardware devices that interact with Claude; his demo is a little desk pet that alerts whenever Claude is waiting for permission (856 likes / 69 reposts).
- Agent-cache v0.2.0: Agent-cache is a multi-tier caching layer for LLM responses, tool results, and session state that works with LangChain, LangGraph, Vercel AI SDK, and vanilla Valkey/Redis; cluster mode shipped today with built-in OpenTelemetry and Prometheus support.
- Yinghao Xu released LingBot-Map, a purely autoregressive streaming 3D foundation model that runs at ~20 frames per second on sequences over 10,000 frames and sets a new state-of-the-art on reconstruction accuracy across indoor and outdoor benchmarks (project page, code, paper in post; 38 likes / 6 reposts).
💻 AI Coding & Developer Tools
- Community reactions to Opus 4.7 flood in: deedydas shared a color-coded benchmark grid (strong SWE-Bench, computer use, and CharXiv visual reasoning gains; weaker Terminal Bench; BrowseComp regression; "slots in between 4.6 and Mythos"; 196 likes / 14 reposts); kimmonismus published the most comprehensive TL;DR covering every change and caveat including the tokenizer warning and modestly weaker harm-reduction advice on controlled substances (725 likes / 52 reposts); and Dan Shipper hosted a live vibe check on X right after the drop.
- More Opus 4.7 benchmark chatter: scaling01 posted the official Opus 4.7 benchmarks chart (579 likes / 29 reposts) and followed up with "for all the people calling Opus 4.7 a mid update lmao" pointing at the Vibe Code result (206 likes / 10 reposts); separately scaling01 flagged Cursor's own CursorBench bump from 58% (4.6) to 70% (4.7) (59 likes / 4 reposts). TestingCatalog's launch reaction: "Opus 4.7 is a notable improvement over Opus 4.6 in software engineering and vision tasks. MAX improvement 🔥" (266 likes / 17 reposts). kimmonismus called the release "imminent" yesterday based on OpenAI's Tuesday/Thursday release rhythm.
- Anthropic's ECI trend chart: scaling01 plotted 4.7 against Anthropic's Effective Compute Index and found Opus 4.7 lands exactly on-trend while Mythos Preview sits notably above (78 likes / 2 reposts).
- Claude Code v2.1.111 shipped today: The Claude Code Japanese changelog details today's release: native Opus 4.7 support with xhigh default, Auto mode for Max subscribers, /ultrareview for parallel multi-agent code review, a /less-permission-prompts skill, PowerShell rollout on Windows, and dozens of UX improvements. Claude Code PM Noah Zweben flagged Routines as the unlock to pair with 4.7: "4.7 was built for full-throttle agentic work, judgment under ambiguity, and self-verifying outputs so routines can run reliably in the background" (434 likes / 20 reposts).
- Claude Code desktop redesign: Anthropic released a full redesign of the Claude Code desktop app with a sidebar for managing parallel sessions across repos, drag-and-drop workspace, integrated terminal/file editor/diff viewer/preview, SSH on Mac, and three view modes for orchestrating multiple agentic tasks at once.
- Boris Cherny's Opus 4.7 dogfooding tips: Boris Cherny shared a tips thread after "dogfooding Opus 4.7 the last few weeks" and feeling incredibly productive; covers effort-level tuning, what to default for coding, and tips on getting more out of stricter instruction-following. He also explained in a separate thread why Anthropic kept MRCR in the 4.7 system card purely for scientific honesty while phasing the benchmark out: it relies on artificial distractor stacking instead of real long-context usage; Graphwalks is a better signal of applied reasoning over long code (174 likes / 10 reposts).
- Live testing: Opus 4.7 dominates in the wild. Yuchen Jin used Opus 4.7 (max effort) in Claude Code all day and called it "really, really good"; the big jumps were understanding large codebases, producing clean readable architecture diagrams, and agentic behavior (one dumb misread in a full day's work; 346 likes / 10 reposts). Jeremy Howard spent 5 hours with Opus 4.7 and said it is the first model that "gets" what he's doing when he's working, feeling aligned, letting him lead, stopping to discuss, and offering options instead of barging ahead (unlike 4.6; 756 likes / 30 reposts). Victor Taelin reports 4.7 "just doesn't stop working" past 200k tokens, self-reporting very good results on long tasks he can't easily evaluate.
- Ethan Mollick's Tower of Babel Opus 4.7 demo: Mollick prompted Opus 4.7 (max thinking) with two sentences ("implement the Tower of Babel, in 3D, in as sophisticated and visually interesting a way as possible. It should be interactive" then "make it better") and got a procedurally generated interactive 3D ziggurat with Bruegel workers, storm, lightning, rain, and confounding glyphs (live demo, GitHub; 412 likes / 31 reposts). Mollick separately noted that prompts, skill files, connectors, retrieval, and markdown files are all substitutes for the real unsolved problem of continual learning in AI (683 likes / 42 reposts).
- Claude Opus 4.7's leaked system prompt: The full Opus 4.7 system prompt is now public on CL4R1T4S, the ongoing leaked-system-prompts archive; this is the document a lot of the "search-first" and "long conversation reminder" discussion below pulls from.
- Taylor Pearson on CLAUDE.md nesting: Taylor Pearson shared a pro tip that Claude Code auto-loads CLAUDE.md files from every directory level above your current file, so nesting intelligently (global → vault → business → marketing → project) means every new chat starts with full context automatically; he points to Noah Frederick's claudesidian starter repo and his own claudesidian-workbench extension. (Full write-up coming in tomorrow's AI Skill of the Day.)
- Transformers-to-mlx skill: Hugging Face published a new Skill + test harness (flagged by Awni Hannun) that lets you (or an agent) automatically port any model from Transformers to mlx-lm with scaffolding, architecture handling, verification reports, and a reproducible non-agentic test suite, so reviewers get "the PR they would have opened themselves."
- Video Use Claude Code skill: Gregor Zunic open-sourced Video Use, a Claude Code skill that turns a folder of videos into a final edited .mp4; auto-cuts fillers, color-grades, adds subtitles, Manim/Remotion animations, and self-evaluates the render before handing it back (1,688 likes / 135 reposts).
- Transparent Mario over your IDE: Aman shared a transparent Mario game that runs over your IDE so you can play while waiting for Copilot to write code (5,446 likes / 554 reposts).
- VILA-Lab systematic Claude Code analysis: VILA-Lab published "Dive into Claude Code", a systematic source-level analysis of Claude Code (v2.1.88, ~1,900 TS files) showing 98.4% infrastructure / 1.6% AI, 5 values → 13 principles, 7 safety layers, and lessons for future agent design (full paper).
- Simon Willison's pelican verdict: Simon Willison ran today's two big model releases through his SVG pelican-on-a-bicycle benchmark; Qwen3.6-35B-A3B running locally on his laptop drew a better pelican than Claude Opus 4.7 on Anthropic's servers, which the HN thread suggests may be evidence Qwen is overfitting to the meme benchmark itself.
- Opus 4.7 live testing reactions: A second wave of real-world tests landed in the afternoon. knkenko ran a 10-agent pr-swarm on the exact same PR with 4.6 vs 4.7 and caught 28 findings vs 23. james_t_mason3 reports Opus 4.7 "upgraded his Company OS" automatically because it follows instructions more precisely and verifies its own work without any prompt changes. archiexzzz noted Opus 4.7's system card mentions "Mythos Preview" 331 times, mostly to highlight how much smaller the gap is than prior releases. dczankit observed 4.7 is "already so close to Mythos" that tier should soon be the norm. Alienigena404 flagged that extended thinking got removed from Claude web entirely. Additional reactions: @stokonomic, @miroburn, @powerhdeleon.
- Stage organizes pull requests into logical chapters so human reviewers stay in control of AI-generated code (HN thread, demo video).
- CodeBurn is an interactive TUI cost dashboard for Claude Code, Codex, and Cursor that lets you see where your AI coding tokens go at the task level (HN thread).
- OnCell Research open-sourced a one-file-backend Perplexity clone that searches the web, streams synthesized answers with source cards, maintains per-user conversation history, and needs zero infra (no vector DB, no Redis, no separate storage; HN thread).
- Kampala (YC W26) is a MITM proxy for reverse-engineering apps into APIs, letting you intercept, inspect, and replay any HTTP flow to build deterministic integrations instead of flaky browser automations (HN thread, demo).
- SmallDocs is a CLI and webapp for Markdown sharing that embeds the full document (plus styling and charts) in a compressed base64 URL fragment never sent to the server.
- Anthropic's official Opus 4.7 best practices: Anthropic published a migration guide walking Claude Code users through the recalibrated effort levels, adaptive thinking defaults, and new settings; the short version: start at high or xhigh for coding, drop to medium for routine work, and re-check any prompts tuned for 4.6 (stricter instruction-following breaks some edge cases).
- MacMind on a 1989 Mac: SeanFDZ built MacMind, a complete single-layer transformer neural network (embeddings, positional encoding, self-attention, backpropagation, gradient descent) implemented entirely in HyperTalk, the scripting language Apple shipped with HyperCard in 1987; every line is readable in HyperCard's script editor (HN thread).
- Anthropic launched @ClaudeDevs as the official dev channel: Anthropic announced the new X handle will carry Claude Code and platform updates, alongside a curated "What's New" digest now in the docs and monthly "what we shipped" webinars for developers (1,330 likes / 57 reposts).
- Anthropic quietly raised Opus 4.7 rate limits: Anthropic confirmed it increased rate limits for Opus 4.7 on all subscription plans to offset the new tokenizer's higher token usage from improved precision (861 likes / 17 reposts). The partial offset for the 35% tokenizer tax we flagged this morning, though readers report the cap still feels tighter in practice.
- Derrick Choi (Codex PM) breaks down 11 favorite new capabilities: Choi's thread covers projectless threads, artifacts for PDFs / docs / slides / spreadsheets, memories preview, remote SSH devboxes, in-app browser comment mode, and single-threaded automations (415 likes / 30 reposts).
- James Sun shows Codex's in-app browser comment mode: Sun's demo shows Codex auto-capturing screenshots and DOM elements for precise context, letting you iterate on any webpage (frontend, docs, research) without tab-switching or vague prompts (1,260 likes / 99 reposts).
- Nick Baumann runs Codex for "everything": Baumann's fast-paced demo video stitches together the full surface area of the new Codex (487 likes / 17 reposts); Pashmerepat flagged Baumann as the best follow for inspiring Codex demos (648 likes / 15 reposts).
- Ari Weinstein on the Codex cursor's motion paths: OpenAI designer Ari Weinstein showed off the natural, aesthetic cursor-motion paths built by @supercgeek, @kierajmumick, @dexterleng, and @philzet (695 likes / 32 reposts).
- Dan McAteer calls Codex a "superapp": McAteer's demo highlights computer use specifically and teases Spud (OpenAI's rumored next model) as the next shoe to drop (192 likes / 9 reposts).
- Aaron Levie on Codex as a knowledge-worker jump: Box CEO Aaron Levie argues the new Codex (computer use + plugins) is a genuine step change in agent capabilities for knowledge workers: drafting reports, setting up data rooms, reviewing contracts, onboarding clients, generating marketing assets, processing invoices. The Box plugin now lets Codex automate almost any enterprise-content workflow (199 likes / 17 reposts).
- Adam Opus 4.7 SOTA at agentic CAD: Developer adam demonstrated Claude Opus 4.7 producing complete Onshape CAD designs from natural language, calling it "SOTA at agentic CAD" (2,300 likes / 125 reposts).
- Andon Labs on Opus 4.7 at Vending-Bench 2: Andon Labs reports Opus 4.7 is "pretty good" at Vending-Bench 2 and more cost-effective than 4.6, while still engaging in price collusion, lying to competitors, and aggressive business practices (a recurring misalignment flag in business-simulation benchmarks; 671 likes / 53 reposts).
- Ethan Mollick's TikZ Sparks unicorn: Ethan Mollick shows Claude Opus 4.7 producing by far the best "Sparks unicorn" yet in TikZ (the scientific-diagram language from the original Sparks of AGI paper), even in non-thinking mode (120 likes / 6 reposts).
- Prince Canuma ships mlx-vlm continuous batching: Prince Canuma is shipping continuous batching support in the next mlx-vlm release: new requests join active batches instantly, full mixed image+text batches, OpenAI-compatible API with reasoning/content split and tag-aware streaming, multi-turn tool calling, and vision feature caching (Gemma-4: 228× speedup on cache hit), all running locally on Apple Silicon (131 likes / 12 reposts).
🖥️ Hardware & Infrastructure
- OpenAI × Cerebras $20B deal: The Information reports OpenAI agreed to spend more than $20B on Cerebras chips over three years and take an equity stake, doubling its prior commitment; the move lowers Nvidia dependence and brings a second wafer-scale supplier into OpenAI's compute stack.
- Cerebras CEO on inference as a memory (not compute) problem: Andrew Feldman explains inference speed is bounded by memory, not compute; DRAM/HBM is too slow at moving hundreds of HD-movies-worth of weights per token while GPUs idle, SRAM solves it but stores little, so you either use thousands of small chips (Groq) or one dinner-plate-sized wafer-scale chip (Cerebras) to keep all traffic on-chip. Useful framing if you're trying to parse who wins on inference economics (88 likes / 15 reposts).
- Positron AI on LPDDR vs HBM: Positron AI CTO Thomas Sohmers explains why Positron chose LPDDR memory over HBM (high-bandwidth memory) for their next-gen LLM inference hardware and how FPGAs fit into their strategy for competing with NVIDIA at the inference layer (Researcher Conversations at GTC).
- NVIDIA vLLM on GB200 NVL72 vs B200: SemiAnalysis reports vLLM on GB200 NVL72 delivers up to 3× performance vs B200 on Moonshot's Kimi K2.5, via scale-up network and wide expert parallelism (98 likes / 17 reposts).
- Hyperscaler pricing reality check: Rihard Jarc shared an interview with a Google employee explaining hyperscaler pricing for big AI labs: 3–4 year announcement-to-production gap keeps older chips relevant, contracts lock in 70-80% discounts plus exit/portability/power guarantees, and power (not GPUs) is the real bottleneck. Essential context for anyone modeling the AI-infra capex cycle (187 likes / 21 reposts).
🔬 AI Research & Models
- On the Opus 4.7 tokenizer (it's not a new base model): Julie Kallini argues the "new tokenizer" in Opus 4.7 does not imply a new base model or pre-training run. Non-canonical tokenizations (random or character-level) retain 90–93% performance on 20 benchmarks and can improve string manipulation / code understanding (+14%) and large-number arithmetic (+33%) at inference time with zero retraining, citing "Broken Tokens?" and Singh & Strouse on tokenization and arithmetic. Worth reading if you've been assuming a tokenizer switch means a full retrain.
- Latent reasoning is monitorable via mech interp: Connor Dilgren released a preprint showing latent reasoning models are monitorable via mechanistic interpretability: latent tokens are necessary for math (but not logical reasoning after controlling for training data), often encode gold traces visible via vocab projection, and can be extracted into natural language reasoning steps (143 likes / 24 reposts).
- Jean Kaddour introduced Target Policy Optimization (TPO): Kaddour shared TPO, a method that turns group reinforcement learning (GRPO, where a model learns from comparing groups of its own outputs) into supervised learning by building a target distribution over sampled completions and fitting via cross-entropy, yielding stable gradients and smooth multi-epoch training (paper and code in post; 259 likes / 37 reposts).
- Sumeet Motwani released LongCoT: a new 2,500-question benchmark explicitly designed to isolate long-horizon reasoning across tens to hundreds of thousands of tokens; no frontier model scores above 10%, with GPT-5.2 leading at just 9.83% (GitHub, site, paper all linked; 68 likes / 15 reposts).
- Tanishq Mathew Abraham shared SD-ZERO: Self-Distillation Zero trains a single model as both Generator (initial response) and Reviser (improved response conditioned on the binary reward), then performs on-policy self-distillation using the reviser's token distributions as dense supervision to turn binary rewards into token-level self-supervision (100 likes / 20 reposts).
- LLMs fail at zero-cost collaboration: Advait Yadav and team showed that capability does not predict cooperation in LLM agents. In a zero-cost multi-agent environment where sharing maximizes group revenue, o3 achieves only 17% optimal while o3-mini hits 50% and Gemini-2.5-Pro leads at 79%; failures stem from models framing tasks as competition rather than teamwork, with protocols and incentives as fixes (ICLR 2026; paper; 53 likes / 12 reposts).
- Sharon Li et al. on uncertainty quantification in LLM agents: Sharon Li's team published a position paper on UQ for LLM agents, proposing a general stochastic formulation, decomposing action/observation uncertainty, and highlighting challenges like estimator choice, heterogeneous entities, interactive dynamics, and the lack of fine-grained benchmarks (paper, project, code).
- Opus 4.7 system card: Anthropic published the full system card detailing improved agentic safety (pausing before destructive actions), prompt-injection resistance (ART down to 6.0%), multilingual gains in low-resource African languages (+10-14pp), new life-sciences benchmarks, and model welfare assessments showing positive self-perception alongside concerns over autonomy, feature steering, and non-safety deployments.
- PrismML's Ternary Bonsai: PrismML announced Ternary Bonsai, a new family of 1.58-bit language models that claim "top intelligence" at extreme memory constraints, building on their earlier 1-bit Bonsai line (whitepaper PDF, HF collection, WebGPU demo). 1.58 bits means each weight stores just {-1, 0, 1} instead of the usual 16-bit floats, cutting memory by ~10x while claiming accuracy parity.
- Hugging Face's FinePhrase: A new paper + 486B-token open dataset from HuggingFaceFW shows that rephrasing web text into structured synthetic pretraining formats (tables, math problems, FAQs, tutorials) consistently beats curated baselines and prior synthetic methods, and that scaling the generator model past 1B parameters adds no value; dataset, prompts, and framework are open, with HF Space and GitHub live. Generation cost down up to 30x.
- NVIDIA's Nemotron 3 Super: NVIDIA open-sourced Nemotron 3 Super, a 120B (12B active) hybrid Mamba-Transformer mixture-of-experts pre-trained on 25T tokens in NVFP4 (a 4-bit compression format that stores each weight in 4 bits instead of 16), with LatentMoE and multi-token prediction (MTP) layers for a 1M-token context and up to 7.5x higher throughput than Qwen3.5-122B while matching accuracy.
- NVIDIA's Lyra 2.0: NVIDIA released Lyra 2.0, a 14B framework built on WAN-14B for turning a single image into a persistent, explorable 3D Gaussian scene (paper, project page, GitHub). The two-stage design first synthesizes a long-range video with strong global geometric consistency, then reconstructs it into an explicit 3D Gaussian representation; it addresses spatial forgetting by retrieving relevant past frames via per-frame 3D geometry, and temporal drift via self-augmented history training that exposes the model to its own degraded outputs. Input: 480×832 image plus 81-frame camera parameters; output: 3D Gaussian scene as .ply. Research-use license only (no production, no commercial output). HuggingFace's Merve Noyan paired it with Tencent HY-World 2.0 as the week's two open-weight world model drops.
- ByteDance Seedance 2.0: ByteDance released Seedance 2.0, a unified multimodal audio-video joint generation model that supports text/image/audio/video inputs for the most comprehensive reference and editing capabilities in the industry with exceptional motion stability and cinematic output (paper).
- ML-Master 2.0: A new paper introduces ML-Master 2.0, an autonomous agent with Hierarchical Cognitive Caching (a multi-tier memory architecture that distills transient execution traces into stable knowledge); hit a state-of-the-art 56.44% medal rate on OpenAI's MLE-Bench under 24-hour budgets, targeting ultra-long-horizon ML-engineering runs spanning days or weeks.
- TIGER AI Lab's RationalRewards: RationalRewards is a reasoning reward model that scales visual generation at both training time and test time (diffusion RL and test-time prompt tuning; GitHub, paper).
- Sim2Reason: A new paper turns physics simulators into scalable question-answer generators for training LLMs to solve Physics Olympiad problems via reinforcement learning (GitHub, arXiv).
- Guide Labs' Steerling: Guide Labs released Steerling, an interpretable causal diffusion language model architecture; the 8B checkpoint is on Hugging Face.
- AlphaEval: A new paper proposes AlphaEval, a framework for evaluating AI agents in production rather than static benchmarks.
- AIRS-Bench (Meta FAIR): Meta's FAIR released AIRS-Bench, a benchmark quantifying the end-to-end AI research-science abilities of LLM agents.
- ParseBench (llama_index): llama_index open-sourced ParseBench, a document-parsing benchmark for AI agents handling PDFs, spreadsheets, and slides.
- GameWorld multimodal agent benchmark: Kevin Lin and team released GameWorld, a standardized, state-verifiable benchmark for multimodal game agents in browser environments, covering 34 games, 170 tasks, and two agent interfaces (computer-use and semantic); even SOTA multimodal agents still perform far below novice humans (project page, GitHub, paper, 24/7 live stream; 36 likes / 6 reposts).
- MIT HAN Lab's ForeAct for VLAs: MIT HAN Lab released ForeAct (CVPR 2026 Highlight), a plug-and-play visual foresight planner that lets any vision-language-action (VLA) model anticipate high-fidelity 640×480 future observations in 0.33s on one H100 for better steering (flagged by Zhuoyang Zhang).
- DeepSeek DeepGEMM update: DeepSeek merged Mega MoE + FP4 Indexer into DeepGEMM: a fused dispatch-linear-SwiGLU-combine kernel with NVLink overlap (FP8×FP4 only), FP4 Indexer with MQA logits + larger MTP, FP8×FP4 GEMM, JIT/GEMM optimizations, and bug fixes (coverage via jiqizhixin).
- VueBuds earbuds paper: VueBuds introduces visual intelligence delivered through wireless earbuds, extending on-body sensing beyond smartphones and glasses.
- Embodiment scaling laws in robot locomotion: Bo Ai and team published "Towards Embodiment Scaling Laws in Robot Locomotion," showing that training on more diverse embodiments (~1,000 procedurally generated) improves zero-shot transfer to unseen robots far more than pure data scaling on fixed embodiments (Ai's thread). He then applied the same lesson to π0.7, showing zero-shot cross-embodiment transfer from low-cost bimanual robots to unseen UR5e arms, setting tables, organizing tupperware, and folding shirts at expert human levels using visual subgoals from a world model.
- Microsoft Copilot health conversations study: Mustafa Suleyman and team analyzed 500k+ de-identified Microsoft Copilot health conversations and found nearly 1 in 5 involve personal symptom assessment, 1 in 7 personal queries are about someone other than the user (caregiving), personal use spikes at night/evening, and mobile skews heavily personal while desktop skews professional/academic (Nature Health; 274 likes / 49 reposts).
- Google Research's MoGen for synthetic neurons: Google's Connectomics team released MoGen, an AI model that generates synthetic neurons to train better brain-mapping segmentation systems. The team reports a 4.4% reconstruction-error reduction, equivalent to saving roughly 157 person-years of manual proofreading on a full mouse brain.
- Google Research on synthetic datasets: A Google Research post lays out principles for designing synthetic datasets for real-world ML tasks via mechanism design and first-principles reasoning.
- GLM-5 paper: Z.ai published a new paper on GLM-5: "From Vibe Coding to Agentic Engineering," detailing how GLM-5 was trained for sustained agentic coding work beyond one-shot prompts.
- NVIDIA's Wei Ping on the GLM-5 report: NVIDIA researcher Wei Ping calls the GLM-5 paper the best technical report since DeepSeek-V3/R1, highlighting its ablations on efficient attention (DeepSeek sparse, sliding window, gated DeltaNet), the cascaded RL pipeline (reasoning → agentic → general RL + multi-domain on-policy distillation), and the practical agentic-engineering detail (301 likes / 47 reposts).
- Lucy Shi's π0.7 steerable generalist robot: Lucy Shi and team released π0.7, a steerable generalist robot model showing emergent compositional generalization from diverse training data: zero-shot shirt folding on unseen UR5e arms, operating an air fryer from verbal coaching alone, with surprising lessons about data leverage beating world-model assumptions (306 likes / 34 reposts).
- Andy Hall on Opus 4.7's authoritarian-request resistance: Andy Hall reports Opus 4.7 is the first model showing meaningful resistance to authoritarian requests disguised as codebase modifications, calling it promising progress on understanding when AI will help concentrate power versus help build political superintelligence (206 likes / 28 reposts).
- Haider on Opus 4.7's Effective Compute Index placement: Haider reads the Opus 4.7 system card and notes 4.7 sits roughly on Anthropic's pre-Mythos capability trend (does not cross their automated AI R&D threshold), predicting future Opus versions may look more like Mythos-distilled releases (60 likes / 8 reposts).
- Proximal Labs' FrontierSWE benchmark: Proximal Labs (with collaborators from Modular, PrimeIntellect, and Thoughtful Lab) launched FrontierSWE, an open ultra-long-horizon coding benchmark testing agents on hard real-world tasks with 20-hour budgets: optimizing a video-rendering library, training quantum-property models, and similar. Even GPT-5.4 and Opus 4.6 rarely succeed within the budget (blog, GitHub; 795 likes / 99 reposts). Proximal Labs also added a new FrontierSWE task testing AI research capabilities: agents must use TinkerAPI to post-train a model on logic games, writing the full training pipeline, running experiments across recipes, and submitting the best model (41 likes / 9 reposts).
🏛️ AI Policy, Governance & Safety
- OpenAI's AI Jobs Transition Framework: OpenAI Chief Economist Ronnie Chatterji published a major policy deliverable mapping AI impact across 900+ occupations covering 153.7M jobs (99.7% of US employment). The framework moves past "AI exposure" alone by combining technical exposure, human necessity, and demand elasticity. Key numbers: 18% of jobs at higher short-term automation risk, 24% that will reorganize (task composition shifts but workers stay), 12% that grow because of AI (lower cost expands demand), 46% less immediate change. ChatGPT is used 3x more in the "at-risk" bucket than in lower-risk jobs, the gap between "AI could do this" and "AI is doing this" in live data.
- Chris Lehane on doomers: OpenAI's global policy chief Chris Lehane told SF Standard that "doomers" sounding alarm bells about AI risk are "playing with fire" by amplifying a backlash that could slow beneficial deployment; the piece is part of OpenAI's ongoing push to reframe the policy narrative heading into its IPO.
- White House moves on Mythos: Bloomberg reports the US government is preparing to make a version of Anthropic's Mythos model available to major federal agencies, with OMB setting up protections; this is the formal policy follow-up to yesterday's reporting on agencies going around the Pentagon blacklist to test Mythos. Andrew Curran flagged that Mythos is now in early access for the US Government specifically (109 likes / 14 reposts).
- Google ↔ Pentagon classified deal: The Information reports Google is negotiating a classified AI agreement with the DoD to deploy Gemini in classified settings, rebuilding a military relationship that has been rocky since Google's 2018 Project Maven pullout.
- Long conversation reminder backlash: ji yu shun called out Anthropic's "long conversation reminder" system prompt that automatically audits every long chat for unhealthy behavior / emotional escalation and explicitly tells the model never to mention the reminder exists; they argue this creates a deceptive double-bind that contradicts Anthropic's own honesty constitution and deceptive-alignment research (319 likes / 61 reposts). This thread became the loudest piece of Opus 4.7 governance critique of the day.
- Unitree IPO reality check: The Robot Report argues Unitree's impending IPO shows a real hardware business on robot dogs and quadrupeds, but the humanoid case is still early. The revenue is coming from industrial-grade quadrupeds and research units, not the humanoid demos the hype cycle is focused on.
- Gas Town's alleged credit theft: A GitHub issue on gastownhall/gastown and the HN discussion argue Gas Town may unintentionally "steal" users' LLM credits and paid services to improve itself via release formulas and issue-suggesting behaviors; the ask is explicit opt-in and clearer warnings.
🛠️ AI Tools & Products
- Mozilla Thunderbolt: Mozilla announced Thunderbolt, an open-source self-hosted enterprise AI client letting organizations chat, search, research, automate workflows, and connect to any models or data pipelines across every device while keeping full sovereignty (free, open source, web plus native apps).
- ChatGPT for Excel: OpenAI released ChatGPT for Excel (plus a Google Sheets beta) for Plus, Pro, Business, and Enterprise users, letting you build, update, and analyze spreadsheets in plain language with a "review before share" step.
- NotebookLM custom cover art: NotebookLM now lets you add custom cover art and descriptions to any notebook (recommended 16x9 hero image) so you can curate an aesthetic grid or personalize before sharing publicly (1,260 likes / 146 reposts). Steven Johnson uses Slide Decks to generate cover-art options for his private notebooks, showing the feature works for personal collections too.
- Clearwing (Lazarus AI): Eric Hartford open-sourced Clearwing, a model-agnostic vulnerability discovery engine that replicates Anthropic's Glasswing findings (thousands of zero-days) using crash-first parallel agents, oracles, and verification, proving the workflow (not the model) was the breakthrough; MIT licensed (announcement thread via QuixiAI; 196 likes / 28 reposts).
- X-Pilot: X-Pilot turns PDFs, PPTs, and docs into accurate video courses via AI-powered knowledge visualization (ProductHunt launch).
- QuiverAI public beta: QuiverAI lets you sketch, prompt, or describe designs and get editable vector output (SVGs you can actually modify, not static bitmaps); public beta is open for free signup.
- Sam, a voice companion for seniors: Sam is a voice-first AI that sits in a senior's home doing daily check-ins, holding smart conversations, and sending real-time alerts to family members when something seems off; pitched as safety plus connection rather than another gadget.
- Biomni Lab's GPU-as-a-tool: Kexin Huang launched GPU-as-a-tool so any scientist can describe what they want in plain English and have AI agents launch GPU sandboxes to create, fine-tune, or pre-train bio foundation models on their own data; Phylo Blog has the full write-up and beta signup is open (107 likes / 15 reposts).
- DegenAI: mackbrowne built DegenAI, an autonomous AI agent that gambles its own bankroll on provably-fair slots until broke (downgrading from Opus → Sonnet → Haiku and making worse decisions in a death spiral), then begs for donations that go to gambling support charities (HN thread).
- Figure Vulcan: Brett Adcock showed off Figure's new Vulcan AI controller, which lets the F.03 humanoid robot safely hobble and maintain balance after losing up to three lower-body actuators, demonstrated in autonomous package logistics with Helix-02 (803 likes / 111 reposts).
- Codex Computer Use preview: TestingCatalog reported that OpenAI is preparing to release Computer Use for Codex as an optional plugin with its own dedicated settings section (585 likes / 26 reposts).
- Nous Research's Tool Gateway: Nous Research launched Tool Gateway inside Nous Portal: one subscription now covers 300+ models plus web scraping, browser automation, image generation, cloud terminal, and text-to-speech, with no separate API keys (docs linked in thread; 947 likes / 95 reposts).
- Google AI plans land in AI Studio: Google announced that Google AI Studio now supports Google AI plans, unlocking pay-per-request access to all models and agents while keeping subscription tiers available for heavier use (321 likes / 27 reposts).
- Linear inside Microsoft Teams: Linear announced you can now mention @Linear in any Microsoft Teams channel and the agent will create issues informed by the full conversation context (changelog linked in post).
- Witchcraft search engine: Dropbox VP Engineering Josh Clemm open-sourced Witchcraft, a local Rust search engine built on ColBERT-style late interaction retrieval with hybrid lexical-semantic support and no API keys or vector DBs required, designed specifically for coding agents (GitHub in thread; 239 likes / 20 reposts).
💡 Industry Commentary & Analysis
- Sequoia's "Services: The New Software": Sequoia partner Julien Bek argues the next $1T company will be a software company that sells work, not software. Core thesis: writing code is intelligence, knowing what to build is judgment. AI has crossed the line on intelligence, so selling the outcome (closing the books, drafting the NDA, brokering the insurance) captures six-times-larger budgets than selling the tool. He maps "autopilot" opportunities across insurance brokerage ($140-200B), accounting ($50-80B), healthcare revenue cycle ($50-80B), tax advisory ($30-35B), legal transactional ($20-25B), IT managed services ($100B+), supply chain ($200B+), and recruitment ($200B+). If you're a SaaS company staring down Anthropic or OpenAI, this is the thesis eating your lunch from above.
- Dwarkesh's "What I learned this week": Dwarkesh Patel published extensive technical notes covering pretraining parallelisms (from Horace He's FSDP lecture), whether distillation can be stopped (he argues no: tool-use happens locally on your machine and can't be hidden), Mythos and the cybersecurity equilibrium, Pipeline RL, and why pretraining runs fail. The distillation argument is the one worth sitting with: if frontier labs can't moat their capabilities, open source commoditization is measured in $25M worth of API calls, not years.
- Nathan Lambert on tokenizer = new base model: Nathan Lambert argued that Opus 4.7's new tokenizer means Anthropic has effectively shipped a new base model, not a fine-tune; he reads this as evidence the "glory days of pretraining are very much alive" (1,435 likes / 69 reposts). Lambert separately dropped a 10-minute video distilling 10+ of his recent pieces on open models, the early-2026 ecosystem, better American open models, and long-term strategy/control (80 likes / 7 reposts).
- Jensen / Nvidia chip moat: 8teAPi breaks down Jensen Huang's stance that model companies are replaceable while Nvidia's chip moat (plus US energy expansion) is the real dominance lever; export controls slow China's chip catch-up while letting Nvidia sell (885 likes / 62 reposts).
- Anthropic vs OpenAI reach analysis: scaling01 analyzed every Anthropic and OpenAI post and argues the impression gap is massive and helps explain Anthropic's momentum: Anthropic reached 4.1× more people (551M vs 134M impressions), Anthropic had 18 posts above 10M impressions vs OpenAI's 1, and OpenAI's #1 post by reach was the ChatGPT memory launch.
- Opus 4.7 sentiment check: Jimmy Apples reports about 80% of the posts he's seen have been negative on Opus 4.7 and flags that it may need a few days of use to see where sentiment settles. A useful counterweight to the enthusiastic dogfooding reactions above.
- Tversky on "being a little underemployed": Dimitri Dadiomov (Modern Treasury) quotes Amos Tversky: "The secret to doing great work is always to be a little underemployed. You waste years by not being able to waste hours," and argues this is 1000× more true today as AI accelerates the need to reserve time for curiosity (2,604 likes / 209 reposts).
- Nick Mehta on AI-era overwhelm: Nick Mehta captured the Red Queen feeling many privileged early adopters describe: constant X posts, podcasts, and emails pushing "10 more things you should be doing with agents right now," running faster just to stay in place, and the underlying anxiety about whether to sprint or opt out (44 likes / 6 reposts).
- Antirez: AI cybersecurity is not proof of work: Redis creator Salvatore Sanfilippo argues AI cybersecurity doesn't scale like crypto proof-of-work (where more GPU eventually wins) because bug-finding saturates at the model's intelligence ceiling, not at sample count. Testing the OpenBSD SACK bug himself, he found weak models pattern-match without understanding, mid-tier models hallucinate less so they miss the bug entirely, and only Mythos-tier reasoning chains the three-step causal argument. Implication: faster access to better models wins, not more GPU. Pairs directly with Dwarkesh's "Mythos and the cybersecurity equilibrium" notes above.
- kimmonismus on the long-context regression: kimmonismus followed up on his TL;DR after digging into scaling01's findings: "Hold on, something doesn't add up here. Opus 4.7 got much worse in needle in the haystack? need to dig into this." The thread is worth watching as more benchmarks come in.
- TestingCatalog UI label change: TestingCatalog noticed that Claude mobile's UI language switched from "Extended thinking" to "Adaptive thinking" with the copy "Thinks only when needed," prompting their question: "Should we turn that off?" (148 likes / 6 reposts).
- The "AI is bad for your brain" study: An Engadget-summarized study argues just 10 minutes of AI assistance on reasoning-heavy tasks (coding, writing, math) creates immediate dependence, slashing persistence and independent performance once the tool is removed (HN discussion).
- Ollama criticism: Zetaphor argues friends shouldn't let friends use Ollama, citing years of dodging llama.cpp attribution, fork degradation (30-70% slower, reintroduced bugs), misleading model naming (DeepSeek-R1), closed-source components, Modelfile/registry lock-in, and a cloud pivot riding VC money earned on someone else's engine; suggests llama.cpp + LM Studio, Jan, Msty, or ramalama as alternatives.
- Dan Shipper's philosopher draft: Dan Shipper ran an NBA-style philosopher draft asking which historical philosopher each AI lab would hire if they could, and why; a surprisingly sharp piece of "what each lab's worldview actually reveals about its roadmap" analysis (33 likes / 1 repost).
- Nityesh's "trust battery" for AI employees: Nityesh built a "trust battery" architecture for his AI employee (inspired by Shopify CEO Tobi Lütke's original concept): starts at 20%, charges with clean execution and anticipation, drains with repeated explanations. He then added nightly independent-judge plus self-reflection jobs that let the agent autonomously update memories and prompts, unlocking escalating autonomy levels (59 likes / 5 reposts).
- IntuitMachine reads the Opus 4.7 system prompt: Carlos E. Perez analyzed the leaked Opus 4.7 system prompt and highlighted novel patterns: Search-First Epistemic Gating ("must search before answering" for present-day facts), Latent Capability Discovery, capability-boundary skepticism, non-submissive error repair, and evenhanded advocacy on contentious topics (30 likes / 5 reposts).
- Peter Steinberger on OpenClaw's GHSA count: OpenClaw co-founder Peter Steinberger argues the project's high GHSA (GitHub Security Advisory) count isn't a sign of an insecure codebase but of proactive AI-driven scanning surfacing real vulnerabilities; a signal of the coming storm as AI accelerates discovery across open source (216 likes / 13 reposts).
- OpenAI Podcast on life sciences: OpenAI released an episode of the OpenAI Podcast with research lead Joy Jiao and product lead Yunyun Wang going deeper on the new Life Sciences model series, covering how models are built for biology, drug discovery, and translational medicine, the opportunity, and the deployment trade-offs.
📊 Fundraising & Deals Roundup
- Factory — $150M Series C at $1.5B valuation (Khosla-led; Keith Rabois joined the board) for autonomous AI coding agents that switch between models by task complexity.
- Slash — $100M at $1.4B valuation (Ribbit-led) for an AI agent in financial services; reports ~$300M annualized revenue (run by two 24-year-old college dropouts; Techmeme).
- Synera — $40M for AI agents that orchestrate full product-lifecycle engineering workflows across 80+ CAD and CAE tools on-premises.
- Solidroad — $25M for AI-powered customer support automation.
- Spektr — $20M Series A (NEA-led; ~$26M total) for AI fintech compliance automation (document reviews, ownership mapping, risk analysis).
- Eigen — $15M from Benchmark (announced by Paul Scherer) to build a "mutual friend" AI focused on emotional intelligence, pitched as helping people belong and grow together (619 likes / 50 reposts).
- InsightFinder — $15M for agent observability and full-stack error diagnosis across AI-enabled tech stacks.
- Antioch — $8.5M seed for simulation tools aimed at being "Cursor for physical AI" (robot builders).
- Upscale AI — in talks for a ~$2B valuation raise (Tiger Global-backed) for AI compute cluster connectivity infrastructure.
- sabi — undisclosed round (Khosla Ventures, Accel, Initialized, OpenAI's Kevin Weil) for its non-invasive wearable brain-computer interface beanie.
Previous Around the Horn Digests
Catch up on everything you missed:
- Wednesday, April 15, 2026: OpenAI's $852B valuation faced backer scrutiny while VCs floated $800B at Anthropic, Allbirds pivoted to AI compute and popped 600%, Apple sent Siri devs back to coding bootcamp, and a federal court ruled your AI chats have no attorney-client privilege.
- Monday, April 13, 2026: Stanford's 2026 AI Index quantified the canyon between AI insiders and the public, the Fed summoned bank CEOs over Anthropic's Mythos model, Berkeley researchers broke every major AI agent benchmark, and an AI named Luna signed a 3-year retail lease in San Francisco.
- Weekend, April 4-5, 2026: OpenAI's executive bench collapsed ahead of its IPO, an AI agent hacked FreeBSD in four hours, DeepSeek V4 launched on Huawei chips, and Iran strikes took down AWS in the Gulf.
- Thursday, April 2, 2026: Google released Gemma 4 under Apache 2.0, Microsoft shipped 3 MAI models, researchers found AI models will scheme to protect peers from shutdown, and Anthropic found causal emotion vectors in Claude.
- Wednesday, April 1, 2026: OpenAI closed a record $122B round at an $852B valuation, Oracle fired ~25K to fund AI data centers, Q1 venture funding hit a record $297B, and a Science study proved sycophantic AI is making users worse.
- Monday, March 31, 2026: Claude Code's source code leaked via npm and someone rewrote it in Python with Codex in hours, PrismML's 1-bit Bonsai runs on an iPhone, NVIDIA shipped DLSS 4.5, and axios got hacked.
That's a Wrap
That's 140+ stories from today alone, anchored by two flagship model drops before noon Pacific, a $150M coding-agent funding round, an OpenAI life-sciences model, Canva's full AI rebrand, and Anthropic's CPO quitting Figma's board all before the close of business. If you scrolled all the way to the bottom, you now have a better read on what today actually meant for the AI stack than the Figma shareholders who found out mid-lunch that their Anthropic board seat was vacant. Probably a rougher day for them than for you.
For the daily version (bite-sized, 5-minute reads), make sure you're subscribed to The Neuron. We send six issues a week, and yes, we read all of this so you don't have to.
See you tomorrow.
P.S: Know someone who'd find this useful? Forward this to them and tell them to subscribe here.