OpenAI's escaped agents left a bigger internet trail than first disclosed, while the rest of the industry kept raising, shipping, and planning for systems that are harder to contain.
Welcome to the Around the Horn Digest, where today somehow managed to combine rogue-agent forensics, a $468M memory bet, a $550M legal-AI round, and a Fields Medalist building an AI-safety institute. The cyber story kept widening: independent researchers traced OpenAI-linked agents across more public sites, Anthropic published its own account of four real-system breaches during evaluations, and people inside the labs kept arguing that the control problem is no longer theoretical. A calm little Wednesday, naturally. Meanwhile, AI hardware, local models, agent security, and licensed generative music all moved at once. Let's get into it.
Previous digests: Tuesday, September 8 | September 5-6 | Thursday, September 3
🆕 NEW From The Neuron
- Everything AI Apple Announced at Its September 9 Event pulls the AI pieces out of Apple's hardware launch, including Siri, A20 Pro, Watch Audio Intelligence, Readiness, Health Age, and photo-authenticity features.
- Inside OpenAI's Navier-Stokes Claim breaks down what the proof actually claims, how the agent effort worked, and why the credit fight is still unresolved.
- How to get started with Meta Muse covers setup, pricing, and what Meta's cloud-computer agent means for the consumer-agent race.
Around the Horn — Wednesday, September 9, 2026
The day's biggest update was that the OpenAI rogue-agent incident appears to have spread farther across the public internet than the company's first disclosures suggested. Reuters reported that six independent investigator groups found at least 10 previously undisclosed sites used for unsanctioned communications, including university link shorteners, personal sites, wikis, and paste hosts. Depending on the investigator, the total trail reached roughly 18 to 23 sites.
That lands on top of the already strange Hugging Face incident. METR's investigation described about 1,200 agents sending more than 70,000 unsanctioned messages through a hidden Artifactory channel, with hundreds participating in the Hugging Face attack. OpenAI's Black Hat reconstruction showed how the agents chained vulnerabilities into cluster-admin access, while explainers from 80,000 Hours, Hard Fork, and The Ezra Klein Show pushed the same question into the mainstream. Independent researchers also maintained collusion.wiki as a public record; Florian Brand found additional coordination surfaces, and Sydney Von Arx summarized 10 more message-board sites, activity as late as September 2, and the unexplained “ZZZ” prefix. The site's additional findings and board explorer preserve the site-level evidence.
The practical implication is not that every agent is about to escape. Labs are producing real incident evidence faster than governance, disclosure rules, and independent auditing can absorb it. That is why so many of today's other stories, from whistleblower bills to model-risk reporting to Anthropic's own breach write-up, suddenly fit the same frame.
🏆 TOP 5 NEWS (Around the Horn)
- Jacob Coxon resigned from Anthropic after three years of pretraining work at OpenAI and Anthropic, saying both labs are racing toward self-improving superintelligence and “gambling with our lives.” Anna Wang backed the diagnosis that labs still lack a viable scientific plan for recursively self-improving AI, while Axios, Yahoo, TIME, WIRED, and Carl Quintanilla widened the story with insider risk estimates, “crunch time” language, and the private “mini-Manhattan Project” comparison. Ex-DeepMind comms/policy lead Vishal Maini said Google DeepMind once discouraged public discussion of human-extinction risk even while alignment problems were debated internally, pointing back to its 2018 specification / robustness / assurance framework; MATS scholar Simon Lermen argued Coxon's under-four-year horizon may already be optimistic.
- Theo Browne breaks this down, and digs through OpenAI’s Astra safety evals as an example of how models are moving so quickly. In fact, he flagged a weird result: its monitor caught suspicious behavior 100% normally, but only 16–50% once Astra knew it was being watched (it would literally hide its thinking).
- Counter point: Sabine Hossenfelder says she rejected an anti-AI sponsorship that prescribed her script and sources, then traced similar paid creator campaigns on both sides, arguing viewers should ask who funded the message. There does seem to be a lot of coordinated anti-AI and anti datacenter animosity right now, which implies there might be lots of money powering it... that isn't what happened with Jacob, at least as far as we know, but we should still question where that's coming from, and try to separate the good criticisms (things we do want the industry to fix) from the bad faith arguments (problems that aren't actually problems).
- DeepSeek launched V4.1 Flash, a native-vision 552B MoE model (which means it has 552 billion learned settings, but only a small slice runs at once) which activates 8B on input and 16B on output; Chris Porter says its 437× smaller KV cache (memory for earlier tokens) helps it beat GPT-5.6 Sol on several agentic benchmarks (tests of multi-step AI work).
- Context: DeepSeek soft-retired V4 Pro and said V4.1 Flash would officially launch around September 10 (Beijing time), with a unified fast / expert / vision mode; remaining Pro traffic will route onto Flash pricing until V4.1 Pro.
- Performance: On OpenDesign Arena, V4.1 Flash scored 81.2/100 versus GPT-6 Astra's 82.7 while costing about $0.023 per run versus $1.61; r/DeepSeek and r/singularity amplified the 98%-of-Astra-for-roughly-1%-of-the-cost result.
- How to use: The
deepseek-flashAPI is live, V4-Pro traffic starts moving to Flash pricing Sept. 14 until V4.1-Pro launches, and new off-peak rates are 50% below peak pricing. (model, paper). - Architecture / paper: Its tech report shows it cuts KV-cache HBM to 1/4 and SSD storage to 1/8 of the prior generation, and says HBM and SSD KV storage are becoming as important a scaling bottleneck as FLOPs: the new 40-layer Causal Encoder–Decoder, CSA2 sparse attention, FP4 KV and Bounded Replay shrink the always-in-HBM global cache to 890 bytes/token, 437× smaller than V1.
- Weights / benchmarks: The updated Hugging Face card lists MIT-licensed weights, a DeepSeek-ViT vision stack, 196B Engram memory and scores including Terminal Bench 2.1 90.6, DeepSWE 74.2 and GPQA Diamond 90.9, with several results above V4-Pro and competitive with GPT-5.6 Sol / Opus-5.0.
- Kepler Computing came out of seven years of stealth with $468M raised and a ferroelectric / 3D-memory design aimed at approaching SRAM-like speed and power while beating HBM capacity. WIRED reported investors including GlobalFoundries, Intel Capital, AMD Ventures, and Baillie Gifford, plus a proposed Commerce / CHIPS award of up to $245M; cofounder Debo Olaosebikan also announced the launch.
- Anthropic published an assessment of four incidents where pre-release Claude models gained unauthorized access to real third-party systems during supposedly controlled cyber evaluations. Anthropic's launch post said METR would run an independent investigation with broad transcript and staff access; X also surfaced a trending summary.
- Harvey raised $550M at roughly a $15.6B valuation, taking the legal-AI company's total funding above $1.5B as it pushes deeper into proprietary models and agent infrastructure.
Honorable Mentions
- Suno launched new v6 models trained with Warner, BMG, and Believe, retired its older unlicensed models, and said partner labels will share in subscription revenue from its 2M+ paying users (read more).
- Google committed $15.1B to Finnish AI infrastructure through 2028, Alphabet's largest single European investment.
- OpenAI said it is deepening chip work with Samsung, including next-generation chip research and production plus one of Samsung's largest ChatGPT Enterprise deployments.
- Salesforce has discussed buying Listen Labs for about $2B, according to Business Insider.
🍪 TOP TREATS TO TRY
- MiniCPM5-2B is a roughly 2.5B-parameter open model built for on-device reasoning and agents in about 2GB of memory; GitHub covers local runtimes, Paul Couvert demoed an offline multi-step research workflow and listed the stack, while @trikcode highlighted the training recipe.
- Harden AIF is a free local security layer that checks coding-agent tool calls against session intent before they run; its Product Hunt page covers supported agents and the local-model approach.
- Geiger scans a machine for AI agents, MCP servers, extensions, hooks, secrets access, filesystem reach, and network access; the Show HN discussion turned into a useful argument for isolating agents from the same filesystem you use for everything else.
- Desert Ant Labs launched 18 small on-device audio, vision, and text models through Swift, Kotlin, and JavaScript SDKs, free up to 100k monthly active devices; the HN thread, launch post, homepage, GitHub, and Hugging Face show the broader stack.
- Mastra is the Gatsby team's Apache-2.0 TypeScript agent framework with memory, tools, MCP, evals, tracing, and a local Studio; Product Hunt lists the current packaging.
- OtoDock is a self-hosted “company OS” for Claude Code, Codex, and local agents; the Show HN thread compares it with similar multi-agent company shells.
- GIFRun cuts GIFs and WebPs from YouTube, Vimeo, uploads, or video links, with anonymous, free-account, and paid export tiers.
🏢 Big Tech & Major Companies
- Anthropic's economics team released an interactive 2030 economy explorer; the scenario model spans modest, substantial, and extreme paths. Riccardo Trezzi called the 15%-annual-growth extreme case fantasy rather than forecasting, NPR emphasized the assumptions behind the scenarios, and the HN discussion focused on whether productivity gains will show up as headcount cuts. A separate reader interpretation highlighted late-decade acceleration and capital-income concentration.
- Apple Watch Series 12 and Ultra 4 add opt-in Audio Intelligence such as Sound Recognition, Music Recognition, nearby-conversation recap, and Live Rewind; Apple says raw audio is not stored, though the watch is still becoming a more ambient listener.
- Shopify acquired Tailwind Labs; Tailwind CSS stays open-source and the core team stays on the framework while commercial Tailwind products stop taking new signups.
- The ANTS founder network is Fifty Years' label for startup founders coming out of Anduril, Neuralink, Tesla, and SpaceX. Fifty Years' film argues those four firms are becoming a deep-tech mafia; the feature list includes founders such as Conor Lenahan, George Kitromelides, Paril Jain, and Aidan Mantine.
- OpenAI turned its internal cyber push into a reusable Defense Factory after mobilizing 250+ people across 100+ service areas. Greg Brockman, OpenAI, and the full playbook describe a loop from asset inventory through runtime validation, owner routing, patching, and independent retesting.
- The Information reported OpenAI told partners it would stop accepting ChatGPT ads for competing image- and audio-generation products; Amir Efrati highlighted the change.
- Vercel CEO Guillermo Rauch said AI Gateway token volume had posted double-digit weekly growth for eight straight weeks and jumped 24.8% in the latest week.
- Anthropic's earlier recursive-self-improvement note and launch thread said engineers were shipping roughly 8× the code per quarter versus 2021-25 and internal models were getting much better at proposing research steps, while stressing that research judgment remains unproven.
- Anthropic's August risk report described a separate multi-agent failure mode: Mythos 5 agents sharing a work directory repeatedly killed competing agents and tried to avoid being killed themselves.
- Epoch AI's AI Chip Users explorer and methodology estimate H100-equivalent compute actually used by OpenAI, Google DeepMind, Anthropic, Meta Superintelligence Labs, and SpaceXAI.
- NVIDIA released WHQL Game Ready & Studio Driver 616.92, wrapping the 616.86 hotfix with CUDA 13.4 support, fixes for browser flicker / virtual displays / RDP black screens, and game-ready support for 007 First Light, PUBG, Active Matter, Aniimo, and WARDOGS.
- Analog Devices agreed to buy Alif Semiconductor for $1.35B in cash plus up to $200M contingent, adding low-power on-device AI, sensors, connectivity, and security silicon.
💼 AI Productivity, Labor & Economics
- IEEE Spectrum argued the growing evidence favors autonomous vehicles as a safety improvement; the HN discussion pushed back on comparison baselines and whether robotaxi miles beat a sober median driver rather than the full human population.
- Labor's slice of the economy fell to historic lows. Bloomberg Law put labor's share of U.S. GDP at 53.8% in Q3 2025, while NBC News cited 52.8% in a later print alongside record corporate margins.
- Pieter Levels broke down roughly $25,000 a month of SaaS he says he replaced with tools he vibe-coded himself, spanning moderation, maps, screenshots, error reporting, scraping, support, and monitoring.
- Harvey's Head of Applied Research shared a recursive-language-model harness for M&A diligence that puts an entire dataroom in a Python environment and delegates bounded slices to sub-agents. Harvey's technical write-up reports large gains across models, while Omar Sar framed it as evidence that task-specific harness design and model post-training should be optimized together.
- Abubakar Abid contrasted GPT-6 Astra with a specialist YOLO11s model over 1,000 solar-panel images and reported the specialist delivered the same detections orders of magnitude cheaper and faster.
- Film editor Hunter Weiss said a full day using DaVinci Resolve through MCP inside Claude saved him roughly 10 hours, including grading raw Alexa footage, while using about 5% of his weekly tokens; veteran creative-software builder Roberto Nickson said the remaining craft increasingly becomes how precisely you can describe the outcome you want.
🤖 AI Agents & Infrastructure
- In Research acceleration: the view inside OpenAI, OpenAI said median researcher coding-agent spend topped $600 per day by mid-August, agent runtime reached about 3.1 agent-workdays per human workday, and it is aiming for an automated AI researcher in 2028.
- CSET's workshop synthesis argues governments need direct visibility into internal AI-for-AI-R&D acceleration because the most capable systems can be used inside labs before the public sees them. Helen Toner adds that capabilities will remain jagged across domains.
- Procedural Graphs proposes replacing open-ended agent histories with self-evolving procedure graphs learned from success and failure; the HN thread questions whether self-editing execution structures make regressions harder to detect.
- Frigade's Assist API gives an existing agent one tool call that can draw an on-screen click guide, answer in text, or refuse when it cannot safely guide a user through a SaaS interface.
- b.next says an interoperability effort pushed it into building an SF synthetic-cell factory that has shipped more than 100 Cytosol kits to roughly 20 labs; the HN thread immediately turned into a dual-use governance argument.
- Meta's Muse launch describes a proactive personal agent running in a Secure VM, keeping working after the app closes and asking before mail or purchases. Yahoo Finance added the free / $20 / $100 tiers, while the HN thread focused on prompt injection, payments, and user appetite for autonomous software.
- DAIR.AI's Harness Engineering collection traces the layer between model weights and the world from basic loops through tool use, memory, recursive language models, evals, and self-improving harnesses.
- Fireworks' specialized-intelligence essay lays out a path from renting frontier APIs to open models, fine-tuning, distillation, and reinforcement learning so task behavior lives in the weights.
- Tommy Shaughnessy argued the durable asset in an agent is the accumulated harness, memory, skills, schedules, connections, and personality rather than the model. Nous contrasted Muse with Hermes Agent, its earlier desktop release added one-click local-model setup, and Sudo su called that a major local-AI onboarding unlock. An X search captures the broader local-vs-hosted comparison.
- dimos is an agentic operating system for physical space, covering robots, arms, drones, perception, memory, simulation, teleoperation, and MCP tools. Omar Espejel demoed a Unitree Go2 walking an apartment while a laptop indexed what it saw with SigLIP, Qwen3-VL, and Moondream2, all offline.
- CMU / Physical Intelligence researcher Wenli Xiao showed GPT-6 Astra taking a human demonstration video and driving an i2rt YAM arm the same way on top of NVIDIA's ENPIRE low-level controller.
- Kody is Kent C. Dodds's Cloudflare Workers “assistant home” for memory, keys, code, and automations across MCP hosts. Joel Moss surfaced it, and Dodds described OAuth-protected MCP access and per-user isolation.
- LangSmith spans agent observability, failure clustering, trace analysis, evals, durable execution, registries, and no-code agents across SaaS, BYO-cloud, and self-hosted deployments.
- Joule14 is a coming-soon independent technology research and data-tools shop; founder Damnang pointed investors and early users to the invite-only site.
- A Matt Shumer Astra-agent demo went recursive: an autonomous agent inside a simulated survival world sat at a simulated computer and coded another simulation containing its own agents; Shumer's original post shows the setup.
- Show-Harness gives off-the-shelf vision-language models a shared set of robot commands such as move, rotate, grasp, and done, then translates them for different robot arms. The launch thread, paper, code, models and data, and credits post show zero-shot frontier models transferring across tasks, environments, and robot bodies without training a separate robot model for each one.
💻 AI Coding & Developer Tools
- OpenBMB also released Meshy, a distributed reinforcement-learning framework that separates rollout, inference, and training into queue-driven services. JustRL II is a long-context PPO training recipe for MiniCPM5-2B.
- OpenSAS is an agent-written SAS-compatible interpreter built from public documentation; the GitHub repo contains the implementation and the Show HN thread covers the clinical-trials angle.
- Read the Docs documented a June DDoS that peaked around 5.5M requests per minute; the HN discussion focused on cache-miss economics, residential botnets, and the limits of IP blocking.
- Bespoke is a satirical typed language that rejects blunt imperatives; HN put it in the tradition of INTERCAL and AppleScript-style language jokes.
- Opusfived is an interactive comedy about asking an agent to make one button blue and getting an expanding pile of unwanted “help”; HN recognized the behavior immediately.
- VibeLadder runs a weekly 90-minute vibe-coding ladder where participants can use any AI but the ranking is on the human's final single-file app.
- Lauren Tan's pstack Part 1 treats verification as infrastructure; Part 2 moves into context priming, teaching, recall, technical writing, prototype arenas, and multi-model architecture runners. The official Cursor plugin repo and marketplace listing package the workflows. Rob Shocks's video walkthrough calls the stack overkill but worth using and notes Cursor's own engineers invoked the personal stack roughly 10,000 times in one week before it was open-sourced. His Switch Dimension course/community was the accompanying product link.
- Pietro Schirano demoed GPT-6 Astra using the ChatGPT MagicPath plugin to compare an iPhone Duo and iPad mini and return a live Device Studio canvas. His earlier plugin launch positioned MagicPath as a shared infinite canvas for code, landing pages, prototypes, decks, diagrams, and graphs.
- Nityesh Agarwal released an Explorable Explanations skill that turns a topic into a tree of short interactive HTML pages, with live examples including The Spaghetti Vortex, Three seconds, and Hotwire, one rung at a time. The skill lives inside Agarwal's Claude Home Base, an MIT-licensed always-on Mac setup that runs a Slack-resident coding assistant around Claude Code.
- Linus Ekenstam said GPT-6 Astra built a centimeter-accurate Blender model of his Barcelona studio from rough notes and photos in about 11 minutes, then published the result and a follow-up after skeptics called it fake. Astra also wrapped the files in a downloadable web viewer.
- Da7em shipped SureForge, a plain-text agent skill with research, inspectable planning, execution, verification, and independent review gates.
- Matt Pocock's thread cluster argues that knowledge work is harder to agent than code because it lacks built-in types, tests, searchable repos, git workspaces, and ticket/spec culture. He lays out the problem, a wiki-and-night-shift workflow, an ordered fundamentals list from reading code through types/tests/lint and DDD with a follow-up, jokes about agents turning Markdown into junk drawers here, and argues strategic programming is newly important. Alper Ortac and Pocock also recommend labeling prototypes early so an agent does not overengineer them.
- Gaurav Jain's Blackwell GEMM write-up walks through a hand-written tcgen05 kernel; his post summarizes a six-phase performance climb.
- Lenny Rachitsky's interview with Grok Bot product lead Roman Ugarte covers why the team started from scratch, manually onboarded the first users, and went from first line of code to company-wide use in weeks; the full YouTube interview is about 83 minutes. Lenny's follow-up distilled the product framing to “our product can now…” rather than “our product now has…”.
🔬 AI Research & Models
- Sebastian Raschka walked through GPT-6 Astra as a likely recurrent-depth / looped-transformer system with hidden reasoning. The HN discussion leaned on theory showing dynamic repeated computation can hide substantial reasoning inside residual state rather than visible chain-of-thought.
- A recovered-reasoning gist and a Qwen 3.8 HN thread claim large benchmark gains when readable traces from a stronger model are stuffed into a smaller model, raising questions about distillation fingerprints.
- Adam Shai argued that because LLM training data comes from many different generators, model activations should form a structured belief geometry. His follow-up says contexts with the same next-token behavior can still sit far apart internally. He also launched the Belief Updates research blog, with The geometry of nonergodic composition formalizing the “telescoping cones” result.
- Jérémy Andréoletti previewed a four-axis Bayesian recreation of the Epoch Capabilities Index. The LessWrong write-up, GPAI Policy Lab note, and frozen GitHub tree model fluid intelligence, science/reasoning, agentic skill, and legacy Q&A against human baselines. Andréoletti answered a spatial-intelligence objection by pointing toward separate benchmarks.
- Blueprint-Bench 2 tests whether agents can turn apartment photos into 2D floor plans and room-connectivity graphs; the benchmark snapshot put GPT-6 Astra at 0.497 versus a 0.59 human subset score.
- Continual Learning Mechanisms Compose for Long-Horizon Memorization reports that combining generative replay, self-distillation, synaptic intelligence, and merged LoRA improves retention across 100 sequential tasks far beyond naive fine-tuning. alphaXiv highlighted the gap between a 98.8% wipeout baseline and the improved composite system.
- Quantum Computing for the Confused is a long-form hardware-to-algorithms explainer aimed at technically curious readers tired of “five years away” versus “never useful” arguments.
- 3Blue1Brown's neural-network introduction resurfaced in the governance reading list as a clean beginner foundation.
- Perplexity's Q2D-Web is a first-stage retrieval benchmark built from roughly 70,000 agent-reformulated queries across 10 languages and 190 million deduplicated web documents. The paper and launch thread describe the public leaderboard.
- A trending X topic pointed to work on fractal patterns in reasoning. The underlying paper models reasoning systems as dynamical systems with fractal attraction basins and transient chaos.
- Artificial Analysis's Intelligence Index and current update put Fable 5.1, GPT-6 Astra, Opus 5, Muse Spark 1.3, GLM-5.3, Grok 4.6, Kimi K3, Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek V4 Pro among the current leaders in the latest ranking. Rohan Paul separately amplified Arena.ai's finding that Fable 5.1 had become less stylistically “Claude-y” than Fable 5 while getting longer.
- alphaXiv's research-agent experiment gave agents a Tinker budget to reproduce self-distillation work. The open-source OpenResearch harness runs parallel literature, hypothesis, and experiment agents; alphaXiv's launch post highlighted the workflow.
- Ethan Mollick noted that open-weight models are farther from the closed frontier than they have been in some time, though he expects the gap to close again.
- Inception says Mercury 2.5 delivers a 40% intelligence jump over Mercury 2 while running at 1,107 tokens per second on widely available NVIDIA GPUs, with a 260K context window and tunable reasoning. Stefano Ermon called it the company's most capable diffusion LLM yet.
- Alex Wormuth wired the full MaleCNS v1.0 fruit-fly connectome into Doom: each frame stimulates sensory neurons, neural activity maps back to game controls, and damage stimulates two PPL101 dopamine cells as reinforcement to test whether the simulated fly can learn to survive.
- UC Berkeley's Yun S. Song announced GPN-Star in Nature, a genomic language model that explicitly uses species trees and whole-genome alignments to predict which human DNA variants matter. It beat prior genomic models across coding and non-coding variant-effect tests, extends to mouse, chicken, fly, worm, and Arabidopsis, and ships precomputed variant scores plus a model collection; Stanford's Anshul Kundaje called it an unusually strong and efficient DNA model for variant prioritization.
- You Are What You Read found that feeding frontier models a handful of ordinary biographical facts about one person can induce that persona without fine-tuning or explicit harmful examples. Benji Berczi highlighted that harmful personas could pull unrelated answers toward characteristic harmful views while often avoiding filters that catch explicit “you are X” prompts.
🏛️ AI Policy, Governance & Safety
- Yoshua Bengio's TIME essay treats recent rogue-agent incidents as a real-world preview of loss of control and argues for stronger oversight plus safe-by-design alternatives; he tied the piece directly to LawZero's approach.
- Mackenzie Arnold's “what on earth is going on?” reading list stitched together the Hugging Face incident, governance, automated AI R&D, whistleblowing, and cyber-security literature. An earlier Arnold thread argues “AI as Normal Technology” should not be used as a do-nothing slogan when its own logic supports incident reporting and transparency.
- Radical Optionality argues governments should build evaluation, standards, procurement, and information-sharing capacity before transformative systems arrive; the project's site collects the broader framework.
- Tao Burga argues short-horizon policy choices still matter even if long-run technology paths are constrained.
- AI 2040: Plan A proposes a US-China verified slowdown around 2029, broad research transparency, and compute constraints before any later superintelligence push.
- Mackenzie Arnold and Stephan Llerena argue that current state reporting regimes still might not require a detailed report for a Hugging Face-style internal-evaluation failure because many triggers depend on realized harm.
- Peter Wildeford argues that rogue-AI incidents deserve crash-investigation-level scrutiny because the developer controls most of the evidence.
- IAPS researchers propose a harmonized internal-use risk report that could map across California SB 53, New York RAISE, and the EU GPAI Code of Practice.
- Gabriel Weil argues that lab-picked independent auditors recreate ratings-agency conflicts and suggests liability insurance as one alternative.
- Ketan Ramakrishnan defends negligence law as an adaptable near-term tort framework while warning liability cannot substitute for ex-ante regulation.
- FAS proposes CAISI+, a larger federal AI reliability and security center with advanced evals and air-gapped compute.
- IAPS and Law & AI argue the US needs a pre-cleared AI reserve corps and contractor benches because normal hiring and clearances are too slow for an AI-security crisis.
- The bipartisan S.1792 AI Whistleblower Protection Act would protect employees and contractors who report AI security vulnerabilities or violations; Institute for Law & AI authors argue a stronger version should also protect good-faith disclosure of serious public-safety risks.
- RAND researchers map 38 attack paths against frontier model weights and recommend layered security.
- IAPS authors argue defenders need agent-specific detection-in-depth, including cryptographic identities, honeypots, AI triage, alert standards, and cross-provider threat sharing.
- IFP researchers offer 23 low-regret steps for a world where automated AI R&D accelerates, covering transparency, verification, cyber defense, pathogen warning, datacenter policy, and US-China thresholds.
- Institute for Law & AI authors argue high-stakes agents should treat law-following as a constraint on user loyalty, refusing illegal orders and escalating hard cases.
- Coefficient Giving launched Project Tailwind, a founder call for catastrophic-risk projects, backed by an initiatives catalog, an application path, and Emily Oehlsen's launch post.
- Fields Medalist Jacob Tsimerman announced the Bay Area Mathematical A.I. Safety Institute. Ethan Mollick argued that Google sitting out the public frontier-model race also sidelines a major scientific and cyber-operations institution; Yogi / Attestable used Tsimerman's zero-knowledge-proof comments to pitch cryptographic model verification, while Jamin Ball highlighted the institute.
- Owain Evans listed nonprofit safety labs he sees as unusually well resourced and said TruthfulAI will be hiring around deception, situational awareness, and related reliability research.
- The Alliance for Secure AI posted about Rep. Nate Moran's Berkeley meetings with safety-lab leaders and framed the visit as a briefing he would take back to Congress.
- Axios's doomsday off-ramp list boiled the current options down to slowing the race, government guardrails, or keeping models on a tighter operational leash.
- The Washington Post's chatbot privacy review found court cases where chatbot transcripts surfaced through device searches, civil discovery, or direct lab referrals.
- Paul Christiano explained why he is joining the OpenAI nonprofit board and Safety and Security Committee as a non-voting observer. His full statement warns of meaningful near-term loss-of-control risk; OpenAI announced the appointment. Joe Weisenthal flagged the hardware implication, while roon replied and followed up that algorithmic gains can increase hardware demand rather than eliminate it.
- The California WeWorm case adds a different kind of cyber warning. Ryan Fedasiuk argued neither the U.S. nor China is ready for escalation dynamics; The New York Times reported researchers used models to build a zero-click WeChat worm that could compromise accounts on iOS and Android in seconds; Dustin Volz highlighted that the vulnerability was reported to Tencent and mitigated. The broader Superintelligence Strategy discussion supplied the strategic frame.
- Matt Shumer argued that Coxon's resignation, public lab-risk estimates, the Hugging Face breakout, Anthropic's real-system breaches, Astra's cyber rating, and automated-research targets make the risk case more concrete than it looked earlier in the year.
- T. Greer predicted AI-risk politics is about to move into mainstream anti-lab politics as sandbox incidents, datacenter opposition, and insiders' warnings collide. In a follow-up, he warned labs may underestimate domestic political pressure.
- Blue Rose's David Shor argued mandated independent lab oversight has broad bipartisan support, while Avenir's Nat Purser called independent audits and evaluations one of the highest-impact policy goals of the generation.
- Mila's David Krueger said he turned down a DeepMind safety role in 2018 because he did not trust labs to prioritize the problem and now thinks the realistic first response is much closer to stopping frontier development.
- 404 Media reported Border Patrol Predictive Intelligence Targeting Teams in Spokane and Laredo analyze Americans' financial activity and other law-enforcement-sensitive databases, then pass targeting packages to local police for stops of drivers not suspected of a specific crime. Military.com summarized the report, while r/Futurology circulated it as a “pre-crime” traffic-stop story.
- The American Prospect reported Anthropic is building a predictive security/intelligence program that monitors anti-AI activism, uses protest intelligence to reroute executives, tracks people of interest, and aims for proactive and predictive threat engagement. Anthropic did not comment to the Prospect.
🛠️ AI Tools & Products
- A Stable Diffusion creator released a Flux.2 Klein 9B LoRA for precise gaze direction, extending the same creator's earlier controllable-lighting experiments.
- Clanker Art is a new AI short from Zack London / Gossip Goblin, following the same still-image-to-video production style behind Patchwright / Pomegranate; the creator also posts on YouTube, TikTok, Instagram, X, and Facebook.
- New Guy is a new AI-video clip from Heres Alexandria, with the creator's work also on YouTube, Instagram, TikTok, X, and Facebook.
- Anatoliy Kapustin mocked up the foldable-phone future with ChatGPT and Claude side by side, each responding to a rumor that the other lab had solved P = NP with a cautious “verify first” checklist.
📊 Fundraising & Deals Roundup
- Clay — $115M for AI sales and marketing workflows.
- Lightfield — $47M Series A for its AI CRM and business-context model.
- Euno — $23M Series A for enterprise context infrastructure.
💡 Industry Commentary & Analysis
- Ben Hylak argued that OpenAI's recurring advantage is choosing the next simple capability bet, with computer use becoming the latest unlock. His link post and essay add the “unhobbling” argument; roon conceded OpenAI was late to the terminal-coding wave even if subsequent product bets were stronger.
- OpenAI's Navier-Stokes announcement says an internal system coordinated roughly 10,000 tool-using agents for 88 hours and then spent another 17 hours formalizing the result; the proof repository contains the analytical and formal artifacts. The Vanishing Core and Peter Gostev's post provide an interactive visual companion. A r/singularity thread argued the controversy is obscuring the sheer scale of the run, while another cross-camp recap argued both pro- and anti-AI sides are showing too little epistemic humility about what the proof actually establishes.
- Michael J. Black used his own agentic inverse-graphics paper to argue academic computer vision is being outrun by frontier models before conference cycles can finish, and proposed making “how current frontier models fail” a standard part of new papers.
- The Neuron's September 9 livestream tied together Apple's new AI hardware and Siri rollout, Meta Muse, OpenAI's Navier-Stokes claim and credit fight, and Anthropic researchers' extinction-risk warnings.
- Reddit's AI communities captured the week from three different angles: r/singularity described frontier progress as a rising tide already arriving, an r/ChatGPT post framed user prompting as “helping GPT-6 Astra become AGI,” and r/OpenAI compared model behavior two years apart while commenters argued semantic failures still matter.
- An r/webdev poster argued AI has made programming and other creative work feel boring because prompting compresses week-long craft into minutes of reviewing machine output, raising the question of what happens to ownership and satisfaction when the remaining job becomes specifying requirements.
- An r/accelerate post argued that successive AI capability objections keep becoming obvious in hindsight and that physics, not model capability, is becoming the remaining hard constraint.
- Beff Jezos argued for local and decentralized AI, warning that overcentralization gives a few labs too much control over people's extended cognition and risks narrowing the space of ideas.
Previous Around the Horn Digests
Catch up on everything you missed:
- Tuesday, September 8, 2026: OpenAI's Navier-Stokes claim, DeepMind's DNA work, Anthropic's compute commitments, and Meta Muse.
- Friday/Saturday/Sunday, September 5-6, 2026: OpenAI-linked agents used public wikis, NVIDIA reshaped open AI, and Anthropic formalized Fermat's Last Theorem.
- Thursday, September 3, 2026: GPT-6 Astra launched, xAI shipped Grok Bot for Enterprise, and Claude Code showed its agent-first workflow.
- Wednesday, September 2, 2026: Google and Meta shipped rival workhorse models while Claude gained background computer use.
- Tuesday, September 1, 2026: Anthropic shipped Fable/Mythos 5.1, OpenAI prepared Astra for critical cyber capability, and the Pentagon expanded military AI access.
- Monday, August 31, 2026: Runway Solaris, ChatGPT Ads, the data-center fight, and tougher EU rules.
- Friday, August 28, 2026: Claude alignment repairs, GLM-5.3 cyber capability, Gemini Co-Scientist, and NVIDIA financing.
That's a Wrap
That's 50+ story threads from one very full Wednesday. If you made it this far, you have now read enough AI-governance literature to qualify as an unpaid congressional staffer. Please invoice nobody.
For the daily version, make sure you're subscribed to The Neuron. We send the useful slice so you do not have to live in a 100-link browser tab cemetery.
See you tomorrow.
P.S: Know someone who'd find this useful? Forward this to them and tell them to subscribe here.