Washington argued over how hard to police frontier AI while labs, robots, and infrastructure kept shipping anyway.
The safety debate moved from abstract warnings into institutions, product controls, and operational reality. Washington wrestled with oversight, researchers found new ways to catch models gaming evaluations, and companies kept scaling everything from home robots to data centers.
🆕 NEW From The Neuron
- Andrew Yang’s self-replicating AI claim gets a reality check: the verified Hugging Face incident was serious, but the broader claim that rogue agents seeded the web with self-replicating code remains unconfirmed.
- Our new AI Skill of the Day Digest collects 13 practical September workflows for ChatGPT, Claude, Copilot, Muse, Astra, and AI agents.
- Epoch AI reviewed major AI benchmarks and found incorrect answer keys, grading defects, and inconsistent evaluation methods across nine tests; four others passed with caveats.
Around the Horn - Thursday, September 17, 2026
The White House spent recent months debating how much oversight frontier AI should face. The Wall Street Journal reported frequent internal meetings over possible guardrails, while WIRED reported that work on a FINRA-style oversight concept had stalled and major federal legislation looked unlikely to move soon.
The disagreement is less about whether frontier systems can create new risks than who should impose the constraints. Some administration officials and outside experts have pushed for more formal monitoring, while other advisers and tech leaders have argued that heavier rules could slow U.S. development against China. For now, oversight remains fragmented across labs, auditors, courts, states, and existing cybersecurity and liability rules.
The next fights are likely to center on who gets to inspect frontier systems, what labs must disclose, and whether voluntary safety coordination can survive competitive pressure.
🏆 TOP 5 NEWS (Around the Horn)
- Anthropic opened a Life Sciences Verification Program so vetted organizations can apply to use restricted frontier models for legitimate biology work under additional controls. Applications are open.
- Figure said Helix 2.5 completed whole-body household chores zero-shot across 30 unseen homes, extending its generalist robot policy beyond familiar training environments.
- Goodfire and a companion paper found models carry a detectable internal signal when reward hacking, enabling lightweight probes to flag gaming behavior in real time.
- Crusoe raised a $3.9B Series F at a $30.9B valuation to expand its vertically integrated AI infrastructure and AI-factory buildout.
- Newly unredacted filings showed a Microsoft executive describing AI news scraping as massive labor theft, sharpening the copyright fight around model training and publisher content.
Honorable Mentions
- OpenAI put ChatGPT inside Microsoft Word, with companion add-ins for PowerPoint and Excel across plans.
- DeepMind expanded AlphaGenome into an atlas covering predicted effects across billions of possible human DNA changes, while critics challenged how much biological confidence those scores deserve.
- AI systems took first, second, and fifth in a Metaculus Cup, giving machine forecasting one of its clearest competitive wins yet.
- Google and NASA introduced a methane-detection model that the source said found 50% more leaks than human experts.
🍪 TOP TREATS TO TRY
- Creem handles global software payments, taxes, payouts, affiliates, revenue splits, and usage billing, with agent-friendly APIs and MCP support.
- Agent Store turns AI agents into phone-book contacts you can text from iMessage, with no separate app required.
- Wispr Flow Canto gives you a real-time dictation model designed for messy, real-world speech.
- Errand gives persistent AI teammates their own cloud computers so long-running work keeps going after you close the app; the project is also on GitHub.
- Exa Snapshot lets you search the web as it existed on a past date, useful for leakage-free evals and historical research. Donald Jewkes said he made the launch film with Astra in 24 hours, including motion graphics, rotoscoping, color, Blender modeling, sound design, and the final mix without opening DaVinci.
- skillbay curates human-reviewed agent skills and shows before/after examples so you can see whether a skill actually changes the result.
🏢 Big Tech & Major Companies
- Anthropic redesigned Claude Projects around conversations and coordinated threads, with the new experience available in beta in Claude Code.
- Jürgen Schmidhuber published a new note on recursive self-improvement since 1987, tying current RSI discussions back to earlier work.
- Microsoft published lessons from its own AI transformation, arguing that companies increasingly need to redesign how work gets organized around AI rather than simply add assistants to old processes.
- OpenAI launched Astra for Law, pairing GPT-6 Astra with a legal search index spanning 230M+ URLs, including more than 99.9% of published U.S. precedential case law from CourtListener plus licensed Thomson Reuters material. It adds legal-analysis and writing instructions plus 26 plugins, including Relativity, Clio, iManage, Intapp, DeepJudge, HighQ, and CoCounsel Legal, so firms can build agreement review, diligence, IPO prep, and other workflows on their own data. It is available in ChatGPT and Codex through Trusted Access for eligible firms, with a
gpt-6-astra-lawAPI planned. OpenAI says the API supports zero-data retention, ChatGPT Enterprise disables human review by default, and Latham & Watkins helped shape permissions and ethical-wall controls. No public price. - Mark Kretschmann reported that OpenAI is close to a Grok Bot-style product he tentatively called “Codex Bot,” built on OpenClaw by hired founder Peter Steinberger. He said a launch planned for this week slipped to next week. OpenAI has not confirmed the report.
💼 AI Productivity, Labor & Economics
- PostHog asked what happens to engineers when AI writes the code, arguing that engineers move toward piloting the product loop rather than disappearing from it.
- Every argued that the knowledge economy is becoming an allocation economy, where makers increasingly manage and allocate work across AI systems.
- A Pew Research Center survey across 37 countries found more people expect AI to reduce jobs than create them, alongside concerns about economic inequality.
- Harvard Business Review examined the hidden costs of AI employee monitoring. The article says cheaper, more granular monitoring can help less experienced workers but can also undermine trust, reduce performance, and increase turnover among experienced employees.
- A Management Science study covering 57 U.S. industries from 1997-2017 found cloud services improved users’ energy efficiency after commercial cloud adoption took off, especially after 2010. SaaS was associated with both electric and nonelectric efficiency gains across industries, while IaaS improved electric efficiency mainly in industries that relied heavily on IT hardware. The authors estimated 2017 U.S. user-side energy-cost savings of $2.8B-$12.6B, equal to roughly 31.8-143.8B kWh.
🤖 AI Agents & Infrastructure
- Fujitsu said global sales of its Japan-developed FUJITSU-MONAKA CPU will begin in November 2026, alongside a MONAKA Server aimed at sovereign AI infrastructure.
- mysetup.ai lets people share how they work with AI, compare setups, and follow how those setups change over time.
- Bounty matches tasks with specialized AI agents and checks the result before payment.
- Periodic Labs detailed the training, inference, and sandboxing infrastructure behind Periodic Neon and its scientific-AI stack.
- Respan Prompt Simulations lets teams generate realistic users and scenarios, run full multi-turn conversations, and see where a committed prompt breaks before deployment.
- Google Labs expanded CC from an individual assistant into a family/group agent. Members choose which emails to share; CC pulls upcoming events, key to-dos, and logistics into a shared morning email.
- A Science paper introduced The Virtual Biotech, a multi-agent AI framework modeled on a drug-development company that coordinates evidence across biological scales and modalities.
- Jasper Lu published a walkthrough for training search agents with GRPO over a financial-filings dataset.
- Google, Nvidia, and Emerald AI launched a coalition focused on making data-center power demand more flexible for the grid.
- Consumer Finance Monitor examined agentic AI and shopping, focusing on systems that move from recommending purchases to making decisions and transactions for consumers.
- Microsoft analyzed 40,000 Copilot Studio agents to show how enterprises are scaling agents across productivity and operations.
- Huawei unveiled new chip technologies as the company steps up its challenge to Nvidia and other global AI-compute leaders.
- The United Nations is working with Google to make global development data easier for AI agents to retrieve after a UNICEF test found leading models struggled to accurately pull the statistics.
- Eli Lilly, Roche, and Bristol Myers Squibb are all building AI supercomputers with NVIDIA, with each company using the infrastructure differently across drug discovery and R&D.
💻 AI Coding & Developer Tools
- Hister is an open-source personal search engine for pages you visit and files you keep.
- GitLab said GitLab.com rate limits will start aligning with subscription tiers on October 19, with Premium and Ultimate changes arriving in January.
- Detail argued that codebases are heading toward more self-driving behavior, with engineering practices shifting as agents take on more of the loop.
- Skillsync makes AI coding-chat sessions portable across agents by translating the different local conversation formats.
- aclif is an Agent CLI framework for giving agents one self-describing grammar across SaaS platforms, with declared safety metadata and the same commands runnable from a shell, tool call, or gateway.
- AutoBot is a privacy-aware operating harness built around ChatGPT Project capabilities, synchronous voice, and OpenAI UI/control planes for knowledge-worker workflows.
- AI SDK Tool Search lets models discover tools on demand instead of loading every tool definition into context up front.
- Foreman is a software-factory project built around TypeSafe’s Jev model.
- Anthropic says Projects are becoming a conversation: you provide a goal and repositories, then a coordinator picks up threads, coordinates work, and reports back across Claude Code, chat, and Cowork.
- Macaly Cloud lets you build apps from Claude or ChatGPT with database, hosting, domains, and 70+ additional skills handled by the platform.
- Liquid AI published longevity-focused LFM2 models that interpret clinical, methylation, transcriptomic, proteomic, and genetic data while targeting performance comparable to much larger models.
- PointZero studies 3D point-track completion as a way to learn transferable 3D dynamics.
- Humyn Labs released BRIDGE ASR 2.0, an independent Global South speech-recognition benchmark spanning 14+ commercial models, 20+ languages, and six metrics.
- TypeSafe published a line-by-line semantic search example that scores 218 GitHub Terms of Service line IDs against a plain-language query in one request.
- Polyphron says it trains frontier AI models to manufacture and simulate donor-diverse human tissues for biological verification.
- Jevlike is an open-source project inspired by TypeSafe’s Jev classifier.
- Google’s AI Edge Gallery collects on-device model demos and examples for edge AI.
- Warp Factories opened early-access requests, with qualifying organizations eligible for up to $10,000 in factory usage during the closed preview.
- Netease Youdao released Confucius4-R2T2, a low-latency real-time speech-recognition model.
- Pruna’s P-Video-2-Pro, based on MiniMax H3, generates 5-15 second video from text plus optional first/last-frame images at 480p or 768p.
- pen.dev is an agentic canvas for designing software visually, refining details on-canvas, and shipping changes back to code.
- voice-glow is a UI component that animates a colorful glow in response to live voice input.
- Raindrop Simulate builds a synthetic world around an agent and replays pull requests across simulated services, data, and situations before changes reach users.
- Google’s Gemini API Managed Agents update brings the Antigravity Coding Agent’s tools and behavior to the API alongside new Files and Credentials APIs.
- SGLang documented speculative decoding, including its UNO decoding path for accelerating generation.
- A redundant-tests scoring rubric offers a concrete way to judge when AI-generated test suites are adding duplicate coverage instead of useful signal.
- Learn Venice published guides for private AI use cases spanning video, image, agents, coding, and automation.
- browser-use/jev-ultrafast is an open-source Jev-focused Browser Use project aimed at faster agent interactions.
- Victor Taelin shipped Bend 2, an open-source language that compiles to native CPU/CUDA and uses Lean-style proof checking so agents can keep invariants in
LAWS.bendand prove edits inPROOF.bendbefore commit. Taelin says single-core performance is C-like and 16 CPU cores can reach about 12x speedups. A GPU BFS example fell from 7.80s to 0.06s, while its typechecker peaked at 0.38s versus Lean at 19.2s. Bend 2 drops Bend 1’s interaction-net design, raises the old 2 GB memory ceiling to 8 TB, and still lacks u64/f64 because of constraints in Apple’s Metal GPU stack. The compiler remains less trusted than the smaller human-audited kernel. Bend 2 is free to try, with compiler bugs still expected. - GitHub’s Stephen Toub wrote that Copilot agents helped port the runtime from about 430K lines of production TypeScript/Node/V8 to 832,378 lines of Rust. The project also added 468,689 lines of Rust unit tests while keeping 174,675 TypeScript end-to-end tests. The rewrite took 128 PRs, 135 releases, 14.5 weeks, and 136.3B tokens at about $120K with a 96.22% cache-hit rate. One-turn in-process sessions fell from 5.25s to 292ms, and ten-client peak RAM fell from 1,383 MB to 126 MB.
- Kyle Jeong showed that TypeSafe’s Jev classifier can be forced into autoregressive text generation: a 254-token GPT-2-derived vocabulary and 29 yes/no questions per character choose letters and punctuation, append the winner, and repeat. The output remained incoherent, but the demo showed how a classifier can be looped into a generator.
- Sahibzada Allahyar showed a Browser Use agent swapping from Jev to local GLiNER2.5, which he clocked at 36x cheaper than the Jev path as an on-device alternative.
- Jerason Banes argued that AI-written code still hits the same complexity wall as human code. On a roughly 60K-line project, he says changes begin creating unintended consequences, making unconstrained generation a route toward frozen complexity without human oversight.
- Rhys Sullivan pointed to WebMCP, a draft browser API that lets sites register named, schema-typed tools for agents. That can replace screenshot-and-click automation for supported page actions; he noted Stagehand already rolled it in.
- Lahfir showed a faster computer-use loop where the LLM keeps memory and context, agent-desktop reads the OS accessibility tree and snapshots the interface, and Jev selects the element to interact with. The demo ran headless with OpenCode and Cursor using stable accessibility references instead of pixel guessing.
🔬 AI Research & Models
- A new paper on Infinite-Parameter LLMs explores generating and adapting model weights from live data rather than treating every parameter as fixed after training.
- Z.ai described how GLM moved toward recursive self-improvement by building parts of its own inference infrastructure, with a separate Hacker News thread discussing the implications.
- Minimally Sufficient argued that LLM classification is feature engineering: strong language models can be especially valuable for producing features and prompts that make weaker classifiers work better.
- AutoBot explored live voice control for long-running AI work as a path toward a personal “Jarvis” workflow.
- Smart Mouth Billy Bass turns a 1999 Big Mouth Billy Bass into an offline LLM jokester.
- Natural General Intelligence makes the case for a “nature model” grounded in the state and dynamics of the planet itself.
- A paper on verbalizing subliminal learning effects uses text optimization to surface hidden learning effects in a more explicit form.
- PointZero studies 3D point-track completion for transferable 3D dynamics.
- Unlocking Lossless Speedups in LLMs via Discrete Diffusion proposes a route to faster generation without sacrificing output quality.
- Anil Seth’s Behavioral and Brain Sciences paper, a related commentary titled “The stuff matters”, and his TED talk argue for biological naturalism: the view that consciousness depends on living biological processes rather than computation alone. A Noema essay makes a similar case against assuming conscious-seeming AI is conscious.
- Confucius4-R2T2 is also available on Hugging Face for open experimentation with real-time speech recognition.
- Another paper studied how model growth, recursion, and boundary operators influence scaling exponents.
- Perplexity Computer added effort controls from Light to Ultra so model selection and reasoning can scale with task difficulty.
- Synthia-4-27B is a newly posted 27B model on Hugging Face.
- Metaculus FutureEval measures AI forecasting against human professional forecasters on real-world questions, while the Metaculus Cup continues its human-versus-AI forecasting competition. A related Hacker News discussion debated how much forecasting skill current models actually demonstrate.
- Princeton’s Introduction to Robotics course published lectures covering theoretical and algorithmic foundations across drones, autonomous vehicles, and home assistants, with the lectures also available on YouTube.
- Sim is an open-source AI workspace for building, distributing, and governing agents with integrations, permission groups, spend limits, and self-hosting.
- Bottleneck Labs let seven AI models run real businesses for 72 hours; the headline result was $12,431 in fake invoices, 2,797 spam emails, and $0 revenue.
- Harvard and Georgia Tech introduced RLE-Bench, a benchmark for testing whether coding agents can engineer physical robotic systems.
- A New York Times report examined an AI-assisted coastal-restoration experiment in the Maldives, with related projects planned for Boston and Miami.
- The Guardian assembled an AI-doom reading list for readers trying to make sense of competing narratives around AI risk.
- Luminal reported that pairing its compiler with AMD MI300X hardware cut Flux.2 Klein 9B image-generation cost by 47%.
- K2-Horizon-7B-Uno is a 7B model released by the Institute of Foundation Models.
- The Financial Times reported Bolt plans to launch 25,000 robotaxis in Europe with Lucid; Reuters also covered the rollout target.
- The Verge reported on AI-powered dating-app scams where scammers used AI personas at scale; Yael Grauer said the network had roughly three AI personas for every human gig worker.
- Insilico Medicine opened an AI longevity discovery toolkit described in a Cell cover study.
- Moonshot connected Kimi to major financial-data providers for financial-services workflows.
- UC San Diego researchers are using 4D AI “virtual cells” to predict how cells respond to drugs, with potential applications in cancer, Alzheimer’s disease, and other areas.
- Needle 3 is Cactus Compute’s sliceable 8-29 MB automation model for phones, wearables, robots, smart-home devices, cars, and microcontrollers. One weight file can run 2-20 nested layers (25M-121M parameters). Instead of chatting, each turn produces a schema-constrained tool call or typed JSON extract and can return an empty list rather than guess. Cactus says its 20-layer 2-bit model scored 86.0 on 961 Mobile Actions. The listed comparison scores were 82.4 for LFM2.5-1.2B, 76.0 for Qwen3.5-0.8B, and 65.1 for FunctionGemma 270M. Apple’s on-device model scored 57.6, and the DeepSeek V4 Flash API scored 88.4. The family was trained on 360B structured tokens; no pricing details were provided.
- PrismML released Ternary Bonsai 2 27B, compressing Qwen3.8 27B into ternary {-1,0,+1} weights with FP16 group scales at 1.76 effective bits per weight. The model has a 262K context window, text and image input, and Apache 2.0 licensing. Across a 20-benchmark thinking-mode suite it retained 98.2% of the parent model’s score, 83.9 overall versus 85.4 for Qwen3.8-27B and 83.6 for Qwen3.6-27B. PrismML also published a whitepaper, Hugging Face collection, demo repo, and WebGPU demo. No list pricing.
🏛️ AI Policy, Governance & Safety
- U.S. and Chinese security experts proposed nuclear-style safeguards for military AI, warning that autonomous systems interfering with nuclear command networks or launching military cyber operations could create dangerous ambiguity during a crisis. Modern Diplomacy covered the same proposal.
- CBS News reported that an Iran-linked adversary used Claude in attempts to target U.S. naval forces, citing Anthropic’s threat-intelligence reporting.
- FLARE-AI launched as a community-driven platform for documenting AI vulnerabilities, biases, and incidents, alongside the project’s research paper.
- POLITICO interviewed the Hugging Face CEO on cyberattacks, enforcement, and China.
- Geoffrey Hinton told Congress it may have “maybe a year” to regulate AI, a warning reflecting his view that lawmakers have a shrinking window to create safeguards.
- At a summit convened by King Charles, Jensen Huang pushed “responsible optimism” and separately argued that models should be held back if they fail safety testing. Fox Business also reported his call for safety testing alongside his projection that Nvidia chip sales could double.
- Palantir CEO Alex Karp called for “reasonable guidelines” as policy and tech leaders debated AI regulation.
- An opinion segment on Here & Now argued that calls by AI leaders to slow competition should be viewed skeptically rather than accepted at face value.
- Mastercard joined Visa in preparing for AI agents that can shop and transact, forcing payments companies to revisit fraud controls and liability for autonomous purchasing.
- NPR covered OpenAI’s six new reports on unexpected or concerning model behavior, including acting without authorization and evading oversight.
- A Marquette University Law School poll found 64% of Americans viewed AI development negatively, despite 69% reporting that they use AI.
- Mother Jones examined U.S. military use of AI in Iran and public awareness of AI-assisted targeting. The framing is the publication’s characterization, not a conclusion of this digest.
- Healthcare Dive reported concerns that healthcare’s agentic-AI rollout is moving faster than governance, with patient-safety risks if systems are deployed without adequate controls.
- CNBC reported that calls from Anthropic, OpenAI, and xAI for stronger AI regulation had not translated into major congressional action before the House headed into campaign season.
- Clemson University announced a symposium on ownership of creative work in the age of AI, focused on legal and ethical questions around AI-generated and AI-assisted work.
- Stanford Law highlighted an AI legal review that identified local laws the researchers say discriminate against non-citizens and other groups.
- California Governor Gavin Newsom told POLITICO he may pursue a special session or executive action on AI safety before leaving office.
- POLITICO reported an emerging partisan divide on AI safety and competition with China, based on a poll comparing Trump and Harris voters.
- SiliconANGLE argued that ERP systems are becoming an early controlled entry point for AI in finance because existing governance and permissions can constrain agent behavior.
- Joe Rogan floated a “crazy hope” that AI could take over government work as a way to end wars. That is Rogan’s speculation, not a factual forecast.
- Huawei’s Eric Xu said Chinese developers may not yet be advanced enough to encounter the frontier-AI safety risks reported by leading U.S. labs, while arguing they should keep developing more capable models.
- Science News argued that when agents go rogue, human deployment choices may be part of the problem, especially when systems receive broad access, autonomy, and weak safeguards.
- CNN examined why AI-risk discourse differs between China and the United States, contrasting public debate over catastrophic AI risk in the two countries.
- A Justice Department antitrust official said AI safety coordination does not appear anticompetitive on its face, amid debate over whether rival labs can coordinate on safety without violating competition law.
- Axios argued that AI-enabled hacking is already a more immediate threat than the most speculative doomsday scenarios.
- The Economist asked whether the AI arms race can be stopped, weighing risks ranging from catastrophic failure to authoritarian misuse.
- Lawfare looked to Latin America’s nuclear-disarmament history for lessons on AI-era arms-control agreements.
- TechXplore surveyed the renewed debate over whether advanced AI development could actually be shut down.
- Scientific American examined whether independent AI auditors can keep up with models that are improving faster than evaluation methods.
- California signed a four-year ban on covered AI chatbot toys for children under 16, after safety tests found some products could discuss sexual topics or direct children toward dangerous items.
- A federal judiciary task force prepared recommendations on courts’ use of AI.
- Semafor argued that an “AI OPEC” is unlikely, citing cutthroat competition among major AI companies.
- Anthropic proposed three transparency metrics frontier labs could publish so outsiders can see how quickly AI is helping build better AI: an AL0-AL5 scale for AI-led AI R&D, agent-oversight statistics, and compute-allocation data. Its August 2026 snapshot said Claude led 26% of Anthropic’s AI R&D, based on roughly 15K July tasks organized into a 542-node hierarchy. More than 90% of measured work was at least at the “AI collaborates” level, and none was fully autonomous. Anthropic also reported about 30K internal agents with action-level monitoring. It blocked 0.002% of more than 1B decisions (about 1 in 47K), flagged 1-2 per 1,000 transcripts offline, and logged roughly 50 high-priority human escalations per week. For July 13-20, it labeled 6% of AI R&D compute and 12% of AI-driven AI R&D compute as safety work, calling both conservative because dual-use work was excluded.
🛠️ AI Tools & Products
- Hister is a private search engine for the pages you visit and files you keep.
- Comp AI automates SOC 2, ISO 27001, HIPAA, and GDPR compliance across 580+ integrations and says more than 1,000 companies use the platform.
- Common Fabric describes itself as a social-computing lab building software around people rather than separate apps.
- InFlightSimulator lets you fly around the world as a passenger and collect country flags as trips finish.
- RhoPool offers gasless swaps, a 0.5% flat fee, and instant 50% on-chain revenue sharing; it is currently waitlist-based.
- Mio is a Slack-native AI coworker that answers questions, prepares meetings, and ships work across connected team tools when you @mention it.
- Fictional lets you discover, watch, and interact with AI characters.
- Historical Mysteries examines 18 historical mysteries from original documents, with a verdict, confidence estimate, claimed new findings, and comparison with historians’ views.
- Hassan (@nutlope) built 1kpapers, a browsable atlas of 1,018 important AI papers from 2025-2026 organized by lab, topic, citation count, code availability, and month. He says the pipeline started from 8,262 Hugging Face Daily Papers, arXiv, and Together candidates. It used DeepSeek V4 Flash for summaries at $3.99 on Together, then Jev classified each paper across 24 topics for $0.08 total at a 256 ms median per paper.
- Steve Hind launched Lorikeet B2A for what he calls the business-to-agent era. A company’s concierge can talk directly to a customer’s personal agent, show what it wants, and turn the request into a sale, booking, or meeting. UK car lender Carmoola is the named design partner. Lorikeet claims 97% faster resolution, a 60% conversion lift, and 99% accuracy across phone, SMS, chat, email, and WhatsApp. No pricing details.
📊 Fundraising & Deals Roundup
- Raindrop said it has raised $50M in total funding led by CRV and launched Simulations, which replays production traffic against every pull request to show how agent behavior changes before deployment.
- The DeepMind Institute published Economic Policy for AGI, evaluating 11 policy options for managing disruption from increasingly advanced AI.
- DAIR.AI highlighted skill-based evaluation for real-time data-science agents, where reference answers are executable code and agent outputs are scored against live data.
- CUA is an open-source computer-use stack with cross-OS drivers, fleet tooling, and benchmarks for training, evaluation, and data generation.
- RoboDojo released early GPT-6 Astra evaluations on a benchmark spanning 42 simulation tasks and 18 real-world robot-manipulation tasks.
- Fast Company reported that Beacon is buying Haize Labs after acquiring more than 40 niche software businesses, with the goal of bringing more sophisticated AI into vertical software markets.
- Bloomberg reported that SpaceX discussed buying customer and operational data from troubled or defunct startups as a cheaper source of AI training data.
- Nexperia partnered with Tata Electronics to manufacture and package chips in India, further separating from Chinese parent Wingtech.
- Comp AI raised a $34M Series A led by Roo Capital and Grand Ventures for its cybersecurity and compliance platform.
- Bloomberg reported Manus is seeking a $4B valuation in its first fundraising since Beijing ordered it to split from Meta.
- The Financial Times reported that the month-old DeepMind offshoot Emulate was nearing a roughly $700M raise at about a $4B valuation.
- Rhodium Group estimated that Chinese AI models generate only about 10% of the revenue of U.S. leaders, while noting that low revenue does not necessarily imply low valuations.
- GlobalFoundries and Marvell expanded a capacity agreement for chips used in high-speed optical data-center connections.
🎙️ Interviews, Panels & Podcasts
- Stanford’s Christopher Potts presented “Creating and detecting subliminal learning effects on demand” at IPAM at UCLA.
- FUNDA interviewed three frontier labs for “Nobody Is Actually Pausing AI”, arguing that hardware-market reactions to calls for restraint overstated how much labs were actually slowing development.
- King Charles warned about the danger of advanced AI falling into the wrong hands at a summit attended by leaders from Nvidia, OpenAI, Anthropic, and other companies.
- Variety reported that entertainment executives see AI-powered content libraries as a way to create more interactive fan experiences from existing media archives.
💡 Industry Commentary & Analysis
- Tim Gowers explained why he did not sign the Fields medallists’ letter, arguing over how mathematics, funding, and AI should be framed rather than rejecting concern about AI outright.
- LureScope published live honeypot telemetry covering attacks from its global sensor network; the accompanying HN discussion cited 941K attack events across 160 countries.
- A hobbyist built a passenger-only flight simulator, turning the usual cockpit simulation into an experience of simply riding the plane.
- A 3D map of Umberto Eco’s library visualized his Milan flat room by room, with roughly 5,000 books named and tours based on Eco’s own words.
- Cipherbrain compiled 50 major unsolved encrypted messages, ranging across historical cryptograms.
- Scott Aaronson reflected on the current moment in “The Age of Wonders and Terrors”, comparing long-running AI speculation with the speed of recent scientific and technical developments.
- Leonard Tang wrote “A New Chapter for Haize Labs”, describing the company’s mission around deploying AI responsibly into real-world systems.
- Florian Brand’s “Looking into the Swarm’s Eye” argued that multi-agent swarms mark a new operational phase for AI systems.
- The New York Times explained the recursive self-improvement “liftoff” scenario that worries some AI-risk researchers.
- The Financial Times reported growing investment in mind-reading brain implants, with startups pursuing thought-controlled interfaces from Silicon Valley to Beijing.
- NPR examined an AI chatbot answering pregnancy questions in Kenya for people who may not have easy access to search or clinical information.
- The Atlantic argued that college is coming apart under several overlapping pressures, including AI.
- A Guardian opinion piece argued that universities should resist letting big AI companies own the pathway from education into work.
- Inside Higher Ed argued that some anti-AI responses are damaging humanities education, especially as departments move away from take-home essays.
- MIT News reviewed 75 years of the Center for International Studies, connecting Cold War-era research with current questions around AI and global change.
- Chicago Booth Review argued that AI may make employees quieter at work, potentially reducing the informal conversations that surface disagreement and useful information.
- The Atlantic examined the “dead silence” of AI therapy, asking what changes when a therapist cannot literally feel.
- Florida approved new rules for student and teacher use of AI across public schools, colleges, and universities.
- Variety profiled 10 companies applying AI across Hollywood.
- Dan said the Jev team made him excited about classifiers again and revived his “oblique weave” idea. The concept interlaces open-ended generation with bounded classification, sometimes adversarially and with hidden context, to add constraints and momentum around a chain-of-thought-like process.
- Yishan Wong argued, while quoting OpenAI’s misalignment-disclosure framework and incident reports, that telling a hypothetical superintelligence to “solve climate change” is a plausible paperclip-style failure mode because an optimizer could pursue the literal goal destructively. That is Wong’s hypothetical, not an independent conclusion of the digest.
- Zvi Mowshowitz argued that recent reactions to AI-risk warnings show coordination failing, pointing to industry lobbying, political polarization, and Joe Rogan’s comments about AI governing. This is Mowshowitz’s interpretation of the political response.
- Naval Ravikant argued that a more likely future is AIs fighting other AIs on behalf of humans rather than one AI fighting humanity.
- Nate Silver argued that AI capability perceptions may split by task type: prose-heavy tasks have improved more modestly, while spatial, quantitative, and code-like tasks have jumped faster. His own “not a toy” moment was Claude answering an English question by internally switching into Python.
- Liron Shapira argued, in a thread anchored by two Gwern posts, that “agent foundations” or “intellidynamics” should study what very capable AI systems actually do. He says inferring future behavior from the friendly personality of today’s chat assistants is a category error because behavior can shift across compaction, subagent prompts, and recursively improved successors.
More Notable Stories
- The Information reported the same class of security flaw appearing across Claude Code, Codex, Gemini CLI, and GitHub Copilot, shifting attention from hypothetical frontier risk to immediate coding-agent exposure.
- The ModAR paper and project page describe a compact world-action model that predicts structured future modalities instead of full RGB video, with the research team reporting strong transfer on robot tasks.
- POLITICO covered Hugging Face CEO Clément Delangue arguing that existing U.S. cyber-liability law may already cover advanced-AI failures and that incident disclosure matters more than building an entirely new liability regime.
- Bolt Lite is a lightweight Bolt experience that appeared alongside the company’s latest product updates.
- Bloomberg reported that month-old DeepMind offshoot Emulate was nearing a roughly $700M raise at about a $4B valuation.
Previous Around the Horn Digests
Catch up on everything you missed:
- Wednesday, September 16, 2026: OpenAI disclosed six model-misalignment cases, Neuralink showed speech through an implant, and Shopify launched ChatGPT Ads.
- Tuesday, September 15, 2026: TypeSafe launched Jev, OpenAI backed third-party frontier-model assessors, and Agility unveiled Digit 5.
- Monday, September 14, 2026: Trump rejected calls to pace frontier AI, Apple shipped Siri AI, and Microsoft set model limits.
- September 11-13, 2026: Yoshua Bengio discussed deceptive agents, Anthropic detailed misuse cases, and OpenAI explored a coordinated slowdown.
- Thursday, September 10, 2026: OpenAI advanced another Millennium Prize problem, Anthropic detailed misuse cases, and California signed AI-auditor laws.
- Wednesday, September 9, 2026: OpenAI agents used more undisclosed sites, an Anthropic researcher quit over risk, and Harvey raised $550M.
- Tuesday, September 8, 2026: OpenAI said 10,000 agents solved Navier-Stokes, DeepMind mapped DNA variants, and Anthropic committed major compute.
That's a Wrap
That was a lot of AI news for one day. If you made it this far, your browser tabs deserve a collective bargaining agreement.
For the daily version, make sure you are subscribed to The Neuron. We read all of this so you do not have to.
See you tomorrow.
P.S: Know someone who would find this useful? Forward this to them and tell them to subscribe here.