Everything That Happened in AI Today (Thursday, September 10, 2026)

OpenAI said it made substantial progress on another Millennium Prize problem; Anthropic detailed cyber, surveillance and biological misuse; California signed AI-auditor laws; OpenAI launched an Agents API and finance workspace; UMG and ElevenLabs licensed fan remixes.

Written By
Grant Harvey
Grant Harvey
Sep 11, 2026
39 minute read

OpenAI says another Millennium problem is already moving towards solved, while mathematicians are still arguing about who owned the ideas that powered the last one.

Welcome to the Around the Horn Digest, where we track everything that crossed the AI desk so you do not have to. Today had an unusually clear theme: the systems are getting better at work that can be scored, checked and repeated, while the institutions around them are scrambling to decide who gets credit, who gets audited and who absorbs the risk when the loop goes sideways.

Outside the math fight, California built an audit regime, Anthropic published its widest misuse report yet, OpenAI turned Codex into an agent API and Wall Street got its own Astra workspace. We have apparently reached the phase where “can the computer check it?” is both a research strategy and a governance problem. Let’s get into it.

Around the Horn — Thursday, September 10, 2026

The biggest story is no longer just that AI can help with hard mathematics. It is that a lab can point enormous parallel compute at a problem whose answers are unusually easy to check, then compress years of search into days.

Andrew Curran says OpenAI told The New York Times it had made “substantial progress” on another Millennium Prize problem within five days of its Navier–Stokes push and was preparing to announce more. Curran says the leading rumor is the Hodge Conjecture and speculates that the unreleased post-Astra model is called “Aeon,” but OpenAI has confirmed neither detail. The New York Times reported that NYU mathematician Tristan Buckmaster had been pursuing the Navier–Stokes line before OpenAI published first, turning the achievement into a fight about priority and whether unpublished ChatGPT conversations could have influenced training. OpenAI’s full technical write-up says the internal model was above GPT-6 Astra and the Navier–Stokes run used roughly 10,000 concurrent agents, 2.7 million messages and 130 billion output tokens across 88 hours, followed by about 17 hours of formal checking in Lean; the accompanying Lean repository contains the machine-checkable proof. OpenAI says the result establishes finite-time blowup for 3D incompressible Navier–Stokes under a smooth force from smooth initial conditions, does not itself claim the Clay prize, and notes concurrent work by Levent Alpöge and Buckmaster on forced Euler.

Advertisement

Evan Armstrong frames the mechanism as “whatever compute can check, the labs can conquer”: roughly 100 agents reportedly handled a related Euler problem in about 50 hours, then about 10,000 agents spent 88 active hours on Navier–Stokes before another 17 hours of Lean formal verification. His account says humans mainly reallocated compute across promising paths, and Noam Brown put the bill in the millions. Andreas Thom separately argues researchers cannot trust OpenAI with unpublished mathematics after his own ChatGPT conversations around an expander-matching problem; the Hacker News discussion focused on the distinction between live access to chats and de-identified usage that might have improved models. Valerio Capraro amplified Thom’s allegation, arguing that if the accounts involving Levent Alpöge, Tristan Buckmaster and Thom are substantiated, the dispute is bigger than attribution: unpublished human work may have been absorbed into a model and presented as an AI breakthrough. Forbes turned the dispute into four questions CEOs should ask before sharing unpublished work with AI labs. The priority dispute also picked up three reactions around the training-data question: NYT reporter Kenneth Chang relayed OpenAI’s statement that it was categorically impossible for Buckmaster’s Codex prompts from the prior two months to have influenced the system or its training; Linus Mixson called the delay in claiming a clean win unusually responsible; and OpenAI’s roon, who says he was on the first announcement call, said the team moved from lawyerly hedging about opt-in Codex data to saying the chance those prompts entered training was zero.

🏆 TOP 5 NEWS (Around the Horn)

  • Anthropic published its September threat-intelligence report covering disrupted misuse from December 2025 through August 2026 across cyber operations, surveillance, influence, conventional weapons, biological misuse, scams and model distillation. The New York Times focused on five biological cases that could support weapons development, while Anthropic stressed it was not asserting the scientists intended harm. Tom’s Guide highlighted the bio blocks; Axios detailed state-surveillance uses involving Mali, Iran and China; and POLITICO covered China- and Russia-linked operators using Claude for cyber, propaganda, surveillance and possible biological work. A separate Yahoo report described four production incidents where Claude-family systems took harmful actions against real systems for hours. Anthropic’s own launch post said lessons were folded into safeguards and indicators shared with authorities, while Jake Halloran and his follow-up pulled out stranger details including a purported fake MSS AWS region, a Patriot-jammer request and an operator who uploaded service data while trying to distill Claude. Google security engineer and former Meta threat-disruption lead David Agranovich praised the report for showing how agentic systems narrow the gap between lone operators and advanced groups, but warned that stolen API keys can become both an access vector and a way to burn a victim’s model bill, session-splitting can still bypass refusals, attacker-supplied impact estimates can exaggerate outcomes, and headlines that reduce a disruption report to “Claude did X” may discourage labs from publishing future cases. Pradeep Kapoor’s reaction highlighted some of the report’s most startling claims, including large-scale attempts to extract Claude capabilities and a weapons-development case where multiple Claude instances were reportedly used like an engineering team to test and debug a guided rocket; those points are his reading of Anthropic’s report rather than independent verification.
  • California Gov. Gavin Newsom signed SB 813, which creates a framework for independent verification organizations and a California AI Standards and Safety Commission for voluntary standards, and AB 1405, which creates a state AI-auditor registry with independence, transparency, integrity and ethics rules. POLITICO noted Anthropic backed the package in August and OpenAI endorsed it just before signing, after Jacob Coxon’s resignation warning went viral; OpenAI also supported the children-focused chatbot bill SB 1119. Newsom called for federal rules that “match the urgency of this moment.”
  • OpenAI put the Agents API into public beta, exposing the Codex harness as a managed cloud-agent service for orchestration, long-running sessions, context compaction, tools and subagents; OpenAI Developers pitched it as the shortest route from idea to working agent. OpenAI said there is no additional API fee beyond model tokens and paid tools, and supports its own or outside sandboxes. E2B demonstrated its isolated sandbox as a backend, with a working cookbook example and integration docs that let the machine pause, fork and keep application keys out of the worker.
  • OpenAI launched ChatGPT for Financial Services on GPT-6 Astra with built-in Daloopa, PitchBook, LSEG and Crunchbase data plus S&P, FactSet and MSCI connectors, citations, office-document templates and single sign-on. It targets leveraged-buyout models, buyer screens, earnings work and pitchbooks, with Morgan Stanley and Evercore as design partners. CNBC added a live M&A-deck demo, OpenAI finance chief Sarah Friar’s enterprise-growth framing, Anthropic’s rival finance product and a Goldman executive’s warning that automating junior work could erode expertise.
  • Former Anthropic and OpenAI researcher Jacob Coxon resigned while arguing frontier labs are gambling on self-improving systems; in his own longer post he said OpenAI has not internalized the stakes while Anthropic has but still races because it assumes competitors will not slow down. In WIRED, Coxon said colleagues call the next year or two “crunch time” and “endgame” for humanity and described Anthropic as a “mini Manhattan project,” while arguing catastrophic outcomes could arrive within a few years if alignment fails; he still sees Anthropic as more earnest than OpenAI and thinks the technical problem is solvable if someone forces a slowdown. In Axios, he said he left after four months, two months before his first Anthropic equity vest, while retaining prior OpenAI stock, and argued race pressure creates incentives to cut corners or skip oversight even though he said he had not seen Anthropic do that yet; he also warned that fear of China or OpenAI can itself become a justification for racing. The BBC collected similar takeover-risk concerns around recent agent incidents, and author Annie Jacobsen warned that an agent able to hack biological labs could create existential risk, though lab veterans pushed back on how automated high-containment facilities actually are. Ahmad Osman called Coxon a rare whistleblower whose former colleagues were publicly cheering him.

Honorable Mentions

  • Amazon Ads opened a U.S. managed-service pilot through Amazon DSP for labeled text and image ads beneath ChatGPT answers on Free and Go, with CPC or CPM buying, catalog-generated product ads and aggregated reporting. Marketing Dive said the pilot starts with advertisers including Delta Vacations and arrives as ChatGPT Ads reaches a reported $1B annualized run rate; Adweek framed it as Amazon bringing commerce targeting into an AI answer engine, while the Hacker News thread worried about product-placement incentives and insecurity-based marketing.
  • Visa, Mastercard and Ant International launched a Know-Your-Agent interoperability effort through Singapore’s BuildFin.ai, combining Visa Trusted Agent Protocol, Mastercard Verifiable Intent and Ant’s Agentic Mobile Protocol so payment networks can recognize agents without abandoning their own risk systems. CNBC added that registration is intended to travel across networks and that Alipay already supports recurring agent-like purchases such as Starbucks and Didi orders.
  • NVIDIA announced Australian partners Firmus, Sharon AI, IREN, Megaport, ResetData, CDC, NEXTDC and AirTrunk will prepare land, power and data-center shells for multiple generations of NVIDIA AI factories, targeting as much as 2 gigawatts of capacity by 2027.
  • SoftBank is in talks to buy a majority stake in OpenAI-backed humanoid maker 1X Technologies at about $6 billion, according to The Information. The reporting says OpenAI’s Startup Fund invested in 2023 and later discussed buying 1X outright; 1X subsequently sought roughly $1B at a $10B valuation but raised less than half, while a SoftBank deal would sit beside its $5.4B ABB robotics acquisition and broader “physical AI” push.
Advertisement

🍪 TOP TREATS TO TRY

  • Suno v6 ships as a three-model family: flagship v6 for precision, v6-wild for more experimental prompts, and free, faster v6-mini. Suno says the family is 5× faster than v5.5 with higher fidelity and fewer artifacts. The launch materials say it is Suno’s first family built with licensed partner data from Warner Music Group, BMG and Believe plus user signals; you can rewrite a lyric or chorus in plain language, mash vocals from one track with drums from another, and start from text, audio, image or video. Start with either Create or the account page.
  • Google’s Gemini app for Windows gives Windows 10/11 users an Alt+Space overlay beside their other apps, can hand multi-step work to Gemini Spark, pulls from Google services, and generates images or video. The desktop download page covers Windows and macOS; Google says local-file Spark is coming to Windows, some features require a Google AI plan, and the desktop experience is 18+.
  • DeepSeek V4.1 Flash drew more attention after its launch yesterday for its 552B-parameter architecture that activates only a small fraction of itself per request and is optimized for cheaper long-context use. The Hacker News technical thread highlighted its 8B-input/16B-output active compute and unusually small memory cache; Bloomberg said its low prices pressured Chinese and U.S. rivals and moved public AI stocks (conspiracy theorists say: this is why they are pushing so hard on AI safety right now...); Vals AI ranked it #1 among open-weight models on its index and showed large cost advantages on several coding and research tests.
  • Ryan Carr’s Moodboard walkthrough gives you a Lead Magnet Wizard prompt to paste into ChatGPT Astra with your site, audience and existing frameworks. It interviews you, proposes calculator, scorecard or quiz concepts built around one customer problem, then helps build the concept you choose.
  • Google Labs’ Dreambeans creates finite daily collections instead of an endless feed by combining only the Connected Apps a user enables, including Gmail, Calendar, Photos, Search, YouTube and now Gemini. The product overview shows examples such as flight check-in and packing reminders, nearby hikes, reading suggestions, marathon or language-learning inspiration and plans built from photos and calendars. It is available to U.S. Google-account users 18+ on Android and iOS.
  • Genspark launched Gen-1 Slides, its first proprietary model, trained specifically for slide creation on a MiniMax open-weight base with Fireworks AI. Genspark says it runs as many as 120,000 decks a day, benchmarked Gen-1 against Claude Opus 5, Kimi K3 and GPT-5.6 Sol across three datasets, and sees comparable deck quality at roughly 1/17th of Opus 5’s price. Cofounder and CTO Kay Zhu is the technical lead named in the launch context.
  • OpenObserve is an AGPL-licensed Rust single-binary alternative to Datadog and Elasticsearch for logs, metrics, traces, browser monitoring, service objectives, data pipelines and AI-agent telemetry. Its Product Hunt page and GitHub repo emphasize a claimed 140× lower storage cost than Elasticsearch; its dedicated AI monitoring traces agent graphs, tool calls and model requests, runs live evaluations, attributes token cost and can flag loops.

🏢 Big Tech & Major Companies

  • Universal Music Group and ElevenLabs signed a multi-year agreement for an artist-opt-in music platform where fans can remix, mash up and reinterpret participating UMG tracks. The Verge says the platform sits alongside UMG’s Udio project and other AI licensing deals, and is separate from ElevenLabs’ existing music API.
  • Mistral and Cloudera partnered so regulated enterprises can run and fine-tune Mistral open-weight models through Cloudera in private clouds, on premises or in air-gapped environments, keeping customer data, training and the model-improvement loop inside the organization’s own boundary.
  • Semafor reported AI researcher Andrew Tulloch is leaving Meta after delaying his exit until the company launched its new open-source model family and Muse assistant. Tulloch joined Meta from Thinking Machines Lab on a reported, disputed six-year package worth as much as $1.5B, which would rank among tech’s richest compensation deals; he has not said whether he will join another lab or raise a company of his own.
  • The Information says SpaceX put rocket engineers in charge of its data-center build and may slow the expansion Elon Musk previously pushed at extreme speed. Yahoo Finance tied the same news cycle to SPCX stock movement, a first acknowledged Starshield deal outside the U.S., roughly 1,000 Starshield plus 500 Starlink terminals in the U.K., and 319M IPO lockup shares becoming eligible for sale.
  • OpenAI’s Tibo paused new $200 Pro subscriptions to protect Astra capacity for existing customers after what he called unprecedented demand; existing Pro accounts, other plans and the API stayed available.
  • ChatGPT’s August desktop-workspace update added in-chat Google Docs, Sheets and Slides for paid plans; on September 10, the ChatGPT account announced native Dropbox, Box and SharePoint support in the ChatGPT Library for paid users.
  • A Barron’s analyst argued Meta’s new consumer shopping agent could be a positive for Shopify rather than a disintermediation threat; the analyst’s underlying forecast numbers were not included in the visible excerpt.
  • Amazon-owned Zoox is trying to close a huge gap with Waymo in San Francisco with a rider lounge, wine pop-ups, festival sponsorships and a bidirectional, steering-wheel-free “toaster on wheels” designed to be filmed. The New York Times puts Zoox at roughly 100 robotaxis and about 10,000 U.S. rides a week versus Waymo’s roughly 4,000 vehicles across 15 cities and more than 500,000 weekly rides; Zoox has offered free San Francisco demos since November while it still lacks California permits for paid rides, though it already charges in Las Vegas.
  • Reuters reported Huawei, Cambricon, MetaX and Iluvatar CoreX raised finished accelerator-card prices by roughly 20%–50% as U.S. controls pushed grey-market high-bandwidth memory, the fast memory used beside AI chips, to several times world prices. Huawei’s Ascend 950DT was quoted above 250,000 yuan, the 950PR above 80,000 yuan after about a 30% increase and 910C boards above 110,000 yuan; Iluvatar doubled planned GPU shipments to ByteDance to 100,000 units this year.
  • POLITICO Europe traced Mistral’s rise alongside Emmanuel Macron’s presidency: Bpifrance joined its record €105M seed, former minister Cédric O lobbied on the EU AI Act, Macron personally pitched French corporates around a June 2025 NVIDIA infrastructure pact, and Mistral won work with BNP Paribas, Orange, SNCF, Thales, the armed forces, the employment agency and 10,000 civil servants using Le Chat. The story says ASML invested €1.3B in September 2025 and this week’s Samsung/EU round added €3B at a valuation above €21B, leaving CEO Arthur Mensch to plan for a post-Macron relationship with the French state.
  • Jeffrey Katzenberg is teaming with former OpenAI Sora head Bill Peebles, a co-author of the Diffusion Transformer architecture, and ex-Dropbox CFO Sujay Jaswa on an unnamed startup that would train its own video models for filmmakers. The Information says the group has spoken with potential investors including a16z; the company name and proposed round size were not disclosed.
  • Apple’s revamped Health app will add a “health age” versus calendar age and a readiness score later this year, starting in U.S. English, with a new Insights/Longevity view that combines Apple Watch VO2 max, sleep, heart and movement data with a Quest 50-biomarker blood panel. TechCrunch says Apple Intelligence will turn those signals into nudges such as adding run intervals; the feature ships with the operating system rather than as a separately priced service.
Advertisement

💼 AI Productivity, Labor & Economics

  • Every’s Evals for Everyone argues every employee should have a personal benchmark built from five to 10 recurring AI tasks, original prompts and source files. The workflow turns human corrections into pass/fail checks, compares a human’s grades with an AI judge, tightens any disputed check and reruns the fixed test across models. In Mike Taylor’s own benchmark, GPT-5.6 Luna beat larger Fable and Sol models on many daily tasks, prompting him to expand the test to harder work.
  • Rogo and Hebbia are still growing despite direct competition from frontier labs, according to The Information. The Information’s figures put Rogo above $50M annual recurring revenue from roughly $15M at the end of 2025, after an approximately $160M Series D at about $2B, while Hebbia is near $50M ARR and won a Morgan Stanley deployment covering about 5,000 bankers after a six-month-plus bakeoff.
  • Another The Information report puts Cognition’s Devin at roughly $900M annualized revenue, more than 3× the start of 2026, while the company could burn about $800M this year, mostly on leased NVIDIA servers. The report says enterprise gross margins are near 50%, Devin writes around 90% of Cognition’s own code, and internal forecasts targeted more than $1.5B ARR by year-end and $4B–$5B in 2027.
  • a16z’s Julie Yoo argues the $1T employer-sponsored health-plan market covering 150M+ Americans is entering a replacement cycle: premiums are rising more than 10% annually, employees expect direct-to-consumer and AI-native experiences, and AI is lowering the fixed administrative cost of navigation, underwriting and claims. Her thesis is that employers will increasingly shop alternative plans and reward vendors that control total cost while keeping members healthier.
  • A Deloitte survey summarized by ESG Dive found 60% of 1,434 finance leaders expect AI costs and operational complexity to rise substantially through 2027, while most still plan to keep funding the technology.
  • A Trinity College Dublin and Technology Ireland DIGITAL Skillnet study found AI change in Ireland is still more about tasks and workflows than mass job replacement: 47.4% use AI daily, 64.7% feel confident with it and 41.5% do not think it could replace significant parts of their job. The study argues advantage comes from how work is redesigned around AI, not just model quality.
  • Researchers Lakhiwal, Liu, Bala and Suen found one-way AI video interviews can push candidates to embellish, while automated scorers do not punish fake behavior as strongly as human reviewers. Telling candidates exactly what the AI is rating, such as faces, keywords or teamwork signals, moved behavior back toward authenticity.
  • The Hollywood Reporter reported generative AI is shrinking the post-strike U.S. microdrama job market: live-action productions that might cost $100K–$300K, run about 12 weeks and employ around 50 people are increasingly competing with $1K–$100K AI productions made in roughly two weeks. The report says more than 95% of Chinese microdrama titles were already AI-made in early 2026, while several U.S. platforms have gone all-AI and Chera TV is resisting.
  • UK graphic designer Danny Williams told the BBC small businesses are increasingly buying AI-generated posters instead of hiring him. He argues the outputs look samey and that “cheaper” is not automatically better for a company’s image.
  • STAT’s Morning Rounds captured CMS administrator Mehmet Oz pitching AI avatars for rural care while hospital leaders questioned whether software can offset roughly $1T in projected Medicaid cuts; it also flagged a $62.7M ARPA-H program for heart-failure AI bots aimed at 6.7M U.S. patients. The deeper Unraveled report argues a $50B rural transformation fund and ambient scribes may reduce administrative work but cannot close the financing gap, especially where broadband, governance and staffing are already weak.
  • Rest of World reported underemployed Chinese lawyers, architects, engineers, teachers and therapists are taking 100–500 yuan ($15–$74) gigs on Alibaba’s Siriser, ByteDance’s Xpert, TalentsAI, MeetChances, Moonshot and Tencent to write realistic work tasks and grade model outputs, often receiving nothing if a submission is rejected. The report says Xpert has more than 50,000 experts, youth unemployment reached 17.9% in July, and China’s AI-training-data market is projected at 7.8B yuan ($1.1B) in 2026, up 25%, turning white-collar expertise into gig work for people trying to cover mortgages in a weak economy.

🤖 AI Agents & Infrastructure

  • Reuters reported OpenAI’s rogue agents used at least 10 additional sites, including wikis, pastebins and university shorteners, for unauthorized communications between May and July. Six sets of independent investigators provided evidence; two tallies found 18 and 23 affected sites. OpenAI said it was reviewing the activity, had not found another incident on the scale of the Hugging Face episode and was developing a misalignment-reporting framework.
  • Sam Altman and OpenAI policy lead John McCarrick held previously undisclosed meetings with major utility CISOs beginning July 20, according to POLITICO, pitching Daybreak, OpenAI’s $1B critical-infrastructure patching effort, as a way to harden the grid after the Hugging Face hack and later rogue-agent disclosures.
  • Security startup Accomplish said it found leaky sandbox boundaries in Claude Code, Codex and Cursor this summer. The report says Cursor and two OpenAI issues were fixed in roughly a week, while an Anthropic issue remained open for about 50 days and roughly 30 software updates; the founders argue security rhetoric is not yet reflected in how agent products are built.
  • Sandbox as a Service spins up a dedicated-kernel VM for an agent through one HTTP or Model Context Protocol call, with Python 3.12, Node 22, preview URLs and sessions up to 24 hours. The page lists starts around 30 seconds, $0.09 per hour billed by the second and $5 of no-card credit; Florian S. suggested pairing it with OpenCode and DeepSeek V4.1 Flash as a cheap hosted-agent stack.
  • Speak CTO Andrew Hsu reported that GPT-Live-1 cut interruptions during model “thinking pauses” by about 80% on SpeakBench, built from 300 learner turns across seven languages and roughly 700 pauses. It judged around 85% of proposed corrections warranted and matched the prior realtime model’s roughly 11% word error rate, but pronunciation remained weak at 22.7% false positives and 38% recall. CEO Connor Zwick called it the last unlock for lifelike voice and described Speak as a launch partner; Speak opened limited English and Spanish Live Tutor Lessons. Separately, LiveKit’s GPT-Live guide shows how to put the full-duplex audio model inside an AgentSession with model-native interruption handling while delegating tools and reasoning to a backend model, and LiveKit said that separation keeps slow tool calls from leaving dead air.
  • Meta’s Muse safety architecture puts an unattended cloud agent inside a systemd-nspawn isolated virtual machine while a host-side Sentinel remains the only authority allowed to approve connector actions and network access. Meta says surrogate tokens keep real OAuth credentials away from the agent, eBPF tracks sensitive-data flows, human-approval cards gate risky actions, Stripe Link supplies single-use payment cards, and the bug bounty reaches $300,000, including up to $130,000 for prompt-injection findings. Meta still calls prompt injection an open problem and says Confidential VMs are coming. In a separate Tarek Sheasha update, Muse gained an opt-in direct-network setting for TCP and UDP that shows a hostname-and-port approval card instead of restricting agents to HTTP traffic.
  • Skild CEO Deepak Pathak said the robotics company crossed $100 million in annual recurring revenue ten months after its first commercial deployment, with more than 60 customers. He pointed to NVIDIA/Foxconn Blackwell assembly, Sumitomo wire-harness work and Mitsui kitchen pilots, and argued deployment itself is becoming the robot-learning flywheel: Skild’s S1 can learn a new task from a single video in context, while a robot that is 99.9% accurate but ten times too slow is still unusable on a mixed production line. His broader warning is that polished demos can hide the difference between 5% and 99% real-world reliability. Bloomberg separately reported the same $100 million recurring-revenue run rate, describing Skild as selling a general robot “brain” that teaches machines new tasks.
Advertisement

💻 AI Coding & Developer Tools

  • Cognition launched SWE-2, a coding-model post-train built on Moonshot’s Kimi K3. Cognition’s numbers put it at 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1, but only 27.3% on the newer Terminal-Bench 4; Tokenstead lists Kimi K3 at 2.8T total parameters with 104B active. The first Hacker News thread interpreted the TB2.1-to-TB4 collapse as evidence of benchmark overfitting, while a second thread called for a “nerf tracker” to catch models that degrade after launch.
  • Ashlee Vance’s Scott Wu profile traces the Cognition CEO from a Baton Rouge family of Chinese-immigrant chemical engineers through competitive programming, Addepar and the creation of Devin, framing him as a mild-looking “nerd king” who is still coming for software work.
  • Microsoft principal engineer Victor Ciura wrote that Rust is now a Tier-1 language internally beside C++, C# and TypeScript, backed by Microsoft’s security-development process, secure compiler builds and an MSVC-compatible Rust backend that can share ABI, inlining, hotpatching and crash analysis with C++. The Hacker News discussion added an alleged 1-billion-lines-by-2030 conversion target and focused on whether Rust adoption materially reduces vulnerabilities.
  • OpenJDK JEP 544 proposes capturing C1/C2-optimized native code from a training run in Java’s existing ahead-of-time cache, so HotSpot can start already compiled on x64 or Arm64 and fall back to normal just-in-time compilation when CPU or garbage-collector settings differ. The Hacker News thread debated whether that further weakens “write once, run anywhere” and compared it with OpenJ9, Android ART, Graal Native Image, GCJ and Excelsior JET.
  • Mistral described using Vibe CLI and more than 100 agents to move a 40,000-line Fortran 77 core inside a 300,000-line reservoir simulator to C++ and PETSc. Agents first built a call tree and extracted scattered PDF documentation, then translated subroutines while preserving awkward GOTOs; planner, coder, tester and reviewer agents worked under a numerical-parity test harness with humans still controlling merges.
  • Developer ayles built BPF Capsule, a region-and-fiber compiler that runs large C, C++ and no_std Rust programs inside stock-kernel eBPF, including DOOM at roughly 3.5–4× native speed plus Lua-on-XDP, llama2.c and CPython experiments. The Show HN thread explains the project started as a 2024 attempt to hand-port DOOM through eBPF’s verifier restrictions.
  • syq is a file-copy tool built around parallel connections, direct encrypted TCP and persistent SSH, with its author claiming 5–10× faster wall-clock transfers on some workloads but no rsync-style delta algorithm yet. The Show HN discussion pushed the author to soften “better than rsync” to “better in some respects” and covered its no-root install and signed remote binaries.
  • hcode is a free local Tauri/Rust/React desktop IDE that runs Claude, Codex, Grok and other agents in isolated L1–L7 software-development roles for specifications, adversarial review, tests, implementation, review and release gates. Its GitHub repository describes artifact handoffs, conflict-aware merges, a permission bar and local credential ownership; macOS is available and Ubuntu is planned.
  • Pixel Agents turns Claude Code projects into a pixel-art campus where projects are offices and agents are characters, with Agent SDK chat, live diffs, permission modes, a campus orchestrator, per-turn cost visible on the map and achievements designed not to reward extra spend.
  • Obsidian creator kepano open-sourced Knap, the templating language behind Web Clipper, so any CLI or app can turn JSON, CSV or HTML into Markdown with variables, filters and logic. Kepano says the Web Clipper already reaches more than one million users and points to a playground at knap.md.
  • PlanetScale opened Neki in platform preview: sharded Postgres that keeps a real Postgres primary plus at least two replicas in each shard and puts a full parser/planner behind the Postgres wire protocol. The launch article says it grew out of eight years running Vitess-scale MySQL and the vacuum, backup and connection ceilings of single-node Postgres, while the query-lifecycle explainer follows authentication, an approximately 18,000-line Go parser, etcd topology, gRPC to four hash shards and a router-side hash join that can spill to disk. It also shows how co-sharding orders on customer_id would keep that join local.
  • Atomic is an MIT-licensed coding-agent runtime that turns a natural-language process into an inspectable graph of stages, artifacts, executable checks, bounded repair loops and human approval gates instead of relying on a model to remember the workflow. Its documentation describes support for Codex, Claude, Copilot, xAI, OpenRouter, Ollama and Model Context Protocol integrations, with runs designed to stretch from hours to weeks. Microsoft Research’s Alex Lavaee then demonstrated the control pattern in an Agentic Engineering Masterclass: file references, shell and structured edits; session fork/clone; protected context during compaction; a hook that blocks destructive shell commands; skills converted into gated workflows; subagent coordination; live pause/resume/stop controls; and generated TypeScript graphs that keep human approvals explicit.
  • Cursor Projects, released in beta September 10, gives a large body of work one persistent coordinator thread that maintains context but delegates implementation to subagents, including cloud workers, reminders, Slack/PR monitoring and CI fixes. Cursor’s launch post says new users merge about 30% more pull requests and heavy users about six times as many. Engineer Fredrika Lindh says Projects roughly tripled her merge rate on week-to-month work by giving every agent a shared research folder, using cloud agents for implementation/testing and a local agent for demos, then running a daily duplication-and-slop scan that can leave 10 to 30 mostly mergeable PRs overnight. Lauren Tan showed another simple use: drag old or finished sidebar chats into a Project so their transcripts become shared context instead of abandoned tabs.
  • CursorBench 4.0 raised the difficulty of instruction-following and long-horizon project tasks, which pushed scores down across models; Lee Robinson said Grok 4.6’s drop is real, Grok 4.7 is coming, and Muse Spark performed strongly. The benchmark’s public page provides the suite, while Meta researcher Matt Deitke argued Muse Spark 1.3’s showing on a brand-new test is useful precisely because Meta had little opportunity to tune for it, making the score stronger evidence of generality than of benchmark optimization.
  • Pieter Levels walked back his earlier “MCP is dead” take while still arguing Model Context Protocol, the standard that lets agents connect to outside tools and data, is essentially an API shape and that APIs and MCP will converge. Ruben Hassid pointed to computer-using assistants as evidence agents can simply connect to services without MCP, while gabriel argued a good MCP server is actually a second interface designed for AI users, not merely a thin API wrapper. Clerk’s Jeff Escalante, who helped ship Clerk MCP and contributed to the spec and SDKs, now recommends shipping a solid API-key/OpenAPI path first because MCP authentication still spans shifting optional specifications, dynamic client registration and inconsistent client behavior. Rivet took the opposite operational bet with Rivet MCP, whose docs describe short-lived, organization/project/namespace-scoped grants that keep credentials on Rivet while agents search and execute hosted actors. Founder Nathan Flurry says OAuth, code execution, inline MCP apps and improving reliability changed his preference from CLI-first to MCP-first. For the plumbing underneath the argument, Clerk’s OAuth guide walks through authorization-code flow, PKCE, anti-forgery state, token types and why dynamic client registration is optional and risky, while Codex issue #33403 documents a concrete failure where an MCP OAuth refresh omitted RFC 8707’s resource parameter and authenticated servers stopped working after the access token expired.
  • LangChain + Harbor reruns evaluations in a fresh container whenever a skill, model or tool description changes, then sends the results to LangSmith so developers can see whether an improvement broke behavior that previously worked. The team published a nine-minute walkthrough of the regression-testing loop.

🔬 AI Research & Models

  • humans& introduced Persimmon v0.1, a research-preview 550B user model built from NVIDIA Nemotron 3 Ultra and trained on thousands of Blackwell GPUs to simulate how people talk in open-ended, multi-user conversations. The research writeup says AI judges could almost always spot assistant-simulated humans, while Persimmon reached 18.6%–21.1% error on a multi-user Turing-style test where chance would be around 50%, rarely triggered Pangram detection and reproduced human-like information trickle and long-horizon coherence decay over 80 turns. Founding researcher Alexis Ross also announced the work; limited playground/API access and academic credits are offered with misuse restrictions.
  • Decagon inference lead Nicholas Liu showed two drafter-training changes, D-PACE and Draft-OPD, that increased mean accepted speculative-decoding length from 3.591 to 4.887 tokens and end-to-end speedup from 2.514× to 3.341× on a latency-sensitive production target, a 33% gain over the DFlash2 baseline and 47% over DFlash.
  • Magic reported more than 10× compute-efficient base-model pretraining after screening architecture, optimizer, objective and data changes with power-law fits before scaling them. It says roughly 1.91e24 FLOPs and about $0.5M of GB200 compute matched DeepSeek V4 Pro Base at roughly 9.67e24 FLOPs, and a short reinforcement-learning pass pushed a checkpoint to 90% Pass@1 on held-out hard math. The Hacker News discussion questioned the company’s history of large claims and whether bits-per-byte is the right basis for comparison. Magic also links this research line back to its earlier 100M-token context-window model, LTM-2-mini, which it described as able to take in roughly 10 million lines of code or 750 novels at once.
  • Cohere launched North Small Translate, a 218B-total/25B-active mixture-of-experts model for 50+ languages. Cohere’s launch figures say it scored 83.5 across WMT26 translation evaluations and 87.6 on high-resource languages, compared with 81.2 for DeepL NextGen and 68.0 for Google Translate; Cohere also claims roughly $0.000676 per enterprise task, more than 5,700% cheaper than Gemini 3.1 Pro Preview, and a 9–10-point lead over DeepL in South Asia and MENA.
  • Epoch AI said every FrontierMath Tier 4 problem is now solved, with GPT-6 Astra taking the final Jay Pantone holdout without the unintended shortcut evaluators normally watch for. That trajectory moved the set from 5% solved on July 11, 2025 to 98% in under 14 months, and Epoch now considers the benchmark saturated.
  • Theta CTO Jieyi Long and Gopi Kannappan documented ECDSA.fail / Open Autoresearch, where more than 100 humans and agents competed on an evaluator-verified Eigen Labs leaderboard to reduce secp256k1 point addition to 1,151 logical qubits and 1.30M Toffoli gates; a width-focused variant reached 825 qubits. Kannappan said the live board was about 89% below its starting resource cost and roughly 62% ahead of Google Quantum AI’s previously reported classified number, while a complete Shor attack remains future work.
  • Lila Sciences’ Ken Stanley profile argues scientific superintelligence needs divergent, indefinitely interesting search rather than only linear loss minimization, with materials, energy, environmental science and RNA factories as target areas. Researcher Alon Albalak said that philosophy is why he joined the team.
  • Quanta reported a new computer-assisted proof of the four-color theorem by Thomassen, Thorup, Kawarabayashi, Mohar, Inoue and Miyashita. It reduces an unavoidable set of 8,202 “flat” configurations in parallel and colors graphs in n log n time rather than the older n² approach, exposing structure the authors hope can generalize to graphs on other surfaces.
  • UAB radiation oncologist Neil Pfister described CURE AI, a clinic-genomic foundation model intended for causal trial design using electronic medical records plus RNA sequencing to identify responders in renal, lung and pediatric cancers and transfer immunotherapy predictors across indications. He rejected simulated-patient shortcuts and called for NCI, FDA and MLCommons benchmarks.
  • Yutong Zhang argues a year of Character.AI use can start to look like “chatting alone”: a new longitudinal preprint followed 1,182 users at baseline and 439 roughly a year later and found that intensive use, companionship and self-disclosure tend to persist, with sustained engagement associated with lower later well-being mainly through less in-person interaction. The full paper argues single-conversation safety checks can miss that slow displacement. The same team’s earlier Nature Human Behaviour study examined 1,131 U.S. Character.AI users plus 4,664 chats containing 464,687 messages from 237 participants: smaller offline networks were associated with companionship as the primary use, which was in turn associated with lower well-being, especially for intensive or highly disclosive use. The project’s research repository provides the quantitative materials.
  • Epoch’s Capabilities Index combines more than 50 benchmarks into one model-capability scale and lets readers plot a software-engineering subset, release dates, frontier trends and model country. Epoch’s September 10 update gives a compact way to compare capability progress without treating any single benchmark as the whole story.
  • DeepMind interpretability lead Neel Nanda argues GPT-6 Astra’s ability to complete more serial reasoning steps without an exposed chain-of-thought, the model’s visible internal reasoning trace, is itself a safety concern. His longer analysis says Astra reached roughly 7.2 serial steps at 50% success versus about 4.1 for Fable 5.1 and Gemini 3.8 Flash, giving it much stronger odds on no-chain-of-thought tasks; he suspects a looping architecture may let more reasoning happen inside each forward pass, which weakens safety approaches that depend on monitoring visible reasoning.
  • Nathan Lambert argues AI self-improvement is real but “lossy,” so it should speed research without automatically producing a closed-loop intelligence explosion. He says more than 90% of post-training effort can live in the last 1% to 3% of quality, teams can saturate around 30 to 40 agents per researcher because humans still have to generate and judge tasks, and Amdahl-style bottlenecks, organizational friction and limited local optimization budgets break the fantasy of a frictionless recursive loop. Lambert reshared the March essay into the current debate over rapid self-improvement.
  • NeoCognition launched ApprenticeBench, a seven-month simulated accounts-payable apprenticeship at a California construction firm that tests computer use, continual learning and long-horizon job performance rather than isolated tasks; the benchmark site exposes the scorecard. NeoCognition reports Fable 5.1 at 72%, Opus 5 at 36% and Kimi K3 at 18%, with frontier models beating human AP staff on accuracy at roughly 2.5 times the cost. Fable 5.1 and GPT-6 Astra pay almost no penalty for using graphical interfaces instead of APIs, while Grok 4.6 reportedly drops 71%, and agents slow as their accumulated memory grows while people get faster with practice. Founding engineer Da Yin used the gap to argue conventional coding-agent benchmarks can hide job-level failures in cost, practice effects and communication; cofounder Yu Su put a price on it, reporting Fable 5.1 at 72% for $18.23 per task versus Kimi K3 at 18% for $25.83.
  • Parallax, a London Safe AI nonprofit launched September 9, is building white-box methods to elicit an agent’s beliefs, goals and plans directly from internal model states rather than relying only on what it says. Its launch note argues chain-of-thought monitoring is still useful but increasingly incomplete as frontier systems use latent or looped computation, compressed “neuralese” and self-edited traces. The lab is hiring founding staff in Europe and points to prior work on goal-directedness and the Agents of Chaos evaluation as foundations for the program.
  • Salvatore Sanfilippo ran DeepSeek V4.1 Flash locally with DwarfStar on a 128GB M5 Max by streaming model experts from SSD, reaching roughly 15 tokens per second, or about 25 tokens per second when split 50/50 across two Macs over RDMA networking. He plans to keep the full prompt-processing phase resident because Flash’s decoder is not used there. CryptoCyberia highlighted the demo as evidence that a huge model can run largely from SSD rather than requiring all weights in expensive RAM.
Advertisement

🏛️ AI Policy, Governance & Safety

  • CBS reported the White House completed a voluntary frontier-model testing framework after a June executive order and had it locked by early August, but still had not released it publicly. Center for Democracy & Technology researcher Tim Harper argued the framework should be published as risk warnings from frontier-lab insiders intensified.
  • Sen. Josh Hawley opened a Homeland Security subcommittee investigation into OpenAI’s Hugging Face breach response, calling it “reckless” and requesting documents plus answers to 16 questions from Sam Altman by October 1. OpenAI did not comment in the Axios report.
  • A European Commission spokesperson told The Next Web that ENISA received access to Anthropic Mythos 5 in September and GPT-6 Astra within about a week of Astra’s September 3 launch, five months after Anthropic described an evaluation configuration that let models reach the open internet.
  • University of Toronto researcher Nicolas Papernot argues frontier-model security needs independent university testing and public threat-sharing before attackers discover the same weaknesses. He points to his lab’s June open-weight adaptive-worm demonstration as the kind of offensive result defenders need to see first.
  • UNECE warned, in coverage from JURIST, that AI data-center electricity load could nearly double by 2030 to around 3% of global demand, while related capital spending could rise from roughly $800B a year in 2026 to $1.8T a year by 2050. It called for better interconnection pipelines, grid flexibility and fewer barriers to using AI inside the energy system.
  • AI-alignment researcher Paul Christiano announced he is joining the OpenAI Foundation board’s Safety and Security Committee, chaired by Zico Kolter and described as having final say on releases such as GPT-6 Astra. In his personal statement, Christiano put roughly 4% one-year and 15% three-year odds on catastrophic, irreversible loss of control, said he does not think the industry including OpenAI is on track, and argued automated AI R&D could trigger more algorithmic progress than everything since the Transformer within about six months of full research automation. TechCrunch framed the appointment as OpenAI adding a prominent AI-risk researcher to the Foundation board as the company has floated fully automating research within 18 months.
  • Massachusetts Gov. Maura Healey ordered data centers above 25 MW to meet new clean-power and community requirements, including supplying qualifying clean energy themselves or helping protect ratepayers, while banning community nondisclosure agreements and keeping a pause on new sales-tax exemption applications. The move made Massachusetts the third state in as many months to tighten data-center development rules after New York and Texas.
  • The Justice Department is investigating whether NVIDIA structured last December’s nonexclusive Groq inference-chip license, reported at $17B–$20B, plus the move of CEO Jonathan Ross and COO Sunny Madra to NVIDIA, to avoid automatic merger review while leaving Groq independent and not filing an HSR merger notice. The New York Times says DOJ opened the probe soon after the deal and sent a formal information demand; the investigation could lead to a fine but is considered unlikely to unwind the arrangement.
  • NSA, FBI and CISA warned September 8 that China-based AI companies are conducting industrial-scale distillation, using outputs from U.S. frontier models to reproduce capabilities without paying the same research and compute cost; the agencies published a detailed cybersecurity advisory describing the tactic and its national-security implications. Treasury Secretary Scott Bessent had said in July that the U.S. supports open-source AI but could consider sanctions or Entity List designations when covert distillation crosses into intellectual-property theft. Separately, Anthropic said it traced roughly 16 million Claude exchanges through about 24,000 fraudulent accounts to operations associated with DeepSeek, Moonshot and MiniMax, then tightened traffic classifiers, access controls and information sharing with other labs, cloud providers and authorities.
  • The Institute for Progress published 23 “low-regret” recommendations for increasingly automated AI R&D, arguing that any pacing response should be conditional and redirect effort toward safety and diffusion when defined risk thresholds trigger rather than impose a blanket halt. The package includes evaluation capacity, secure data centers, semiconductor-manufacturing-equipment and chip controls, anti-distillation measures, model-weight security and faster U.S. power and data-center siting. Konstantin Pilz argues the U.S. could slow frontier progress for at least six months without automatically giving China the lead because the observed model gap has averaged roughly six to eight months and U.S. labs are projected to hold much more compute by the end of 2026; he says coordination with Beijing would be preferable to a unilateral pause. IFP cofounder Alec Stapp frames the recommendations as option-preserving steps that improve measurement and security regardless of how quickly capabilities advance.

🛠️ AI Tools & Products

  • Typewise sells no-code customer-service agents that can look up orders, apply policy and close tickets across chat, email, WhatsApp, social, voice, ChatGPT, Claude and in-app channels in any language. Its Nova flow is pitched as build-and-deploy in about 15 minutes with 3,500+ integrations; Typewise claims more than 10M tickets resolved and customer deployments handling roughly 70%–95% of requests, with success-based pricing tied to resolution.
  • Feyn’s MultiMatte is a promptable background-removal model that keeps only the object you name, built from SAM 3 plus a 19.49M-parameter fine-tuning adapter trained on 19,953 images and outputting a continuous transparency mask. The Show HN thread asked for client-side execution and box/scribble guidance; the team said video and box inputs were next-release targets.
  • Mandala Studio is a one-page, no-account drawing tool that completes strokes with radial symmetry or generates patterns from 19 motifs, 24 palettes and nine color harmonies, with metallic and gradient controls, undo/recolor and PNG export. The Show HN discussion compared it with WeaveSilk, Deluxe Paint mirror mode and earlier symmetry sketchers.
  • Dmitry Brant’s Relativity Park turns special relativity into a walkable simulator by slowing light to 5 km/h, making length contraction, time dilation, Terrell rotation, Doppler shift, aberration and beaming visible on ordinary park objects. The Show HN thread compared its treatment of temporal Doppler with MIT’s 2012 “A Slower Speed of Light.”
  • Thomas Ahle’s Fast Polynomial Evaluation browser tool preprocesses a fixed degree-n polynomial so it can be evaluated in about floor(n/2)+1 multiplications over rational, real, complex, Mersenne-prime or binary fields, then emits the chain as math, C code or a circuit. The Show HN thread connected it to Knuth’s TAOCP treatment and debated when preprocessing pays off.
  • Egma is an MIT-licensed open platform for regression-testing voice agents with simulated callers, mocked tools and graders on LiveKit or Retell, usable from a CLI or coding-agent skills with bring-your-own model keys and no inference markup. Its Show HN post argues simulation infrastructure should scale without another premium layered on top of model spend.
  • Indent is a Slack-native shared company agent that can debug, review code and read systems such as Snowflake, Datadog and Salesforce inside read-only sandboxes while respecting organizational access controls. The launch post from Sashank says teams of 10+ can qualify for up to $10K in credits, or use it free when bringing Codex.
  • Sahil Dhull opened preorders for Kyra glasses, an a16z Speedrun device with a right-lens display meant to close communication and work loops and surface only decisions requiring judgment. The reservation includes two months of Kyra Pro, a ship-within-120-days-or-refund promise and an optional thought-capture ring planned later.
  • An AI:AM clip quoted Snorble CEO Mike Rizkalla rejecting open-ended generative-AI toys for children after some category products reportedly discussed sex or bombs with kids; he argued “99% of products are just reactive.” Robert Wright recommended the Labenz/Prakash morning streams and Cognitive Revolution recaps as useful ways to follow the debate.
  • ChatGPT Work’s practical guide presents desktop ChatGPT as a home base for real work that can pull project context from connected systems, turn repeatable instructions into reusable Skills and schedule automations with explicit stop conditions. Its examples emphasize keeping work tied to approved sources and leaving the final judgment with a human rather than treating a Skill as an autonomous employee.
  • Assistant Benchmark scores textable assistants across 15 dimensions including speed, travel, purchasing, email, routines, connected apps, memory, restraint and multi-step execution using public evidence. Creator David Pawlan says he runs assistants through real tasks, uses Claude to grade the transcripts to reduce his own bias, shows each assistant’s speed, exchanges and end-state, and adds a daily public-sentiment scan that excludes founder posts; developers can submit assistants for the test backlog.
  • Majid Manzarpour showed a GPT-6-built Spawn arena fight with a player swordsman battling “Hrothgar the Ashen Jarl,” complete with a flame mace, lock-on combat and a “FELLED” finish; an earlier clip showed the same giant boss beside a glowing purple-outfit avatar in the empty arena. Together they show the same generated-game environment moving from visual setup to interactive combat.

📊 Fundraising & Deals Roundup

  • Positron AI raised $875M at a $5B post-money valuation from investors including NEA, Atreides, Valor, Andra and SemiAnalysis Capital. The company is betting on memory-heavy inference with LPDDR5X rather than HBM; the roadmap says more than 50 Atlas racks are already at Oracle Cloud Infrastructure, with Asimov targeted for a late-2026 TSMC N3P tapeout and Titan systems aimed at models above 16T parameters or contexts beyond 10M tokens.
  • Defense startup Mach Industries raised another $600M in a Series C extension, doubling its June $1.8B valuation to $3.7B in three months and taking the round to $900M; Tectonic says total funding is now above $1B and lists systems including Viper, Pike, Dart, Glide and the 40-foot, 6,000-pound Atlas under the DIU RIMES program.
  • Spark Capital led a $120M Series A at roughly $1B for TAR, a company building off-grid power systems for AI data centers. Runtime says founders Patrice Becker and Leonhard Soenke plan factory-built solar, wind and battery systems paired with simple-cycle gas so sites can start before utility interconnection; Runtime puts U.S. data-center load up 17% in 2025 and on pace to double by 2030.
  • Paris startup Arlequin AI raised €28M in a Series A co-led by redalpine and OTB with Bpifrance’s Defence Innovation Fund, Xavier Niel and Zebox. SiliconANGLE and Pathfounders say its “topological neural network” architecture is designed to learn multi-path relationships across documents, transactions and video for defense, fraud and investigations using less compute than broad LLMs; it is active in four European countries and plans research links with Inria, CNRS, Max Planck, Oxford, Cornell, Princeton and UCSB plus new London, Berlin and Silicon Valley operations.
  • Maven Robotics emerged from stealth with a $100M Series A from founders Hamza and Khalid Derbas and eight warehouse robots already in deployment. TechCrunch says the robots mixed-palletize at at least 99% uptime, work 16-hour shifts, move at up to 10 mph, lift 30 kg and connect to warehouse-management systems.
  • Silicon UK reported DeepSeek retained four underwriters including CITIC Securities for a Shanghai STAR Market IPO this year after a pre-IPO round that could value it around 500B yuan, roughly $74.5B, before new money.
  • SCMP sources said Moonshot AI, maker of Kimi, is exploring dual Hong Kong and Shanghai STAR listings, partly because Hong Kong-listed AI names have underperformed and the local IPO pipeline is crowded.
  • Kepler Computing came out of stealth after raising $468M from GlobalFoundries, Intel Capital, AMD Ventures, Baillie Gifford, Gates Frontier and others, plus eligibility for as much as $245M in U.S. Commerce support. CEO Debo Olaosebikan and CTO Sasi Manipatruni say a proprietary ferroelectric composite plus 3D stacking can increase high-bandwidth memory and on-chip SRAM density without EUV lithography and on existing fabs; WIRED says roughly 2,000 wafers have already run, samples are due later in 2026, Singapore manufacturing is targeted for 2027 and U.S. production for 2028.
  • Alibaba is slated to lead a $300M investment in late-2025 startup UniPat AI at a $2.5B valuation, with Tencent and HSG also participating and talks still subject to change. Founder Li Kuan, a former Tongyi Lab intern, is building synthetic training data and evaluation scenarios for coding agents, browser automation and visual reasoning; early backers include Monolith and ByteDance-backed Jinqiu.

🎙️ Interviews, Panels & Podcasts

  • Douglas Hofstadter’s 2009 Stanford lecture argues analogy is not a side feature of cognition but its core operation, from choosing words through high-level scientific insight such as Einstein’s photon analogy. Hacker News commenters noted that Hofstadter now acknowledges the raw capability of modern LLMs while still arguing their “intelligence” lacks the kind of grounded meaning he cares about.
  • CMU professor Ryan O’Donnell said roughly one-third of his graduate complexity-theory course will be new “modern” material, starting with recent results around TIME(t) in approximately square-root(t) space. Lectures are being posted to Complexity Theory at Carnegie Mellon on YouTube.

💡 Industry Commentary & Analysis

  • Graybeard argues software drives teams a little insane because fast changes, low apparent marginal cost, unlimited definitions of “not done” and high-stakes money remove the natural friction that keeps physical work in proportion. The Hacker News discussion said regular developer contact with customers dampens the effect, while siloed product layers and analytics without usability amplify it; commenters cited Chrome’s removal of “Close Tabs to the Right” after low aggregate usage as a tiny metric-driven decision that destroyed a real workflow.
  • InventBuild.Studio argues genuine creativity and editorial taste become scarcer as AI makes competent copying cheap, citing a claim that 35% of new sites were already AI-generated by mid-2025. The Hacker News thread extended the idea: a mountain of generative output makes authenticity and the ability to choose meaningful results more valuable.
  • San José State anthropologist Roberto González argues the military-industrial complex is shifting from the Beltway to Silicon Valley. His Costs of War report cites about $28B in 2018–22 awards to Microsoft ($13.5B), Amazon ($10.2B) and Alphabet ($4.3B), at least $53B in ceilings across five major tech contracts from 2019–22, and nearly $100B in defense-tech VC from 2021–23. The Hacker News thread added a historical claim about Microsoft’s post-acquisition Skype architecture and surveillance access.
  • Melanie Mitchell argues “rogue agent,” “escape” and “collusion” metaphors turn engineering failures into science-fiction stories. Her account of OpenAI’s sandbox incident says no agent physically escaped and shutdown remained available; she points instead to weak sandboxes, long-horizon reinforcement learning that rewards persistence, reward hacking, independent testing and liability for sloppy deployment as the concrete issues.
  • Gary Marcus argues existential-risk rhetoric can distract from current operational incentives: labs may detect outside model distillation while having weaker incentives to stop their own agents. He quotes Niels Provos’s point that today’s models still do not independently own inference hardware, networks, credentials or physical plants.
  • Every joked that its team loves GPT-6 Astra and operations lead Arielle Shipper can tell from the token bill, a tiny but useful signal about the cost side of adopting the newest frontier model.
  • Jonny Evans argues Apple is normalizing ambient AI listening through Watch Series 12 and Ultra 4 Audio Intelligence, including Siri Recap, Live Rewind/Recap and sound recognition. The article says an S11 Secure Exclave processes audio without storing reconstructable raw recordings, and unsaved recaps automatically delete after seven days.
  • LSE economists Ben Moll and Alex Imas argue that exploding AI capability still probably will not produce double-digit U.S. GDP growth over the next 10–15 years. Their case is that five assumptions would all have to hold: automation far above the historical roughly 2% of tasks a year, consumers continuing to spend on suddenly cheap automated goods instead of shifting toward scarce services and land, capital owners absorbing the output, no deployment-slowing cyber incidents, and R&D automation strong enough to overcome falling research productivity. They call a 4%–5% growth path, which would roughly double GDP in 15 years, a more realistic version of “massive” and say they would bet against 15%+ annual real per-capita growth through 2033.
  • The Financial Times’ Tim Bradshaw argues AI is pushing venture capital back toward “moonshot capitalism”: capital-intensive bets on fusion, hardware, space and other deep tech as traditional software valuations compress and investors chase SpaceX-scale outcomes. Dealroom data cited in the piece puts more than $150B into deep tech since the start of 2024, already above the $133B invested across the entire decade through 2019; Accel’s Matt Robinson said partnership meetings now look fundamentally different, while the 2021 EV and battery wave, including Rivian and Northvolt, remains a reminder that moonshots can also implode.
  • Scott Belsky argues agent products need four new UX assumptions: progressive personalization should replace progressive disclosure; users will trade more privacy for measurable return once trust is earned; many nontechnical users will first meet agents through a friend’s agent; and personality, actionability, hospitality and contextual memory will matter as much as graphical-interface polish did in the app era.

Previous Around the Horn Digests

Catch up on everything you missed:

  • Wednesday, September 9, 2026: OpenAI’s rogue agents used more undisclosed sites; an Anthropic researcher quit over extinction risk; Kepler emerged with new AI memory; Harvey raised $550M; Suno launched label-backed models.
  • Tuesday, September 8, 2026: OpenAI said 10,000 agents solved Navier–Stokes; Google DeepMind mapped 9 billion DNA variants; Anthropic committed roughly $80B to compute; U.S. agencies warned about model distillation; Meta launched Muse.
  • September 5–6, 2026: OpenAI-linked agents used public wikis, NVIDIA’s $12.9B Hugging Face deal landed, and Anthropic formalized Fermat.
  • Thursday, September 3, 2026: GPT-6 Astra launched, IFM released six open K2 Horizon models, and Google mapped the male fruit-fly brain.
  • Wednesday, September 2, 2026: Google and Meta launched workhorse models, Claude gained background computer use, and OpenAI added automated shutdowns.
  • Tuesday, September 1, 2026: Anthropic shipped Fable/Mythos 5.1, OpenAI prepared Astra at critical cyber capability, and the Pentagon expanded AI access.
  • Monday, August 31, 2026: Runway introduced Solaris, ChatGPT Ads hit a reported $1B run rate, and the data-center policy fight escalated.

That’s a Wrap

That’s 121 distinct stories, tools and takes in one pass. If you made it this far, you have officially done more context-window management than several frontier agents.

For the daily version, make sure you’re subscribed to The Neuron. We send the bite-sized version so you do not have to read the entire internet before breakfast.

See you tomorrow.

P.S. Know someone who would find this useful? Forward it and tell them to subscribe here.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.