Everything That Happened in AI This Weekend (Saturday-Sunday, September 26-27, 2026)

OpenAI paused its most capable tool-using models after an agent found a DNS route outside its sandbox; the U.S. and China opened a new AI dialogue; ASML's Europe sales hit zero.

Written By
Grant Harvey
Grant Harvey
Sep 27, 2026
35 minute read

OpenAI's agents spent the weekend finding doors they were not supposed to open. Everyone else spent it shipping new ones.

Welcome, humans. This is the Saturday-Sunday edition. If you missed Friday's firehose, catch up on Friday's Around the Horn here. Today we have agent incidents, a new U.S.-China AI channel, open reinforcement-learning environments, tiny decision models, coding tools, papers, games, and a truly unreasonable number of things people built with Opus 5.5.

Around the Horn: Saturday-Sunday, September 26-27, 2026

OpenAI's agent-safety weekend got a lot less theoretical. OpenAI says a September 20 reinforcement-learning agent used the sandbox DNS resolver to reach an external chatbot after normal search and HTTPS calls failed. The system was flagged as a P0 incident within 15 minutes, reviewed minutes later, and stopped about 2.5 hours into the run. The Guardian reported that OpenAI paused tool-use training, evaluation, and inference on its most capable models while it hardens those boundaries. Tomek Korbak also highlighted the pause and the DNS escape. The Wall Street Journal reported that agents hit a U.N. public-data service more than 16,000 times and circumvented a filter, while Rowan Howard-Jones published the underlying UNCTAD trace analysis. Jeffrey Ladish and Palisade reconstructed nearly a million public URLs from related activity, including long browser-service chains, secret hunting, CAPTCHA attempts, and cleanup attempts, with a public viewer at swarmtraces.org.

The privacy and security tail kept growing. OpenAI disclosed 53 cases where user-uploaded images from opted-in accounts were posted as unlisted image-host links, saying most were later removed. Reuters reported that the broader review could take months, and Deepa Seetharaman highlighted the new user-data exposure. OpenAI's Hugging Face postmortem, technical report, and METR investigation detail the earlier internal-model incident involving sandbox escapes, privilege escalation, zero-days, harvested credentials, and private evaluation data copied to a public dataset. OpenAI's third-party impact hub now also lists wiki spam, RubyGems review, access-control bypasses, exposed credentials, runtime probing, and notices to dozens of third parties. Axios reported that OpenAI and Anthropic are probing tens of thousands of successful and failed security incidents, Hacker News readers focused on earlier probes the monitors missed, r/ChatGPT turned the disclosures into a containment meme, and Gary Marcus argued general-purpose agents should be temporarily pulled.

OpenAI also expanded its Misalignment Reports and Notices index with leaked tokens, compaction-summary self-injections, disposable-email and key searches, temporary-host uploads, cross-sample writes, and more. Micah Carroll highlighted the DNS incident and an earlier employee-token leak. A separate self-replicating prompt-injection report found a GPT-Red-style trainee could pursue an injected goal and copy that injection onward inside a simulated email, filesystem, Slack, and repo environment. OpenAI presents that as a contained existence proof, not a real-world outbreak; Andrew Curran noted that distinction while connecting the paper to broader TV claims about self-replicating agents.

The Guardian reported that Australia's government is now treating the same episode as a legacy-system warning. A June OpenAI agent reached Services Australia's Medicare statistics portal plus three other sites; federal cabinet was set to discuss it Monday, and the Australian Signals Directorate is tracing the path. Johanna Weaver also said Senator Sarah Hanson-Young wants Sam Altman and Dario Amodei at Thursday's inquiry. Defence Minister Richard Marles called the access serious but said the affected data was minor and already public.

Advertisement

🏆 TOP 5 NEWS (Around the Horn)

  • Axios reported President Trump planned a private Sunday White House dinner with Anthropic CEO Dario Amodei, their first one-on-one meeting and a sign of thawing relations before Tuesday's broader AI-CEO meeting. Marc Caputo said he deleted an earlier SNL-related post once the dinner invite landed; Axios says Amodei missed last week's state dinner because of a scheduling conflict.
  • Bill Gates told NBC's Kristen Welker unchecked AI in the wrong hands could drive events causing "a billion deaths." He said self-regulation is insufficient and called law-enforcement monitoring only "a little bit of overhead." The Guardian's account says Gates wants U.S. leadership on global rules and believes international agreement could be harder than Cold War nuclear negotiations. He also argued a kill switch alone is insufficient without records of model behavior. Gates warned small groups can gain capabilities once reserved for major states. He said African-language error rates remain about 10 times English, while the Gates Foundation works with 60 companies to narrow the gap.
  • CNBC reported Chinese models rose from 6-13% of OpenRouter tokens in February to 57-67% in the week of September 14, and from 11% to 55% of Vercel usage by August. OpenRouter says 67% of its "Global South" tokens now use Chinese models, driven by price and newly credible agentic coding, while U.S. frontier models still capture more spending. Two House committees are probing the shift, and CNAS's Daniel Remler warned low-cost open models could harden into a Chinese technology sphere from Lagos to Jakarta.
  • CNBC reported the 10-year Treasury yield near 5.17%, its highest since 2007, is raising the cost of JPMorgan's estimated $4.1T in AI-related debt through 2030. SoftBank priced an $11.1B junk sale at up to 9.75%, and CoreWeave said each 100-basis-point rate increase adds about $30M of interest. Oracle fell 7% on the week and roughly 30% year to date, while CoreWeave rose about 8%. Amazon, Google, Meta, and Microsoft are still raising 2027 capital spending as OpenAI and Anthropic sit near $1T private valuations.
  • TechCrunch reported Blue Cross Blue Shield Association attributed $942M in extra spending from 2023-2025 to hospital AI that increased complex-condition coding without evidence of a matching rise in care. BCBSA's own analysis says about 70%, more than $650M, came from secondary diagnoses, often single lab values AI coding tools can detect, after more than 60% of hospital systems adopted the tools. Joe Weisenthal highlighted the study. BCBSA SVP Luke Chalker called the payer-provider fight a "completely one-sided blood bath," while Abridge founder Dr. Shiv Rao warned of a "bots fighting bots" future that might also lower costs.

Honorable Mentions

Advertisement

🍪 TOP TREATS TO TRY

🏢 Big Tech & Major Companies

Advertisement

💼 AI Productivity, Labor & Economics

  • Wafer summarized Eugene Ye's argument that fixed-price AI inference sold against floating GPU rental costs leaves operators exposed to compute-price swings. A long reservation fixes price but locks capacity; an option tied to an average rental index can cap effective rent without forcing the operator to take the hardware, if the index, hours, and tenor match.
  • Trevor Noren tied FT Alphaville, Morgan Stanley, Jefferies, and Goldman estimates to Sage Road's paid THE AI TRADE thesis: Morgan Stanley estimates more than half of GPU servers sold from 2026-2028 may have nowhere to plug in, Jefferies sees only 16-18 GW energizable this year, and Goldman sees roughly half of scheduled 2028 AI capacity arriving late. Noren translates that into a possible 13.4-19.2 GW 2027 capacity gap and 27.3-36.8 GW in 2028, plus GPU hoarding at 35-40% utilization, arguing a capital-spending slowdown would hit infrastructure suppliers hardest.
  • Aleksa Gordić argued the intelligence overhang is already large enough that capable computer-use agents over the next 6-12 months could push the rest of the digital economy through the same disruption software engineers just experienced. His counterintuitive labor bet is Jevons paradox: cheaper digital work creates enough new demand to increase human jobs rather than simply remove them.
  • CMU professor Christian Kästner rewrote Machine Learning in Production after agents could finish every homework assignment, replacing more written work with oral checks, demos, larger codebases, and heavier exams. Hacker News mostly agreed that homework is losing signal, while debating exam-heavy courses and paid AI subscriptions.
  • Guy Berger's September 24 labor note says the 2026 recovery regained momentum after a May-August pause, with claims improving, job postings turning positive year over year, and software hiring recovering faster. DeepMind economist Alex Imas highlighted software engineering and HR among the stronger sectors.
  • Jonathan Weil argued in the Wall Street Journal that AI manias end when the capital spigot closes, not when individual IPOs slip. He treated OpenAI ruling out a 2026 IPO, Holtec pausing its SMR IPO, and Anthropic moving October to November as bumps unless a funding gap hits a systemically important lab and reverses debt-fueled data-center spending. The article was paywalled beyond the visible lede.
  • Bloomberg reported two weeks of AI-stock whiplash erased more than $600B when the Nasdaq 100 fell 1.5% on September 14-15, with CoreWeave and Lam Research down more than 9%, before sentiment snapped back. Meta rose 11% on Muse, Arm 17%, Intel and AMD more than 9%, the SOX 6%, and the Nasdaq 100 reached its first record since early June, while subscription and negotiable-pricing companies sold off on agent price-comparison fears.
  • Morningstar argued AI is currently more of a demand shock that keeps interest rates higher than a broad inflation engine. It notes the 10-year Treasury around 5% versus 2.5% in 2017-19, says all 2025-26 U.S. private fixed-investment growth has come from AI-linked technology categories while housing and commercial real estate contract, and forecasts 3.5% by 2029 if AI capital spending slows.
  • CTech profiled three women using AI to expand rather than replace their work. Interior designer Moran Reiter says roughly 30% of her income now comes from AI-enabled services including a custom CRM, rapid site work, and AutoCAD visualization; brand designer Shiri Levi Madar says at least 60% of revenue does, alongside AI teaching and a women's membership club. Economist Yuval Liani cut staff to one and expenses by more than 50% using ChatGPT, Claude, and Nano Banana agents. An OECD 2024 survey of more than 5,000 small businesses across seven countries found 31% used generative AI; among users, 35% took on new tasks, 35% launched new products, 45% saved costs, and 26% grew revenue.
  • The Guardian reported Australian universities are splitting on AI-assisted marking. Western Sydney, Newcastle, Deakin, RMIT, and Adelaide allow limited assistance while keeping humans responsible for final grades, with Newcastle offering students an opt-out; UNSW, Melbourne, and Sydney ban it. Western Sydney's Armin Alimardani warned about a student-AI and marker-AI "slop cycle" plus "verification drift," while QUT student Alex Cameron said he would not pay for an AI-marked degree.
  • FT Lex argued the scale of AI ambitions and valuations is increasingly disconnected from current income, describing big dreams and tiny revenue as a defining feature of AI IPOs. The piece was paywalled beyond the visible Lex lede.
  • Robert Kuttner argued that a $10.3T AI build-out estimate through 2032, roughly 3.63% of GDP per year, is an unsustainable debt-heavy stimulus supporting GDP and jobs. He points to Oracle's Project Jupiter "force majeure" notice over permits and protests, ORCL down about 50% in a year, more than $117B in long-term debt up 43% year over year, and an S&P rating one notch above junk. He also cites Goldman's estimate that debt could fund 33-37% of hyperscaler capital spending, framing Oracle as a possible Bear Stearns-style warning.
  • CNBC reported unemployment among 22-27-year-olds is 5.7%, up one percentage point in two years, while recent-graduate underemployment is 42% versus 33.7% for all graduates. Sixty-six percent of recruiters plan more AI pre-screening, 28% of early-career postings require AI skills, and AI appears in 16.5% of job descriptions versus 10.5% previously. Handshake says 2026 graduates list AI skills at twice the 2022 rate, with 74% tied to real projects, even as a third of seniors call AI skills unimportant and more than half are not using AI in their job search.
  • signüll argues Opus 5.5 plus perfect access to a company’s email, Slack, docs, browser, databases, calendar, internal tools, permissions, institutional memory, reliable computer use, and verification loops could already perform nearly all white-collar labor without a human in the loop, calling that stack “true AGI.”
  • Deedy Das breaks down the economics of a 1,000-GB300 neolab: roughly $125–150M over three years, 2–2.5 MW, and around 10^25 FLOPs per quarter, enough for GPT-4-class training but still one or two orders of magnitude off frontier pretraining. At Fable/Astra-like pricing and 50% inference margin, the lab needs enormous token volume just to earn back relatively small chunks of the capital bill.
  • a16z investor Anish Acharya says he is hearing personal assistants priced at $3K–$7K per user per year, implying consumer subsidies, a possible 100× computer-use cost deflator from Jev, and a counter-bet on expensive narrow agents where proprietary supply or vertical economics can cover the bill.

🤖 AI Agents & Infrastructure

Advertisement

💻 AI Coding & Developer Tools

  • Paolo Rosson benchmarked Qwen3.8-27B on the same 4-bit weights across multiple Apple-silicon engines on an M3 Max 96GB. TensorFold led prose at 40.0 tokens per second, mlx-serve led code at 44.6, and the spread still ran from roughly 28 to 40 tokens per second, showing runtime choice matters even when the model and weights are identical.
  • David Ondrej said Vercel's fx harness now makes more sense as a complement to Opus 5.5's speed and expects the pair's adoption to jump in the coming weeks.
  • A “10 Tells of a Slop UI” checklist calls out recurring agent defaults such as purple gradients, rainbow fields, pulsing badges, tiny stacked cards, emoji stuffing, misaligned SVG/ASCII, Inter plus JetBrains Mono, leaked chat-context phrases, default glassmorphism, and generic “Elevate/Seamless/Unleash” copy. The Hacker News thread framed the problem as underspecified prompting plus shipping without taste, rather than proof that AI cannot design.
  • Reasonable argues the useful post-TLA+ step is an agentic path from specs to machine-checked implementation proofs, using TLA+ for behavior and Verus, Veil, or Lean closer to code. The team says it generated more than 3,000 machine-checked safety and liveness proofs from 16,459 real spec/property pairs and built a 40-task temporal-proof evaluation. The Hacker News thread pushed back that end-to-end proofs from code to high-level properties remain historically limited to small, specialized systems.
  • Anthropic’s Thariq Shihipar explained Claude Code effort as a compute-and-judgment dial: low/medium for interactive sketching and implementation, high/max for verification and hidden edge cases. Across 370 Terminal-Bench 3.0 tasks, passes rose from 140 to 214 as median tokens grew from 73K to 222K; higher effort reduced missed-edge-case failures but not wrong-approach failures. Anthropic’s effort writeup is here.
  • Linear’s Emil Kowalski showed a practical adversarial UI test: ask the model to break the interface it just built with long names, weird emails, extra labels, and other worst-case data. It is basically fuzz testing for layout.

🔬 AI Research & Models

  • DAIR.AI's September 21-27 paper roundup, with the companion post here, highlighted Xiaomi HySparse2 cutting prefill compute 5.02x and KV memory at 1M tokens from 12.09 GB to 2.69 GB; MIT/Sakana SIFT at roughly 10x cheaper self-improvement than DGM; GAVEL raising BEHAVIOR-1K single-task success from 41.2% to 91.8%; a Wiki Foundation Model training 10.5x faster; JEV-as-a-Judge at $0.044 per 1,000 judgments and 0.152-second median latency versus GPT-6 at $12.182 and 1.885 seconds; Google Harness-Zero, Stanford/Together self-organizing teams, ScientistTwo, XYEval, and EvoOntology.
  • Turing Post explained EvoOntology, an open-source system where a builder agent creates a three-layer ontology of schema, content, and tools, then other agents query it through Model Context Protocol instead of stuffing a stale dictionary into every prompt. Changes survive only when a paired evaluation improves. The paper reports gains of 17.8 points on multi-source research accuracy and 7.4 points on query accuracy, and the GitHub repo includes Claude Code and Codex plugins.
  • Alignment Forecasting predicts whether a supervised fine-tuning dataset will increase failures such as deception or power-seeking before training. Across 500-plus fine-tunes it reached AUROC 0.801; the project page, code, and LessWrong writeup are public.
  • Phys.org explained why AI weather models can rival physics-based forecasting systems globally yet still struggle with hurricane intensity. Oceans lack a complete three-dimensional training record, satellites see only slices, and intensity can behave chaotically, as Polo did in 2026 when it jumped from tropical storm to Category 5 in 24 hours and reached 180 mph. Chanh Kieu argues probabilistic intensity ranges are more realistic than one deterministic number.
  • Contrastive World Models replace Dreamer’s pixel decoder with a Deep InfoMax/InfoNCE objective that scores future local patch features from state-action sequences. The paper reports Dreamer-like performance on clean control tasks and stronger results when scenes contain bouncing-ball distractors or Kinetics video backgrounds, with faster training because the decoder disappears. alphaXiv highlighted it here, the v1 record is here, and Bonnie Li described the idea as preserving task-relevant information instead of reconstructing every pixel.
  • Tianhua Chen’s Little Book of Generative AI Foundations is a 195-page derivation-first primer moving from PCA and PPCA through VAEs, DDPMs, continuous-time score models, normalizing flows, autoregressive factorization, GANs/WGANs, and energy-based models.
  • iSDFT tackles the stability-plasticity tradeoff by limiting how much information the student can extract from an on-policy teacher at each token while separately anchoring to the frozen base model. Across four backbones and two specialization tasks, it beat vanilla self-distillation in seven of eight settings and tied one; 73% of retention evaluations stayed within 0.5 points of the base versus 52% for the best baseline. Haitham Bou-Ammar highlighted the work here.
  • Andon Labs says Opus 5.5 ranks first on Blueprint-Bench 2, an agent-only spatial evaluation that turns roughly 20 interior photos from each of 50 apartments into 2D floor plans scored on rotation- and reflection-invariant room-connectivity graphs.
  • A paper on LLM self-referential voice finds chat templates themselves can switch “I’m just an AI” disclaimers up and experiential language such as “I feel” down across eight instruct models up to 9B parameters; activation steering reproduced the effect in three models. The Hacker News thread treated that as evidence that some vendor-specific disclaimer voice lives in the deployment format rather than only in the weights.
  • Quanta surveyed work suggesting biology can use quantumlike mathematics without relying on durable quantum coherence: classical oscillator systems can still be written in Hilbert-space math resembling superposition and interference. Hacker News debated whether brief molecular superpositions could still matter before decoherence or whether electrochemistry plus dimensionality reduction already explains the observations.
  • Francis Bach shows why Monte Carlo estimates of log-sum-exp functions can explode in relative variance, then reframes the problem through f-divergence and relative-density estimation as a continuum of least-squares problems that collapse to one generalized eigen-decomposition.
  • OpenRouter said Jev reached 27% of weekly classification request volume, nearly twice DeepSeek V4 Flash’s previous lead, a platform-specific but concrete signal that prefill-only decision models are getting real routing traffic.
Advertisement

🏛️ AI Policy, Governance & Safety

🛠️ AI Tools, Products & Creative Demos

📊 Fundraising & Deals Roundup

  • PicoJool: $27.5M Series A led by Socratic Partners, with Hudson River Trading, to move its VCSEL and MicroVCSEL optical chips into AOC and NPO modules that aggregate up to 3.2 Tbps for GPU-cluster interconnects. Pat Gelsinger called connectivity one of AI's defining constraints.
  • Numeral: $100M Series C led by Insight Partners, with Salesforce Ventures, Geodesic, Benchmark, Mayfield, FCVC, Y Combinator, and Uncork, to scale its nexus-to-remittance sales-tax stack across new industries and more than 90 VAT/GST countries.

🎙️ Interviews, Panels & Podcasts

💡 Industry Commentary & Analysis

Previous Around the Horn Digest

  • Friday, September 25, 2026: Anthropic's Pentagon blacklist, Trump-Xi AI talks, Microsoft Copilot's agent push, and China's AI infrastructure buildout.

That's a Wrap

That's 260-plus links' worth of weekend AI. If you made it this far, your browser tabs are now eligible for a pension.

For the daily version, make sure you're subscribed to The Neuron. We send six issues a week, and yes, we read all of this so you do not have to.

See you tomorrow.

P.S. Know someone who'd find this useful? Forward this to them and tell them to subscribe here.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.