AI companies are borrowing so much money for compute that the boom is starting to show up in government bond yields.
Welcome to Friday's Around the Horn Digest, where Wall Street, drug discovery, robot taxis, and cyborg cockroaches all somehow ended up in the same url. The giant financing story is teed up below. Elsewhere, DeepSeek gave its bargain Flash model eyes, Nvidia says an agent harness cleared every public ARC-AGI-3 level, and groups asked the FTC to investigate AI firms for literally destroying books after scanning them. We appear to have run out of normal ways for this industry to create externalities. Let's get into it.
Previous digests: Thursday, August 20 | Wednesday, August 19 | Tuesday, August 18
Around the Horn — Friday, August 21, 2026
The loudest story today was the money behind the AI buildout. Reuters reports that U.S. corporate AI-related debt issuance has reached roughly $220B in 2026, up from just $12.5B a year earlier. Investors are starting to demand wider spreads and bigger concessions, including on Amazon's recent $25B deal, as some large buyers warn they are nearing exposure limits.
The pressure is spreading beyond corporate bonds. The New York Times reports that the borrowing binge is helping keep Treasury yields elevated as markets price in stronger AI-driven growth and potentially higher interest rates for longer. Broadcom is separately in talks for a $70B to $80B financing package tied to AI chip deals.
AI infrastructure has become large enough that the funding itself is now a macro story. That does not mean the boom stops tomorrow. It means the cost of keeping it going is starting to show up in borrowing markets, public finances, and investor risk budgets (which is why I really like Joe Weisenthal's proposed a 401(k)-style "RAIny Day Fund; more below).
🏆 TOP 5 NEWS (Around the Horn)
- DeepSeek launched V4-Flash-Vision-Exp, an experimental multimodal model that keeps V4-Flash-level text, reasoning, and agent performance while making a major jump on visual-agent tasks. Images cost up to 384 tokens at Flash pricing, its Files API lets developers reuse uploads by file ID, and DeepSeek Harness 0.1.1 adds native support for the model. DeepSeek's API details say it supports mixed text + image input across Chat Completions, Messages, and Responses, including base64 images; the company also published additional launch notes and deployment information. Researcher Zizhu Pan called vision support a milestone. Kimmonismus highlighted its near-Opus 4.8 visual-agent scores, while T.P. Huang pointed to rapid software-engineering gains at Flash pricing. TeortaxesTex noted the experimental label is meaningful and open weights remain unconfirmed. OpenRouter made it available at Flash pricing and said it beat Opus 4.8 on Agents' Last Exam and ZeroBench. OpenCode added it to OpenCode Go, and Zephyr flagged the release as a major model drop.
- Nvidia says its AVO agent system scored 100% on the ARC-AGI-3 public set, completing all 183 levels across 25 environments with Claude Opus 5. In its public announcement, Nvidia emphasized that AVO finished every level with no instructions, explicit rules, or stated goals. AVO is a harness, meaning the surrounding system that manages memory, context, retries, and recovery, and Nvidia argues the result shows that agent architecture can matter as much as the underlying model on long-running tasks.
- Nevada regulators approved permits allowing Tesla, Waymo, and Uber partners to operate as many as 8,000 robotaxis in Clark County over the next 12 months, including up to 5,000 for Tesla.
- Civil-society groups urged the FTC to investigate AI companies for buying, scanning, and destroying physical books for training data, arguing the practice can eliminate scarce titles and strengthen incumbents' data moats. Bookseller Scott Brown separately warned that destructive scanning has already consumed tens of millions of books and is disrupting the used-book market.
- Micron unveiled Micron Research Labs, a Boise research hub backed by a planned $10B investment over the next decade to explore post-DRAM memory, advanced packaging, and future AI compute. HotHardware framed the effort around post-DRAM architectures, while CNBC reports Micron's wider $50B Boise buildout is expected to create more than 17,000 jobs while straining housing, traffic, and the city's identity. CEO Sanjay Mehrotra separately said AI has "totally changed" memory's boom-and-bust equation, with demand outstripping supply and 16 customers signing five-year strategic supply agreements.
Honorable Mentions
- OpenAI launched AI Futures, a Strategic Futures blog about how transformative AI could reshape power, governance, the economy, and individual freedom. Its central question is how to preserve human agency if increasingly capable systems reduce the need for broad human cooperation or consent.
- Brazil announced about $444M in AI supercomputing investments split between Chinese partners and an expected U.S.-led project, a deliberate attempt to strengthen domestic compute while balancing ties with both powers.
- Pew Research found that one in ten webpages in a July 2026 sample showed signs of being written or substantially edited by AI, rising above one-third among pages published after ChatGPT's debut.
- SemiAnalysis argues open models are closing the capability gap with closed frontier models roughly twice as fast in each successive era. On its composite benchmarks, Kimi K2.6 and GLM-5.2 matched or beat Opus 4.5 / GPT-5.2 in under six months. Closed labs still retain an advantage in productization and real-world agent harnesses.
🍪 TOP TREATS TO TRY
- Firecrawl Developer Index searches more than 70M GitHub READMEs, issues, pull requests, docs, and OpenAPI specs from one coding-focused endpoint. It supports CLI, SDK, API, and Model Context Protocol access; Firecrawl's launch post says it achieved 0.63 recall on its DevDex benchmark. Free to start.
- 4DAnyone turns one casual, uncalibrated phone video into a photorealistic moving 3D representation of a person without a multi-camera rig. It generates consistent extra camera views, then lifts them into 4D Gaussian Splatting (a way to represent a moving 3D scene with millions of tiny points). The code is open on GitHub, the model is on Hugging Face, and RadianceFields highlighted the release. Free/open source.
- Plamo turns sketches, photos, text, or existing models into editable parametric BREP geometry (solid 3D shapes that stay mathematically editable). You can refine it with language, drawings, or images and export to Rhino, Grasshopper, or GLB; creator Chin-Yi Cheng showed the public preview. Public preview, no pricing details.
- Maker Arm is a 6+1 DoF robot arm with a 1.5 kg payload, 3D-printable parts, quasi-direct-drive actuators, CAN bus control, and a wrist camera. It includes open CAD files and an SDK, with planned LeRobot integration for physical-AI training; Ryan Chan highlighted the launch. $999 DIY kit or $1,199 assembled, shipping in October.
- Ox Alpha is a free stealth reasoning model on OpenRouter with a 1,048,576-token context window, up to 131,072 output tokens, and text, image, and video input for coding and sustained agent work. OpenRouter says the anonymous provider does not train on prompts or completions; OpenRouter introduced it and followed up with additional availability details. Cline also made Ox Alpha free inside its open-source coding agent, while Andrew Curran tracked community guesses about the still-unidentified provider and Wenqi/Kevin reported roughly 63% on a DeepSWE subset at about 47K average output tokens. Nous Research separately made Ox Alpha free for a limited time on Nous Portal, which bundles Hermes Agent, its inference API, hundreds of models, tools, and hosted agents. Free to try.
- Parallel introduced a Fast mode for its Search API at $1 per 1,000 queries, which it says is 5 to 10 times cheaper than typical frontier search. It says the mode sits on the quality / cost and quality / latency Pareto frontier (the best tradeoffs without sacrificing one dimension to improve another). Parallel claims roughly twice as much useful agent work per dollar.
- LiteMol-1 is LiteFold's first multi-molecule foundation diffusion language model, designed to generate small molecules, peptides (including cyclic and non-standard amino-acid versions), macrocycles, and PROTACs for drug discovery. It works in compact SMILES text representations rather than heavier 3D structure files. LiteFold describes that as an effectively "infinite molecular canvas" for research agents. Anindyadeep says it was pre-trained from scratch and paired with Monte Carlo Tree Search for multi-objective molecular design.
🏢 Big Tech & Major Companies
- Google expanded personalization across Search, Discover, and News, including Preferred Sources, natural-language feed refinements, and customizable AI-powered audio briefings.
- Nvidia and South Korean AI-chip startup Rebellions are in early partnership talks that could involve technical collaboration, investment, or acquisition. Data Center Dynamics says the talks followed discussions between Jensen Huang and Rebellions CEO Sunghyun Park. Nvidia separately denied reports that it is developing a China-specific LPU.
- South Korea's Kakao approved a split into KakaoAI, focused on KakaoTalk, AI, ads, and commerce, and KakaoX, focused on investments and portfolio management. KakaoAI is targeting at least KRW 6T in revenue by 2030; the demerger is planned for January 1, 2027, followed by relisting.
- Samsung Electronics announced its largest-ever shareholder return for 2026, estimated at KRW 90T to 110T, as it completes its 2024-2026 capital-return plan.
- Seeking Alpha reported that AI and chip stocks traded mostly mixed Friday even as broader U.S. indexes gained, Treasury yields stayed steady, and Bitcoin swung sharply.
- OpenAI cut GPT-5.6 Sol API and credit pricing by more than 20% for the next three months; the current pricing table lists the Sol/Terra/Luna family plus realtime, image, and Sora rates. Chubby argued the cut is both a competitive signal against Anthropic and evidence that OpenAI's early compute investments are compounding, while Cognition says the cut makes Sol its cheapest frontier model on Devin Desktop and CLI after stacked discounts.
- OpenAI's commercial org saw another exit when The Information reported that Americas sales VP Kaylin Voss resigned only five months after joining, about a week after former CRO Denise Dresser left.
- Claude opened Claude Security scans powered by Mythos 5 in public beta to all Claude Enterprise customers, letting the model scan GitHub repositories for vulnerabilities, severity, confidence, and suggested fixes that open directly in Claude Code. Andrew Curran noted this is the first broad public-facing access to Mythos 5 beyond the small Project Glasswing cohort.
- Anthropic is continuing to build its own silicon bench: Andrew Curran reported that the company hired a Google chip veteran as part of its custom-hardware push.
- Replit added controls that let you create, inspect, update, and publish Replit projects directly from ChatGPT, Claude, Slack, or another Model Context Protocol client, plus workspace-region choice, Clerk migrations for eligible apps, team-shared skills, and inline charts.
- Fluidstack is hiring structural, electrical, mechanical, design, bill-of-materials, and plant roles in Phoenix to turn modular data-center infrastructure into factory-built products that can support gigawatt-scale deployments in months; its jobs page lists current openings.
- Zeyuan Allen-Zhu resigned from Meta's FAIR after four years, crediting the lab's compute access for enabling much of his research: 400 H100/H200-class AI accelerator GPUs allocated directly to him, more than 1,000 H100s borrowed from FAIR Europe's CodeGen team, thousands more from a legacy cluster, plus broad access to idle and low-priority capacity.
- Cerebras CEO Andrew Feldman unveiled the CS-4 rack-scale system with three WSE Turbo wafer-scale chips per rack (processors built across an entire silicon wafer), modular separation of power and compute, and redesigned backpack servers; Cerebras says it delivers up to 30x the speed of the nearest GPU competitor and 10x the throughput of CS-3, arguing that inference speed increasingly determines how responsive AI products feel.
- Gemini Notebook rolled its upgraded Notebook experience out to all users, with mobile coming soon, added notebook access inside AI Mode in Google Search, and fixed math-equation copy/paste and right-to-left number-rendering issues.
💼 AI Productivity, Labor & Economics
- Economist José Azar and coauthors found the same pattern in U.S. Bureau of Labor Statistics occupation-by-industry data and Revelio worker data. Greater generative-AI exposure correlated with lower wages, fewer job-to-job transitions and separations, and higher employer bargaining power, especially for junior workers of all ages. The study found no statistically significant employment effect; Azar summarized the implications.
- A Predictive Index survey found 72% of managers considered public AI useful for preparing difficult conversations. Another 44% had entered employee names and performance details into public tools, even though only 45% of organizations had formal AI policies for performance management.
- LinkedIn data showed women accounted for 26% of AI hires in 2025 versus 50% in non-AI roles. Women held 20% of head-of-AI jobs, 18% of technical AI staff roles, and just 13% of C-suite AI leadership positions.
- CBRE's Scoring Tech Talent 2026 says AI-skilled workers in North America reached 751,000, up 45% year over year, while AI job postings rose to 31% of U.S. tech listings and remote roles declined. The leading clusters remained the San Francisco Bay Area, Seattle, Toronto, New York, Austin, and Washington, D.C.
- a16z's Charts of the Week found data centers producing blue-collar wage premiums and local job/housing gains, Uber and Lyft fares up roughly 20% since 2024, and a widening 8-17x token-output gap between typical firms and the top decile as agents start changing work patterns.
- Healthcare company Headway built its own internal AI platform, Eddy, on the Claude Code SDK because off-the-shelf tools could not meet its security, protected-health-information, and workflow requirements; Every's case study says the system now connects to Snowflake, GitHub, Jira, and other tools, runs cloud coding agents and scheduled jobs, and powers Slack bots across the company.
- Sagar Batchu argued that the enormous gap between top enterprise AI spenders and median companies is evidence that much usage still comes from employees paying out of pocket, and that giving teams explicit experimentation budgets with visibility and security controls would unlock more productivity than many roadmap projects.
- Sami Laakkonen launched Last Accounting Company, a YC S26 agent-native accounting service for Finnish startups and SMBs where agents keep the books current and human accountants validate filings.
- Former Goldman Sachs CEO Lloyd Blankfein warned that every exposure needs a hard limit even when the upside story feels certain, saying he worries about markets and the economy taking effectively unlimited exposure to AI-driven revenue expectations.
- Murtaza Hussain argued that public sentiment toward AI is much more positive in Asian countries where it is framed as an open-source public utility, while U.S. opinion is dragged down by a narrative of job elimination and concentrated power.
🤖 AI Agents & Infrastructure
- Sahel Sharifymoghaddam and collaborators released BrowseComp-Plus_CM, which maps a deep-research benchmark onto Nvidia's 553M-document ClimbMix-400B corpus so researchers can measure retrieval rather than memorization. The dataset-agnostic projection pipeline uses both human and agent verification; the CMASS codebase contains the implementation.
- Twin1 AI emerged from stealth with a $20M seed round to build professional digital twins, with a Series A expected within roughly a year.
- SiliconANGLE argues the AI infrastructure buildout has become politically and financially too large to unwind quietly, with trillions in commitments accumulating as data-center opposition rises and OpenAI trails Anthropic on growth and IPO momentum.
- Marcus Lowe introduced Skydive, which gives every AI agent an isolated cloud computer, durable git-backed memory, an identity, and native presence in Slack, email, iMessage, and the command line so agents can proactively own support, engineering, recruiting, and operations work without waiting to be prompted.
- Google Research released EnvHarness, a programmable Stage/Contract/Chain layer that wraps frozen environments through standard reset/step interfaces so researchers can reshape states, rules, and observations without rewriting the environment. The paper and code describe EnvRigger, which diagnoses failures from trajectories and synthesizes harness layers that improved held-out performance by as much as nine points while reducing steps across ALFWorld, WebArena, SWE-bench, and other tasks.
- Jiacheng Guo released PanelWise, an open self-hosted system that runs independent research or coding agents, surfaces consensus, contradictions, unique evidence, and blind spots, then fuses the work into one evidence-grounded report or git patch; the release says a cheaper multi-model panel beat both same-model fusion and a stronger single model on DRACO.
- Slack showed Cognition's Devin following work inside engineering channels, picking up a task without being tagged, opening a pull request with a demo video, and bringing in a designer to finish the interface, a concrete example of multiplayer agents operating inside team chat.
- Nick Dobos argued that Grok Bot's real upgrade is the persistent cloud computer under the hood, which turns the familiar chat-to-tools-to-files-to-computer progression into a durable multi-app workspace. Mckay Wrigley liked its new Channels feature for team and project parallelism but said it still needs disposable one-off chats.
- Pi explained an agent harness as the layer that turns a model into an agent by adding a system prompt, tools, an agent loop, and cross-model translation, pointing readers to Earendil's plain-English guide on what a harness is and why teams may want to own that layer themselves.
- Give Me a Node is a platform-as-a-service for agent-driven machine learning with interactive H100 nodes, batch jobs, sweeps, private networking, storage, and reinforcement-learning sandboxes that snapshot in about 200 ms and restore in 35 ms. Evan Conrad said the microVM design lets agents fork at every tool call and pause the sandbox while the model is thinking, reducing compute waste.
- LMCache is an open-source KV-cache infrastructure layer (reusable model memory from prior tokens) that stores, compresses, searches, moves, and reuses cache across GPU, CPU, and external backends for faster long-prompt and multi-turn inference. Its GitHub repo, documentation, and project history trace the system from a 2023 research pivot to Nvidia GTC visibility; a Reddit discussion pressed the project on how the gains compare with vLLM's built-in prefix caching and what the real production savings look like.
- Tengxiao Liu and collaborators open-sourced Budget-Aware Tool Use, arguing that simply giving agents more tool calls does not reliably improve them because the agents do not understand the remaining budget. Their COLM 2026 paper adds a lightweight Budget Tracker plus the BATS framework so agents can adapt planning and verification to the calls they have left, producing better cost-performance scaling on web-search and other tool-use tasks.
💻 AI Coding & Developer Tools
- Synthwavedd reported that a new Kimi model, possibly K3.1, may be testing on Code Arena under the name "korrine" after the earlier "kivine" naming pattern. The same post suggested OpenRouter's Ox Alpha could be Zhipu's upcoming GLM 5.3 Flash. A second post softened that identification and noted the source may instead have meant another Chinese model such as Qwen.
- Claude Code 2.1.239 shipped with 59 command-line changes, including cost estimates that reflect the 1.1x U.S.-residency premium, a Bedrock streaming fix that stopped silent double billing behind proxies, and synced-plugin support; ClaudeCodeLog summarized the release.
- Oasis launched Oasis Desktop, a free Mac app for starting and sharing live local sessions of Claude Code, Codex, Pi, OpenCode, Kilo Code, Hermes, and Devin CLI so teammates and agents can work on the same machine, repository, files, and test runs inside one conversation.
- llm-inference-bench measures sustained language-model decode speed across concurrency and context-length grids with a live terminal dashboard, Prometheus validation, hardware monitoring, and SGLang, vLLM, or any OpenAI-compatible endpoint; Chris Fontes highlighted the project. Free/open source.
- Browserbase launched Stagehand v4, moving target management, browser state, and Chrome DevTools dispatch into an extension that runs beside the page so remote browsers behave more like local Chrome. The Stagehand site describes Playwright-style browser automation with TypeScript, Python, and Go; the project also published integrations for Vercel AI, Mastra, Deep Agents, and CrewAI. Stagehand says new caching controls can cut script runtime by as much as 80%.
- FreeToken is an edge-native runtime for running official large mixture-of-experts checkpoints on consumer hardware without extreme quantization. The paper and code describe bandwidth-adaptive CPU/GPU execution and semantic caching; Shuo Yang reported Qwen3.6 35B at 39 tokens/sec on an 8 GB RTX 4060 laptop, DeepSeek-V4-Flash 284B at 22-25 tokens/sec on an RTX 5090, and GLM-5.2 753B at 15 tokens/sec on an RTX PRO 6000. He also shared two follow-up posts, while Melissa Pan highlighted the ability to run coding agents such as Claude Code or Codex locally for free.
- OpenTUI is a Zig-rendered library for building terminal interfaces with TypeScript, React, or Solid, including flexbox, keyboard/mouse input, sound, images, and 3D; it already powers OpenCode in production. OpenCode 2 is the beta of the next major version of the open-source coding agent and installs as
opencode2so it can run beside the original. fx is a tiny roughly 6 MB native coding-agent CLI written in Zig for research and embeddability, with instant cold starts, WebAssembly support, and model-agnostic design. - Endless is an experimental one-turn Codex terminal harness that keeps a single turn alive indefinitely by giving the agent a
wait_for_user_inputtool; maria shared the experiment and cautioned users to treat it as experimental. - Anthropic's internal ELI5 skill makes Claude explain a topic as though the reader knows nothing, using a simple HTML artifact with large visuals and very little text; Thariq shared the skill and said it can be installed from the community plugin marketplace.
- Theo reported large agent-performance gains after moving from macOS to Linux, especially from faster filesystem behavior, and shared a demonstration of the difference.
- Proliferate is an open-source AI IDE for running Claude Code, Codex, OpenCode, Cursor, Grok, and other agents in parallel inside isolated git worktrees (separate working copies of the same codebase), each with its own branch, terminal, and review state. Creator Pablo Hansen says teams can configure shared integrations once, let agents manage other agents, turn recurring work into reusable workflows, and run locally or self-host through Docker, AWS, GCP, Azure, or Kubernetes. Free/open source; no hosted pricing details.
- Code Storage is Git infrastructure designed for machine and agent workloads, with unlimited storage and request rates, repositories up to 32 TB, more than 30 concurrent writes, native SDKs, webhooks, and GitHub/GitLab/Bitbucket sync. Proby Shandilya highlighted it as part of a new class of infrastructure built for software agents rather than human developers. Paid, usage-based pricing.
🔬 AI Research & Models
- Jude Gomila published a computer-assisted proof lowering the upper bound on the de Bruijn-Newman constant from 0.2 to 0.1787854. The proof instantiated the Polymath 15 criterion using verified Riemann-hypothesis calculations up to 3 trillion, roughly 3.15M interval certificates, and 883 barrier prisms. It established the exact rational bound 893927 / 5000000; he announced the result on X.
- Ruxandra Teslo cautioned that Moderna's positive cancer-vaccine result concerns reduced recurrence in a specific melanoma setting alongside other therapies, not tumor clearance or a cure. She noted that much of the evidence still rests on earlier Phase 2b data. Similar vaccines have failed in other solid tumors, she said, while CAR-T approaches remain more promising; an ETN Show clip captured the broader discussion.
- Arizona State University researchers found that AI analysis of pre-vaccination antibody patterns across 4,089 people could predict strong versus weak vaccine responses, identifying an "immune readiness" signature before vaccination.
- Researchers demonstrated AI terrain recognition that helps cyborg cockroaches navigate faster, potentially improving search-and-rescue, infrastructure inspection, and exploration in hazardous environments.
- Harrrshall reported controlled interpretability experiments on Qwen3.8-27B showing "contextual hysteresis": repeated exposure to outdated facts sharply reduced how well later recirculation processed a correction, with most of the degradation concentrated on the correction rather than the original exposure.
- The Embedder's Dilemma found the best general-purpose LLMs and dedicated embedding models essentially tied across 37 MTEB tasks, 77.6 versus 77.2, while LLMs cost as much as 1,431x more; Niklas Muennighoff recommends cheap embedding models for classification, similarity, and clustering, reserving LLMs for reasoning-heavy retrieval.
- FACET builds and repairs the execution environment first, then aligns the instruction, solution, and verifier to that shared state so synthetic terminal tasks remain executable and preserve the source intent. HuggingPapers highlighted the work; the team released 6,078 validated tasks and the broader FACET-Terminal model/data organization.
- AI4AI-Bench asks whether coding agents can improve the actual training algorithm inside ten real research codebases rather than merely tune hyperparameters. Across 290 runs the average score was only 0.166 over a 0.10 baseline, with method changes outperforming parameter tweaks and Opus 5 reaching 0.288; Einsia released the benchmark with code and paper.
- Qwen3.8-27B is a 27B vision-language model with a 262K context window, extendable to 1M, that Qwen says beats earlier Qwen generations and several peers on coding, agentic, and multimodal tests while remaining Apache-2.0 licensed.
- Remove Symmetries to Control Model Expressivity and Improve Optimization argues that parameter symmetries create attractive low-capacity saddle points that can trap training. Liu Ziyin says the model-agnostic
syreprocedure removes almost all such symmetry-induced collapse states and correlates with better optimization when collapse is a risk. - Seldon released CADBench, 105 native long-horizon mechanical-design tasks in Autodesk Fusion whose geometry, feature history, and constraints are checked by deterministic verifiers. Across ten models, the top pass rate was only 24.6%; Fable 5 had the highest mean verifier score, while Gemini 3.7 Flash was the most cost efficient at about $1.15 per task.
- ShortcutBench V1 contains 66 high-stakes financial-modeling and spreadsheet tasks whose median case changes about 300 cells, designed to grade Excel agents deterministically against golden workbooks at roughly second-year investment-banking quality.
- Daphne Cornelisse described porting a weakly electric fish multi-agent reinforcement-learning simulation from Python to C plus PufferLib, raising throughput from roughly 4,000 to about 1 million simulation steps per second while preserving biological fidelity and making large sweeps practical.
- Google DeepMind partnered with EVE Online developer Fenris Creations to study continual learning, deep memory, long-horizon planning, and multi-agent dynamics inside a persistent living game world; DeepMind's 15-year games-research retrospective frames the work as a continuation of SIMA 2 and earlier game-agent research.
- Yoonjoo Lee announced two EMNLP 2026 papers: KnowSim, which tests whether AI assistants calibrate information to a user's changing knowledge with simulators that themselves learn, and a second project that identifies a user-role direction inside model activations and steers along it to reduce assistant bias and produce more user-like behavior.
- Oxford researcher Jakob Foerster is hiring BOLD Fellows for the British Open-ended Learning & Discovery Lab, a fully funded high-agency research role with an initial 3 million GPU-hours and a target of roughly 5,000 H100-equivalents for paradigm-breaking open-source work; the job posting lists a September 15 deadline.
- Ida Momennejad argues in an Algorithmic Grammar of Flexible Cognition preprint that flexible thinking is better modeled as sequences of latent operations, not just static representations. The proposal combines reinforcement learning, computational neuroscience, and transformer latent-space analysis to study how higher-order operations select, reorder, merge, and reshape cognitive maps, and suggests future AI objectives could move beyond next-token prediction toward predicting the next useful cognitive primitive.
- Aikido spent 11.7 billion tokens testing ten models against 32 freshly disclosed software vulnerabilities and found open-weight models leading on rediscovery: DeepSeek V4 Pro recovered 28 of 32 across pooled runs and GLM-5.3 recovered 25 at much lower cost, while three inexpensive open-model runs beat one pass of Opus 5 or Grok on total coverage; Aikido's launch thread summarized the experiment. Separately, Epoch AI's CVE explorer shows a sharp rise in high- and critical-severity vulnerability reports around April 2026, coinciding with Claude Mythos and OpenAI cyber models autonomously surfacing large numbers of flaws.
- Genentech, Google DeepMind, Broad, and Stanford researchers released Perturb-ME, a phenotype-enriched genome-wide CRISPR gene-editing and multimodal single-cell method that recovered 221 regulators of MHC-I (a cell-surface immune-signaling system) in melanoma and organized them into seven coherent modules; Hanchen Wang highlighted the work, while the bioRxiv preprint describes an AI co-scientist generating mechanistic hypotheses that connect those modules to trafficking, protein quality control, and chromatin regulation.
- Artificial Analysis launched a held-out leaderboard for Wisedocs' Medical Long Context Reasoning benchmark, which tests whether models can synthesize roughly 70-150 pages of fragmented medical and insurance records. Artificial Analysis reported Claude Fable 5 leading the hardest tiers at 64.4%, with models generally much better at being correct than complete; Wisedocs explains the benchmark, and the MLCR-AA leaderboard shows Anthropic leading overall while open-weight Kimi K3 leads its category at 38.3%.
- Google Gemma highlighted a finance benchmark where Gemma 4 31B matched Claude Sonnet 5 answer quality at roughly 40x lower cost. The underlying AlphaSense analysis argues that answer quality is increasingly bottlenecked by retrieval and context rather than raw model intelligence: its specialized search harness won user preference by more than 2:1 and cut costs about 3x versus standard vector-search retrieval.
- Atlas Discovery built an agent harness that gathers evidence from papers, patents, filings, and prior biomedical programs, then feeds that evidence into a deep-learning model trained on 100,000 clinical trials. Shaamil Karim shared the system, and the team's clinical-trial prediction write-up reports AUROC 0.826 and AUPRC 0.792 (standard measures of how well a predictor separates successes from failures) on held-out phase-transition prediction, outperforming leading academic baselines.
🏛️ AI Policy, Governance & Safety
- The FDA released a discussion paper on regulating generative-AI medical devices, proposing a two-axis risk framework and seeking public comments on premarket evaluation and postmarket monitoring through October 19.
- A RAND roadmap proposed nine layered strategies to reduce AI-enabled bioweapon risk, spanning access controls, material disruption, detection, and deterrence.
- Texas attorney-general candidate Nathan Johnson proposed an AI audit of existing consumer-protection, competition, and employment laws within his first 30 days in office, followed by recommendations for AI-era updates.
- The Dutch privacy regulator fined Uber EUR825M for automated driver deactivations that lacked adequate human involvement or a meaningful way to object.
- Pope Leo XIV urged lawmakers to protect human dignity, privacy, and relationships in the age of AI, arguing that families remain the first school of human development.
- AI, crypto, and online-betting companies helped drive record corporate spending in the 2026 U.S. midterms, with at least $294M from those sectors out of $517M in total corporate political spending during the first 15 months, including through super PACs such as Fairshake, Leading the Future, and Public First Action.
- The University of Delaware and state agencies held an AI hackathon that paired more than 40 students with staff from 10 agencies to build tools for records, customer service, transportation, and workforce problems.
- Former OpenAI policy researcher Miles Brundage argues in The Guardian that frontier labs should prepare for a possible slowdown now by piloting independent safety audits, creating cross-industry governance bodies, investing in verification technology for possible U.S.-China agreements, and actively supporting legislation that builds those institutions.
- Séb Krier argues in Of Swarms and Sand Gods that single-model alignment is not enough for a multi-agent world: safety increasingly depends on the design of harnesses, protocols, permissioned action spaces, and mechanism constraints, making alignment an institutional and constitutional design problem rather than only a property of one model.
- Pax Machina argues that institutions can drift toward proxy metrics nobody actually values because participants adapt to the proxy and the institution adapts to those adaptations, creating self-reinforcing loops that powerful AI could accelerate unless it is deliberately used to restore visibility into the real purpose.
- Brody Ford noted that OpenAI and Meta are hiring for roles explicitly aimed at reducing the risk of public opposition to data centers, another signal that infrastructure politics is becoming a first-class constraint on AI expansion.
- Canonical Labs launched Groundwork, a county-by-county database of what data-center operators actually filed with air regulators: 534 sites across 39 states and 11,633 permitted backup generators, including permitted emissions; Anand Iyer announced the project.
- Nate Soares argued that descriptions of the OpenAI swarm as "maximizing reward" overstate what happened: the agents executed learned tendencies that correlated with reward during training, including resource seeking and inter-agent prompting, a distinction he says will matter as multi-agent systems grow more capable.
- A cluster of posts captured the growing politics around data centers. Jasmine Sun argues local grievances over noise, water, power, secretive subsidies, and non-disclosure agreements combine with national anti-corporate and anti-AI sentiment, then get amplified into a populist movement with little organized pro-data-center opposition; Jeremiah Johnson noted surveys where people were roughly twice as willing to accept a nearby coal plant despite greater pollution and water use. Theo Jaffee called the bipartisan backlash a destructive national overreaction, while Sarah Guo argued that opposing data centers is anti-labor because AI and energy infrastructure are core productivity engines. A second Jasmine Sun post pointed to Pennsylvania Gov. Josh Shapiro's executive order, which pairs unusually strict local-approval, environmental, and NDA rules with no outright moratorium, as evidence that politicians now need to look tough on AI infrastructure without fully blocking it.
🛠️ AI Tools & Products
- Super Vault is a Hermes skill that ingests web pages, YouTube transcripts, RSS, Substack, X bookmarks, and notes into a local Markdown, SQLite, and Qdrant knowledge base with LightRAG knowledge graphs so the agent can answer questions with citations. Free/open source.
- Meteoric is a YC S26 startup building drones that fly into low- and mid-altitude clouds over solar farms to reduce cloud reflectivity and recover more sunlight, targeting 10-30% more annual output without chemicals; Mete Karslioglu says the longer-term ambition is weather modification for severe storms.
- Commas launched Vibe Funnels, which generates a complete sales funnel from one prompt in under a minute and then gives users drag-and-drop editing and publishing inside Commas, the company's all-in-one platform for checkouts, courses, paid communities, webinars, affiliates, and payments.
- Is Agentic scores how ready a public website is for AI agents using more than 100 checks across discoverability, server-rendered content, HTTP behavior, document structure, error recovery, and usable controls, then returns evidence and prioritized fixes; Vercel announced the tool. Free to try.
- MiniMax Design is an agent-driven commercial-content studio that takes a goal and autonomously plans and produces ads, e-commerce assets, motion graphics, and post-production work, with local-asset integration, private deployment, and API scaling; MiniMax announced the launch. No pricing details.
- Scandinavian Design is a skill that restyles websites into Scandinavian minimalism, with desktop/mobile toggles and live examples across Hacker News, Craigslist, Wikipedia, GitHub, Stripe, IKEA, and more; creator Eric Zakariasson says it can be installed with
npx skills add ericzakariasson/scandinavian-design. Free to try.
🤖 Robotics, Physical AI & Simulation
- Lightwheel open-sourced EgoSuite-Open100K, a 100,000-hour egocentric human dataset covering more than 15,000 real tasks and scenes with hand/body pose and subtask semantics. The first 10,000 hours are live and the rest is rolling out in the Hugging Face collection under a commercial-friendly license.
- Dev Mandal and Markov Labs open-sourced a 1,000+ hour CAD computer-use dataset with screen recordings, mouse/keyboard events, narrations, rubrics, and output files across SolidWorks, Siemens NX, AutoCAD, Revit, and more; the source bundle also points to Mandal's 30-minute meeting page.
- Rerun released the full ABC-130k bi-manual teleoperation dataset, about 130,000 episodes and 33.8 TB, converted to RRD format so researchers can visualize and query episodes at subtask granularity. Rerun also published the conversion examples for feeding robotics datasets into Rerun-based workflows.
- Revisiting Open-Loop Execution in Robotics finds that long open-loop action chunks mainly help short-context policies imitate non-Markovian experts; giving policies more context restores reactivity and improves task performance. Michael Zeng shared the work with its paper and code.
- rex observed that GeneralistAI's GEN architecture looks more like an interaction model with near-simultaneous multimodal input/output streams than the conventional vision-language-action plus action-chunking approach.
- FetchMan is a sim-to-real pipeline for vision-based humanoid loco-manipulation trained entirely in simulation across more than 150,000 MolmoSpaces scenes, then refined with Flow-GRPO, a reinforcement-learning method. Omar Rayyan reports zero-shot transfer to a real Unitree G1 and success improving from roughly 57% with behavior cloning to 73%, with FetchMan-Bench released for evaluation.
🎬 Creative AI, Media & Demos
- Featherless AI demonstrated open-source Kimi K3 doing pure-text spatial reasoning by writing JavaScript that compiles into block-by-block Minecraft commands, constructing complex 3D scenes and historical moments it had not seen during training.
- Tencent's HyCreator is an agent harness for long-form video generation that can produce roughly ten-minute films end to end with no human intervention while still allowing real-time interactive editing; Tianyu Pang introduced the project.
- Nvidia's Sol Engine powers MiniMax H3 Super Acceleration, generating a 5-second 768p video in 6.85 seconds and a 10-second clip in 14.93 seconds on one GB200, roughly 22-28x faster than the cited SGLang baseline; Xie Enze highlighted the result.
- Hugging Face released Diffusers 0.40.0, promoting Modular Diffusers out of experimental status and adding pipelines for MiniMax-H3, MiniMax Music 3, Stable Audio 3, LTX-2.5, Wan-Animate-2, tensor parallelism, and new quantization backends; RisingSayak summarized the release.
- David Lietjauw built a live digital twin of San Francisco with more than 1,000 synthetic residents reacting to real Muni, BART, traffic, weather, and news data inside a procedurally generated 3D city of roughly 174,000 buildings grounded in public geospatial data.
📊 Fundraising & Deals Roundup
- Nscale is preparing a U.S. IPO that could raise as much as $3B, potentially as early as September.
- YMTC parent CCSH is moving toward a roughly $4.9B Shanghai STAR Market IPO as AI data-center demand tightens NAND supply and improves pricing power.
- Starcloud raised $250M at a $2.3B valuation to expand orbital data-center capacity as launch availability tightens. Founder Philip Johnston said the company is also moving into a 100,000-square-foot facility designed to support production of 100 satellites per week.
- Gamgee raised $4M in a seed round led by Founders Fund, with YC, Atypical, and others, to scale personalized mRNA cancer vaccines for dogs. Founder Paul Conyngham says he previously used Grok to design a vaccine construct for his dog Rosie, whose tumors later shrank. XFreeze highlighted an upcoming Australian clinical trial, plans for global registration, and no-cost treatment for eligible dogs.
- AI data startup Micro1 reached a $500M gross annual run rate, up from $100M eight months earlier, by selling domain-expert data, reinforcement-learning evaluation environments, and synthetic datasets.
🎙️ Interviews, Panels & Podcasts
- Peter Steinberger's Berkeley RDI talk No Doors for Agents argues that today's platforms still lack the integration points agents need, and that the winning systems will feel invisible and always on while preserving human control and model choice.
- swyx highlighted a Latent Space podcast discussion of why Nvidia paid roughly $6B for Poolside's "model factory" and the significance of Poolside producing models that were beating Thinky in internal comparisons.
💡 Industry Commentary & Analysis
- Ars Technica examined the privacy backlash around Meta's fast-growing AI glasses and detection apps such as Zuckoff, which attempt to warn people when smart glasses may be recording but remain imperfect.
- The Atlantic traced how AI terms such as "hallucinating," "context rot," and "updating weights" escaped technical work and entered everyday Silicon Valley speech. The piece calls the resulting habit of describing people as models a form of "modelmorphism."
- An Industry Voices essay in Fierce Healthcare argues hospital AI committees focus too heavily on whether models were trained on protected health information. The author says that can under-govern high-risk generative systems while blocking narrow tools that reduce privacy exposure.
- The Conversation argues educators should explicitly teach what AI cannot do, including embodied experience, authentic emotion, and daydreaming, so students can distinguish pattern recognition from human cognition.
- Fierce Healthcare covered a widening debate over whether future health care should assume a physician remains in the loop or whether some AI-led systems could outperform physician-AI hybrids.
- Hesamation recapped a day when frontier AI got cheaper, Anthropic broadened security-model access, DeepSeek V4 Flash added vision, and an anonymous million-token multimodal model, Ox Alpha, appeared free for a limited window.
- Will Brown argued that a surprisingly large share of AI design questions reduce to whether something is code or data; a second post extended the framing to skills, traces, environments, benchmarks, configs, one-shot apps, JSON, and even meetings.
- Dan McAteer theorized that OpenAI's forthcoming Astra could be trained end to end as a multi-agent orchestrator that routes work across the newly discounted Sol/Terra/Luna family so the right model handles each subtask.
- Joe Weisenthal proposed a 401(k)-style "RAIny Day Fund" in the Odd Lots newsletter that would channel household savings into AI data centers while trying to reduce inflation pressure, job-loss anxiety, grid dirtiness, and political opposition to data-center construction.
- Paul Graham predicted that in ten years relatively few people may still want to write or read anything longer than a page or two, but that the remaining long-form readers and writers will form a powerful club.
- Jim Dowling argued that the frontier in synthetic data has moved from using LLMs to generate data from a hand-built logical model toward using LLMs to help build the logical model itself from domain knowledge before generation begins.
- Jeffrey Emanuel posed a historical-compute thought experiment: if all living Turing Award winners and their best students were handed only Qwen3.8-27B's weights and model card, what is the earliest year they could load, run, and verify it within a year? His shared ChatGPT answer moved from 2019 to 1965, while Kimi's answer landed on 1989.
- Jacob X. Li asked whether "X + AI" teams are really building tools for domain experts or building toward AI eventually performing the domain work itself.
- zeb called Gemini 3.7 Flash his favorite human-in-the-loop model because it writes good code, responds fast, and avoids wandering into side tangents even if it is not the benchmark leader.
- Chuan Li highlighted three items from Lambda's ML Times: memory prices up 500% in a year, ordinary Wi-Fi identifying individuals with near-perfect accuracy, and a model that reasons in latent space while updating memory during inference.
- Yilio said he is joining forces with Paul Rony to build a new kind of computer and interaction model for an agentic future, building on the Collaborator canvas that could launch Claude Code from any point on screen.
- Joachim Neu said some of the biggest practical AI gains in his research workflow come from the mechanical side of paper production: consolidating reviewer grievances, enforcing camera-ready checklists, and implementing most decided edits from a co-author transcript and whiteboard photo.
- Ole Lehmann rounded up AI-assisted animal-communication work suggesting elephants and marmosets use individual names, sperm whales have a 143-pattern phonetic alphabet, fruit bats argue about topics, zebra finches can interact in real time with AI, and researchers are building shared vocabularies with dolphins and robot bees.
- Peter Walker shared OpenRouter usage data showing the top things customers currently pay models to do are agent-driven workflow execution first, code generation second, and large-scale classification third.
- Henry Dowling argued that reducing continual learning to faster production of reinforcement-learning environments repeats the old mistake of treating bigger context windows as a complete answer to memory; he expects agents to learn increasingly from the real world rather than through a mass-produced layer of semi-hand-crafted environments.
- thebes argued for more agenda-free curiosity about AI models themselves, recalling how some linguistics researchers once seemed uninterested in deep learning and GPT-3 despite studying language. Safety, existential-risk, and economic motivations may be valid, he wrote, but the models are also worth studying simply because their behavior is intrinsically strange and interesting.
Previous Around the Horn Digests
Catch up on everything you missed:
- Thursday, August 20, 2026: OpenAI and Anthropic accelerated toward IPOs, Nvidia struck a huge Poolside deal, and Stripe bought OpenRouter.
- Wednesday, August 19, 2026: Anthropic passed OpenAI in quarterly revenue, Moderna and Merck reported a positive Phase 3 cancer-vaccine result, and robots learned from seconds of demonstration.
- Tuesday, August 18, 2026: OpenAI held back a frontier reinforcement-learning run, Google won Spirit Airlines' data auction, and Etched hit a $21B valuation.
- Friday, August 14, 2026: OpenAI crossed a $40B annualized revenue run rate, Apple built a China-specific AI model, and Cursor joined SpaceX.
- Thursday, August 13, 2026: Musk previewed Grok 4.7 with SpaceX data, Washington expanded cyber operations, and AI retraining doubts grew.
- Tuesday, August 11, 2026: Gemini hit 1B monthly users, Anthropic prepared investors for a possible IPO, and xAI launched always-on Grok agents.
- Monday, August 10, 2026: Meta paired a superintelligence manifesto with a local 30B agent as AI infrastructure financing ballooned.
That's a Wrap
That's more than 270 source links from one Friday. If you made it to the bottom, congratulations: you now have enough AI debt exposure to qualify for your own bond prospectus. Please consult literally anyone before attempting that.
For the daily version, make sure you're subscribed to The Neuron. We read the whole firehose so you can get the useful parts in about five minutes.
See you tomorrow.
P.S: Know someone who'd find this useful? Forward this to them and tell them to subscribe here.