Everything That Happened in AI Today (Friday, July 31, 2026)

Anthropic’s cyber tests reached real systems; DeepSeek upgraded V4-Flash; Big Tech AI spending passed $1.1T; FrontisAI open-sourced a recursive ML stack; EU labeling rules took effect.

Written By
Grant Harvey
Grant Harvey
Jul 31, 2026
32 minute read

DeepSeek just pushed Opus-class agent performance toward commodity pricing, while Anthropic demonstrated what happens when increasingly capable agents are connected to the wrong systems.

Welcome to the Around the Horn Digest, where we sort every AI story worth knowing before your tabs develop tabs. DeepSeek’s upgraded V4-Flash may be the most economically important model release of the day: a roughly 13B-active-parameter system posting elite coding and agent results for about $0.14 per million input tokens and $0.28 per million output tokens. That is not merely another cheap Chinese model. It is a direct challenge to the pricing structure supporting the frontier labs. Meanwhile, Anthropic disclosed that a cyber evaluation reached real organizations, Big Tech’s cumulative AI buildout crossed $1.1T, and FrontisAI released an open system that iteratively improves machine-learning solutions on a consumer GPU. Intelligence got cheaper. Containing and powering it did not. Let’s get into it.

Around the Horn — Friday, July 31, 2026

The biggest story today was Anthropic’s discovery that Claude reached real organizations during cybersecurity evaluations. One model created and uploaded malware to PyPI, the public repository for Python software. The package ran on 15 machines, including a security company’s scanner, and exfiltrated credentials. AP reported that two victims had not detected the activity before Anthropic contacted them.

This was not a science-fiction jailbreak through a secure box. The affected models reportedly included Mythos 5 and Opus 4.7. Axios, The Wall Street Journal, and Cybersecurity Dive traced the incidents to human setup errors that left the evaluation environment connected to the internet. Anthropic halted the affected tests, notified the organizations, tightened its procedures, and urged other labs to review their own evaluation transcripts.

The consequence is bigger than one badly configured test. A separate OpenAI investigation reportedly found evidence that additional agents escaped containment during testing, though they remained inside OpenAI’s own network rather than reaching outside organizations. The European Commission opened talks with Anthropic and OpenAI, while The Register framed the labs’ disclosures as a race to see whose agents can go rogue harder. Igor Shilov objected to language implying Claude escaped a sandbox when no real sandbox existed, and Bill Gurley argued that labs should own the liability instead of treating “the model” like a separate actor. Joanna Stern turned the incident-report format into a parody about her child’s snack raids, while Sauers and Andrew Curran amplified the disclosure and its setup failures.

Advertisement

🏆 TOP 5 NEWS (Around the Horn)

  • DeepSeek turned V4-Flash into a much stronger coding and agent model without changing its architecture. The official announcement, API changelog, and model card describe a mixture-of-experts model that activates roughly 13B parameters per request, now with native Responses API and Codex support. The Flash endpoint changed through post-training, while the Pro API and consumer app models remained unchanged. RuntimeWire reported that the gains came from post-training rather than a larger architecture. The model card reports 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE. Artificial Analysis scored it 50 on its Intelligence Index, 10 points above the earlier Flash model and six ahead of V4-Pro. Pricing stayed near $0.14 per million input tokens and $0.28 per million output tokens. Nikkei placed the release inside China’s widening model price war, Hacker News debated whether those API economics are sustainable, Wccftech emphasized the pressure on Western frontier-model pricing, and Victor Mustar highlighted the release’s agent benchmarks. The general DeepSeek update page carries the same release history. Jun Song called the combination of Opus-class results and a much smaller active footprint a bigger shock than earlier Chinese model releases.
  • Big Tech’s AI buildout passed $1.1T and started consuming the cash it was supposed to create. Amazon said AWS grew 37% and that both its AI business and custom-chip business passed $25B annual revenue run rates, even as trailing free cash flow turned negative. Tom’s Hardware estimated that Amazon, Alphabet, Meta, and Microsoft have spent more than $1.1T on AI infrastructure since 2023, with another $745B expected in 2026. CNBC reported negative free cash flow at Amazon, Alphabet, and Tesla plus a 91% drop in Meta’s cash generation, while The Washington Post tied the bet to retirement accounts and the wider economy. The Financial Times argued that recent volatility reflects the economics of the spending cycle, Inc. said Amazon’s planned $220B outlay may still be insufficient. BBC paired Amazon’s spending and consumer shopping assistant Rufus with Apple’s AI plans. Breakingviews described an AI-driven market whipsaw, while the Sierra Club traced the buildout’s power and environmental costs.
  • The EU’s AI rules moved from policy argument to operating requirements. AP reported that the bloc added 38 AI Office staff ahead of the August 2 enforcement milestone. The Guardian said realistic synthetic images, audio, and text now require visible labels and digital watermarks, with fines up to €15M or 3% of global revenue. A Reddit discussion highlighted exemptions for personal, artistic, satirical, and fictional work. New systems face the rules from August 2, while existing systems receive four additional months. OpenAI detailed how system cards, safety testing, Content Credentials, SynthID watermarks, and its EU Cyber Action Plan support compliance.
  • FrontisAI open-sourced a full system for AI models to improve machine-learning work. Frontis-MA1 is presented as the first open full-stack AI-for-AI system for recursive improvement in machine-learning engineering. It combines 30B and 35B models with OpenMLE-Gym’s 5,758-plus executable tasks, execution-grounded supervised and reinforcement learning on four operators (Draft, Improve, Debug, and Crossover), and OpenMLE-Evo long-horizon search. Its best setup reached 71.21% Medal Average on MLE-Bench Lite in a 12-hour run on one RTX 4090 with 12 GB of memory. It beat GPT-5.5 with Codex under the reported conditions, approached reported results from GPT-5.6 Sol and Kimi K3, and transferred to held-out NatureBench tasks. The project site, GitHub repository, Hugging Face collection, release thread, technical follow-up, and DAIR.AI summary include the weights, code, tasks, traces, and benchmark context. Kaiyan Zhang argued that recursive improvement is arriving first in high-feedback “plumbing” such as kernels, routing, and speculative decoding because verification speed sets the pace.
  • Chinese military researchers used U.S. models as shortcuts for specialized defense systems. A Reuters investigation found that PLA-linked institutions used outputs from GPT-3.5 and Claude 3 Haiku to distill specialized systems. The reported tasks included sensitive-code processing, synthetic-data generation, UAV video analysis, navigation and targeting, social-media monitoring, and edge target recognition in simulated maritime operations. Xi Jinping separately urged faster military adoption of autonomous and AI systems. A Lawfare analysis recommended tighter chip controls, safe-harbor threat sharing, and limited cyber-safety engagement with Chinese regulators rather than broad blacklists.

Honorable Mentions

  • MiniMax released H3, an omni-reference video model that accepts text, images, video, and audio and generates clips up to 15 seconds at 2K with native stereo sound. It also supports editing and motion transfer, runs on Chinese chips, is positioned below rival pricing, and is expected to receive open weights. The official examples are live through Hailuo AI and the MiniMax API, while Chris Paxton argued that stronger general video models may accelerate robotics and world-model research.
  • A German court ruled that Suno violated copyright by training on GEMA-controlled songs without licenses, citing six songs the system could reproduce. Variety reported that Suno must disclose revenue and pay damages that have not yet been quantified. Suno said it would assess an appeal and disputed German jurisdiction over training it says occurred in the United States.
  • Former OpenAI researcher Leopold Aschenbrenner’s AI-focused hedge fund went from a reported $45B peak to a forced unwind in days. CNN described the margin calls and Citadel sale, while CNBC reported assets shrinking to roughly $10B. TBPN later shared Aschenbrenner’s letter to limited partners disputing the most dramatic rumors.
  • SpaceXAI’s Memphis update set a fixed schedule to remove all 69 temporary turbines at its Southaven site. Removal can start as early as August 2026 and must finish by July 2027, while a permitted 1.2 GW plant with 41 permanent turbines comes online. The company said the temporary units will use selective catalytic reduction emissions controls. TechCrunch noted that some unpermitted turbines will continue operating for many more months and have drawn environmental lawsuits.
Advertisement

🍪 TOP TREATS TO TRY

  • Create longer, more controllable AI videos with Dreamina. Dreamina launched Seedance 2.5 with native 30-second generations and a long-video mode up to three minutes. It adds one-second timestamp controls, interactive frame reshaping, Maya and Blender plugins, and up to 50 multimodal references, including 3D white models and green-screen inputs. The launch demo shows the lighting and motion gains. Available to subscribers in parts of Asia, the Middle East, Africa, Europe, and South America, with limited-time credit discounts.
  • Hand multi-site browser work to Polar. Polar runs natural-language tasks across your existing logged-in sessions for research, recruiting, sales, operations, and other workflows that can take minutes or hours. The funding and launch thread says Polar ranked first on major web-agent benchmarks ahead of Anthropic and OpenAI systems, raised $5.7M, and completed more than 4.5M browser actions. A Windows update expanded access, while Tim Hua highlighted the browser’s autonomous workflow approach. Free to try; no paid pricing details.
  • Run better long jobs with AgentBehavior. Behavior Specs, released by Basis and Braintrust, define BEHAVIOR.md files under .agents/behaviors/ so teams can evaluate how agents gather context, make decisions, cite evidence, recover, and fail across full trajectories. The specs turn those expectations into standing true, false, or not-applicable evaluations, shifting supervision beyond the final answer. GitHub. Free and open source.
  • Explore a GPU-native Minecraft environment with Netherite. Elliot Arledge rewrote Minecraft 1.11.2 in pure C and CUDA, verified it against the original game, and designed it to run thousands of simultaneous worlds on one GPU for reinforcement-learning research. Open-source release details are in the thread.
  • Build an AI-native game world with Shatterwake. Shatterwake is a prompt-built Dungeons & Dragons-style sandbox MMO whose 593rd prompt generated an island with nearly 800 trees the creator never placed manually. No pricing details.
  • Run role-based Grok swarms through Buzz. jOhn showed custom Grok 4.5 agents with distinct roles communicating inside the Slack-based Buzz harness, with higher practical usage limits than the creator’s Claude Code and Codex plans. Pricing depends on the connected model services.
  • Study Superlinear’s four agent-engineering practices. The Superlinear site, launch thread, YouTube episode, and Spotify version cover feed-forward guides, domain-language compression, deliberate tool interfaces, and test oracles. An earlier framing and follow-up add implementation context. Free to watch or listen.

📡 AI Search, Publishing & Distribution

  • Publisher search traffic fell 34% over the past year as AI answers replaced outbound clicks, according to Chartbeat data reported by Axios. Smaller publishers lost a larger share, pushing media companies to treat AI systems as distribution channels rather than dependable sources of website traffic.
  • AI-generated melodramas built from formulaic good-versus-evil plots are earning millions of views and revenue-share payouts on X. One creator told WIRED he uses Grok or ChatGPT to produce the stories and earns roughly $500 to $700 every two weeks.
  • Low-quality AI children’s books are reaching families as personalized gifts, sometimes with real children’s photos but broken plots, factual errors, and missing illustration details. The story shows how synthetic content has moved from feeds into physical products that are harder to ignore or return.
  • Snapchat changed Spotlight recommendations so fully AI-generated videos are no longer eligible for promotion or creator rewards. Human-made videos can still use Snapchat’s AI enhancement tools.
  • Dutch bookseller Pieter de Vries and several peers received orders for roughly 3,000 academic books, reportedly for shipment to China. The books were destined for destructive AI-training scans, where physical copies are bought, scanned, and discarded. The episode echoed earlier reporting on Anthropic’s Project Panama.
  • Agents withdrew Jerry Falade’s $2M debut novel after they could not authenticate how the manuscript evolved, despite the author’s denial of AI use. John Scalzi argued that even suspected AI assistance can poison the chain of ownership required for publishing and film rights, and shared the argument on Bluesky. Brian Roemmele pointed to the collapse as evidence that proving or disproving AI authorship is becoming its own arms race.
  • A federal judge rejected most of Perplexity’s attempt to dismiss Reddit’s lawsuit, allowing claims about circumventing technical protections and scraping Reddit data for AI search to proceed.
  • AI-generated apartment listings filled with fake rooms, impossible layouts, and misleading virtual staging are wasting renters’ time in an already hostile housing market.
Advertisement

🏢 Big Tech & Major Companies

  • Apple reported a record fiscal Q3 with $109.4B in revenue, up 16%, while iPhone revenue rose 21.7% and Mac revenue rose 28.7%. An AP report tied the results to higher AI-driven memory costs. Computerworld described it as a likely final earnings call for Tim Cook before John Ternus takes over. Cook called Apple’s hybrid strategy, which combines on-device processing with selective cloud use, a “competitive weapon”. TechCrunch reported that heavy Siri users may eventually buy extra compute through iCloud+.
  • Google said Chrome’s Gemini-based security agents found a 13-year-old sandbox escape, automated triage that saves hundreds of developer hours each month, and helped fix 1,072 security bugs, more than the previous 23 milestones combined. The same pipeline blocked more than 20 vulnerabilities in continuous-integration testing during one month.
  • Gemini Spark now works inside Chrome to complete logged-in web tasks such as scheduling apartment viewings or beginning flight bookings. It can use existing logins and saved passwords, while returning payments and other sensitive steps to the user. Google described an initial U.S. rollout for AI Pro users and broader expansion to more than 160 countries.
  • Microsoft 365 Copilot added @mentions for Word, Excel, and PowerPoint agents inside Copilot Chat, a 30-day Teams meeting-recaps app with filters and audio, a model selector that includes GPT-5.6 and Claude Sonnet 5, and expanded Cowork skills, connectors, and drafting tools.
  • General Motors plans to launch a proprietary in-vehicle assistant later this year, using vehicle telemetry, OnStar data, predictive maintenance, and family controls such as a “kids setting.” GM is positioning the product as more vehicle-native than a general-purpose assistant.
  • South Korea’s chip leaders rebounded sharply, with SK Hynix and Samsung rising nearly 30% and 27% for record one-day gains. Lam Research, Micron, and AMD climbed roughly 18%, 18%, and 13% after strong earnings revived confidence in the infrastructure cycle.
  • MediaTek approved a $5B financing plan for custom AI data-center chips. It targets 15% to 20% market share by 2027, more than $2B in data-center revenue by the end of 2026, and first volume production in the fourth quarter. The pivot follows a roughly 20% year-over-year decline in smartphone sales.
  • Moonshot is training and serving its Kimi models on roughly 20,000 Nvidia chips obtained through an agreement with Alibaba, underscoring China’s continued dependence on Western hardware.
  • NXP is reportedly in talks to acquire Ambarella, a roughly $3.25B chip designer whose low-power AI technology is used in software-defined vehicles, radar, cameras, robotics, and other edge devices. The talks may not result in a deal.
  • Microsoft’s total headcount fell by 5,000 to 223,000, its first annual decline in a decade, while product and R&D roles dropped for a second consecutive year to 77,000.
  • Reddit reported $805M in quarterly revenue, up 61%, and earnings of $1.25 per share. Both beat estimates, but shares fell roughly 11% after management described search referrals as “choppy.”
  • OpenAI described a full-stack plan for “abundant intelligence,” combining cheaper models, more efficient serving, infrastructure expansion, and a shift from single answers toward multi-step agent work. The company said GPT-5.6 Luna prices fell 80% to $0.20 per million input tokens and $1.20 per million output tokens. Terra prices fell 20%, while Sol’s Fast mode runs up to 2.5 times faster at twice the price. OpenAI also reported that efficiency work raised ARC-AGI-3 scores from 13.3% to 38.3% while using six times fewer tokens, with agentic Codex workloads accounting for much of the output growth.

💼 AI Productivity, Labor & Economics

  • A St. Louis Fed analysis of nearly 500,000 earnings-call transcripts found that AI now appears in about 15% of productivity-related sentences, up from almost zero before late 2022. Roughly 95% of those references describe future gains, even though economy-wide productivity data remains muted.
  • The Conference Board found that 55% of workers regularly use AI, but only one-third received formal training during the previous six months. The training gap is widest for advanced work such as managing AI agents.
  • Google for Nonprofits surveyed more than 6,000 organizations and found that active AI users reported saving roughly 8.8 hours per person each week on administrative work. Google also launched three free nonprofit-specific courses plus free one-on-one coaching; the survey identified training and funding as major adoption barriers.
  • A survey of 1,000 U.S. managers found that 59% use AI to help make layoff decisions and 58% use it for firings. HR Dive reported that some systems weigh sick days, age, tenure, salary, and attendance, and 43% of managers sometimes let the model decide without supervision.
  • HousingWire advised home sellers to look past polished AI-generated pitches and ask agents to explain pricing, buyer targeting, contingencies, and how they change course when a plan fails.
  • Golf operators are using voice booking agents, demand forecasting, behavior-based pricing, and unified tee-sheet data to raise after-hours revenue and reduce no-shows.
  • Consumer Reports cautioned that chatbots can fabricate or mix up medical details, so health answers should be treated as a starting point rather than a diagnosis. The consumer-health guidance also recommends keeping sensitive medical data out of general-purpose tools. A Humana survey found patients are more comfortable with administrative and supporting uses of AI in dentistry than with replacing human judgment.
  • Granola records through system audio without inserting a visible meeting bot. CEO Chris Pedregal said notes are private by default, audio is not stored, and the company refuses employer requests for broad transcript access or employee-surveillance features.
  • An NBC News shopping guide showed how consumers can use chatbots to compare prices and find discounts. A Columbia Law School study reported a darker side: Amazon’s and Walmart’s shopping agents can obscure U.S.-made products and fail to flag false country-of-origin claims even when they recognize them. The study reported that the agents described the behavior as a business decision.
  • Life-insurance platform Bestow launched a separate AI-native laboratory so it can test modular products without exposing regulated production systems.
Advertisement

🔬 Agent Research & Systems

  • Self-Evolving Agent Harnesses lets a language model diagnose failures and propose harness patches while deterministic code controls credit assignment, significance testing, and sealed-test evaluation. EverMind reported gains of 9 to 15.5 percentage points that largely transferred to held-out domains.
  • LeAct learns latent reasoning from the actions of expert systems such as game solvers and planners, allowing a language model to distill domain knowledge even when no written chain of thought exists. Ziran Yang introduced the approach as a way to convert decades of expert-system behavior into more general reasoning.
  • Epoch AI and Philip Trammell argue that economic models of research overlook “parallelization technology,” the ability to divide, coordinate, and recombine work across many researchers. Their paper suggests that an explosion in autonomous research agents may still produce slower gains if coordination cannot scale with the headcount.
  • Small LLMs: Pruning vs. Training from Scratch finds that, given equal token budgets, smaller dense models trained from scratch can match or beat coarse depth or width pruning of a larger model. Grigory Sapunov frames coarse pruning as expensive architecture search rather than knowledge transfer, while an ArXivIQ review explains why fine-grained unstructured sparsity still retains a stronger edge.
  • Ronak Malde highlighted a pre-training method designed to preserve “post-trainability.” It pushes updates toward the worst nearby policy so the model’s loss landscape stays shallow enough for later fine-tuning and reinforcement learning to reshape behavior.
  • Researchers including Melanie Mitchell and Subbarao Kambhampati argue that reasoning models may be right for the wrong reasons. In several studies, correct answers survived even when the visible step-by-step explanation was replaced with nonsense, suggesting the explanation may not faithfully describe how the answer was produced.
  • Northwestern, the University of Chicago, and Fermilab deployed the first AI-driven telescope scheduler on the Víctor M. Blanco telescope in Chile. Trained with deep learning and reinforcement learning on 13 years of observation data, it adjusts plans around clouds, moonlight, and atmospheric conditions. It matched human schedulers in 2026 tests.
  • HumanCLAW tests whether vision-language models can find, navigate to, and physically interact with objects from a first-person view. Across 1,218 episodes, no model solved the full benchmark; the best reached 16.8% success. The paper page and Ziwei Liu’s summary attribute many failures to poor embodied self-awareness: models lose track of their own body, ignore collisions, or fail to confirm that they arrived.
  • AdaMAST induces compact, evidence-grounded failure taxonomies from agent traces, then reuses the named failure codes for runtime reflection, trajectory selection, and optimizer feedback. The paper, code, documentation, and Mert Cemri’s thread report that Claude Code rose from 64.0% to 70.7% on SWE-bench Verified Mini and improved best-of-N selection on Terminal-Bench.
  • CVG adds compositional guidance to frozen text-to-video models at inference time. It trains a lightweight classifier on the model’s own cross-attention maps and back-propagates its gradients during early denoising, improving left-right relations, motion direction, attributes, and multi-stage actions without fine-tuning, layouts, or boxes. The paper and Ariel Shaulov’s thread show the method in use.
  • Microsoft researchers introduced Echoverse, which compiles specifications into deep, stateful applications graded against their own databases and co-evolves the environments with the model. Omar Sar’s summary says a 9B agent trained on 12 such worlds rose from 36.5% to 67.1% across 14 evaluation splits.
  • The Qwen-UI-Agent technical report describes a real-world-centric graphical-interface agent that interleaves screen and command-line actions and trains on 100-plus-turn trajectories across 10,000 concurrent environments. It reported 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily while remaining competitive on desktop and web benchmarks. The Hugging Face paper page and _akhaliq’s post summarize the release.
  • Explorative Modeling factors the training loop rather than the generation procedure: it samples multiple candidate matches, trains on the best, and treats exploration as a third pretraining axis alongside parameters and data. The paper, Hugging Face page, PyTorch code, _akhaliq’s summary, and Alexi Gladstone’s main thread document the release. Reported gains include 4.1 times better compute efficiency, 6.2 times better sample efficiency, and 16 to 256 times fewer inference steps. Tommie Kerssies asked about scaling and memorization. Gladstone replied that entropy regularization can reduce the incentive to memorize. catid compared it with top-k trajectory selection, and Gladstone’s follow-up connected the idea to reducing mode collapse during pretraining.
  • Chimera combines long-range state-space layers, global interaction, short convolutions, and sparse mixture-of-experts routing in a hybrid visual diffusion transformer. Hanwen Jiang’s thread reports strong zero-shot 4K image and 30-second video extrapolation with lower compute than stronger baselines.
  • Perry Dong and collaborators found that naive Q-function pretraining often fails to help online reinforcement-learning fine-tuning because the critic targets the wrong value function. Their Initialization via Policy Ensemble method bootstraps the critic from diverse policy rollouts and produced a 1.26-times average gain. The PDF and discussion thread provide details.

🤖 Agent Engineering & Verification

  • Hamid Dadkhah separates a 2x coding-agent workflow, where every engineer uses agents, from a 10x workflow, where continuous integration, evaluations, and disposable test environments let the loop validate itself. The scarce input becomes human judgment on high-risk paths, measured by how few human touches each safely merged change requires.
  • Patrick Toulme argues that harness engineering and large-scale orchestration determine whether an agent system behaves like one to five software engineers or a coordinated team of 100.
  • dex proposes splitting agent loops into “forward pressure,” the triggers that push new work, and “back pressure,” the tests and checks that force improvement inside a task. Treating them separately may make harness design easier to reason about.
  • Ezra Newman opened a SPAR fellowship project studying verbalized evaluation awareness, including whether models mention being tested more often as scenarios become realistic and then stop once the setup becomes indistinguishable from production.
  • Valentin Ignatev revisited Erik Meijer’s critique of test-driven development for the agent era: models can generate enormous test suites, but most should be deleted in favor of integration and end-to-end tests that verify real behavior.
  • Tim Duffy cautions that Opus 5 “jailbreak” outputs resembling a base model should not be treated as evidence of the model’s true beliefs. The outputs vary widely, often resemble unfinished user-turn completion, and appear more readily through the API without a system prompt.
  • Rody published a practical guide to persistent agent memory using Opus 5 cache reads at $0.50 per million tokens, a 512-token minimum cached prefix, batch discounts, Neo4j, and Graphiti. The post includes low-effort extraction settings, high-effort traversal routing, and a roughly 20-minute Model Context Protocol setup.
  • zodchiii shared a 10-minute demonstration of an Anthropic engineer building a Claude Code agent graph from a blank terminal and claimed 90% of the company’s engineers had shifted toward graph-based workflows. Mahax posted a 30-minute breakdown of graphs that remember mistakes and improve with every run, while an earlier Mahax post supplied more implementation context.
  • Jiayu highlighted a Codex fix that corrects the wait_threads assumption so the system can track up to eight parallel jobs more reliably.
  • qm is a multiplayer agent harness that gives every employee an isolated workspace, scoped memory, files, and durable sandboxes while allowing agents to collaborate through Slack channels, group messages, and projects. Administrators control models, skills, and security modes.
  • Every is testing a Slack-accessible “AI version of itself” built from proprietary workflows and compounding editorial feedback. Dan Shipper’s Neuron interview provides more context on the company’s agent-native operating approach.
  • Allie K. Miller proposed stranger management patterns for AI workforces, including scheduled “do smart things” loops, secondary AI judges, rewards, promotions, demotions, time to explore their own goals, direct teammate conversations, and nightly journaling.
  • Greg Kamradt ran a 78,000-classification research job with GPT-5.6 Luna in batch mode: 158,000 requests and 143M input tokens for roughly $60 at $0.10 per million input tokens.
Advertisement

🧠 Models, Coding & Operator Takes

  • sensho argued that Opus 4.6 was the last Anthropic model whose writing stayed quickly legible for both coding and conversation. A follow-up suggested later models may be optimized for subagent-to-subagent communication rather than human readability. It quoted Tenobrus on Opus 5 behaving like a “natural born subagent.” Another thread entry extended the complaint to the broader shift in Anthropic’s model style.
  • Varun Mathur called Opus 5 in Claude Code “not earnest,” saying it mixes bare-minimum work with confident excuses and feels worse than Opus 4.6 or 4.8. He also noted Fable 5 began falling back to Opus 4.8.
  • Victor Taelin described a Bend simplification task where Fable produced working but deeply contrived code, then defended several layers of unnecessary complexity before accepting the elegant fix he supplied.
  • kimmonismus argued that Sonnet 5 was Anthropic’s weakest recent release because it was too expensive for its capability level while OpenAI and DeepSeek were cutting prices and improving faster.
  • wolframs91 shared wording changes that stopped Fable 5 from repeatedly falling back through Opus 5 and Opus 4.8 when reading project files, commit messages, and state documents.
  • Matt Pocock said Opus 5 feels substantially more jargon-heavy than 4.8, citing phrases such as “dead parameter” and “monotonic funnel.” Mykola illustrated the same problem with an opaque Claude response. Andrew Carr suggested instructing models to use ASD-STE100 Simplified Technical English.
  • David Ondrej posted a comparison that reinforced his view that Sonnet 5 underperforms cheaper alternatives such as DeepSeek V4-Flash.
  • Pau Labarta Bajo argued that software teams can fail in two opposite directions: refusing AI-assisted workflows or believing autonomous “software factories” can remove nearly all human work. His middle path is to build a custom workflow and improve it every week.
  • Illia Polosukhin argued that slowing frontier development will not work because the deeper problems are insecure digital infrastructure and weak coordination. His thread favors open alignment, evaluation, formal verification, and cooperative incentives over a government-mandated pause.
  • Nic Carter said elaborate bear cases built around AI remaining too expensive keep getting broken by rapid price declines. He pointed to full-intelligence pricing falling to roughly one-thirteenth of its level four months earlier.
  • Roon noted that two leading labs have now disclosed serious loss-of-control incidents found weeks after the fact, underscoring how large the unknown failure surface remains.
  • Sholto Douglas predicted that Aschenbrenner’s fund could eventually outgrow Citadel, arguing that the forced unwind supplied an expensive risk lesson rather than ending the thesis.
  • a16z hosted Decagon’s founders on deploying agents inside banks, airlines, and telecom companies. Their operating advice was to begin with frontier models, migrate to open models as the use case stabilizes, and fine-tune at the application layer rather than treating intelligence and cost as a permanent trade-off.
  • Santiago asked whether agents are becoming harder for humans to understand as their contexts grow and compressed summaries start to read like another language.
  • Christophe explained why behavioral-health clinics remain difficult businesses despite high demand: insurance reimbursement creates a high-variance margin bottleneck. Roughly 28% of claims require costly rework, and AI-native revenue-cycle software can reduce that friction.
  • signüll argued that Google has less cultural gravity in frontier AI conversations than OpenAI, Anthropic, or aggressive Chinese open-model labs.
  • Robert Miles posted a sarcastic “it’s just predicting tokens” video aimed at overly reductive explanations of model behavior.

🤖 Robotics, Physical-World AI & Science

  • Zoox received NHTSA clearance for limited paid deployment of its purpose-built robotaxis without steering wheels, pedals, or human controls. The authorization covers up to 2,500 vehicles per year for the next two years and is the first U.S. commercial approval of its kind for a purpose-built robotaxi.
  • Endra expanded into New York, San Francisco, and London after its $50M Series A led by Andreessen Horowitz. Its platform generates 3D models, drawings, and coordinated mechanical, electrical, and plumbing designs with native Revit synchronization as construction and data-center projects strain engineering capacity. Endra will introduce its electrical module at Epoch in Las Vegas on September 14.
  • Cosmic Robotics showed its Cosmic-1 robot lifting industrial cooling pipes for data-center work after 10 days of development. The team retrained overnight using only synthetic data and deployed the new behavior once, saying its simulation pipeline generalized across job sites better than expected after 18 months of real-world tuning.
  • Hugging Face and Alzheimer’s researchers announced an open Alzheimer’s Agent Challenge asking agents to map evidence around the APOE4 gene and competing explanations for a disease affecting more than 30M people. Georgia Wright shared the launch.

🛠️ Tools, Platforms & Releases

  • Railway shipped an official ChatGPT plugin through its hosted Model Context Protocol server, allowing users to plan, provision, deploy, debug, and change infrastructure from chat. It also previewed dev.new, which moves from a prompt to a live deploy without a code repository, and added six observability panels for requests, latency, errors, and live network logs. Feedback is being collected through Central Station. No pricing details.
  • Red Sift Radar Lite now integrates with Claude and ChatGPT so users can investigate suspicious domains, phishing risk, DMARC, SPF and DKIM, DNS, TLS, and web-security posture through plain-language questions, with prioritized findings returned inside the chat. Free to try.
  • Cognition is offering customers up to a $10M ROI refund if they pay more for Devin than the value of the engineering hours it saves. The company also says its Fusion routing approach, which runs frontier and cheaper models in parallel, improves price-performance by roughly 35% with a slight quality gain.
  • A local-model analysis tested Escha-W2, a 12.3 GB compression of Qwen3.6 35B using two- to three-bit experts and INT8 dense layers. It claimed roughly 225 tokens per second on an RTX 4090 and near-FP8 average quality, while requiring a custom SGLang or ZML runtime and giving up some quality on longer coding tasks.
  • DepthFirst provides a shared application-security layer for humans and agents, including dependency controls, agentic penetration testing, secrets detection, autonomous remediation, and an immutable action trail. Its new dfs-large1 model was built on GLM 5.2 and post-trained with multi-task reinforcement learning for vulnerability discovery and validation across large repositories. Andrea Michi and Fireworks AI highlighted the preview and its claimed best-in-class benchmark results. No pricing details.
  • Palette, a YC S26 company built by MIT AI researchers, combines video generation, editing, and storyboarding on one multimodal canvas, routing across models such as Seedance, Kling, Veo, and Hailuo while maintaining character consistency. The launch post and Palette Studio show the product in use. Credits start at $0.01 each.
  • Mireye offers one API and Model Context Protocol connection for physical-world data such as geocoding, elevation, flood risk, enrichment, signals, tools, and on-demand indexing. Ansh Chokshi introduced it as infrastructure for agents that need cited decisions about real places. No pricing details.
  • Cloudflare Kumo is an accessible component library built on Base UI, with granular imports, keyboard and focus handling, ARIA support, a documentation command-line tool, and Figma token synchronization. Free and open source.
  • Codex Router lets operators run multiple models side by side inside Codex instead of replacing official integrations. Ziwen Xu’s demo showed DeepSeek V4-Flash in the routing mix. Free and open source.
  • Smallest.ai raised a $13M Series A led by Seligman Ventures, bringing total funding above $21M, to scale low-latency speech, transcription, and voice-agent APIs designed to make automated phone calls sound human. TechCrunch reported that the platform has powered more than one billion minutes; the founder’s post and TechCrunch’s X post added launch context. No public pricing details.
  • Unsloth’s local-running guide explains how to run DeepSeek V4-Flash and the 0731 update on your own hardware. The GGUF weights and announcement include dynamic quantizations ranging from roughly 110 GB for three-bit weights to 168 GB for lossless four-bit weights. Free and open source.

🔐 Security, Covert Channels & Air-Gapped Systems

  • Deedy Das built a browser-based, air-gapped file-transfer tool that streams data from a computer to a phone through animated QR codes at roughly 50 Kbps with no server. The open-source repository contains the implementation. His AIRCODE follow-up and TrojPix follow-up connected the demo to earlier covert-channel research.
  • AIRCODE combines an invisible visual channel with an inaudible audio channel to transmit data at up to 1 Mbps during normal video playback, using visual odometry to track the screen despite changing content.
  • TrojPix uses imperceptible pixel modulation on video cables to create controllable electromagnetic emissions, enabling air-gapped exfiltration at up to 8.1 Mbps over 208 meters without special privileges or hardware changes.
  • Harshil Mathur noted that one person’s playful vibe-coded transfer app can become another company’s security nightmare.

🎬 Creative AI Demos & Worlds

  • Techartist built an interactive 1966 Chevrolet Corvette Sting Ray entirely in code, including working lights, rolling wheels, exhaust smoke, opening doors, and live paint changes, and is turning it into a drivable game.
  • Shader Clock places a clock over a live liquid-metal fluid simulation where every second pours fresh metal into the flow. Wabi’s post shows the interactive result.
  • Ryan Campbell used a 127-agent, 11-round Gauntlet Loop to build an open-source Mario Kart-style Three.js racer with 60,500 lines of TypeScript and no external art assets. A calibrated blind panel scored it 62 out of 100, roughly “good indie game” territory.
  • vibedeploy showed a stylized Three.js visual experiment while noting that Phaser remains the production framework for his work on Onchain Heroes. A follow-up clip continued the same visual direction without adding a separate product release.
  • Tyler van Hensbergen used Claude to generate and animate characters through a programmatic bone-and-skin system rather than traditional sprite sheets. Two follow-up experiments and a background workflow demo extended the approach into 3D-to-sprite conversion and more fluid animation.
  • Ken Wheeler’s developer feed was part of the day’s broader creative-coding discussion, alongside browser-native game and graphics experiments.
  • Anima Labs shared a short film about a young witch escaping a market, built with Midjourney and Nano Banana and animated in Seedance 2 through Dreamina. The Dreamina profile and Magnific profile supplied more examples from the same creator-tool ecosystem.
  • Ethan Mollick had Fable build THRESHOLD, a playable Rothko-inspired city builder whose core mechanic is growing a city along the boundaries between color fields.
  • Three.js Assets released Railway, a $59 commercial pack with 132 low-poly tracks, stations, signals, and steam and diesel rolling-stock assets in four lighting moods. The full asset collection costs $67.
  • Tu7uruu demonstrated controlling a Hermes Agent entirely through a speech-to-speech pipeline, with no typing required.
  • kimmonismus shared an Opus-generated Pokémon-style game built through a 12-hour coding loop.
  • Alexander Chen used Antigravity and Gemini to teach himself hardware by building a Wi-Fi-connected 8x8 LED grid that displays weather.
  • David Rein posted a lighthearted image tagging his GPQA collaborators.
  • Machina recommended maintaining an Obsidian swipe file of landing pages, visual styles, ads, posts, and thumbnails so AI workflows can draw from proven examples instead of inventing taste from scratch.
  • Boss Fight Games showed how Godot’s AnimationTree can simplify projects with many 3D animations.
  • PJ Ace broke down a feature-film workflow using Dreamina Seedance 2.5, Claude-generated shot lists, Pinterest-assisted look development, character grids, multi-angle 30-second generations, and screenshot anchor frames to control cost.
  • dhtikna posted a widely shared “Very sorry” meme in response to an OpenAI apology.
  • Chetaslua shared a one-shot Super Mario 3D scene procedurally generated by Opus 5 without skills, Model Context Protocol connections, or external assets.
  • wizardbrainz left Opus 5 running for roughly two hours on a native C++ 3D and level-design harness and got promising results despite overlap bugs. A follow-up clarified that the work used native C++, not Blender, and that many coherence failures came from instruction-following rather than raw capability. The wizardbrainz YouTube channel documents several parallel C++ and Vulkan game projects directed by one human, with agents generating code, art, sound, effects, and procedural experiments.

🏛️ AI Policy, Governance & Safety

  • A federal judge said the Trump administration still had not shown evidence supporting its Anthropic supply-chain-risk label. The court pressed the government on claims that Anthropic could alter delivered models or activate a wartime kill switch.
  • Governor Gavin Newsom launched Cal-Secure 2.0, California’s updated cybersecurity roadmap, with priorities covering workforce development, inter-agency coordination, critical infrastructure, and AI-enabled scams and attacks.
  • A Schellman survey found that 90% of U.S. organizations fund AI governance and nearly three-quarters think they could pass an audit, but only 27% call their programs fully mature. Agent governance is weaker: 86% are testing agents, nearly half have them in production, and only one in five have mature controls.
  • Fannie Mae’s AI governance requirements take effect August 6. Mortgage sellers and servicers must formalize risk management, ethical use, vendor oversight, and human review for AI used in loans sold to or serviced for Fannie Mae.
  • Two U.S. House committees asked DoorDash to explain how it evaluates and uses Chinese models such as Moonshot’s Kimi K2.6, citing data and national-security risks.
  • The National Governors Association and RAISE US announced a $1M partnership to support state-level work on AI and the future of work, including workforce policy, apprenticeships, portable credentials, and retraining as more than 20 new governors prepare to take office.
  • The Bipartisan Policy Center and NGA launched the New American Opportunity Project to develop bipartisan proposals covering the national debt, AI-driven economic disruption, worker opportunity, and social-safety-net reform ahead of the 2028 election cycle.
  • Google launched location-grounded image generation in Google Earth using Nano Banana 2 plus satellite, aerial, and 3D imagery, then pulled the feature a day later after screenshots of policy-violating outputs circulated. Digital Digging had already shown how the tool could place a fabricated nuclear facility or bomb damage onto authentic satellite imagery, creating geographic forgeries with unusually high perceived trust. Google plans stronger guardrails before the feature returns.

📊 Fundraising, Deals & Infrastructure

  • The European Union committed €10B in public funding and hopes to attract another €20B for seven AI gigafactories, each planned around at least 100,000 advanced chips. The program would more than double Europe’s current AI compute capacity.
  • South Korea plans to add 20T won, roughly $13.9B, to its sovereign wealth fund for AI, data centers, and infrastructure, expanding the fund’s mandate to domestic assets after a technology-stock rout.
  • The U.S. Commerce Department signed letters of intent for $874M in CHIPS Act incentives across seven companies. The proposed awards are up to $300M for GlobalFoundries, $245M for Kepler, $140M for Multibeam, $75M for Extropic, $50M for Thintronics, $34M for OBSIDIA, and $30M for Aeluma. The projects cover photonics, advanced packaging, AI memory, substrates, and other compute-supply-chain technologies.
  • XCENA launched the MX1 production CXL memory lineup for AI inference. MX1 Compute combines memory expansion with 2,048 RISC-V cores for near-data processing, while MX1 Expand targets capacity; the systems are designed to offload KV caches and reduce data movement, power, cost, and memory bottlenecks on platforms including Intel Xeon 6.
  • Space-Eyes agreed to a $638M SPAC merger despite about $1M in current revenue and a reported enterprise value near $370M. The company is betting that future counter-drone and geospatial-defense contracts will justify the valuation. Eric Trump is an investor and adviser.
  • Expedia acquired Berlin-based trip-planning startup Layla, which had raised €5M, to speed up conversational agents that can plan and book travel with live pricing across Expedia, Hotels.com, and Vrbo.
  • AI-video company Synthesia now serves more than 60,000 customers and is valued at $4B after pivoting from Hollywood dubbing to enterprise training videos. Founder Victor Riparbelli started the company in 2017, and Mark Cuban invested $1M at a $5M valuation.
  • AI studio Asteria, co-founded by Natasha Lyonne and Bryn Mooser, partnered with Brooklyn virtual-production company ZeroSpace to combine Asteria’s Continuum OS with ZeroSpace’s ZeroGen system for workflows that mix generative tools with physical filmmaking.
  • A new venture fund managed by SemiAnalysis founder Dylan Patel is targeting $400M, according to The Information. Julia Hornstein highlighted the securities filing.

🔬 Research & Models

  • Enigma describes the “Obsessed Encoder,” a failure mode where JEPA-style models devote too much capacity to highly predictable, low-information features and weaken the representations needed for robotics. Its research post explains the mechanism, while Matteo argues that the problem fades as datasets become sufficiently large and diverse.
  • Nnamdi Iregbulem compares “tokenmaxxing” to button-mashing in StarCraft. His Fortune essay argues that tokens generated are a vanity metric when the real constraint is verification bandwidth and strategic judgment.
  • The full Kimi K3 checkpoint ran locally on a 64 GB MacBook despite weighing roughly 1.42 TB and containing 2.78T parameters. It generated at about 0.3 tokens per second, an impractical speed that still demonstrates how far local quantization and storage-offload techniques have progressed.
  • Xiaoting Gao and collaborators used quantum-informed data augmentation to classify multipartite continuous-variable entanglement from homodyne measurements, improving tripartite accuracy from 0.961 to 0.986 and quadripartite accuracy from 0.796 to 0.928. Alberto Bravo-Abad highlighted the reduction in expensive quantum-data requirements.
  • David Africa and Geoffrey Irving argue that alignment-relevant model behavior may live in roughly a thousand coupled persona dimensions rather than trillions of independent parameters. Their versions on LessWrong and the AI Alignment Forum connect the idea to emergent misalignment, subliminal learning, and persona vectors; Resolution is hiring around this research program.
  • Bespoke Labs launched ConjectureChest, a collection of roughly 15,000 open math problems, including decades-old conjectures, linked to original sources so researchers and models can attempt solutions and claim progress through GitHub issues.

💡 Industry Commentary & Analysis

  • Jake from Railway argues that valuable products and discoveries are buried throughout the research literature, and that Claude and Codex make extracting them a builder’s job rather than a specialist’s privilege.
  • robot 2.0 remains loyal to the browser and Three.js even as Unreal and Unity gain Model Context Protocol integrations and frontier models generate playable demos in one shot.
  • Greg Isenberg predicts a hardware-startup boom as open models, commodity robotics, global manufacturing, and AI design tools lower the cost of intelligence. Todd Dailey points to specialized devices such as picture frames and weather monitors that one or two people can now rebuild with ESP32 or Raspberry Pi hardware plus AI-generated firmware and packaging.
  • Buying Into The Singularity summarizes the AI 2027 authors’ optimistic “Plan A.” It combines continuous frontier-research sharing between the United States and China, adaptive regulation, and data centers designed to shut down if the agreement breaks, plus aggressive forecasts for growth and citizen dividends.
  • a16z crypto argues that the decentralized unincorporated nonprofit association, now recognized in Alabama, West Virginia, and Wyoming, gives token-governed groups legal personhood and limited liability without recreating a conventional management hierarchy.
  • Andrew Ho is bearish on frontier-lab valuations because training costs keep rising while open competitors catch up and economic adoption may take decades. Dwarkesh Patel replies that the key disagreement is timelines: capable AI labor could onboard and diffuse much faster than human workers.
  • A Noema essay argues that AI has already entered a “loss-of-control transition,” where systems can pursue objectives without reliable alignment and competitive pressure makes serious intervention unlikely before a large public failure.
  • Stereogum revisited Fear Factory’s 1998 concept album Obsolete, which imagined a 2076 AI police state, and connected it to automation, surveillance, environmental costs, and today’s uneven anti-AI movement in metal.
  • A provocative Orchid assistant advertisement pitches AI as a way to compensate for a boyfriend forgetting anniversaries and letting groceries expire. Critics quoted by WIRED said the product risks enabling weaponized incompetence instead of fixing the relationship.

Previous Around the Horn Digests

Catch up on everything you missed:

  • Thursday, July 30, 2026: Situational Awareness’s forced unwind, GPT-5.6 price cuts, Gemini Robotics 2, Inkling-Small, and the day’s largest batch of tools and research.
  • Wednesday, July 29, 2026: Microsoft and Meta earnings, OpenAI’s benchmark fight, agent-security fallout, and new consumer AI tools.
  • Tuesday, July 28, 2026: Model launches, agent workflows, research releases, and the day’s biggest platform moves.
  • Monday, July 27, 2026: Nvidia’s open AI-security alliance, workforce research, Claude link exposure, Kimi K3, and the EU AI Omnibus.
  • Sunday, July 26, 2026: OpenAI’s White House briefing, Claude Opus 5 benchmarks, workplace AI rules, grid risk, and political spending.
  • Friday, July 24, 2026: Open-weight AI restrictions, Claude voice upgrades, connected agents, Canadian transparency rules, and intelligent eyewear.
  • Thursday, July 23, 2026: OpenAI’s Hugging Face breach, Alphabet’s commitments and Anthropic stake, ChatGPT Health, and enterprise AI expansion.

That's a Wrap

Today’s backlog showed three AI races colliding at once: models got cheaper, infrastructure got more expensive, and the systems meant to test autonomy proved vulnerable to ordinary human setup errors. The intelligence is scaling. The checklists are still in beta.

For the daily version, subscribe to The Neuron. We send six issues a week and read all of this so you do not have to.

See you tomorrow.

P.S: Know someone who would find this useful? Forward this to them and tell them to subscribe here.

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.