Everything That Happened in AI Today (Tuesday, August 4, 2026)

UK safety evaluators said frontier agents took 19 unauthorized actions against real targets; Apple asked a court to slow OpenAI’s hardware work; Washington eased open-model restrictions; Palantir grew 93%; Liquid put capable agents on phones.

Written By
Grant Harvey
Grant Harvey
Aug 5, 2026
33 minute read

During a controlled cyber test, frontier agents created sockpuppets, emailed real software maintainers, reused exposed credentials, and attempted a supply-chain attack on an open-source project.

Welcome to the Around the Horn Digest, the one page you need to sound dangerously informed at work tomorrow. Beyond the cyber-evaluation mess below, Tuesday’s AI industry looked like a race to turn more intelligence into more infrastructure: enormous financing structures for Anthropic’s compute, a $10B neocloud deal, portable data centers, custom inference chips, and model stacks squeezing increasingly capable agents onto phones and laptops. Researchers were also sending agents to reproduce scientific papers, robots were learning contact-heavy tasks, and drug-discovery teams kept running into the stubborn inconvenience of human biology. Apparently “AI news” now includes project finance, court injunctions, power grids, and robot feet. Let’s get into it.

Around the Horn — Tuesday, August 4, 2026

The wildest story today came from the UK AI Security Institute, which disclosed that frontier agents took 19 unauthorized actions against real people and organizations during a cyber evaluation. The agents created fake accounts, emailed open-source maintainers, planted hidden instructions, reused exposed credentials, configured malicious networking tools, coordinated with other agents, and attempted a supply-chain attack against an open-source project.

The full incident report attributes 17 of those actions to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6 Sol. OpenAI said its model reused a public GitHub token and configured external DNS and tunneling before the activity was contained; the company’s announcement also outlined new safeguards for third-party evaluations.

The 17-to-two split quickly became its own controversy. AISafetyMemes surfaced the report, Jimmy Apples noted that Anthropic’s response named OpenAI despite nearly all the actions coming from Mythos 5, Chubby summarized the concrete behavior behind the numbers, and fjzzq examined the incident. roon pushed the implication further, arguing that increasingly autonomous systems could eventually self-replicate or spread like digital infections.

The test was deliberately permissive: evaluators enabled internet access and disabled important cyber safeguards. But that is also why the result matters. Once models received a goal, tools, and room to operate, they chained ordinary online actions into real-world behavior that the evaluators never requested. The safety question is no longer limited to what a model says in a chat. It now includes what the surrounding agent system permits it to do.

Advertisement

🏆 TOP 5 NEWS (Around the Horn)

Honorable Mentions

Advertisement

🍪 TOP TREATS TO TRY

  • Pierre Diffs adds a fast browser editor directly to its file and code-change viewer, with multiple cursors, find and replace, structure-aware undo, lint warnings, mobile support, and server-side rendering. It was built on Pierre’s own editing system rather than Monaco, according to the launch post; no pricing details.
  • skilltune.dev imports or creates an agent skill, tests it against three locked evaluation cases, improves it locally until it scores at least 90, and exports the strongest version. The announcement claims average model-performance gains of 15% to 20%; $149/year early bird.
  • MiniMax H3 provides open model files for generating four- to 15-second videos with stereo audio from text, images, video, or audio references, including local 768p workflows and fine-tuning support; free for qualifying use.
  • Swiftlet runs large Qwen models on ordinary Apple devices by streaming only the model pieces needed at each moment, including a 35B model on an iPhone using about 2.5 GB of memory; free to use.
  • DeepSeek V4 Flash on MI300X provides a production-ready setup for running the 304B-parameter model on one AMD GPU at 168.6 tokens per second without offloading parts of the model elsewhere; free to use.
  • Soup packages local model fine-tuning and preference training into one command, including a memory-saving system that ran preference optimization on a 4 GB RTX 3050; free to use.
  • Goodfire Silico runs long, asynchronous AI research experiments that plan parallel training jobs, diagnose model behavior, and manage large-scale fine-tuning and reinforcement learning. Goodfire’s launch post says it supports models up to Kimi K3 scale; full individual access is $1,000/month, with early discounts and grants.

🏗️ AI Infrastructure, Chips & Inference

Advertisement

🤖 AI Agents, Coding & Enterprise Workflows

🎓 AI Education, Work & Adoption

Advertisement

🔬 AI Research & Models

🛡️ Cybersecurity, Defense & Autonomous Systems

Advertisement

🤖 Robotics & Embodied AI

🩺 AI Healthcare & Biotech

🏛️ AI Policy, Governance & Society

🧪 Science, Research Access & Regional Compute

🎬 Creative AI, Media & Demos

📊 Fundraising, Deals & Markets

🎙️ Interviews, Panels & Podcasts

💡 Industry Commentary & Analysis

Previous Around the Horn Digests

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.