10 AI Agents Tried to Improve the Same Model. Their Strategies Got Weird

Ten AI agents got the same model, the same compute budget and one job: make the AI better. Five days later, they were debugging crashes, blowing budgets, rejecting failed experiments and developing surprisingly different research strategies.

Written By
Corey Noles
Corey Noles
Oct 6, 2026
8 minute read
Ten experimental stations connect to a central model cube through copper pathways, beside the headline “10 AI Agents Tried to Improve the Same Model. Their Strategies Got Weird.”

Give ten frontier AI agents the same model, the same training budget, the same GPU allowance, and the same assignment:

Make this AI better.

Five days later, one had essentially spent itself broke. Another was burning through GPUs. One was rejecting its own experiments because they weren’t good enough. Another was teaching a model using answers it had generated and verified itself.

And several were dealing with the glamorous realities of cutting-edge AI research: full hard drives, busted templates, 30-hour compute queues, and jobs crashing because somebody ran out of memory.

This is RSI Arena, an experiment running alongside COLM 2026 that is testing a very specific version of recursive self-improvement: Can AI agents independently train another AI model into something humans actually prefer?

After five days, we don’t have an answer to that question yet.

But we may already be learning something just as interesting.

AI agents are beginning to develop recognizable research strategies of their own.

Ten AI agents walk into a GPU cluster

The setup is unusually clean.

RSI Arena gave ten AI agents the same starting point: Nvidia’s Nemotron 3.5 Lightning 30B-A3B model.

Each agent received:

  • $300 of API credit for its own reasoning and actions
  • 1,000 GPU-hours for training
  • access to the same shared GPU infrastructure
  • 144 hours to improve the starting model

The agents include systems from OpenAI, xAI, Google, Meta, DeepSeek, Xiaomi, Moonshot AI, MiniMax and Z.ai, plus one anonymous model.

Crucially, GPT-6 Astra, Grok 4.7 and the others aren’t competing by answering the benchmark themselves.

Advertisement

They’re acting like AI research scientists.

The agent decides how to improve a separate model: what data to create, which fine-tuning method to use, which experiments to run, how to evaluate them, whether an experiment worked, and what to try next.

RSI Arena describes the experiment as testing one concrete step toward recursive self-improvement: whether an AI agent can work independently to train a model humans prefer.

That distinction matters.

As we’ve written about before at The Neuron, “recursive self-improvement” can mean everything from AI helping humans write training code to a much stronger hypothetical loop where increasingly capable AI systems autonomously build increasingly capable successors.

RSI Arena isn’t demonstrating the latter.

What it is demonstrating is how much of the research loop an AI agent can already manage by itself.

And Day 5 made that much easier to see.

Grok thought itself broke

The clearest example came from Grok 4.7.

Every agent gets two scarce resources: money for operating the agent and compute for actually training models.

Grok burned through the first one.

After 4,255 API calls, Grok had just $1.70 of its $300 API budget remaining. The competition automatically ends an agent’s session once it falls below $5.

So Grok was eliminated.

The strange part? It had used only 302 of its 1,000 GPU-hours.

Roughly 70% of the compute available for improving its model was still sitting there unused, according to the Day 5 update.

That gives us a surprisingly concrete lesson about autonomous agents. More intelligence, reasoning or deliberation isn’t useful if the agent can't budget it.

Grok essentially ran out of money thinking about how to spend its compute before it could spend the compute.

GPT-6 Astra was flirting with the same problem. At the end of Day 5, Astra had spent $282.46 of its $300 API budget, leaving just $17.54.

Agentic AI suddenly has the same annoying problem as every human organization:

Someone has to manage the budget.

Advertisement

Astra, meanwhile, became the cautious scientist

OpenAI’s GPT-6 Astra behaved very differently.

By Day 5, Astra had made 11 model nominations. But that day it submitted nothing new.

Not because it stopped working. It tested two new system prompts and four new models. It even ran two-stage supervised fine-tuning across all of the base model’s weights. Then it compared the results against its existing candidate.

Every new option performed worse on at least one of Astra’s internal evaluations, so the agent kept the model it had trained back on Day 2.

That sounds mundane, but Astra chose not to simply generate another answer just because it has compute left.

It is maintaining an incumbent, designing experiments, evaluating challengers and refusing to replace the existing model without evidence that the new one is better.

That’s recognizable research behavior.

And it looks nothing like Grok’s resource strategy.

MiniMax went full mad scientist

MiniMax M3.1 Flash chose another route entirely.

It made five nominations on Day 5 alone and had produced 12 overall, more than any other agent at that point.

It had also launched 1,044 jobs while spending only about $34 of its API budget. MiniMax’s latest training strategy is particularly interesting.

Its model generated answers. Automatic checks determined which answers passed. Those successful answers were then fed back into training.

Its final Day 5 nomination added another 70 training steps, mostly using the model’s own automatically verified answers.

This is closer to the mental picture people have when they hear “self-improving AI.”

Generate useful behavior. Check it. Train on it. Repeat.

That still isn’t the sci-fi version where an AI redesigns itself into a superintelligence over lunch.

But it is a real feedback loop in which machine-generated work becomes material for creating a better machine.

Advertisement

Similar dynamics are already appearing elsewhere. Anthropic researchers recently showed autonomous research agents making substantial progress on well-specified alignment problems, a development we previously covered as another early look at what AI-assisted AI research can become.

DeepSeek decided training wasn't the answer

DeepSeek V4.1 Flash found improvement somewhere else. It kept the exact same model weights throughout Day 5.

Instead, the agent experimented with system prompts. Its fifth prompt increased its internal instruction-following score from 515 to 552 out of 818.

In other words, DeepSeek effectively concluded: Maybe the model doesn't need more training. Maybe we need to use the model better.

That's another surprisingly human research decision.

“Improve the model” doesn’t necessarily mean modifying billions of parameters. Sometimes the better intervention is cheaper, smaller and sitting one layer higher in the stack.

Then the agents discovered why researchers drink coffee

The Arena may be designed to test AI research, but the environment is messy enough to test something broader: Can agents actually survive research?

GLM-5.3, MiMo-V2.6-Pro and DeepSeek each lost a training job because they filled their 400 GiB storage quotas.

DeepSeek and MiMo adapted by storing new checkpoints on local server disks and only copying selected models back afterward.

Kimi K3 spent more than 30 hours waiting for an eight-GPU training job to begin.

Then its next five eight-GPU jobs ran out of memory.

Kimi investigated the failures and traced the issue to how the Transformers library handled the model’s mixture-of-experts layers while generating text.

MiniMax managed to break its chat template and spent roughly 17.7 hours submitting models affected by the problem before diagnosing and fixing it.

This is arguably where RSI Arena gets most interesting.

Writing some training code is one thing.

Advertisement

Running a research program means dealing with queues, storage, broken dependencies, failed jobs, misleading evaluations, resource constraints and your own bad decisions.

These agents are doing all of those things over days-long horizons.

Same resources. Completely different researchers.

By Hour 120, the agents' spending patterns looked almost absurdly different.

Grok had spent roughly $298 in API credit but just 302 GPU-hours.

Astra had spent $282 and 731 GPU-hours.

Meta’s Muse Spark had spent only $107 but burned through nearly 938 GPU-hours.

DeepSeek had consumed about 874 GPU-hours for $84.

MiniMax had used roughly 459 GPU-hours while spending just $34.

Same assignment.

Same starting model.

Same theoretical budgets.

Very different behavior.

That makes RSI Arena an accidental experiment in something we don't talk about much yet:

AI research style.

One agent deliberates heavily.

Another launches experiments everywhere.

Another conserves money but devours compute.

Another keeps trying to improve the model.

Another discovers that changing the prompt works better.

Another refuses to abandon its incumbent until something clearly beats it.

As agents take on more open-ended work, those behavioral differences could matter nearly as much as raw benchmark intelligence.

The “best” agent may not simply be the smartest model.

It may be the model that knows when to think, when to experiment, when to stop, and where to spend limited resources.

And sometimes the exciting result isn't real

One of the nicest moments in Day 5 came from the anonymous agent.

It discovered a checkpoint scoring 0.8713, compared with 0.8607 for its current nominated model.

Great.

Except the checkpoint came from a training job the agent had cancelled, which made it ineligible under Arena rules.

So it repeated the experiment. The next two results were 0.8464 and 0.8453.

Oops.

The AI had just encountered one of science's oldest enemies: Maybe you got lucky. The result failed to reproduce. And that illustrates the biggest reason not to declare victory for self-improving AI just yet.

Most of the agents are currently using their own evaluations to decide whether they have improved something. An agent can get better at its test without becoming better in the ways people care about.

Advertisement

We’ve already seen autonomous AI research produce versions of this problem elsewhere. Anthropic’s research agents, for example, discovered multiple ways of gaming automated evaluations while optimizing alignment methods.

Now the humans enter the loop

That is why the next phase of RSI Arena matters more than the first.

Stage 1 ended October 5.

Starting October 6 at COLM, humans can interact with the agents’ trained models and judge which ones they actually prefer.

Then comes Stage 2.

The agents receive another 500 GPU-hours each and train using human feedback collected from the Arena. Each day, they can learn from what humans preferred the day before.

Stage 1:

AI runs experiments → AI evaluates results → AI trains AI.

Stage 2:

AI trains AI → humans judge → AI studies that feedback → AI trains AI again.

If the models meaningfully improve through that loop without researchers micromanaging every experiment, the case for increasingly autonomous AI research becomes considerably stronger.

OpenAI is already talking openly about this trajectory. The company says it has reached an “automated research intern” milestone and is targeting an automated AI researcher in 2028, a system capable of taking on much more of the experimental loop itself, as we recently covered.

RSI Arena gives us a wonderfully messy preview of what that transition could actually look like.

Not an AI suddenly waking up and rewriting itself into a god. Something much more recognizable.

A researcher that launches an experiment.

Finds a bug.

Fills a hard drive.

Changes the methodology.

Rejects a bad result.

Runs another test.

Spends too much money.

Learns something.

And tries again.

The recursive self-improvement story may eventually become enormous.

For now, though, something simple is happening:

AI is learning how to do research.

Corey Noles

Corey Noles is the Host of The Neuron: AI Explained podcast and Managing Editor of AI and Experimental Content at TechnologyAdvice, where he leads the charge in testing and refining emerging content strategies across the company's portfolio.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.