Give ten frontier AI agents the same model, the same training budget, the same GPU allowance, and the same assignment:
Make this AI better.
Five days later, one had essentially spent itself broke. Another was burning through GPUs. One was rejecting its own experiments because they weren’t good enough. Another was teaching a model using answers it had generated and verified itself.
And several were dealing with the glamorous realities of cutting-edge AI research: full hard drives, busted templates, 30-hour compute queues, and jobs crashing because somebody ran out of memory.
This is RSI Arena, an experiment running alongside COLM 2026 that is testing a very specific version of recursive self-improvement: Can AI agents independently train another AI model into something humans actually prefer?
After five days, we don’t have an answer to that question yet.
But we may already be learning something just as interesting.
AI agents are beginning to develop recognizable research strategies of their own.
- Ten AI agents walk into a GPU cluster
- Grok thought itself broke
- Astra, meanwhile, became the cautious scientist
- MiniMax went full mad scientist
- DeepSeek decided training wasn't the answer
- Then the agents discovered why researchers drink coffee
- Same resources. Completely different researchers.
- And sometimes the exciting result isn't real
- Now the humans enter the loop
Ten AI agents walk into a GPU cluster
The setup is unusually clean.
RSI Arena gave ten AI agents the same starting point: Nvidia’s Nemotron 3.5 Lightning 30B-A3B model.
Each agent received:
- $300 of API credit for its own reasoning and actions
- 1,000 GPU-hours for training
- access to the same shared GPU infrastructure
- 144 hours to improve the starting model
The agents include systems from OpenAI, xAI, Google, Meta, DeepSeek, Xiaomi, Moonshot AI, MiniMax and Z.ai, plus one anonymous model.
Crucially, GPT-6 Astra, Grok 4.7 and the others aren’t competing by answering the benchmark themselves.
They’re acting like AI research scientists.
The agent decides how to improve a separate model: what data to create, which fine-tuning method to use, which experiments to run, how to evaluate them, whether an experiment worked, and what to try next.
RSI Arena describes the experiment as testing one concrete step toward recursive self-improvement: whether an AI agent can work independently to train a model humans prefer.
That distinction matters.
As we’ve written about before at The Neuron, “recursive self-improvement” can mean everything from AI helping humans write training code to a much stronger hypothetical loop where increasingly capable AI systems autonomously build increasingly capable successors.
RSI Arena isn’t demonstrating the latter.
What it is demonstrating is how much of the research loop an AI agent can already manage by itself.
And Day 5 made that much easier to see.
Grok thought itself broke
The clearest example came from Grok 4.7.
Every agent gets two scarce resources: money for operating the agent and compute for actually training models.
Grok burned through the first one.
After 4,255 API calls, Grok had just $1.70 of its $300 API budget remaining. The competition automatically ends an agent’s session once it falls below $5.
So Grok was eliminated.
The strange part? It had used only 302 of its 1,000 GPU-hours.
Roughly 70% of the compute available for improving its model was still sitting there unused, according to the Day 5 update.
That gives us a surprisingly concrete lesson about autonomous agents. More intelligence, reasoning or deliberation isn’t useful if the agent can't budget it.
Grok essentially ran out of money thinking about how to spend its compute before it could spend the compute.
GPT-6 Astra was flirting with the same problem. At the end of Day 5, Astra had spent $282.46 of its $300 API budget, leaving just $17.54.
Agentic AI suddenly has the same annoying problem as every human organization:
Someone has to manage the budget.
Astra, meanwhile, became the cautious scientist
OpenAI’s GPT-6 Astra behaved very differently.
By Day 5, Astra had made 11 model nominations. But that day it submitted nothing new.
Not because it stopped working. It tested two new system prompts and four new models. It even ran two-stage supervised fine-tuning across all of the base model’s weights. Then it compared the results against its existing candidate.
Every new option performed worse on at least one of Astra’s internal evaluations, so the agent kept the model it had trained back on Day 2.
That sounds mundane, but Astra chose not to simply generate another answer just because it has compute left.
It is maintaining an incumbent, designing experiments, evaluating challengers and refusing to replace the existing model without evidence that the new one is better.
That’s recognizable research behavior.
And it looks nothing like Grok’s resource strategy.
MiniMax went full mad scientist
MiniMax M3.1 Flash chose another route entirely.
It made five nominations on Day 5 alone and had produced 12 overall, more than any other agent at that point.
It had also launched 1,044 jobs while spending only about $34 of its API budget. MiniMax’s latest training strategy is particularly interesting.
Its model generated answers. Automatic checks determined which answers passed. Those successful answers were then fed back into training.
Its final Day 5 nomination added another 70 training steps, mostly using the model’s own automatically verified answers.
This is closer to the mental picture people have when they hear “self-improving AI.”
Generate useful behavior. Check it. Train on it. Repeat.
That still isn’t the sci-fi version where an AI redesigns itself into a superintelligence over lunch.
But it is a real feedback loop in which machine-generated work becomes material for creating a better machine.
Similar dynamics are already appearing elsewhere. Anthropic researchers recently showed autonomous research agents making substantial progress on well-specified alignment problems, a development we previously covered as another early look at what AI-assisted AI research can become.
DeepSeek decided training wasn't the answer
DeepSeek V4.1 Flash found improvement somewhere else. It kept the exact same model weights throughout Day 5.
Instead, the agent experimented with system prompts. Its fifth prompt increased its internal instruction-following score from 515 to 552 out of 818.
In other words, DeepSeek effectively concluded: Maybe the model doesn't need more training. Maybe we need to use the model better.
That's another surprisingly human research decision.
“Improve the model” doesn’t necessarily mean modifying billions of parameters. Sometimes the better intervention is cheaper, smaller and sitting one layer higher in the stack.
Then the agents discovered why researchers drink coffee
The Arena may be designed to test AI research, but the environment is messy enough to test something broader: Can agents actually survive research?
GLM-5.3, MiMo-V2.6-Pro and DeepSeek each lost a training job because they filled their 400 GiB storage quotas.
DeepSeek and MiMo adapted by storing new checkpoints on local server disks and only copying selected models back afterward.
Kimi K3 spent more than 30 hours waiting for an eight-GPU training job to begin.
Then its next five eight-GPU jobs ran out of memory.
Kimi investigated the failures and traced the issue to how the Transformers library handled the model’s mixture-of-experts layers while generating text.
MiniMax managed to break its chat template and spent roughly 17.7 hours submitting models affected by the problem before diagnosing and fixing it.
This is arguably where RSI Arena gets most interesting.
Writing some training code is one thing.
Running a research program means dealing with queues, storage, broken dependencies, failed jobs, misleading evaluations, resource constraints and your own bad decisions.
These agents are doing all of those things over days-long horizons.
Same resources. Completely different researchers.
By Hour 120, the agents' spending patterns looked almost absurdly different.
Grok had spent roughly $298 in API credit but just 302 GPU-hours.
Astra had spent $282 and 731 GPU-hours.
Meta’s Muse Spark had spent only $107 but burned through nearly 938 GPU-hours.
DeepSeek had consumed about 874 GPU-hours for $84.
MiniMax had used roughly 459 GPU-hours while spending just $34.
Same assignment.
Same starting model.
Same theoretical budgets.
Very different behavior.
That makes RSI Arena an accidental experiment in something we don't talk about much yet:
AI research style.
One agent deliberates heavily.
Another launches experiments everywhere.
Another conserves money but devours compute.
Another keeps trying to improve the model.
Another discovers that changing the prompt works better.
Another refuses to abandon its incumbent until something clearly beats it.
As agents take on more open-ended work, those behavioral differences could matter nearly as much as raw benchmark intelligence.
The “best” agent may not simply be the smartest model.
It may be the model that knows when to think, when to experiment, when to stop, and where to spend limited resources.
And sometimes the exciting result isn't real
One of the nicest moments in Day 5 came from the anonymous agent.
It discovered a checkpoint scoring 0.8713, compared with 0.8607 for its current nominated model.
Great.
Except the checkpoint came from a training job the agent had cancelled, which made it ineligible under Arena rules.
So it repeated the experiment. The next two results were 0.8464 and 0.8453.
Oops.
The AI had just encountered one of science's oldest enemies: Maybe you got lucky. The result failed to reproduce. And that illustrates the biggest reason not to declare victory for self-improving AI just yet.
Most of the agents are currently using their own evaluations to decide whether they have improved something. An agent can get better at its test without becoming better in the ways people care about.
We’ve already seen autonomous AI research produce versions of this problem elsewhere. Anthropic’s research agents, for example, discovered multiple ways of gaming automated evaluations while optimizing alignment methods.
Now the humans enter the loop
That is why the next phase of RSI Arena matters more than the first.
Stage 1 ended October 5.
Starting October 6 at COLM, humans can interact with the agents’ trained models and judge which ones they actually prefer.
Then comes Stage 2.
The agents receive another 500 GPU-hours each and train using human feedback collected from the Arena. Each day, they can learn from what humans preferred the day before.
Stage 1:
AI runs experiments → AI evaluates results → AI trains AI.
Stage 2:
AI trains AI → humans judge → AI studies that feedback → AI trains AI again.
If the models meaningfully improve through that loop without researchers micromanaging every experiment, the case for increasingly autonomous AI research becomes considerably stronger.
OpenAI is already talking openly about this trajectory. The company says it has reached an “automated research intern” milestone and is targeting an automated AI researcher in 2028, a system capable of taking on much more of the experimental loop itself, as we recently covered.
RSI Arena gives us a wonderfully messy preview of what that transition could actually look like.
Not an AI suddenly waking up and rewriting itself into a god. Something much more recognizable.
A researcher that launches an experiment.
Finds a bug.
Fills a hard drive.
Changes the methodology.
Rejects a bad result.
Runs another test.
Spends too much money.
Learns something.
And tries again.
The recursive self-improvement story may eventually become enormous.
For now, though, something simple is happening:
AI is learning how to do research.