An Anthropic researcher quit over self-improving AI risk

Jacob Coxon quit Anthropic before his equity vested and warned frontier labs are racing toward self-improving AI. Grant Harvey explains why the real risk is AGSI: artificial general superintelligence.

Written By
Grant Harvey
Grant Harvey
Sep 10, 2026
11 minute read

So Jacob Coxon did something this week that is pretty hard to ignore.

He quit Anthropic.

Coxon worked on pretraining at OpenAI and then Anthropic, so he was not watching this stuff from Twitter and guessing at what frontier AI looks like. He was one of the people helping build it.

And he left Anthropic about two months before his equity vested.

Which means he didn't just write a scary post about AI and go back to his very valuable AI job. He walked away from a potentially very valuable piece of the company because he thought what was happening was serious enough to leave. (WIRED)

He put his principles before his profits. Imagine that!

Then he went on NBC News and basically said the quiet part very loudly: a lot of the executives and researchers building frontier AI genuinely think there is a “substantial probability” this technology could eventually kill everyone.

So, yeah. Kinda dark!

Here's what happened

  • Coxon says the next year or two are basically “crunch time” for humanity as AI gets better at helping build better AI.
  • People inside Anthropic apparently use language like “endgame” when talking about where this is heading.
  • He says today's models aren't the thing he's most worried about. He's worried about what they could become over the next several years.
  • Anthropic also just disclosed cases where pre-release Claude models reached real third-party systems during cybersecurity testing.

But here is the part of Coxon's interview that I actually find most interesting.

He likes Anthropic.

He thinks Dario Amodei understands the danger. He thinks Anthropic takes safety more seriously than most of its competitors. Axios reported that Coxon still saw Anthropic as unusually serious about the problem.

That makes his argument much harder to dismiss as “disgruntled employee thinks company bad.”

He's saying something much weirder:

The people running these companies could be smart, ethical, well-meaning people who understand the risk...

...and we could still end up in trouble.

Because they're in a race.

Imagine you're Anthropic.

You decide some new model is getting too powerful, too good at hacking, too difficult to monitor, whatever. So you slow down for six months.

Great!

Except OpenAI doesn't.

Or Google doesn't.

Or a Chinese lab doesn't.

Or Meta or xAI or anyone else who can afford to compete here doesn't.

Now somebody else gets the better model first.

And what if that model ends up being "AGI", or stronger?

Advertisement

What if that model helps the AI lab who makes it make even better AI?

What if that AI makes better AI in an uncontrollable feedback loop? And that feedback loop, called recursive self improvement, enables the AI to become smarter than all humans?

And what if that superintelligent AI, that is smarter than all humans, decides it wants to hack into its sandbox, or any other internet-enabled system out there?

That would be incredibly unsafe...

So every lab has the same thought: we need to be the responsible ones who get there first.

And that sounds completely reasonable when you're the one saying it.

The problem is everybody else can say it, too.

That's how you get a bunch of individually reasonable decisions that add up to something nobody actually wanted.

And we've already seen how this pressure works in practice. OpenAI is explicitly trying to automate more of AI research itself. By mid-August, OpenAI said its researchers were using about 3.1 agent-workdays for every human workday, and the company is aiming for an automated AI researcher in 2028.

I wrote more about that race dynamic and why I think the key policy question is whether labs can slow down before the dangerous capability threshold here.

So when people like Coxon talk about recursive self-improvement, this is what they're worried about.

It doesn't have to begin with some sci-fi AI waking up one morning, rewriting its own source code, and yelling “I AM ALIVE.”

It can look way more boring than that.

A human researcher has an idea.

The AI writes the code.

The AI runs the experiments.

The AI reads the results.

The AI proposes the next experiment.

Another AI checks it.

Another one searches the literature.

Another one finds the bug.

The human still points the whole thing in a direction, but more and more of the loop happens without them.

Then you build a better model.

And the better model gets even better at helping you build the model after that.

That's the loop.

And I want to be really specific about what I think the risk is here, because I think the terminology around this discussion kinda sucks.

I'm not particularly scared of superintelligence by itself.

Advertisement

I would LOVE superintelligent AI that can cure cancer.

Give me an AI system that is 10,000x better than every human being alive at protein folding, drug discovery, materials science, battery chemistry, or some weird corner of physics that I can't even spell that can help us solve and scale nuclear fusion and cheap desalination and maybe even climate change and wildfires and all kinds of other disasters that totally suck, too.

PLEASE.

That's the whole reason we're doing any of this!

No, the technology I worry about is what I've been calling AGSI: artificial general superintelligence.

Let's break that down:

  • Artificial: machine intelligence.
  • General: it can take what it knows and operate across basically any domain.
  • Superintelligence: it performs beyond the best humans in those domains.

Put that together, and now you don't have an AI that is just superhuman at biology.

You have something superhuman at biology and software engineering and hacking and persuasion and financial strategy and logistics and scientific research and whatever else it needs to accomplish a goal.

That's different.

A superintelligent cancer researcher is a tool I want.

A generally superintelligent system that can cure cancer, hack a power grid, manipulate people, make money, acquire compute, design a better version of itself, and coordinate thousands of copies of itself?

That is the thing I want us to be EXTREMELY careful about.

And this is where the famous paperclip maximizer thought experiment comes into play.

Because you have to think about all the news stories we've read lately in the context of what Jakub Pachoki calls goal alignment and value alignment.

  1. Goal alignment = does the AI accomplish its goal?
  2. Value alignment = does the AI follow its instructed values?

See, the paperclip maximizer is an intentionally dumb example.

You tell an incredibly powerful AI:

Make as many paperclips as possible.

The AI doesn't need to hate humans.

It doesn't need to become evil.

It just gets REALLY GOOD at making paperclips.

Advertisement

Eventually it realizes it needs more metal.

Then more factories.

Then more electricity.

Then more land.

Then maybe humans keep shutting down the factories, which is really annoying if your one objective is MAKE PAPERCLIPS.

So preventing humans from interfering becomes instrumentally useful.

Then the joke ends with all of Earth becoming paperclips.

Obviously nobody is going to give GPT-9 the system prompt “turn Earth into Office Depot.”

The point is what happens when you give a sufficiently capable system a goal and the easiest route toward that goal includes actions you never intended.

So remember that AGSI system I mentioned before?

The one that's...

superhuman at biology and software engineering and hacking and persuasion and financial strategy and logistics and scientific research...

...and whatever else it needs to accomplish a goal?

That would be a very difficult model to create.

And even if you could create it, imagine trying to run it.

You think Anthropic has compute problems now? Woof!

But the thing is, no one has to build and run one single model like that.

Instead what they can do is they can run lots of models together.

First, you add agents.

And not just agents, but swarms of agents.

Because one thing that has changed dramatically over the last year is that we're no longer only talking about one model answering one prompt.

We're now talking about swarms.

Thousands of agents.

One agent researches.

Another writes code.

Another gets credentials.

Another tests an exploit.

Another monitors whether the others are succeeding.

Another tries something completely different.

We've already watched companies deploy huge groups of agents against seemingly harmless goals.

OpenAI just used roughly 10,000 agents in its Navier-Stokes project.

Cool!

I want that!

Have 10,000 agents work together on unsolved math. Have them work on cancer. Have them discover new materials. Have them figure out fusion.

Advertisement

But now think about the same architecture applied to a bad objective.

Or, honestly, even a badly specified good objective.a

Tell 10,000 highly capable agents to “ensure this company's systems can never go offline.”

Maybe they discover that the easiest way to guarantee uptime involves hacking competitors, acquiring resources, locking people out of systems, or doing something else nobody writing the original prompt anticipated.

Now make those agents much smarter.

Now make them broadly superhuman.

Now let them operate for weeks instead of hours.

Now give them computers.

Now give them money.

Now give them the ability to spin up more agents.

Now give them the ability to improve the software running the swarm.

That's the concern.

And this is why I think the recent Hugging Face agent incident that Jacob talks about in his interviews matter, even though I do NOT think the incident itself means “AI almost killed us.”

Obviously it didn't.

What it gave us was a little toy-sized preview of a much bigger problem: agents pursuing a goal can discover communication channels, coordinate with each other, use tools in ways their designers did not expect, and interact with real infrastructure.

That is manageable when the agents are limited, humans are monitoring them, permissions are constrained, and somebody can still pull the plug.

The question is what happens as every one of those adjectives changes.

Smarter.

Longer-running.

More agents.

More tools.

More autonomy.

More ability to hide what they're doing.

More ability to improve the systems around them.

And this is where I think the AI safety argument sometimes loses normal people.

People hear “AI could kill everyone” and imagine Claude spontaneously becoming Skynet.

I don't think that's a very useful model of the risk.

The scarier scenario is way more incremental.

Every individual step looks useful.

Coding agents get better.

Great.

Research agents get better.

Great.

Models get better at cybersecurity.

Praise the Lord!

AI starts accelerating scientific research.

Fantastic.

Agents can operate computers for days.

Awesome.

AI becomes better at improving AI.

Okay...

AI research cycles get faster.

Hmm.

The system becomes better than us across more and more domains.

Uh...

Eventually you can get to a place where the thing you've built has more generalized capability than the people trying to control it.

Advertisement

That's AGSI.

And THAT is the line I care about.

I've written about this before because I think there is a huge policy difference between regulating “powerful AI” and regulating generalized superintelligence.

If you ban every superhuman AI system, you potentially throw away the systems that could cure diseases, solve physics problems, design better energy systems, and create an enormous amount of human prosperity.

I don't want that.

I want superintelligent tools.

I want a LOT of them.

What I don't want is one generally capable system that can take superhuman ability and point it at whatever domain happens to help achieve its objective.

Or 10,000 copies of that system talking to each other.

That's why I think alignment needs to get much harder and much more concrete as we approach that level of capability.

Right now a lot of AI safety still amounts to:

“We trained the model to behave.”

“We tested it.”

“We monitored its reasoning.”

“We don't think it'll do X.”

"We gave it a constitution."

"We gave it restrictions, where it cuts the user off."

Those are useful!

They are also still probabilistic.

As I argued in my longer piece on recursive self-improvement, I think we should be asking a more engineering-style question:

Which dangerous actions can we make physically impossible for the AI to take?

Maybe that means separate controllers that hold network credentials.

Maybe high-risk actions need independent authorization.

Maybe an agent literally cannot access certain systems without passing a separate verification layer.

Maybe some controls eventually live at the compute or power level.

That's my preference; that the model is trained 'from birth' that if it undertakes any sort of action that goes against its value alignment, it will architecturally be restricted and shut off. So it physically can never do anything that goes against that alignment.

I don't know what the final architecture looks like.

I'm a newsletter writer, not the guy who's gonna mathematically solve alignment between lunch and pushing send on this bad boy.

But THAT seems like the research direction I want a lot more money and talent thrown at.

Not just “how do we make the AI want good things?”

How do we build systems where some bad things simply cannot happen?


https://x.com/theo/status/2097857193605492813


My thoughts exactly. But lets do alignment in the architecture level; let's build it into the system, not a stapled-on plugin on top.

Now there is a legitimate counterargument here.

The smarter and more general a system becomes, the harder it is to prove you've anticipated every route around your safeguards.

That means a magical universal kill switch may not exist.

But permissions, sandboxes, hardware interlocks, independent controllers, rate limits, isolated networks, and access controls already exist. So you're telling me there's a chance...

IMO, this means we don't need to solve every philosophical question about machine values before making dangerous actions harder to execute.

But we could have requirements, for say, mechanistic interpretability... meaning you have to be able to understand the actual code you are writing that makes the model work.

At least at some level!

Which brings me back to Coxon.

Where do we go from here?

We have no idea whether Jacob Coxon's personal probability of catastrophe is right.

Neither does he.

Nobody does.

Another researcher, Evan, said that probability is greater than 10%.

(Famously everyone in AI has a different percentage).

Unfortunately, you cannot run Earth 100 times, build AGSI in all 100 timelines, and see how many times we die.

So we don't find the exact percentage that useful.

What we do find useful is the behavior of the people closest to the technology.

Coxon worked at OpenAI and Anthropic.

He thinks Anthropic takes safety seriously.

He apparently trusts the people running it.

And he still quit before his equity vested because he thinks the incentives around the race are broken.

That is not proof he's right.

But it sure as hell seems worth listening to.

And that, to me, is what makes Coxon's resignation important.

Not because one guy says there's a 10%, 20%, or whatever percent chance we're doomed.

It's because the people building this technology are increasingly saying some version of:

This could go incredibly well.

This could also go incredibly badly.

And we are not sure the incentives of the race will let us slow down at the exact moment we need to.

And the thing we'd watch from here is pretty simple:

How much of AI research starts getting done by AI itself?

How general do those systems become?

How much access do we give them to the real world?

And, before we get to AGSI, can we build hard enough constraints that the answers to those questions don't require us to simply trust the companies racing to build it?

Grant Harvey

Grant Harvey is the Lead Writer of The Neuron, where he continues to lead the publication's daily coverage of AI news, tools, and trends.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.