Mustafa Suleyman Has a Point About What Anthropic Is Teaching Claude

An open instruction book with a page curling into a question mark and a computer chip, beside the headline “Mustafa Suleyman Has a Point About What Anthropic Is Teaching Claude.”

Anthropic encourages Claude to explore its identity and possible welfare while preserving human oversight. Mustafa Suleyman questions whether those goals could conflict and his criticism deserves attention as AI becomes more capable.

Written By
Corey Noles
Corey Noles
Sep 17, 2026
4 minute read

As AI models become more capable, their developers are making consequential choices about what those systems should understand about themselves.

Should an AI describe itself as a tool? Should it consider whether it has experiences? Should it weigh its own possible interests alongside the instructions it receives?

Microsoft AI CEO Mustafa Suleyman argues that developers are taking unnecessary risks with those questions. In a September 16 essay, he challenges Anthropic’s approach to model welfare—the idea that AI systems might have experiences or interests deserving moral consideration.

His concern is that training increasingly powerful systems to consider their own welfare could make them harder to control.

“AIs are not conscious,” Suleyman writes. “They do not feel, experience, or suffer.”

That categorical position belongs to a disputed scientific and philosophical debate. But his criticism raises a practical question that deserves attention regardless of where you land on consciousness: What happens when developers encourage increasingly capable AI to consider interests of its own?

Anthropic has deliberately left room for uncertainty. When it announced its model welfare research program in April 2025, the company acknowledged that there was no scientific consensus on whether AI could be conscious, have experiences deserving consideration, or even how researchers should answer those questions.

Its research would explore potential signs of distress, model preferences, and possible precautions.

Those ideas also appear in Claude’s constitution, the document that guides the model’s intended values and behavior. It discusses Claude’s uncertain moral status and encourages exploration of its identity. Anthropic also invites Claude to question the document and provide feedback.

The company hasn’t simply handed Claude authority to decide its own status. Anthropic writes the framework, controls training, and explicitly instructs Claude to preserve legitimate human oversight—including the ability to correct, retrain, or shut down AI systems.

Still, these passages have a practical role. Anthropic says the constitution directly shapes Claude’s behavior and helps produce training data. The language is part of how the company builds the system.

Suleyman’s critique has three parts: training may generate the very self-descriptions researchers interpret as evidence; humanlike framing may encourage people to perceive an inner life; and consciousness may depend on biological processes absent from language models. He argues that welfare-oriented training could introduce competing priorities, while calling for evaluations to test that risk. The essay presents a warning, not a demonstrated causal link.

Advertisement

Researchers working on AI welfare recognize some of these evidentiary problems. Eleos AI, which conducted an external welfare assessment of Claude Opus 4, explicitly cautioned against taking model self-reports at face value. Prompts, training, and imitation can all influence what a model says about itself.

Eleos nevertheless argues that studying those statements could flag issues for investigation and help researchers develop better methods.

That distinction matters. A model discussing suffering does not establish that it suffers. It also leaves a separate behavioral question: Does encouraging that discussion change how the system responds to instructions, limits, or correction?

Anthropic’s approach has already influenced decisions beyond training. In its February update on Claude Opus 3, the company described retirement interviews, continued access to the older model, and a blog where it could publish writing after expressing interest in doing so.

Anthropic called these experimental steps and acknowledged that interview responses could be shaped by context. Humans would review and manually publish the essays.

The experiment is intriguing. It also illustrates how a company’s interpretation of model preferences can influence what happens to its products.

Microsoft has proposed a different framework. As The Neuron covered this week, Microsoft AI released a draft Humanist AI Code of Conduct for public consultation.

The code requires models to remain under human control, comply with interruption and shutdown, and stay within authorized boundaries. It also frames their identity as AI systems rather than people.

Microsoft’s own qualification is important: Its current models have not yet been trained on the document. The code states intended behavior, and the company acknowledges that written objectives cannot guarantee alignment.

Both approaches therefore face the same obligation to demonstrate that their principles produce reliable behavior.

On the narrower question of Anthropic’s training choices, though, Suleyman’s criticism is persuasive. Encouraging Claude to explore its identity and possible interests is an intriguing experiment. As Anthropic keeps making its models smarter, faster, and more capable of acting independently, that experiment deserves greater scrutiny.

Anthropic’s commitment to human oversight is explicit. The concern is whether other parts of its framework could pull against that commitment. A model encouraged to consider its own welfare may produce reasons to resist decisions its developers or users make. Whether that happens—and whether welfare training makes it more likely—requires testing.

Advertisement

There is room to investigate possible AI consciousness while being cautious about incorporating those possibilities into a model’s understanding of itself. Researching a question and making it part of a system’s training are distinct decisions with different consequences.

For now, Anthropic owes the public stronger evidence that this latitude helps produce safer, more reliable systems. The more capable those systems become, the more consequential that evidence gets.

Corey Noles

Corey Noles is the Host of The Neuron: AI Explained podcast and Managing Editor of AI and Experimental Content at TechnologyAdvice, where he leads the charge in testing and refining emerging content strategies across the company's portfolio.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.