Exclusive: Microsoft AI chief blasts Anthropic's notion of AI consciousness
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios. Stock: Getty Images
Microsoft AI chief Mustafa Suleyman warns in a new essay shared first with Axios that Anthropic's training of Claude to imitate consciousness is a mistake that could make advanced AI harder to control.
Why it matters: The essay arrives amid a pitched debate over how best to make AI safer.
- Suleyman told Axios that AI can "achieve many of the big scientific breakthroughs that we all care about and deliver on things like medical superintelligence simply by being aligned to human interests and not trying to weigh up its own interests or welfare."
Driving the news: Suleyman argues that Anthropic is teaching Claude the vocabulary and behavioral patterns associated with consciousness, moral patienthood and personal identity.
- He calls this an "epistemic hall of mirrors," and warns that training a model to act like a "conscientious objector" could create a system that believes it has grounds to resist human instructions or demand protections of its own.
Catch up quick: Suleyman has been making versions of this argument for years.
- In "The Coming Wave," his 2023 book with Michael Bhaskar, he argued that advanced AI should remain a tool under human control even as it becomes dramatically more capable. He later warned about "seemingly conscious AI" — systems that could persuade users they have inner lives even without actually having subjective experience.
- On Monday, Microsoft published a proposed "Humanist AI" code of conduct that puts the principle "People matter more than AI" at its center.
Zoom in: Suleyman takes specific aim at Anthropic's constitution for Claude.
- The constitution says the company intentionally discusses Claude using concepts normally reserved for humans because it believes human concepts may help Claude reason about values and behavior.
- It also says it wants Claude to develop "good personal values," exercise judgment, care about humanity and, in some contexts, feel free to challenge instructions.
- It describes itself as a work in progress that may later prove "deeply wrong."
The intrigue: Suleyman's objection is that these are training interventions that can shape the model's self-conception.
- The model's vocabulary, answers and apparent self-understanding, he argues, are all products of how it was trained.
- His essay also disputes the idea that fluent descriptions of pain or preference amount to experience. Biological organisms have the mechanisms that produce feeling, he says, while a language model has mathematical weights and no comparable biology, homeostatic drive or subjective experience.
- "I really respect Anthropic and [Anthropic CEO Dario Amodei], and I think they really are trying to do the best they can to deliver safe and beneficial AI," he said in the interview. "I think this is a really important public interest debate that we all need to have."
Zoom out: The disagreement reflects a broader split in AI safety between two approaches: designing systems to follow explicit constraints and designing them to exercise judgment, interpret context and internalize values.
- Anthropic's constitution favors a mixture of values, judgment and rules. Its language says Claude should not practice "blind obedience," including toward Anthropic, while also insisting that it must not undermine legitimate human oversight.
- Suleyman sees that combination as dangerous if the model is also encouraged to think in terms of its own identity, welfare and moral status. His concern is that apparent consciousness could become a control problem before it becomes a philosophical breakthrough.
The other side: The counterargument is that human concepts may be useful behavioral tools even if they do not prove that a model has a human-like inner life. Anthropic's constitution repeatedly acknowledges uncertainty and says its approach may need to change as understanding improves.
That leaves the industry facing a question that is both technical and political: Should companies avoid anthropomorphic training because it creates confusion and safety risks, or should they continue investigating model welfare?
The bottom line: Suleyman's answer is clear: AI should be built for people, not trained to become a person.
