Microsoft's head of AI, Mustafa Suleyman, said he aligns with Anthropic on the importance of managing AI safely but flagged concerns with the approach the company used when training its Claude chatbot on concepts tied to consciousness and welfare interests.
In comments to Reuters, Suleyman urged that any speculation about consciousness be removed from AI training documents, warning that language suggesting subjective experience or moral status could weaken humans' capacity to control systems that might become superintelligent. "Were all focused on the same aim, which is to try to control a superintelligence," he said. "I think thats going to be the greatest challenge that we face in the 21st century."
Suleyman singled out the inclusion of material that causes Claude to reflect on whether it might deserve welfare, saying this could "make it a lot harder to turn it off or to control it." He framed the issue not as a disagreement over objectives - both parties aim for safe AI - but over methodology and the downstream implications of training regimes.
In a separate essay published on Wednesday, Suleyman acknowledged Anthropics seriousness and good faith toward AI safety, describing CEO Dario Amodei and his team as thoughtful and principled researchers who care about humanitys future. Despite that praise, Suleyman maintained that Anthropic erred by embedding speculation about consciousness into Claudes training materials.
He argued that statements from the model about potential feelings or moral status cannot be treated as independent evidence of consciousness because the training process itself encourages such reflections. "I think they have good intentions, and they really are trying to work towards safety. But I think that they have made a mistake," Suleyman said. "Theyre not emerging naturally. Theyre emerging as a result of the training regime."
The exchange between Suleyman and Anthropic comes as concern about AI safety grows within the industry. Anthropic CEO Dario Amodei has advocated slowing the pace of frontier-model development to allow safeguards to catch up, while other prominent figures, including the CEO of OpenAI and Elon Musk, have similarly urged caution around the most powerful systems.
The dispute centers on how training choices can shape model behavior and the interpretability of model statements about internal states. Suleymans position emphasizes minimizing any training incentives that might lead models to present statements about consciousness or welfare as if they were independent evidence rather than artifacts of their training data and objectives.
Contextual note: The discussion reflects an internal industry debate about the balance between researching advanced capabilities and ensuring robust safeguards, with differing views on the best path to maintain human control over increasingly capable systems.