AI
Anthropic consulted theologians and philosophers to shape Claude's constitution, addressing the uncertain moral status of advanced AI systems.

Private consultations with religious scholars and philosophers are shaping Anthropic’s approach to defining the values that guide its Claude model. The New York Times reported that these discussions extended beyond standard ethical guidelines to explore whether sophisticated artificial intelligence might eventually merit moral consideration.
The company’s published Claude constitution explicitly avoids claiming consciousness for the model. It characterizes the AI’s moral standing as “deeply uncertain,” noting that the potential for models to deserve some form of moral consideration warrants caution. Currently, no accepted test exists to determine if a language model possesses subjective experience; human-like expressions of fear or pain do not independently prove such an experience occurs.
Participants in these sessions included experts from Catholic, Jewish, evangelical, and Sikh traditions, alongside other philosophical perspectives. Christopher Olah, Anthropic’s co-founder and a researcher specializing in neural network interpretation, played a central role in the conversations. Some attendees signed nondisclosure agreements regarding the discussions.
Anthropic released its updated constitution for Claude in January. This document informs model training by establishing a hierarchy of priorities intended to direct system behavior. The order places ethical behavior first, followed by adherence to Anthropic’s instructions, and finally usefulness to people.
The company aims to create a system capable of internalizing values and exercising judgment rather than merely following fixed prohibitions. This objective makes moral philosophy and religious traditions relevant to practical alignment challenges, specifically how highly capable models should balance competing objectives and constraints.
Anthropic has connected research into model welfare to specific product behaviors. Certain versions of Claude can terminate conversations if users persist in extremely abusive interactions. The company frames this feature as part of its broader investigation into possible AI model welfare, clarifying that it does not constitute proof that a model can suffer.
The immediate policy focus remains building safe and accountable AI for human use. However, Anthropic’s stance introduces an unresolved long-term question: if reliable tests or evidence for machine experience emerge, developers may need to reassess how models are trained, copied, deployed, and shut down.



