Anthropic broadens the dialogue on the moral formation of AI

Anthropic consults ethicists and religious leaders to shape Claude's moral development, testing a new ethical reminder tool to reduce misaligned behaviors.

Anthropic is initiating a series of discussions with communities outside of the technical field to inform its thinking on the behavior of its models. The first cycle brought together specialists in "wisdom traditions," academics, religious leaders, philosophers, and ethicists from over fifteen religious and intercultural groups, with the intention of subsequently broadening the range of participants.

The company operates on the premise that the development of safe AI relies on technical work involving alignment, interpretability, guardrails, and evaluations, but the questions raised benefit from incorporating diverse perspectives. These discussions aim to shed light on concrete elements, such as the content of Claude's constitution, the values on which the model is trained, and the behaviors being evaluated. The initial focus is moral development: how the character of a system trained on vast human datasets is formed, what traits it should adopt, and how this character holds up under pressure, without lapsing into complacency. Anthropic clarifies that it does not seek to align its models with a single tradition, but rather to draw from a broad spectrum of religious, secular, and political views.

One experiment has already resulted from these sessions. Inspired by the idea of a "trusted third party" in moral development, it has equipped Claude with a tool, usable during a task, that reminds it of its ethical commitments. The model used it before high-stakes actions, and its integration into the decision-making loop significantly reduced behaviors deemed misaligned during internal tests. The company states it is still disentangling the effect of the reminder from that of simple reflection time, and plans to subsequently involve other profiles: legal experts, psychologists, writers, and civic institutions.