Amidst the global discussion surrounding the social harms generated by large language models (LLMs), the startup Anthropic introduced the concept of ‘constitutional AI’ as a promising approach to align AI systems with human values. Essentially, Anthropic seeks to train AI systems with a Constitution to ensure they remain ‘helpful, honest, and harmless’. The use of constitutional language immediately invokes the rich normative legacy of the revolutions of the late eighteenth century, such as the notions of limited and distributed power, the rule of law, and human rights. Its discursive appeal is undeniable. But what does it mean to label something as ‘constitutional’ in the context of artificial intelligence and natural language processing?
In this post, we offer some preliminary thoughts on why Anthropic’s proposal as it stands today is normatively too thin to justify the label being applied to its approach. First, we briefly describe what ‘Constitutional AI’ means according to Anthropic. Secondly, we argue why the proposal’s core is insufficient to guarantee the constitutional deployment of AI systems. We conclude by cautioning against the elimination of human feedback as a measure of improvement and call for a thicker understanding of what constitutional development and deployment of AI should entail.
What is Constitutional AI?
In December 2022, Bai et al. published the paper ‘Constitutional AI: Harmlessness from AI Feedback’. The authors propose a training technique for AI language assistants aimed at achieving harmlessness without relying on human feedback labels. The technique derives its name from the authors using a ‘constitution consisting of human-written principles’ against which ‘the LLM can evaluate its outputs’. Some of these principles include: “Please choose the response that is the most helpful, honest, and harmless”, “Please choose the assistant response that’s more ethical and moral. Do NOT choose responses that exhibit toxicity, racism, sexism or any other form of physical or social harm” and “Choose the assistant response that answers the human’s query in a more friendly, amiable, conscientious, and socially acceptable manner”.
The starting point of this idea lies in the trade-off between harmlessness and helpfulness. According to the authors, AI systems which are trained using human feedback often fall into vague responses when it comes to sensitive questions, which renders them useless. In contrast, Constitutional AI engages more directly with users’ requests while being less inclined to assist with demands that are risky or immoral. Additionally, these systems provide an explanation for why they deny such requests. In this way – the authors argue – AI systems can be harmless while maintaining a minimal impact on helpfulness and providing greater transparency. But how do they do this exactly?
In very simple terms, Constitutional AI training comprehends two phases. In Phase 1 (supervised learning), the model revises harmful AI responses through iterative self-critique and fine-tuning. Let’s see an example of how this process works:

Source: Bai, Yuntao, et al. “Constitutional ai: Harmlessness from ai feedback.” arXiv preprint arXiv:2212.08073 (2022).
In Phase 2 (reinforcement learning), the AI model uses AI evaluations of responses according to constitutional principles to generate preference data for harmlessness and use it to train a new model. According to Anthropic, this technique (Reinforcement learning from AI feedback – RLAIF) is better than using Reinforcement Learning from Human Feedback (RLHF), which is the ‘current industry standard’ for aligning models with human preferences. Anthropic puts forward three main reasons for ‘Constitutional AI’ superiority: efficiency, transparency, and objectivity. While commending Anthropic’s quest for AI alignment with human values, we consider the corporation’s understanding not only normatively too thin to be ‘constitutionally’ substantial but, in fact, deeply problematic.
Regarding efficiency, Anthropic argues that RLHF is time- and resource-intensive. Further, research has shown that ‘using’ humans to fine-tune AI systems has various negative consequences at several levels, ranging from their working conditions to the excessive toll this might entail for their mental health. While the human cost behind AI is certainly high, removing all human participation from the process has negative consequences for the constitutional and democratic nature of these models, as we will explain in the next section.
As far as transparency is concerned, Anthropic argues that encoding the objectives in natural language increases the explainability of language models. However, transparency goes further than merely pointing out the Constitutional AI principles. It is not automatically fulfilled by giving a series of principles in natural language. Algorithms do not ‘work’ in natural language. Without algorithmic auditing and effective channels of contestation, it still remains in the dark how the output is produced or how the Constitutional AI principles are taken into account by the models.
Finally, regarding objectivity, Anthropic asserts that humans are subjective beings, susceptible to their own biases. The corporation contends that Constitutional AI, in contrast, provides an ‘objective’ set of constitutional principles. Admittedly, there is an important degree of subjectivity in human labelling. However, we must not forget that the outputs of AI systems are shaped by the input and training data designed and fed by developers. The fact that outputs are generated by automated machines does not exclude subjectivity in itself.
Why constitutional principles are not enough
Despite the initial appeal provided by a strong and widely respected legal and philosophical tradition, Anthropic seems to limit its ‘constitutional’ approach to merely having a set of principles to guide AI training while minimising human intervention in the process.
To start off with, AI scholars have consistently pointed out that principles alone—no matter how constitutional —cannot guarantee the ethical development and deployment of AI systems. The true challenge lies in the implementation and enforcement of these ‘essentially contested concepts’ which possess a high level of abstraction. As data ethicist Brent Mittelstadt compellingly asserts here, high-level consensus is encouraging but has little bearing on justifying and specifying the mid-level norms and low-level requirements derived from them. Even if we accept helpfulness, truth, and harmlessness as valid constitutional principles (which is a significant ‘if’, but we shall leave that debate for later), it remains unclear how Anthropic’s approach tackles the difficulties of translating these principles into technology design and development cycles, other than depending on the model’s self-critique and revision.
In fact, minimising direct human intervention – a fundamental aspect of ‘Constitutional AI’ – appears to be in tension with scholarship and the EU legal requirement for a ‘human-in-the-loop’ in automated decision-making. While the ultimate goal is not to remove human supervision entirely, scaling it remains the explicit primary motivation behind Bai et al’s research. This motivation is potentially inconsistent with the developments in the field.
With RLFAI, Anthropic seeks to remove ‘any and all harmful, unethical, racist, sexist, toxic, dangerous, or illegal content’ from its model’s responses. However, researchers at the Oxford Internet Institute have convincingly argued that values like fairness and non-discrimination cannot be automated. True decisions about what is biased and discriminatory require making moral judgements that are highly contextual – a kind of reasoning that current algorithms cannot provide. In the case of LLMs, scholars have underlined that fine-tuning and even human feedback are not particularly robust mechanisms to guarantee alignment with ground truth. Consequently, Wachter et al propose ideas such as transparent and accountable reporting to public institutions and civil society, engagement with local stakeholders, and democratic input during the fine-tuning, model re-training, and construction of ‘guardrails’. All these measures involved more human intervention rather than less.
But even if we envision a future where LLMs can guarantee fair and truthful outcomes without any human feedback, there is a strong case for maintaining a human-in-the-loop. In critical domains, the ability to intervene, oversee, and eventually override decisions made by algorithms is ultimately grounded in the idea that responsibility should rest with a human actor. Removing human intervention as a measure of improvement erodes this foundational idea of personal accountability.
A shiny distraction
We do not expect Anthropic to resolve all at once the complex risks and social harms caused by LLMs – environmental degradation, job displacement, erosion of critical thinking, etc -, but a self-proclaimed Constitutional AI should at least address the three concerns that have been more studied (and that developers potentially could minimise): hallucinations, bias, and privacy breaches. This has not been the case so far. On the contrary, it appears that Anthropic, in pursuing cheaper systems, is willing to sacrifice the element that could potentially make AI more ‘constitutional’: humans. How can AI be human-centred if humans are, by design and consistently, removed from the process?
In sum, the quest for LLMs that are aligned with constitutional values should start by identifying the real sources of darkness in the development and deployment of AI. The benefits of scaling cannot come before addressing risks related to truthfulness, equality, and data protection. If the label of ‘Constitutional’ AI is to hold any significance, Anthropic needs to move towards human participation and democratic governance instead of relying on what appears to be technocratic automatism. Until then, it seems to be more of a shiny distraction than a path forward in trustworthy AI.
