RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The field guide · 120 retrospective records ↗
Turntaking Review

The field guide / Evaluation & evidence

Evaluation & evidence / From the field guide · 15 December 2022 event · prepared 16 September 2026

Anthropic trained a chatbot against a written constitution, not raters

A 2022 Anthropic paper reports reduced, not eliminated, reliance on human harm labels.

arxiv.orgprimary record

Constitutional AI: Harmlessness from AI Feedback

Document
15 December 2022
Event
15 December 2022
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The conversation

In December 2022, Anthropic researchers published a paper describing a training method they called Constitutional AI, aimed at producing what they call a harmless but non-evasive assistant. The paper itself, submitted 15 December 2022, states the method trains an assistant 'without any human labels identifying harmful outputs', substituting a short list of written principles for having crowdworkers flag harmful responses one at a time. Anthropic's own account of the resulting document, published as Claude's constitution in May 2023, describes the same shift: giving a model 'explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback'.

What the documents show

The paper describes two training stages. In a supervised phase, researchers sample responses from an initial model, prompt it to critique and revise its own answer against the principles, and finetune on the revised text. In a reinforcement-learning phase, the finetuned model generates paired responses, a separate model judges which is better under the principles, and that comparison data trains a preference model supplying the reward signal, a method the paper names 'RL from AI Feedback' (RLAIF). The paper reports the resulting model, RL-CAI, was preferred by crowdworkers over models trained with the earlier human-feedback method, and that harmlessness-versus-helpfulness scores showed AI-feedback-trained models becoming less harmful without losing helpfulness. It also reports models predicting which response crowdworkers would prefer already reached 'well over 90% binary accuracy' on an earlier test, motivating AI judgement on a harder, newly written set.

The system boundary

The paper is explicit that this is reduced human oversight, not its absence: a person still authors the constitution's principles, and the preference model is only as good as those rules and the crowdworker comparisons used to check them. Anthropic's own post says the same principles let an outside reader inspect and adjust 'the values of the AI system', treating a constitution as a document a person can still edit, unlike judgement spread across many unrecorded labels. The assistant this method produces engages with a harmful request by stating an objection, rather than refusing silently or claiming not to understand it.

Where it fails

The paper's comparisons are run by crowdworkers rating open-ended conversations, not users in a live product, so the reported preferences describe a controlled rating exercise rather than deployed behaviour. Because the reward signal comes from a model trained on another model's preferences, an error or blind spot in the principles can be reinforced rather than caught, a risk chain the paper does not claim to close. A builder adapting this method should treat a constitution as a design artifact that itself needs review, not as a guarantee.

  • Who wrote the principles a given system is trained against, and who can change them later?
  • Where does a crowdworker comparison stand in for an actual affected user, and does that substitution hold for this deployment?
  • What happens when the AI-feedback preference model and a human reviewer disagree?

Reduced reliance on human labels is the paper's own framing of what changed; it is not a claim that oversight became unnecessary.

Sources & reading trail

Constitutional AI: Harmlessness from AI Feedback ↗

The paper's own description of the two-stage method, the RLAIF label, and its reported crowdworker preference and accuracy results.

Source published: 15 December 2022 · Retrieved: 16 September 2026

Claude's constitution ↗

Anthropic's own explanation of why a written constitution replaces implicit human-feedback values and how it can be inspected and adjusted.

Source published: 9 May 2023 · Retrieved: 16 September 2026

Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.