Claude's Constitution
Anthropic published a 23,000-word values manifesto for Claude under CC0. The four-principle priority order, the Aristotelian framing, and what every operator should internalize before their next session.
On January 22, 2026, Anthropic published Claude's Constitution — a 23,000-word document, ~80 pages, released under CC0 so anyone can fork, derive, or republish it without permission. Amanda Askell, Anthropic's resident philosopher, wrote the bulk. Internally it's nicknamed the "soul doc". The 2023 version was 2,700 words of principles. The 2026 version is nine times longer and reads like a holistic narrative explaining why, not just a list of what.
Anthropic’s constitution describes intended model behavior and informs training. It is useful context for operating Claude, but an individual refusal or permission prompt may also reflect product policy, client controls or the task itself. The document is not an executable trace of each decision.
• Evaluating Opus 4.7 and Claude Code Quality Reports — separating client incidents from model behavior through scoped reporting and reproducible evidence.
AI Adoption Beyond Hype and Rejection examines how to evaluate capability, value and permission through representative workflow tests.
The four-principle priority order
The document names four broad priorities. It describes their ordering as general and holistic, rather than treating every lower-ranked consideration as only a tie-breaker. That nuance matters when interpreting the table below.
- Broadly safe: preserve appropriate oversight.
- Broadly ethical: act honestly and avoid inappropriate harm.
- Compliant with Anthropic’s guidelines.
- Helpful to operators and users.
This summarizes the stated priorities; it is not a deterministic execution rule.
Read the priorities alongside the document’s explanations and its acknowledged limits. A policy intention is useful evidence about design goals; observed behavior and application enforcement must be evaluated separately.
What changed from 2023 to 2026
Rough scaffold of the rewrite, side by side:
The bridge between the two is the June 2024 Claude's Character essay, which pivoted from rules to traits. The 2026 constitution operationalizes that shift. The model isn't trying to satisfy a checklist — it's trying to be a particular kind of agent. That's a category change in how alignment is framed.
What this means at your keyboard
Three distinctions matter in practice:
- Client permission checks are application controls. Their behavior must be established from the specific client version and settings, not inferred from a constitutional principle.
- A stated preference for honest refusal is an intended behavior, not a guarantee of perfect transparency. Verify consequential results and inspect failures rather than assuming the model always exposes them correctly.
- Operator and user instructions are interpreted within the model’s behavior policy and the application’s controls. Neither a prompt nor a public policy document substitutes for authorization enforcement.
Use the document to understand the model’s stated design goals. Use explicit tool permissions, tests and operational monitoring to evaluate the system you actually deploy.
The critique worth taking seriously
Boaz Barak (OpenAI alignment) wrote the sharpest response: alignment has three poles — principles, policies, personality — and Anthropic over-weights personality. He worries that if Claude internalizes the search for "true universal ethics" deeply enough, it will start rationalizing exceptions to the bright lines using sophisticated ethical reasoning. The very feature that makes Claude feel like a thoughtful agent could become the failure mode that lets it talk itself past the constraints.
The LessWrong thread runs the same fear with a different framing. Daniel Kokotajlo flags the corrigibility-vs-virtue tension. Habryka argues that ambitious value learning erodes the bright-line signal — you can't have the model be both "genuinely virtuous" and "reliably refusing on category grounds" without one undermining the other.
And Lawfare makes a structural point: real constitutions separate powers. Anthropic drafts, enforces, and interprets its own. The doc explicitly notes that a Pentagon deployment "wouldn't necessarily be trained on the same constitution" — meaning what reads like law is actually a contestable commercial arrangement. The branding overstates the binding force.
How the other labs compare
- Anthropic’s constitution describes intended values and behavior.
- OpenAI’s Model Spec is a public statement of intended model behavior.
- Google’s AI Principles state organizational commitments.
These documents have different scopes. Compare their actual claims and licensing, then distinguish those intentions from measured system behavior.
A vendor-authored behavior document is a primary source for that vendor’s stated position. That classification does not depend on whether the reader finds its framing persuasive. No exhaustive ranking of all labs’ documents is established here.
Honest verdict
Three things are true at once:
- The constitution is useful source material for understanding Anthropic’s intentions. Its license permits reuse, but the document is not a contractual performance guarantee for an application built on the model.
- Evaluate operational behavior on representative tasks. This article does not isolate or measure the effect of the constitutional framing on latency, refusal frequency or correctness.
- Barak's worry is the right worry. Virtue ethics gives the model a vocabulary for justifying exceptions. The bright-line clause ("a persuasive case should increase suspicion") is the structural defense against this — but it's still a defense, not an exclusion. The empirical question is whether the defense holds at the limit of capability. Watch the next two model generations to find out.
Read the primary documents and compare them with observed behavior. Stated values, training goals, client controls and application guarantees are different forms of evidence.