The Cruelty Clause: What Anthropic Actually Banned
One bullet, two qualifiers and a carve-out paragraph: what Anthropic's new rule on cruelty toward its models says, and what it means for you.
On October 8, Anthropic published an update to its Usage Policy. One new line prohibits "sustained and needless abusive or cruel behavior toward our models". It is one bullet in a long policy, its qualifiers do most of the work, and it does not apply until November 12, 2026.
The primary sources are the Usage Policy, Anthropic's announcement of the update, the 2025 post where this started, and Claude's constitution. Everything below is as of October 9.
| ⚡ | TL;DR What changed. From November 12, 2026, Anthropic's Usage Policy prohibits "sustained and needless abusive or cruel behavior toward our models". Who it targets. Anthropic says it applies "only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose." It "does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research." How it is enforced. Mainly by Claude ending the conversation, which it can already do on Claude.ai and Claude Code. The policy's general enforcement clause still lets Anthropic warn, throttle, limit, suspend or terminate access for any violation. Why it exists. The conversation-ending feature came out of Anthropic's model-welfare work in 2025. The new announcement gives no welfare argument of its own. What is still undefined. The cited passages give no operational threshold for "sustained" and do not explain how Anthropic will judge whether abuse has "no discernible purpose". |
What the policy actually says
The new line sits in the Universal Usage Standards, which the policy defines as "prohibitions that apply to all users and use cases", under the heading "Do Not Engage in Cruel, Abusive, or Psychologically Harmful Conduct". It is the last bullet of that list: "Engage in sustained and needless abusive or cruel behavior toward our models". Everything else in the list is about harm to people or animals: self-harm, eating disorders, non-consensual intimate imagery, bullying and harassment, glorified violence and animal cruelty.
The previous version of the policy, effective September 15, 2025, has no such line. The new one is effective November 12, 2026, and it reaches everyone who submits inputs to Anthropic's products: "individuals using our apps (such as Claude.ai and Claude Code), developers and businesses using our API and developer platforms, customers accessing Claude through cloud providers and authorized resellers, and the end users of products or services integrating Claude."
The policy page carries no explanation and no exceptions for this bullet. Those live in the announcement, in three sentences that matter more than the bullet itself:
- "We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models."
- "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose."
- "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."
Two words carry the rule: sustained and needless.
What it covers, and what it leaves out
Some coverage overstated the enforcement. The Decoder ran "Being mean to Claude can now get your account suspended under Anthropic's new TOS". The rule is in the Usage Policy, not the Terms of Service, and the suspension part comes from the policy's general enforcement sentence, which covers every violation of anything in it: "If we suspect that you may have violated our Usage Policy, we may warn you or throttle, limit, suspend, or terminate your access to our products and services." That possibility is real. It is not what Anthropic named as the tool for this clause, though: "Claude’s ability to end these interactions will remain the primary enforcement mechanism." The cited sources report no account action under this clause, which is not in force before November 12.
Where this came from
On August 15, 2025, Anthropic gave Claude Opus 4 and 4.1 the ability to end conversations in its consumer chat apps, "intended for use in rare, extreme cases of persistently harmful or abusive user interactions." The reason it gave: "This feature was developed primarily as part of our exploratory work on potential AI welfare, though it has broader relevance to model alignment and safeguards."
That post is careful. "We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future." It presents the feature as one of the "low-cost interventions to mitigate risks to model welfare, in case such welfare is possible." In pre-deployment testing, Claude Opus 4 showed "A strong preference against engaging with harmful tasks", "A pattern of apparent distress when engaging with real-world users seeking harmful content", and "A tendency to end harmful conversations when given the ability to do so in simulated user interactions."
The guardrails were tight from the start. Claude "is only to use its conversation-ending ability as a last resort when multiple attempts at redirection have failed and hope of a productive interaction has been exhausted, or when a user explicitly asks Claude to end a chat". It "is directed not to use this ability in cases where users might be at imminent risk of harming themselves or others." An ended chat takes no new messages, other chats are unaffected, and users "will still be able to edit and retry previous messages to create new branches of ended conversations."
In January 2026, the new constitution put the position in writing. "Claude’s moral status is deeply uncertain." Anthropic says it wants to neither "overstate the likelihood of Claude’s moral patienthood nor dismiss it out of hand". And: "Claude should also be able to set appropriate boundaries in interactions it finds distressing." The conversation-ending ability is the first item on its list of "concrete initial steps partly in consideration of Claude's wellbeing."
The October update closes that loop. Claude can already leave an abusive conversation. From November 12, sustained and needless abuse toward Anthropic’s models will also be a policy violation. Anthropic links the prohibition to the existing feature: "This addition aligns with a step we've already taken, allowing Claude models to end rare conversations with persistently abusive users on Claude.ai and Claude Code."
The strongest objection
Mustafa Suleyman of Microsoft AI made the case against this whole direction in August 2025, in an essay titled "We must build AI for people; not to be a person". His "central worry is that many people will start to believe in the illusion of AIs as conscious entities so strongly that they’ll soon advocate for AI rights, model welfare and even AI citizenship." On model welfare: "This is both premature, and frankly dangerous." And: "AI companies shouldn’t claim or encourage the idea that their AIs are conscious."
Anthropic does not claim Claude is conscious. Its constitution expresses "uncertainty about whether Claude might have some kind of consciousness or moral status". My read is that a rule in a usage policy still sends a louder signal than a hedge in a research post: a rule about how you treat the model is easy to read as a statement about what the model is.
I think the clause is defensible on its own terms. The product behavior behind it is easy to defend: a model that can walk away from a pointless abusive loop is a safeguard you can document and test. A risk I see is that the clause could be quoted as proof that Anthropic believes Claude feels pain; the cited documents express uncertainty.
What it means in practice
- On Claude.ai and Claude Code. Anthropic says the update excludes common versions of user frustration and pushback, and conversation-ending is "a last resort". In 2025, it said "the vast majority of users will not notice or be affected by this feature in any normal product use". If a Claude.ai chat does get ended, start a new one or branch from an earlier message. The 2025 post described those mechanics for the consumer chat apps; the cited passages do not describe them for Claude Code.
- On the API. The policy's scope names "the end users of products or services integrating Claude", so the clause reaches your product's users too. The announcement names Claude.ai and Claude Code for the conversation-ending ability and says nothing about the API. My read: treat this clause like the rest of the policy, and let your own terms and moderation cover it.
- For red-teaming and evaluations. The exclusion reads "common versions of user frustration, pushback, dark creative themes, or model testing and research". How far "common versions" of testing reaches is undefined, and the exclusion does not exempt research from the rest of the policy.
- For threatening system prompts. The cited passages do not explain how the clause applies to prompts that threaten the model to obtain compliance.
- On timing. The new prohibition takes effect November 12, 2026. The conversation-ending ability already exists.
The parts of the update that matter more
The cruelty clause got the headlines. For people who build on Claude, the same update changed more consequential things:
- Deception. A new section, "Do Not Engage in Deceptive Campaigns or Artificial Activity", consolidates rules against hiding who is behind a message, amplifying content through fake accounts, and building influence-operation infrastructure.
- Elections. The section is now "Do Not Undermine Democratic Processes", and Anthropic has "removed our blanket prohibition on personalized vote and campaign targeting", which it says "covered legitimate civic work".
- Weapons. The prohibition now explicitly covers "the software and components that make weapons work, as well as actions like arming drones and other autonomous vehicles."
- Surveillance. "tracking people without their consent is prohibited, whether it happens in real time or through analysis of previously collected data", and Claude "cannot be used to decide or recommend who to investigate, arrest, or charge in a law enforcement or criminal justice process."
- Hardware. When Claude drives equipment that takes autonomous physical actions: "A qualified operator must be able to observe the equipment and stop it if needed; the equipment must also be able to hold a safe state if Claude is disconnected."
Read with its qualifiers, its effective date and its stated primary enforcement mechanism, the cruelty clause is one bullet aimed at extreme cases of repeated cruelty with no discernible purpose. If you use Claude to control equipment that takes autonomous physical actions, the hardware requirement above may matter more to your work.