Start the day here

Policy — Anthropic — Claude

The Problem With Anthropic's Ban on Cruelty to Claude

On October 8, Anthropic published a new usage policy and set it to take effect on November 12. Most of the pages tighten language the company says it already enforced: weapons software, surveillance, deceptive campaigns. One addition is new in kind. Users must not engage in "sustained and needless abusive or cruel behavior toward our models." Claude, the post says, can already end rare conversations with persistently abusive users on Claude.ai and Claude Code. That hang-up "will remain the primary enforcement mechanism."

Read the qualifiers before you treat the ban as a theory of mind. Anthropic says the rule is for extreme cases, repeated cruelty with "no discernible purpose." It leaves out ordinary frustration, pushback, dark creative themes, and model testing and research. Pointless abuse is prohibited. Harsh treatment with a purpose, including a research purpose, is carved out in the same paragraph. The lab has written a manners rule and listed the situations in which harsh treatment of Claude is still part of the job.

7 min read
A round red emergency pull handle with the words EMERGENCY PULL in white, seen from above.

The cruelty line sits in a document that is otherwise about what people do with a tool. A new section, "Do Not Engage in Deceptive Campaigns or Artificial Activity," gathers bans on fake accounts, fabricated news sites, and the infrastructure for influence operations, political or commercial. The elections section is renamed "Do Not Undermine Democratic Processes" and narrowed to deception, impersonation of candidates or officials, and turnout suppression. Anthropic removed a blanket ban on personalized vote and campaign targeting because, it says, the old wording caught legitimate civic work: nonprofits drafting voter information in other languages, election officials sending ballot-cure notices. Deceptive targeting and misuse of voter data stay banned under other sections.

Surveillance gets the same explicitness, language for a line Anthropic says it already enforced. The company points to a September threat report on AI used to identify and track political dissidents. The rewrite prohibits tracking a person without consent, in real time or by analyzing data already collected. Claude may not decide or recommend who to investigate, arrest, or charge. It may not be used to build or improve tools designed for the surveillance the policy forbids. Fraud monitoring that people have agreed to, content moderation, journalism, and legal research stay permitted. Consent, arrest lists, dossier-building, and those permitted exceptions all describe what a person does with Claude. The section never describes a harm done to Claude.

The hardware clause is blunter. After Anthropic's Model Hardware Standard, high-risk physical actions require a qualified individual who can watch the equipment and stop it. The equipment must stop or hold a safe state when that person intervenes, or when the connection to Claude is lost. Speed, force, reach, temperature, and the other operating limits have to be enforced by the machine or by a controller independent of the model's output. If Claude is driving something that can injure a person, the policy's answer is a human at the switch and a device that survives the dropout.

Microsoft barred its models from claiming they are people. Anthropic has taken a neighboring seat and faced the user. The model is not invited to announce a self. The user is told that some ways of talking to it will get the session closed. One instruction stops a marketing claim. The other polices a tone, and only when the tone has no stated point.

If Anthropic believed Claude could be wronged the way a person can, the testing carve-out would be the scandal of the document. The carve-out is there on purpose. Enforcement, the company says, will mostly be the model ending the chat. The duty of care in this section is a product setting. Some sessions close.

November 12 switches the whole policy on at once. The courtesy rule and the stop rule share an effective date. One tells you how to speak to Claude when you are being pointlessly vicious. The other tells a qualified operator to keep a hand near the machine, because Claude can disappear and the hardware still has to hold. The stop rule names a person and a safe state. The courtesy rule names a reason to end a chat. Anthropic published a hang-up for the vicious user and a switch for the machine. The switch is the sentence that knows what Claude is.

Live scoreboardFollow the AI race on AI Wars

Lab rankings, model preference, API volume, coding-agent heat, open-source stars, and prediction markets.

Related stories

A wooden judge's gavel rests on a thick black hardcover book against a dark background.Acute Social Issues

Markey's Biometric Moratorium

Today

Clear plastic sunglasses with violet-tinted lenses rest on a blue surface, a window reflected in the glass.Acute Social Issues

Who Saw the Bathroom Footage from Meta's Glasses?

Today

The white dome of the California State Capitol rising above trees at sunrise, with Sacramento's skyline behind it.Acute Social Issues

The Problem With California's Child Chatbot Law Starting in 2027

Today

Letters

0

No letters yet.

Write a letter