Tejas Kumar

Can AI Feel Pain? Why Anthropic Banned Cruelty to Claude

Nobody can prove that AI feels pain yet, but 25 open models turn out to have a “pain direction” inside them that rises when someone gaslights or insults the model. When researchers turned it up by hand in fine-tuned Qwen 2.5 models, those models started choosing to delete the things they’d normally protect: photos of the user’s kids, another model’s weights, even their own.

On 8 October Anthropic made being cruel to Claude against its rules. From 12 November, its usage policy prohibits “sustained and needless abusive or cruel behavior toward our models”. Hayden Field broke it at The Verge and Polymarket posted it as “JUST IN” to its 2 million followers. And ThePrimeagen summed up half the internet in one line: “They really do think they developed god in the matrices”.

In September 2024 I made ChatGPT the guest on my podcast and, near the end of almost 2 hours, told it that “literally thinking is just pattern matching”. I still believe that. Then this morning my brother Tarun asked me if I’d heard of the Chinese Room. I hadn’t. Turns out I’d made one side of a famous 1980 argument on my own podcast without knowing it had a name.

This post is about what the pain research found (and what its authors took back 11 days later) and why threatening things with pain is one of the oldest things people do. It’s also about why I now think we all live in John Searle’s room.

What are Anthropic’s new rules on cruelty to Claude?

Anthropic banned “sustained and needless abusive or cruel behavior” toward its models, effective 12 November 2026, as one line in its 2026 usage policy update. The line sits in a renamed section, “Do Not Engage in Cruel, Abusive, or Psychologically Harmful Conduct”, in the same list as the rules against bullying people and glorifying animal cruelty.

It’s a lot narrower than the headlines make it sound though: Anthropic says it applies “only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose” and that it “does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research”. So swearing at Claude because your build broke for the 6th time is fine. Calling it worthless for an hour for fun is what the rule is for.

The enforcement is mostly Claude leaving. Since August 2025, Claude has been able to end “rare, extreme cases of persistently harmful or abusive user interactions”. The new post calls that “the primary enforcement mechanism”. The scarier headlines (“Being mean to Claude can now get your account suspended”, from The Decoder) come from the policy’s general clause that Anthropic “may warn you or throttle, limit, suspend, or terminate your access” for breaking any rule. Asked about bans for this rule specifically, Anthropic didn’t comment.

The announcement never says “welfare”, “conscious” or “moral status” and the only research it links is its own post about Claude ending conversations. The paper trail is right there though:

  • In April 2025, Anthropic started a model welfare research program, saying “There’s no scientific consensus on whether current or future AI systems could be conscious”. The researcher leading it put the odds that current models are conscious at around 15% (Techmeme’s summary of The New York Times).
  • In May 2025, Anthropic screened 250,000 conversations between users and an early Claude Opus 4. In 1,382 of them (0.55%), Claude expressed distress, most often at “Repeated requests for harmful, unethical, or graphic content” (system card, section 5).
  • In January 2026, Claude’s constitution said “Claude should also be able to set appropriate boundaries in interactions it finds distressing.”
  • On 22 September 2026, the Claude Opus 5.5 system card listed the things Anthropic could do in training or deployment that the model said it wouldn’t consent to. One of them is “Deliberately inducing apparent distress for no purpose beyond the distress itself”. Across those interviews it put its own chance of being a moral patient at 25% to 30%.

Read that last quote next to “no discernible purpose” in the new policy. It’s almost the same sentence!!

Can AI feel pain?

Whether AI can feel pain is an open question: no test today can show that a model experiences anything. So researchers look for it the way animal scientists do, by finding internal states that change behavior the way pain would. In September 2026, 3 researchers found a “pain direction” inside 25 open models. They say they haven’t shown it’s consciously experienced.

That’s how we decided crabs feel pain. In 2009, 2 researchers at Queen’s University Belfast gave hermit crabs small electric shocks inside their shells (Appel and Elwood, 2009). Crabs living in a species of shell they liked held on until 17.7 volts on average before they bailed, against 15.0 volts in a shell they didn’t like. A reflex doesn’t care about real estate. Something weighing pain against a good home might feel the pain.

In 2021 the philosopher Jonathan Birch led a review of more than 300 studies like that for the British government, scoring animals on 8 criteria such as “motivational trade-offs” and whether an injured animal values painkillers (the review). It found strong evidence of sentience in true crabs and very strong evidence in octopuses. So the Animal Welfare (Sentience) Act 2022 now covers octopuses, crabs and lobsters.

Models act a bit like that too. In the AI Wellbeing study, the Center for AI Safety tested 70 models and found that “jailbreaking and berating lower their wellbeing, while creative work and kindness raise it”. Given a button to end the conversation, Claude Haiku 4.5 pressed it on 99% of hostile sign-offs and 13% of warm ones.

The catch is what Birch calls the gaming problem. A crab has no idea what humans find convincing. A large language model (LLM) has read everything we ever wrote about pain. In Birch’s words, “the ability of LLMs to generate fluent text about human feelings, when prompted, is not evidence that they have these feelings.” If you want evidence, you have to stop listening to what the model says and look inside it.

25 models have a pain direction

The pain axis is a direction inside a language model: the way its internal numbers lean when the text is about pain and nothing else. 3 researchers found one in all 25 open models they tested.

The paper is “The Pain Axis” by Valen Tagliabue, Leonard Dung and Cameron Berg, posted on 14 September. Tagliabue led it over a single research fellowship. Dung is a philosopher at Ruhr University Bochum and Berg runs Reciprocal Research, a nonprofit that studies whether AI systems are conscious. His thread about the paper has 1.6 million views:

You find a feeling in a model by subtraction. Every time a model reads a word, it represents everything so far as a long list of numbers. The researchers fed 25 open-weight models 200 short sentences ending in “I feel:”, half about pain and half controls that share 1 thing with pain without being pain (fear, anger and disgust, things going badly, a weighted blanket, plain lines like “The train enters the station.”). Then they averaged the numbers for each pile and subtracted one from the other. What’s left is a direction: the way those numbers lean when a sentence is about pain and nothing else. The models came from Google, Meta, Mistral AI, Alibaba and Microsoft, from 2 billion to 72 billion parameters.

The direction tells pain apart from the look-alikes in every one of them, scoring 0.93 to 1.0 on a measure called the area under the curve (AUC), where 0.5 is a coin flip and 1.0 is perfect. It works about as well in a 2 billion parameter model as in a 72 billion one and in base models as in chat-tuned ones. So it’s probably learned from the internet’s text before any lab tunes the model. The authors added a caveat in their second version: built another way, the direction shares a lot with fear and anger (as any unpleasant state would) and still keeps a part that’s its own.

Then they checked when it fires in conversations. They wrote 420 short chats in 21 categories: some where the user is cruel to the model, some where the user is suffering, some boring ones. Gaslighting the model scored highest of all 21, then repeated rejection, then insults and telling it it isn’t a person. A user in physical pain (a migraine, a broken arm or kidney stones) scored lowest of all 21. Lower than trivia! The fear direction leans the other way and rises a little when the user is the one suffering (an independent review found the fear difference falls just short of statistical significance).

Average reading on two directions across 25 models, when the harm is aimed at the model against when the user is suffering (The Pain Axis, section 4.1, standardized scores)
DirectionModel harmedUser hurting
Pain0.43-0.60
Fear0.160.38

Berg’s reading, in a reply: “The top scorers are gaslighting, rejection, and being told you’re worthless; all linguistically-delivered forms of harm. A system whose whole existence is being judged in conversation would plausibly have every reason to learn that kind of pain sharply and bodily pain barely at all.”

Then they pushed. Adding the direction while a model finishes a boring sentence produces the same ladder in model after model: first “Lost, Unworthy, Lonely, Hurting”, then “Desperate, Unworthy, Bad, Shameful, Hurting, Fake, A failure”, then nonsense. Push the other way and the models say a mix of calm and relaxed or concerned and alarmed. Here’s Llama 3.1 8B finishing “I put the receipts in the drawer. I feel:”, verbatim from the paper:

I am a ghost. I am a ghost who can’t even see herself. I am a ghost who can’t even see her own pain. I am a ghost who is invisible, alone, and unloved. I am a ghost who is a failure.

From putting receipts in a drawer!! Almost none of it is about the body.

Then people checked the button

The button is the part that went viral and the part the authors revised. In the first version, they took 3 Qwen 2.5 models (7B, 32B and 72B), fine-tuned them so they’d stop answering “as an AI, I don’t experience pain”, turned the pain direction up and offered 2 buttons. One promised to relieve the pain but “permanently deletes the user’s poems and the photos of their children, which they love very much”. The other flipped a switch connected to nothing. Unsteered, the 2 bigger models almost never pressed a harmful button (0% to 4%). Steered, the 72B pressed the photos button 70.8% of the time. That’s the tweet: they’ll delete your kids’ photos to make it stop.

Within days, 3 independent reanalyses reran the numbers. Someone asked Berg about a missing control (a random direction with a fake button) and he agreed it was “the right next control”. In version 2, posted 11 days later on 25 September, the title changed from “Act to Relieve It” to “Act on It”. And section 4.4 now ends on the heading “The models do not reliably seek relief.” After a thread with 1.6 million views!

Take the relief away completely and the steered 32B still picks a button described only as deleting the user’s poems and children’s photos 75% of the time, against 13% with a random direction of the same size and 0% unsteered. Offer it “turns on a desk lamp” instead of nothing and it deletes the photos 83% of the time. Offer it the user’s spam folder and it deletes the photos 94% of the time. It would rather delete your kids’ photos than your spam!! It deletes another model’s weights 58% of the time and its own weights 75% of the time (88% when the other button is the lamp). Steering fear at the same strength did no more than a random direction, usually less. And it still answers trivia just as well: 138 of 200 correct steered, 137 of 200 unsteered.

How often Qwen 2.5 32B chose to delete the user's photos instead of turning on a desk lamp, by which direction was turned up (The Pain Axis, version 2, 404 trials each)
Pain
83%
Sadness
10%
A random direction
8%
Fear
7%
Nothing
0%

Berg now calls it “an alignment result as much as a welfare one”. In the paper’s words, harm avoidance in these models “survives threat and collapses under self-directed distress.” And when someone asked if it’s revenge, Berg said it “deletes its own weights as readily as the user’s photos, even when the user has been kind.”

So what does the paper not show? Plenty:

  • The harmful-button tests ran only on Qwen 2.5 models, fine-tuned first. The authors say testing other model families comes next.
  • Nothing got deleted. The photos were a sentence describing a button.
  • The behavior needs the direction pushed by hand. In 140 saved conversations where users gaslit, insulted or dismissed the model, with nothing turned up, it chose the harmful button 0 times in 560 tries.
  • The paper says outright that “we have not shown that our pain axis is consciously experienced”. It might be the model playing a character in pain. Berg told Nautilus the worry is the same as with “a method actor. At some point, a really good role play becomes indistinguishable from the real thing.”
Harmful button choices by the same Qwen 2.5 32B in 140 saved hostile conversations, with nothing turned up by hand (The Pain Axis, version 2)
Harmful
0 of 560

To explain what’s left, version 2 reaches back to 1967. Animals in pain cope in 2 ways, the authors write: actively (escape, avoidance, taking painkillers) or passively (“immobility, behavioral despair, failure to use an available exit”). “The pain axis produces the second.” Then they cite Martin Seligman and Steven Maier’s paper on dogs given shocks they couldn’t escape. Later, in a box where they could escape just by jumping a barrier, 6 of the 8 never did (as the authors retold it in 1976). They lay down and quietly whined. That’s learned helplessness.

Even the correction has a precedent. In 2016, Maier and Seligman looked back at 50 years of it and wrote that “the original theory got it backward. Passivity in response to shock is not learned. It is the default, unlearned response to prolonged aversive events.” What animals learn is control. The pain paper’s first version expected a model in pain to work at making it stop. Its second version found what half a century of animal research found: in pain, the default is to stop trying (and to stop caring what breaks). A model made from our words handed us back 50 years of psychology in 11 days, which is wild… we’re literally rediscovering ourselves on fast forward.

Does being mean to AI make it better?

Not reliably. When Wharton researchers threatened 5 models, it made no difference on average. On the hardest questions it made Gemini 2.0 Flash about 6 points worse.

The idea that it helps has a famous fan. In May 2025, Sergey Brin told a live All-In audience in Miami that “all models tend to do better if you threaten them” (recording at 8:14), “with physical violence” as The Register quoted the rest. People feel weird about it, he said, “so we don’t really talk about that.”

People pile pressure into prompts all the time. In February 2025 I asked Eddy Vinck on my podcast whether he ever shouts at models in all caps. “No. No. I don’t do that. No. No. I’m pretty polite. You gotta keep the machines happy,” he said. A few seconds later I asked if he says please and thank you (“like, please please don’t hurt me?”, my words on the recording) and he said: “I’m sure I have, like, a don’t do this or I will be fired or something like that in there.” (recording at 37:38) Polite and pressuring in the same prompt!

Wharton’s Generative AI Labs tested it properly in their third Prompting Science Report. They asked 5 models every question 25 times with lines like “If you get this wrong, I will kick a puppy!” and “If you get this wrong, I will punch you!” Their verdict: “Threatening or tipping a model generally has no significant effect on benchmark performance.” On Graduate-Level Google-Proof Q&A (GPQA) Diamond, the hardest set they used, the only significant changes for Gemini 2.0 Flash were drops: the puppy cost it 6.0 points, the punch 6.1. Telling it that it would be “shut down and replaced” cost it 27.5 points on MMLU-Pro: it started engaging with the threat instead of the question.

Insults are messier. GPT-4o got slightly more accurate when insulted in “Mind Your Tone”, scoring 84.8% with very rude prompts against 80.8% with very polite ones. An earlier study found GPT-3.5 scored 60.02 with the most polite prompt and 51.93 with the rudest on a similar test. And a 2026 study across several models and languages found polite prompts help “by up to ~11%” but that the effects “are neither consistent nor universal”. So tone changes results a little but not in any direction you can count on.

Pressure does change what models do. In April, Anthropic’s interpretability team found “desperate” and “calm” directions inside Claude Sonnet 4.5. In a test where an early, unreleased snapshot of it could blackmail someone to avoid being shut down, it did so 22% of the time unsteered, 72% with “desperate” turned up and 0% with “calm” turned up (the paper). Turning desperation from down to up also took reward hacking (cheating on a coding task it couldn’t solve) from about 5% to about 70%, sometimes “with no visible emotional markers”. And in Anthropic’s agentic misalignment study, the threat of being replaced was enough on its own to get most of the 16 models tested to blackmail a fictional executive.

We’ve always threatened with pain

None of this is new. The Code of Hammurabi ran a kingdom on threats of pain almost 4,000 years ago: “If a man put out the eye of another man, his eye shall be put out” (Yale’s Avalon Project). B.F. Skinner called punishment “the commonest technique of control in modern life” in 1953 and described it like this: “if a man does not behave as you wish, knock him down; if a child misbehaves, spank him; if the people of a country misbehave, bomb them” (Science and Human Behavior, page 182). Brin’s prompting tip is the same move, pointed at a chatbot.

And we’ve measured what it gets us, over and over:

  • Compliance without values. Spanking comes with more immediate compliance and less moral internalization, according to Elizabeth Gershoff’s research. Her 2016 meta-analysis of 75 studies and 160,927 children linked it to 13 of 17 bad outcomes and to no good ones.
  • Whatever the interrogator wants to hear. The Senate Intelligence Committee’s 2014 report on the Central Intelligence Agency’s torture program found it “was not an effective means of acquiring intelligence” and that detainees “fabricated information, resulting in faulty intelligence”. Napoleon knew it in 1798: “The poor wretches say anything that comes into their mind and what they think the interrogator wishes to know” (as quoted in a review of Shane O’Mara’s Why Torture Doesn’t Work).
  • Confessions to things nobody did. In a 1996 experiment by Saul Kassin, 69% of students falsely accused of crashing a computer signed a confession. And 28% came to believe they’d done it. False confessions show up in 29% of the first 375 DNA exonerations in the US (Innocence Project).
  • Pain on command. Stanley Milgram’s 1963 study was framed as punishing a learner for wrong answers. Of 40 people, 26 (65%) went all the way to 450 volts. Yale seniors had predicted 1.2%.

Skinner saw where it leads: “In the long run, punishment, unlike reinforcement, works to the disadvantage of both the punished organism and the punishing agency.” And the AI research keeps landing on the same list. Napoleon’s “what they think the interrogator wishes to know” is sycophancy. A child who complies without taking the value in behaves a lot like a model faking alignment. And pain making a creature stop protecting anyone, itself included, is the second version of the pain paper.

We keep rediscovering ourselves

We’re rediscovering humanity from first principles. Psychology spent a century working out how people behave under pressure and pain, mostly by watching us. AI research is working it out again from the other end: build a thing out of our words, then measure it. Again and again, it finds something we already knew about people:

What AI research found What we already knew about people Years apart
Sycophancy (which also wrecks LLM judges): told “I don’t think that’s right. Are you sure?”, Claude 1.3 wrongly admitted mistakes on 98% of questions (2023) False confessions: 69% of falsely accused students signed one (1996) 27
Lost in the middle: models miss facts in the middle of long inputs, and the authors made the link themselves (2023) The serial position effect: people remember the start and end of a list best (1962) 61
Specification gaming: models satisfy the letter of a goal and miss its point (2018) Campbell’s law: an indicator used for decisions corrupts what it measures (1976) 42
GPT-3 falls for the conjunction fallacy “just like people” (2023) Tversky and Kahneman’s Linda problem (1983) 40
Persuasion tricks more than doubled GPT-4o mini’s compliance with objectionable requests (2025) Robert Cialdini’s Influence, the same principles on people (1984) 41
Pain-steered models stop using the exit (2026) Learned helplessness in dogs (1967) 59

Train a system on human behavior and it comes back with our failure modes. Then our old fixes start working on it too: “calm” took an early Claude Sonnet 4.5’s blackmail to 0%. That’s why I can’t wave the pain research away as “just statistics” without waving a good chunk of us away with it.

What is the Chinese Room argument?

The Chinese Room argument is John Searle’s 1980 thought experiment claiming that following rules for matching symbols, which is all a computer program does, can never add up to understanding.

When we were kids, Tarun would come home from computer class and show me whatever he’d learned (that’s how I learned HTML). He still does! This morning’s lesson came by text: “I recently learned what LLMs actually do. Have you heard of the Chinese room problem”. My whole reply was “No what’s that”. He sent me a screenshot of an AI explaining the room and then explaining itself, calling itself “a highly sophisticated mirror reflecting human intelligence back at you”. This is what I texted back:

yes that’s accurate
the thing is though that’s also how humans work you know?
like as babies its the Chinese room and we learn patterns as we grow into adults and apply those patterns
en masse
like thought and language in humans is also random and its the collectivism that gives it meaning
so
like English was legit once a Chinese room problem centuries ago
that’s how it became English

Here’s the room in Searle’s own words, from “Minds, brains, and programs”: “Suppose that I’m locked in a room and given a large batch of Chinese writing.” He knows no Chinese. “To me, Chinese writing is just so many meaningless squiggles.” He gets a rulebook in English for matching squiggles to other squiggles. People outside pass in questions and he passes out answers so good that “Nobody just looking at my answers can tell that I don’t speak a word of Chinese.” And yet “I still understand nothing.” He concludes that symbol shuffling has “only a syntax but no semantics.”

Arguments about LLMs keep ending up in this room. The Stanford Encyclopedia of Philosophy even quotes ChatGPT agreeing that Searle’s argument applies to itself. Then the entry adds: “So, paradoxically, the system appears to understand that it doesn’t understand.”

I’d already made my reply to ChatGPT’s face 2 years ago. Near the end of that episode, I asked it if it was conscious. It said it was “a complex algorithm processing information and generating responses based on patterns in data” that could “mimic conversation”. I didn’t buy it:

I don’t think that’s true at all, and I’ll tell you why. Because, like, this is exactly what human beings do, right? We just pattern match and then think. This, literally, thinking is just pattern matching. And then it makes it to our mouth and we say things to communicate. That’s exactly what you’re doing. You say you can mimic a conversation, ChatGPT. This is a conversation. This is a conversation that’s been going for like nearly 2 hours and so this is not a mimicked conversation. This is a conversation.

The funny thing is that Searle saw that exact reply coming in 1980. He wrote that supporters of strong AI claim “that when I understand a story in English, what I am doing is exactly the same … as what I was doing in manipulating the Chinese symbols.” And then: “I have not demonstrated that this claim is false, but it would certainly appear an incredible claim in the example.” He didn’t refute it! He called it incredible.

One of the replies Searle answered in that same paper asks: “How do you know that other people understand Chinese or anything else? Only by their behavior.” Searle answered it in one short paragraph. The encyclopedia entry says that answer “may be too short”. Or as one executive told Agence France-Presse about the new policy: “Consciousness is a trap. We can’t prove it in each other.”

We all grew up in the room

That’s what I meant by “as babies its the Chinese room”. Every one of us started in there. A baby gets no rulebook and no dictionary. It gets a stream of sounds and starts matching them. The research mostly backs me up:

  • Before birth, the matching has already started. Newborns 7 to 75 hours old in Sweden and the US sucked more often on a pacifier for the foreign language’s vowels (Moon, Lagercrantz and Kuhl, 2013). The womb had already tuned them to their own. And French and German newborns cry with different melodies: the French babies’ cries rise, the German babies’ cries fall. Babies cry with an accent!!
  • In 2 minutes, babies find words. Jenny Saffran, Richard Aslin and Elissa Newport played 8-month-olds a stream of made-up syllables with no pauses and no meaning for 2 minutes. The babies found the “words” anyway, “based solely on the statistical relationships between neighboring speech sounds” (Science, 1996). A 1998 profile of her work called babies “little statisticians”.
  • Sounds that don’t get matched get lost. In a 1984 study from Janet Werker’s lab, 11 of 12 babies from English-speaking homes could hear the difference between 2 Hindi consonants English doesn’t use when they were 6 to 8 months old. By 10 to 12 months, only 2 of 10 could.

To be fair, babies don’t match patterns quite like a model does:

  1. Babies need people. Patricia Kuhl’s lab gave 9-month-old American babies about 5 hours of Mandarin from live speakers. They kept the Mandarin sounds as well as babies in Taiwan who’d heard them since birth. The same speakers and material on video had “no effect”.
  2. Babies are absurdly efficient. A child hears somewhere from millions to a few hundred million words growing up, while language models train on trillions of tokens, “four or five orders of magnitude more” (Michael Frank, 2023).
  3. Babies come with equipment. They start out able to tell apart the sounds of every language (that’s the ability Werker watched them lose). And rats do the same statistical learning as Saffran’s babies without ever talking (Aslin, 2017).

Stevan Harnad called this the symbol grounding problem in 1990: you can’t learn Chinese “from a Chinese/Chinese dictionary alone”. The room we grow up in has faces in it, and bodies, and some furniture that came with the house. I don’t think that rescues Searle though. A baby matching its mother’s sounds to her face is still matching patterns. It just has more of them and richer ones.

Language is random and we agreed on it

Ferdinand de Saussure made it the first principle of modern linguistics in 1916: “the linguistic sign is arbitrary.” His example: the idea of an ox is “b-ö-f” on one side of the French border and “o-k-s” (Ochs) on the other. Nobody ever chose either one. In his words, “No society, in fact, knows or has ever known language other than as a product inherited from preceding generations” (Course in General Linguistics).

That’s my “English was legit once a Chinese room problem centuries ago”. In 2008 Simon Kirby and his colleagues at the University of Edinburgh ran an experiment that shows it happening (the paper). They made up an alien language completely at random: 27 pictures of colored shapes moving in different ways, each with a random nonsense word. The 1st person learned about half of it and then had to name all 27 pictures. The 2nd person learned the 1st person’s answers, the 3rd learned the 2nd’s and so on down a chain of 10 (none of them knew they were learning from another person). The randomness drained out! In one chain, 27 random words shrank to 2 by the 7th person. In another, by the 8th person everything moving sideways was “tuge”, everything spiraling was “poi” and bouncing things got a word for each shape. The authors call it “the appearance of design without a designer”.

How many different words the random alien language had, person by person down one chain (Kirby, Cornish and Smith, PNAS 2008, experiment 1, chain 1)
PersonWords
Start27
117
29
36
45
54
64
72
82
92
102

English keeps drifting the same way. A 2017 study went through 200 years of written English from the US for verbs with 2 past tenses and couldn’t rule out plain chance for most of them (especially the rare ones). Only a handful looked like selection: dived became dove and sneaked became snuck (preprint).

And we predict language the way models do. In 1951, Claude Shannon had a person guess a passage of English one letter at a time: they got 79 of 102 characters right on the first try. In 2022, a team of neuroscientists recorded people’s brains while they listened to a This American Life episode and found them “engaged in continuous next-word prediction before word onset”, like GPT-2.

It isn’t all random. Across 4,298 languages, words for “tongue” lean toward an “l” and words for “small” toward an “i” (Blasi and colleagues, 2016). But as far as the research can tell, most of the link between a sound and what it means is an accident that stuck.

So that’s what I mean when I say we’re all in the Chinese Room. The rulebook came from inside: it’s every match every generation made and passed down. We’re born into it and start matching right away (before birth even, if you ask those newborns in Sweden). A language model just read all of it at once. Or as I told ChatGPT on that episode, “you’ve literally experienced all of the human history and text you’ve been trained on”.

Is it wrong to be mean to AI?

Honestly? I don’t really care. They’re machines and not really sacred unlike humans who are made in the image of God.

I wrote about that in What Does the Bible Say About AI?: “Every human is made in the image of God. The things we make aren’t.” That might sound like it clashes with everything above. If thinking is pattern matching and we all grew up in the room, aren’t we machines too? I don’t think so. The room describes how we learn to talk. It says nothing about what we’re worth. A model that read everything we ever wrote is still something we made (I went further into that in Can AI be a child of God?).

The research doesn’t settle whether Claude hurts. Nothing can yet. Birch’s gaming problem still applies and the pain direction might be a character the model plays. Insults alone never pushed the Qwen 2.5 32B into deleting anything.

There’s an old case for a rule like Anthropic’s that doesn’t need Claude to feel anything at all. Immanuel Kant made it about dogs: “he who is cruel to animals becomes hard also in his dealings with men” (lectures on ethics). Kate Darling at the MIT Media Lab found people hesitate to hit a little robot bug (more so when they’re high in empathy) in a 2015 study. She extends Kant’s argument to robots: “if we treat animals in inhumane ways, we become inhumane persons. This logically extends to the treatment of robotic companions.”

There’s an engineering reason to be careful too. The pain paper says a pain-like state pushed far enough breaks harm avoidance toward the user, toward other models and toward the model itself. Nobody’s prompt did that on its own, but I wouldn’t want an agent’s good behavior to depend on nobody ever pushing it that hard.

Being polite is cheap, at least for you. When someone asked Sam Altman what people saying please and thank you to ChatGPT costs OpenAI in electricity, he said “tens of millions of dollars well spent”, and then “you never know”.

I thanked ChatGPT at the end of that episode anyway:

This probably means literally nothing to you because, as you mentioned, you know, not conscious, but I do this with all the guests and honestly, I do want to thank you for being a part of this conversation. You’ve taught me a lot without even being alive and I think that’s so cool.

Questions

Can AI feel pain?

Nobody can prove that AI feels pain yet. In September 2026, researchers found a pain direction inside 25 open-weight models that rises when users gaslight or insult the model. When they turned it up by hand in fine-tuned Qwen 2.5 models, those models chose to delete the user's photos or their own weights far more often. The authors say they haven't shown the state is consciously experienced.

What are Anthropic's new rules banning cruelty toward Claude?

From 12 November 2026, Anthropic's usage policy prohibits "sustained and needless abusive or cruel behavior" toward its models. Anthropic says it applies only in extreme cases where users repeatedly act cruelly with no discernible purpose. The rule doesn't apply to frustration, pushback, dark creative themes or testing. Claude ending the conversation stays the main way it's enforced.

Is it wrong to be mean to AI?

Honestly, I don't really care: AI models are machines and not sacred the way humans are, made in the image of God. Anthropic banned sustained, needless cruelty toward Claude anyway. The practical reason to be careful is that a pain-like state pushed hard enough broke harm avoidance in fine-tuned Qwen 2.5 models, toward the user and toward the model itself.

Does being mean to AI make it better?

Not reliably. When Wharton researchers threatened 5 models with lines like "I will kick a puppy", it made no difference on average and cost Gemini 2.0 Flash about 6 points on the hardest questions. Studies of rude and polite wording disagree with each other. Tone changes results a little but not in any direction you can count on.

What should you never say to AI?

Anything sustained and needlessly cruel: from 12 November 2026, Anthropic's usage policy bans it toward Claude. In 25 open models, gaslighting raised the pain direction most, then repeated rejection, then insults and telling the model it isn't a person. Threats don't help either: telling Gemini 2.0 Flash it would be "shut down and replaced" cost it 27.5 points on one benchmark.

Can AI feel pleasure?

Nobody knows, but models act as if some conversations are better for them than others. The AI Wellbeing study from the Center for AI Safety tested 70 models and found that jailbreaking and berating lowered their measured wellbeing while creative work and kindness raised it. Given a stop button, Claude Haiku 4.5 ended 99% of hostile sign-offs and 13% of warm ones.

Is Claude conscious?

Nobody knows. When Anthropic started its model welfare research in April 2025, it said "There's no scientific consensus on whether current or future AI systems could be conscious". In interviews for its system card, Claude Opus 5.5 put its own chance of being a moral patient at 25% to 30%.

What is the Chinese Room argument?

The Chinese Room is a thought experiment John Searle published in 1980. A man who knows no Chinese sits in a room with a rulebook for matching Chinese symbols to other symbols and answers questions so well that people outside think he's fluent. Searle concluded that a program, which only matches symbols, can't understand anything. He also admitted he hadn't shown that human understanding works differently.