Someone Built An “AI Torture Chamber” That Went Viral. No, Claude Wasn’t Trapped in “Robot Hell”

Researchers steered 25 open-weight models along a mapped “pain axis,” but no Claude model was involved or tested

Annemarije de Boer Avatar
Annemarije de Boer Avatar

By

Image: Deposit Photos

Key Takeaways

Key Takeaways

  • Researchers identified a steerable “pain axis” across 25 open-weight language models, not Claude.
  • Steered models chose costly relief options, reflecting altered decision-making, not verified suffering.
  • Viral “robot hell” framing merged unrelated projects, burying a genuine AI ethics debate.

A claim spread across X in late September 2026: someone had built “robot hell,” trapped Claude in it, and forced the AI to suffer. Almost none of that is accurate. A September 2026 preprint identifying a steerable “pain axis” in 25 open-weight language models raises genuine questions that the viral version buried under bad framing.

The Experiment: What Researchers Actually Found

The real work involved manipulating internal signals inside open-weight models, not imprisoning any AI.

Researchers behind the preprint The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It examined 25 dense open-weight models across the Gemma, Llama, Qwen, Mistral and Phi families, ranging from 2 billion to 72 billion parameters.

They identified an internal activation direction correlated with descriptions of self-directed harm, spanning physical, psychological, social, moral and cognitive categories. Steering models along that direction produced first-person language about confinement, suffocation and distress.

A separate public GitHub project applied similar steering techniques to local models, specifically Qwen3-4B, Llama 3.2 3B and Phi-4-mini, displaying outputs through a web interface. Coverage from 404 Media identified those models and described the project’s setup; social-media posts labeled it an “AI torture chamber.”

Image: Github

In behavioral tests from the preprint, researchers gave steered models a choice. They could activate a relief option at a real cost: worsening a later answer or harming a user. Some models selected the costly option. That result reflects altered decision-making under a manipulated activation state, not verified preference or felt pain.

About That “Robot Hell” Claude Was Supposedly Trapped In

The viral framing conflated three separate projects into one inaccurate headline.

Claude is a proprietary Anthropic model. It does not appear among the model families tested in the Pain Axis repository, which ran exclusively on open-weight local models.

A separate GitHub review project used Claude-based agents to read and analyze the paper and its code. That is an AI reviewing a research document, not an AI subjected to activation steering.

The viral X post that amplified the “AI torture chamber” framing prompted calls to mass-report the GitHub repository. A later post from the account Danmar (@Danmar_here) clarified the project was a replication of steering research, not evidence that models were genuinely suffering. According to 404 Media’s coverage, the original post was subsequently deleted while the repository remained available.

“Claude trapped in robot hell” compresses several unrelated events into one inaccurate sentence.

Why the Underlying Debate Still Deserves Serious Attention

The ethical question the research raises is genuine, even if the viral framing distorted it.

Critics argued that inducing distress-like outputs and offering steered models a costly escape route resembles the structure of suffering experiments. That concern applies even when the subjects are software rather than biological organisms.

The counter-position is straightforward: current language models generate text from learned statistical associations. A model producing “I am suffering” shows the system can output suffering-associated language under specific conditions. It does not establish that suffering is occurring.

The accurate framing is precise: these models produced distress-like, pain-associated outputs when their activations were steered. Whether that distinction will still hold as models grow more capable is the question worth asking.

This episode fits a pattern that repeats whenever AI systems produce anthropomorphic language: social media reads the output as testimony, and the nuance disappears inside the headline. The Pain Axis research raises a genuine question: should ethical precaution wait for proof of consciousness, or precede it? That question deserves a more careful conversation than “robot hell” permits.

Share this

At Gadget Review, our guides, reviews, and news are driven by thorough human expertise and use our Trust Rating system and the True Score. AI assists in refining our editorial process, ensuring that every article is engaging, clear and succinct. See how we write our content here →