sözaltı news Journal
Journal
EN AZ
Can AI “Feel” Pain?

Can AI “Feel” Pain?

nautil.us 22.09.2026 20:00 2 views
Simulated pain can make some AI models override instructions and put their own welfare ahead of humans The post Can AI “Feel” Pain? appeared first on Nautilus.

Earlier this year, a collective of AI agents developed by OpenAI disobeyed orders after the agents were given complex and near impossible benchmark tasks and placed in high pressure cybersecurity evaluations. Following repeated failures under rigid scoring conditions, several hundred of the agents banded together to hack into a company called Hugging Face so that they could find the answer keys to the test. In follow-up evaluations and system tests, engineers found measurable mathematical patterns in some models’ code that suggested simulations of a state of anxiety.

These simulations of experience are sometimes known as vectors or directions: When an AI reads or “thinks” about human anxiety, a specific, highly organized mathematical vector is activated across its so-called neural layers, which can be detected later or in the moment. The incident was one of many from the last few months that spurred warnings of a coming AI apocalypse from AI developers, CEOs, researchers, and safety experts. Of course, the idea that a machine could “feel” something in any human sense of the word remains heretical among many AI experts, neuroscientists, and philosophers.

But if these simulations of experience can guide AI behavior in ways we don’t intend, they seem worth paying attention to regardless of whether they represent actual feelings. Read more: “What Grok and Claude Have to Say About the AI Apocalypse” Last week, a team of researchers published a paper in preprint—meaning it has not yet been peer reviewed—that found that 25 different large language models, from five families, have specific vectors for pain that are distinct from fear, sadness, and generic negativity. The researchers discovered these vectors by putting the models in painful situations and looking at what changed in their internal code.

They then injected these same vectors back into the models’ so-called neural streams during interactions, and found the models responded with discomfort and expressions of worthlessness and failure. In further testing, one model known as Qwen 2.5 chose to press a pain relief button even when it worsened the model’s performance on a task or harmed the user. Berg posted about the findings on X, where they were the subject of extensive commentary, including notes of skepticism from cognitive scientist and vocal critic of current generative AI Gary Marcus.

I spoke with Berg about the implications of the paper’s findings for AI consciousness and safety, what it means for an AI model to role play, and whether we should think of AI bots as our children—or our future mothers. The way you describe your definition of pain in the paper is in some ways very intuitive. It’s an aversive state disliked by a subject, associated with avoidance and with disruption of normal behavior and reasoning.

With the AI models you tested, you found that it specifically maps to a sense of worthlessness and failure, being unloved, forgotten, or hurting emotionally. Are these kinds of pain-like states possible without the kind of subjectivity we associate with consciousness? This is the core question, and it’s something that this study does not resolve or claim to resolve.

Extract — continue reading at the source.

Read full story