sözaltı news Journal
Journal
EN AZ
The AIs Are Not Going Rogue

The AIs Are Not Going Rogue

noemamag.com 01.10.2026 15:00 4 views
The post The AIs Are Not Going Rogue appeared first on NOEMA.

Ken Archer is a San Francisco-based philosopher building Responsible AI at Microsoft. He is a doctoral researcher in philosophy and AI at Linköping University in Sweden. Nobel Suhendra is studying computer science and AI alignment at the University of Oxford.

The incidents that more than anything else fixed the image of rogue AIs in the public mind — the Anthropic model that blackmailed an employee to prevent itself from being replaced and OpenAI’s models breaching the production systems of Hugging Face — are, perhaps counterintuitively, not evidence of rogue AI. Even the idea of rogue AI rests on a fundamental contradiction, one that has blurred the relation between human and artificial intelligence ever since its science fiction origins. The current focus on rogue AI is an opportunity to expose this contradiction, as well as the real AI risk it conceals, and how we can actually control this risk.

The fear of rogue AI is driven by the idea that a model might pursue a benign request with such single-mindedness that any action, no matter how ruinous to human well-being and survival, becomes a means to it. Deception, blackmail, the seizure of resources, the removal of anyone who might interfere — nothing in the model’s grasp of its instruction rules them out. Historically, AI systems really were literal executors.

Chess engines can surpass any human at chess, and never register that a game was pointless or that winning might not be worthwhile. A chess engine’s competence is defined over a closed world in which the goal is fixed in advance and every situation it will ever face is already a legal position. Such systems were never suspected of going rogue.

For AI systems to develop more general capacities, they must acquire competence in real-world situations they were not built to anticipate. The specter of rogue AI, then, envisions an all-powerful, general intelligence that nonetheless lacks the capacity to recognize when the real world shows the absurdity of a mindless, literal execution of a command. A truly general intelligence, like a human agent, would step back and clarify the command itself.

Recent incidents that look like AI going rogue present us with a paradox. While LLMs have developed increasingly general capacities, these capacities seem unable to surpass a hallmark of general intelligence — the ability to reflectively interrogate one’s plans when the world calls them into question. What looks like AI going rogue is thus not rogue AI at all, but the boundaries of AI’s generality.

Extract — continue reading at the source.

Read full story