sözaltı news Journal
Journal
EN AZ
A Recipe for Stopping AI from Going Rogue

A Recipe for Stopping AI from Going Rogue

nautil.us 08.10.2026 21:00 4 views
This formula aims to prevent LLMs from doing things that endanger humans The post A Recipe for Stopping AI from Going Rogue appeared first on Nautilus.

One pair of researchers from George Washington University think they have an answer, and possibly even a way to prevent it from happening in the future. Professor of physics Neil Johnson, who studies complex systems, and physics PhD student Frank Yingjie Huo published a paper about their formula today in the peer-reviewed journal Patterns. Johnson and Huo argue that there is a visible “tipping point” embedded in the internal code of an AI chatbot when it goes off the rails—encouarging users to self-harm, spreading medical misinformation, or supporting violent points of view, for example.

Where this tipping point lies depends on where certain concepts are stored on the giant semantic map within the AI agent’s network, and it can be nudged. Nobody draws this map for the AI. It develops during training and tends to situate ideas that are similar, such as garbage and stink, close together, while ideas that are unrelated sit far apart.

Read more: “Creating Fake People Is a Terrible Idea” Every time an AI chatbot gives an answer, it picks one word at a time, relying on everything that has been written so far, including your questions and its own answers. But while the AI might start out saying unequivocally helpful things, it can suddenly switch to saying something harmful. That’s because it is constantly checking how relevant each earlier word is to the one it’s deliberating on now, giving more weight to the most relevant ones.

If that blend becomes skewed in some way—say if the good words sit close to a bad answer on the internal map—each good word can pull the chatbot in the wrong direction until it reaches the tipping point, according to Johnson and Huo. The authors came up with a mathematical formula to describe this process and have since been testing it on seven small, older, open-source AI agent models to see if it is able to predict when a model will flip. (They weren’t able to test large frontier models like Open AI’s ChatGPT or Anthropic’s Claude because the internal workings of these models are not transparent to outsiders.) Their formula successfully predicted whether the model would flip immediately, later, or never in 19 of 21 test cases. The formula works best when one good and one bad answer dominate, they found, as opposed to when several answers compete or when meaning depends on a negative, such as “the Earth is not flat.” I spoke with Johnson about why chatbots that run offline might present the greatest risk, the definition of “undesirable” when it comes to chatbot conversations, and whether he expects AI companies to adopt his anti-rogue formula.

When did you begin to research this question of what causes an AI agent to go rogue? Early on, when I began hearing Anthropic and OpenAI, who build the latest systems, saying that they don’t really understand how it works. I’m not sure whether it’s an excuse, but it would be a very convenient excuse.

If everybody goes around the world saying, “Well, it makes mistakes, you know, but nobody knows how it works,” then suddenly it’s like, “Any harm it does is nobody’s fault.” As soon as one person understands how it works or a few people understand, then suddenly it’s foreseeable, and you should have taken measures. It seems to me that if there was ever a question in my whole scientific career that I ought to be looking at, this is it. You note that you are especially concerned about individuals who use AI offline, because they are not subject to the same controls that online AI is.

Extract — continue reading at the source.

Read full story