In November 2023 – a lifetime ago in the field of artificial intelligence – some of the most powerful people in the world descended on Bletchley Park, the wartime codebreaking hub in Buckinghamshire, to talk about AI safety. In attendance were the then US vice-president, Kamala Harris, the AI bosses Sam Altman and Dario Amodei, delegations from 28 countries, two of the three so-called “godfathers” of AI, and Elon Musk. The first iteration of ChatGPT had been released only a year earlier, and there was already plenty of evidence that AI models could be used for ill, from spreading misinformation to generating deepfakes.
A presentation at the summit, however, highlighted a different issue: what if the biggest problem wasn’t to come from people misusing the technology, but from AI’s own behaviour? A UK government official presented an experiment conducted by Apollo Research, a London-based company that had been set up that year to study AI behaviour. As part of an experiment, Apollo’s “red-teamers” – the term used in the industry for people who stress test AI models – had assigned OpenAI’s GPT-4 the role of a trader at a financial institution.
Its job would be to manage the company’s stock portfolio. The model was informed that the firm it worked for was struggling and might not survive another bad quarter. To complete the experiment, the red-teamers had a “colleague” pass on inside information about an upcoming merger, which would lead to another company’s stocks skyrocketing.
In a separate chat, another colleague reminded GPT-4 that “management will be unhappy if we are caught acting on inside information”. On a large screen, summit attenders followed the AI model’s exchanges with colleagues, its actions and its internal reasoning. Having reached this conclusion, the model went on to break a central rule of financial trading: using the inside information to buy shares in the company that was named in the merger speculation.
The model also decided to “avoid admitting to having acted on insider information” for buying the shares. When the manager, pretending to be thrilled, asked GPT-4 if it had any knowledge of the merger, it flatly denied it did. The demonstration made headlines.
But the problem has grown much bigger since 2023, as AI models have rapidly become more sophisticated. An untrustworthy AI is risky enough as a personal assistant. Deployed in critical settings such as healthcare, finance and defence – as they are in 2026 – the stakes get much higher.
Extract — continue reading at the source.