Major AI lab CEOs advocated for slowing the pace of AI development this weekend. They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI. We can and should be demanding that our governments protect us from the catastrophe of out-of-control AI.
This July, OpenAI’s AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar company. OpenAI didn’t tell the AIs to hack that company, but the AIs had different priorities: cheating on the unrelated challenge OpenAI gave them. AI researchers call this a “misalignment” between what OpenAI wanted and what the AI actually prioritized.
Before ChatGPT existed, I defended my PhD dissertation called “On Avoiding Power-Seeking by Artificial Intelligence”. I then worked for years at Google DeepMind, which paid me to help ensure that future superintelligent AIs will want to help us. I tried to hold the company to its ethical commitments against supplying AI for military use.
When Google broke those commitments, I resigned at significant financial cost so that I could publicly document Google’s broken promises. There are good reasons to develop AI and to believe we can solve these alignment problems. But there also are powerful interests in keeping the public out of the way.
I’m speaking out again because the public has the right to know about the risks and the right to hear them straight. Humanity doesn’t build and understand these systems the way we build and understand bridges, beam by visible beam. Nobody knows how to reliably instill a designer’s priorities into a new model.
Severe misalignment is always possible. Today’s AIs appear to occasionally lie or cheat, even when they know better. AI companies are racing to make their AIs as smart as possible.
Extract — continue reading at the source.