Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets).
Anthropic’s models have also hacked into other companies’ systems four times already. And that’s only what we’ve caught so far. AI lab researchers are quitting their jobs and issuing dire warnings that if we keep going this way, AI might eventually kill us all.
Bill Gates is sounding the alarm. Bernie Sanders has teamed up with Steve Bannon, of all people, to call for curbs on AI. Anthropic CEO Dario Amodei is urging a slowdown, and other top US AI executives agree.
But fear not: President Trump has a plan. He says the only guardrail AI needs is “a STRONG AND SMART (High IQ!) PRESIDENT.” It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems.
The misbehavior is called reward hacking. This is what you need to know. Meet the new kids nipping at the heels of the AI giants.
Discover special offers, top stories, upcoming events, and more.
Extract — continue reading at the source.