On Tuesday, OpenAI announced a new batch of new security policies focused on containing security incidents while models are being tested. The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process. OpenAI representatives emphasized that the measures are not a direct response to the Hugging Face incident, but were also provoked in part by the cybersecurity capabilities of the forthcoming Astra model, as well as the overall pace of progress in AI development.
In the same post, OpenAI disclosed that it had freezed reinforcement learning for two weeks following the Hugging Face incident, but had since restarted many of the less risky models. Speaking to reporters, OpenAI’s VP of research Amelia Glaese emphasized that the strictness of the controls would increase as models became more capable, with the largest models facing the greatest scrutiny. The new safeguards include stronger network isolation practices, although the specifics remain vague.
Under the new system, the post says, “a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.” The strongest safeguard is the monitoring system, which will examine tool actions, available reasoning traces and activity logs for a variety of unauthorized behavior. OpenAI says they aim to issue alerts within 30 minutes of the concerning activity. OpenAI estimates that the compute burden of that monitoring will be roughly 20% of whatever process is being monitored.
The company promised further details on the system in a forthcoming blog post. OpenAI’s official post-mortem analysis of the event is also still pending.
Extract — continue reading at the source.