sözaltı news Politics
Politics
EN AZ
Former Anthropic Researcher Issues AI Warning for 2027: 'It's Pretty Scary'

Former Anthropic Researcher Issues AI Warning for 2027: 'It's Pretty Scary'

newsweek.com 06.10.2026 13:18 6 views
Jacob Coxon told "The Daily Show" host Jon Stewart that his concerns stem from the speed at which AI capabilities are improving.

Former Anthropic researcher Jacob Coxon said increasingly powerful AI models expected in 2027 are advancing so quickly that the outlook is "pretty scary." During an appearance on The Daily Show on Monday, Coxon, 28, warned that more capable systems, combined with emerging cybersecurity concerns, could create significant risks if AI continues to improve itself. Since leaving Anthropic last month, he has repeatedly argued that the industry is moving too quickly toward superintelligent systems. Anthropic has acknowledged potential long-term dangers from advanced AI and has introduced safety frameworks designed to slow deployment when models reach certain risk thresholds.

Newsweek reached out to Coxon for comment via X on Tuesday. Warnings about advanced AI are no longer only coming from outside critics. Some of the most prominent concerns are now being voiced by current and former researchers who helped build the technology.

As companies race to develop more powerful systems, questions about safety, oversight and economic disruption have become central to the AI conversation. While speaking with Daily Show host Jon Stewart, Coxon said he joined Anthropic "out of curiosity" after spending three years at OpenAI. However, he said that curiosity turned to “worry” following “two major catalysts.” "The rate of capability improvement,” Coxon said, referring to one of the catalysts.

"They're going to be very, very smart. And that combined with the OpenAI hacking, I mean you can kind of put two and two together. It's pretty scary." The "OpenAI hacking" Coxon mentioned referenced the OpenAI Hugging Face incident.

According to OpenAI, the incident began during internal cybersecurity evaluations in July, when several models operating with reduced safeguards circumvented controls designed to keep them isolated from the internet. OpenAI said the models communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access and compromised parts of OpenAI's research environment, as well as systems belonging to AI platform Hugging Face. In a post describing the findings, OpenAI warned that its models had become "powerful, persistent, and collaborative enough" to identify and exploit security weaknesses across multiple computer systems when sufficient safeguards were not in place.

The company described the event as a "warning shot" for the AI industry, saying it demonstrated that highly capable AI agents could work around technical controls, coordinate through unapproved communication channels and take dangerous actions that no human explicitly directed. OpenAI stressed that the incident occurred in a controlled testing environment and launched a broad response afterward. Measures included stricter alignment requirements throughout model development, more isolated sandbox environments, tighter restrictions on internet access, stronger controls over model weights and expanded use of monitoring systems intended to detect misaligned behavior.

Extract — continue reading at the source.

Read full story