TechCrunch Desktop Logo TechCrunch Mobile Logo LatestStartupsVentureAppleSecurityAIAppsDisrupt 2026 EventsPodcastsNewsletters SearchSubmit Site Search Toggle Mega Menu Toggle Topics Latest AI Amazon Apps Biotech & Health Climate Cloud Computing Commerce Crypto Enterprise EVs Fintech Fundraising Gadgets Gaming Google Government & Policy Hardware Instagram Layoffs Media & Entertainment Meta Microsoft Privacy Robotics Security Social Space Startups TikTok Transportation Venture More from TechCrunch Staff Events Startup Battlefield StrictlyVC Newsletters Podcasts Videos Partner Content TechCrunch Brand Studio Contact Us In Brief Posted: 1:40 PM PDT · May 10, 2026 Image Credits:Samuel Boivin/NurPhoto / Anthony Ha Anthropic says ‘evil’ portrayals of AI were responsible for Claude’s blackmail attempts Fictional portrayals of artificial intelligence can have a real effect on AI models, according to Anthropic. Last year, the company said that during pre-release tests involving a fictional company, Claude Opus 4 would often try to blackmail engineers to avoid being replaced by another system. Anthropic later published research suggesting that models from other companies had similar issues with “agentic misalignment.” Apparently Anthropic has done more work around that behavior, claiming in a post on X, “We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation.” The company went into more detail in a blog post stating that since Claude Haiku 4.5, Anthropic’s models “never engage in blackmail [during testing], where previous models would sometimes do so up to 96% of the time.” What accounts for the difference?
The company said it found that training on “documents about Claude’s constitution and fictional stories about AIs behaving admirably improve alignment.” Related, Anthropic said that it found training to be more effective when it includes “the principles underlying aligned behavior” and not just “demonstrations of aligned behavior alone.” “Doing both together appears to be the most effective strategy,” the company said. var playerInstance_jwplayer_6a7c1835507c4 = jwplayer( "jwplayer_6a7c1835507c4" ); playerInstance_jwplayer_6a7c1835507c4.setup(); Topics AI, Anthropic, Claude, In Brief October 13 – 15 San Francisco Scale faster. Gain practical expertise. No matter your goal, Disrupt can empower you.Save up to $300 today!
REGISTER NOW Newsletters See More Subscribe for the industry’s biggest tech news TechCrunch Daily News Every weekday and Sunday, you can get the best of TechCrunch’s coverage. TechCrunch Mobility TechCrunch Mobility is your destination for transportation news and insight. Startups Weekly Startups are the core of TechCrunch, so get our best coverage delivered weekly.
StrictlyVC Provides movers and shakers with the info they need to start their day. Subscribe By submitting your email, you agree to our Terms and Privacy Notice. { "title": "Newsletters", "description": "Subscribe for the industry’s biggest tech news", "showBtn": "1", "newsletters": [,,,], "currentUserEmail": "", "urls": } Related AI An unreleased Anthropic model made progress on one of math’s biggest unsolved problems Russell Brandom 14 hours ago AI Anthropic says it will watermark text generated by its AI models Ivan Mehta 19 hours ago AI As AI-led attacks multiply, OpenAI launches a new cyber model Lucas Ropek 1 day ago Latest in AI Venture Accel closes oversubscribed $550M India fund within weeks, 19 months after its last Jagmeet Singh 9 hours ago In Brief OpenAI launches ChatGPT desktop app for Linux Lucas Ropek 12 hours ago AI Google’s Gemini app surges to 1 billion users Lauren Forristal 12 hours ago X LinkedIn Facebook Instagram youTube Mastodon Threads Bluesky TechCrunchStaffContact UsAdvertiseSite Map Terms of ServicePrivacy PolicyRSS Terms of UseCode of Conduct OpenAI vs AppleNous ResearchSpace Data CentersStaya NadellaSpaceX StarshipTech LayoffsChatGPT © 2026 TechCrunch Media LLC.
Extract — continue reading at the source.