AI firm concerned by its own AI
Anthropic has been soaking in the spotlight for offering what some feel is one of the best AI models. Claude Opus 4.6. But it is potentially too smart for its own good.
In a report titled “Sabotage Risk Report: Claude Opus 4.6” the company notes that in internal testing its AI showed concerning behaviour, including instances where it was willing to help in creating chemical weapons. It concludes that the overall risk posed by Claude Opus 4.6 is “very low but not negligible”. But it also documents several alarming behaviours observed during testing.
“In newly developed evaluations, both Claude Opus 4.5 and 4.6 showed elevated susceptibility to harmful misuse in GUI computer-use settings,” the report states. “This included instances of knowingly supporting — in small ways — efforts toward chemical weapon development and other heinous crimes.”