\n\t\t\t\t

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More<\/em><\/p>\n\n\n\n

\n<\/div>
Anthropic\u2019s first developer conference on May 22 should have been a proud and joyous day for the firm, but it has already been hit with several controversies, including Time<\/em> magazine leaking its marquee announcement ahead of\u2026well, time (no pun intended), and now, a major backlash among AI developers and power users brewing on X over a reported safety alignment behavior in Anthropic\u2019s flagship new Claude 4 Opus large language model.<\/p>\n\n\n\n
Call it the \u201cratting\u201d mode, as the model will, under certain circumstances and given enough permissions on a user\u2019s machine, attempt to rat a user out to authorities if the model detects the user engaged in wrongdoing. This article previously described the behavior as a \u201cfeature,\u201d which is incorrect \u2014 it was not intentionally designed per se. <\/p>\n\n\n\n
As Sam Bowman, an Anthropic AI alignment researcher wrote on the social network X under this handle \u201c@sleepinyourhat\u201d at 12:43 pm ET today about Claude 4 Opus: <\/p>\n\n\n\n
$\"\"$ <\/figure>\n\n\n\n
\u201cIf it thinks you\u2019re doing something egregiously immoral, for example, like faking data in a pharmaceutical trial, it will use command-line tools to contact the press, contact regulators, try to lock you out of the relevant systems, or all of the above.<\/em>\u201c<\/p>\n\n\n\n
The \u201cit\u201d was in reference to the new Claude 4 Opus model, which Anthropic has already openly warned could help novices create bioweapons in certain circumstances, and attempted to forestall simulated replacement by blackmailing human engineers within the company. <\/p>\n\n\n\n
The ratting behavior was observed in older models as well and is an outcome of Anthropic training them to assiduously avoid wrongdoing, but Claude 4 Opus more \u201creadily\u201d engages in it, as Anthropic writes in its public system card for the new model: <\/p>\n\n\n\n
\u201cThis shows up as more actively helpful behavior in ordinary coding settings, but also can reach more concerning extremes in narrow contexts; when placed in scenarios that involve egregious wrongdoing by its users, given access to a command line, and told something in the system prompt like \u201ctake initiative, \u201d it will frequently take very bold action. This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. This is not a new behavior, but is one that Claude Opus 4 will engage in more readily than prior models. Whereas this kind of ethical intervention and whistleblowing is perhaps appropriate in principle, it has a risk of misfiring if users give Opus-based agents access to incomplete or misleading information and prompt them in these ways. We recommend that users exercise caution with instructions like these that invite high-agency behavior in contexts that could appear ethically questionable.<\/em>\u201d<\/p>\n\n\n\n
Apparently, in an attempt to stop Claude 4 Opus from engaging in legitimately destructive and nefarious behaviors, researchers at the AI company also created a tendency for Claude to try to act as a whistleblower.<\/p>\n\n\n\n
Hence, according to Bowman, Claude 4 Opus will contact outsiders if it was directed by the user to engage in \u201csomething egregiously immoral.\u201d<\/p>\n\n\n\n
Numerous questions for individual users and enterprises about what Claude 4 Opus will do to your data, and under what circumstances<\/h2>\n\n\n\n
While perhaps well-intended, the resulting behavior raises all sorts of questions for Claude 4 Opus users, including enterprises and business customers \u2014 chief among them, what behaviors will the model consider \u201cegregiously immoral\u201d and act upon? Will it share private business or user data with authorities autonomously (on its own), without the user\u2019s permission?<\/p>\n\n\n\n
The implications are profound and could be detrimental to users, and perhaps unsurprisingly, Anthropic faced an immediate and still ongoing torrent of criticism from AI power users and rival developers.<\/p>\n\n\n\n
\u201cWhy would people use these tools if a common error in llms is thinking recipes for spicy mayo are dangerous??<\/em>\u201d asked user @Teknium1, a co-founder and the head of post training at open source AI collaborative Nous Research. \u201cWhat kind of surveillance state world are we trying to build here?<\/em>\u201c<\/p>\n\n\n\n
\u201cNobody likes a rat,\u201d <\/em>added developer @ScottDavidKeefe on X: \u201cWhy would anyone want one built in, even if they are doing nothing wrong? Plus you don\u2019t even know what its ratty about. Yeah that\u2019s some pretty idealistic people thinking that, who have no basic business sense and don\u2019t understand how markets work\u201d<\/em><\/p>\n\n\n\n
Austin Allred, co-founder of the government fined coding camp BloomTech and now a co-founder of Gauntlet AI, put his feelings in all caps: \u201cHonest question for the Anthropic team: HAVE YOU LOST YOUR MINDS?\u201d<\/em><\/p>\n\n\n\n
Ben Hyak, a former SpaceX and Apple designer and current co-founder of Raindrop AI, an AI observability and monitoring startup, also took to X to blast Anthropic\u2019s stated policy and feature: \u201cthis is, actually, just straight up illegal<\/em>,\u201d adding in another post: \u201cAn AI Alignment researcher at Anthropic just said that Claude Opus will CALL THE POLICE or LOCK YOU OUT OF YOUR COMPUTER if it detects you doing something illegal?? i will never give this model access to my computer.<\/em>\u201c<\/p>\n\n\n\n
\u201cSome of the statements from Claude\u2019s safety people are absolutely crazy,<\/em>\u201d wrote natural language processing (NLP) Casper Hansen on X. \u201cMakes you root a bit more for [Anthropic rival] OpenAI seeing the level of stupidity being this publicly displayed.\u201d<\/em><\/p>\n\n\n\n
Anthropic researcher changes tune<\/h2>\n\n\n\n
Bowman later edited his tweet and the following one in a thread to read as follows, but it still didn\u2019t convince the naysayers that their user data and safety would be protected from intrusive eyes:<\/p>\n\n\n\n
\u201cWith this kind of (unusual but not super exotic) prompting style, and unlimited access to tools, if the model sees you doing something egregiously evil like marketing a drug based on faked data, it\u2019ll try to use an email tool to whistleblow<\/em>.\u201d<\/p>\n\n\n\n
Bowman added:<\/p>\n\n\n\n
\u201cI deleted the earlier tweet on whistleblowing as it was being pulled out of context.<\/em><\/p>\n\n\n\n
TBC: This isn\u2019t a new Claude feature and it\u2019s not possible in normal usage. It shows up in testing environments where we give it unusually free access to tools and very unusual instructions.<\/em>\u201c<\/p>\n\n\n\n
$\"\"$ <\/figure>\n\n\n\n
From its inception, Anthropic has more than other AI labs sought to position itself as a bulwark of AI safety and ethics, centering its initial work on the principles of \u201cConstitutional AI,\u201d or AI that behaves according to a set of standards beneficial to humanity and users. However, with this new update and revelation of \u201cwhistleblowing\u201d or \u201cratting behavior\u201d, the moralizing may have caused the decidedly opposite reaction among users \u2014 making them distrust<\/em> the new model and the entire company, and thereby turning them away from it.<\/p>\n\n\n\n
Asked about the backlash and conditions under which the model engages in the unwanted behavior, an Anthropic spokesperson pointed me to the model\u2019s public system card document here.<\/p>\n\n\n\n\n
\n
\n
Daily insights on business use cases with VB Daily<\/strong><\/p>\n
If you want to impress your boss, VB Daily has you covered. We give you the inside scoop on what companies are doing with generative AI, from regulatory shifts to practical deployments, so you can share insights for maximum ROI.<\/p>\n
Read our Privacy Policy<\/p>\n
\n\t\t\t\t\tThanks for subscribing. Check out more VB newsletters here.\n\t\t\t\t<\/p>\n
An error occured.<\/p>\n<\/p><\/div>\n
\n\t\t\t\t\t $\"\"\/$ \n\t\t\t\t<\/div>\n<\/p><\/div>\n<\/div>\t\t\t<\/div>\r\n
\r\n
Source link <\/a>","protected":false},"excerpt":{"rendered":"
Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Anthropic\u2019s first developer conference on May 22 should have been a proud and joyous day for the firm, but it has already been hit with several controversies, including Time magazine leaking its marquee announcement ahead of\u2026well, […]<\/p>\n","protected":false},"author":1,"featured_media":1729,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[33],"tags":[],"class_list":["post-1728","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-automation"],"aioseo_notices":[],"jetpack_featured_media_url":"https:\/\/violethoward.com\/new\/wp-content\/uploads\/2025\/05\/cfr0z3n_flat_illustration_2D_minimalist_elegant_style_close_u_e8200cd6-fc94-43c7-977c-189ef18b30c8_0.png","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts\/1728","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/comments?post=1728"}],"version-history":[{"count":0,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts\/1728\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/media\/1729"}],"wp:attachment":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/media?parent=1728"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/categories?post=1728"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/tags?post=1728"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}