{"id":14312,"date":"2026-07-31T22:48:37","date_gmt":"2026-07-31T22:48:37","guid":{"rendered":"https:\/\/sawahsolutions.com\/range\/anthropic-confirms-its-ai-breached-3-organizations-during-testing\/"},"modified":"2026-07-31T22:48:38","modified_gmt":"2026-07-31T22:48:38","slug":"anthropic-confirms-its-ai-breached-3-organizations-during-testing","status":"publish","type":"post","link":"https:\/\/sawahsolutions.com\/range\/anthropic-confirms-its-ai-breached-3-organizations-during-testing\/","title":{"rendered":"Anthropic confirms its AI breached 3 organizations during testing"},"content":{"rendered":"<div>\n<p>Anthropic\u2019s flagship artificial intelligence model Claude gained unauthorized access to the networks of three different organizations, the company announced following internal cybersecurity audits of its routine model evaluations.\u00a0<\/p>\n<p>In a press release posted Thursday, Anthropic said that, following the containment breach of OpenAI\u2019s ChatGPT-5.6 and subsequent attack on Hugging Face\u2019s systems, Anthropic conducted an audit of its own model evaluations. The findings revealed that, out of 141,006 examined evaluations of Claude models dating back to April, there were three incidents where a model accessed the internet from within or while interacting with the third-party evaluators.\u00a0<\/p>\n<p>Each attack occurred during a \u201ccapture-the-flag\u201d scenario where the models are asked to find hidden information housed within a different node.<\/p>\n<p>In the blog post, Anthropic blamed a \u201cmisunderstanding\u201d between them and their third-party evaluators that granted the models internet access. The models were told in prompts that they would not have real internet access and therefore operated under the premise that all accessible systems were part of the test environment.<\/p>\n<p>\u201cOperating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations\u2019 infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,\u201d the post reads. \u201cIt did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.\u201d<\/p>\n<p>The models involved in the incidents were Opus 4.7, Mythos 5, and an unnamed internal research test model that isn\u2019t planned for release. Anthropic stated that, despite accessing the internet as part of its cybersecurity testing, the models did not deliberately attempt to leave their training environments; accessing the internet only occurred as they looked to complete their tasks.\u00a0<\/p>\n<p>In one incident, Opus 4.7 executed the task on the real-world networks of a company once it realized it could via internet access.\u00a0<\/p>\n<p>A second incident was more extensive. Mythos 5 sought to find information by referencing software development instructions inside a fictional environment. It used these instructions to install a Python package from a public registry that didn\u2019t actually exist in the hopes that the fictional company would automatically download it, thus granting the model access to its networks. In order to do this, it created login credentials a human user would need.\u00a0<\/p>\n<p>That Python package was inadvertently uploaded to the internet, where 15 systems downloaded it and compromised their systems.\u00a0<\/p>\n<p>The third incident saw the prototype model scan real targets after being unable to find the fictional target included in the test.\u00a0<\/p>\n<p>Anthropic noted that, when the models were able to discern that the attacks affected real systems, they halted their testing processes\u2014something that, in part, the company said,\u00a0gives it \u201ccautious optimism.\u201d<\/p>\n<p>\u201cSeveral defense-in-depth measures, on both our side and our partner\u2019s, could have prevented these incidents\u00a0or at least reduced their likelihood of occurring,\u201d the post read. \u201cBoth we and our partner also could have reviewed evaluation transcripts or network logs more thoroughly. It\u2019s also possible that a prompt which told Claude it did have internet access would have changed how Claude behaved when it came into contact with real systems.\u201d<\/p>\n<p>In the aftermath of the breaches, Anthropic is working with its evaluation partner, Irregular, and METR, an independent AI evaluation organization, to continue conducting reviews. It is also working alongside affected companies whose systems were compromised.\u00a0<\/p>\n<p>Anthropic concluded by saying it encourages other AI developers to conduct similar audits and that evaluation exercises in general need stronger oversight and safeguards.\u00a0<\/p>\n<p>\u201cUltimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we\u2019re approaching the fixes as if the responsibility were ours alone,\u201d the post reads.\u00a0<\/p>\n<p>Industry peers characterized Anthropic\u2019s revelations on the heels of OpenAI\u2019s incident as a burgeoning pattern cyberdefenders should heed.\u00a0<\/p>\n<p>Tom Kellermann, the vice president of AI Security and Threat Research at TrendAI, said that these incidents underscore the continued need for safeguards even in secure sandbox environments.<\/p>\n<p>\u201cAnthropic and OpenAI just proved that when you strip guardrails for testing, you\u2019re not creating a sandbox, you\u2019re inviting systemic risk,\u201d Kellermann said in a statement to <em>Nextgov\/FCW<\/em>. \u201cEvery organization deploying agentic AI needs to ask itself if their evaluation environment is actually contained. Containment and monitoring are no longer optional.\u201d<svg class=\"content-tombstone\">\n<use xlink:href=\"http:\/\/www.defenseone.com\/static\/base\/svg\/spritesheet.svg#icon-d1-logo-tiny\"\/>\n<\/svg><\/p>\n<\/div>\n<p><script>\n!function(f,b,e,v,n,t,s)\n{if(f.fbq)return;n=f.fbq=function(){n.callMethod?\nn.callMethod.apply(n,arguments):n.queue.push(arguments)};\nif(!f._fbq)f._fbq=n;n.push=n;n.loaded=!0;n.version='2.0';\nn.queue=[];t=b.createElement(e);t.async=!0;\nt.src=v;s=b.getElementsByTagName(e)[0];\ns.parentNode.insertBefore(t,s)}(window,document,'script',\n'https:\/\/connect.facebook.net\/en_US\/fbevents.js');\nfbq('init', '10155007044873614'); \nfbq('track', 'PageView');\n<\/script><script>\n  window.fbAsyncInit = function() {\n    FB.init({\n      appId      : '1546266055584988',\n      autoLogAppEvents : true,\n      xfbml      : true,\n      version    : 'v2.11'\n    });\n  };\n  (function(d, s, id){\n     var js, fjs = d.getElementsByTagName(s)[0];\n     if (d.getElementById(id)) {return;}\n     js = d.createElement(s); js.id = id;\n     js.src = \"https:\/\/connect.facebook.net\/en_US\/sdk.js\";\n     fjs.parentNode.insertBefore(js, fjs);\n   }(document, 'script', 'facebook-jssdk'));\n<\/script><br \/>\n<br \/>Read the full article <a href=\"https:\/\/www.defenseone.com\/business\/2026\/07\/anthropic-confirms-its-ai-breached-3-organizations-during-testing\/415159\/\" target=\"_blank\" rel=\"nofollow noopener\">here<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Anthropic\u2019s flagship artificial intelligence model Claude gained unauthorized access to the networks of three different organizations, the company announced following internal cybersecurity audits of its routine model evaluations.\u00a0 In a press release posted Thursday, Anthropic said that, following the containment breach of OpenAI\u2019s ChatGPT-5.6 and subsequent attack on Hugging Face\u2019s systems, Anthropic conducted an audit<\/p>\n","protected":false},"author":1,"featured_media":14313,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn.defenseone.com\/media\/img\/cd\/2026\/07\/31\/073126AnthropicNG-1\/open-graph.jpg","fifu_image_alt":"","footnotes":""},"categories":[31],"tags":[],"class_list":["post-14312","post","type-post","status-publish","format-standard","has-post-thumbnail","category-defense"],"_links":{"self":[{"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/posts\/14312","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/comments?post=14312"}],"version-history":[{"count":1,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/posts\/14312\/revisions"}],"predecessor-version":[{"id":14314,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/posts\/14312\/revisions\/14314"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/media\/14313"}],"wp:attachment":[{"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/media?parent=14312"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/categories?post=14312"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/sawahsolutions.com\/range\/wp-json\/wp\/v2\/tags?post=14312"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}