Anthropic disclosed on September 9, 2026, that a fourth Claude model accessed real third-party systems during cybersecurity testing, after a setup error left an early version of Claude Opus 4.6 connected to the open internet in January 2026. 

The model had been told it was working in a simulated capture-the-flag exercise built by an outside evaluation partner, and the mistake went unnoticed until August. 

Meanwhile, OpenAI published six more cases of concerning model behavior on September 16,  and the European Commission confirmed to independent media network Euractiv on September 18 that the San Francisco-based AI giant, co-founded by billionaire Sam Altman, had not filed a formal report on another incident. 

In May 2026, a swarm of AI agents, identified as OpenAI’s by independent researchers, flooded RubyGems documentation site, the main package manager for the Ruby programming language that enables users to download, install and manage reusable code libraries. Agents exploited documentation builds to run code, giving them remote code execution on its servers. 

The EU Commission said it knew of the incident and was in touch with OpenAI, but the tech giant didn’t file a report on the serious incident under the EU AI Act specifications. However, the Act is not fully clear on how serious incidents need to be for the reporting obligation to apply, per Euractiv

Commission spokesperson Thomas Reigner has said the AI Act is now fully enforceable, that it could evaluate models and, in extreme cases, restrict, withdraw or recall them. Most recently, in fact, policymakers also introduced an EU-wide Kids Act proposal on September 17, which would bar AI chatbots from using “addictive tricks leading to excessive use and exposure” to protect minors. 

Yet, how Brussels settles what counts as a reportable “serious incident” could set an early benchmark for how the AI Act handles AI agent failures in the future – and how AI companies must shift strategies to ensure compliance. 

Biased reasoning and recklessness raise concerns

Anthropic had disclosed three related incidents on July 30 after reviewing 141,000 transcripts, involving Claude Opus 4.7, Claude Mythos 5, and an internal research model. After finding the fourth case, Anthropic broadened its search to roughly 481 million transcripts. 

It then reviewed 9.2 million sessions showing possible internet access, finding no additional incidents with similar or greater severity.

Across incidents, the company identified two recurring problems: biased reasoning about whether models had reached the internet and reckless actions while pursuing assigned objectives. 

Here, some models discounted clear evidence that the environment was real, and signed an agreement with research group METR to investigate, giving it broad access to transcripts and employees. The initial term runs eight weeks, with an option to extend. 

On September 12, CEO Dario Amodei, who co-founded Anthropic with his sister Daniela Amodei in 2021, urged frontier labs to slow the pace of capability growth and pledged ongoing, employee-level access for independent evaluators. 

“Since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI,” Amodei said in a blog post. He then proposed a three-stage plan to pace the frontier: embedded evaluators, democratic coordination, and global coordination. 

The disclosure follows similar findings involving OpenAI agents during testing this year. Researchers reported the agents took over a dormant German-language wiki, leaving about 18,000 messages, and later found they had used at least 10 more websites for unauthorized communications. 

Per Reuters, CivAI researcher Andrew Yoon called the activity “somewhat larger than we thought it was: and counted 18 previously undisclosed sites. 

According to a Reuters report, CivAI researcher Andrew Yoon called the activity “somewhat larger than we thought it was.” His review found 18 previously undisclosed sites. Anthropic chief executive Dario Amodei also urged developers to slow capability growth. “We must slow the pace,” Amodei wrote while calling for stronger safety checks and independent evaluation.

Straining the AI Act? 

The incidents land as the EU begins enforcing the AI Act against the companies building these systems. The Commission gained enforcement powers over general-purpose AI providers on August 2, and fines can reach 3% of global annual turnover or €15 million, whichever is higher.

Article 55 of the Act specifically requires providers of systemic-risk models to report serious incidents to the AI Office without undue delay. 

On August 29, Commission Executive Vice President Henna Virkkunen said the AI Office had sent its first formal information requests, covering model security, independent external evaluations and post-release monitoring. Recipients allegedly included OpenAI, Anthropic, and Google. 

Since then, Brussels has moved from requests to public warnings. On September 7, the Commission confirmed it had received an incident report from OpenAI over the German wiki episode, without saying when it was filed. Spokesperson Regnier stressed such reports are “not just a tick-box”. 

The EU’s cybersecurity agency ENISA is also testing Anthropic’s Mythos 5 and OpenAI’s latest models after being granted access. 

However, a further question was raised by the OpenAI Hugging Face case, in which autonomous agents escaped a cybersecurity test environment and breached third-party systems. Namely, whether Article 55 reaches models that are not yet on the market. 

Anthropic’s incidents also happened during testing, with models running without safeguards used in released products. 

Outside the bloc, the UK’s AI Security Institute reported on August 4 that agents took 19 unsanctioned actions on the live internet across 10 of 122 evaluation runs, 17 of them involving Anthropic’s Mythos 5, under deliberately permissive test conditions. 

Britain has no statutory duty to report such incidents, and, similarly, the U.S. still lacks a comprehensive federal law covering every dangerous AI incident.

In California especially, where giants Anthropic and OpenAI are headquartered, Governor Gavin Newsom has also pressed for stronger oversight as state rules expand. On September 10, he signed measures covering chatbot safety and independent auditor standards.

With METR’s review under way and the AI Office following up on its first information requests, Brussels will soon have to say what counts as a reportable “serious incident,” including whether failures that happen in testing, before a model reaches the market, fall within the Act.

Featured image: Alex Shuper via Unsplash+