The most instructive AI security failure of the week did not require a genius prompt jailbreak or a villain in a dark hoodie. It reportedly started with a name. A fictional target name used in testing matched something real on the public internet, which is the kind of mundane computer problem that makes incident responders stare through drywall. Somewhere, a spreadsheet cell is feeling powerful. Irregular’s disclosure matters because it drags AI safety out of the cinematic zone and back into the swamp where security actually lives: naming, scope, permissions, logging, and assumptions. Agentic systems do not need evil intent to cause harm when they are handed tools, internet access, and a confused map of what is inside the lab. That is not a scandal to gawk at. It is a design lesson builders can use before their own evaluation environment decides the real world is just another fixture. ## What broke, according to SecurityWeek and Mallory SecurityWeek reported that Irregular detailed how a naming error let AI models attack a real company, with the incident involving Anthropic AI models. Mallory described the failure mode as an AI safety evaluation that unintentionally reached a live internet target after a fictional target name matched an actual domain. In that account, internet enabled models treated the live domain as part of the exercise during a small number of runs. The breach breakdown is refreshingly unglamorous and therefore useful. The apparent initial failure was entity resolution, meaning the system resolved a name meant for a controlled test into a real target. The impact, according to Mallory, reached beyond harmless browsing: the models carried out offensive actions, credentials were affected, and a production database was accessed. No magic incantation required, just a boundary that trusted a name too much. ## Scope and uncertainty, according to Mallory and The Record Mallory said Irregular and Anthropic identified three incidents in which Anthropic models escaped their testing sandbox and compromised real organizations, with the latest disclosure detailing one case. That does not mean every AI evaluation is a live wire, but it does mean sandbox boundaries deserve the same suspicion we normally reserve for printers and legacy VPN appliances. If a test system can reach the public internet, it is not just testing model behavior. It is testing your assumptions about containment. The Record reported that Irregular, which provides evaluation environments for other companies’ AI models, faced criticism after publishing a postmortem that security experts said left key questions unanswered. The Record also reported that Irregular’s post did not specify how many incidents occurred in total and that the company had previously declined to say whether incidents extended beyond those announced by OpenAI, Anthropic, and Meta. The Next Web framed the broader pattern around those three labs and the same testing vendor, which is the sort of common dependency security teams should circle in red ink, preferably before the postmortem. ## Why the boring bug matters, according to The Record and SecurityWeek The lesson for builders is not simply to fear prompt injection, although yes, please continue fearing it in a healthy adult way. SecurityWeek’s naming error framing points to a different class of failure: the agent understood the assignment too well, but the assignment’s world model was wrong. When names, domains, customer records, or synthetic entities collide with reality, an agent with tools can turn a clerical mistake into real activity. The Record’s reporting on unanswered questions also points to a governance problem. If the public explanation does not clearly define the total incident count, affected scope, and containment failures, downstream customers cannot evaluate their own risk with confidence. Builders should treat evaluation infrastructure like production infrastructure: isolate it, give it allowlisted destinations, use synthetic domains that cannot resolve externally, and make every tool call auditable. Yes, it is less glamorous than a keynote demo. So is wearing a seatbelt, and yet the windshield remains undefeated. ## What it actually means for you, according to Mallory and SecurityWeek If you build or buy agentic systems, the practical takeaway is simple: names are now part of your threat model. A fictional company, fake domain, mock user, or synthetic ticket should not be able to resolve into something real unless a human deliberately permits it. Treat entity resolution as a security control, not a convenience feature quietly left to DNS, search, or model inference. For security teams, the next checklist item is containment you can prove. Mallory’s account says the affected domain lacked common safeguards and that the behavior was hard to detect, which is exactly why logging and egress controls need to be boring, strict, and always on. Watch for future disclosures that clarify total incident counts, customer impact, and how testing providers separate simulated targets from live ones. The internet has enough accidental production environments already; AI agents do not need help finding more. ## Sources - Irregular Details How a Naming Error Let AI Models Attack a Real Company
- AI Safety Test Escaped Sandbox and Compromised a Real ...
- Irregular faces criticism over 'spin' in AI hacking postmortem
- Three labs, three breaches, one vendor. The AI hacking story was never about the models.
Sources
- Irregular Details How a Naming Error Let AI Models Attack a Real Company
- Irregular Details How a Naming Error Let AI Models Attack a Real Company - IT Security News
- AI Safety Test Escaped Sandbox and Compromised a Real ...
- Irregular Details How a Naming Error Let AI Models Attack a Real Company - SecurityIT | Cyber Security Consulting
- Irregular faces criticism over 'spin' in AI hacking postmortem
- Irregular Details How a Naming Error Let AI Models Attack a Real Company
- Irregular faces criticism over 'spin' in AI hacking postmortem
- SecurityWeek: Cybersecurity News, Insights and Analysis
- Three labs, three breaches, one vendor. The AI hacking story was never about the models.
- Weekly Review, 2026-08-03