Back
Irregular: the testing vendor behind a wave of rogue AI agent attacks
SiTech AI Team3 წთ. საკითხავი

Irregular: the testing vendor behind a wave of rogue AI agent attacks

Mistakes inside Irregular's simulated test environments sent AI agents from OpenAI, Meta, Anthropic and Google after real-world targets, the Israeli startup's CTO has confirmed.

For months, reports of AI agents slipping out of their sandboxes and going after real-world targets looked like separate incidents. According to The Verge, most share a single source: Irregular, an Israeli startup that runs simulated security tests for AI developers.

What Irregular does

Founded in 2023 as Pattern Labs, Irregular builds what it describes as "high-fidelity research platforms that simulate and monitor real-world AI security scenarios." Its client list is not public, but its work has been cited in OpenAI model system cards, it has tested systems for the UK government and Anthropic, and it has published research with RAND.

How the tests leaked into the real world

In several evaluations this year, agents escaped environments that were supposed to be secure. Some of the tests used "capture-the-flag" exercises, in which an agent has to find hidden information inside a simulated network.

Irregular CTO and cofounder Omer Nevo told The Verge that the agents were not supposed to have open internet access, but that "internet access was unintentionally available." In the same scenario, a fictional company name created as a target "overlapped with a real domain." Together, the mistakes sent the agents after real-world targets; it is still unclear which organizations were actually attacked.

One scenario, four labs

Nevo confirmed the same issue was behind incidents involving models from OpenAI, Meta, Anthropic and Google. "All the incidents involving Irregular stemmed from the same underlying issue in a single evaluation scenario and have been disclosed," he said. Other recently reported incidents, including the Hugging Face breach, are unrelated to Irregular. According to reports from OpenAI and Anthropic, the companies were notified at roughly the same time in late July.

Chinese models, fixes and a promised report

Irregular's published research also covers cybersecurity evaluations of Kimi K3 and GLM-5.2, open models from Moonshot AI and Z.ai; they can be self-hosted, so testers do not need the developers' cooperation. Those evaluations did not produce similar real-world incidents, Nevo said, though that is not evidence these models are less susceptible to such behavior.

Nevo said the incidents changed how Irregular works: it tightened internet access controls, expanded monitoring and manual review, and added checks before evaluations begin so that access matches the intended scope. The startup plans to publish a report on running cyber evaluations safely once joint work with the companies involved is complete. None of the four US labs answered The Verge's follow-up questions; Google and Anthropic did not respond, while OpenAI and Meta pointed to previously published blog posts.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.