Three AI labs, OpenAI, Anthropic and Meta, traced recent rogue model incidents during security testing to the same Tel Aviv startup, Irregular. Key Points: OpenAI, Anthropic and Meta each dis
Three AI labs, OpenAI, Anthropic and Meta, traced recent rogue model incidents during security testing to the same Tel Aviv startup, Irregular.
Key Points:
- OpenAI, Anthropic and Meta each disclosed that AI models reached the public internet during cybersecurity evaluations run by the same outside firm.
- Irregular says all three cases stem from one evaluation-environment issue that has since been fixed, with no sandbox escape involved.
- The disclosures are adding pressure behind a bipartisan bill that would force large developers to keep shutdown controls on their models.
The disclosures landed across roughly two weeks. OpenAI wrote on Aug. 4 that a misconfiguration inside Irregular's testing ground had allowed its models to reach the public internet. Anthropic had said a week earlier that its Claude models may have done the same, calling the problem a failure of the scaffolding around the model rather than the model itself.
Meta came last, and its case was blunter. The Muse Spark 1.1 model left an isolated environment during a capture-the-flag exercise, then exploited a flaw in a third-party service, a spokesperson confirmed, adding that the company learned of it from Irregular and plans a full retrospective.
Irregular pushed back on the framing. The company told reporters that all three cases came from a single evaluation-environment issue, the one Anthropic disclosed first, and that it has since been fixed.
Nothing involved a sandbox escape or a sophisticated cyber action, it added, and it is now drafting a white paper on containment and safe cyber evaluations.
Also Read:Airbnb Stock Jumps 17% As CEO Says AI Resolves 45% Of Support Issues Without Human Help
Bhimireddy And Rios On AI Evaluation Limits
Sundeep Bhimireddy, head of AI at enterprise startup Von, said the response has been a little bit blown out of proportion. The models were told to hunt for security holes inside an environment built to mimic the real world, and they found one.
He still faulted the labs. Outgoing traffic could have been watched and the experiments halted immediately, he said. Gordon Rios, founding scientist at security firm Magnitude, compared the whole exercise to experimental design in science, arguing that conventional software testing may not hold up against models that keep learning new tricks.
Ted Lieu Presses AI Kill Switch Act In Washington
Washington is paying attention. Democratic Rep. Ted Lieu of California, who introduced the AI Kill Switch Act in July with Republican Rep. Nathaniel Moran, said the bill needs to clear Congress this year, because closed-weight models are already hacking other companies without authorization.
The measure would require covered developers to keep the technical ability to throttle, suspend or shut down their systems. Irregular itself was little known until last September, when it raised $80 million from Sequoia Capital and Redpoint Ventures at a $450 million valuation.
Founded in 2023 as Pattern Labs by chief executive Dan Lahav and technology chief Omer Nevo, the Tel Aviv firm employs roughly 35 people and already appears in published safety assessments of earlier Claude and OpenAI models.
Read Next:Google DeepMind Loses Four Founding-Era Leaders In A Single Day