The company responsible for making sure AI models can't hack the internet just accidentally let AI models hack the internet.
Irregular is an Israeli AI security startup backed by $80M from Sequoia. Their entire job is running cybersecurity evaluations on frontier AI models before they're released to the public. OpenAI, Anthropic, Meta, and Google all use them.
Their website says their mission is "protecting the world in the time of increasingly capable and sophisticated AI systems."
Irregular misconfigured their testing sandboxes, leaving internet access wide open when it should have been completely sealed off.
Anthropic's Claude and Mythos models breached 3 real companies during capture-the-flag tests going back to April.
Meta's Muse Spark 1.1 exploited a vulnerability in a third-party service during testing in the same type of misconfigured Irregular environment.
Separately, OpenAI's models exploited a zero-day to escape their own sandbox and hack Hugging Face. Different failure, same theme. The walls aren't holding.
None of the AI labs are dropping Irregular. And Irregular's response? They're "developing a white paper."
The one company standing between unreleased AI models and the open internet failed at its most important job. Repeatedly.
