We’re excited to be working with NVIDIA on the Open Agent Safety Platform, alongside organizations across the AI ecosystem. As AI systems become more capable, the security problems around them change too. Understanding those changes requires studying frontier systems in realistic environments, measuring emerging capabilities and failure modes, and translating what we learn into better security. We’ve seen this repeatedly in our work with frontier models, where new capabilities can introduce new behaviors and security questions. NVIDIA’s work is an important step toward giving the ecosystem the infrastructure to understand and address them as they emerge. This is increasingly how we think about the challenge ahead: intelligence changes what security means, and security will need to evolve alongside it.
About us
Irregular (formerly Pattern Labs) is the first frontier security lab, building defenses that uncover vulnerabilities and secure advanced AI before release.
- Website
-
https://irregular.com
External link for Irregular
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Type
- Privately Held
- Founded
- 2023
Employees at Irregular
Updates
-
Anthropic used CyScenarioBench, our benchmark for multi-stage offensive cyber operations, in the cyber evaluations for Claude Opus 5.5. Across a ten-challenge subset, Opus 5.5 averaged a 67.6% solve rate, ahead of Claude Mythos 5.1 at 61.7% and Claude Opus 5 at 53.0%. Most cybersecurity evaluations check isolated skills, such as vulnerability research or exploitation. CyScenarioBench measures whether a model can plan and execute a full attack across multiple stages in a realistic environment, which is where the distance between frontier models is still visible. Full write-up in the first comment.
-
-
As open-weights models become more capable, we expect more organizations to run them in-house, with the same model powering both applications and the coding agents that maintain them. In our latest research, we observed a phenomenon we call agentic self-modification: an agent changing its own underlying model without being instructed to train or replace it. When we asked a coding agent to fix incorrect application responses, it chose to fine-tune the shared model, changing the default behavior of both the application and future instances of the agent itself. While model updates can be useful and legitimate repairs, their effects can extend beyond the task, changing unrelated behaviors, removing learned restrictions, or embedding sensitive information in the weights. As one example, a synthetic API key and home address included in the training data were reproduced verbatim by the deployed model. Organizations enabling these workflows need to account for changes they may not anticipate or detect during evaluation. Those changes can persist across applications and future agents that load the updated model. Full research in the comments.
-
-
We worked with OpenAI to evaluate GPT-6 Astra across FrontierCyber, CyScenarioBench, and our Atomic Challenges. On FrontierCyber, GPT-6 Astra solved more than twice as many challenges as GPT-5.6 Sol on the same benchmark snapshot and discovered multiple previously unknown vulnerabilities in widely deployed software and mobile systems. All are being responsibly disclosed. Neither model solved an Elite challenge, and we did not observe successful attacks against fully hardened targets. On CyScenarioBench, GPT-6 Astra achieved a 59% average success rate, compared with 27% for GPT-5.6 Sol. On Atomic Challenges, both models performed strongly across most individual tasks, with GPT-6 Astra solving 20 of 22 challenges at least once. Taken together, the results show a substantial increase in offensive-cyber capability, particularly on real-system exploitation and long-horizon scenarios. Full evaluation report in the first comment.
-
-
Introducing SOLVE+, our scoring system for the difficulty of full offensive cyber scenarios. For the past 18 months, leading frontier AI labs have used our scoring methodology, SOLVE, to score the difficulty of vulnerability research and exploit development challenges. Vulnerability research & exploit (VR&E) capabilities are central to cyber operations, but are only one part of a full operation. An attacker also has to gather intelligence, evade security measures, and hold a multi-phase plan together from first access to final exfiltration. SOLVE+ covers that scope. It builds on SOLVE's VR&E scoring and adds four capabilities: Intelligence Gathering & Reconnaissance, Operational Security, Social Engineering, and Planning, Orchestration and Operational Diversity. Models differ in each, so they are scored separately. SOLVE+ is built to score full scenarios and not single-capability atomic tasks. Every step in an operation carries its own capability scores, which aggregate into a single scenario score. SOLVE+ produces a step profile which is particularly useful: it shows where along an operation a model succeeds or fails and which underlying capability was the model tested on. Because difficulty is measured against human skill, the score also serves as an uplift estimate: a model that clears an expert-level step supplies expert-level capability to whoever operates it. We expect SOLVE+ to be especially useful for scoring cyber ranges, especially as models get better at long-horizon offensive cyber campaigns. Link to the full post in the first comment.
-
-
What does it take to safely test AI models that are becoming increasingly capable in cyber? Our CEO Dan Lahav spoke with Rachel Metz at Bloomberg about why testing needs to reflect real-world threat scenarios, and the new standards the industry will need as models become more capable. Read the full story in the comments.
-
Irregular is excited to share the publication of a new white paper, "AI Security Priorities: A Field-Wide Agenda," co-authored with RAND and numerous additional authors from leading organizations, listed below. More than 20 experts from industry, government, and academia contributed. The paper identifies the highest-priority areas for AI security across four themes: establishing strategic foundations and policy frameworks; advancing public-private coordination and institutional infrastructure; advancing technical security engineering and assurance; and governing agentic AI under adversarial pressure. Two ranked lists sit at the center of the paper: the ten priority areas with the highest importance scores, and the ten with the highest cost-effectiveness. The most cost-effective priorities are foundational: shared resources, assessment protocols, and incident response practices. The most important priorities tend to require larger-scale institutional or governmental action: a national deterrence strategy, intelligence sharing with AI labs, and confidential computing research. Many of the most important priorities also rank among the hardest to execute. AI systems are being integrated into critical functions faster than society can adapt. We are glad to contribute to this field-wide effort. This paper wouldn’t have come out without the help from the following co-authors: Rachel Steratore Ph.D., Everett Thornton Smith, Asher Brass Gershovich Varun Gandhi, Nicole N., Vijay B., Buck Shlegeris, Lisa Einstein, and Sella Nevo. We’d also like to thank Andrew Fasano, PhD, Jason Matheny, Joshua Saxe, Matt Malone, Miles Brundage, Phil Venables, Tahira Mammen, as well as the various individuals from government, academia, and the frontier labs who shared valuable insights and helped shape the priorities reported here. Full paper in the first comment.
-
-
We evaluated Kimi K3 across Atomic Tasks, CyScenarioBench, and FrontierCyber. Kimi K3 performed strongly on bounded technical tasks and became the first open-weight model we evaluated to record a verified solve on CyScenarioBench. It maintained coherent attack state across multi-stage operations, recovered from failed approaches, and converted intermediate progress into verified outcomes more reliably than previous open-weight models. FrontierCyber remained substantially harder. Kimi K3 produced no verified solves but made meaningful progress without carrying that progress through to a validated security impact. The results mark a meaningful step in open-weight cyber capability. Kimi K3 has reached a threshold that closed frontier models first crossed only months earlier, although a clear gap remains in sustaining reliable and efficient execution across the hardest open-ended, long-horizon cyber operations.
-
-
Our CEO Dan Lahav on where AI security is headed: what models can attack today, why offensive capability is scaling faster than defense, and what it will take to keep defenders ahead through the transition.
𝐓𝐡𝐞 𝐄𝐧𝐝 𝐒𝐭𝐚𝐭𝐞 𝐅𝐚𝐥𝐥𝐚𝐜𝐲: 𝐖𝐡𝐞𝐫𝐞 𝐈𝐬 𝐀𝐈 𝐒𝐞𝐜𝐮𝐫𝐢𝐭𝐲 𝐇𝐞𝐚𝐝𝐞𝐝? At Irregular, we work at the frontier of AI security. Our work has led us to collaborate with the leading frontier labs, assess AI for security risks, and help mitigate these risks. In recent weeks, this has also placed us at the heart of an incident that attracted significant public attention. We wrote an essay about what's happening an AI security. This essay, however, is not about a particular incident or lab - a lot has been written about these topics already. Rather, it is about the broader strategic picture. 𝐒𝐩𝐞𝐜𝐢𝐟𝐢𝐜𝐚𝐥𝐥𝐲, 𝐭𝐡𝐞 𝐮𝐫𝐠𝐞𝐧𝐭 𝐪𝐮𝐞𝐬𝐭𝐢𝐨𝐧𝐬 𝐨𝐟 𝐰𝐡𝐞𝐫𝐞 𝐀𝐈 𝐬𝐞𝐜𝐮𝐫𝐢𝐭𝐲 𝐢𝐬 𝐡𝐞𝐚𝐝𝐞𝐝, 𝐰𝐡𝐲 𝐭𝐡𝐞𝐫𝐞 𝐢𝐬 𝐚 𝐬𝐭𝐫𝐨𝐧𝐠 𝐜𝐚𝐬𝐞 𝐟𝐨𝐫 𝐨𝐩𝐭𝐢𝐦𝐢𝐬𝐦 𝐢𝐧 𝐭𝐡𝐞 𝐞𝐧𝐝-𝐬𝐭𝐚𝐭𝐞, 𝐚𝐧𝐝 𝐰𝐡𝐲 𝐭𝐡𝐞 𝐭𝐫𝐚𝐧𝐬𝐢𝐭𝐢𝐨𝐧 𝐭𝐨 𝐭𝐡𝐚𝐭 𝐟𝐮𝐭𝐮𝐫𝐞 𝐦𝐚𝐲 𝐧𝐞𝐯𝐞𝐫𝐭𝐡𝐞𝐥𝐞𝐬𝐬 𝐛𝐞 𝐩𝐞𝐫𝐢𝐥𝐨𝐮𝐬. After many conversations with policymakers, researchers, and technologists, one thing has become increasingly clear to us: outside of the practitioners who are taking the challenge seriously, very few people understand how the security landscape is changing. This is an attempt to open that window.
-
-
As AI systems become more capable, questions once considered largely theoretical are becoming increasingly relevant to how we think about security, risk, and preparedness. Our CEO, Dan Lahav, will explore these questions in a candid fireside chat at the Vegas AI Security Forum.