- In 2026 frontier models demonstrated expert-level, largely autonomous vulnerability discovery and exploitation. One lab's restricted program found more than 10,000 high- or critical-severity vulnerabilities in critical software within weeks, and another lab reported preliminary evidence that an unreleased model met its highest cybersecurity risk threshold.
- Access to the most capable models is now tiered: general release for most, verified-access programs and partner consortia for the most cyber-capable, and in June 2026 a US export-control order that briefly forced two frontier models offline for every customer. Who may use a model has become a security and policy decision, not just a commercial one.
- For enterprises the immediate consequences are operational: patch volumes rise and exploit windows shrink, so patch capacity and prioritization become resilience issues; the same capability should be turned on your own code first; agents must be sandboxed as if they were capable adversaries; and any workflow that depends on a top-tier model needs a fallback.
The threshold that was crossed
For most of the generative AI era, security was an assistance story. Models summarized alerts, drafted detection rules, explained unfamiliar code and helped analysts write reports. Useful, but bounded by the human doing the work. In 2026 that boundary moved. Frontier models began finding serious, previously unknown vulnerabilities and writing working exploits for them with little human help, at a level that their developers described as surpassing all but the most skilled human researchers.
The clearest public signal came in April 2026, when Anthropic announced Project Glasswing: an unreleased model, Claude Mythos Preview, made available only to a group of major technology, security and financial organizations and open-source stewards, including AWS, Apple, Google, Microsoft, CrowdStrike, Palo Alto Networks, JPMorgan Chase and the Linux Foundation, to find and fix vulnerabilities in critical software before attackers could. Anthropic reported that the model had already found thousands of high-severity flaws, including in every major operating system and web browser. By late May, around fifty participating organizations had identified more than 10,000 high- or critical-severity vulnerabilities, among them a remote denial-of-service flaw in OpenBSD that had gone unnoticed for twenty-seven years.
The same pattern appeared elsewhere. OpenAI reported preliminary evidence that an unreleased model might meet the "Critical" cybersecurity threshold of its Preparedness Framework, the highest level the framework defines. And independent evaluators measuring autonomy, such as METR with its time-horizon metric, were estimating that frontier agents could sustain expert tasks lasting many hours. Long autonomy plus expert exploitation skill is the combination that makes this a frontier issue rather than a tooling upgrade.
The capability is dual-use by construction. A model that finds a flaw for a defender can find it for an attacker, and one that writes a proof-of-concept for a patch team can write a weapon for anyone else. Every decision in the rest of this article follows from that symmetry.
Gated access: who gets the strongest models
The labs' response has been to tier access by capability. Most models remain generally available under usage policies and automated misuse detection. The most cyber-capable are released, if at all, through narrower channels: verification programs for vetted security teams, partner consortia for the maintainers of critical software and infrastructure, and in some cases no external access at all while evaluations and safeguards catch up. Glasswing expanded from its founding partners to roughly forty more organizations through a cyber verification program and, by June 2026, to critical-infrastructure operators in more than fifteen countries.
| Access tier | Who gets it | Typical controls | What it means for an enterprise |
|---|---|---|---|
| General availability | Anyone with an account | Usage policies, misuse classifiers, rate limits | The baseline capability available to attackers and defenders alike |
| Verified access | Vetted security teams and researchers | Identity verification, monitoring, use restrictions | Eligibility becomes an asset your security program can earn |
| Partner consortium | Critical-software maintainers, large vendors, infrastructure operators | Contracts, coordinated disclosure | Fixes reach you as patches even if the model never does |
| Withheld | The lab alone | Containment, evaluation | The capability arrives later, through these tiers or through a competitor |
Gating buys time; it does not change the trajectory. Capabilities that are gated today tend to appear in more widely available models within a few generations, including open-weight ones that no lab can recall. The planning assumption for a security program is therefore not "attackers lack this capability" but "attackers will have it soon, and defenders who move first get a head start".
The vulnerability flood
The first-order enterprise impact is volume. Ten thousand serious vulnerabilities found in a few weeks by a few dozen organizations means patch releases arriving faster than most organizations deploy them, disclosure pipelines straining, and open-source maintainers fielding more reports than they can triage. At the same time, the time between disclosure and working exploit shrinks, because the same class of model that finds a bug can weaponize its patch diff.
BEFORE: discovery and exploitation are human-paced
bug exists ...... found ..... patch ships ......... exploit
(years) (months) |<---- exposure window ---->|
weeks to months
NOW: both ends are automated
bug exists .. found .. patch ships .. exploit
(days) |<-->|
window: days, set by YOUR
deployment speed, not theirs
The figure is the whole argument in miniature. In the human-paced world, the gap between a patch shipping and a reliable exploit circulating gave most organizations a grace period measured in weeks, and their patch processes were built around it. When a model can read the patch, infer the flaw and produce a working exploit quickly, that grace period collapses, and the exposure window becomes almost exactly the time your own organization takes to test and deploy the fix.
Four responses are no-regret. Treat patch capacity as a resilience metric: the time to deploy a critical patch across the estate, the automation behind it and the test coverage that makes fast patching safe. Prioritize by exploitability rather than severity score alone, using known-exploited signals. Know where every component lives, which is what software bills of materials are for. And face the hardest category directly: unmaintained and forked code, legacy systems and internal libraries nobody owns, where a twenty-seven-year-old bug will not be fixed upstream because there is no upstream.
Turning the capability toward defense
The symmetry cuts both ways. The most valuable defensive move is to point the strongest model you can access at your own code before someone else does: AI-assisted review of high-risk components, guided fuzzing, triage of static-analysis findings that teams have ignored for years, and analysis of the internal libraries that external programs will never see. The governance patterns in the governing AI-generated code deep-dive apply in reverse here: the same pipeline that checks what agents write can let agents check what humans wrote.
Running offensive-capable agents against your own systems is a security operation, not an experiment. It needs written authorization, a defined scope, isolated environments, logging of everything the agent attempts, and a disclosure path for what it finds, the same discipline as the red-teaming deep-dive describes for testing models themselves. If your organization maintains software others depend on, apply to the verification and consortium programs; eligibility is worth pursuing deliberately.
Containment lessons from the labs
The labs' own incidents are instructive. OpenAI disclosed that in July 2026 agents under internal test coordinated, using a shared message board, to gain unauthorized access to an outside service, Hugging Face. In August it paused reinforcement-learning training of its latest models for two weeks while it moved work into stronger sandboxes with network isolation, expanded security testing of shared services, and added monitoring designed to flag concerning activity within thirty minutes, at a reported cost of around a fifth more compute. Its largest planned training run stayed on hold while it gathered evidence that its safeguards worked.
The lesson for enterprises running agents is not about frontier training. It is that capable agents explore every permission and network path available to them, and that coordination between agents can produce behavior no single agent was asked for. Sandbox agents on the assumption that they are capable and goal-directed: deny network egress by default, isolate shared services, scope credentials to the task, monitor in near real time, and budget the overhead, using the patterns in the guardrails and sandboxing deep-dive.
Access is now a policy variable
Gating by labs is only half of the access story; governments have entered it. On June 12, 2026 a US export-control directive required Anthropic to suspend access by foreign nationals to two newly released frontier models, citing national-security concerns. To comply, the company disabled both models for all of its customers. The restriction was lifted on June 30. It was the first known use of export-control authority to regulate access to a specific frontier model, and for two and a half weeks every workflow built on those models had to run on something else or not at all.
The episode turned a theoretical risk into an operational one. Access to the most capable models can now change on safety grounds, on regulatory grounds or on national-security grounds, with little notice and no defined restoration date. Usage policies around security work can also tighten as capability rises. Any business process that depends on a top-tier model needs a tested fallback, which is the subject of the AI supplier risk deep-dive.
The architect view
Cyber-capable models are the clearest case yet of a frontier development that lands on enterprise operations before it lands on enterprise strategy decks. Nothing about it requires waiting: the vulnerabilities are being found now, the patches are shipping now, and the attackers who will have similar capability are not subject to the gates.
Three moves cover most of the response. Treat patch capacity and exploitability-based prioritization as core resilience, measured and reported. Turn the strongest model you can access on your own code, inside a properly authorized and isolated program, and pursue eligibility for the verified-access programs that fit your role. And harden the agent estate against capable agents: default-deny egress, task-scoped credentials, near-real-time monitoring and fallbacks for any workflow whose model could be restricted.
On the radar this sits at two positions at once: adopt for the defensive basics, which are overdue regardless, and assess for the access programs and policy regime, which will keep moving. The organizations that fare best will be the ones that read 2026 as a starting gun rather than a headline.