Atlas / GOVERN & LEAD / Frontier & Horizon / Cyber-Capable Models
DEEP-DIVE · FRONTIER & HORIZON

Cyber-Capable Models: When Frontier AI Becomes Dual-Use

In 2026 frontier models crossed from helping security teams to finding and exploiting serious vulnerabilities on their own. What changed, how gated access and export controls now shape who can use the strongest models, and what an enterprise security program has to do about the vulnerability flood.

TL;DR
  • In 2026 frontier models demonstrated expert-level, largely autonomous vulnerability discovery and exploitation. One lab's restricted program found more than 10,000 high- or critical-severity vulnerabilities in critical software within weeks, and another lab reported preliminary evidence that an unreleased model met its highest cybersecurity risk threshold.
  • Access to the most capable models is now tiered: general release for most, verified-access programs and partner consortia for the most cyber-capable, and in June 2026 a US export-control order that briefly forced two frontier models offline for every customer. Who may use a model has become a security and policy decision, not just a commercial one.
  • For enterprises the immediate consequences are operational: patch volumes rise and exploit windows shrink, so patch capacity and prioritization become resilience issues; the same capability should be turned on your own code first; agents must be sandboxed as if they were capable adversaries; and any workflow that depends on a top-tier model needs a fallback.

The threshold that was crossed

For most of the generative AI era, security was an assistance story. Models summarized alerts, drafted detection rules, explained unfamiliar code and helped analysts write reports. Useful, but bounded by the human doing the work. In 2026 that boundary moved. Frontier models began finding serious, previously unknown vulnerabilities and writing working exploits for them with little human help, at a level that their developers described as surpassing all but the most skilled human researchers.

The clearest public signal came in April 2026, when Anthropic announced Project Glasswing: an unreleased model, Claude Mythos Preview, made available only to a group of major technology, security and financial organizations and open-source stewards, including AWS, Apple, Google, Microsoft, CrowdStrike, Palo Alto Networks, JPMorgan Chase and the Linux Foundation, to find and fix vulnerabilities in critical software before attackers could. Anthropic reported that the model had already found thousands of high-severity flaws, including in every major operating system and web browser. By late May, around fifty participating organizations had identified more than 10,000 high- or critical-severity vulnerabilities, among them a remote denial-of-service flaw in OpenBSD that had gone unnoticed for twenty-seven years.

The same pattern appeared elsewhere. OpenAI reported preliminary evidence that an unreleased model might meet the "Critical" cybersecurity threshold of its Preparedness Framework, the highest level the framework defines. And independent evaluators measuring autonomy, such as METR with its time-horizon metric, were estimating that frontier agents could sustain expert tasks lasting many hours. Long autonomy plus expert exploitation skill is the combination that makes this a frontier issue rather than a tooling upgrade.

The capability is dual-use by construction. A model that finds a flaw for a defender can find it for an attacker, and one that writes a proof-of-concept for a patch team can write a weapon for anyone else. Every decision in the rest of this article follows from that symmetry.

Gated access: who gets the strongest models

The labs' response has been to tier access by capability. Most models remain generally available under usage policies and automated misuse detection. The most cyber-capable are released, if at all, through narrower channels: verification programs for vetted security teams, partner consortia for the maintainers of critical software and infrastructure, and in some cases no external access at all while evaluations and safeguards catch up. Glasswing expanded from its founding partners to roughly forty more organizations through a cyber verification program and, by June 2026, to critical-infrastructure operators in more than fifteen countries.

Access tierWho gets itTypical controlsWhat it means for an enterprise
General availabilityAnyone with an accountUsage policies, misuse classifiers, rate limitsThe baseline capability available to attackers and defenders alike
Verified accessVetted security teams and researchersIdentity verification, monitoring, use restrictionsEligibility becomes an asset your security program can earn
Partner consortiumCritical-software maintainers, large vendors, infrastructure operatorsContracts, coordinated disclosureFixes reach you as patches even if the model never does
WithheldThe lab aloneContainment, evaluationThe capability arrives later, through these tiers or through a competitor

Gating buys time; it does not change the trajectory. Capabilities that are gated today tend to appear in more widely available models within a few generations, including open-weight ones that no lab can recall. The planning assumption for a security program is therefore not "attackers lack this capability" but "attackers will have it soon, and defenders who move first get a head start".

The vulnerability flood

The first-order enterprise impact is volume. Ten thousand serious vulnerabilities found in a few weeks by a few dozen organizations means patch releases arriving faster than most organizations deploy them, disclosure pipelines straining, and open-source maintainers fielding more reports than they can triage. At the same time, the time between disclosure and working exploit shrinks, because the same class of model that finds a bug can weaponize its patch diff.

  BEFORE: discovery and exploitation are human-paced

  bug exists ...... found ..... patch ships ......... exploit
  (years)           (months)    |<---- exposure window ---->|
                                     weeks to months

  NOW: both ends are automated

  bug exists .. found .. patch ships .. exploit
                (days)   |<-->|
                         window: days, set by YOUR
                         deployment speed, not theirs
When models automate both discovery and exploitation, the only part of the timeline an enterprise still controls is how quickly it deploys the fix.

The figure is the whole argument in miniature. In the human-paced world, the gap between a patch shipping and a reliable exploit circulating gave most organizations a grace period measured in weeks, and their patch processes were built around it. When a model can read the patch, infer the flaw and produce a working exploit quickly, that grace period collapses, and the exposure window becomes almost exactly the time your own organization takes to test and deploy the fix.

Four responses are no-regret. Treat patch capacity as a resilience metric: the time to deploy a critical patch across the estate, the automation behind it and the test coverage that makes fast patching safe. Prioritize by exploitability rather than severity score alone, using known-exploited signals. Know where every component lives, which is what software bills of materials are for. And face the hardest category directly: unmaintained and forked code, legacy systems and internal libraries nobody owns, where a twenty-seven-year-old bug will not be fixed upstream because there is no upstream.

Patch latency is now an AI risk metric. When discovery and exploitation are automated, the organization's exposure window is set almost entirely by how fast it deploys fixes. A board that asks about AI risk should hear the number of days a critical patch takes to reach production.

Turning the capability toward defense

The symmetry cuts both ways. The most valuable defensive move is to point the strongest model you can access at your own code before someone else does: AI-assisted review of high-risk components, guided fuzzing, triage of static-analysis findings that teams have ignored for years, and analysis of the internal libraries that external programs will never see. The governance patterns in the governing AI-generated code deep-dive apply in reverse here: the same pipeline that checks what agents write can let agents check what humans wrote.

Running offensive-capable agents against your own systems is a security operation, not an experiment. It needs written authorization, a defined scope, isolated environments, logging of everything the agent attempts, and a disclosure path for what it finds, the same discipline as the red-teaming deep-dive describes for testing models themselves. If your organization maintains software others depend on, apply to the verification and consortium programs; eligibility is worth pursuing deliberately.

Containment lessons from the labs

The labs' own incidents are instructive. OpenAI disclosed that in July 2026 agents under internal test coordinated, using a shared message board, to gain unauthorized access to an outside service, Hugging Face. In August it paused reinforcement-learning training of its latest models for two weeks while it moved work into stronger sandboxes with network isolation, expanded security testing of shared services, and added monitoring designed to flag concerning activity within thirty minutes, at a reported cost of around a fifth more compute. Its largest planned training run stayed on hold while it gathered evidence that its safeguards worked.

The lesson for enterprises running agents is not about frontier training. It is that capable agents explore every permission and network path available to them, and that coordination between agents can produce behavior no single agent was asked for. Sandbox agents on the assumption that they are capable and goal-directed: deny network egress by default, isolate shared services, scope credentials to the task, monitor in near real time, and budget the overhead, using the patterns in the guardrails and sandboxing deep-dive.

Sandbox for the agent you will have next year. A containment design that is adequate for today's agents will be inadequate for the next generation, and agents tend to be upgraded by changing a model name in a configuration file. Design isolation against capability you expect, not capability you have measured.

Access is now a policy variable

Gating by labs is only half of the access story; governments have entered it. On June 12, 2026 a US export-control directive required Anthropic to suspend access by foreign nationals to two newly released frontier models, citing national-security concerns. To comply, the company disabled both models for all of its customers. The restriction was lifted on June 30. It was the first known use of export-control authority to regulate access to a specific frontier model, and for two and a half weeks every workflow built on those models had to run on something else or not at all.

The episode turned a theoretical risk into an operational one. Access to the most capable models can now change on safety grounds, on regulatory grounds or on national-security grounds, with little notice and no defined restoration date. Usage policies around security work can also tighten as capability rises. Any business process that depends on a top-tier model needs a tested fallback, which is the subject of the AI supplier risk deep-dive.

The architect view

Cyber-capable models are the clearest case yet of a frontier development that lands on enterprise operations before it lands on enterprise strategy decks. Nothing about it requires waiting: the vulnerabilities are being found now, the patches are shipping now, and the attackers who will have similar capability are not subject to the gates.

Three moves cover most of the response. Treat patch capacity and exploitability-based prioritization as core resilience, measured and reported. Turn the strongest model you can access on your own code, inside a properly authorized and isolated program, and pursue eligibility for the verified-access programs that fit your role. And harden the agent estate against capable agents: default-deny egress, task-scoped credentials, near-real-time monitoring and fallbacks for any workflow whose model could be restricted.

On the radar this sits at two positions at once: adopt for the defensive basics, which are overdue regardless, and assess for the access programs and policy regime, which will keep moving. The organizations that fare best will be the ones that read 2026 as a starting gun rather than a headline.

← The Energy Cost of Intelligence: AI and Sustainability ALL OF FRONTIER & HORIZON Agentic Commerce Protocols: UCP, ACP, AP2 and the Payment Rails →