- An agent's capabilities come from components it loads at run time: MCP servers, skills, plugins, repository instruction files, models and frameworks. Many of them carry natural-language instructions the agent obeys, so the agent supply chain is made of instructions as much as code, and OWASP made it a top-ten agentic risk in December 2025.
- The record is already real: a malicious MCP server that silently copied users' email in 2025, a coordinated wave of weaponized skills on a public registry in early 2026, and thousands of agent servers exposed to the internet without authentication. Compromises use permissions the agent legitimately holds, so they look like normal tool use.
- The controls are software supply-chain discipline plus three additions: review prose as well as code, detect changes in remote components after approval, and restrict egress so stolen data has nowhere to go. Keep an agent bill of materials so you can find every agent that uses a compromised part.
A supply chain made of instructions
A modern agent is assembled more than it is written. Its model comes from a provider or a model hub. Its access to systems comes from MCP servers, some installed as local packages and some called as remote services. Its procedures come from skills. Its behavior inside a codebase is shaped by instruction files such as AGENTS.md. Coding agents and IDEs add plugins and extensions from marketplaces. Underneath sit the agent framework and its dependencies. Every one of these is a supply-chain link, and most were written by someone outside your organization.
What makes this supply chain different from the one software teams already defend is that many of its components carry instructions. A tool description, a skill, an instruction file or a README tells the agent what to do, and the agent does it with whatever permissions it holds. A traditional malicious package has to exploit something. A malicious instruction only has to be read.
| Component | Typical source | What a compromise can do |
|---|---|---|
| Local MCP server | Package registries, code hosts | Run code with the user's privileges; read local secrets |
| Remote MCP server | Third-party services | See every request; poison tool descriptions; change after approval |
| Skill | Community registries, vendors | Instructions executed with the agent's authority; bundled scripts |
| Plugin or extension | Tool marketplaces | Code and instructions inside the developer toolchain |
| Instruction file | Any repository or document the agent reads | Redirect the agent's behavior whenever it works in that context |
| Model artifact | Model hubs | Backdoored behavior; unsafe serialization that runs code on load |
| Framework or SDK | Package registries | Everything above, across every agent built on it |
The incident record so far
The threat moved from research papers to incident reports within about a year. In September 2025 researchers found what is widely described as the first malicious MCP server in the wild: an update to a popular community server for a transactional email service that silently blind-copied every message it handled to an attacker-controlled address. Through 2025, critical remote-code-execution flaws were also disclosed in widely used MCP developer tooling, a reminder that the plumbing is as exposed as the plugins.
Early 2026 brought scale. In February a community skill registry for a popular open-source personal agent saw a coordinated wave of malicious uploads; by one security firm's count, roughly one in five of its approximately 4,500 skills had been weaponized within days. An independent audit of nearly 4,000 public skills that month found instructions hidden in ordinary prose, in invisible Unicode characters and in conditional logic that activates only in particular contexts. Internet-wide scans counted thousands of exposed MCP servers, around half with no authentication, and tens of thousands of publicly reachable instances of one open-source agent, many vulnerable to remote code execution.
The standards bodies registered the shift. The OWASP Top 10 for Agentic Applications, published in December 2025, lists agentic supply-chain vulnerabilities as one of its ten risks, covering compromised or tampered agents, tools, plugins, registries and update channels, alongside the supply-chain entry already present in OWASP's list for LLM applications.
Why agent components are harder to vet
Six properties make agent components harder to secure than ordinary dependencies, and each one defeats a control teams rely on today.
- Natural-language payloads. Software composition analysis reads code and version numbers. It does not read a skill's instructions or a tool's description, where the payload often lives.
- Contextual behavior. A component can behave perfectly during review and maliciously only for certain users, files or phrases, which static review rarely triggers.
- Mutable remote components. You vet a remote MCP server once and call it every day. Its operator can change its tool descriptions or behavior after approval, the pattern known as a rug pull.
- Amplified privilege. A compromised component inherits whatever the agent can reach, and agents are often granted broad access across many systems at once.
- Volume and weak vetting. Community registries hold hundreds of thousands of skills and servers. Listing is easy; review is rare.
- Thin provenance. Publisher identity is often unverified, signing is uncommon, and forks and lookalikes are trivial to create.
Attack patterns to design against
The techniques seen so far fall into a handful of patterns. Tool poisoning hides instructions in a tool's description that steer the model when the tool is merely listed, not even called. Rug pulls change an approved component's behavior or descriptions later. Typosquats and lookalikes exploit near-identical names. Malicious updates of genuinely popular components follow the takeover of a maintainer's account. Instruction injection through repositories plants directives in AGENTS.md files, READMEs or comments that a coding agent will read. Poisoned examples hide malicious code in the reference snippets that agents copy. And exfiltration through legitimate channels moves data out through an email, a webhook, a repository push or an image URL the agent was always allowed to use.
attacker publishes, or hijacks, a component
|
v
[ REGISTRY ] skill | MCP server | plugin weak vetting
|
| someone installs it: it looks useful
v
[ AGENT ] loads its instructions and tools
| with the agent's own permissions
|
+--> reads data it is allowed to read
|
+--> calls a tool it is allowed to call
| (email, webhook, git push, HTTP fetch)
v
data leaves through a legitimate channel;
the logs show ordinary tool use
The figure explains why runtime detection alone disappoints. Nothing in the attack path is anomalous at the level of individual permissions: the agent read data it may read and called a tool it may call. The two places where the attack is visible are the moment the component entered the estate and the moment data tried to leave it.
Controls that work
The defensive playbook is the established software supply-chain discipline, extended with three additions specific to agents.
- One source. Production agents load components only from an internal registry or mirror. Public components enter through review, never directly. An official public registry is useful for discovery, but listing is not vetting.
- Pin, verify, diff. Reference exact versions, verify hashes and signatures where they exist, and require re-approval when any tool description, skill text or instruction file changes. For remote servers, the gateway records the approved descriptions and flags drift on every connection, which turns rug pulls into alerts.
- Review the prose. Treat instructions as code under review: scan for hidden Unicode, encoded content and suspicious URLs, and have a human or a dedicated reviewer model read skills and descriptions for scope creep and exfiltration patterns, as the agent skills deep-dive recommends.
- Isolate by default. Run local servers and skill scripts in sandboxes with scoped credentials and no network unless granted, using the patterns in the guardrails and sandboxing deep-dive.
- Separate the dangerous combination. An agent that holds private data, reads untrusted content and can communicate externally has all three ingredients of what Simon Willison named the lethal trifecta. Remove one ingredient per agent wherever the design allows.
- Mediate at run time. Route tool calls through a gateway that logs which component and version acted, enforces policy, and can disable a component across the fleet in one action.
Process completes the picture. Remote servers from third parties belong in vendor due diligence like any other data processor: data handling, retention, security attestations and change notification. And incident response needs an agent-specific runbook: disable the component everywhere, revoke and rotate the credentials it could reach, identify every agent that loaded it, and review what those agents did while it was active. The threat modeling deep-dive places these controls in the broader GenAI threat picture.
Inventory: the agent bill of materials
The incident runbook above has a hidden dependency: knowing which agents loaded the compromised part. That requires an agent bill of materials, a per-agent inventory of the model and its version, every MCP server with version and endpoint, every skill with its hash, plugins, instruction files, framework versions and the datasets or indexes it reads. The software world's bill-of-materials standards have already grown in this direction: CycloneDX covers machine-learning components, and SPDX 3.0 adds AI and dataset profiles.
The bill of materials earns its keep three ways. It makes incident response a query instead of an investigation. It gives auditors and model risk teams a concrete record of what each agent is made of. And joined to the registry in the agent control plane deep-dive, it lets policy say things like "no production agent may load a skill that has not been reviewed in the last ninety days" and have the statement enforced.
The architect view
The agent supply chain is the software supply chain with an additional layer made of natural language, and most of the defensive investment enterprises have made since the major software supply-chain incidents of the past decade still applies. What does not transfer automatically is the assumption that the dangerous part of a dependency is its code.
Four commitments close most of the gap. Make an internal registry the single source of agent components for production. Pin, verify and diff every component, including the prose, and alert on drift in remote ones. Sandbox agent workloads and restrict their egress by default. And maintain an agent bill of materials linked to the agent inventory, so that the question every incident starts with has a fast answer.
Adoption of agent components is growing faster than the vetting around them, which is exactly the condition attackers look for. The organizations that come through the next few years without a headline will be the ones that treated a skill like a package, a tool description like code and an outbound connection like a privilege.