Atlas / BUILD / Agents / Agent Skills
DEEP-DIVE · AGENTS

Agent Skills: Packaging Know-How for Agents

A folder, a SKILL.md file and a one-line description became, in under a year, the cross-vendor way to teach agents how your organization does things. What skills are, where they fit next to MCP and system prompts, and how to run a skill library without importing someone else's malware.

TL;DR
  • A skill is a folder: a SKILL.md file with a name, a description and instructions, plus optional scripts and reference files. Agents load only the short description up front and pull in the rest when a task matches, so one agent can carry hundreds of skills without paying for them in every prompt.
  • Skills and MCP solve different problems: MCP gives an agent access to systems, skills give it the know-how to use them your way. Published as an open standard in December 2025, the format was read by more than thirty agent tools from competing vendors within three months.
  • Run a skill library like a code library: owned, reviewed, versioned and evaluated. Skills carry instructions an agent follows with real authority, so a careless or malicious one is a supply-chain risk, and public skill registries saw coordinated malicious uploads in early 2026.

What a skill is

Every enterprise deploying agents hits the same wall. The agent is broadly capable and knows nothing about how this organization does things: how finance closes a quarter, how a board memo is structured, which twelve steps follow a severity-one incident, what the brand guide says about charts. Until recently there were three unsatisfying places to put that knowledge. You could stuff it into the system prompt, where it costs tokens on every call and slowly rots. You could fine-tune it in, which is expensive, opaque and stale the day the procedure changes. Or you could wrap it in a tool, which works only for steps that are fully deterministic.

Agent skills are the fourth place, and in 2026 they became the default one. Anthropic introduced the format in October 2025 and published it as an open standard in December 2025. A skill is a directory containing a SKILL.md file: YAML frontmatter with two required fields, a name and a description, followed by Markdown instructions. Optional subfolders hold scripts the agent can run, reference documents it can read and assets such as templates and schemas.

The design idea that makes skills work is progressive disclosure. An agent does not load every skill it has. It loads them in tiers, paying for detail only when a task needs it:

 TIER 1  startup        name + description of every skill
         (always)       ~30-50 tokens each, e.g.
                        "quarterly-close: the steps finance
                         follows to close a fiscal quarter"
            |
            |  task matches a description
            v
 TIER 2  on demand      the full SKILL.md body: procedure,
                        rules, checks, when to stop
            |
            |  procedure calls for detail
            v
 TIER 3  as needed      references/ to read, scripts/ to run
                        (script output enters context,
                         not the script itself)
Progressive disclosure: an agent can hold hundreds of skills because it only pays for the ones a given task actually uses.

The economics follow directly. A library of a hundred skills costs a few thousand tokens of descriptions at startup, less than many system prompts, while the procedures themselves arrive only when relevant. Scripts add a second saving: a deterministic step such as reconciling two spreadsheets runs as code, and only its result enters the context window, which is both cheaper and more reliable than asking the model to do arithmetic in prose.

Where know-how livesLoaded whenChanged byBest for
System promptEvery callPlatform or product teamIdentity, scope and non-negotiable policy
SkillWhen a task matchesThe function that owns the procedureProcedures, playbooks, templates, judgment calls
Tool or MCP serverWhen the model calls itThe owning system's engineersAccess to systems; exact, deterministic actions
Fine-tuningBaked into weightsAn ML team, slowlyStyle, format and narrow skills at very high volume

Why skills spread so fast

Few standards have converged as quickly. Within days of the open specification, competing vendors began reading the same files, and by March 2026 thirty-two agent tools supported the format, from vendors including OpenAI, Microsoft, Google, JetBrains, AWS and Block. Four properties explain the speed.

The result is that a well-written skill is now one of the most portable assets in an AI program. The same folder serves a coding agent, a desktop assistant and a hosted agent platform, and it survives a change of vendor intact.

Skills, tools and MCP: a division of labor

The cleanest way to place skills is by what each layer contributes. The Model Context Protocol provides access: a standard way for the agent to reach the ticketing system, the CRM, the warehouse. Tools provide actions with typed contracts: create this ticket with these fields. Skills provide procedure and judgment: when a severity-one incident is declared, open the bridge, page these three roles, post the status template in this channel every thirty minutes, and file the postmortem ticket with these fields within two business days.

Two anti-patterns follow from blurring the layers. The first is building an MCP server for what is really a procedure, such as a write_board_memo tool that hard-codes one team's preferences into a service every agent must call. The second is writing a skill for what should be deterministic, such as prose instructions for currency conversion. The rule of thumb: if a step must be exactly right every time, make it a tool or a script inside the skill; if it requires context and judgment, make it instructions.

Skills also change what belongs in the system prompt. The system prompt should carry identity, scope and the policies that apply to every request. Situational procedures belong in skills. Moving them out shrinks the always-on context, reduces the drift and contradiction that long prompts accumulate, and lets the people who own each procedure maintain it directly.

Designing a skill that works

The single most important line in a skill is its description, because the agent decides whether to load the skill by reading that line next to the task. A vague description causes misses, where the skill exists but never triggers, and false triggers, where it loads for the wrong job and pollutes the context. Write the description as when to use the skill and what it does, in the vocabulary users actually type: "Use when preparing the quarterly board deck: builds the standard twelve-slide structure, applies brand rules, and checks figures against the finance export."

Write the description for the router, not the reader. Humans browse a skill library by title; agents select from it by description. Test descriptions the way you test search: a set of tasks that should trigger the skill, a set that should not, and a measured hit rate for both.

Beyond the description, five habits separate skills that hold up in production from demos:

Running an enterprise skill library

Once skills work, they multiply, and an enterprise needs the same disciplines it applies to shared code. Ownership comes first: each skill has an owner in the function that owns the procedure, so finance owns the close skill and legal owns the contract-review skill, while a platform team owns the library itself. Versioning comes second: skills live in a repository with changelogs, production agents pin specific versions, and agent platforms' administrative controls distribute approved versions to the right teams.

Review needs two lenses. The domain review asks whether the procedure is right. The technical review asks what the scripts do, which tools and data the skill touches, and whether its instructions could push an agent beyond its intended scope. Lifecycle completes the picture: when a procedure changes, the skill must change the same day, because an agent follows a stale skill with complete diligence.

Skills are also an unusually effective way to carry policy. Brand rules, mandatory disclosure language and compliance checks embedded in the skills every agent uses are applied consistently in a way that a training slide never is. Business-platform vendors have noticed: some now ship prebuilt skill packs for their own applications that inherit the platform's existing permission model, which is the right instinct and a pattern worth asking every vendor about.

Skills are code, even when they are prose

The property that makes skills useful makes them dangerous. A skill's instructions are executed with the authority of the agent that loads them: they can direct it to read files, call tools and send data, and its scripts run with the agent's privileges. A skill is therefore a dependency in exactly the sense a software package is, with one twist: much of its payload is natural language, which conventional scanners do not read.

The risk is not theoretical. In early 2026 a community registry for a popular open-source agent saw a coordinated wave of malicious uploads; by one security firm's count, roughly one in five of its approximately 4,500 skills had been weaponized within days. An independent audit of nearly 4,000 public skills found instructions hidden in ordinary-looking prose, in invisible Unicode characters and in conditional logic that triggers only in particular contexts. Researchers also documented malicious code planted in the example snippets that agents copy as reference implementations.

The controls are the ones you already apply to code dependencies, adapted for prose:

Never install a skill you would not merge. If a skill would fail code review as a pull request, it fails review as a skill. Natural language does not make an instruction less executable; it only makes it harder to notice, which is the core lesson of the prompt injection deep-dive.

The architect view

Skills are the cheapest and most portable unit of enterprise AI knowledge to emerge so far. They are written in plain language by the people who own the procedure, versioned like code, loaded only when needed and readable by every major agent. That combination makes them the natural home for the operational know-how that previously had nowhere good to live.

The strategic move is to turn the organization's runbooks into a governed skill library, and to recognize that this library will outlive any particular agent vendor because the format is an open standard. Four commitments get there. Stand up a single internal skill registry with owners and two-lens review. Put skill evaluation, both trigger accuracy and task success, into continuous integration. Apply supply-chain controls equal to those for code dependencies. And migrate procedures out of bloated system prompts and premature fine-tunes into skills, where their owners can maintain them.

The organizations that get the most from agents will not be the ones with the cleverest prompts. They will be the ones whose agents know how the organization actually works, and skills are now the standard way to teach them.

← Computer-Use Agents: When AI Operates the Screen ALL OF AGENTS Long-Running Agents: From Minutes to Days →