What are agent skills, and how do they work?
An agent skill is a folder of instructions that teaches an AI assistant how your organization does a task. Its core is a single text file that names the skill, says when to use it, and explains how the work is done. The format became an open standard in December 2025, and most major AI assistants and coding tools now load it.
The assistant reads each skill’s description first and loads the rest only when a task matches, so one line decides whether the skill is ever used.
Questions as you read? Ask, or take the context files with you.
Where should we send it?
Sent ✓ Download it now →
A skill is how your company does something, written down for an AI
Open an agent skill and there is not much to see. There is a folder, and in it a short text file. At the top sit a name and a sentence saying when the skill should be used. Below that come the instructions: how this piece of work gets done here, by the people who do it well. Sometimes a few examples sit alongside, or a template, or a small script for a step that has to run the same way every time.
The modesty is the point. A skill contains no model and no integration, and no code at all unless you choose to add some. It holds something your organization already knows: how a proposal gets structured, which checks a contract gets before it goes out, what a month-end summary for the board has to contain.
What has changed is where that file can go. The format became an open standard, and most of the major AI assistants and coding tools now load the same folder. A company can write a piece of its expertise down once and hand it to whichever assistant its people already use.
The file is the easy part, though. Whether a skill makes the AI better at the work depends on two things the format cannot supply. One is the expertise inside it, which has to come from the people who hold it. The other is that single line at the top, the description, which the AI reads to decide whether the skill applies to the task in front of it. Write that line badly and the skill never runs. Nothing tells you it didn’t.
Those three numbers carry the argument of this article. The format has spread widely. Written from the knowledge of people who do the work, a skill makes a measurable difference, and written without it, a skill can make things worse. And an agent can ignore a perfectly good skill more often than it uses it.
What is inside a skill
Every skill starts as a folder, and the only thing the folder must contain is a file called SKILL.md, according to the Agent Skills specification. It can also hold three optional subfolders: scripts for code the agent can run, references for longer documentation it can consult, and assets such as templates.
The file itself has two parts. The top is a short block of labeled fields that software can read. The rest is plain instructions for the AI, written the way you would brief a capable new colleague.
Only two of those fields are required by the specification. The name is a short identifier of up to 64 characters, in lowercase letters, numbers and hyphens, and it has to match the folder’s name. The description can run to 1,024 characters and has one job: to say what the skill does and when to use it. Four optional fields cover a license, the environment the skill needs, free-form metadata, and an experimental list of tools the skill may use.
Here is what a small skill looks like, in the structure the specification recommends. The folder is named for the skill, reviewing-proposals, and the file inside opens with its two required fields:
name: reviewing-proposals
description: Reviews a draft client proposal against the firm's scope, pricing and tone standards. Use when someone asks to check, review or tighten a proposal, statement of work or quote.
Below those lines come three short sections in plain text. First the steps: check that every deliverable has an owner, a date and an acceptance test; compare the price with the rules in a reference file; tighten the executive summary. Then one worked example of a weak phrase and its replacement. Then a short list of gotchas, such as work mentioned in the summary but missing from the price.
The standard’s own documentation calls a gotchas list the highest-value content in many skills, and the description in that sample does two jobs in two sentences. The first says what the skill does, in the third person, as Anthropic’s authoring guidance asks. The second starts with “Use when” and names the situations that should trigger it, which is the pattern Google recommends for the Gemini app. Compare the specification’s own example of a poor description, three words long: “Helps with PDFs.”
The instructions below those fields are where the value sits. The specification recommends keeping them under about 5,000 tokens and the whole file under 500 lines, and moving longer material into reference files the agent opens only when it needs them. Anthropic’s engineers compare writing a skill to “putting together an onboarding guide for a new hire,” which makes a good test of any draft. If a capable new colleague could follow it, an AI probably can too.
Scripts deserve a note of caution. They let a skill do more than advise, such as fill in a spreadsheet the same way every time, and they are also where most of the risk sits. The section on managing skills comes back to that.
The AI Briefing
Tuesdays. 500+ leaders. No hype, just what works.
How an AI decides to use a skill
An agent with a large library does not read every skill before it starts work. The standard works in three steps, which its specification calls progressive disclosure. At startup the agent loads only each skill’s name and description, roughly 100 tokens apiece. When a request matches one of those descriptions, it loads that skill’s full instructions. When the instructions point to a reference file or a script, it opens that only at the moment it needs it.
That design lets a library grow without crowding the AI’s working memory, and it is also why the description carries so much weight. The agent chooses a skill by reading those short summaries. It never sees the instructions of a skill whose description failed to match. A vague description means the skill gets skipped, and two overlapping descriptions mean the agent may pick the wrong one.
Anthropic’s own guidance for enterprise deployments warns that with too many skills active, Claude “may fail to select the right Skill or miss relevant ones entirely.” Vercel measured the same problem from the other side. In evaluations on its own framework, published in January 2026, the agent never invoked the available skill in 56% of cases.
A larger study found the same thing at scale. When Liu and colleagues had agents search a pool of 34,198 public skills instead of handing them the right ones, the strongest model’s gain over using no skills shrank from about 20 points to 3 (arXiv 2604.04323, April 2026). The authors traced part of the loss to agents failing to recognize relevant skills from their names and descriptions alone.
That pool was far larger than any company’s library, but the lesson scales down. Inside a company, the description is the part of a skill you control that decides whether the rest ever gets read.
A skill the AI never selects is a skill you do not have. The description decides.
Tools also differ on whether a skill may fire on its own. As of September 2026, OpenAI lets skills trigger implicitly by default, with a setting to turn that off per skill. Claude Code has a field that stops a skill from loading automatically. Gemini CLI asks the user to confirm before a skill activates.
Microsoft’s Agent Framework requires approval by default before an agent loads a skill, reads one of its resources or runs one of its scripts. Same file, four different answers to the question of who pulls the trigger.
Where skills run
Anthropic introduced skills in October 2025 and published the format as an open standard that December, arguing that skills should be portable across tools and platforms. We build on Claude ourselves, so we should say plainly what that authorship means today. Anthropic still maintains the specification and holds its name, its website and its code repository. In September 2026 it proposed handing it to the Agentic AI Foundation, a Linux Foundation body, where the proposal was awaiting a vote at the time of writing.
The other major vendors took the format up. As of September 2026, this is where a skill can run:
| Where | How a skill gets there | Keep in mind |
|---|---|---|
| Claude: the app, Claude Code, the API | Upload in the app, a folder in the project for Claude Code, or an upload through the API | Skills in a Claude account now load in signed-in Claude Code sessions, one way only; API skills stay separate |
| ChatGPT and Codex | Workspace skills once an owner switches them on; skill folders for Codex | Still labeled an initial beta, off by default |
| GitHub Copilot and VS Code | A skills folder in the repository or in the user’s home folder | Also reads the folder Claude Code uses |
| Cursor | Project or user skill folders | Also loads Claude’s and Codex’s skill folders |
| Microsoft Copilot Studio | Upload a SKILL.md or a zip file to an agent | Only for agents on the GitHub Copilot harness, managed one agent at a time |
| Gemini Enterprise | Create or upload in the web app | Sharing works only if an administrator enables it |
| The Gemini app | Upload a folder with SKILL.md at its root | Personal Google accounts only |
The five context files.
Fill-in templates for what your AI should know, and the page that says where each one loads.
On its way. Check your inbox, or download it now →
The standard’s own showcase lists 46 products that support the format, 44 of them built outside Anthropic, and it leaves out several of the adopters in that table. Treat the number as a floor.
What carries over is the core: the folder, the name, the description and the instructions. What does not carry over is everything each vendor adds around it.
Claude Code accepts extra fields that make an upload to the Claude app fail outright. OpenAI keeps its invocation settings in a separate file of its own. The folders each tool looks in differ, although several now read one another’s. And the same instructions can behave differently on different models, which is why Anthropic tells enterprises to test skills on every model they use.
So write to the core, and treat each vendor’s extras as local settings. If your company is choosing an assistant, the skill format matters less to that decision than it did a year ago, which takes one item off the list we work through in ChatGPT vs Claude vs Copilot vs Gemini. And if the worry is lock-in, the core file is now one of the easier things to carry across when switching AI models.
One limit on our own evidence belongs here. Our library is written for Claude, and we have not moved it to another vendor. The portability described above rests on the vendors’ documentation, not on a migration of our own.
Skills and the things you already use
Most people meet skills after they have already used something that sounds similar. The differences are simple once they are laid side by side.
Prompts and custom instructions
A prompt lives in one conversation and disappears with it. Custom instructions, or a system prompt, apply to every conversation whether they are relevant or not. A skill sits between the two. It is written once, like instructions, but it loads only when the task calls for it, which is what lets a company keep a large library without crowding every conversation. GitHub puts it practically: use custom instructions for simple rules that apply to almost every task, and skills for detailed instructions that matter only some of the time.
Custom GPTs, Projects and Gems
A Project, in Claude or in ChatGPT, is a shared workspace: files, background knowledge and conversations gathered around one goal and loaded every time you work inside it. Anthropic describes Projects as “static background knowledge that’s always loaded” in the chats within them. A custom GPT or a Gem is an assistant you configure and then open on purpose. A skill is neither a place nor a persona. It is a procedure the assistant picks up by itself, anywhere, when a task matches its description.
The vendors are now folding the older forms into the new one. When a company migrates a custom GPT, OpenAI turns the GPT’s instructions into a skill and moves its files and connected apps into a plugin. Google says Gems on personal accounts will become skills from November 2026. Microsoft draws the cleanest line: an agent has instructions for its general behavior, knowledge for its data, tools for its actions, and skills for reusable, task-specific capabilities.
A custom assistant is a place you go. A skill is expertise the AI brings to wherever you are working.
Tools and MCP
A tool is a single action the AI can take, such as running a search or updating an entry in a system. MCP, the Model Context Protocol, is the standard way to connect an AI to outside systems and the tools they expose. A skill supplies the know-how for using them: which tool, in what order, to what standard.
The talk that introduced skills put the division simply, with MCP “providing the connection to the outside world, while skills are providing the expertise.” OpenAI’s documentation draws the same line, describing the skill as the part that “explains how to complete the workflow.” OpenAI also packages skills and MCP servers together in installable bundles it calls plugins.
Agents
An agent is the runtime: the AI that takes a task, decides what to do and carries it out. Skills are what it knows how to do well. One capable agent with a good library of skills is easier to maintain and improve than a fleet of specialized agents, an argument we make at length in AI skills versus agents.
The five context files.
Fill-in templates for what your AI should know, and the page that says where each one loads.
On its way. Check your inbox, or download it now →
How to write a skill that actually gets used
Most of what makes a skill work happens before anyone opens a text editor. The clearest evidence comes from SkillsBench, a preprint benchmark built by a group of 77 researchers and last updated in June 2026. Across 87 tasks, skills curated from real material raised the average pass rate from 33.9% to 50.5%. Skills the agent wrote for itself, with no human material, scored between 8 and 12 points below using no skill at all.
The model can write the format. It cannot supply the expertise. SkillsBench also found that compact skills beat exhaustive ones, with comprehensive documentation adding almost nothing, and that even curated skills made 13 of its 87 tasks worse. A skill is not automatically an improvement, which is why the testing rule below matters most.
Start with one task, and the right kind. Pick a piece of recurring work that one person does well and others do inconsistently. Google’s guidance for the Gemini app names the cases where a skill is the wrong tool: one-off jobs, tasks the assistant already handles well, and steps that change too often to keep current.
Build it from your own material. The standard’s documentation advises starting from real expertise, such as your runbooks, incident reports and review comments, because a skill generated from general knowledge comes out vague. Write down only what the AI would get wrong without it.
Write the description before the instructions. Say what the skill does, then when to use it, with the situations named in the words people actually use. Put the main use case first, because tools cut descriptions short when many skills are installed: as of September 2026, Codex shortens descriptions once its skill list outgrows 2% of the context window, and Claude Code starts dropping them at 1%. Give each skill one job.
Samuel Berthe, a software engineer and Go open-source maintainer, puts the point more sharply than any vendor: “Skills aren’t docs. They’re APIs.” In his framing the description is the function signature, and it fails in three ways. The skill does not trigger when it should, triggers when it should not, or competes with another skill for the same request.
Keep the instructions short and move the detail out. Anthropic’s guidance keeps the main file under 500 lines, with longer material in reference files one level deep, and the standard’s documentation adds that the file should say when to open each one. Resist turning every step into a mandatory checklist. A study that traced 307 failures back to specific skills found that they turned checklists and recipes into mandatory extra work.
Put the steps that must never vary into scripts, and keep the judgment in words. A script fills in the same fields the same way every time. The words carry what a script cannot: when an exception applies, what good looks like, when to stop and ask.
Test before you share. Run the same task with and without the skill and compare the results. Then test the trigger: the standard’s documentation suggests about 20 realistic requests, roughly half that should load the skill and half near misses that should not. Repeat on every model your people use, since the same file can behave differently on each.
Give every skill an owner. A skill carries someone’s judgment, and when that person’s practice changes and the file does not, the AI goes on doing the old thing with complete confidence. Anthropic’s own enterprise guidance tells companies to record each skill’s purpose, owner and version. The fix is not technical. Each skill needs a named person who keeps it current, which is the practical meaning of Humans First AI here: the AI extends what your people know, and your people remain the ones who own it.
The model can write the format. It cannot supply the expertise.
What skills are good for in a company
The clearest uses are the ones where the same expertise gets applied again and again, by different people, with results that vary more than anyone would like.
The vendors’ own examples cluster in a few places. Documents come first: Anthropic ships ready-made skills for creating and editing PowerPoint, Excel, Word and PDF files. Then standards that have to hold across people. OpenAI’s examples for business teams include a monthly close narrative, a finance memo standard and a brand voice polish, and Microsoft suggests packaging expense policies and legal workflows.
Then multi-step work across spreadsheets. In Anthropic’s launch post, Rakuten’s general manager for AI said skills now run parts of its management accounting and finance workflows, and that “what once took a day, we can now accomplish in an hour.” That is Rakuten’s own claim, published by the vendor and not independently audited.
In the talk that introduced skills, Anthropic’s engineers said Fortune 100 companies were using them to teach agents their organizations’ best practices. They also described developer productivity teams, some serving tens of thousands of developers, using skills to teach coding agents the house style for code.
Closer to home, the candidates tend to be ordinary and valuable at once. These are our examples of where skills fit well: the structure and pricing logic of a proposal, the checklist a contract goes through before it is signed, the format of a month-end report, the answers a new hire needs in the first month.
We run our own firm this way. When we write an article, a skill carries our voice, our structure and our citation rules. When we draft a proposal, a skill carries our positioning and pricing logic. Month-end numbers go through a skill that encodes how we categorize what we spend. None of those skills is clever. Each one is simply how we do that piece of work, written down well enough that the AI does it our way.
Managing skills across a team
The major platforms each give administrators real control over skills. Use what is there before building anything of your own. As of September 2026:
On Claude’s Team and Enterprise plans, owners can provision a skill to everyone, switch off the skills users create themselves, require an owner’s review before anything is published to the shared library, and see sharing in the audit log. Enterprise plans also scan uploaded skills, by default from October 2, 2026, though not skills added through the API.
In ChatGPT’s business workspaces, skills are still an initial beta: off until a workspace owner switches them on, with owners deciding who may create, share and install them for others. Gemini Enterprise lets people share a skill only if an administrator has enabled sharing, and share requests wait for approval unless the administrator waives it. Administrators there can suspend, disable or delete skills. Microsoft manages skills inside Copilot Studio one agent at a time, and governs the agents themselves centrally through Agent 365.
Our article on AI agent governance covers who answers for what an agent does, and the short version holds here: the platforms have largely solved identity and logging. Skills add one decision of their own, which is which skills your people may install. The consumer Gemini app shows why that matters. It accepts skills on personal accounts only, so a skill installed there sits outside every control above, which is shadow AI with a new file type.
The riskiest skills are the ones nobody in your company wrote. A study of 31,132 skills from two public marketplaces found that 26.1% had at least one vulnerability, and that skills bundling scripts were 2.12 times as likely to have one (Liu and colleagues, January 2026).
A scan by the security vendor Snyk found 76 confirmed malicious payloads among 3,984 marketplace skills. And a July 2026 paper showed that a simple packing technique got past each of eight skill scanners more than 90% of the time, so scanning alone does not settle it. GitHub, for its part, states that it does not verify the skills people add.
Those are marketplace numbers, not a base rate for skills your own people write. They still set the rule for anything that arrives from outside.
Treat a skill you did not write the way you would treat code you did not write: read it before you run it.
Where outside help fits
Most of this a company can do on its own. Writing a first skill needs no consultant, and the controls above are settings, not projects.
The library for a whole team is harder from the inside. Someone has to decide which expertise gets written down first, whose version of a process becomes the skill when two departments do it differently, and who keeps each one current. Those are questions about how many people have to agree, and they are political before they are technical.
That is the work we do when we build an AI operating system around a company’s skills. Whoever you bring in, apply one test. At the end, does your team own the library and know how to keep it current, or does the outside firm?
Write Your First Skill
Pick one task that someone on your team does well and others do inconsistently. Paste this into your AI assistant. It interviews you and drafts a skill file you can test right away.
Context: I want to write my first agent skill: a SKILL.md file that teaches an AI assistant how our organization does one recurring task. The task is [DESCRIBE IT]. Interview me one question at a time, push back when my answers are generic, and draft only from what I tell you.
Step 1. The trigger: Ask how people phrase this request when they need it done, and which similar requests should not use this skill.
Step 2. The standard: Ask what a good result looks like, with one example of a result that would be sent back, and why.
Step 3. The method: Ask for the steps in order, the checks, and the exceptions a newcomer usually misses.
Step 4. The materials: Ask which templates, examples or reference documents the skill should point to.
Output: A SKILL.md with a lowercase, hyphenated name; a one or two sentence description that says what the skill does and when to use it; instructions short enough to follow; a list of reference files still to create; and three test requests, two that should trigger the skill and one that should not.
Test the trigger before you share it: open a fresh conversation, make one of the test requests without naming the skill, and see whether it loads. A library that a whole team relies on is a different job, and that is where the editorial work starts. See where you stand →
Sources
- Agent Skills specification (agentskills.io, read September 2026): required and optional fields, folder structure, loading tiers, size guidance and the poor-description example.
- Optimizing skill descriptions and Best practices (agentskills.io, maintained under Anthropic, read September 2026): the description as the trigger, trigger testing, gotchas, real expertise.
- Agent Skills client showcase (agentskills.io, read September 2026): the 46 products listed as supporting the format.
- Introducing Agent Skills (Anthropic, October 16, 2025): the launch, and Rakuten’s vendor-published claim.
- Skills for organizations, partners, the ecosystem (Anthropic, December 18, 2025): the open standard and organization-wide management.
- Proposal: Agent Skills (Anthropic filing to the Agentic AI Foundation, September 24, 2026): how the specification is governed today, and the pending move.
- Skill authoring best practices (Claude Platform documentation, read September 2026): third-person descriptions, the 500-line guidance, reference files one level deep.
- Skills for enterprise (Claude Platform documentation, read September 2026): selection degrading with too many skills, testing across models, recording owners.
- Agent Skills overview (Claude Platform documentation, read September 2026): the pre-built document skills.
- Equipping agents for the real world with Agent Skills (Anthropic Engineering, October 16, 2025): the onboarding-guide comparison.
- What are skills? and Provision and manage skills for your organization (Claude Help Center, read September 2026): skills versus Projects, and organization controls.
- Extend Claude with skills (Claude Code documentation, read September 2026): description budgets, automatic loading, sync from a Claude account.
- Build skills and Migrate custom GPTs (OpenAI, read September 2026): the open standard, invocation settings, description guidance, and GPT instructions becoming skills.
- Skills (OpenAI Academy, updated September 17, 2026): the initial beta, owner controls, business examples.
- Skills (OpenAI Developers, read September 2026): skills versus MCP servers.
- About agent skills and Add skills (GitHub Docs, read September 2026): the folders Copilot reads, custom instructions versus skills, no verification of added skills.
- Agent Skills (Cursor documentation, read September 2026): Cursor’s skill folders, including Claude’s and Codex’s.
- Skills overview for agents (Microsoft Copilot Studio documentation, read September 2026): SKILL.md upload, the four parts of an agent, example uses.
- Write effective skills, Gems and Create and manage skills (Google, read September 2026): description guidance, when not to write a skill, Gems becoming skills, Gemini Enterprise controls.
- SkillsBench (Li and 76 co-authors, arXiv 2602.12670, version of June 14, 2026, preprint): curated versus self-generated skills, length effects.
- How Well Do Agentic Skills Work in the Wild (Liu and colleagues, arXiv 2604.04323, April 2026, preprint): skill selection from a pool of 34,198 public skills.
- Dong and colleagues (arXiv 2608.11888, August 2026, preprint): failures traced to specific skills.
- AGENTS.md outperforms skills in our agent evals (Vercel, January 27, 2026): the skill was never invoked in 56% of cases.
- Agent skills are an API design problem (Samuel Berthe, March 24, 2026): the description as the skill’s interface.
- Security analysis of public agent skills (Liu and colleagues, arXiv 2601.10338, January 2026, preprint; a different group from the April paper): 31,132 marketplace skills, 26.1% with at least one vulnerability.
- ToxicSkills (Snyk, a security vendor, February 5, 2026) and scanner evasion study (arXiv 2607.02357, July 2026, preprint): malicious payloads, and packing that evaded eight scanners.
- Don’t Build Agents, Build Skills Instead (Barry Zhang and Mahesh Murag, AI Engineer Code Summit, November 2025): the enterprise uses and the MCP distinction.
Frequently Asked Questions
What are agent skills?
Agent skills are folders of instructions that teach an AI assistant how your organization does a specific task. Each one centers on a text file, SKILL.md, with a name, a description of when to use it, and the instructions themselves, plus optional examples, reference files and scripts. The assistant loads a skill only when a task matches its description, so a company can keep a large library without slowing every conversation.
Are agent skills an open standard?
Yes. Anthropic published the skill format as an open standard in December 2025, and it is documented at agentskills.io. Anthropic still maintains the specification, and in September 2026 it proposed moving it to the Agentic AI Foundation under the Linux Foundation, pending a vote. OpenAI, GitHub, Microsoft, Google and Cursor all document support for the format.
Do skills work across ChatGPT, Claude, Copilot and Gemini?
The core file does. As of September 2026, the same SKILL.md folder loads in Claude, ChatGPT and Codex, GitHub Copilot and VS Code, Cursor, Microsoft Copilot Studio for agents on its GitHub Copilot harness, Gemini Enterprise, and the Gemini app for personal accounts. Each vendor adds its own settings around the file, such as invocation rules and folder locations, and those do not carry over.
How is a skill different from a custom GPT or a Gem?
A custom GPT or a Gem is an assistant you set up and then open on purpose. A skill is a procedure any compatible assistant can pick up by itself when a task matches its description, and several skills can work together in one task. The two are converging: OpenAI turns a custom GPT’s instructions into a skill when you migrate it, and Google says Gems on personal accounts will become skills from November 2026.
How do I write a good skill description?
Say what the skill does in one third-person sentence, then list the situations that should trigger it, starting with “Use when” and using the words people actually type. Put the main use case first, because some tools shorten descriptions when many skills are installed, and give each skill one job. Then test it with realistic requests, including a few near misses that should not trigger it.
Are third-party skills safe to install?
Treat them like code from an unknown source. A study of 31,132 skills from public marketplaces found that 26.1% had at least one vulnerability, and skills that bundle scripts were more than twice as likely to. Read a skill before you install it, prefer skills you wrote or got from a vendor you trust, and use your platform’s review and provisioning controls so that installing one is a decision rather than a habit.
How many skills can an AI agent handle?
There is no single number. Each skill costs the agent roughly 100 tokens at startup, and Claude’s API accepts up to 20 skills per request as of September 2026. The practical limit is selection: tools such as Codex and Claude Code shorten descriptions when many skills are installed, and Anthropic warns that with too many skills active, the agent may miss the right one. Clear, non-overlapping descriptions matter more than the count.
Can AI write a skill for me?
It can draft the format, but it cannot supply the expertise. In SkillsBench, a 2026 benchmark, skills an agent wrote for itself without human material scored below using no skill at all, while skills curated from real material raised the average pass rate by 16.6 points. Use the AI to interview you and structure what you know, then test the result against doing the task without it.
What is the difference between a skill and an agent?
An agent is the runtime: the AI that takes a task, decides what to do and carries it out. A skill is what it knows about doing a specific kind of work your organization’s way. One capable agent with a library of skills is easier to maintain and improve than a fleet of specialized agents, because every improvement to a skill reaches every future use of it.


