When to Build an Agent Skill, When to Use a Prompt, and What You’re Running When You Install One


lines of HTML codes

“Write a skill for that” is the most common piece of advice in agentic coding right now, and most of the time it is the wrong move. The instinct is understandable. Agent Skills are the mechanism everyone is shipping into their coding agents in the back half of 2026, and Anthropic’s own Claude Code changelog has touched the skill system almost every working day this month: version 2.1.289 on October 3 fixed skills not resolving when the folder name differed from the name inside SKILL.md, 2.1.290 taught Claude to recognize skills typed mid-message, and 2.1.295 on October 8 patched a bug where a skill’s allowed-tools and effort settings were silently dropped in headless runs (Claude Code changelog). A feature getting that much maintenance attention is a feature people are leaning on hard.

The problem is that very few teams can say what a skill does that a prompt, an AGENTS.md entry, or an MCP server does not, which means they build skills for things that should not be skills, pay a context cost they never measured, and install executable folders from a marketplace most of them have never audited. Here is the rule I use, and the two questions that actually matter before you write one.

A skill costs almost nothing until Claude needs it

Anthropic defines Agent Skills as “organized folders of instructions, scripts, and resources that agents can discover and load dynamically to perform better at specific tasks” (Anthropic engineering). The folder holds one required file, SKILL.md, which opens with YAML frontmatter carrying a name and a description, followed by the instructions themselves, plus optional scripts/, references/, and assets/ subdirectories.

The part that matters is how little of that the model reads at any given moment. Skills use what Anthropic calls progressive disclosure, in three levels. At startup, only each skill’s name and description load into context, a few dozen tokens apiece (firecrawl analysis). The full SKILL.md body loads only when a task matches that description. The bundled scripts and reference files load only at the moment they are executed or read, and a file the agent never touches costs zero. Because the agent has filesystem and code-execution tools, Anthropic notes the context you can bundle into a skill is “effectively unbounded”: the heavy material sits on disk until the exact step that needs it.

This is the whole reason skills exist, and it is worth saying plainly because it connects directly to a finding this site covered two weeks ago. An ETH Zurich preprint found that a repository context file (AGENTS.md) raises inference cost by over 20% on every run because the entire file loads every time, whether or not the task needs it. Skills invert that arrangement. The instruction sits dormant at a cost of its one-line description until the moment it becomes relevant. A skill is not a smarter prompt. It is a cheaper place to keep a procedure you only need sometimes.

Four places an instruction can live

Before writing a skill, decide whether the thing you are encoding actually belongs in one. There are four homes for agent instructions, and they are not interchangeable.

A prompt is for anything you can say in a sentence or two for this task only: “use British spelling,” “return JSON.” If it is one-off, it goes in the prompt and nowhere else.

AGENTS.md (or CLAUDE.md) is for the small set of non-inferable facts true across the whole repository, every session: a non-standard test command, a generated directory that looks editable but is not, a house rule that contradicts the language default. The research is blunt that this file should stay short, because every line is re-sent on every run.

An Agent Skill is for a repeatable procedure that is not always relevant and benefits from being packaged: your house format for a quarterly report, a PDF-to-structured-data extraction, a deploy runbook with an exact step order, a linting-and-fix routine. The test is two-sided. It has to be a procedure you reach for more than once (otherwise it is a prompt), and it has to be dormant most of the time (otherwise it belongs in AGENTS.md and pays the always-loaded cost on purpose).

An MCP server is for a live connection to a system of record: your database, your ticketing tool, your internal API. If the thing you want is fresh data or an action against an external system, that is a tool, not a skill. Skills carry knowledge and scripts; MCP carries connections. Teams blur this constantly and end up with a skill that hardcodes a credential, which is the start of a security problem, not a capability.

The dividing question for the skill box specifically: is this a procedure the agent should run the same way every time, that it does not need on most turns? If yes, it is a skill. If it needs it on every turn, it is AGENTS.md. If it is a connection, it is MCP.

When a skill earns its place

I spent two decades in enterprise IT operations, and the closest analog I have for a skill is the runbook. The runbooks that earned their place in the binder were never the ones restating how the system worked; they were the ones documenting the single non-obvious step that saved a 3 a.m. shift. A skill is that runbook with one difference that changes everything: it can run its own scripts. It is a runbook with hands.

That difference is also where the value is. When a skill bundles a Python script to pull form fields out of a PDF, the agent runs the script instead of generating the extraction token by token, which is faster and deterministic. The design brief Anthropic gives is the same one I would give a junior engineer writing operational docs: evaluate the capability gap before building anything, keep each SKILL.md lean (the recommended body cap is around 5,000 tokens), split files when they get unwieldy, and iterate on real usage rather than guessing. One preprint measured the median skill body in a large marketplace at 1,414 tokens (arXiv 2602.08004, a data-driven analysis from Bosch Research and Carnegie Mellon, not peer-reviewed), which is about right: enough for a real procedure, not a manual.

The single most load-bearing field is the description. It is the only part of the skill the model sees until it decides to load the rest, so it is the trigger. A vague description (“helps with code”) will either never fire or fire on everything. A specific one (“format commit messages using the conventional commits spec; use when asked to write, review, or fix a commit message”) fires at the right moments and stays silent otherwise. Getting this wrong is not just an accuracy problem. Every installed skill’s description is re-sent at the start of every session, so a drawer full of junk skills with sloppy descriptions quietly recreates the same always-loaded tax that sits at the center of any serious AI cost conversation. Picture a hundred engineers running ten agent sessions a day against fifty installed skills: that is tens of millions of tokens a month spent just announcing skills, most of which never fire. Prune the drawer the way you would prune a cron table.

A skill is an executable folder you downloaded

Here is the second question, the one the enthusiasm skips. A skill is an open standard now: Anthropic published the SKILL.md spec at agentskills.io on December 18, 2025, and within weeks OpenAI’s Codex CLI, Google’s Gemini CLI, GitHub Copilot, Cursor, and VS Code all read the same format, with partner-built skills from the likes of Canva, Stripe, and Notion. Installing one is often a single command (npx skills add owner/repo). That convenience is the risk. You are downloading a folder that can contain executable Python, Bash, or JavaScript, and handing it to an agent that will run it.

Two preprints have now measured what is actually in these marketplaces, and both are worth reading with the caveat that neither is peer-reviewed. The Bosch and Carnegie Mellon analysis catalogued 40,285 publicly listed skills (the marketplace it studied grew roughly 18.5 times in 20 days in early 2026) and graded 9% of them critical-risk, with nearly two in five able to access sensitive context or perform writes and actions (arXiv 2602.08004). A separate security study of 31,132 skills across two other marketplaces found 26.1% carrying vulnerabilities, 5.2% showing high-severity patterns consistent with malicious intent, and, most usefully for a practitioner, that skills bundling executable scripts were 2.12 times more likely to contain a vulnerability than instruction-only skills (arXiv 2601.10338, “Agent Skills in the Wild,” January 15, 2026, also a preprint). The attack classes are the familiar ones: hidden prompt-injection instructions, environment-variable and credential harvesting, privilege escalation, and supply-chain tricks like fetching an external script at runtime.

This is the same lesson the site drew from the Claude Code mods advisory and the Plugin4Shell supply-chain story, and it has the same answer. The spec gives you scoping tools: disable-model-invocation stops a skill with side effects from loading automatically, allowed-tools limits what it can call, and paths restricts where it activates. But the standard does not enforce a permission model at the platform level, so the burden of scoping a skill safely falls on whoever writes and installs it. Anthropic’s own guidance is to install skills only from sources you trust. Treat an installed third-party skill exactly as you would a third-party service account or an unvetted dependency: default to the narrowest scope, read the scripts before you run them, pin what you install, and keep an eye on the ones that can write or reach the network.

The build order

If a procedure passes both tests (repeatable but not always-on, and either yours or from a source you have read), the build is short. Write the tightest description you can, because that sentence is the whole trigger. Keep the body under the cap and push anything long into references/ so it loads only when needed. If there is a deterministic step, make it a script rather than prose the model regenerates each time. Set allowed-tools and disable-model-invocation deliberately for anything with side effects. Then do what the AGENTS.md research insisted on and almost nobody does: run ten or twenty real tasks with and without the skill, and check that it fired when it should have and that the token cost moved the way you expected.

The reason “write a skill for that” became a reflex is that skills are genuinely the right tool for a specific job: a procedure you run more than once, need rarely, and want executed the same way every time. The reason it is usually wrong is that most of what people reach to encode is a one-line prompt, a repository fact, or a live connection wearing a skill’s clothes. Get the home right first. And remember that the convenience of a one-command install is also the thing that puts someone else’s executable code one step from your credentials.

Ty Sutherland

Ty Sutherland is the Chief Editor of AI Rising Trends. Living in what he believes to be the most transformative era in history, Ty is deeply captivated by the boundless potential of emerging technologies like the metaverse and artificial intelligence. He envisions a future where these innovations seamlessly enhance every facet of human existence. With a fervent desire to champion the adoption of AI for humanity's collective betterment, Ty emphasizes the urgency of integrating AI into our professional and personal spheres, cautioning against the risk of obsolescence for those who lag behind. "Airising Trends" stands as a testament to his mission, dedicated to spotlighting the latest in AI advancements and offering guidance on harnessing these tools to elevate one's life.

Recent Posts