Table of Contents
A good skill description answers two questions immediately: what does this skill do, and when should an agent use it? If the wording is vague, the agent may miss relevant requests. If it is too broad, the skill may load for unrelated work.
The reliable way to improve a description is to treat activation as a classification problem. Define the intended boundary, test realistic prompts on both sides of that boundary, revise the wording, and verify the change on prompts you did not use during editing.
What the description controls
In the Agent Skills specification, a SKILL.md file has YAML frontmatter containing a required name and description. The description can contain 1–1,024 characters and should state both the skill’s function and when it applies.
---
name: spreadsheet-analysis
description: Analyze and transform spreadsheet data, including formulas, summaries, cleanup, and charts. Use when a user wants to inspect or modify a CSV, TSV, XLS, or XLSX file, or describes a tabular-data task without naming the file type.
---
Many agent clients use the name and description to decide whether a skill is relevant before loading its full instructions. Exact behavior varies by client, configuration, and model, so confirm activation using the logs or tool history available in your own environment.
Write the first description
Start with the user’s goal rather than the skill’s implementation. A description should normally include:
- Actions: what the skill helps accomplish, such as extracting tables, editing formulas, or creating charts.
- Objects or formats: the file types, systems, or domains it handles.
- Trigger situations: how a user might describe the need, including indirect wording.
- Important boundaries: closely related tasks that belong to another skill, when confusion is likely.
Use concrete verbs and recognizable terms. “Handles data” is too vague; “cleans CSV data, calculates summary statistics, and creates charts” gives the agent a useful decision boundary.
Keep it concise
The description is routing information, not the full operating manual. Put step-by-step procedures, commands, templates, and edge-case handling in the body of SKILL.md or its referenced resources. This preserves context and makes the description easier to classify.
Build a trigger evaluation set
Create a list of realistic user prompts and label each one with the expected activation result. A practical first set contains roughly equal numbers of positive and negative examples.
[
{
"query": "Add a profit-margin column to q4_results.xlsx and highlight rows below 10%.",
"should_trigger": true
},
{
"query": "Write a Python service that imports CSV rows into PostgreSQL.",
"should_trigger": false
}
]
The first prompt needs spreadsheet editing. The second happens to mention CSV, but its primary intent is application and database development. This near-boundary negative example is more informative than an unrelated prompt such as “What is the weather?”
Positive examples to include
- Direct requests that name the format or domain.
- Indirect requests that describe the task without using the obvious keyword.
- Short prompts and detailed prompts with filenames, column names, or constraints.
- Informal wording, abbreviations, and minor spelling mistakes.
- Multi-step requests in which the skill is only one part of the work.
Negative examples to include
- Adjacent tasks that share keywords but need a different capability.
- Requests that involve the same file type but a different goal.
- Tasks covered by another installed skill.
- Requests that should be handled directly without loading this skill.
Avoid filling the set with obviously unrelated prompts. They can make accuracy look high without testing the difficult routing decisions.
Measure activation
For each query, run the agent and record whether it loaded the skill. Count a result as correct when a positive query activates the skill or a negative query does not.
| Expected | Observed | Result | Likely issue |
|---|---|---|---|
| Trigger | Triggered | Correct | None |
| Trigger | Not triggered | False negative | Description may be too narrow or omit user language |
| Do not trigger | Triggered | False positive | Description may be too broad or overlap another skill |
| Do not trigger | Not triggered | Correct | None |
Agent behavior can vary between runs. If the client and budget allow, repeat each prompt several times and record the activation rate rather than relying on a single result. Use the same model, tools, installed skills, and configuration when comparing description versions.
Separate revision prompts from validation prompts
Split the evaluation set before editing:
- Revision set: use these results to identify missing scope and unwanted activations.
- Validation set: keep these prompts hidden from the editing process and use them only to compare candidate descriptions.
Both sets should contain positive, negative, direct, and near-boundary examples. If you repeatedly rewrite the description to fix every individual prompt, it may memorize the wording rather than represent the skill’s true scope.
Use a controlled optimization loop
- Save the current description as the baseline.
- Run the revision and validation sets under the same conditions.
- Group false negatives by missing intent, format, or terminology.
- Group false positives by overly broad language or overlap with adjacent skills.
- Make one general change to the description.
- Run the same evaluation again and compare the results.
- Keep the version with the best validation performance, not automatically the latest version.
If a positive prompt is missed, add the general task category or natural wording it represents. If a negative prompt triggers, clarify the skill’s boundary. Do not paste entire failing prompts or a long list of special cases into the description.
Diagnose problems before adding words
Activation is not determined by the description alone. A skill can appear to fail because it is installed in the wrong directory, hidden by client settings, invalid at startup, manually invoked only, or overshadowed by another skill with similar scope. Before rewriting, confirm that:
- the skill is registered and visible to the agent;
- the frontmatter parses correctly;
- the name and description meet the client’s requirements;
- automatic invocation is enabled where applicable;
- the observation method really detects a loaded skill.
Before-and-after example
Weak:
description: Handles spreadsheet files.
Improved:
description: Analyze and edit CSV, TSV, XLS, and XLSX data, including formulas, cleanup, summaries, and charts. Use for spreadsheet or tabular-data tasks even when the user does not name the file format.
The improved version names the supported work, relevant formats, and an indirect trigger. It remains short enough for routing and leaves the actual workflow in the skill instructions.
Final checks
- The description states what the skill does and when to use it.
- Its scope does not unnecessarily overlap neighboring skills.
- It uses language that appears in real user requests.
- Detailed procedures stay outside the description.
- It remains within the 1,024-character specification limit.
- It performs well on fresh positive and negative prompts.
The official description optimization guide provides additional evaluation guidance. Re-run a small validation set whenever you change the model, client, installed skill set, or routing configuration, because those changes can alter activation behavior even when the description stays the same.
Reader Comments 0
Sign in with email or Google to join the discussion.