In short

A Claude Code skill is a Markdown file that the AI reads before acting on a given task. It replaces long documentation with a self-contained instruction that carries the rules, the process, and the guardrails all in one. Analysis of 116 skills from the Everything Claude Code community reveals that 93% adopt a monolith architecture: a single SKILL.md file, with no external dependencies. The observed median size is 240 lines — a useful benchmark for calibrating your own skills without under-specifying or over-documenting.


What a skill is, concretely

In plain terms : a skill is a short manual that the AI reads before acting — not a plugin, just structured text that says “when, how, what to avoid”.

A skill is not a plugin, nor a tool in the technical sense. It is a text file — typically SKILL.md, typically 5 to 15 kilobytes [estimated order of magnitude] — that Claude Code loads into context before executing a category of tasks. It encodes the expected behavior: when to activate (trigger), how to proceed (process), what errors to avoid (anti-patterns), how to verify the result (quality gate).

This mechanism relies on a fundamental property of LLMs: if instructions are present in context, they steer generation. A skill is not a system configuration — it is a structured textual instruction. The AI reading it is sufficient to activate its rules.

The difference from a simple prompt: a skill is reusable, versionable, and shareable. It integrates into a git repository, evolves with the project, and can be injected by an orchestrator into a sub-agent’s context.

In practice, a skill loads in two ways: either via a reference in CLAUDE.md (systematic loading), or injected dynamically into an agent’s prompt at orchestration time. This second modality is particularly useful in multi-agent mode: the orchestrator decides which skills are relevant for each sub-agent, without loading all of them into the global context.


Canonical anatomy: the four components

In plain terms : an effective skill answers four questions — when to use it, how to proceed, what to deliver, how to verify it is good.

Analysis of the 116 skills surfaces a canonical four-block structure, present in virtually all well-rated community skills.

1. Trigger — The activation condition. Describes when the skill should be used. Can be a direct phrase (“when the prompt contains X”) or a list of situations. Without an explicit trigger, the skill risks being under-used or applied out of context.

2. Rules / Process — The operational core. Describes the steps in order, with actionable instructions. Effective skills number their steps, use imperative mode, and avoid narrative prose. One concept per instruction. ASCII diagrams are useful for workflows with multiple branches.

3. Output format — Explicitly specifies the structure of the expected deliverable: file type, organization, length, naming conventions. Without this block, the AI improvises the format on every execution.

4. Quality gate / Checklist — The verification criteria before delivery. A skill without a quality gate cannot know when it is done. The checklist (- [ ] criterion) is the most widespread pattern because it is directly verifiable.

A fifth component appears frequently in more developed skills: the decision matrix or explicit anti-patterns. This block lists known errors to proactively inhibit them. In practice, explicit anti-patterns significantly reduce recurring drift on long tasks.


Monolith vs. multi-file: when to add dependencies

In plain terms : a single file is enough in 93% of cases — adding companion folders is only justified when the skill triggers an external process (script, hook, agent).

Of the 116 skills analyzed, 108 (93%) are pure monoliths: a single SKILL.md file, with no dependencies whatsoever. The remaining 8 skills (7%) use sub-directories — agents/, hooks/, scripts/ — for specific system needs (pre-commit hooks, autonomous agents, shell scripts called at execution time).

ContextRecommendationWhy
Skill < 100 lines, atomic taskMonolith, 1 recipeSimplicity of model reading, zero portability friction.
Skill 100–300 lines, cognitive behaviorMonolith with 4 canonical blocksOptimal size (55% of corpus), instructions followed in full.
Skill triggers a bash script or system hookMulti-file with scripts/Executed code has no place in instruction prose.
Skill orchestrates several distinct agentsMulti-file with agents/Isolation of sub-agent prompts, but portability cost.
Skill > 500 lines single-themeRefactor into several short skillsDilution of critical instructions, the AI skims beyond 500 lines.

The practical rule that emerges: the monolith is sufficient as long as the skill encodes a behavior. Additional files only appear when the skill triggers external processes — a bash script, a separate agent, a system hook. For everything that is “how the AI should reason and produce,” a single file is sufficient and preferable.

The multi-file architecture has a cost: it creates dependencies, complicates portability, and requires documentation of the directory tree. The continuous-learning-v2 skill (365 lines + 9 files, with agents/, hooks/, scripts/ sub-directories) and the videodb skill (11 files, with reference/ and scripts/) are functionally powerful but difficult to transfer or adapt without knowing their complete ecosystem.

The limit is clear: as soon as you remove a companion file from a multi-file skill, behavior changes in unpredictable ways. A monolith degrades gracefully — it produces a less precise but coherent result. A multi-file skill with a missing dependency can produce inconsistent behavior without an explicit error signal.


Optimal size: 100–300 lines

In plain terms : aim for 100–300 lines — less is vague; more becomes a document the AI skims instead of applying.

The distribution observed across 116 skills:

SizeProportionAssessment
< 100 lines~10%Often too vague
100–300 lines~55%Optimal zone
300–500 lines~25%Acceptable if structured
> 500 lines~10%Encyclopedic risk

A skill under 100 lines can work for simple tasks, but generally lacks an output format and quality gate. Beyond 500 lines, the skill becomes documentation that the AI skims rather than an instruction it follows.

The 100–300 line range corresponds to skills that cover a complete behavior without drowning critical instructions in accessory content. The community’s article-writing skill (85 lines) and typical domain skills (writing, collection, extraction) fall in this range and produce consistent results on repeated tasks.

The observed extremes illustrate the two opposite drifts. The shortest in the corpus: nanoclaw-repl at 33 lines — a functional skill for a very narrow task (launching a REPL), with no output format or quality gate. The longest: kotlin-testing at 824 lines — a complete reference document on Kotlin testing, covering all possible patterns, but whose length dilutes the critical instructions. These two cases are exceptions justified by their context; they are not generalizable models.


Best practices: what makes an effective skill

In plain terms : an effective skill lists what to do and what to avoid, with concrete examples rather than prose.

Minimal but complete frontmatter. All 116 skills systematically use three fields: name, description, origin. The description must be precise enough for an orchestrator to decide whether to inject this skill without reading its full content. A vague frontmatter (“skill for writing”) is unusable in automatic dispatch.

Concrete examples, not prose. Skills with structured examples (tables, named lists, snippets) produce more homogeneous results than those that describe expected behavior in prose. The AI extracts constraints better from a list than from a paragraph.

Explicit anti-patterns. Listing what the skill must not do is as important as listing what it must do. A well-formulated anti-pattern (“never summarize raw data”) is a guardrail the AI activates on every execution. The table format (forbidden | examples) makes prohibitions more salient than bullet lists.

Decision matrix for behavioral dispatch. When a skill covers several situations (simple vs. complex task, data present vs. absent), a decision matrix prevents the AI from choosing arbitrarily. The community’s search-first skill is a documented example: two columns (criterion / behavior), no prose.

One skill = one responsibility. Skills that cover several distinct domains drift toward the encyclopedic. The right granularity: one task category, one deliverable type, one activation context.

The skill replaces documentation, not reflection. A skill encodes a known and proven procedure. It is not designed to handle unanticipated edge cases — faced with an out-of-scope situation, the AI should signal the ambiguity rather than improvise.


The skill as orchestration interface

In plain terms : in multi-agent mode, the skill becomes a job description — its frontmatter is used by the orchestrator to choose who does what, without reading the content.

An advanced use of skills, observed in multi-agent configurations, is to treat them as a dispatch interface rather than simple behavior guides. The orchestrator reads the frontmatter description to decide which skill to inject into each sub-agent’s prompt. This pattern requires the frontmatter to be precise and discriminating: two skills with similar descriptions create dispatch ambiguity.

In practice, this leads to a naming rule: {domain}-{action} (e.g., redaction-article, collecte-web, extraction-faits). The skill name should allow selection without reading the content. This is a design constraint — not a stylistic convention.


What to remember

  • A Claude Code skill is a self-contained Markdown file that encodes a repeatable behavior: trigger, process, format, quality gate.
  • 93% of community skills are monoliths. Additional files only appear for external system needs (hooks, scripts, agents).
  • The optimal size is between 100 and 300 lines: complete enough to be precise, short enough to be followed in full.
  • Explicit anti-patterns and verification checklists are the most under-used and most effective components.
  • A skill encodes a proven procedure, not a generic one: narrow scope is a strength, not a limitation.

Skill signals

No major deficiency identified. A minor ambiguity: the skill prescribes “minimum 3 sources, target 7–12” but the T-sources provided for this topic are primarily internal observations and a community corpus without individual URLs. The [UNVERIFIED] format has been applied to the relevant point.