The Head -100 Problem: Why Your Skills Suddenly Fail
โถ 0:00When Claude opens a long reference file inside one of your skills, it does not always read the whole thing. It runs a head -100-style command โ reading roughly the first 100 lines โ to decide whether the file will add information it actually needs. The consequence is blunt: if your important rules sit after line 100, as far as Claude is concerned they do not exist.
Skills launched early this year, and the original playbook was simple โ good descriptions, total lines under 200, reference files separated out, written like a set of instructions. Those rules have now changed. Anthropic updated its skill-building best-practices guide with a new set of rules that apply to every skill you have already built. Get them right and your skills give far more consistent results regardless of which model is running them.
Rule 1: Put a Content List at the Top of Long Files
โถ 0:53The direct counter to the 100-line problem: add a contents list โ an index โ at the top of any reference file over 100 lines, so Claude can either read the whole file deliberately or jump straight to the section it needs. Anthropic's own example is an API reference: contents at the top, then the sections beneath it โ authentication and setup, core methods, advanced features, error handling patterns, and code examples.
The practical task is mechanical: open Claude Code or Chat, walk through every skill in your claude/skills folder, find any reference file over 100 lines, and add a contents list at the top that matches its actual headings. If a reference file is over 100 lines and has no contents list, the chances of it being fully used are genuinely low.
Rule 2: Set the Right Degrees of Freedom
โถ 1:38The rule that changes how you think about skills: the detail in a skill should depend on how fragile the task is and how much it varies โ not on a blanket assumption that more precision is better. Anthropic calls this "appropriate degrees of freedom," and there are three levels.
High freedom is plain instructions. A code-review skill says "check the structure, look for bugs, suggest improvements for readability and maintainability, follow project conventions" โ and leaves the how to the model, because there are many good ways to do a code review and the right one depends on the actual code. Medium freedom is a template with settings: a preferred shape with some variation allowed (include charts true/false, format markdown/HTML) โ a weekly client report is the classic case. Low freedom is an exact command: a database migration where the instruction is "run exactly this script, do not change the command, do not add flags." Anything touching money or deletion โ raising invoices, issuing refunds, removing members โ deserves this strict level.
Three takeaways: one skill can mix levels (an invoicing skill's writing step is high-freedom while its create-invoice step is locked down); the question for each step is "what happens if Claude does this differently?" โ if nothing much changes, loosen it, and if something consequential changes, tighten it; and low freedom usually means a script, not more plain text, because a script runs the same way every time.
Rule 3: Test on Every Model You'll Use
โถ 4:25Obvious and yet nobody does it: a skill's result depends on the model underneath it, so test it on every model you plan to run it with. The right amount of detail has drifted a lot. Six to eight months ago, older models skipped steps, so skills carried numbered lists and CAPITAL-LETTER emphasis. Now it is the opposite โ the newer reasoning models find over-prescriptive instructions actively harmful, and Anthropic tells you to remove older instructions when the model does better without them.
The best-practices page gives one question per model tier: for Haiku โ does the skill provide enough guidance? For Sonnet โ is the skill clear and efficient? For Opus โ does the skill avoid over-explaining? Run your most-used skill on the same task with Haiku, then Sonnet, then Opus. If Haiku misses a step, that step needs to be clearer or turned into a script (low freedom, runs identically everywhere). If Opus does worse with the skill than without it, start cutting instructions until it improves. And if you plan to share the skill, document which models it is intended for in the YAML frontmatter.
Rule 4: Keep References One Level Deep
โถ 6:30Keep the SKILL.md body under 500 lines and treat it as the contents page for everything else, splitting content into separate files as you approach the limit. A well-structured skill directory looks like a PDF skill with SKILL.md plus a forms.md guide, a reference.md API reference, and a scripts/ file โ Claude only opens the reference and forms files when a step actually needs them. Crucially, scripts get run, not read into memory, so they do not consume any context.
When a skill covers several areas, split the references by area โ a data skill gets separate files for finance, sales, product, and marketing, so a question about revenue never loads marketing.md. A client-reporting skill gets one file per client.
Then the trap the opening 100-line rule sets up: nested references. If SKILL.md points to advanced.md which points to details.md, Claude is likely to only preview the file at the end of the chain โ the details at the bottom get read partially, if at all. The fix is to make every reference file link directly from SKILL.md, exactly one level deep. Audit your skill by listing every file it points to and finding any file only reachable through another file โ then link it directly.
Rule 5: Give Ordered Work a Checklist
โถ 8:43For longer, complex processes, break the work into steps and give Claude a checklist it copies into its reply and ticks off as it goes. This sounds like the opposite of the newer prompting advice to "give the whole job, not prescriptive steps" โ but it is not. Steps are for when the order matters, like checking the data before building the report on it; if order does not matter, leave the checklist out.
Anthropic's example is a research-synthesis workflow โ a five-line, deliberately non-prescriptive checklist: read all source documents, identify key themes, and so on, each step followed by a short paragraph. The important line is the last one: "verify citations โ check that every claim references the correct source document, and if citations are incomplete, return to step three." That "go back to step three" instruction is what stops Claude from ticking a failed step and moving on.
Rule 6: Build in a Self-Correction Loop
โถ 9:58Anthropic's pattern for validating skill output is a loop: run a check, fix the errors, repeat until it passes. The check does not have to be code โ the first example is a style guide. Draft, review the draft against the guide, note each issue with the section it breaks, revise, review again โ and only finalize when everything passes. For you, that style guide might be your brand-voice document.
This is also how the skill starts improving itself: when a draft fails for a reason that is not yet in the voice document, have Claude suggest the new rule at the end of the run. You approve it, it goes into the document, and the next run is checked against it. This compounds โ the guide notes that newer models are genuinely good at updating skills from what they learn on a task.
Rule 7: Spell Out Dependencies
โถ 10:50The final rule is about portability. A skill works on your laptop because you installed its dependencies six months ago and forgot which ones โ but on a teammate's machine or in a fresh session, those libraries will not be there. The rule is never assume a tool is installed.
Instead of silently using a PDF library, every skill should say something like "install the required package" and name it โ because if it is already installed, Claude is smart enough to skip the step. Put the install line next to every script in the skill. If you share or sell your skills, this is the difference between a skill that works on day one and one you have to hand-hold into existence.
Key Takeaways
- Claude only reads the first ~100 lines of a reference file unless there is a contents list โ rules after line 100 are invisible.
- Match detail to fragility, not to a reflex for precision. High freedom = plain instructions, medium = a template with settings, low = an exact script (money and deletions).
- Low freedom means a script, not more prose โ scripts run identically on every model and cost no context.
- Test every skill on Haiku, Sonnet, and Opus โ the right level of detail has flipped, and over-prescriptive skills now hurt reasoning models.
- Keep every reference file one level deep from SKILL.md. Nested references (aโbโc) mean the deepest file gets only partially read.
- Checklists are for when order matters โ and a "return to step N" line is what stops a failed step being ticked off.
- A self-correction loop (check โ fix โ repeat) is the skill's engine โ and have it propose new rules when a failure is not yet covered.
- Never assume a dependency is installed โ name the package and let Claude skip it if it is already there.