Skills without evals are just vibes

Tribal knowledge encoded as an AI skill is still just text until you evaluate it. Ablation baselines, routing regression tests, trajectory autoraters, and the gotchas flywheel keep encoded knowledge from rotting.

· Updated September 12, 2026

Every organization has people who know how things actually get done. Not the official process. The real one.

They know which edge case the template doesn’t handle. Which stakeholder has to be looped in first. Which step everyone quietly skips, and why. This is tribal knowledge, and it lives in people’s heads until they leave.

So you write it down. A SKILL.md with the right sequence, the right tools, the judgment calls spelled out. Now the expert scales past one person and one timezone.

Then most teams ship it and call it done. That is the part I want to argue with. An unevaluated skill is not encoded knowledge. It is a text file that resembles encoded knowledge, and the resemblance holds right up until it doesn’t.

Write the eval first

The conventional workflow is to write the skill, then write evals to check it worked. This is backwards.

Writing the evals first forces you to say what success looks like before you have written a single instruction. Specificity that early is the difference between a skill that does what you intended and one that does what you wrote.

Perplexity’s engineering team, which maintains a large skill library behind its Computer product, calls evaluation “Step 0” and puts writing the body at Step 2. [3] Their cases come from three places: production queries as users actually typed them, the known failures that made you want the skill in the first place, and the boundary cases where a neighbouring skill might load instead.

That last category is the one everyone skips and the one that pays. If you have both a “vendor response” skill and a “procurement inquiry” skill, you need cases that confirm each loads for its own queries and stays quiet for the other’s. Writing those cases before the skill exists forces the distinction into the description, where the router can see it.

Negative examples matter as much as positive ones. Perplexity’s guide says they “can matter more.” [3] An eval that only checks “the skill works when it loads” leaves half the failure space untested, and the untested half is the one where a skill fires on the wrong task and does something confidently wrong with a full context window behind it.

The same guide says that a skill written in five minutes will “almost certainly be subpar.” The eval cases are where the time goes. Spending it up front is cheaper than discovering the failure mode in production.

Run the ablation: does the skill earn its place?

Before writing a single instruction, run your eval cases against the base model with no skill loaded.

This is the ablation baseline. If the base model handles the task at 80 percent and your skill gets it to 83, you have not built a capability. You have built three points of lift that you will pay index tokens for on every session, maintain indefinitely, and debug whenever the routing shifts. That skill should not exist.

Ablation answers two questions. The first is whether the skill is needed at all, and the honest answer is often no. A lot of what teams reach for skills to encode is already in the weights: sequential git operations, standard document formatting, the common API patterns. Perplexity’s guide is blunt about it: “If it’s easy to explain, the model already knows it. Delete it.” [3] Wrapping knowledge the model already has is not free. It is a cost you stopped noticing.

The second question is which part of the skill is doing the work. Run the eval with the full skill, then strip it down section by section. If removing the examples drops performance by 15 points and removing the step-by-step instructions drops it by 2, you know where to invest. A lot of skill content is padding around a small number of high-signal lines. Ablation makes that visible.

from statistics import mean

def run_ablation(eval_cases, skill_name):
    baseline, with_skill = [], []
    for case in eval_cases:
        baseline.append(grade(agent.run(case.query, skills=[]),           case.expected))
        with_skill.append(grade(agent.run(case.query, skills=[skill_name]), case.expected))

    lift = mean(with_skill) - mean(baseline)
    print(f"Baseline:   {mean(baseline):.0%}")
    print(f"With skill: {mean(with_skill):.0%}")
    print(f"Lift:       {lift:+.0%}")
    return lift  # if small, question whether the skill should exist

SkillsBench, a benchmark of 87 tasks with curated skills and deterministic verifiers, is the largest public version of this experiment. Curated skills lifted the average pass rate from about 34 percent to about 50 percent, but the gain varied wildly by domain, focused skills with at most three modules beat exhaustive bundles, and skills the model wrote for itself provided no benefit on average. [7] Every one of those findings is an argument for measuring before shipping.

The honest conclusion of a well-run ablation is sometimes that the skill should not be built. The tribal knowledge you thought needed encoding was already in the weights, or the gap was too small to justify the infrastructure. That is a good outcome. A skill that never gets written is a skill that never rots.

How harnesses load skills, and what breaks with too many

To write good routing evals you need to understand what you are testing against. The loading mechanics are similar across the major harnesses, and the failure mode is the same in each.

All of them use progressive disclosure. A compact index of every installed skill, name and description only, sits permanently in context. The full SKILL.md loads only when a skill is invoked. Bundled files load only when read. [4] This keeps the per-session baseline cost manageable. It does not make it zero.

Claude Code caps each skill’s combined description and when_to_use text at 1,536 characters in the listing. The listing as a whole has a character budget that scales at 1 percent of the model’s context window. The listing always keeps every skill’s name, but when it overflows, Claude Code drops descriptions starting with the skills you invoke least, and the only notice is a line in the debug log. After a compaction, invoked skills are re-attached at up to 5,000 tokens each within a shared 25,000-token budget, most recent first, so older skills can be dropped silently. [5]

Codex scans .agents/skills from the working directory up to the repository root, then $HOME/.agents/skills and /etc/codex/skills. Its initial roster gets at most 2 percent of the context window, or 8,000 characters when the window is unknown. When that overflows, descriptions are shortened first. Past that, skills are omitted from the list with a warning. [6]

Hermes Agent from Nous Research does the same thing with three explicit levels: a skills_list() call that costs about 3,000 tokens for the whole index, a skill_view(name) for the body, and skill_view(name, path) for a reference file. It pulls skills from eight hub sources, including ClawHub and skills.sh. [8] The tradeoff is identical. Every skill in the index costs space whether it is ever used or not.

The shared failure mode is silent truncation. You add a new skill, descriptions get shortened across the board, a routing phrase disappears mid-sentence, and a skill that routed fine yesterday starts failing today. You did not change that skill. You changed the space around it.

Perplexity calls this action at a distance. [3] It is the reason routing regression tests have to run after every addition to the library, not only when you edit a skill directly.

The practical ceiling depends on description length. On a 200,000-token model, Claude Code’s 1 percent budget is roughly 2,000 tokens, about 8,000 characters. At the 50-word descriptions Perplexity recommends, that is room for something like 20 to 25 skills before truncation starts. The ablation discipline matters here too: every skill that does not earn its place is spending index budget a needed skill could use.

Skill loading: three tiers across all major harnesses

Test loading, not just execution

Most teams evaluate the wrong thing.

They check whether the skill completes the task once loaded. That is necessary and not sufficient. Skills have a failure mode ordinary software doesn’t: the agent has to decide whether to load the skill at all.

This is a routing problem with the usual two failure modes. A false positive is the skill loading when it shouldn’t. The agent is doing something unrelated, the description matched closely enough, and now the context is full of instructions for the wrong task. A false negative is the skill not loading when it should, which is worse in a quiet way. The capability exists, the agent never reaches for it, and the expert’s knowledge sits in a file nobody opens.

Both failures are invisible without explicit routing evals: cases that assert whether the skill loaded, not only how it performed afterwards. Perplexity runs these as a separate suite that checks precision, recall, and forbidden loads for every skill. [3]

This is where action at a distance bites. Two descriptions that read as distinct on their own can be ambiguous in combination, and a skill that routed correctly in isolation starts losing traffic to the new arrival without anyone editing a line of it.

So the unit of evaluation is the skill library, not the skill. Test them together or you are testing a configuration that does not exist in production.

import pytest

@pytest.mark.parametrize("query, skill, should_load", [
    # positive cases: the right skill loads
    ("draft a response to this vendor email",  "vendor-response",    True),
    ("reply to the supplier about the delay",  "vendor-response",    True),
    ("we received a new procurement request",  "procurement-inquiry", True),
    # boundary cases: the wrong skill must NOT load
    ("draft a response to this vendor email",  "procurement-inquiry", False),
    ("we received a new procurement request",  "vendor-response",    False),
])
def test_routing(query, skill, should_load):
    session = run_session(query)
    loaded = get_loaded_skills(session)
    if should_load:
        assert skill in loaded, f"{skill!r} should load for: {query!r}"
    else:
        assert skill not in loaded, f"{skill!r} must not load for: {query!r}"

Evaluate the path, not just the answer

Assume the skill loaded correctly. The next failure mode is a correct-looking output from a broken process.

An agent that produces a correctly formatted expense report may have skipped receipt validation. An agent that returns a polished vendor response may have queried the wrong contract database. The output hides the process.

This matters because tribal knowledge is usually procedural. It is about how to do something, not only what the result looks like. Encoding the how and evaluating only the what misses the point.

Trajectory evaluation grades the sequence of tool calls rather than the final output. [2] For a skill that is supposed to:

  1. Look up the client’s account tier
  2. Check the relevant SLA policy
  3. Draft a response using the approved template

you can verify each step fired, in order, with the right parameters. A model-based grader, an autorater, reads the transcript and checks whether the tool sequence matches the expected pattern. It catches failures that output grading never would.

RUBRIC = """\
Evaluate whether the agent followed the required procedure.

Required steps, in order:
  1. get_account_tier   (must be called first)
  2. check_sla_policy   (must reference the result of step 1)
  3. draft_response     (must be the final action)

Tool calls observed (chronological):
{tool_calls}

Answer yes or no:
- Was get_account_tier called before any other tool?
- Was check_sla_policy called after get_account_tier?
- Was draft_response the last call?
- Did the agent follow the correct procedure overall?
"""

def evaluate_trajectory(transcript):
    calls = extract_tool_calls(transcript)
    formatted = "\n".join(f"  {i+1}. {c['name']}({c['args']})" for i, c in enumerate(calls))
    response = judge.complete(RUBRIC.format(tool_calls=formatted))
    return parse_yes_no(response, key="overall")

An autorater for tool calls does not need a framework. A rubric like “did the agent call get_account_tier before drafting the response?” grades cleanly with a capable model. Pair it with exact-match checks on the critical parameters and you have a grader that runs at scale without a human on every case.

The caveat: grade paths when the path matters, not by reflex. Anthropic’s eval guide warns that checking for a specific sequence of tool calls is “too rigid” and produces brittle tests, because agents regularly find valid routes the eval designer did not anticipate. [1] A skill with one valid execution path needs trajectory evaluation. A skill that can reasonably reach the same outcome several ways should be graded on the outcome.

One metric worth understanding early: pass^k, the probability that all k trials succeed, not just one. A skill that works 75 percent of the time looks fine in a single run. Its pass^3 is about 42 percent. [1] For skills encoding knowledge that has to hold every time, that is the number you want.

Outcome verification is the north star

Trajectory evaluation tells you the path was right. Outcome verification tells you the world changed correctly.

A flight-booking agent can produce a perfect transcript, correct tool calls, right parameters, a professional confirmation, and still fail to create a reservation. The outcome is the row in the database, not the sentence in the chat. [1]

For skills encoding tribal knowledge, the outcome check is often the most objective signal you have. It needs no rubric and no calibrated judge. It needs the environment state before and after the skill ran, and a check that the right thing changed.

It is harder to set up than output grading. You need a real or realistic environment, ground truth state, and automated verification. That investment is what separates teams that know their skills work from teams that assume they do.

Start with the highest-stakes skills, the ones where a failure is expensive, visible, or both. Get outcome verification working for those before expanding.

Mock the boundary, not the decision

Outcome checks need real environments, which makes offline testing slow and expensive. The natural fix is mocking: simulated tool responses so evals can run without live APIs.

The risk is that aggressive mocking hollows out the eval. Bypass the model call entirely and you learn nothing about the skill. Mock the tool results but leave the agent’s tool selection real and you get something useful: fast, repeatable evals that still test whether the agent calls the right tool with the right arguments.

Mock at the environment boundary, not inside the agent. Intercept the tool response, never the tool call.

# Wrong: bypass the LLM entirely. Tests nothing about the skill's actual behavior.
def test_wrong(monkeypatch):
    monkeypatch.setattr("skill.run", lambda _: "Your issue has been escalated.")

# Right: mock tool results, keep the LLM call real
def test_correct(monkeypatch):
    monkeypatch.setattr("tools.get_account_tier", lambda _: "Starter")
    monkeypatch.setattr("tools.check_sla_policy",  lambda t: SLA[t])
    # The model still runs with the real skill instructions and real LLM.
    # Only what comes back from the environment is controlled.
    result = agent.run("draft a response to this vendor email about their delay")
    assert "5 business days" in result  # Starter tier SLA, not the 48-hour Enterprise figure

Brittle evals come from under-constrained mocks. If your mock for get_account_tier always returns “Enterprise,” the skill never gets tested on the “Starter” case that causes half the real failures. Stratified mocks, covering the case distribution that actually shows up in production, are what make offline evals worth running.

Start with real failures, not hypothetical ones. Anthropic’s guide suggests 20 to 50 simple tasks drawn from the bug tracker and the support queue. [1] Those surface more signal than two hundred synthetic cases. And carry the boundary cases over from your routing evals. Mock the adjacent scenarios, not only the clean ones.

Skills rot. The gotchas flywheel is the fix.

A skill encoding tribal knowledge at a point in time is a snapshot. The organization keeps changing. Policies update. Tool APIs evolve. Models improve in ways that shift behavior without announcement.

Most maintenance advice says to rewrite: update the instructions, restructure the flow, improve the examples. Perplexity’s production experience points the other way. Skills, they say, are “append-mostly,” and the section that accrues the most value is the gotchas. [3]

The mechanism is a running record of real failures, translated into warnings. Every time the skill fails in a way you did not anticipate, the failure becomes a gotcha. The gotcha becomes an eval case. The eval case prevents the regression.

This flywheel is how tribal knowledge accumulates in a skill over time. The initial SKILL.md captures what the expert knew when the skill was written. The gotchas capture what the skill revealed it did not know once it started running. The second layer is often the more valuable one, because it holds the lessons nobody thought to write down the first time.

Perplexity calls gotchas “extremely high-signal content.” [3] A line that says “do not skip receipt validation even if the submitter says it’s already approved” does more work than three paragraphs on how to process expense reports correctly.

## Gotchas

- Do not skip receipt validation even if the submitter says "already approved."
  Approvals older than 30 days must be re-confirmed in the system before filing.
- The 48-hour SLA applies only to Enterprise accounts with the priority add-on.
  Starter tier is 5 business days. Always call get_account_tier before citing any figure.
- If get_account_tier returns null, the account is in trial mode.
  Use the trial SLA policy, which differs from the free-tier policy.
- Do not draft the response before both get_account_tier and check_sla_policy have
  returned. A tier value from an earlier turn may be stale; always call fresh.

The gotchas flywheel: how failures become permanent improvements

Ownership is what makes the flywheel turn. Someone has to observe the failure, extract the pattern, write the gotcha, and add the eval case. Without that person the flywheel stops. Failures happen and do not get encoded. The skill stays at whatever quality it shipped at while the world keeps moving.

Ownership also covers the routing layer. When a new skill lands in the registry, the owner of every adjacent skill needs to check that their routing has not shifted. This does not happen without explicit responsibility.

The re-evaluation cadence should be explicit. High-stakes skills: after every significant model update and every change to the tools they call. Stable, low-stakes skills: on a longer cycle, but a defined one. “Whenever someone gets around to it” is not a cadence.

Two loops, not one

Offline evaluation and online monitoring complement each other. Neither replaces the other.

Offline evals run in a controlled environment against a curated dataset. They are fast and repeatable, and they tell you whether the new version beats the old one before you deploy. The catch is that the dataset needs maintaining. Stale test cases mislead more effectively than no test cases, because they come with a number attached.

Online monitoring runs against real traffic. It surfaces the failures that never appeared in your dataset, because users do things eval designers do not think of. Production logs, feedback signals, and outcome audits all feed it.

Teams that get this right run both. When online signals diverge from offline metrics, the move is to update the eval dataset, not to wave production away as noise.

One thing to watch in the offline loop is saturation. When performance on the eval set stops moving and sits near perfect, the set has usually stopped covering the hard cases, and it has stopped being useful as an iteration signal. [1] The gotchas flywheel is what keeps the set from saturating. Each production failure becomes a new case.

Test across models explicitly. Perplexity runs its skills against GPT, Claude Opus, and Claude Sonnet as orchestrators, and reports that “Sonnet and GPT behave quite differently when it comes to Skills.” [3] If your agents run on more than one backend, the eval suite has to cover all of them as part of the standard run.

The skill that works vs. the skill you think works

Encoding tribal knowledge is worth doing. The knowledge outlives the person who had it, it scales past what one expert can staff, and it runs at 2am when that person is asleep.

The catch is that the encoding only holds if something keeps checking it. A skill nobody has evaluated in six months is not a capability. It is a belief about a capability, and the two are indistinguishable from the outside until the moment they aren’t.

Which is why I keep landing on the same unglamorous answer. Evaluation is not a gate you pass on the way to shipping. It is the thing that keeps the file connected to reality after you ship. Write the eval before the skill, because it forces you to say what success means while you can still change your mind cheaply. Test whether the skill loads before you test how it performs, because a skill that never fires is a skill you do not have. Verify what changed in the environment, not what the transcript said. And let production failures become gotchas, because that is the only mechanism that puts new knowledge into the file after the expert has moved on.

None of that is hard. It is just nobody’s job by default, which is the real failure mode. Tribal knowledge rotted in people’s heads for the same reason. Not because it was wrong, but because nobody was responsible for noticing when it stopped being right.


References

[1] Anthropic Engineering, "Demystifying evals for AI agents," January 9, 2026. anthropic.com

[2] LangChain, "LLM Evaluation Framework: Trajectories vs. Outputs," April 8, 2026. langchain.com

[3] Perplexity Engineering, "Designing, refining, and maintaining agent skills at Perplexity," May 1, 2026. perplexity.ai

[4] Anthropic Engineering, "Equipping agents for the real world with Agent Skills," October 16, 2025. anthropic.com

[5] Anthropic, "Extend Claude with skills," Claude Code documentation, sections on the skill listing budget and compaction. code.claude.com

[6] OpenAI, "Build skills," Codex documentation, skill locations and the initial skills list budget. learn.chatgpt.com

[7] X. Li et al., "SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks," arXiv:2602.12670, February 2026. arxiv.org

[8] Nous Research, "Skills," Hermes Agent user guide. hermes-agent.nousresearch.com

A sunlit Brooklyn brownstone street in autumn: parked cars, iron railings, stoops, and a canopy of turning leaves.