A neutral benchmark for agent skills, answering the question everyone is shouting about caveman / ponytail / mempalace: "Does this skill actually make things better?"
It runs every task twice (no-skill baseline vs. with-skill) and prints a "nutrition label" along three axes:
- Cost — input / output / total tokens + dollar cost.
- Behavior — number of turns + tool-calls. This is what catches the ponytail case: "turn the skill on → more tool-calls & higher cost."
- Quality (
--judge) — LLM-as-judge scoring 0–10. Detects skills that are cheap but make the result worse.
🇻🇳 A Vietnamese version of this document is kept at
README.vi.md.
Skills are actually used inside Claude Code / Codex, not over a bare API — so that is the default measurement plane:
--harness |
What it is | When to use |
|---|---|---|
claude-code (default) |
Runs the task via claude -p ... --output-format json. This is how skills are really used; it reads total_cost_usd / num_turns / usage as reported by the CLI itself. No ANTHROPIC_API_KEY needed (uses Claude Code's own auth). |
Trustworthy conclusions |
api |
A cheap, provider-agnostic proxy (Claude API / Bedrock): a system prompt stands in for the skill + an internal tool-loop. | Quick checks, model comparisons |
Why the plane matters (observed in real runs): a one-shot Q&A task costs ~1,100 input tokens on
apibut ~20,900 inclaude-code(the CLI's system prompt + tool defs). A few-line "terse" skill cuts cost ~30% on theapiproxy but almost 0% in real Claude Code, because harness overhead swallows the savings. The proxy is easy to be lulled by —claude-codeis the truth.
# Verify the logic without calling anything (deterministic mock):
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --judge --dry-run
# Measure inside real Claude Code (default) — installs SKILL.md, lets the agent fire it itself:
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --judge
# ... or force the skill always-on via the system prompt:
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --inject system --judge
# AGENTIC tasks on a fixture repo — forces the agent to use real tools (measures the real tool-loop/turns):
python skillbench/bench.py --skill skillbench/skills/example-skill.md \
--tasks skillbench/tasks-agentic.json --repo skillbench/fixture --judge
# API proxy (Claude API):
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --harness api --tools --judge
# API proxy (AWS Bedrock) — Bedrock model-id + AWS creds in env:
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json \
--harness api --provider bedrock --model global.anthropic.claude-opus-4-6-v1 --tools --judge--harnessclaude-code(default) |api.--inject(claude-code only)skill(default — installs.claude/skills/<name>/SKILL.mdand lets the agent fire it) |system(injected via--append-system-prompt, always on).--repo(claude-code only) — fixture repo directory for agentic tasks. Each run executes in its own isolated tmp copy (fair baseline/treated, never touches the source repo, not contaminated by the current project's CLAUDE.md/skills).--repeat N— repeat each cell N times to average out noise. The report takes the median for tokens/cost/turns and a pass-rate for the oracle (e.g.4/5). The cost verdict is driven by a 95% bootstrap CI on the DIFFERENCE in total cost (skill − base), paired by task: CI contains 0 ⇒↔️ noise; entirely below 0 ⇒ ✅ genuinely cheaper; entirely above 0 ⇒ ❌/⚠️ genuinely more expensive. The CI shrinks as N grows (unlike the IQR band, which is purely descriptive). It also prints a 25–75% IQR band per task for reference on spread (robust to outliers).--model— claude-code: leave empty to use the CC-configured model. api: defaults toclaude-opus-4-8; Bedrock usesanthropic.-prefixed ids (global.anthropic.claude-opus-4-6-v1,us.anthropic.claude-haiku-4-5-20251001-v1:0).--provider(api only)anthropic(usesANTHROPIC_API_KEY) |bedrock(uses AWS creds).--tools(api only) — enables the tool-loop to count tool-calls.--judge— LLM-as-judge 0–10 (the claude-code harness judges withclaude -pitself, no API key needed).--dry-run— mock, calls nothing.
tasks.json— plain-text Q&A (no tools needed) →turnsis always 1, measures token cost only.tasks-agentic.json+--repo skillbench/fixture— forces the agent to do real work (runrun_tests.py, read/edit files) → lets us measure turns/tool-loop. This is where the ponytail case shows up. Includes hand-written tasks oncalc.py/utils.py(fix-bug, add-median, robust-average, add-fizzbuzz) plus 3 real tasks borrowed from Exercism (exercism-wordy,exercism-grade-school,exercism-phone-number) — hard, unsaturated, withpython -m unittestas the oracle.
A task may declare "check": "<command>" — after the agent finishes, that command runs in the agent-modified workdir through a shell; exit 0 = task pass. This is an objective quality measure for agentic tasks, more reliable than the LLM-judge (which only sees the answer text, not the repo/test result → for agentic work it tends to floor the score and become meaningless).
- The table gets an Oracle (task pass) row,
<pass>/<total>for baseline vs. treated. - The verdict prioritizes the oracle: if a skill reduces the number of passing tasks → ⛔ strong warning (cheap-but-wrong is useless); the judge (if enabled) is then just a reference line.
- PASS_TO_PASS (borrowed from SWE-bench): a
checkcan run a&&chain, forcing the entire old suite to keep passing + the new test to pass — a skill can't "win" by fixing this task while breaking something else. Example foradd-median:"check": "python run_tests.py && python check_median.py".
- Real measurement is noisy run-to-run (agents are non-deterministic). Use
--repeat N→ the verdict is based on a bootstrap CI of the cost difference; the default N=1 (one sample — reference only, no inference possible). A CI that contains 0 is not yet trustworthy. We use a bootstrap CI (+ a descriptive IQR band) instead of min–max because a single outlier run (e.g. a base of $1.01 across 8 runs) keeps min–max overlapping forever even when the median is very stable; a bootstrap CI shrinks as N grows. The current task set is ~10 (incl. 3 Exercism) — still small; the CI is wide when there are few tasks. - Headless runs use an
--allowed-toolswhitelist (Bash Edit Write Read Glob Grep);--dangerously-skip-permissionsis NOT used because it is blocked when running as root. Tools outside the whitelist are denied rather than prompted. --inject skillonly installs the skill; the agent may not fire it → this measures "available vs. not".--inject systemforces it always-on → a clearer signal, but not the real progressive-disclosure mechanism.- The
claude-codeharness reportsnum_turnsbut does not yet break out tool-call counts (would need to parse--output-format stream-json— out of scope for the MVP). - Cost on
apiis approximated from Claude API list prices per model family; Bedrock may differ (regionalus./eu.+10%). The % Δ A/B part is still accurate. Theclaude-codeharness uses the CLI's self-reported cost (including cache). --judgespends its own tokens (a measurement cost, not counted toward skill cost); judging with the same model may be biased; the sample task set (5) is illustrative only.- Codex is not yet supported (a similar runner can be added once
codex execis available).
4 real skills currently trending on GitHub, measured with the same setup (--inject system, 10 tasks, oracle, repeat 5, bootstrap CI). Details + limitations: benchmarks/trending-2026-06-28.md; sources/licenses: skills/trending/SOURCES.md.
| Skill | Δ cost | 95% CI (Δ cost) | Oracle | Verdict |
|---|---|---|---|---|
| caveman (token-cutting) | +14.9% | [+10.8%, +17.6%] | 25/35 → 25/35 | |
| ponytail (minimalist) | +10.7% | [+8.8%, +15.5%] | 25/35 → 25/35 | |
| karpathy-guidelines | +6.6% | [+5.1%, +12.4%] | 25/35 → 25/35 | ❌ genuinely more expensive |
| systematic-debugging | +40.5% | [+36.1%, +43.7%] | 25/35 → 25/35 | ❌ genuinely more expensive |
Result: all 4 trending skills INCREASE cost (every CI sits entirely above 0), and 0/4 improve the oracle (which is unsaturated, so there was room to improve). caveman, which claims to "cut 65-75% of tokens", actually costs +15% — its "saved tokens" are measured on output prose, while in a real harness input dominates. The real cost lever is how heavy the skill is, not the marketing label. (Fairness note: this task set leans fix-bug with a binary oracle — a skill may help on other, better-suited tasks.)
The same 4 skills re-run across 3 model tiers (real-measured, --model, repeat 5). Details: benchmarks/models-2026-06-28.md.
| Skill | Opus 4.8 | Sonnet 4.6 | Haiku 4.5 |
|---|---|---|---|
| caveman | +14.9% | +14.6% | +21.4% |
| ponytail | +10.7% | +10.5% | +33.1% |
| karpathy-guidelines | +6.6% | ||
| systematic-debugging | +40.5% | +25.2% | +26.9% |
(bold = 95% CI sits entirely above 0;
Cross-tier conclusion: 9/12 cells are genuinely more expensive, 0/12 improve quality — a model-independent finding. The skill's behavioral effect is model-dependent (ponytail +33% on Haiku vs +10% on Opus, because the smaller model rambles). The light skill (karpathy) is cost-neutral but still useless on pass-rate.
⚠️ skillbench does NOT claim to be the first benchmark for agent skills. As of early 2026 several academic preprints already measure whether skills help, inside real harnesses, with paired designs (see Related work below). skillbench is not competing for "first" on any of those axes.
What skillbench does offer, against that prior art:
- It targets the viral community prompt-skills — caveman, ponytail, superpowers/systematic-debugging, karpathy-guidelines — which the academic benchmarks (SkillsBench, SkillTester, AgentSkillOS, …) do not specifically evaluate. They test "official" SKILL.md packages or self-generated skills; they do not measure the GitHub-trending skills people are actually arguing about.
- A joint cost × quality verdict in a single harness, with statistical error bars — paired-by-task bootstrap CI on the cost difference + a deterministic, SWE-bench-style
PASS_TO_PASSoracle (not LLM-judge pairwise). - An independent, low-cost, fully reproducible probe — the academic benchmarks report their own self-defined numbers and have not been externally replicated; skillbench is small enough that anyone can re-run it.
- A concrete demonstration of the measurement-plane effect (api proxy vs real harness) on the same skills — corroborating, not discovering, a point SkillsBench also raises.
Honest scope: this is a narrow, contrarian probe, not a comprehensive leaderboard. Treat it as a complement to — and an independent sanity-check on — the works below.
By early 2026 the "do agent skills help?" question is an active research area. The papers most directly adjacent to skillbench (all arXiv preprints; figures are author-reported, not independently verified at time of writing):
| Work | What it does | Relation to skillbench |
|---|---|---|
| SkillsBench (arXiv:2602.12670, BenchFlow) | Self-described "first benchmark to systematically evaluate how Agent Skills improve agent performance": paired vanilla-vs-skill, ~84 tasks / 11 domains / 7,308 trajectories, run inside Claude Code / Gemini CLI / Codex CLI; +16.2pp avg, 16 tasks negative; explicitly notes harness mediation. | Closest prior art; owns the "first / general benchmark" claim. skillbench differs by targeting community prompt-skills + cost×quality + CI. |
| How Well Do Agentic Skills Work in the Wild (arXiv:2604.04323, UCSB/MIT) | 34,198 real SKILL.md skills, retrieval-at-scale inside native harnesses; benefit degrades toward baseline as conditions get realistic (some models drop below baseline). | Strong evidence skills don't reliably help; focuses on retrieval-at-scale, not cost. |
| SkillTester (arXiv:2603.28815, Peking Univ.) | Paired baseline-vs-skill utility; a skill counts as useful only if it enables a new pass or reduces token cost / elapsed time. | Closest on the token-cost dimension. |
| SkillReducer (arXiv:2603.29919, HKUST) | 55,315 public skills; directly measures token-cost-vs-quality of skill content; "less-is-more" (heavy compression, +2.8% quality). | Adjacent: optimizes skill content rather than benchmarking outcomes. |
| AgentSkillOS (arXiv:2603.02176, Shanghai AI Lab) | Skills benchmark of 30 tasks, LLM-judge pairwise (Bradley-Terry) vs vanilla Claude Code. | Quality-only; reports no token/cost — a contrast skillbench fills. |
- Sypherd et al. — "Token Cost" (arXiv:2505.14880): establishes evaluating prompting strategies by efficiency (performance per token); "increased token usage → drastically diminishing returns". The conceptual backbone, but measured on bare benchmarks, not a harness or skills.
- Vercel — "AGENTS.md outperforms skills in our agent evals" (Jan 2026): measured pass-rate no-docs 53% / Skills 53% (+0pp, the skill didn't fire 56% of the time) / Skills+instructions 79% / static AGENTS.md 100%. A blog, not a benchmark. https://vercel.com/blog/agents-md-outperforms-skills-in-our-agent-evals
- Scott Logic — "Ponytail, YAGNI, and the problem with prompt benchmarks" (Jun 2026): the "ponytail" benchmark is misleading because the baseline was artificially handicapped (the hook fired on the baseline too). The motivation for skillbench's objective oracle + fair baseline. https://blog.scottlogic.com/2026/06/16/ponytail-yagni-and-the-problem-with-prompt-benchmarks.html
- Anthropic — "Adding Error Bars to Evals" (Miller 2024): the lesson → a paired-difference test on the same task (skillbench runs the same task in both arms). https://www.anthropic.com/research/statistical-approach-to-model-evals
- SWE-bench (MIT): the FAIL_TO_PASS / PASS_TO_PASS concept + resolve-rate. https://github.com/SWE-bench/SWE-bench
- Aider polyglot benchmark (Exercism, Python track MIT): source of the 3 real fixture tasks. https://github.com/Aider-AI/polyglot-benchmark
- Chroma — "Context Rot" (2025): "more context ≠ better" — quality degrades as the context window fills. https://www.trychroma.com/research/context-rot