Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

skillbench

A neutral benchmark for agent skills, answering the question everyone is shouting about caveman / ponytail / mempalace: "Does this skill actually make things better?"

It runs every task twice (no-skill baseline vs. with-skill) and prints a "nutrition label" along three axes:

  • Cost — input / output / total tokens + dollar cost.
  • Behavior — number of turns + tool-calls. This is what catches the ponytail case: "turn the skill on → more tool-calls & higher cost."
  • Quality (--judge) — LLM-as-judge scoring 0–10. Detects skills that are cheap but make the result worse.

🇻🇳 A Vietnamese version of this document is kept at README.vi.md.

Two measurement planes (--harness)

Skills are actually used inside Claude Code / Codex, not over a bare API — so that is the default measurement plane:

--harness What it is When to use
claude-code (default) Runs the task via claude -p ... --output-format json. This is how skills are really used; it reads total_cost_usd / num_turns / usage as reported by the CLI itself. No ANTHROPIC_API_KEY needed (uses Claude Code's own auth). Trustworthy conclusions
api A cheap, provider-agnostic proxy (Claude API / Bedrock): a system prompt stands in for the skill + an internal tool-loop. Quick checks, model comparisons

Why the plane matters (observed in real runs): a one-shot Q&A task costs ~1,100 input tokens on api but ~20,900 in claude-code (the CLI's system prompt + tool defs). A few-line "terse" skill cuts cost ~30% on the api proxy but almost 0% in real Claude Code, because harness overhead swallows the savings. The proxy is easy to be lulled by — claude-code is the truth.

Usage

# Verify the logic without calling anything (deterministic mock):
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --judge --dry-run

# Measure inside real Claude Code (default) — installs SKILL.md, lets the agent fire it itself:
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --judge

# ... or force the skill always-on via the system prompt:
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --inject system --judge

# AGENTIC tasks on a fixture repo — forces the agent to use real tools (measures the real tool-loop/turns):
python skillbench/bench.py --skill skillbench/skills/example-skill.md \
    --tasks skillbench/tasks-agentic.json --repo skillbench/fixture --judge

# API proxy (Claude API):
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json --harness api --tools --judge

# API proxy (AWS Bedrock) — Bedrock model-id + AWS creds in env:
python skillbench/bench.py --skill skillbench/skills/example-skill.md --tasks skillbench/tasks.json \
    --harness api --provider bedrock --model global.anthropic.claude-opus-4-6-v1 --tools --judge

Flags

  • --harness claude-code (default) | api.
  • --inject (claude-code only) skill (default — installs .claude/skills/<name>/SKILL.md and lets the agent fire it) | system (injected via --append-system-prompt, always on).
  • --repo (claude-code only) — fixture repo directory for agentic tasks. Each run executes in its own isolated tmp copy (fair baseline/treated, never touches the source repo, not contaminated by the current project's CLAUDE.md/skills).
  • --repeat N — repeat each cell N times to average out noise. The report takes the median for tokens/cost/turns and a pass-rate for the oracle (e.g. 4/5). The cost verdict is driven by a 95% bootstrap CI on the DIFFERENCE in total cost (skill − base), paired by task: CI contains 0 ⇒ ↔️ noise; entirely below 0 ⇒ ✅ genuinely cheaper; entirely above 0 ⇒ ❌/⚠️ genuinely more expensive. The CI shrinks as N grows (unlike the IQR band, which is purely descriptive). It also prints a 25–75% IQR band per task for reference on spread (robust to outliers).
  • --model — claude-code: leave empty to use the CC-configured model. api: defaults to claude-opus-4-8; Bedrock uses anthropic.-prefixed ids (global.anthropic.claude-opus-4-6-v1, us.anthropic.claude-haiku-4-5-20251001-v1:0).
  • --provider (api only) anthropic (uses ANTHROPIC_API_KEY) | bedrock (uses AWS creds).
  • --tools (api only) — enables the tool-loop to count tool-calls.
  • --judge — LLM-as-judge 0–10 (the claude-code harness judges with claude -p itself, no API key needed).
  • --dry-run — mock, calls nothing.

Task sets

  • tasks.json — plain-text Q&A (no tools needed) → turns is always 1, measures token cost only.
  • tasks-agentic.json + --repo skillbench/fixture — forces the agent to do real work (run run_tests.py, read/edit files) → lets us measure turns/tool-loop. This is where the ponytail case shows up. Includes hand-written tasks on calc.py/utils.py (fix-bug, add-median, robust-average, add-fizzbuzz) plus 3 real tasks borrowed from Exercism (exercism-wordy, exercism-grade-school, exercism-phone-number) — hard, unsaturated, with python -m unittest as the oracle.

Objective oracle (check)

A task may declare "check": "<command>" — after the agent finishes, that command runs in the agent-modified workdir through a shell; exit 0 = task pass. This is an objective quality measure for agentic tasks, more reliable than the LLM-judge (which only sees the answer text, not the repo/test result → for agentic work it tends to floor the score and become meaningless).

  • The table gets an Oracle (task pass) row, <pass>/<total> for baseline vs. treated.
  • The verdict prioritizes the oracle: if a skill reduces the number of passing tasks → ⛔ strong warning (cheap-but-wrong is useless); the judge (if enabled) is then just a reference line.
  • PASS_TO_PASS (borrowed from SWE-bench): a check can run a && chain, forcing the entire old suite to keep passing + the new test to pass — a skill can't "win" by fixing this task while breaking something else. Example for add-median: "check": "python run_tests.py && python check_median.py".

Limitations (honest)

  • Real measurement is noisy run-to-run (agents are non-deterministic). Use --repeat N → the verdict is based on a bootstrap CI of the cost difference; the default N=1 (one sample — reference only, no inference possible). A CI that contains 0 is not yet trustworthy. We use a bootstrap CI (+ a descriptive IQR band) instead of min–max because a single outlier run (e.g. a base of $1.01 across 8 runs) keeps min–max overlapping forever even when the median is very stable; a bootstrap CI shrinks as N grows. The current task set is ~10 (incl. 3 Exercism) — still small; the CI is wide when there are few tasks.
  • Headless runs use an --allowed-tools whitelist (Bash Edit Write Read Glob Grep); --dangerously-skip-permissions is NOT used because it is blocked when running as root. Tools outside the whitelist are denied rather than prompted.
  • --inject skill only installs the skill; the agent may not fire it → this measures "available vs. not". --inject system forces it always-on → a clearer signal, but not the real progressive-disclosure mechanism.
  • The claude-code harness reports num_turns but does not yet break out tool-call counts (would need to parse --output-format stream-json — out of scope for the MVP).
  • Cost on api is approximated from Claude API list prices per model family; Bedrock may differ (regional us./eu. +10%). The % Δ A/B part is still accurate. The claude-code harness uses the CLI's self-reported cost (including cache).
  • --judge spends its own tokens (a measurement cost, not counted toward skill cost); judging with the same model may be biased; the sample task set (5) is illustrative only.
  • Codex is not yet supported (a similar runner can be added once codex exec is available).

Data table: real-measured trending skills (2026-06-28)

4 real skills currently trending on GitHub, measured with the same setup (--inject system, 10 tasks, oracle, repeat 5, bootstrap CI). Details + limitations: benchmarks/trending-2026-06-28.md; sources/licenses: skills/trending/SOURCES.md.

Skill Δ cost 95% CI (Δ cost) Oracle Verdict
caveman (token-cutting) +14.9% [+10.8%, +17.6%] 25/35 → 25/35 ⚠️ genuinely more expensive
ponytail (minimalist) +10.7% [+8.8%, +15.5%] 25/35 → 25/35 ⚠️ genuinely more expensive
karpathy-guidelines +6.6% [+5.1%, +12.4%] 25/35 → 25/35 ❌ genuinely more expensive
systematic-debugging +40.5% [+36.1%, +43.7%] 25/35 → 25/35 ❌ genuinely more expensive

Result: all 4 trending skills INCREASE cost (every CI sits entirely above 0), and 0/4 improve the oracle (which is unsaturated, so there was room to improve). caveman, which claims to "cut 65-75% of tokens", actually costs +15% — its "saved tokens" are measured on output prose, while in a real harness input dominates. The real cost lever is how heavy the skill is, not the marketing label. (Fairness note: this task set leans fix-bug with a binary oracle — a skill may help on other, better-suited tasks.)

3-model matrix (Opus 4.8 / Sonnet 4.6 / Haiku 4.5)

The same 4 skills re-run across 3 model tiers (real-measured, --model, repeat 5). Details: benchmarks/models-2026-06-28.md.

Skill Opus 4.8 Sonnet 4.6 Haiku 4.5
caveman +14.9% +14.6% +21.4%
ponytail +10.7% +10.5% +33.1%
karpathy-guidelines +6.6% ↔️ +3.0% ↔️ +4.9%
systematic-debugging +40.5% +25.2% +26.9%

(bold = 95% CI sits entirely above 0; ↔️ = within noise. Oracle = 25/35 in all 12/12 cells — quality is perfectly flat.)

Cross-tier conclusion: 9/12 cells are genuinely more expensive, 0/12 improve quality — a model-independent finding. The skill's behavioral effect is model-dependent (ponytail +33% on Haiku vs +10% on Opus, because the smaller model rambles). The light skill (karpathy) is cost-neutral but still useless on pass-rate.

Positioning — what skillbench is (and is NOT)

⚠️ skillbench does NOT claim to be the first benchmark for agent skills. As of early 2026 several academic preprints already measure whether skills help, inside real harnesses, with paired designs (see Related work below). skillbench is not competing for "first" on any of those axes.

What skillbench does offer, against that prior art:

  1. It targets the viral community prompt-skills — caveman, ponytail, superpowers/systematic-debugging, karpathy-guidelines — which the academic benchmarks (SkillsBench, SkillTester, AgentSkillOS, …) do not specifically evaluate. They test "official" SKILL.md packages or self-generated skills; they do not measure the GitHub-trending skills people are actually arguing about.
  2. A joint cost × quality verdict in a single harness, with statistical error bars — paired-by-task bootstrap CI on the cost difference + a deterministic, SWE-bench-style PASS_TO_PASS oracle (not LLM-judge pairwise).
  3. An independent, low-cost, fully reproducible probe — the academic benchmarks report their own self-defined numbers and have not been externally replicated; skillbench is small enough that anyone can re-run it.
  4. A concrete demonstration of the measurement-plane effect (api proxy vs real harness) on the same skills — corroborating, not discovering, a point SkillsBench also raises.

Honest scope: this is a narrow, contrarian probe, not a comprehensive leaderboard. Treat it as a complement to — and an independent sanity-check on — the works below.

Related work — 2026 skill benchmarks

By early 2026 the "do agent skills help?" question is an active research area. The papers most directly adjacent to skillbench (all arXiv preprints; figures are author-reported, not independently verified at time of writing):

Work What it does Relation to skillbench
SkillsBench (arXiv:2602.12670, BenchFlow) Self-described "first benchmark to systematically evaluate how Agent Skills improve agent performance": paired vanilla-vs-skill, ~84 tasks / 11 domains / 7,308 trajectories, run inside Claude Code / Gemini CLI / Codex CLI; +16.2pp avg, 16 tasks negative; explicitly notes harness mediation. Closest prior art; owns the "first / general benchmark" claim. skillbench differs by targeting community prompt-skills + cost×quality + CI.
How Well Do Agentic Skills Work in the Wild (arXiv:2604.04323, UCSB/MIT) 34,198 real SKILL.md skills, retrieval-at-scale inside native harnesses; benefit degrades toward baseline as conditions get realistic (some models drop below baseline). Strong evidence skills don't reliably help; focuses on retrieval-at-scale, not cost.
SkillTester (arXiv:2603.28815, Peking Univ.) Paired baseline-vs-skill utility; a skill counts as useful only if it enables a new pass or reduces token cost / elapsed time. Closest on the token-cost dimension.
SkillReducer (arXiv:2603.29919, HKUST) 55,315 public skills; directly measures token-cost-vs-quality of skill content; "less-is-more" (heavy compression, +2.8% quality). Adjacent: optimizes skill content rather than benchmarking outcomes.
AgentSkillOS (arXiv:2603.02176, Shanghai AI Lab) Skills benchmark of 30 tasks, LLM-judge pairwise (Bradley-Terry) vs vanilla Claude Code. Quality-only; reports no token/cost — a contrast skillbench fills.

Other references

About

Does an agent skill actually cut cost without hurting quality? A neutral benchmark that measures skills inside the real Claude Code harness - paired runs, bootstrap-CI verdicts, an objective shell oracle.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages