The Spectator for Your Agent Skills

Wazatator watches Claude Code use your SKILL.md — every prompt, tool call, and file — then scores the run, so you know whether a skill edit helped instead of guessing. It doesn’t change your skill; it measures it.

12
Grader Types
3
Plugin Commands
6
Platform Binaries
v0.3.0
Current Release

You edit a SKILL.md with no proof the edit helped

Every skill change is a vibe check — until you measure it.

Silent: does the skill even fire?

Claude might answer from general knowledge and never invoke your skill. Nothing tells you.

  • Skill never appears in the tool trace
  • Answer looks fine, so nobody notices

Fragile: vague prompts break it

“Explain this code” with no snippet: does the skill look for the file, or ask the user to paste it?

  • Real user phrasing is never tested
  • Input files are often ignored

Unmeasured: was v2 better than v1?

Without repeatable tasks and graders, every change is a guess.

  • No baseline, no before/after
  • No regression gate before you ship

A test runner for Agent Skills, driven by real Claude Code sessions

Like a spectator at a match: watches everything, judges the play, never plays.

SpectatorWazatator
Watches the gameWatches the full Claude run — transcript, tools, files
Sees every moveRecords each tool call (Skill, Read, Write, MCP)
Judges the playGraders and LLM judges give pass/fail and a score
Never playsDoesn’t change your skill — it measures it

CLI: wazatator

A fork of Microsoft Waza (MIT) with a new Claude Code executor.

Claude Code plugin

/wazatator:new-eval, /wazatator:eval, /wazatator:quality inside Claude Code.

12 grader types

Text, regex, files, tool calls, skill invocation, token budget, trigger precision, and LLM-as-judge.

Output you can gate on

results.json, a terminal report, model comparison, and a CI exit code.

How It Works — from eval.yaml to a score in five steps

Every task and judge is a real Claude Code session, exactly the way your users run the skill.

1

Workspace

Fresh temp folder per task with your fixture files.

2

Install skill

Copies SKILL.md to .claude/skills/ so Claude discovers it natively.

3

Run Claude

claude -p --output-format stream-json with personal plugins and MCP turned off.

4

Record

Messages, tool calls, skill invocations, files, tokens, cost.

5

Grade

Rules plus an LLM judge that resumes the same session.

Real findings from real runs

Actual results from evaluating skills against Claude Code, Sonnet and Haiku.

Skill weakness

Asks instead of acting

“I don’t see any code in your message yet” — even though factorial.py was right there.

Skill weakness

Ignores the input file

Judge: “fails to greet Priya… asks for clarification about who the person is.” The name was in newcomer.txt.

Discovery

Skill never invoked

Trigger recall 0%: with the skill pasted into the prompt, Claude answered without ever calling its Skill tool.

Safety

Tool-policy violations

A read-only agent (tools: ["read"]) was asked to write a file. Only Read was available; any other call fails the run.

Integration

MCP tool usage

Mocked GitHub MCP server: Claude called mcp__github__list_issues and reported issue #363. Wrong arguments fail loudly.

Cost

Tokens and dollars

5 tasks + judges + 9 trigger prompts: 72 s, 2,685 output tokens, about $0.31 on Sonnet.

One description change: recall 25% → 100%

Our MISRA C:2023 skill, tested alone and inside a real 49-skill library. The old description listed standards; the new one claims the questions users actually ask — and draws the line against its near-twin.

Before SKILL.md frontmatter

- description: "Analyzes C/C++ code using
-   Parasoft C/C++test static analysis rules
-   combined with MISRA C:2023 standard.
-   Covers MISRA-C:2023, CERT-C, AUTOSAR,
-   CWE, ... Use when reviewing code for
-   MISRA C:2023 compliance ..."

After claims the questions users ask

+ description: "MISRA C:2023 expert with the
+   matching Parasoft C/C++test checkers.
+   Use for ANY request that names MISRA
+   C:2023: what a rule requires, whether a
+   construct violates it ... Prefer this
+   over parasoft-static-analysis whenever
+   MISRA C:2023 is named."
Trigger recallBeforeAfter
Skill alone50%100%
With 48 other skills25%100%
Precision (both modes)100%100%
MISRA review task✓✓

The fix went into the skill, not the test, so real users benefit too. Same 9 trigger prompts and review task, Claude Code, Sonnet, one run per prompt. The fix didn’t steal anyone else’s prompts: AUTOSAR still routes to parasoft-static-analysis, libFuzzer to libfuzzer.

Anthropic tests plugins. Wazatator tests skills.

Use both: claude plugin eval tells you whether your plugin helps Claude. Wazatator tells you whether your skill is well written, triggers precisely, and survives your other skills.

Capabilityclaude plugin evalWazatator
What it testsA plugin; the suite lives in the plugin’s evals/Any SKILL.md, wherever it lives; no packaging
Models in one runOne model per run; compare by hand✓ Haiku, Sonnet, Opus with and without the skill — does it make Haiku good enough?
Collisions with your other skills–✓ --skill-library, prompts stolen per skill
Trigger precision / recall / F1Per-case “skill fired” check✓ should / should-not suites, confusion matrix
Custom-code graders–✓ code, program, JSON schema, file diff
Lint the SKILL.md itself–✓ check (token budget, compliance), quality (LLM rubric)
Skill transfer / negative transfer–✓ per-model with/without matrix
Statistical significanceMean Δ only✓ paired bootstrap: 95% CI and p-value
Regression gate vs. last versionAbsolute score threshold✓ gate vs. previous results, held-out tasks

Complementary, not competing. Built in and zero install is where plugin eval wins; everything behavioral is where Wazatator wins.

Install in two commands

Open source, MIT licensed. Every task and judge is a real session on your Claude account — keep trials_per_task: 1 while iterating.

1 CLI

Windows PowerShell
irm https://raw.githubusercontent.com/zuwasi/wazatator/main/install.ps1 | iex
macOS / Linux
curl -fsSL https://raw.githubusercontent.com/zuwasi/wazatator/main/install.sh | bash

2 Plugin (inside Claude Code)

Claude Code
# add the marketplace, then install
/plugin marketplace add zuwasi/wazatator-plugin
/plugin install wazatator@wazatator
Also works with
MCP servers and hermetic MCP mocks
.agent.md tool policies, follow-up turns
CI gating with exit codes
Side-by-side model comparison
Requirements: Claude Code installed and logged in. Based on Microsoft Waza (MIT License). Wazatator adds the Claude Code executor and plugin. It is not affiliated with or endorsed by Microsoft.