Files
wshobson-agents/plugins/plugin-eval
Jakob Michael Werner 6fdefba05d fix: monte carlo layer uses nonexistent sdk.stream() instead of sdk.query()
Replace the sdk.stream() call (which doesn't exist in claude-agent-sdk)
with the query() + ClaudeAgentOptions + ResultMessage pattern used by
the judge layer. This was causing every simulation to silently fail,
reporting 100% failure rate and tanking composite scores.

Fixes #477

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-15 01:46:18 -05:00
..

plugin-eval

Three-layer quality evaluation framework for Claude Code plugins.

Quick Start

cd plugins/plugin-eval
uv sync

# Evaluate a skill (static only, instant)
uv run plugin-eval score path/to/skill --depth quick

# Evaluate with LLM judge (~30s)
uv run plugin-eval score path/to/skill --depth standard

# Full certification (all layers, ~5 min)
uv run plugin-eval certify path/to/skill

Layers

  1. Static Analysis — Structural checks, anti-pattern detection. Instant, free.
  2. LLM Judge — Semantic evaluation (triggering, orchestration, output, scope). ~30s, 4 calls.
  3. Monte Carlo — Statistical reliability via 50100 simulated runs. ~25 min.

Commands

CLI Claude Code Description
plugin-eval score /eval Score a plugin or skill
plugin-eval certify /certify Full certification with badge
plugin-eval compare /compare Head-to-head comparison
plugin-eval init Build corpus for Elo ranking

Documentation

See docs/plugin-eval.md for the full reference: layers, dimensions, scoring formula, anti-patterns, statistical methods, and project structure.