skill-eval-harness
Enables systematic, quantifiable assessment of AI agent behavior across variants, reducing uncertainty in production deployments and identifying performance …
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
- Evaluate code changes across variant implementations to measure performance differences
- Generate trace artifacts showing execution paths for debugging production issues
- Compare runner adapters side-by-side to validate engineering decisions before deployment
Enables systematic, quantifiable assessment of AI agent behavior across variants, reducing uncertainty in production deployments and identifying performance regressions before release.
ML teams evaluating agentic systems and engineering leads standardizing AI quality gates.
https://github.com/adewale/skill-eval-harness
By adewale
How to Get It
claude plugins install adewale/skill-eval-harness
Tip: Paste this into a Claude Code conversation. Verify command matches your Claude Code version.
Auto-generated from the tool's public listing — not hands-on verified. Cross-check against the source repo's README before running.
After installing, paste this into Claude:
Help me evaluate code changes across variant implementations to measure performance differences
Trust Signals Auto-scanned
Community Pulse New
No community discussions found yet. This doesn't mean the tool isn't good — it may be new or serve a niche use case.
Reviewer notes
Auto-scanned review. These are observations, not a security certification.
Scored from trust signals (evidence-eval-v1): 59 GitHub stars; contributors unknown; last commit 0d ago; license MIT.
Things to check
- Scanned, not hands-on tested — this entry was auto-scanned from public metadata (GitHub metrics, license, security flags). No reviewer has run it, and no tool-specific limitations have been documented yet.
How to evaluate tools before deploying →
Data shown here comes from public APIs and automated scanning. Reviewer notes reflect one person's experience. This is not a security certification or legal recommendation. Always evaluate tools according to your own organization's policies.
Evaluation
Scored from trust signals (evidence-eval-v1): 59 GitHub stars; contributors unknown; last commit 0d ago; license MIT.