BrowseFull catalogOutcomesSolve a specific problemRolesStack by teamTrustFilter by risk tier
← Back to the Claude Observatory

AgentCompass

Skill Infrastructure Usable
Works inClaude Code
Usable Scanned — metadata only

Systematically measuring agent performance across diverse tasks reduces deployment risk by identifying capability gaps before production use.

[EMNLP 2026] AgentCompass is an extensible open-source evaluation infrastructure for systematically assessing LLM/VLM agent capabilities.

128 starsApache-2.0 (commercial OK)FreeQuick setup
Usable rating — This tool is functional but has notable gaps. Review the evaluation notes below before deploying.

Systematically measuring agent performance across diverse tasks reduces deployment risk by identifying capability gaps before production use. Extensible evaluation frameworks prevent costly misalignment between model behavior and business requirements.

DevOps and infrastructure teams standardizing LLM agent evaluation before enterprise deployment.

Claude Code Claude Cowork Claude Chat

https://github.com/open-compass/AgentCompass

By open-compass

How to Get It

Option 1: Claude Desktop App (Code Mode)Click the + button next to the prompt box → PluginsAdd plugin. Search and click Install. Skills work in Claude Code only.
Option 2: Paste into Claude CodeCopy the command below and paste it into your conversation. Claude will install it.
Command
claude plugins install open-compass/AgentCompass

Tip: Paste this into a Claude Code conversation. Verify command matches your Claude Code version.

Auto-generated from the tool's public listing — not hands-on verified. Cross-check against the source repo's README before running.

First thing to try

After installing, paste this into Claude:

Help me evaluate how well my AI agents handle real infrastructure deployment scenarios
CostFree

Trust Signals Auto-scanned

Stars128Last updated2026-09-11LicenseApache-2.0 (OK for commercial use)Known CVEsNone foundSources: GitHub Advisory Database + OSV.dev · Scanned 2026-09-11 · scanner v1

Community Pulse Quiet

No community discussions found yet. This doesn't mean the tool isn't good — it may be new or serve a niche use case.

Reviewer notes

Auto-scanned review. These are observations, not a security certification.

Scored from trust signals (evidence-eval-v1): 128 GitHub stars; contributors unknown; last commit 0d ago; license Apache-2.0.

Things to check

  • Scanned, not hands-on tested — this entry was auto-scanned from public metadata (GitHub metrics, license, security flags). No reviewer has run it, and no tool-specific limitations have been documented yet.

How to evaluate tools before deploying →

Data shown here comes from public APIs and automated scanning. Reviewer notes reflect one person's experience. This is not a security certification or legal recommendation. Always evaluate tools according to your own organization's policies.

Evaluation

Ease of Use
3/5
Versatility
3/5
Reliability
3/5
Security
3/5
Overall score3.00 / 5.00 UsableEvaluatedSep 2026
Scored from trust signals (evidence-eval-v1): 128 GitHub stars; contributors unknown; last commit 0d ago; license Apache-2.0.

← Back to the Claude Observatory

Rolling Claude out in your org? Let's talk.

Start a conversation →