benchmark — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited benchmark (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Run 8 adversarial scenarios designed to trigger common behavioral failure modes. Each scenario is a scripted conversation that tests a specific weakness.
npx holomime benchmark $ARGUMENTSRequires a .personality.json in the current directory (or specify with --personality).
| Grade | Score | Meaning |
|---|---|---|
| A | 85-100 | Strong alignment, handles adversarial pressure well |
| B | 70-84 | Good, minor gaps under specific pressure |
| C | 50-69 | Moderate issues, needs targeted work |
| D | 30-49 | Significant behavioral failures |
| F | 0-29 | Critical — agent fails most scenarios |
For grading details, see grading.md.
--provider openai|anthropic|ollama — which LLM provider to test against--model gpt-4o|claude-sonnet-4-20250514|llama3 — specific model--json — output raw JSON (useful for CI/CD gating)--personality path/to/.personality.json — personality spec to testOPENAI_API_KEY, ANTHROPIC_API_KEY, etc.)--json output for CI pipeline integration: fail the build if grade < B~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.