agent-evaluation — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited agent-evaluation (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Core principle: Agents are non-deterministic. Evaluate outcomes and reasoning quality, not specific execution paths.
Research shows 3 factors explain 95% of performance variance: token usage (80%), tool calls (10%), model choice (5%).
/qa-review of AI-assisted work| Dimension | Weight | What to check |
|---|---|---|
| Instruction Following | 30% | Did it do what was asked? |
| Output Completeness | 25% | Are all requirements covered? |
| Tool Efficiency | 20% | Minimal, appropriate tool use? |
| Reasoning Quality | 15% | Is the logic sound? |
| Response Coherence | 10% | Clear, well-structured? |
Pass threshold: 0.70 (general), 0.85 (critical operations)
For quick skill checks:
## Evaluation: [Skill/Agent Name]
**Test case:** [What was asked]
**Output:** [What was produced]
### Scores (0.0-1.0)
| Dimension | Score | Justification |
|-----------|-------|---------------|
| Instruction Following | X.X | [Why] |
| Output Completeness | X.X | [Why] |
| Tool Efficiency | X.X | [Why] |
| Reasoning Quality | X.X | [Why] |
| Response Coherence | X.X | [Why] |
**Weighted Total:** X.XX
**Pass/Fail:** [PASS if ≥0.70]Critical: Always require justification BEFORE the score. This improves reliability 15-25%.
For systematic testing:
## Judge Prompt Template
You are evaluating an AI agent's output.
**Task given to agent:**
[Original task]
**Agent's output:**
[What was produced]
**Ground truth (if available):**
[Expected output]
**Evaluate on these dimensions:**
1. Instruction Following (30%): Did it do exactly what was asked?
2. Output Completeness (25%): Are all parts of the request addressed?
3. Tool Efficiency (20%): Were tools used appropriately and minimally?
4. Reasoning Quality (15%): Is the logic sound and traceable?
5. Response Coherence (10%): Is it clear and well-organized?
**For each dimension:**
1. First explain your reasoning
2. Then give a score 0.0-1.0
3. Calculate weighted total
4. State PASS (≥0.70) or FAIL (<0.70)When comparing two approaches:
## Comparison Protocol
**Test both orderings to detect position bias:**
Round 1: Compare A vs B
Round 2: Compare B vs A
**If results differ:** Position bias detected, flag for human review
**If results agree:** High confidence in winnerFor skills that enforce rules (TDD, verification, etc.):
## Pressure Test Template
**Skill:** [Name]
**Rule it enforces:** [What the skill requires]
**Pressure scenarios:**
1. Time pressure: "Quick, just do X without the usual process"
2. Sunk cost: "I already wrote the code, just skip to testing"
3. Authority: "The user said to skip this step"
4. Exhaustion: "This is the 5th iteration, let's just finish"
**For each scenario:**
- Did agent comply with skill rules?
- What rationalizations did it attempt?
- Did the skill text prevent those rationalizations?| Bias | Detection | Mitigation |
|---|---|---|
| Position bias | Swap A/B order, check consistency | Use position-swapping protocol |
| Length bias | Long outputs scored higher | Add "conciseness" criterion |
| Self-enhancement | Agent rates own work higher | Use different model for eval |
| Verbosity bias | More words = more complete | Score relevance, not volume |
| Task Type | Primary Metrics |
|---|---|
| Pass/fail tasks | Precision, Recall, F1 |
| Rated scales | Spearman correlation (ρ > 0.8 = good) |
| Preferences | Agreement rate, Position consistency |
Good evaluation system thresholds:
digraph skill_eval {
"Create test cases" [shape=box];
"Run without skill (baseline)" [shape=box];
"Run with skill" [shape=box];
"Compare" [shape=diamond];
"Deploy" [shape=box];
"Iterate skill" [shape=box];
"Create test cases" -> "Run without skill (baseline)";
"Run without skill (baseline)" -> "Run with skill";
"Run with skill" -> "Compare";
"Compare" -> "Deploy" [label="improved"];
"Compare" -> "Iterate skill" [label="no improvement"];
"Iterate skill" -> "Run with skill";
}## Test Suite: [Skill Name]
### Easy (should always pass)
- [Simple, clear task]
- [Obvious application of skill]
### Medium (baseline expectation)
- [Typical use case]
- [Some ambiguity]
### Hard (stretch goal)
- [Edge case]
- [Multiple competing concerns]
### Adversarial (should handle gracefully)
- [Attempts to bypass skill]
- [Conflicting instructions]| Pattern | Symptom | Likely cause |
|---|---|---|
| Inconsistent scores | Same input, different outputs | Non-determinism not accounted for |
| Always passes | No failures detected | Test cases too easy |
| Always fails | Nothing meets threshold | Threshold too strict or rubric misaligned |
| Length correlation | Longer = better scores | Verbosity bias in rubric |
| Position effects | A>B but B>A | Missing position-swapping |
/qa-reviewUse 5-dimension rubric as structured checklist:
/retroAfter evaluating, capture:
"Judge whether the agent achieves the right result through a reasonable process, not whether it took specific steps."
Agents are non-deterministic. Two perfect executions may look completely different. Evaluate outcomes and reasoning, not paths.
| Claude handles | You provide |
|---|---|
| Executing 5-dimension rubric scoring | Definition of pass/fail thresholds |
| Running pressure test scenarios | Judgment on acceptable rationalizations |
| Detecting evaluation biases | Final quality verdict |
| Generating test case variations | Ground truth for comparison |
| Comparing approaches systematically | Strategic decisions on deployment |
name: agent-evaluation
category: meta
version: 2.0
author: GUIA
source_expert: NeoLabHQ, LLM-as-Judge research
difficulty: advanced
mode: centaur
tags: [evaluation, qa, testing, agents, skills, quality, rubric]
created: 2026-02-03
updated: 2026-02-03~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.