eval-refine — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited eval-refine (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Founder says:
Or dashboard pipeline Stage 8 triggers.
docs/ideas/*/eval.md (idea-eval historical 6-dimension scores)tests/{skills,agents}/results/*.md (skill / agent eval historical pass rate)docs/learnings/candidates/<YYYY-MM-DD>-rubric-refine.md, structured as follows:
---
type: rubric-refine
generatedAt: 2026-04-25
generatedBy: [email protected]
scope:
- idea-eval
- skill:lead-intake
- skill:idea-eval
analyzedRecords: 14
candidatePatches: 3
status: candidate # candidate / approved / rejected
---
# Rubric Adjustment Candidates · 2026-04-25
## Candidate 1 · idea-eval "Ability to Pay" dimension too loose
**Evidence**:
- 14 ideas · 12 scored ≥4 on this dimension · but only 3 actually generated revenue within 1 year
- Inference: scores didn't distinguish "can pay" from "will pay"
**Suggested Change**:
- Modify prompt to add: "Evaluate not just **ability to pay**, but **probability of paying this year** · ≥4 requires evidence of 'already spending on similar products'"
- Add to rubric should_have: "Interview Q3 must confirm 'what tools are you currently using' with specific numbers"
**Impact**: Re-evaluate all 14 ideas · estimated 4 will drop from 4 to 3 · total score change ≤ 4 points
**Atomic Check**:
- trigger: idea-eval runs "ability to pay" dimension
- action: add should_have check · add prompt wording
- domain: skills/idea-eval
- confidence: 0.7 (based on real 14 data points)
---
## Candidate 2 · skill eval all 100% PASS
**Evidence**:
- 5/6 skills first run 100% · only lead-intake at 75%
- Real problem exposed by lead-intake (Q10 signal strength) drove v1.1 upgrade
- Inference: the other 5 skills' rubric should_have is too broad
**Suggested Change**:
For each 100% skill, add 1-2 picky should_have checks in yaml. Example:
- proposal-gen add: "Quotation section must show 3 tiers (basic/standard/flagship)"
- content-blitz add: "30 pieces of content must maintain brand voice consistency · sample 5 must 100% pass brand-voice check"
**Impact**: Retest 5 skills · estimated 2-3 will land in 75-90% range (healthy signal)
**Atomic Check**:
- trigger: any skill eval pass rate = 100% for ≥2 consecutive runs
- action: add picky should_have (not new features · stricter checks)
- domain: tests/skills/*.yaml
- confidence: 0.85 (based on real 5/6 lesson)
---
## Candidate 3 · ...Scan historical eval → find anomaly patterns → draft candidate
↓
Candidate goes into docs/learnings/candidates/
↓
/learn-review weekly review · founder decides promote / skip / reject
↓
promote → update corresponding SKILL.md / agent.md / yaml
↓
Retest affected cases · verify change is effective~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.