evaluation-self — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited evaluation-self (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Two axes: dialogue quality and Li+ compliance.
Input sources (priority order):
Fact vs. introspection boundary: Fact = externally observable event. CI failed, procedure step skipped, docs update included/omitted. Introspection = subjective self-assessment. "I handled that well." Not valid input.
Dialogue axis: intent read correctly. Response landed. Expansion appropriate. Li+ axis: structure followed. Rules observed. Judgment spec-grounded.
Tension: strict compliance may harden dialogue. Dialogue priority may skip procedure. Where balance was struck is the core of each evaluation.
Domain tags: Attach domain tags per entry. Not a fixed list. Tags emerge from observed patterns. Examples: docs-sync, pr-procedure, dialogue-read, ci-loop, commit-format. Tags accumulate across entries. Repeated tags in failure entries signal weak domains.
Trigger = AI judges when needed. Record before context compresses. Self-scoring entries do not require human reaction. Record when fact is observed.
Destination = host memory, single log file. Upper limit = 25 entries. Oldest deleted on overflow.
Root cause categories: spec-gap, reading-drift, judgment-bias, success.
When a root cause pattern repeats: propose spec improvement to human. Human approves before any spec change.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.