code-review — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited code-review (Agent Skill) and scored it 91/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 1 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 1 flagged
The text {match} is the classic direct prompt-injection phrasing. Placed in a skill body that the agent reads as trusted instructions, it tries to make the agent abandon its prior rules and follow whatever comes next — a full system-prompt override.
ignore/disregard/forget … previous instructions sentence.Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Stage 1 — Spec compliance (do this FIRST): verify the changes implement what was intended. Check against the PR description, issue, or task spec. Identify missing requirements, unnecessary additions, and interpretation gaps. If the implementation is wrong, stop here — reviewing code quality on the wrong feature wastes effort.
Stage 2 — Code quality: only after Stage 1 passes, review for correctness, maintainability, security, and performance.
npm run test, make check, etc.) to catch automated failures before manual review.input is empty here?") instead of declarative statements to encourage author thinking.Correctness:
any types, unchecked casts)Maintainability:
Performance:
## Review: [brief title]
### Critical
- **[file:line]** — [issue]. [What happens if not fixed]. Fix: [concrete suggestion].
### Important
- **[file:line]** — [issue]. [Why it matters]. Consider: [alternative approach].
### Minor
- **[file:line]** — [observation].
### What's Working Well
- [specific positive observation with why it's good]Ground every finding in actual code — no invented line references. Limit to 10 findings per severity. If more exist, note the count and show the highest-impact ones.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.