pm-metrics-critic — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited pm-metrics-critic (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Reviews a set of metrics — a dashboard, a success-criteria section, a North Star proposal — against decision-making/metrics.md. The core test: does each metric, if it moved, tell us something strategic, or does it just tell us the product is being used? Most metric sets fail this test and the skill's job is to say so concretely.
Use this when:
Don't use this when:
pm-prd-drafterdecision-making/metrics.md § metric trees and walk it manuallydecision-making/metrics.md § "Match metrics to strategy assumptions." For the work in question, list the load-bearing strategic assumptions. If the user can't articulate them, surface that gap before evaluating any metric — you can't grade metrics without strategy.decision-making/metrics.md § counter-metrics: every primary metric should have a counter-metric that catches the obvious gaming path. Engagement up + retention down is a different story than engagement up + retention flat. Flag any metric that ships without a counter.## TL;DR
[One-paragraph honest take. Are these metrics measuring the strategy or the activity?]
## Strategic assumptions in play
1. [Assumption 1 — load-bearing for this work]
2. [Assumption 2]
3. ...
## Metric-by-metric
| Metric | What it would falsify | Verdict |
|---|---|---|
| [Metric A] | [Assumption / nothing — vanity] | 🟢 strategic / 🟡 weak / 🔴 vanity |
| [Metric B] | ... | ... |
## Failure modes flagged
- **Aggregate masking segment failure:** [Which metric, what segment is hidden]
- **No counter-metric:** [Which metric, what the obvious gaming path looks like]
- **Lagging-only:** [Which metric, what leading indicator should sit alongside it]
- **Decoupled from strategy:** [Which metric, what assumption it claims to validate but doesn't]
## Recommended replacement set
1. **[Primary metric]** — falsifies *[assumption]*. Segment: [specific]. Threshold: [specific]. Counter-metric: [specific].
2. **[Primary metric 2]** — ...
3. ...
## What an exec will press on
[The 1-2 metrics questions the team will be asked that they're not yet ready for.]~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.