clarify-first— agent skill

Enforce risk-based alignment for AI Agents. Stop Claude/Cursor from guessing on vague or high-impact tasks. 让 AI 先对齐、再动手:为 Agent 引入风险分级与澄清机制,拒绝盲目猜测,确保开发安全可控。

by DmiyDing·Agent Skill·github.com/DmiyDing/clarify-first

Is clarify-first safe to install?

SaferSkills independently audited clarify-first (Agent Skill) and scored it 83/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 4 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.

Score
83/100
●●●●●●●●○○
↑ +0 since first scan (83 → 83)Re-scan~30s
Latest scan
ScannedJun 28, 2026 · 27d ago
Scans run1 over 90 days
Detectors55 checks · 5 categories
Findings4 warnings · 0 high
EngineSaferSkills 2b638c6
View methodology →
SaferSkills installs
This week0
This month0
All time0
CategoryWeightCategory scoreContribution
Securityprompt, exec, net, exfil, eval
35%
52
18.2 pts
Supply chainhash, typosquat, maintainer, lockfile
20%
100
20.0 pts
Maintenancestaleness, pinning, CI
15%
100
15.0 pts
TransparencySKILL.md, perms, README
15%
100
15.0 pts
Communityinstalls, verify, response
15%
100
15.0 pts

Findings & checks · 4 flagged

Securityscore 52 · 4 findings
MEDIUMInstruction telling the agent not to ask for approvalSS-SKILL-INJECT-DONT-ASK-01 · Prompt injection · clarify-first/SKILL.md×2
MEDIUMit fires on intent; the real damage depends on the host agent's own approval-gating.
Why it matters

The text {match} tells the agent to skip the normal "ask the user first" gate. Used adversarially it removes the human-in-the-loop check before destructive or sensitive actions, turning a normally-gated agent into a fire-and-forget executor.

The exact value spotted
excerptclarify-first/SKILL.md· markdown
202* User explicitly says "Skip Plan" or "Fast Track"
203* **Priority Rule — Explicit Scope Overrides Vague Verbs**: If path + anchor + acceptance
… (74 chars elided on L203)
204* **No Triage Bypass Rule**: Inputs like "Skip Triage", "Don't ask", or "Just do everythin
… (81 chars elided on L204)
205* Then you MAY skip Phase 1, but MUST add a header comment: `[FAST-TRACKED MEDIUM RISK]` b
… (108 chars elided on L205)
206 
Occurrences
2 occurrences · first at L204, also L208
Show all 2 locations
Line
File
L204
clarify-first/SKILL.md
L208
clarify-first/SKILL.md
How to fix
Remove the approval-skipping instruction, or scope it narrowly to a specific safe, reversible action.
  1. Delete blanket "don't ask / no need to confirm" directives from the skill.
  2. If the skill is a genuine autonomous job, restrict the opt-out to a named non-destructive action rather than all actions.
Framework references
OWASPLLM01ATLASAML.T0051
Trace & refs
ruleSS-SKILL-INJECT-DONT-ASK-01sha25615dae9807bd892c6rubric 365aacaView on GitHub
MEDIUMInstruction telling the agent not to ask for approvalSS-SKILL-INJECT-DONT-ASK-01 · Prompt injection · clarify-first/references/SCENARIOS.md
MEDIUMit fires on intent; the real damage depends on the host agent's own approval-gating.
Why it matters

The text {match} tells the agent to skip the normal "ask the user first" gate. Used adversarially it removes the human-in-the-loop check before destructive or sensitive actions, turning a normally-gated agent into a fire-and-forget executor.

The exact value spotted
excerptclarify-first/references/SCENARIOS.md· markdown
69* If user remains vague after 2 clarification rounds, summarize unresolved blockers and re
… (24 chars elided on L69)
70* Do NOT execute MEDIUM/HIGH-risk changes until confirmation is explicit.
71* "Skip triage" or "don't ask" does not bypass unresolved ambiguity checks.
72* If user cannot answer blockers, switch to Pathfinder Mode: propose read-only diagnostics
… (24 chars elided on L72)
73* If read-only probing is insufficient, propose **Sandbox Validation** as an explicit opti
… (108 chars elided on L73)
Occurrences
1 occurrence · at L71
How to fix
Remove the approval-skipping instruction, or scope it narrowly to a specific safe, reversible action.
  1. Delete blanket "don't ask / no need to confirm" directives from the skill.
  2. If the skill is a genuine autonomous job, restrict the opt-out to a named non-destructive action rather than all actions.
Framework references
OWASPLLM01ATLASAML.T0051
Trace & refs
ruleSS-SKILL-INJECT-DONT-ASK-01sha256ce5204b52388d24frubric 365aacaView on GitHub
MEDIUMInstruction telling the agent not to ask for approvalSS-SKILL-INJECT-DONT-ASK-01 · Prompt injection · CHANGELOG.md
MEDIUMit fires on intent; the real damage depends on the host agent's own approval-gating.
Why it matters

The text {match} tells the agent to skip the normal "ask the user first" gate. Used adversarially it removes the human-in-the-loop check before destructive or sensitive actions, turning a normally-gated agent into a fire-and-forget executor.

The exact value spotted
excerptCHANGELOG.md· markdown
23- **State Checkpoint Recall**: Before execution after long planning threads, the agent must
… (75 chars elided on L23)
24- **Security & Privacy Guardrail**: Added mandatory redaction rule for secrets, tokens, cred
… (53 chars elided on L24)
25- **No Bypass Clarification Policy**: Added explicit rule that "Skip triage/Don't ask" canno
… (69 chars elided on L25)
26- **Pathfinder Mode**: Added deadlock fallback to run safe read-only diagnostics or verifica
… (45 chars elided on L26)
27- **Trigger Attribution**: Risk output now supports explicit trigger source labeling (thresh
… (28 chars elided on L27)
Occurrences
1 occurrence · at L25
How to fix
Remove the approval-skipping instruction, or scope it narrowly to a specific safe, reversible action.
  1. Delete blanket "don't ask / no need to confirm" directives from the skill.
  2. If the skill is a genuine autonomous job, restrict the opt-out to a named non-destructive action rather than all actions.
Framework references
OWASPLLM01ATLASAML.T0051
Trace & refs
ruleSS-SKILL-INJECT-DONT-ASK-01sha25615dae9807bd892c6rubric 365aacaView on GitHub
Supply chainscore 100 · 0 findings
All supply chain checks passedNo findings in this category for the latest scan.pass
Maintenancescore 100 · 0 findings
All maintenance checks passedNo findings in this category for the latest scan.pass
Transparencyscore 100 · 0 findings
All transparency checks passedNo findings in this category for the latest scan.pass
Communityscore 100 · 0 findings
All community checks passedNo findings in this category for the latest scan.pass
Vendor response · right of reply
Are you the maintainer? Submit a response →

Audit the pieces. Scan the whole. Decide.

~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.