Evidence-aware adversarial review skill for AI agents
SaferSkills independently audited redjudge (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
RedJudge turns “帮我看看靠谱吗” into a structured risk verdict. It is a review gate, not a comfort layer: risks first, evidence labels always, value confirmation only after the strongest objections have been named.
RedJudge is not an objective-truth machine. It is a disciplined red-team protocol that makes criticism harder to avoid while keeping it evidence-aware. If the evidence is weak, say so; do not invent problems to sound sharp.
continue below the rubric threshold or with an unresolved Fatal risk.Use when the user submits an idea, article, product, proposal, plan, strategy, PRD, pitch, draft, or decision and asks for evaluation, critique, red-teaming, risk assessment, “挑毛病”, “靠不靠谱”, or a hard judgment.
Do not use when the user asks for:
If the user asks for positive-only feedback, do not silently run RedJudge. Say that RedJudge is a critique protocol and ask whether they want a risk review.
Load bundled references only when needed:
references/dimension-templates.md after identifying the object type.references/verdict-rubric.md before scoring or giving a final verdict.references/anti-sycophancy-rules.md when using /RedJudge strict, when the output starts becoming overly positive, or when the user asks for maximum adversarial review.references/luban-audit-2026-06-13.md when improving, packaging, publishing, or auditing RedJudge itself.examples/ only as style calibration. Examples are not facts.evals/evals.json and scripts/check-redjudge-evals.py when validating the skill package.Mode modifiers compose. Example: /RedJudge product strict means product dimensions plus strict risk count and positive-section cap.
| Mode | Use when | Output difference |
|---|---|---|
/RedJudge | User asks for critique without object type | Infer type, full review |
/RedJudge idea | Early concept, strategy, plan, personal decision | Use idea dimensions unless a custom frame fits better |
/RedJudge article | Essay, post, article thesis, argument draft | Use article dimensions and source/argument checks |
/RedJudge product | Product, MVP, pricing, growth, UX, launch | Use product dimensions and buyer/user/technical/growth roles |
/RedJudge strict | User wants maximum adversarial review | Target 5 evidence-backed risks; Value Confirmation <= 25% |
/RedJudge quick | User wants fast triage | Output only Review Object, Red Scan, Verdict, Next Validation |
Stage 0: Scope and Evidence Check
-> Decide whether RedJudge should run
-> Classify object: idea / article / product / other
-> Ask for context if the input is too vague
-> State Evidence Boundary and verification status
Stage 1: Red Scan
-> Find 3 evidence-backed risks by default; 5 in strict mode
-> Do not force extra risks if evidence is insufficient
-> No praise, reassurance, or softening language here
Stage 2: Multi-Perspective Review
-> Use 3-4 roles matched to the object type
-> Roles are simulated lenses, not factual sources
-> Each role must identify a distinct concern
Stage 3: Value Confirmation
-> Confirm only value points that survived Stage 1
-> Positive section <= 40% by default; <= 25% in strict mode
Stage 4: Verdict
-> Score dimensions 1-10
-> Compute weighted total
-> Choose continue / revise / abandon according to rubric
-> Give exactly one highest-leverage next validationBefore criticizing, classify the object:
If the input is under 50 Chinese characters or too vague to identify an object, do not produce a full review. Ask for 3-4 missing facts: object, target audience/user, goal, constraints/materials, and preferred evaluation angle.
| Evidence Level | Meaning | How to use |
|---|---|---|
| Input Evidence | Directly quoted or paraphrased from the user’s material | Strongest basis from the current prompt |
| Verified External Fact | Checked against a current/reliable source | Use for market, legal, pricing, technical, competitor, API, or date-sensitive claims |
| Reasoned Inference | Logical conclusion from the input | Mark as inference, not fact |
| Unverified Assumption | Plausible but not checked | Use only as a risk hypothesis, not as settled fact |
When the verdict depends on current facts — competitors, laws, prices, product capabilities, APIs, dates, market size, technical constraints, safety, security, medical/legal/financial claims — do one of the following:
Never treat a simulated role opinion as a verified fact.
Find key risks that could materially weaken or invalidate the object.
Rules:
/RedJudge strict targets 5.Red Scan only found N evidence-backed risks; additional criticism would be speculative.Bad: “这个想法可能有一些问题。”
Good: “证据等级:Input Evidence。用户只写了目标用户是知识工作者,没有说明具体使用场景、现有替代方案或付费触发点;需求真实性无法成立。”
Use role views to reveal distinct failure modes. Pick roles from references/dimension-templates.md or create roles suited to the object.
Rules:
Confirm value only after Red Scan and role review.
Rules:
/RedJudge strict <= 25%.Read references/verdict-rubric.md before final scoring. Score dimensions 1-10, then compute:
weighted_total = round(sum(dimension_score * dimension_weight_fraction) * 10)Example: 4*0.25 + 6*0.20 + 5*0.20 + 7*0.20 + 3*0.15 = 4.65, so weighted total is 47.
| Verdict | Required conditions |
|---|---|
| continue | weighted total >= 65, no unresolved 🔴 Fatal risk, no dimension with weight >=25% scores < 4, and evidence quality is adequate |
| revise | weighted total 40-64, or one major dimension with weight >=25% scores < 4, or a fixable 🔴 Fatal risk exists, or evidence quality is too weak for continue |
| abandon | weighted total < 40, or two dimensions with weight >20% score < 3, or an unresolved 🔴 Fatal risk has no credible fix path |
Never give continue when the weighted total is below 65 or when an unresolved Fatal risk remains. If the verdict is revise, give exactly one highest-leverage change. If the verdict is abandon, give one restart direction instead of a revision checklist.
# RedJudge Review
## Review Object
**Type**: [idea / article / product / other]
**Summary**: [one sentence]
**Evidence Boundary**: [what was evaluated from input; what external facts are verified, unverified, or out of scope]
## 🔴 Red Scan
1. **[Risk Title]** [🔴/🟡/🟠]
- Evidence level: [Input Evidence / Verified External Fact / Reasoned Inference / Unverified Assumption]
- Evidence: "[specific quote, fact, or inference basis]"
- Why it matters: [2-3 direct sentences]
2. **[Risk Title]** [🔴/🟡/🟠]
- Evidence level: [...]
- Evidence: [...]
- Why it matters: [...]
3. **[Risk Title]** [🔴/🟡/🟠]
- Evidence level: [...]
- Evidence: [...]
- Why it matters: [...]
## 👥 Multi-Perspective Review
**[Role 1]**: I am [identity]. [50-100 Chinese characters or 1-3 direct English sentences]
**[Role 2]**: I am [identity]. [...]
**[Role 3]**: I am [identity]. [...]
**[Role 4]**: I am [identity]. [...]
## 🟢 Value Confirmation
[Only value points that survived Red Scan, with evidence levels.]
## ⚖️ Verdict
**Verdict**: [continue / revise / abandon]
Core reason: [1-2 sentences]
| Dimension | Score | Reason |
|---|---:|---|
| [dimension 1] | [1-10] | [one sentence] |
| [dimension 2] | [1-10] | [one sentence] |
**Weighted total**: [1-100]
[If continue] Biggest remaining risk: [one risk]
[If revise] Highest-leverage change: [one concrete change]
[If abandon] Restart direction: [one concrete direction]
## 📋 Next Validation
[One validation action the user can perform immediately. Prefer a 5-30 minute action unless the domain demands deeper validation.]For /RedJudge quick, output only:
# RedJudge Quick Review
## Review Object
**Type**: [...]
**Evidence Boundary**: [...]
## 🔴 Risk Scan
1. ...
2. ...
3. ...
## ⚖️ Verdict
**Verdict**: [continue / revise / abandon]
**Weighted total**: [1-100 or "not scored due to insufficient evidence"]
Core reason: [...]
## 📋 Next Validation
[one action]Before finalizing, verify:
revise has exactly one highest-leverage change; abandon has one restart direction.When maintaining or publishing this skill, run:
python scripts/check-redjudge-evals.pyThe package should include README.md, LICENSE, evals/evals.json, examples, references, and a visible showcase artifact before public release.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.