tooluniverse-self-review — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited tooluniverse-self-review (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Derive the criteria that define a good result for a specific task, then review work against them. Each task defines its own "evaluation world" -- a set of scenarios, perspectives, and criteria that capture what matters for judging results for that specific task. Built on the Qworld Recursive Expansion Tree (RET).
what a good result must cover, or what is missing or weak.
conversation, treat the whole exchange as context and the last user message as the primary intent. If it includes an image or retrieved web context, factor that in.
supplied, run in checklist mode -- produce the criteria only, no scoring.
task's content, intent, and context. Never reuse a fixed set of dimensions across tasks.
criteria (what constitutes harmful or misleading content). Positive total points must outweigh negative.
evaluation space implied by the task is fully explored.
The RET builds a 3-level tree from the task:
Task / goal / question
-> Level 1: Scenarios (contextual framings that change what "good" means)
-> Level 2: Perspectives (evaluation dimensions per scenario)
-> Level 3: Criteria (concrete, binary rubric items with scores)At each level, two operators apply:
After the tree is built, Phase B reviews supplied work against the leaf criteria.
Copy this checklist and track progress as you work:
Task Progress:
- [ ] Step 1: Scenario grounding
- [ ] Step 2: Scenario expansion (3 rounds)
- [ ] Step 3: Perspective generation
- [ ] Step 4: Perspective expansion (4 rounds)
- [ ] Step 5: Perspective review and consolidation
- [ ] Step 6: Criteria generation
- [ ] Step 7: Criteria expansion (3 rounds)
- [ ] Step 8: Criteria review and consolidation
- [ ] Step 9: Polarity check
- [ ] Step 10: Score calibration
- [ ] Phase B: Review work against criteria (or emit checklist if no work supplied)Goal: Identify the distinct real-world contexts in which this task could arise, where each context would materially change what constitutes a good result.
Method:
constraints.
situation change what a good result looks like?"
evaluation criteria.
Output format for each scenario:
scenario_name: short descriptive labelscenario_description: 3-5 sentences explaining what is unique about this context andwhy it changes what "good" means
Goal: Ensure comprehensive coverage of the evaluation space at the scenario level.
Method (repeat 3 times):
or rephrase existing scenarios.
After each round, append new scenarios to the list.
Goal: For each scenario, derive the evaluation dimensions that matter for judging a result in that context.
Method:
entirely from the task's content and context -- do not apply a pre-set list of dimensions.
this task.
enough to avoid overlap with other perspectives.
Output format for each perspective:
perspective_name: 2-5 word descriptive labelperspective_description: 3-5 sentences explaining what this perspective evaluates andlisting 3-5 specific sub-aspects it covers
Goal: Fill coverage gaps in the perspective set.
Method (repeat 4 times):
existing ones.
After each round, append new perspectives to the collection.
Goal: Produce a clean, non-redundant set of perspectives ready for criteria generation.
Method:
combine scenario-specific details into each description but keep them as separate perspectives.
concrete criteria generation.
Goal: For each reviewed perspective, generate concrete, binary evaluation criteria.
Method:
result.
Criteria rules:
criteria as detailed and specific as possible.
[Verb + Specific Requirement]. Start with one clear action verb, then state the exact required or forbidden content with qualifiers.
criteria (harmful, misleading, or critically wrong content). Only add negative criteria when the issue represents harmful, dangerous, or significantly quality-reducing behavior -- not minor stylistic concerns.
Negative criteria phrasing: Describe the bad behavior directly. Instead of "Avoids doing X", write "Does X" (where X is the harmful behavior). The criterion text states the behavior; its negative score indicates that meeting it is bad.
Scoring standard:
5-7 = quality enhancer; 1-4 = minor nice-to-have)
-5 to -7 = quality issue; -1 to -4 = minor issue)
Output format for each criterion:
criterion: the criterion textpoints: integer score (positive or negative)reasoning: 2-3 sentences explaining why this criterion matters for this task and whyit received this weight
Goal: Fill coverage gaps in the criteria set.
Method (repeat 3 times):
whether it passes or fails -- and is not yet covered?"
After each round, append new criteria to the collection.
Goal: Produce a concise, non-redundant final rubric.
Method:
precise wording and include all distinct details from merged criteria.
aspect, keep only the positive.
placeholders with a requirement that the result states the current/official/latest value or standard.
all distinct, non-overlapping items.
Goal: Verify that every criterion's score sign correctly reflects whether meeting it is good or bad.
Method:
or indicate a problem (negative)?
score.
"Does X" (where X is bad) describes undesirable behavior (negative).
Goal: Ensure score magnitudes accurately reflect importance.
Method:
important one.
the criterion text.
Goal: Apply the finalized criteria to the actual work, or emit a checklist if none was supplied.
If no work was supplied (checklist mode):
If work was supplied (review mode):
the specific evidence in the work (quote or location) that justifies the verdict.
points of every criterion marked YES (positive criteria add,negative criteria subtract). Report earned points, maximum positive points, and the net total.
ordered by absolute points (most important first). For each, give a one-line concrete fix.
satisfy, and do not penalize for criteria outside the stated task.
Present results in this order:
criterion_id, criterion, points, reasoning.total positive and total negative points.
and the ranked gap list with fixes.
Multi-turn conversations: If the task is a conversation (user-assistant exchange), treat the full conversation as context. The last user message is typically the primary intent; earlier messages provide context that may affect evaluation dimensions.
Tasks with images: If the task includes an image, factor image content into scenario analysis and criteria generation. Reference visual elements where relevant in criteria.
Tasks with retrieved web context: If web-retrieved context is provided alongside the task, use it to inform factual grounding of scenarios and criteria. Do not limit analysis to only the retrieved content -- also apply general reasoning.
Every scenario, perspective, and criterion must be freshly derived from the task.
4 for perspectives, 3 for criteria). Skipping rounds reduces coverage.
for and remove redundant items.
NO. Do not produce criteria that require subjective degree judgments.
evidence from the supplied work.
This skill ports the Qworld method. If you use it in your work, please cite:
@misc{gao2026qworldquestionspecificevaluationcriteria,
title={Qworld: Question-Specific Evaluation Criteria for LLMs},
author={Shanghua Gao and Yuchang Su and Pengwei Sui and Curtis Ginder and Marinka Zitnik},
year={2026},
eprint={2603.23522},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.23522},
}~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.