deep-research — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited deep-research (Agent Skill) and scored it 92/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 2 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 2 flagged
The text {match} tells the agent to skip the normal "ask the user first" gate. Used adversarially it removes the human-in-the-loop check before destructive or sensitive actions, turning a normally-gated agent into a fire-and-forget executor.
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Deep Research is a Codex-compatible AI agent skill for source-backed research, evidence synthesis, literature reviews, fact-checking, market research, policy research, technical research, academic research, and PRISMA-style systematic review planning.
Use this skill to run disciplined research, not generic web browsing. Prefer it when accuracy depends on source quality, citation traceability, contradiction handling, or an explicit research process.
Use Codex tools directly. Do not assume external pipeline hooks, slash commands, passport files, or skill-specific scripts exist unless they are present in the current environment.
socratic, fact-check, quick, lit-review, full, or systematic-review.lite, standard, or thorough.standard/thorough, uses subagents, or will materially consume time.socratic: User has a vague idea, asks to be guided, or is unsure what to research. Ask questions; do not jump to a report.fact-check: User provides specific claims to verify. Output claim-by-claim verdicts.quick: User needs a concise source-backed brief. Use fewer sources, but still verify them.lit-review: User needs literature mapping, annotated bibliography, themes, gaps, or source matrix.full: User has a clear question and wants a comprehensive report.systematic-review: User asks for PRISMA, systematic review, meta-analysis, formal screening, risk of bias, or GRADE-style evidence synthesis.For details, read references/modes.md.
Choose the smallest tier that can answer the question honestly:
lite: fast answer, narrow fact-check, quick brief, or low-stakes overview. Target 3-6 strong sources and verify central claims against at least one source route.standard: default for meaningful research, comparisons, policy/market questions, and literature reviews. Target 6-12 sources and use at least two independent source routes for central claims.thorough: use for high-stakes, systematic-review-like, disputed, technical, legal/medical/financial, or decision-critical research. Target 12+ sources where available, use multiple source families, actively search for counterevidence, and verify central claims independently.Escalate the tier when the user asks for more rigor, the evidence is contradictory, the topic is high-stakes, or early searching shows weak/fragmented sources. Downgrade only with an explicit note about what rigor is being traded away.
Produce a brief research frame:
Research question:
Scope:
Out of scope:
Key definitions:
Assumptions:
Mode:
Tier:
Planned source types:For vague topics, switch to socratic and ask 1-2 focused questions at a time until the user can state what they want to know.
Before collecting sources, present a compact plan:
Plan:
Mode:
Tier:
Sub-questions:
Source routes:
Verification plan:
Estimated effort:Ask for approval before continuing unless all are true: tier is lite, no subagents are planned, source collection is small, and the user already clearly asked for immediate execution. If the user says to adjust, revise the plan once and ask again. If the user says to proceed, execute without relitigating.
Build a reproducible search plan:
Databases/sources:
Search strings:
Date range:
Languages:
Inclusion criteria:
Exclusion criteria:Use at least two independent source routes for non-trivial research, such as official data plus scholarly literature, or primary documentation plus reputable reporting.
Tier defaults:
lite: 1-2 source routes, targeted searches, stop when the answer is adequately supported.standard: 2-3 source routes, explicit counterevidence search, source matrix for central claims.thorough: 3+ source routes, broader keyword variants, formal exclusion notes, deeper contradiction handling.Before synthesis, create a compact evidence table:
Source:
Type:
Verified by:
Evidence grade:
Relevant claim(s):
Limitations/conflicts:
Use / caveat / reject:Read references/source-verification.md when source quality, citations, DOI checks, policy claims, or academic evidence grading matter.
Tier defaults:
lite: verify every central claim used in the bottom line.standard: verify central claims and the strongest counterclaim.thorough: verify central claims, counterclaims, key numbers/dates, and citation metadata.Group findings by answer-relevant themes, not by source. Explicitly handle:
Before finalizing, run a Devil’s Advocate pass:
Strongest counterargument:
Weakest supported claim:
Likely bias in source selection:
Evidence overreach:
What would change the conclusion:If a critical flaw invalidates the conclusion, stop and explain the flaw instead of polishing a bad report.
Match the output to the mode:
fact-check: claim table with verdicts: supported, unsupported, mixed, unverifiable, or misleading.quick: concise brief with key findings, source notes, caveats, and next research steps.lit-review: themes, annotated bibliography, gaps, and source matrix.full: executive summary, method, findings, discussion, limitations, references.systematic-review: protocol, search strategy, screening counts, study table, risk-of-bias plan/results, synthesis approach, limitations.socratic: research plan summary only when the user asks to summarize or move forward.For standard and thorough, include a short method note with mode, tier, source routes, and verification limits.
When subagents are available and the task is large enough, split work like this:
Do not let subagents write the final synthesis independently. Their outputs are evidence inputs, not conclusions.
Read references/failure-paths.md when a workflow stalls, sources are weak, the question is too broad, evidence contradicts the conclusion, or verification fails.
Core defaults:
references/modes.md: mode definitions, output shapes, and transition rules.references/source-verification.md: source grading, citation verification, evidence hierarchy, and predatory-source checks.references/failure-paths.md: recovery rules for common research breakdowns.~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.