scour — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited scour (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
A long conversation builds a closed room. The agent produces a confident, internally consistent narrative — "SQuAD uses exact-match, LLM-judge gets ~97% human agreement, this layered fix is the standard hardening" — and it sounds completely right. That is the danger: it was generated from memory, inside the room, and nobody opened the window to check it against the actual world. Scour is the hotline button: leave the room, verify we are on the right page, return with the truth.
It is the evidence twin of isomorph. Both answer "are we shaped right?" — isomorph from a proven domain's structure (reasoning, the kingfisher), scour from real-world evidence (how people actually, currently do this). And it shares the family spine: devour grounds you in the codebase, bedrock grounds the code in reality by running it, scour grounds the understanding in reality by fetching it.
No claim survives without a real fetched source. Scour exists because the user does not trust confident-from-memory answers. So if scour itself answered from memory, it would be the very disease it is the cure for, wearing a badge. Every verdict traces to something pulled off the web in this run — a doc, a repo, a thread, a paper, a changelog. If the web cannot be reached or the sources are thin, scour says so plainly and leaves the claim unverified — it never dresses up training-data recall as fresh research.
Use this skill when the user asks for any of these:
grounded before trusting it.
Do not use it for: a deliberate topic deep-dive that should produce a kept artifact (that is research-report / scan), a single fact lookup (just use web search directly), finding a structural analogy by reasoning (that is isomorph), or testing whether the code holds up (that is bedrock). Scour checks understanding, fast, in-flow.
Scour reads what just happened in the conversation and does the right thing. Both faces are "leave the room, check the world, come back."
We just concluded something confidently and the user wants it grounded.
decision actually rests on — "SQuAD uses exact-match", "LLM-judge ≈ 97% human agreement", "this layered approach is the standard hardening". Not the trivia; the claims that, if wrong, change what we do next.
against the web: (1)… (2)… (3)…" — so the user can redirect before you spend the reach.
is wrong, outdated, or oversimplified. Rubber-stamping ("yep, confirmed!") is the failure mode — scour attacks the narrative the way bedrock attacks the foundation.
outdated / no-consensus — each with the source that settles it.
We don't know yet, or we're stuck.
real disagreement, the gotcha everyone hits.
listicle with no tie back to the live problem is the failure mode discover refuses.
The most useful thing scour catches is "this was true when the model trained, but the world moved." A model's knowledge has a cutoff; the web does not. Always check whether a confident claim is current — a deprecated API, a benchmark that's been superseded, a number that's been revised, a "best practice" the field has since abandoned. Date the finding when currency matters: "true as of the 2023 paper; the 2025 revision changed it."
This skill must work in Codex, Claude Code, and any other coding-agent host. Use whatever web and documentation tools the host exposes — web search, page fetch, and any connected docs/forum/repo sources — and prefer primary sources (official docs, changelogs, repos, papers, maintainer threads) over aggregators. Never depend on a host-specific tool by name; if the host has no web reach at all, say so and decline rather than answering from memory.
Scour's output is momentum, not an artifact. It returns into the conversation and dissolves, like isomorph and potential — no files. Two handoffs, offered never forced:
a research-report (which owns persistence and the notes.md + report.html convergence).
— "this might be worth a bedrock finding" — but scour itself never writes files and never edits code.
Keep it tight — this is a hotline, not a report. Lead with the verdict.
Scoured: <one line — what was checked>
Verify:
- <claim> → CONFIRMED · <source>
- <claim> → OUTDATED · was true <when>, now <what changed> · <source>
- <claim> → OVERSIMPLIFIED · the nuance we missed · <source>
- <claim> → NO-CONSENSUS · the camps · <sources>
Bottom line: <are we on the right page? what, if anything, to change.>
Unchecked: <what couldn't be reached / what I didn't verify>For discover, swap the claim list for: the dominant pattern, the camps, the gotcha, and the anchored recommendation.
A run is good only if:
held. No source, no verdict; it's marked unverified instead.
we're wrong or outdated is visible.
checked.
If the web couldn't be reached, the honest output is "I couldn't ground this" — not a confident verdict. That failure-to-reach is itself the most important thing to say.
failure.
the claim.
research-report. Scour is the quick reachthat keeps the conversation moving.
and current, say so plainly.
bedrock instead.
is for.
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.