radar-self-eval — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited radar-self-eval (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
The radar must measure whether it is winning, not assume it. Three parts: metrics (every week), retrospective (monthly), curator proposals (every week).
Compute from the week's daily reports (primary source — they list ledger changes) and, where needed, git log -p -- TRENDS.md since the previous weekly commit:
days and still unverified (stale)
logs/source_rotation.md linecontains a venue-exploration entry ÷ daily runs executed
strategy_notes (judgment call — name them)
its evidence-line date and the date it entered the ledger (commit date via git log -p -- TRENDS.md, or the daily reports). Report the median, split by channel — exploration finds vs queue promotions (backfill) — plus the backfill share of all new evidence. This is the daily-ness KPI.
for each "swept every run" list in SOURCES.md (lab blogs, YouTube curators, pointer/digest blogs, discovery venues), diff the list against the week's logs/source_rotation.md lines and classify every listed source as opened, degraded, or MISSING (in SOURCES.md but never in any log line this week). MISSING = a coverage lie — NAME them. A source MISSING or degraded for the whole week is a heal-or-REMOVE candidate (see Amendments): the registry must be honest about what it actually sweeps.
axis yet absent from that trend's evidence — on-axis primaries hoarded in the queue instead of routed to evidence (the UltraQuant failure class). Count and name them; each is a routing miss to correct this week.
Append ONE dated line to logs/calibration.md (the externalized self-eval log; the ## calibration section of TRENDS.md is now only a pointer): - YYYY-MM-DD — W<nn>: queue +a/→p/−d/stale s · evidence +e · moves m · exploration c/r · off-axis o/a · lag expl Xd / backfill Yd (Z%) · coverage <opened>/<listed> (miss <n>, degr <d>) · routing-leak <n> and include the same numbers, readable, in the weekly report.
Interpretation thresholds (act, don't just log):
daily prompt/skill refinement.
propose narrowing or a verification-only day.
tunnel-vision check passed.
queue is cleared) → the radar is doing literature review, not daily detection: propose rebalancing scan time toward exploration.
date +%d ≤ 7)Goal: ground truth — did the radar see early what later became big?
(open the pages — evidence rules apply): HF papers trending, two major lab blog indexes, one arXiv listing. Pick what is everywhere, not what is interesting.
git log -S"<term>" --oneline -- TRENDS.md plus the queue and reports.
HIT-late, MISS.
observation_queue now, and name the venue oraxis that would have caught it — feed that into the proposals.
Log in logs/calibration.md: - YYYY-MM-DD — retro M<mm>: <item> — HIT-early|HIT-late|MISS (first seen <date or never>), …
The radar runs unattended and may change its own operating instructions — under the autonomy contract in AGENTS.md. Every week, up to 3 proposed amendments: routine edits (exact replacement text), axis drops/merges/ supersessions, exploration-budget changes. Each must cite the metric or retrospective result that motivates it.
The new coverage/routing metrics feed amendments directly (this is how the radar self-corrects the failure classes the curator used to catch by hand):
radar-source-healattempt; if it was already healed and still fails, a heal-or-REMOVE amendment (remove it from the SOURCES.md "swept every run" list, or escalate the secret/allowlist it needs) — a first-class amendment, because a listed-but- unswept source is a coverage lie.
hoarded on-axis queue items to their trend's evidence now; if it recurs, amend the routing instruction.
Lifecycle:
calibration and the weekly report.veto appeared in strategy_notes (silence is consent). One dedicated commit per amendment. If the signal vanished, drop the proposal and say so.
git revert the amendment and log the rollback.
Scope-axis changes are recorded as "radar-adopted" dated entries in strategy_notes, explicitly naming what they supersede. Curator entries are never deleted or edited — vetoes and mission input belong to the curator alone. The immutable sections listed in AGENTS.md are out of bounds: an amendment touching them is invalid.
logs/calibration.md is append-only: never edit or delete existing lines.refinement per the maintenance policy in AGENTS.md (dedicated commit, explained in the report).
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.