aaai-experiments — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited aaai-experiments (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Use this before submission to ensure empirical evidence supports the AI contribution. AAAI reviewers may come from adjacent AI subfields, so experiments must be interpretable beyond one benchmark community.
when relevant.
and ethics/IRB status.
risk mitigation, and scope.
Because an AAAI reviewer from an adjacent subfield must trust your numbers quickly, classify each experimental block by how much weight it can bear and what would strengthen it.
| Block | Carries the claim when | Reviewer doubt | Cheap reinforcement |
|---|---|---|---|
| Headline benchmark | beats tuned recent baselines | "lucky seed" | seeds, variance bars |
| Ablation | isolates one mechanism | "joint removal" | single-factor toggles |
| Robustness | holds across split/shift | "one setting" | extra split or perturbation |
| Human eval | protocol is documented | "rater bias" | IRB note, inter-rater agreement |
insight.
A planning paper reports a single-seed win on one domain. Audit: the headline block "needs robustness" and "needs variance", so the fix before the deadline is five seeds with confidence intervals plus one extra IPC-style domain. Because new results cannot rescue this in rebuttal, the team runs both before submission and aligns the checklist's seed answer to the supplement.
[Claim] <paper claim>
[Evidence status] sufficient / needs baseline / needs ablation / needs robustness / unclear
[Fairness issue] <compute, tuning, data, prompt, metric, human eval>
[Checklist dependency] <what checklist answer this supports>
[Fast fix] <experiment or analysis feasible before deadline>~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.