iclr-experiments — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited iclr-experiments (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Use this before submission or during a revision pass to stress-test empirical claims. ICLR experiments should answer the scientific question, not merely assemble a leaderboard.
OpenReview/arXiv papers.
sensitivity, or compute scale when those affect the claim.
where relevant.
or ethics status when needed.
ICLR's empirical culture prizes honest ablations and mechanism over leaderboard position. A clean ablation that explains why a representation works often outscores a larger raw number.
| Claim type | Evidence that convinces ICLR reviewers | Common reject trigger |
|---|---|---|
| New objective helps | Ablate the objective with everything else fixed | Gains confounded with extra tuning |
| Method scales | Several model sizes/tasks with a trend | One large run, no scaling curve |
| Robust representation | Tests across shifts, seeds, prompts | Single-seed peak on one benchmark |
| Beats prior method | Tuned, current, open-source baseline | Stale or under-tuned baseline |
A paper claims a new self-supervised pretext task yields better linear-probe accuracy. Reviewers ask whether the gain is the pretext task or simply longer pretraining. The author audit: hold total pretraining compute fixed, swap only the pretext objective, and report linear-probe accuracy with error bars over five seeds. The compute-matched ablation isolates the mechanism and is small enough to post inline during discussion, where the table becomes part of the permanent public record.
[Claim] <paper claim>
[Experiment evidence] sufficient / needs baseline / needs ablation / needs robustness
[Fairness issue] <compute, tuning, data, prompt, metric>
[Fast fix] <experiment or analysis feasible before deadline>
[Appendix placement] <what can move out of main text>~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.