ijcai-experiments — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited ijcai-experiments (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Use this before submission when the experimental story is not yet locked. IJCAI reviewers can score novelty, correctness, clarity, significance, impact, presentation, ethics, and reproducibility.
reviewers ask.
selection criteria, random seeds or repeats, and compute infrastructure.
differences could change the conclusion.
fairness, misuse, and deployment limits.
the supplementary material.
IJCAI draws reviewers from symbolic AI, search, planning, constraint satisfaction, KR, multi-agent systems, game theory, ML, NLP, and vision, so the experimental section must read across subcommunities. Calibrate evidence to the claim type rather than copying an ML-only template.
| Contribution type | Decisive evidence | Common reject trigger |
|---|---|---|
| Search / planning | Coverage, anytime quality, expansion counts, time/memory cutoffs, per-domain breakdown | Single suite, no domain table, missing strong planner baseline |
| Constraint / SAT | Cactus plots, instances within timeout, solver versions | No virtual-best comparison |
| Multi-agent / game theory | Welfare/equilibrium metrics, agent-count scaling, seeds | Claims hold at one population size only |
| Learning method | Strong current baselines, core-mechanism ablations, variance | Cherry-picked seeds, weak baselines |
| Theory-plus-experiment | Experiments confirming the proven bound | Empirics outside the theorem's regime |
A submission proposes a learned heuristic for cost-optimal classical planning and reports a single aggregate "12% fewer expansions" number. Apply the decision rules:
table, and add a strong admissible-heuristic baseline that shows optimality is preserved.
credited to engineering, and report multiple training seeds since the heuristic is stochastic.
meaningless without a stated timeout.
cross-section reviewer distrusts single-suite claims.
comparisons.
response, since no new results may be added later.
[Experiment readiness] strong / adequate / weak
[Claim -> evidence map] <claim: section/table/figure>
[Missing baseline or ablation] <item>
[Reproducibility gaps] <hyperparameters/seeds/compute/data/code>
[Decision-critical next run] <one experiment>~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.