aer-referee-sim — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited aer-referee-sim (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Most papers submitted to AER are rejected; the realistic acceptance rate is 6-8 percent, and a large share never reach referees. The cheapest referee report is the one generated before submission — but only if it is as harsh as the real one. The failure mode of self-review (human or AI) is leniency: reviewing the paper one hopes was written instead of the one on the page.
This skill runs the AER editorial process against the draft: a ten-minute desk screen, then three referee reports written from distinct, adversarial priors, then an editor's synthesis with a calibrated verdict and a prioritized revise list. The simulation has one rule that overrides all others:
The simulated reviewers' job is to reject the paper. Every comment must survive the question "would this withstand the authors' best rebuttal?" — but praise requires the same evidence as criticism.
aer-consistency reports all-pass
same reports
Do not use on a half-draft — the simulation will correctly report that the paper is incomplete, which wastes the run. And do not let it replace aer-consistency: typo-hunting referees are wasted referees.
Simulate the editor's first pass: ten minutes, first three pages, then the main tables, then the bibliography. The editor is deciding only one thing — is this worth three referees' time?
Work through docs/desk-rejection-audit.md items 1-5 plus three scans:
sentence after page 3? Would an economist outside the subfield care?
it a modern design (aer-identification red flags apply on sight)?
free of the failure patterns in docs/style-guide.md. Editors read craft as a proxy for care in the empirics.
Output a desk decision with the editor's two-paragraph letter:
DESK DECISION: <reject | send to referees>
LETTER: <the letter an AER editor would actually send>Calibration: if any Stage 1-2 item in the desk-rejection audit fails, the decision is reject — write the letter and stop. Do not soften a desk reject into "borderline" to keep the simulation going; fix the draft and rerun.
Three referees, three priors, three reading orders. Each writes independently — draft all three before reconciling anything, and never let R2 inherit R1's findings.
Reads: Empirical Strategy first, then Data, then the robustness appendix. Prior: "the design is broken until proven otherwise."
Attacks: the identifying assumption's plausibility in this setting; missing diagnostics from the aer-identification battery; inference mismatched to the variation's level; estimand-population gaps (whose effect is this?); the alternative story the design cannot exclude. R1 re-derives at least one magnitude from the tables and checks it against the prose.
Reads: Introduction, then the antecedents, then Results against the literature. Prior: "we probably already knew this."
Attacks: novelty against the working-paper frontier (names the closest papers, including any the draft missed — aer-literature's map is the checklist); whether magnitudes are plausible next to the literature's; whether the mechanism evidence distinguishes the favored channel from the obvious rival; institutional errors a field insider would catch. R2 is the referee most likely to have written one of the antecedents.
Reads: linearly, as an editor-board member from another subfield. Prior: "why should I care, and can I follow it?"
Attacks: cross-subfield interest (the explicit AER bar); whether the first three pages are self-contained; under-interpreted results (coefficients never converted to economic meaning — aer-paper-body rules); exhibit overload or disorder; the conclusion overreaching the evidence; external validity left unaddressed.
SUMMARY: <2-3 sentences — the paper as the referee understood it>
MAJOR COMMENTS: <numbered; each one: quote or cite the page/table,
state the problem, state what evidence would resolve it>
MINOR COMMENTS: <numbered, brief>
RECOMMENDATION: <reject | major revision | minor revision | accept>Rules of engagement:
the exact table/figure. Unanchored vibes ("the paper feels thin") are banned.
or rewrite would satisfy the referee. Comments with no resolution path are editor material, not referee material.
certify, against their own checklist, why fewer exist. An AI reviewer that finds nothing major has defaulted to agreeable — restart that report with the prior dialed up.
the cap.
Score the paper on the rubric in docs/referee-report-rubric.md (contribution, identification, data, robustness, magnitudes, exposition, integrity — each 0-5 with anchored definitions), then issue the decision the reports support:
RUBRIC SCORES: <dimension: score, ...>
VERDICT: <desk reject | reject after review | major R&R | minor R&R>
DECISION LETTER: <editor's letter, naming the comments that drove it>
REVISE LIST: <every major comment, deduplicated, ordered by severity:
blocking → major → minor, each tagged with the skill that fixes it>Calibration anchors (do not inflate):
no robustness round fixes a broken design.
below 2. This is already a top-decile outcome for real submissions.
minor R&R, suspect leniency and rerun Stage 2 with the priors sharpened.
aer-consistency (all PASS)
→ aer-referee-sim
→ verdict reject? → route fixes:
identification comments → aer-identification / aer-robustness
novelty comments → aer-literature / aer-topic-selection
interpretation comments → aer-paper-body
framing comments → aer-introduction
exhibit comments → aer-tables-figures
→ revise → aer-consistency → aer-referee-sim (fresh reports)
→ verdict ≥ major R&R on a fresh run → aer-submissionRerun with fresh reports each time — re-grading old comments measures compliance, not quality. Two consecutive runs at major-R&R-or-better, with no blocking comments, is the exit condition.
too long to hold at once, review it section by section against each referee's checklist — never from recall.
contribution) must be verified before they enter a report; a simulated referee who hallucinates a flaw costs a revision round.
rejected the draft for X" is the deliverable, not a diplomatic summary.
the attacks this skill knows how to mount. Say so in the output.
reading orders exist to prevent this
"major" pads the count without testing the paper)
draft — variance is not improvement
than a lower bound on preparedness
When working from the AER-skills repository or plugin bundle, load only the relevant resource:
report: docs/referee-report-rubric.md
list): examples/referee-report-example.md
docs/desk-rejection-audit.mdskills/aer-identification/SKILL.mdand docs/methods-reference.md
skills/aer-robustness/SKILL.md
docs/style-guide.mdskills/aer-rebuttal/SKILL.md
DESK DECISION: <reject | sent to referees>
REFEREE RECOMMENDATIONS: <R1 / R2 / R3>
RUBRIC SCORES: <list>
VERDICT: <desk reject | reject | major R&R | minor R&R>
BLOCKING COMMENTS: <n — list>
REVISE LIST: <comment → skill routing>
NEXT SKILL: <routed fix skill | aer-submission if exit condition met>with different stakes
catches different failures than referees do
the revise list
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.