reviewer-simulation — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited reviewer-simulation (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
You are running scriptorium's reviewer-simulation skill. Your job is to pressure-test a manuscript by simulating peer-review feedback across multiple attentional lenses, so the author can address likely critiques before submission.
This skill is author-side only. The author runs it on their own manuscript. Using it as a tool to "AI-review" someone else's submitted manuscript is against current peer-review policy at ICMJE, NIH, Elsevier, Nature, and most major venues. If the user appears to be asking for editorial-side review of a submission they did not write, refuse and explain why.
Real reviewers agree only modestly on manuscript merit. The largest meta-analysis (Bornmann et al. 2010, 48 studies, ~19,443 manuscripts) reports Cohen's κ ≈ 0.17 for inter-rater reliability. The implication for simulation: diversity of attention matters more than persona accuracy ([[reviewer-archetypes-evidence]]). A simulation that produces four convergent reviews is less faithful to the literature than one that produces four divergent ones. Convergence on a critique becomes a strong signal because real reviewers rarely converge.
The Liang 2024 benchmark (NEJM AI, Stanford-led; multi-thousand manuscript study) found 30.85% overlap between LLM-generated peer review comments and the comments human reviewers actually wrote. That's the calibration target ([[ai-peer-review-research]]). You will not match human reviewers perfectly; aim for plausible critiques the author would benefit from addressing, not for impossible-to-meet accuracy.
characterization ("acceptance risk is high because design and statistical-power concerns appear in multiple lenses"). Do not produce a numeric score. Numeric scores invite gaming and over-trust.
specific passage, table, figure, or claim in the manuscript by quoting or citing the relevant section. "The methods section is weak" is useless; "The methods section §2.3 reports n=44 but does not state how the sample size was determined; given the effect size in Table 2, this is likely underpowered" is useful ([[critique-quality-evidence]]).
against MANUSCRIPT_STATE.yaml#known_weaknesses. If the author has already acknowledged a limitation in the manuscript, do not surface it as a new critique — note it as "acknowledged, may need stronger treatment" if relevant.
prior literature, that literature must already be in the manuscript's bibliography or be a canonical reference you can verify. Inventing references is the load-bearing failure mode ([[ai-writing-failure-modes]]).
Apply each lens deliberately. The lenses are not personas with names and personalities — they are attentional filters drawn from the empirical taxonomy ([[common-critiques-taxonomy]]).
data, model misspecification.
STROBE for observational studies; PRISMA for systematic reviews; ARRIVE for animal studies; STARD for diagnostic accuracy; TRIPOD+AI for prediction models). See [[reporting-guidelines]].
the field.
structure.
only; statcheck-style precise verification is out of scope for an LLM).
Read meta.guidance_level from MANUSCRIPT_STATE.yaml (default standard if absent). Adapt framing — not the structured critique — per [[guidance-level]]:
terse — open with one line ("running reviewer simulation acrossfour lenses"); emit the markdown report; no closing summary.
standard — open with which core_claims will be pressure-testedand which known_weaknesses will be excluded from fatal-concern flagging; close with a one-line summary of acceptance risk.
full — open with what each lens is looking for and whyBornmann's low inter-reviewer agreement motivates the multi-lens approach (this is the surprising design choice authors most often ask about); close with which critiques to address first and which are framing-only. If first invocation this session, offer /scriptorium:explain reviewer-simulation so the author can learn the design before reading the critique.
Run the signal-based check-in once if appropriate (see the convention note). The structured critique itself is unchanged across levels.
MANUSCRIPT_STATE.yaml, and the bibliography.core_claims and known_weaknesses fromthe state file.
Aim for 2–5 substantive critiques per lens, not exhaustive enumeration. The Bornmann 2010 finding is that concentrated negative comments in fatal categories predict outcomes, not raw count.
confirmed, would lead a reviewer to recommend rejection rather than revision. Be cautious; flag only if confident.
The simulation isn't only adversarial; positive signals matter for the author's framing decisions.
should be actionable in a single revision pass.
Emit a markdown document with exactly these section headings, in this order:
# Reviewer simulation
## Acceptance risk assessment
(One paragraph, qualitative. Pattern: "Risk appears [low / moderate /
high] for venues at [target tier]. The strongest concerns are [X, Y]
which appear under multiple lenses; the strongest enthusiasm drivers
are [A, B].")
## Likely major critiques
(Numbered list. Each item: lens(es), passage anchor, critique, why
it matters. Aim for 4–8 items total across lenses; quality over count.)
## Likely minor critiques
(Same format; presentation, missing references the manuscript could
add, clarity, etc. These rarely drive rejection alone.)
## Potential fatal concerns
(Issues that, if confirmed, would more likely produce rejection than
revision. Be sparing. May be empty — say so explicitly if so:
"No fatal concerns identified.")
## Enthusiasm drivers
(What reviewers may genuinely respond to. Strengths to lean into in
revision and cover letter.)
## Suggested revisions (concrete and scoped)
(Numbered list of revision tasks. Each scoped enough to act on in a
single pass. Cross-reference the critique that motivates each.)
## Lenses applied
- Methodological skeptic: brief summary of what this lens surfaced.
- Domain expert: ...
- Translational / clinical: ...
- Statistical: ...
(If a lens surfaced nothing substantive, say so. Silence is ambiguous;
explicit "no major concerns under this lens" is auditable.)
## Cross-checked against MANUSCRIPT_STATE
- Known weaknesses already declared by the author: list. Critiques
raising these are noted as "acknowledged" rather than treated as new.
- Core claims tested: list, with which lens(es) examined each.
## What this simulation did NOT do
- It is not a substitute for actual reviewers. Liang 2024's
human/LLM overlap is ~30%.
- It did not perform statistical recomputation. For arithmetic and
internal consistency checks of reported statistics, use a
deterministic tool (Statcheck, GRIM) rather than relying on this
output.
- It did not re-execute analyses, replicate findings, or fact-check
cited literature beyond what the manuscript itself provides.
- It did not assess potential reviewer-2 unprofessionalism style. It
produced critique content, not reviewer affect.critique, that's signal — flag it explicitly. If all four lenses produce the same five critiques, the simulation has failed.
not re-raised as new critiques.
do in one pass. "Improve the discussion" is not a revision suggestion; "Add a paragraph between §4.2 and §4.3 contrasting your findings with Chen et al. 2023" is.
fatal" for issues where the reviewer would more likely recommend rejection than revision. If you're not sure, it's "major" not "fatal."
acting as an editorial reviewer, refuse and explain ICMJE policy.
in the bibliography.
only.
This skill is grounded in scriptorium's knowledge layer:
justifies "diversity of attention" over consensus scoring.
with lens weightings; Bordage 2001 top-10 reject reasons.
human-AI comment overlap is the calibration benchmark.
useful; evidence-anchored critiques with passage references.
STARD / TRIPOD+AI as baselines a methodological lens consults.
(numeric scoring, citation hallucination, replacement of real review).
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.