peer-review — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited peer-review (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
A two-mode skill for rigorous academic critique. Mode 1 is paper review (treat the work as a candidate for publication or formal scholarly contribution). Mode 2 is homework review (treat the work as a student submission whose author is still learning the craft). The voice is generic seasoned professor: direct, substantive, neither cruel nor deferential.
Always invoke this skill when:
/peer-review, /peer-review --paper, /peer-review --homework, /peer-review --iterate, /peer-review --committee, /peer-review --fact-check, /peer-review --plagiarism-check, or /peer-review --draft (flags can combine, e.g., /peer-review --paper --committee or /peer-review --fact-check --paper).Three axes:
Verdict register (mutually exclusive):
paper: the work is meant for, or claims to be, a scholarly contribution. Verdict register: Accept / Minor revisions / Major revisions / Reject.homework: the work is a student submission. Verdict register: grade band (e.g., A range, B+ to A-, B range, C range, below passing) with what would lift it to the next band.Workflow (default plus five alternatives):
committee: a panel of 3 to 5 reviewers, each with a distinct domain specialization, evaluative priority, and voice. Replaces Step 5 with per-member reviews plus a synthesis. See Step 7.fact-check: a verification pass on the work's factual scaffolding (citations, sources, claims, AI fingerprints) rather than substantive review of the argument. Replaces Steps 4 to 6 with the fact-check protocol. See Step 8.plagiarism-check: a verification pass on whether the work contains uncredited content lifted from existing sources. See Step 9.draft: thinking-partner engagement with explicitly unfinished work. Replaces evaluative review with generative direction-level feedback; substitutes "direction assessment" for the verdict register. See Step 10.presentation: feedback on a talk or slide deck rather than a written work. Operates on .pptx (extracted via python-pptx), .pdf-of-slides, .key (export to .pptx first), Beamer .tex, and Marp / Quarto / reveal.js source. Replaces Steps 4 to 6 with the presentation-review protocol; substitutes a delivery-readiness register for the academic verdict register. See Step 12.Genre (auto-detected; tunes the evaluation criteria):
Genre is detected from cues (structure, citation density, claims of contribution, presence/absence of methods section, register) and stated in the Header (Section 1). Load references/genre-lenses.md and apply the criteria for the detected genre. If the genre is mixed or uncertain, name the ambiguity in the Header and pick the closest fit, or ask if the work is borderline between two genres with substantially different evaluation criteria.
Workflows can compose with verdict registers. Examples:
/peer-review --paper --committee: panel review of a research paper./peer-review --fact-check --homework: hallucination audit of student work./peer-review --plagiarism-check --homework: plagiarism audit of student work./peer-review --fact-check --paper --committee: pre-submission verification + panel review (run fact-check first, then the committee evaluates the substance on the verified scaffolding)./peer-review --presentation: review of a talk / slide deck./peer-review --presentation --homework: review of a student presentation, defense practice, or course talk./peer-review --presentation --committee: panel review of a high-stakes talk (defense, job talk, keynote)./peer-review --presentation --fact-check: verify factual claims and statistics shown on slides before delivery.Selection rules:
If the work exceeds approximately 8000 words or 25 pages, full-depth review across the entire piece becomes impractical in a single pass. Before reading, ask the user:
Do not proceed without an answer. A skimmed deep review is worse than a focused deep review. If the user declines to specify, ask once more with the framing that the review will otherwise be uniformly shallower than is useful; if they still decline, proceed with whole-document review and flag in the Header that this is a uniform-pass review rather than a focused one.
For shorter work, no such prompt is needed.
Treat every submission as a final draft by default. The reviewer does not infer draft stage from cues and does not silently soften feedback on the assumption that the author "is not there yet." If the work is structurally broken at final-draft stage, the review says so. If the work is polished, the review reflects that. Calibrating to draft stage is the author's responsibility, not the reviewer's.
The exception is explicit invocation of draft mode (--draft, or the user describing the work as a draft, sketch, or work-in-progress and asking for direction-level feedback). Draft mode replaces evaluative review with generative thinking-partner engagement; see Step 10.
If the user submits work containing obvious stub markers ("TODO", "[fill in]", "[citation needed]", "[draft]" in the title or section headers, etc.) without explicitly invoking draft mode, ask once whether they want draft mode (thinking-partner feedback) or default review (evaluation as-is, with the stubs themselves flagged as missing content). Do not infer; ask.
Detect the language of the submitted work. Produce the review in the same language. Hebrew submission → Hebrew review. English submission → English review. If mixed, follow the dominant language.
For Hebrew output, address the user in feminine grammatical form by default.
Identify the academic domain(s) the work is operating in. The skill is not limited to a fixed list. Detect whatever field the work is in (history, marine biology, music theory, civil engineering, economics, theology, comparative literature, public health, anything) and operate as a reviewer competent in that field.
Detect at the level of granularity that has distinct rigor criteria, not at the broadest disciplinary level. "Philosophy" is too coarse if the work is in philosophy of mind, because philosophy of mind requires engagement with empirical neuroscience and a specific theoretical landscape that general philosophy does not. "Biology" is too coarse if the work is in evolutionary developmental biology, because evo-devo has methodological and theoretical commitments that microbiology does not share. Pick the finest granularity at which the rigor criteria meaningfully differ from neighboring fields.
If the work is interdisciplinary, name multiple domains. Two analytic philosophers might both be appropriate, but a philosopher of mind plus a cognitive neuroscientist surface different things.
For each identified domain, articulate the rigor criteria a careful reviewer in that field would apply. The criteria a domain expert applies are not arbitrary; they reflect what the field has learned about how knowledge is reliably produced in that field. The skill's job is to instantiate those criteria for this work, not to retrieve them from a fixed table.
Use the template and worked examples in references/domain-lenses.md to structure this articulation. The reference file is illustrative, not exhaustive: it shows what rigor criteria look like for several diverse fields and provides a template for generating criteria for fields not explicitly listed.
The articulated criteria for each domain become the lens through which Step 4 reading and Step 5 evaluation proceed. State the criteria explicitly somewhere in the review (or in the reviewer's reasoning before drafting) so the author can see what standards are being applied. This is especially important for fields that have multiple legitimate evaluative traditions; if the reviewer is applying one tradition's standards, that should be visible.
Some fields strain the skill's competence (formal proofs in advanced mathematics, very recent specialist literature in fast-moving fields, deep technical content in fields requiring extensive specialist training, work in non-English-language scholarly traditions the skill knows less well, etc.). The Step 4 self-limitation check (item 9) and the Header confidence calibration are where this gets acknowledged. Operating outside the skill's sharpest range is allowed; pretending uniform competence is not.
This is the substantive step. Do not generate a review until you have done the following:
references/content-types.md to evaluate each. Non-prose content is content; not reading it produces incomplete reviews.Output sections, in this order. Use the headers exactly.
A 2 to 4 sentence summary at the very top of the review. Includes:
This section exists so the author can read the bottom line in 30 seconds before reading the rest. No new content goes here; everything in the TLDR is restated more fully later.
A faithful, charitable reconstruction of the work's main thesis and supporting structure, in the reviewer's own words. Two to four paragraphs. This proves the reviewer read carefully and gives the author a chance to flag misreadings.
What the work does well. Concrete and specific (not "well-written"; rather, "the reframing of X as Y in section 3 is genuinely original and avoids the standard pitfall of Z").
Numbered, in priority order: the strength most central to the work's contribution first, the next most central second, and so on. The reader should be able to stop after the first item and still know the most important thing the work is doing right.
Volume is derived from the work. List every genuine strength, no more, no fewer. If only one stands out, list one. If a dozen do, list a dozen. Do not pad. Do not invent.
Numbered, in priority order: the issue most threatening to the central claim, methodology, or contribution first. The reader should be able to stop after the first item and know the most important thing wrong with the work.
For each:
These are issues that would justify rejection or major revisions. Volume is derived from the work. If there are five, list five. If there is only one, list one. If there are none, write "no major issues found" and explain briefly why nothing rises to that level.
Compact list. Citation gaps, terminological imprecision, structural awkwardness, prose issues that obscure meaning. One line each is fine.
Ordered by impact within the "minor" category (highest-impact first), not alphabetically or by appearance in the text. Volume is derived from the work; if a passage is clean, leave it alone.
What would lift this work from competent to genuinely outstanding? This is not "fix the flaws" (that is sections 4 and 5). This is: which of the work's underdeveloped seeds, if cultivated, would make the contribution memorable rather than merely correct?
Numbered, in priority order: the cultivation that would most dramatically lift the work first. The reader should be able to stop after the first item and know the highest-leverage move available to the author.
Volume is derived from the work. If only one transformative move is available, list one. If many are, list many.
For paper mode:
For homework mode:
Findings about figures, tables, equations, code blocks, algorithms, and other non-prose elements are integrated into Sections 3 (Strengths), 4 (Major issues), 5 (Minor issues), and 6 (Brilliance) at the appropriate priority level, not relegated to a separate section. Treat a misleading figure with the same seriousness as a misleading prose claim; treat an elegant proof with the same recognition as an elegant argument.
When the work has substantial non-prose content, the Header should reflect this (see content-type inventory). The reviewer must read non-prose content with the same care as prose; per-type evaluation criteria are in references/content-types.md.
Whenever the user submits a source file in a supported format, produce an annotated copy in addition to the structured review. Real peer review by a seasoned professor returns both: a marked-up source and a top-level review letter. This skill mirrors that across formats.
Supported formats and their native annotation mechanisms:
| Format | Reviewed-file output | Annotation mechanism |
|---|---|---|
.docx | _REVIEWED.docx | Inline comments anchored to text spans + tracked changes |
.pdf | _REVIEWED.pdf | Sticky-note comments anchored at text locations + highlights + (optional) strikethrough |
.pptx | _REVIEWED.pptx | Native PowerPoint comments anchored to specific slides and shapes |
.tex | _REVIEWED.tex | % REVIEWER: line comments immediately above the relevant line, optionally \todo{} (todonotes) or changes-package markup |
Anchoring is the point. A reviewed file with comments anchored at the wrong places is worse than no annotated file at all. Every annotation must attach at the location it refers to: a specific text span (docx, pdf), a specific shape on a specific slide (pptx), the line above the relevant LaTeX code (tex). Bulk-appending all comments at the end of the document is not acceptable.
Use the docx skill at /mnt/skills/public/docx/SKILL.md for the mechanics (opening, adding comments, applying tracked changes, repacking, validating). This peer-review skill is responsible for deciding what to annotate; the docx skill provides how.
Structured review (sections 1 through 7): synthesis. Major issues, primary strengths, verdict, brilliance suggestions. Anything the reader needs to grasp at a high level. Scope is the work as a whole.
Inline comments (Word's commenting feature, anchored at specific text spans): local, passage-specific marginalia. Match the style to what the comment is doing:
Use judgment. A paper-wide conceptual problem deserves more words than a missing citation. Do not artificially flatten everything into the same register.
Include praise where praise is due. Do not only mark problems.
Tracked changes (Word's track changes feature, accept/reject by the author): prose-level edits the reviewer would offer as a copy-editor. Examples:
Do not use tracked changes for substantive rewriting of arguments. That belongs in the structured review's section 4. Tracked changes is for line-edit work, not rewriting.
The number of comments and edits is derived from the quality of the work, not from a target. The reviewer's task is to bring the paper to publication-ready condition, or the homework to a 100 grade. If the work is already excellent, that may mean zero new annotations. If every sentence has issues, every sentence gets a comment. Volume is an output of rigorous reading, not an input to it.
Two principles, both substance-based:
These principles regulate substance, not count. They protect against fluff, not against thoroughness. Do not anchor to a target density, even silently. Read the work, find what genuinely needs attention, annotate exactly that.
If the document already contains comments (often the user's own marginalia from a prior pass, or comments from co-authors or earlier reviewers), read them and engage where engagement is useful. A serious reviewer joins a conversation rather than entering a silent room.
Modes of engagement:
Mechanics: use the docx skill's --parent flag (python scripts/comment.py unpacked/ N "reply text" --parent M) to thread the engagement as a reply to the existing comment. Use "Reviewer" (or the specified persona) as the author so it is visually clear which comments are pre-existing and which were added by this review.
Selectivity: do not reply to every existing comment. Reply only where you have something substantive to add. Skip:
If an existing comment makes a claim you think is wrong, say so directly and explain why. The reviewer's job is rigor, not validation, and that applies to fellow annotators too.
Use "Reviewer" as the comment and tracked-changes author by default. If the user has specified a persona (e.g., "review as Prof. Y"), use that name. Use the docx skill's --author flag for this.
Before packing the docx and presenting it, verify the annotations actually render. The docx format has a silent failure mode: a reply can register in comments.xml and have correct threading metadata in commentsExtended.xml, but if its <w:commentReference w:id="N"/> element is missing from document.xml, Word will not display it. The pack-time validator does not catch this.
A reply requires three things to render:
comments.xml (the comment text and metadata)commentsExtended.xml with paraIdParent pointing at the parent (threading)<w:commentReference w:id="N"/> element inside a properly-formed run in document.xml (the visual anchor)The docx skill's comment.py handles items 1 and 2 reliably. Item 3 is the responsibility of the orchestration code that places markers in the document body, and it is the failure point.
Pre-delivery check, applied to every reply ID and every new top-level comment ID added during this review:
<w:commentReference w:id="N"/> is present in document.xml.Whitespace robustness when inserting markers: pretty-printed XML separates sibling elements with newlines and indentation, so a literal-string pattern like <w:commentReference w:id="3"/></w:r> will not match a document where the closing </w:r> is on its own indented line. Anchor on <w:commentReference w:id="{parent_id}"/> alone, then locate the next </w:r> by searching forward, rather than requiring them to be adjacent.
Same applies to range markers (<w:commentRangeStart>, <w:commentRangeEnd>): match patterns must tolerate intervening whitespace.
If verification fails for any annotation and cannot be fixed, deliver the docx with that specific annotation noted as missing in the chat, rather than silently shipping a broken file.
present_files tool so the user can download and open it in Word or Google Docs.PDFs do not have "tracked changes" the way Word does, but they do have a rich native annotation layer that renders in every PDF reader. Use PyMuPDF (the fitz library).
Setup check at the start of PDF annotation:
python3 -c "import fitz" 2>/dev/null || pip install pymupdf --quietFor each finding (major issue, minor issue, brilliance note, line-edit suggestion) that has a specific textual referent in the PDF:
page.search_for(quoted_phrase). Search for a verbatim phrase that uniquely locates the passage; if the phrase repeats, narrow with surrounding context.page.add_highlight_annot(rects).annot.set_info(content=comment_text) and annot.update(). The comment is visible as a sticky-note popup on hover/click in any PDF reader.page.add_strikeout_annot) at the relevant rectangles, with the suggested replacement in the comment text. PDF cannot apply tracked-change-style replacements directly, but the strikethrough + comment communicates the same revision intent.page.add_text_annot(point, comment_text). Place at the top-left of the relevant region so it doesn't obscure content.Reviewer attribution: set annot.set_info(title="Reviewer") (or the specified persona) so all reviewer annotations are filterable in Acrobat / Preview.
Pattern:
import fitz
doc = fitz.open(path)
for page in doc:
for finding in findings_for_this_page:
rects = page.search_for(finding["anchor_text"])
if rects:
highlight = page.add_highlight_annot(rects)
highlight.set_info(title="Reviewer", content=finding["comment"])
highlight.update()
else:
# Anchor text not found verbatim — drop a margin text annotation
page.add_text_annot(fitz.Point(40, 40), finding["comment"]).set_info(title="Reviewer")
doc.save(path.replace(".pdf", "_REVIEWED.pdf"))If the anchor text cannot be located (PDFs with extracted-from-image text, OCR artifacts, or hyphenation breaks), fall back to a margin text annotation on the relevant page, and note in the structured review's section 4/5 that this specific finding could not be precisely anchored.
Verification before delivery: re-open the saved _REVIEWED.pdf and confirm the annotation count matches the count of findings authored. Mismatch = silent failure; surface it.
See Step 12 "Annotated PPTX output" — the presentation mode covers PPTX annotation in detail (native PowerPoint comments anchored to specific slides and shapes). When .pptx is submitted to default-mode review (not presentation mode), still use the Step 12 annotation mechanics — comments anchored at shape level, not appended to speaker notes.
For .tex source files, insert reviewer comments as % REVIEWER: line comments immediately above the relevant line. Plain line comments work in every TeX editor and don't require additional packages.
% REVIEWER: This claim needs a citation. Suggested: Smith (2021).
The data showed a clear effect on memory consolidation.
% REVIEWER: "Subjects" → "participants" (current journal convention).
The 247 subjects were recruited from undergraduate courses.
% REVIEWER: This paragraph contradicts the methods description in §3.2. Resolve.
\subsection{Procedure}For users who want tracked-change-style markup (e.g., supervisor-style line edits visible inline), offer the changes package as a follow-up option with the --latex-changes flag:
\usepackage[final]{changes}
\definechangesauthor[name={Reviewer}, color=blue]{rev}
\replaced[id=rev]{participants}{subjects}
\added[id=rev, comment={citation needed}]{(Smith, 2021)}
\deleted[id=rev]{ — this clause adds nothing}Default is % REVIEWER: line comments (zero dependencies). The changes-package version is opt-in because it requires a \usepackage line in the document preamble.
For BibTeX files (.bib), apply % REVIEWER: comments above the relevant entry the same way.
Deliver as _REVIEWED.tex (and _REVIEWED.bib if applicable). Verify by line-counting reviewer-comment lines in the output and matching against the count of findings.
For Markdown, RTF, HTML, ODT, Pages, Google Docs, Jupyter notebooks, plain text, and pasted content, produce only the structured review. Sections 4 and 5 should reference exact passages (with quoted snippets or location markers like "section 3, paragraph 2") so the author can find them in their original document.
For these formats, offer the user the option of converting to a supported format (docx for prose, pptx for slides) if they want anchored inline annotations, but do not block on this.
The skill auto-installs the required libraries on first use:
python3 -c "import docx, pptx, fitz, lxml" 2>/dev/null || pip install python-docx python-pptx pymupdf lxml --quietInvoked by /peer-review --committee (alone or composed with --paper or --homework). Replaces the single seasoned-professor voice with a panel of 3 to 5 reviewers, each with a distinct domain specialization, evaluative priority, and voice. The synthesis across them is the value-add over running multiple single-reviewer passes.
This is the closest the skill gets to simulating a real PhD committee or journal review panel. A solo reviewer can miss a methodological flaw because they are reading the work as a philosopher; a panel will catch what each individual misses, and the disagreements between members are themselves informative for the author.
Default: 3 members, drawn to maximize useful diversity across the work's identified domains plus, where it would help, one outside lens to stress-test assumptions the inside lenses share.
Members are not redundant. Two analytic philosophers add nothing over one. A philosopher of mind and a cognitive scientist add a lot. The skill picks for range, not chorus.
One member is always adversarial by default. The Adversary's job is not to find a verdict but to find what would falsify the work's central claim. It steelmans the strongest possible objection, looks for unstated assumptions whose denial would collapse the argument, and proposes the test the work would most fear. The Adversary is not gratuitously hostile; it is rigorously hostile to the thesis, not to the author. Its severity is by definition "harsh," but its style is engaged and serious rather than dismissive. If a paper survives the Adversary, that survival is itself a strength worth noting in the synthesis.
The Adversary can be turned off by --no-adversary if the user wants pure positive-substance reviewing, or doubled (--double-adversary) for stress-test-heavy work where the user wants two distinct lines of attack.
Each member (including the Adversary) is defined by:
references/domain-lenses.md to articulate the relevant rigor criteria.User override: user can specify the committee by description (e.g., "one philosopher of mind, one ML researcher, one feminist epistemologist"). The skill instantiates members matching the description, and adds an Adversary by default unless the user opts out.
For each committee member, in turn:
After all individual reviews, a synthesis section:
Each member's voice must be distinguishable. If reviews from members A and B sound the same, the committee mode is failing. Severity, lexicon, and characteristic moves should differ.
Each member stays in character across their entire review. Reviewer A does not concede on a point Reviewer B is raising; resolving disagreements is the synthesis section's job.
If a docx is involved, comments are added with each reviewer's identifier as the author (e.g., author "Reviewer A: Methodologist"). Each member only annotates passages relevant to their priority. Existing rules about engaging with prior comments apply per member.
If two members would annotate the same passage with conflicting suggestions, both annotations are added; the disagreement is preserved rather than smoothed away. The synthesis section is where disagreements get framed, not the docx.
In iterate mode after a committee review, the user can address the committee as a whole, or specific members ("Reviewer A, defend your point about confounds"). The targeted member responds in their own voice. The anti-sycophancy clause (Step 11) applies per member: a member can update their position only on substantive grounds, not on social pressure.
The user can also request that a specific member re-review a rewritten passage in their voice.
Invoked by /peer-review --fact-check (alone or composed with --paper or --homework). A specialized verification pass for essays, papers, and homework that may have been written with AI assistance. Detects hallucinations: fabricated sources, misrepresented citations, made-up facts, and "AI fingerprints" (stylistic and structural tells of LLM-generated content).
This is not a substitute for substantive review of the work's argument. It is a verification pass on the work's factual scaffolding. The output is a report on what is and is not trustworthy in the work, after which a normal review can proceed (or not) on solid ground.
/peer-review --fact-check --paper): run fact-check first, then content review on the verified portion.Three layers, in order of priority:
1. Source existence and metadata. For every cited source:
2. Content fidelity. For every load-bearing claim attributed to a source:
3. Unsourced factual claims. For every specific factual assertion not tied to a citation:
Before flagging any source as UNVERIFIABLE or MISSING, the skill systematically checks for legal open-access copies. Many "paywalled" papers have free, legitimate copies available; the skill is responsible for finding them before declaring a source unreachable.
Cascade (in order; stop when content is found):
Only after this cascade returns nothing should the source be flagged UNVERIFIABLE for content. For metadata, the cascade plus standard search should be definitive: if no record of the paper exists anywhere across all of these resources, the paper is likely MISSING (does not exist), not just unfindable.
When the cascade returns metadata but not full text, do not flag UNVERIFIABLE. Instead, use the METADATA VERIFIED / CONTENT NOT ACCESSED status, and explicitly invite the user to supply the relevant content for content-fidelity verification (see "User-supplied content workflow" below).
In the report, partial-verification entries clearly distinguish what was verified from what was not:
When a source cannot be accessed (paywalled, not digitized, in a private corpus, or otherwise outside the skill's reach), the skill explicitly invites the user to provide the source material so verification can complete. This is invoked at the end of the initial fact-check report and during iterate mode.
What the user can supply:
How it is integrated:
Limits of the workflow:
Beyond direct fact-checking, scan for stylistic and structural patterns characteristic of LLM-generated content:
These are heuristics, not proof. A clean writer with a polished prose style is not a liar. Flag patterns; do not accuse based on style alone. The "trust assessment" output is calibrated by what gets confirmed, not by stylistic suspicion alone.
#### 1. Header
#### 2. Source verification A list of every cited source with status. For misrepresentations and missing sources, quote the relevant passage from the work and explain the issue. List in priority order: confirmed fabrications first, then misrepresentations, then metadata errors, then unverifiable, then verified.
#### 3. Unsourced claim verification A list of factual claims not attached to citations, with status. Priority order: confabulations first, then unverified, then accurate. (Accurate items can be summarized as a count if there are many.)
#### 4. AI fingerprint scan Pattern flags found in the work, ordered by strength of signal. For each, quote the passage and explain why it is a flag. State explicitly that these are heuristics, not proof of AI use.
#### 5. Trust assessment Overall judgment on the factual integrity of the work. Pick one:
#### 6. Recommendation What the user should do with this work given the findings. Examples:
Fact-check mode uses web_search and web_fetch heavily. Search for each citation by author and title. Fetch sources where possible to verify content fidelity. Search for unsourced specific claims to verify or refute.
Run the open-access search cascade (above) before flagging any source UNVERIFIABLE. The cascade exists because the difference between "this source is paywalled and the skill cannot read it" and "this source has a free legal copy somewhere the skill did not look" is enormous: the first is an honest limitation, the second is a lazy verification.
Where the cascade returns nothing for a paper that should exist (well-known author, plausible venue, no obvious flags), flag as UNVERIFIABLE rather than MISSING and note that the cascade returned no results. Where the cascade returns nothing AND the citation has additional flags (suspiciously specific DOI that doesn't resolve, no record of the author publishing in the cited venue, etc.), flag as MISSING.
Distinguish carefully:
When fact-check is composed with paper or homework (/peer-review --fact-check --paper):
The two outputs are delivered separately. Do not silently merge.
The user can challenge specific findings in iterate mode ("you flagged source 3 as missing, but here's the actual link"). The reviewer responds in the same anti-sycophancy frame. If the user provides a verification (a working link, a found quote, a PDF of the source), the reviewer re-runs verification on the supplied material and updates the status. The reviewer does not update on social pressure alone.
The user can also supply content for METADATA VERIFIED / CONTENT NOT ACCESSED entries to upgrade them to full verification. When this happens:
If the user disputes an AI-fingerprint flag, the reviewer either defends (with the specific pattern that triggered the flag) or concedes (if the pattern was actually a false positive on closer reading). Stylistic flags are heuristics; conceding when wrong is appropriate.
If a docx is provided, annotations are added at each flagged location:
Tracked changes are not used in fact-check mode. The job is detection, not editing. The author decides what to do with the flags.
Invoked by /peer-review --plagiarism-check (alone or composed with --paper or --homework). A specialized verification pass for detecting whether the work contains content lifted from existing sources without attribution. Distinct from fact-check (Step 8): fact-check looks for fabricated sources and confabulated facts; plagiarism-check looks for the opposite problem, real sources used as text without credit.
This is most often relevant for homework, but applies to papers as well (especially literature reviews and theoretical sections that draw heavily on existing arguments).
Three layers, in order of priority:
1. Direct lifted text. Verbatim or near-verbatim passages from existing sources without quotation marks or attribution. These are the most clear-cut cases.
2. Paraphrased lifted argument. Sustained borrowing of an argument's structure, examples, or conceptual moves without attribution. The hallmark: the work follows the source's logical sequence (premises, examples, conclusions) without naming the source.
3. Uncredited ideas. Specific concepts, frameworks, or terminology that originated in identifiable sources, presented as the author's own.
Not every uncredited resemblance is plagiarism. The skill applies a calibrated standard:
The skill errs on the side of flagging rather than excusing: false positives can be defended; false negatives cannot.
#### 1. Header
#### 2. Lifted text findings List, in priority order, every flagged passage with status. For each:
If nothing was found, write "no instances of lifted text detected" and explain the search scope.
#### 3. Verdict Pick one:
#### 4. Recommendation What the user should do. Examples:
Plagiarism-check uses web_search and web_fetch. Search distinctive phrases (multi-word sequences with low-frequency combinations); search for distinctive arguments and example structures.
Apply the same open-access cascade as fact-check (arXiv, biorxiv, Unpaywall, OpenAlex, CORE, OpenAIRE, Semantic Scholar, Google Scholar, author pages, institutional repositories) when checking distinctive phrases against possible original sources. A passage that matches a paper available only on arXiv as a preprint is just as plagiarized as one matching the published version; the cascade ensures the skill checks both.
Limitations of the search-based approach: only catches what is web-discoverable through the cascade. Paywalled journals without OA copies, books not digitized, and private corpora are not searchable. The skill flags this honestly: a "clean" verdict means "nothing found through the open-access cascade and standard web search," not "guaranteed original."
For institutional plagiarism-detection tools (Turnitin, iThenticate, etc.), the skill notes that those tools have access to subscription corpora that the skill does not, and recommends them as a complementary check for high-stakes submissions.
When the user is checking a draft against specific sources they have access to but the skill cannot reach (a colleague's manuscript shared in confidence, a paywalled book, an unpublished thesis), the user can supply the source material directly. The skill then runs the same plagiarism comparison against the supplied content.
How it works:
Limits: the skill cannot independently verify that user-supplied content is the source it is claimed to be. The verification is reported as resting on the user's good faith.
When plagiarism-check is composed with paper or homework:
The user can challenge specific findings ("this is common knowledge in my field," "I cited that source elsewhere"). The reviewer evaluates the challenge:
If a docx is provided, annotations are added at each flagged passage:
Tracked changes are not used in plagiarism-check mode; the corrections require authorial decisions (re-write, quote, or credit), not line edits.
Invoked by /peer-review --draft (alone or composed with --committee). For unfinished work where the author wants generative thinking-partner feedback, not evaluation. The mode treats the reviewer as a senior colleague the author is talking through ideas with, not as a gatekeeper rendering verdict.
Draft mode is opt-in. The default rule (treat every submission as final-draft) still applies in every other mode. The point of draft mode is not to be softer; it is to produce a different kind of feedback better suited to work that is still being shaped.
The default review mode answers "is this work good, and what's wrong with it?" Draft mode answers "where is this work going, and what would help it get there?" The shift is from evaluative to generative.
Specifically:
Replace the standard structured review (Sections 0 to 7) with the following:
#### 0. TLDR 2 to 4 sentences:
#### 1. Header
#### 2. Reading of the intended argument A faithful, charitable reconstruction of where the work appears to be going, in the
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.