aerj-research-design — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited aerj-research-design (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
AERJ accepts many methodologies but is demanding about each. The design must credibly connect the framework (aerj-theory-and-framework) to evidence and meet the relevant AERA reporting standards. This skill is mode-aware: name the dominant education-research lens and defend it against the strongest alternative explanation.
aerj-literature-positioningspecify levels, random effects, and cluster-correct inference. Report the design effect / ICC.
IRT/factor evidence. Validity is a design issue, not an afterthought.
quasi-experimental (DID/event study with modern estimators, RD, IV, matching) — defend identifying assumptions, don't assert them. Map to What Works Clearinghouse-style expectations when claiming effects.
sampling), not convenience. Say what the case is a case of.
audit trail, researcher positionality/reflexivity.
were warranted by evidence (hand off to aerj-data-analysis).
the rationale for mixing — what integration buys you that one strand cannot.
For the single strongest rival explanation, write one sentence: "If the rival were true rather than my account, the evidence would look like ___; instead it looks like ___." If you cannot, the design does not yet identify the contribution.
Estimate and audit the design, don't only describe it. Full map: execution-with-mcp. AERJ is empirical education research — field experiments and observational school data; multilevel inference and many-outcome corrections are central.
detect_design → recommend → fit with as_handle=true → audit_result.callaway_santanna / sun_abraham +bacon_decomposition + honest_did_from_result); IV (effective_f_test + anderson_rubin_ci); RDD (rdrobust + mccrary_test).
romano_wolf for many-outcomefamily-wise control, and mediate for mediation (not naive controlling-away).
oster_delta / sensemakr for observational claims.Report the effect size in interpretable units; route the full battery to the appendix/supplement. A run end-to-end (synthetic data, real returns) is in the JF execution walkthrough.
AERJ judges each methodology on its own terms, so the credibility bar differs by mode. Use this matrix to locate the assumption a referee will press hardest.
| Mode | Core thing the design must establish | The assumption referees attack |
|---|---|---|
| RCT | Power/MDE, balance, fidelity, low differential attrition | Attrition or non-compliance undoing randomization |
| Quasi-experimental | A credible counterfactual | Parallel trends / continuity at the cutoff / exclusion |
| Multilevel descriptive | Correct nesting and measurement | Cluster level mis-specified; validity unaddressed |
| Qualitative | Trustworthiness and case logic | Convenience sampling dressed as theoretical |
| Mixed | A real point and method of integration | Two strands never actually joined |
An AERJ team evaluates a peer-tutoring program with a regression-discontinuity design on an eligibility test score. The credibility case states the estimand (effect at the cutoff), shows a density test with no manipulation, reports a bandwidth-robust estimate of an illustrative 0.21 SD on the outcome, and writes the adjudication sentence: if selection rather than the program drove the jump, covariates would also jump at the cutoff; instead they are smooth. That single sentence rules out the strongest rival. A weak version would assert "the program caused gains" with no continuity evidence — exactly the move a methodological referee rejects.
claim to description with a mechanism hypothesis.
say what the case is a case of.
confirm method-specific expectations against the journal's current submission guidelines.
【Mode】quant / qualitative / mixed
【Estimand or claim】what is being identified/shown/understood
【Key assumption(s) / trustworthiness】and how each is defended
【Rival ruled out】the adjudication sentence
【Standards】which AERA reporting standard the design meets
【Next】aerj-data-analysis../../resources/external_tools.md — multilevel/IRT/causal packages and CAQDAS for qualitative work../../resources/official-source-map.md — AERA reporting standards + preregistration notes~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.