jedpsych-study-design — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited jedpsych-study-design (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
The Journal of Educational Psychology expects designs that are adequately powered for their nesting structure, measure learning constructs well, and have ecological validity for real educational settings. Because JEP studies are usually students nested in classes nested in schools, the single most consequential design decision is matching the unit of randomization, power, and analysis to the level at which the treatment and mechanism operate. This skill hardens the design before data collection.
experiment's effective N is the number of clusters, not students. Power at the cluster level using the intraclass correlation (ICC) and number/size of clusters; plan the matching multilevel analysis up front (see jedpsych-data-analysis).
their size — a power analysis for the smallest educationally meaningful effect, given the ICC and a pretest covariate that absorbs cluster variance. State the assumed effect size and its source.
construct; justify their reliability and that they capture transfer/learning, not just teaching to the test. Pre/post designs should plan for measurement at the right grain.
baseline equivalence on covariates; address selection, attrition, contamination across conditions, and teacher/implementation fidelity.
support the educational claim; a stripped lab analog weakens fit at JEP.
exclusion/attrition rules, covariates, and the model. Preregistration is encouraged here.
differences, or RD where assignment is on a cutoff) and state the identifying assumption. For longitudinal/growth designs, plan the timing, attrition handling, and the growth model in advance.
For a teacher-delivered reading-comprehension trial, justify the number of classrooms before recruiting, tied to the smallest educationally meaningful effect — not a round student count.
Smallest meaningful effect: d = 0.20 (a defensible learning gain for a
classroom literacy intervention).
Nesting: students nested in classrooms; assumed ICC = 0.15; ~23 students
per classroom; pretest covariate (r ≈ .6) absorbs cluster variance.
Power: target 80% power, two-sided alpha .05 → ~48 classrooms
(24 per arm), ~1,100 students; design effect handled via the ICC,
not by counting students as independent.
Stopping: fixed number of clusters; no optional addition of schools.
Covariate: baseline comprehension at student and classroom level.State the assumed effect size and its source (prior trial, meta-analytic estimate, or a smallest- meaningful-effect argument). Powering on an inflated lab effect, or on student N alone, is the classic JEP design failure.
| Degree of freedom | Lock before data? | Where it lives |
|---|---|---|
| Hypotheses + direction (at the right level) | yes | preregistration / analysis plan |
| Unit of randomization + number of clusters | yes | preregistration |
| Full measure list (all outcomes) | yes | preregistration (prevents cherry-picking) |
| Exclusion / attrition rules | yes | preregistration, with expected attrition |
| Covariates + multilevel model form | yes | analysis plan |
| Fidelity / implementation measures | yes | protocol |
| Exploratory analyses | allowed, but labeled | reported separately, post hoc |
clusters as the effective N.
ecological validity.
(handoff to jedpsych-data-analysis).
Estimate and audit the design, don't only describe it. Full map: execution-with-mcp. JEdPsych mixes field/lab experiments and observational school data; multilevel (student-in-class-in-school) inference and many-outcome corrections matter most.
detect_design → recommend → fit with as_handle=true → audit_result.callaway_santanna / sun_abraham +bacon_decomposition + honest_did_from_result); IV (effective_f_test + anderson_rubin_ci); RDD (rdrobust + mccrary_test).
romano_wolf for many-outcomefamily-wise control, and mediate for mediation (not naive controlling-away).
oster_delta / sensemakr for observational claims.Report the effect size in interpretable units; route the full battery to the appendix/supplement. A run end-to-end (synthetic data, real returns) is in the JF execution walkthrough.
【Unit】randomization / power / analysis level (matched?) [Y/N]
【Sample size】# clusters + size + ICC + smallest meaningful effect
【Measures】validated learning outcome + reliability + transfer? [Y/N]
【Baseline + confounds】equivalence, attrition, fidelity addressed?
【Ecological validity】setting / delivery supports the educational claim?
【Preregistration】confirmatory core locked? where?
【Next】jedpsych-data-analysis../../resources/external_tools.md — PowerUpR, simr, Optimal Design, multilevel/SEM software, preregistration templates../../resources/official-source-map.md — JARS reporting standards and preregistration policy~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.