joap-study-design — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited joap-study-design (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
JAP holds measurement and design to an exacting standard. The recurring killers are common-method variance (CMV), weak causal warrants (cross-sectional single-source data), unmodeled nesting, and construct validity gaps. This skill hardens the design before data collection, where most of these problems can actually be solved.
new or contested, provide validity evidence (CFA, convergent/discriminant, measurement invariance across groups/time). A weak measure dooms an otherwise good design.
with temporal separation (multi-wave), multiple sources (self + supervisor + objective), experimental or quasi-experimental legs, or a field experiment.
protected anonymity) and plan statistical checks; declare the strategy up front. Post hoc Harman's single-factor test alone is treated as insufficient at JAP.
ICC(1)/ICC(2) and r_wg for aggregated constructs, and use multilevel models — do not ignore dependence or aggregate away the structure without justification.
cross-level interaction or indirect effect), not just the total N; for multilevel designs, the L2 sample size usually constrains power.
| Remedy | Type | Note |
|---|---|---|
| Temporal separation (multi-wave) | procedural | predictor and outcome at different waves |
| Source separation (self + other/objective) | procedural | the strongest single defense |
| Measurement/context separation | procedural | different scales/formats for predictor vs outcome |
| Protected anonymity, balanced items | procedural | reduces consistency and acquiescence bias |
| Marker variable / CFA marker technique | statistical | plan a theoretically unrelated marker in advance |
| ULMC (unmeasured latent method construct) | statistical | report alongside, not instead of, procedural remedies |
For the servant-leadership package, justify N at the level the hypotheses live, before collecting.
Multilevel field study (2-2-2 / 2-1-2 mediation):
Constraint: 74 teams (L2) drives power for the team-level indirect effect.
Power target: 80% for the indirect effect (Monte Carlo power for multilevel
mediation), assuming a path ≈ .25, b path ≈ .30, ICC(1) ≈ .15.
Result: target ≥ 70 teams, ~8 members each → ~560–620; we collect 612 in 74.
Lab experiment (causal leg):
Between-subjects, two conditions; power for the interaction (H3 boundary),
N ≈ 240 at 80%, alpha .05; fixed-N, no optional stopping.
Aggregation: report ICC(1), ICC(2), r_wg(j) to justify team-level aggregation
of psychological safety; preregister exclusion rules.| Degree of freedom | Lock before data? | Where it lives |
|---|---|---|
| Hypotheses + direction + level | yes | preregistration |
| Measures (all scales, all items) | yes | preregistration (prevents scale cherry-picking) |
| CMV remedies (procedural + planned statistical) | yes | design + preregistration |
| Aggregation rules (ICC/r_wg thresholds) | yes | analysis plan |
| Exclusion rules (careless responding, attrition) | yes | preregistration |
| Covariates / model form | yes | analysis plan |
| Exploratory analyses | allowed, labeled | reported separately, post hoc |
experimental leg; declare procedural remedies, not just a Harman's test.
power analysis (handoff to joap-data-analysis).
Estimate and audit the design, don't only describe it. Full map: execution-with-mcp. JAP is organizational psychology — multilevel survey/field data and experiments; cluster at the right level and apply mediation/moderation discipline.
detect_design → recommend → fit with as_handle=true → audit_result.callaway_santanna / sun_abraham +bacon_decomposition + honest_did_from_result); IV (effective_f_test + anderson_rubin_ci); RDD (rdrobust + mccrary_test).
romano_wolf for many-outcomefamily-wise control, and mediate for mediation (not naive controlling-away).
oster_delta / sensemakr for observational claims.Report the effect size in interpretable units; route the full battery to the appendix/supplement. A run end-to-end (synthetic data, real returns) is in the JF execution walkthrough.
【Construct validity】reliability + CFA/invariance evidence? [Y/N]
【Causal warrant】temporal / multi-source / experimental leg present? [Y/N]
【CMV】procedural remedies + planned statistical check declared? [Y/N]
【Nesting】levels, ICC/r_wg, multilevel model justified? [Y/N/NA]
【Sample size】powered for the carrying effect at the right level? [Y/N]
【Next】joap-data-analysis../../resources/external_tools.md — Mplus/lavaan/lme4, Monte Carlo power, CMV-marker and invariance tools../../resources/official-source-map.md — measurement, design, and reporting expectations~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.