aeja-robustness — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited aeja-robustness (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
AEJ: Applied referees probe whether the headline number is stable, honestly inferred, and not the product of researcher degrees of freedom. Robustness here is not a wall of regressions — it is a targeted set of checks each tied to a specific threat to the design. Map every plausible objection to the one check that answers it, and report the checks so the reader sees the estimate barely moves.
| Threat to the result | The check that answers it |
|---|---|
| Omitted confounders | Oster δ / coefficient-stability bounds; added controls in steps |
| Specification search | a specification curve / multiverse; pre-registered primary spec |
| Functional form | levels vs logs, alternative outcome definitions, nonparametric version |
| Sample selection | drop influential units, alternative inclusion rules, balanced vs unbalanced panel |
| Inference too narrow | clustered SEs at the right level, wild-cluster bootstrap (few clusters), randomization inference |
| Design-specific fragility | DID: honest-DID bounds; RD: bandwidth/donut; IV: weak-IV-robust set |
| Multiple outcomes/subgroups | Romano–Wolf / List–Shaikh–Wooldridge MHT adjustment |
Run the battery, don't just enumerate it. Full map: execution-with-mcp. AEJ: Applied is applied microeconomics — labor, health, education, and development field settings where a clean research design is the entry ticket.
romano_wolf (step-down FWER, accounts forcross-test correlation) or benjamini_hochberg — report the adjusted threshold.
oster_delta / sensemakr — the confounder strength that wouldoverturn the headline.
wild_cluster_bootstrap (few clusters), twoway_cluster / conley.audit_result(result_id) lists the missing checks and theexact suggest_function for each — no guessing the battery.
etable / did_summary_to_latex from the handle — no retyped numbers.Keep the decisive checks in the body and the exhaustive (now actually-run) battery in the appendix. See the executed chain in the JF execution walkthrough.
An IV estimate of the return to a training program is 0.11 (s.e. 0.04). The robustness suite: (i) effective F of 23 rules out weak instruments; (ii) the Anderson–Rubin 95% set is [0.04, 0.19], so inference is not weak-IV-fragile; (iii) Oster δ implies selection on unobservables would need to be 1.8× selection on observables to nullify it; (iv) wild-cluster bootstrap with 14 clusters keeps the CI away from zero; (v) dropping the largest region moves the estimate to 0.10. The point estimate barely moves — the AEJ: Applied target.
specification curve in which the point estimate barely moves.
bootstrap or randomization-inference p-value.
unobservables would have to be (relative to observables) to nullify the result.
【Primary spec】declared / pre-registered? [Y/N] — estimate: ___ (s.e. ___)
【Threat → check map】selection: ___ | spec-search: ___ | form: ___ | sample: ___ | inference: ___ | design: ___
【Inference】clustering level: ___; few-cluster/randomization: ___
【Design sensitivity】honest-DID / RD bandwidth / weak-IV set: ___
【Estimate stability】range across checks: [___, ___]; checks that move it: ___
【Next step】aeja-tables-figures~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.