tooluniverse-diagnostic-test-evaluation-627c0e — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited tooluniverse-diagnostic-test-evaluation-627c0e (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Judge how well a test or biomarker discriminates disease — at a fixed cutoff (2×2) or across all cutoffs (ROC) — and turn a result into a probability of disease.
| You have… | Go to |
|---|---|
| A 2×2 table (TP/FP/TN/FN) at a fixed cutoff | Step 1 (Epidemiology_diagnostic) |
| A continuous biomarker score + true labels | Step 2 (ROC / AUC / Youden, Python) |
| A test's sens/spec + a patient's pre-test probability | Step 3 (Epidemiology_bayesian) |
tu run Epidemiology_diagnostic '{"operation":"diagnostic","tp":90,"fp":10,"tn":180,"fn":20}'Returns sensitivity, specificity, PPV, NPV, accuracy, LR_pos, LR_neg, and the sample prevalence.
| Metric | Question it answers | Depends on prevalence? |
|---|---|---|
| Sensitivity = TP/(TP+FN) | Of those WITH disease, what fraction test positive? | No |
| Specificity = TN/(TN+FP) | Of those WITHOUT disease, what fraction test negative? | No |
| PPV = TP/(TP+FP) | If positive, what's the chance of disease? | Yes — strongly |
| NPV = TN/(TN+FN) | If negative, what's the chance of being disease-free? | Yes |
| LR+ = sens/(1−spec) | How much a positive raises the odds of disease | No |
| LR− = (1−sens)/spec | How much a negative lowers the odds | No |
The PPV/NPV trap. Sensitivity and specificity are properties of the test; PPV and NPV depend on the disease prevalence in the tested population. A test with great sens/spec has poor PPV in a low-prevalence (screening) setting. Never quote PPV/NPV from a case-control design (its 50/50 prevalence is artificial) — compute them for the real-world prevalence with Epidemiology_bayesian (Step 3). Report sensitivity, specificity, and likelihood ratios as the prevalence-independent summary.When the test is a continuous score, evaluate across all thresholds:
Prefer the `ROC_analysis` tool — one call returns structured JSON (AUC + bootstrap 95% CI, Youden-optimal cutoff with its sens/spec, optional metrics at a fixed cutoff, and the ROC curve), and works under the MCP server without a shell:
ROC_analysis(scores=[...], labels=[0,1,...]) # inline arrays
ROC_analysis(csv_path="scores.csv", cutoff=0.6) # or a CSV (cols: label, score)The bundled script is the equivalent CLI form:
python skills/tooluniverse-diagnostic-test-evaluation/scripts/roc_analysis.py --input scores.csv
# scores.csv columns: label (1=disease, 0=healthy), score (continuous biomarker)Both report AUC (with a bootstrap 95% CI), the Youden-optimal cutoff (max sensitivity+specificity−1) and its sens/spec.
| AUC | Discrimination |
|---|---|
| 0.5 | no better than chance |
| 0.7–0.8 | acceptable |
| 0.8–0.9 | excellent |
| >0.9 | outstanding |
Turn a result into the probability of disease for a given pre-test probability/prevalence:
tu run Epidemiology_bayesian '{"operation":"bayesian","prevalence":0.10,
"sensitivity":0.90,"specificity":0.95,"test_result":"positive"}'Returns pre_test_odds, the LR, and post_test_probability. This is how you get the real-world PPV: plug the true prevalence in. (Example: a 90%/95% test at 10% prevalence gives a post-positive probability of only ~67%, not 95%.)
tooluniverse-statistical-modeling — logistic regression that produces the score, ORs.tooluniverse-epidemiological-analysis — population-level risk, screening program metrics.tooluniverse-meta-analysis — pool diagnostic accuracy across studies.~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.