aperivue/uncertainty-imaging
> Design or audit the uncertainty-quantification, out-of-distribution (OOD) detection, and selective-prediction layer of a medical-imaging model framed for deployment — so a clinical-use claim carries calibrated per-case uncertainty (MC-dropout / deep ensemble / conformal / Bayesian), an OOD guard validated on a held-out OOD set, an abstention rule at a pre-specified operating point, and uncertainty checked under distribution shift. Emits an uncertainty manifest and a deterministic gate that flags a deployment claim built on point predictions, conformal intervals with unmeasured coverage, and an OOD claim with no held-out OOD data. Integrates MAPIE / captum / pretrained OOD scorers; it does not reimplement them and never runs a model on real patient data.
npx skills add https://github.com/Aperivue/medsci-skills --skill uncertainty-imaging
A medical-imaging model framed for deployment must say more than "class 1, 0.87". It needs a
calibrated uncertainty on each case, an out-of-distribution (OOD) guard validated on data known
to be out-of-distribution, and — if it abstains — a pre-specified operating point. The failures are
predictable and reviewer-visible: a clinical-use claim built on point predictions, conformal intervals
quoted without ever measuring their coverage, an "OOD detector" evaluated only on in-distribution data,
a deep ensemble whose members share a seed, and uncertainty validated only in-distribution when
deployment sees scanner/site/case-mix shift. This skill designs that layer and audits an existing one
(Gal 2016; Lakshminarayanan 2017; Angelopoulos & Bates; Ovadia 2019; DECIDE-AI).
It is the deployment-safety companion in the model-engineering lane: /model-evaluation computes the
held-out metrics and calibration, and uncertainty-imaging covers the uncertainty / OOD / abstention
machinery a deployment claim rests on. It integrates MAPIE (conformal), captum, and pretrained OOD
scorers; it does not reimplement them and never runs a model on real patient data.
unsure, or off-distribution?"
shift checks right before submission.
/model-evaluation then/analyze-stats.
/model-scaffold (+ /model-validation)./explainability./radiomics-ml + /analyze-stats.all — add MC-dropout, a deep ensemble, conformal prediction, or a Bayesian estimate.
can fail on clinical data — measure achieved coverage on a held-out calibration/test set.
until you evaluate on data known to be out-of-distribution (different scanner / site / pathology).
underestimates epistemic uncertainty.
pass is identical and the estimate collapses to a point prediction.
pre-specify the coverage / risk operating point.
degrades under shift, so report it on shifted / external data.
the strongest default when a calibration set is available. Validate empirical coverage.
best-quality epistemic uncertainty, at K× cost.
See references/uncertainty_guide.md.
on a held-out OOD set** (different scanner/site/pathology) and report detection AUROC + the operating
point.
target coverage or risk; report the risk–coverage curve.
Report calibration / coverage on shifted or external data, not in-distribution only (Ovadia 2019).
{
"task": "classification",
"deployment_claim": true,
"uncertainty_method": "conformal",
"coverage_target": 0.90,
"coverage_validated": true,
"ood_method": "mahalanobis",
"ood_heldout_set": "external-ood-cohort",
"selective_prediction": true,
"selective_target": 0.95,
"calibration_under_shift": true
}
python3 scripts/check_uncertainty_reporting.py --manifest uncertainty_manifest.json --strict
Verdicts: POINT_PREDICTION_NO_UNCERTAINTY, CONFORMAL_NO_COVERAGE_VALIDATION, OOD_NO_HELDOUT_SET
(Major); ENSEMBLE_NOT_INDEPENDENT, MCDROPOUT_DISABLED_AT_INFERENCE, SELECTIVE_NO_TARGET,
NO_CALIBRATION_UNDER_SHIFT (Minor). Audits the declared spec at design/report time; it complements
/model-evaluation's executed calibration/subgroup metrics.
/model-evaluation — the point predictor's held-out metrics + calibration this layer sits on top of./analyze-stats — calibration curve / risk–coverage plotting for the report./check-reporting — TRIPOD+AI / DECIDE-AI deployment-monitoring items./model-validation — the DECIDE-AI monitoring seam (the deployment-time counterpart of the splitaudit).
reported number comes from the researcher's executed code — never invented. This skill designs and
audits the uncertainty spec; it does not run a model on real patient data.
clinical data (CONFORMAL_NO_COVERAGE_VALIDATION).
check_uncertainty_reporting.py. Theverdict is reproduced deterministically, never asserted from prose.
or OOD library or claim results for one.
scripts/check_uncertainty_reporting_challenge/ ships a synthetic weak/strong uncertainty-manifest pair
with a network-free verify.sh wired into the skill's validation commands.
Take aperivue/uncertainty-imaging from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.