Glaucoma Screening · Explainable Multimodal

Fundus → probability + attention map · image-only & +demographics · Grad-CAM / Integrated Gradients

⚠ Research prototype — not a medical device

Upload a fundus image

🔬

Drop a fundus image or click to browse · JPEG/PNG/BMP

Age / sex / IOP feed the +demographics fusion models; leave blank to mark them “unknown” (missingness-masked). The heat-map always shows the image evidence.

Running inference + attribution…

👁

Upload a fundus image, then hit Analyse.

What this tool demonstrates

A rigorous evaluation of explainable, multimodal glaucoma screening from retinal fundus photographs. The models here reach ~0.95 AUROC on pooled colour-fundus data; the demo shows their prediction, an optic-disc-focused attention map (two attribution methods), and — for annotated samples — the measured fraction of attention that lands on the optic disc.

Headline research findings

RQ1 · Fusion vs image-onlyStatistically equivalent on colour fundus across 10 seeds per architecture — ResNet50 0.943 vs 0.954, Swin-Tiny 0.959 vs 0.957; neither difference significant.
Metadata completenessFusion does help when demographics are complete: +0.023 AUROC (DeLong p=0.0007), and the advantage scales with coverage.
RQ2 · GeneralisationStrong cross-dataset domain shift (leave-one-dataset-out drops up to −0.29).
RQ3 · ExplanationsAttributions are 2–6× more disc-focused than chance; but “well-localised ≠ always faithful”.
FairnessSubgroup gaps exist and are dataset-dependent; demographic signal is largely age-driven.

Safe-use disclaimer

This is a research prototype for demonstration only. It is not a medical device, must not be used for diagnosis or patient care, and its outputs can be wrong — especially on images from cameras or populations unlike its training data. Any clinical decision requires a qualified ophthalmologist.