[2607.19372]

Beyond Confidence in AI-Assisted Colonoscopy: A Spatial, Temporal and Quality-Aware Audit Framework for Endoscopic AI Review


AI-assisted colonoscopy systems commonly report frame-level confidence scores, but confidence alone does not show whether a prediction is spatially plausible, temporally persistent or reliable under degraded image quality. We propose EndoExplain, a lightweight and reproducible audit framework for endoscopic AI review. The framework integrates classification confidence, lesion segmentation, CAM-style visual attribution, attribution-mask alignment, frame-quality indicators and temporal event summarisation, with the aim of separating signals that are often conflated in computer-aided detection pipelines. On HyperKvasir, the selected EfficientNet-B0 classifier reaches 0.9280 test accuracy over ten endoscopic classes and 0.9969 ROC-AUC for the polyp_family versus rest view. The selected U-Net++ EfficientNet-B1 segmenter reaches Dice 0.9318 and IoU 0.8826 on the segmented test split. A strict top-20% multi-method attribution audit shows that attribution method strongly changes explanation-mask agreement: Eigen-CAM gives the strongest overlap, whereas Grad-CAM++ remains weakly aligned. A frozen-model external sanity check on ETIS-LaribPolypDB and CVC-ClinicDB preserves this attribution-method ranking while showing dataset-dependent overlap. A human-reviewed clip-level temporal benchmark over 60 HyperKvasir videos reaches event F1 0.8081 at the pre-specified threshold 0.85 using one-to-one event matching with overlap >= 1 s. A clinician-informed external plausibility review supported the clinical readability of separating confidence, localisation, attribution, quality metadata and temporal context. The resulting cockpit-style review layer presents these signals as distinct auditable outputs. This is a retrospective research prototype and not a clinically validated medical device.