Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses

| Source: arXiv AI

Tags: medical imaging, lymphoma, PET/CT, deep learning, segmentation, arXiv

A clinical evaluation of four deep learning networks for lymphoma segmentation across 611 multi-institutional PET/CT cases finds that AI errors closely resemble inter-observer physician variability — small, faint lesions remain equally challenging for both humans and AI.

Details

This study evaluates four deep neural networks (ResUNet, SegResNet, DynUNet, SwinUNETR) for lymphoma lesion segmentation and quantification from PET/CT images across 611 cases from multi-institutional datasets covering diverse lymphoma subtypes and imaging conditions. Beyond standard Dice similarity coefficients, the evaluation framework adds lesion-specific measures, detection criteria, and a new Detection Criterion 3 based on metabolic lesion characteristics. Out-of-distribution testing is included — a notable gap in most existing deep learning studies in this space. The core clinical finding: AI networks perform well on large, metabolically active lesions but struggle with small, faint ones — errors that closely resemble those made by expert physicians in intra- and inter-observer variability analyses. The remaining gap between AI and human performance is concentrated in biologically hard cases, not systematic model failures. Ahamed et al. argue that this implies current evaluation metrics (focused on overall segmentation accuracy) may overstate AI performance limitations by conflating biologically hard cases with model deficiencies. Code is publicly available. Now at v5, originally submitted in 2023.