Volume
Volume 6, Issue 2 (June, 2026) – 7 articles
Aim: To evaluate deep learning models for anatomical structure and peritoneal metastasis (PM) detection and segmentation during staging laparoscopy (SL) using a phase-independent dataset, and to quantify how annotation strategy and spatial representation relate to predictive performance.
Methods: A checklist covering 25 anatomical structures, one surgical instrument, and PM was defined. Detection models (YOLOv9, Co-DETR) and segmentation models (SegFormer, Mask2Former) were trained under two label configurations. Videos were split at the video level (60/20/20). To quantify annotation distribution and spatial representation, two class-level descriptors were derived from the training set: object count and area fraction (percentage of image area occupied by each class). Class-level associations between these descriptors and test-set performance [F1-score, Intersection over Union (IoU)] were evaluated using Spearman correlation.
Results: Thirty SL videos yielded 2,309 annotated frames (1,304/433/572 for training/validation/testing). YOLOv9 reached mean mAP@50 of 0.52 and 0.61; Mask2Former achieved mean IoU of 0.51 and 0.61 and F1-scores of 0.65 and 0.73 for Sets A and B, respectively. Despite 4,094 annotations, PM remained difficult to segment (IoU 0.29-0.30; F1-score 0.45-0.46), due to low area fraction and high heterogeneity. For IoU, area fraction showed stronger correlations with performance than object count (ρ up to 0.66 vs. 0.48). Similar differences were observed for F1-score.
Conclusions: Anatomical detection and segmentation during SL are feasible but limited by small-target representation and heterogeneous intra-abdominal context. Spatial representation is more closely associated with segmentation performance than annotation frequency, supporting annotation strategies that address sparse pixel coverage in phase-independent intra-abdominal models.






