Systematic analysis of evaluation metrics for deep learning-based surgical phase recognition in laparoscopic cholecystectomy

Evaluation metrics are essential for assessing the performance of AI models and enabling comparability across different approaches. However, in surgical phase recognition, their application remains inconsistent. This study provides a systematic analysis of evaluation metrics used in this field,...

Evaluation metrics are essential for assessing the performance of AI models and enabling comparability across different approaches. However, in surgical phase recognition, their application remains inconsistent. This study provides a systematic analysis of evaluation metrics used in this field, aiming to support standardization and encourage more explicit explanation of metric selection. A systematic literature review was conducted following the PRISMA framework, covering publications from 2016 to 2025. In total, 50 studies were identified and analyzed with respect to the evaluation metrics applied, with particular focus on relaxed boundaries and the distinction between online and offline models. The results show that accuracy, precision, and recall are the most frequently used metrics, followed by the jaccard index. Since around 2023, an increasing use of segment-based metrics can be observed, reflecting a growing emphasis on temporal dynamics. No significant differences in metric selection between online and offline approaches were identified. Furthermore, many studies do not provide explicit explanation for their choice of evaluation metrics. Instead, they often rely on commonly used metrics or practices adopted from previous work. Additionally, results obtained with relaxed boundaries tend to exhibit higher standard deviations compared to those without relaxed boundaries. Based on these findings, it is recommended to more explicitly align evaluation metrics with model objectives and to further investigate the distinction between online and offline settings. Adaptations of the f1-score, such as the fβ-score, may provide a more flexible evaluation framework. Furthermore, the Matthews Correlation Coefficient represents a promising complementary metric, particularly for multi-class settings, if appropriately normalized. Future work should focus on clearly explaining metric selection and consistently reporting relaxed boundaries to improve transparency, comparability, and interpretability in surgical phase recognition.

Source: Frontiers AI — Published — Category: Research

🔗 Read full article on Frontiers AI →