Representation learning for medical image segmentation quality

Marciano, Vincenzo
Thesis

Deep learning has made medical image segmentation remarkably accurate, yet state-of-the-art models share a dangerous property: they fail silently, returning confident, well-formed masks even when those masks are anatomically wrong, and they do so precisely on the out-of-distribution data encountered in real clinical practice. Because the ground-truth annotations used to validate a model do not exist for the next patient, these failures go undetected at deployment, which undermines the trust required to use automated segmentation in the clinic. This thesis addresses the problem through a single geometric idea: the quality of a segmentation can be encoded as the distance from a quality manifold, a structured region of a learned latent space in which anatomically valid segmentations live and failures do not. From this idea we develop three contributions that give a segmentation system the ability to estimate, correct, and evaluate its own outputs without ground truth.

First, nnQC learns the quality manifold of an organ with a conditional latent diffusion model and reads quality from the gap between a prediction and its reconstruction. This self-configuring, metric-agnostic estimator recovers plausible references even from severely corrupted masks and outperforms prior quality-control methods across fifteen datasets and seven organs.

Second, GRACE removes the need to retrain a separate model per organ. It retrieves anatomically similar cases from a multimodal knowledge base and conditions a generative refiner on them, so that a single model corrects masks across ten organs and generalises zero-shot to unseen datasets. The retrieval mechanism is also found to recover patient-specific priors across longitudinal visits, without any temporal supervision.

Third, MID lifts quality assessment from the single prediction to the population. Inspired by the Fréchet Inception Distance, it compares the joint distribution of image-mask pairs against a reference distribution in a medical feature space, yielding a reference-free score that tracks segmentation quality, detects distribution shift, and ranks segmentation models without any per-sample ground truth.

Together, these contributions trace one arc, from learning the quality manifold, to navigating it, to measuring distance from it, and bring automated medical image segmentation closer to the reliability that clinical deployment demands.


Type:
Thesis
Date:
2026-09-25
Department:
Data Science
Eurecom Ref:
8947
Copyright:
© EURECOM. Personal use of this material is permitted. The definitive version of this paper was published in Thesis and is available at :
See also:

PERMALINK : https://www.eurecom.fr/publication/8947