SARFA: Segment Anything with Radiomic Feature Alignment

Tyler WardAbdullah Imran
tyler.ward@uky.edu``aimran@uky.edu

Computer Science Department
University of Kentucky, Lexington, KY, USA


Abstract

The Segment Anything Model (SAM) has demonstrated strong generalizability across a variety of segmentation tasks. However, SAM often struggles in situations where the target to be segmented is ambiguous. This poses a problem in medical imaging, where accurate delineation of targets such as tumors is vital, but even expert radiologists can disagree on the appropriate boundary for a target. Addressing this, we propose SARFA (Segment Anything with Radiomic Feature Alignment), a novel framework for improved medical image segmentation. Via probabilistic prompting, SARFA generates a diverse set of plausible masks for each input image and optimizes them with a radiomics-driven training objective based on Fréchet Radiomic Distance (FRD) and Direct Preference Optimization (DPO). By minimizing the FRD between masked predicted and ground truth regions within each image, SARFA encourages segmentation outputs whose anatomical and textural characteristics align with clinically meaningful ground truth representations, without relying solely on pixel-level overlap. Evaluated on computed tomography (CT) and magnetic resonance imaging (MRI) benchmarks, SARFA outperforms existing ambiguous segmentation methods, demonstrating the effectiveness of radiomic feature alignment and DPO-style candidate mask ranking as a training objective. Our code is available at https://github.com/tbwa233/SARFA.

1 Introduction↩︎

In medical image analysis, segmentation is commonly used to isolate regions of interest from the background, which helps to identify lesions as well as their sizes, locations, and relationships with surrounding tissues [1]. Despite great reductions in the time it takes to perform medical image segmentation due to advances in deep learning, training models to recognize difficult targets can still be challenging due to the limited availability of expert annotations and the inherent ambiguity that exists in many medical imaging tasks [2].

Ambiguity in medical images can arise from numerous sources, which we lump into three different categories: boundary ambiguity, inter-rater ambiguity, and aleatoric. Based on established literature [3], we define boundary ambiguity as occurring when lesion margins are difficult to distinguish from surrounding tissue; inter-rater ambiguity as occurring when multiple experts provide differing yet clinically accepted annotations; and aleatoric uncertainty as occurring when the image itself contains insufficient information to uniquely determine the target boundary. As a result of these potential sources of ambiguity, there may exist multiple plausible segmentations for a single image [4]. This poses a problem for conventionally trained segmentation models, as these models are largely deterministic in the sense that they produce just one output mask [5]. As a result, there has been much research into the development of probabilistic, ambiguity-aware segmentation models for medical imaging that explicitly model distributions of plausible annotations [4], [6][15].

Recently, advances in large-scale data training have enabled foundation models like the Segment Anything Model (SAM) [16] to achieve strong generalization across a wide range of segmentation tasks. However, applying such models in medical imaging is difficult due to the substantially different characteristics of medical images vs. natural images, including weak differentiation between targets as a result of subtle contrast changes [17]. There remains a need for methods leveraging the representational power of foundation models while simultaneously accounting for the ambiguity inherent to medical image segmentation.

Figure 1: Overview of the proposed SARFA framework. A probabilistic prompt generator produces diverse prompt embeddings that enable LoRA-adapted SAM to generate K candidate segmentation masks for each input image. Masked regions from both ground truth and predicted candidates are passed through a radiomic feature extraction pipeline, and their similarity is quantified using FRD. The radiomic distances are used to rank masks, defining preferred and rejected candidates for our DPO loss, which encourages anatomically and texturally consistent medical image segmentation.

A largely unexplored opportunity lies in the use of radiomic information as a supervisory signal during model training. Radiomic features are quantitative descriptors that characterize image intensity distributions, texture patterns, morphological properties, and higher-order image characteristics that may not be readily apparent to human observers [18]. Because these features capture information beyond simple spatial overlap, they offer a potentially valuable representation for comparing alternative segmentation hypotheses. Specifically, we hypothesize that training a model to minimize differences in ground truth and predicted radiomic feature distributions can lead to more accurate segmentation compared to methods that rely solely on maximizing pixel-level overlap. To this end, we introduce SARFA (Segment Anything with Radiomic Feature Alignment), whose architecture is seen in Fig. 1. Our specific contributions can be summarized as follows:

  • We address annotation unavailability and ambiguity by leveraging SAM’s multimask decoding mechanism to generate multiple plausible segmentation candidates;

  • Integration of radiomics-based supervision into the segmentation pipeline using PyRadiomics [19]-based extraction and calculation of the Fréchet Radiomic Distance (FRD) [20];

  • Optimization of SAM predictions using a novel mask ranking system and direct preference optimization (DPO) [21]-style loss function;

  • Validation on both computed tomography (CT) scans and magnetic resonance imaging (MRI) demonstrates the superiority of SARFA over existing state-of-the-art (SOTA) baselines. This also indicates that radiomic feature alignment through the minimization of FRD and utilization of DPO-style candidate mask ranking can be an effective training objective.

2 Related Work↩︎

2.1 Ambiguous Medical Image Segmentation↩︎

There is an inherent level of ambiguity that exists within medical image segmentation problems that is not adequately addressed by deterministic segmentation methods. Because of this, several research pathways have been developed over the years in an effort to solve this problem. Early work was dominated by conditional variational autoencoder (cVAE)-based approaches, with Probabilistic U-Net [10] first demonstrating that multiple coherent segmentation hypotheses could be generated from a learned latent space, while later methods such as the Hierarchical Probabilistic U-Net (HPU-Net) [11] and PHiSeg [6] introduced hierarchical latent representations to better capture ambiguity across spatial scales.

Subsequent research explored alternative uncertainty modeling strategies, including correlated probabilistic representations [12], adversarial refinement [9], autoregressive generation [14], diffusion models [4], and mixtures of stochastic experts [7]. Collectively, these methods improved the diversity, calibration, and realism of generated segmentation hypotheses while highlighting the importance of modeling ambiguity as a structured distribution rather than as independent pixel-wise uncertainty.

More recent works have explored the capabilities of foundation models. SAMed [15] demonstrated that SAM can be effectively adapted for medical image segmentation, while P\(^2\)SAM [8] and Probabilistic SAM [13] extend SAM to ambiguity-aware segmentation by introducing probabilistic prompting and latent-variable modeling, respectively.

2.2 Prompting the Segment Anything Model↩︎

A defining characteristic of SAM is its reliance on prompts. While this enables great flexibility, it poses a problem for medical imaging tasks, where accurate prompting would require expert knowledge. As a result, a growing body of research has focused on reducing or eliminating the need for manual prompting. One such approach, employed by techniques such as AutoSAM [22], MaskSAM [23], SAM-SP [24], and Sam2Rad [25], relies on the learning or prediction of prompts directly from image representations, greatly enhancing the autonomy of SAM-based segmentation frameworks. Other methods rely on the automatic generation of prompts from auxiliary task information [26], [27].

From the last method, the question naturally arises: how to best utilize such preference data within a SAM-based segmentation pipeline. This is largely under-explored in the literature, although there have been some efforts. Konwer et al. [28] combined automated prompting with DPO, demonstrating that segmentation models can be improved through ranked candidate outputs rather than explicit reward functions. While effective, this approach relied on a complex pipeline comprising multiple external networks and components, leaving room for efficiency improvements.

2.3 Radiomics Information as a Supervisory Signal↩︎

Recognizing the limitations of segmentation methods that rely solely on pixel-level overlap for assessment of segmentation quality, there exists a growing body of work that focuses on leveraging radiomic features within deep learning pipelines. One of the prominent methodologies towards this end is multi-task/dual-branch architectures. For example, Hambarde et al. [29] map semantic segmentations alongside statistical radiomic representations to ensure boundaries enclose distinct microstructures. Similar joint optimization approaches fuse radiomic features with deep semantic embeddings to supervise complex clinical downstream prognostic tasks, such as glioma sub-volume segmentation and breast cancer neoadjuvant chemotherapy response prediction [30], [31]. Such approaches force models to preserve clinical shape, contrast, and phenotypic tumor signatures, yielding superior classification calibration over networks optimized purely on standard imaging features [32], [33].

Alternatively, radiomics can be applied directly via auxiliary regularization and alignment loss functions during backpropagation to enhance robustness. Instead of training separate prediction branches, these networks introduce structural loss functions that penalize discrepancies between the ground-truth and predicted regions’ sub-visual gray-level textures. For example, Yang et al. [34] implement a radiomics-based support vector machine (SVM) alongside a Siamese network with contrastive and cross-entropy losses to evaluate the geometry and texture of 3D volumes, effectively isolating true lesions from false positives. Similarly, Bhattacharya et al. [35] introduced a novel radiomics-visual attention loss (RVAL) to enforce consistency across a shared, domain-invariant radiomic metric space. This effectively calculates the distance between predicted and reference feature distributions, regularizing spatial focus and making the network less sensitive to scanner-specific variations.

This quantitative feature encoding and structural relationship modeling has seen expanding utility across modern architectures, including graph neural networks for lung CT profiling [36], myocardial infarction mapping on cine-CMR [37], adaptable Gabor and Laplacian of Gaussian filtered transformers for abdominal organ segmentation [38], pelvic adnexal mass ultrasound classification [39], global-to-voxel parametric map injections for pancreatic lesions [40], and adversarial networks for volumetric lesion generation [41]. However, existing methods typically apply these signatures post hoc or within rigid, deterministic architectures. To the best of our knowledge, SARFA is the first framework to leverage radiomics-driven supervision within a foundation model architecture like SAM.

3 Methods↩︎

Let \(\mathcal{D} = \{ (x_i, y_i) \}_{i=1}^{N}\) denote a dataset of medical images, \(x_i \in \mathbb{R}^{H \times W}\), and corresponding ground-truth segmentation masks, \(y_i \in \{0,1\}^{H \times W}\). Our goal is to learn a segmentation model, \(G_{\theta}\), parameterized by \(\theta\), such that \(G_{\theta}(x_i) \rightarrow \hat{y}_i\), where \(\hat{y}_i\) approximates \(y_i\) while preserving clinically relevant radiomic characteristics.

3.1 Probabilistic Prompt Generation↩︎

Recent work in fine-tuning SAM for medical image segmentation, particularly in low-label or ambiguously labeled settings, has explored various methods to optimize prompt generation. One such method lies in the generation of multiple segmentation candidates by thresholding mask probabilities and subsequently applying preference optimization over these proposals [28]. While effective, this strategy constructs candidates post hoc from a single deterministic mask distribution. An alternative method relies on the introduction of a prior probabilistic space for prompts [8]. This allows for the generation of “one-to-many” segmentation mappings by sampling from a learned prompt distribution. In this work, we modify this probabilistic prompting strategy.

The SAM architecture consists of an image decoder, \(\text{Enc}_I\), a prompt encoder, \(\text{Enc}_P\), and a mask decoder, \(\text{Dec}_M\). From \(\text{Enc}_I\), we obtain:

\[F_I = \text{Enc}_I(x),\]

where \(F_I \in \mathbb{R}^{h \times w \times c}\) is the image embedding. Simultaneously, \(\text{Dec}_M\) produces \(K\) mask tokens,

\[\{ \hat{y}^{(k)} \}_{k=1}^{K} \text{Dec}_M(F_I, T_M),\]

where \(T_M\) denotes the set of learned tokens.

Following the probabilistic prompting strategy of \(\text{P}^2\text{SAM}\) [8], we interpret these multiple masks as samples from an implicit segmentation distribution,

\[\tilde{Y} \sim P_{\theta}(\tilde{Y} \mid x),\]

where the distribution \(P_{\theta}\) is induced by the interaction between \(F_I\) and \(T_M\). Unlike \(\text{P}^2\text{SAM}\), which explicitly models a Gaussian prompt embedding distribution and samples prompt embeddings, our approach leverages SAM’s built-in multimask decoding to approximate this sampling process,

\[\mathbb{E}_{\tilde{Y} \sim P_{\theta}(\cdot \mid x)} \left[ \tilde{Y} \right] \approx \frac{1}{K} \sum_{k=1}^{K} \hat{y}^{(k)}.\]

From here, we introduce a lightweight learnable mask weighting mechanism that operates directly over SAM’s multimask outputs. This can be expressed mathematically by allowing \(w_k \in \mathbb{R}\) to denote learnable weights such that:

\[\sum_{k=1}^{K} w_k = 1.\]

The final segmentation prediction in this case can be expressed as,

\[\tilde{y} = \sum_{k=1}^{K} w_k \hat{y}^{(k)}.\]

Such a formulation preserves SAM’s intrinsic ambiguity while allowing the model to learn task-specific scale and structural preferences without explicitly modeling a prompt distribution.

3.2 Radiomic Feature Extraction↩︎

FRD [20] is a task-independent metric for comparing medical image distributions using radiomic features instead of learned natural-image embeddings. In this work, we extract radiomic features from masked regions of the input images, \(x_m\), where \(x_m\) is formed from the element-wise multiplication of \(x\) and \(m\). Here, \(m\) can represent either the ground truth mask or one of the predicted candidate masks. Extracting the features from just the region of the image overlapped by the ground truth/predicted mask ensures that the features correspond specifically to the predicted anatomical structure. The extraction process is handled by a standard PyRadiomics [19] pipeline, and extracted features include first-order statistics, texture features, Shape2D features, and wavelet-transformed feature maps. After extraction, we are left with the following feature vector:

\[\Phi(x,m) \in \mathbb{R}^d.\]

Following the normalization strategy used by Konz et al. [20], we compute dataset-level statistics over the ground truth masks, then z-score normalize each feature vector,

\[\bar{z}=\frac{\Phi(x,m)-\mu}{\sigma}.\]

As our model produces \(K\) candidate masks for each image \(x\), we computed these normalized radiomic features for the ground truth and each candidate mask:

\[z_{\text{gt}} = \Phi(x, y), \qquad z^{(k)} = \Phi(x, \hat{y}^{(k)}).\]

After normalization, we measure the distance between the radiomic features of each candidate mask and the ground truth masks’ radiomic features. We do this by calculating the mean squared distance between them in the feature space:

\[d^{(k)} = \left\| \bar{z}^{(k)} - \bar{z}_{\text{gt}} \right\|_2^2.\]

Once the radiomic distances are known, we select a “preferred” and “rejected” mask, like so:

\[k^+ = \arg\min_{k} d^{(k)}, \qquad k^- = \arg\max_{k} d^{(k)}. \label{eq:preferred95vs95rejected95masks}\tag{1}\]

where \(k^+\), the preferred mask, is the candidate mask that has the lowest radiomic distance to the ground truth radiomic features, and the rejected mask, \(k^-\), has the highest. Both \(k^+\) and \(k^-\) are used in the calculation of our DPO loss, discussed in the next section.

In addition to per-sample mask ranking, we evaluate model performance at the epoch level using FRD. Given radiomic feature sets \(D_{gt}\) and \(D_{pred}\) for the ground truth and predicted masked image regions, we compute their empirical means and covariances:

\[(\mu_{\text{gt}}, \Sigma_{\text{gt}}), \qquad (\mu_{\text{pred}}, \Sigma_{\text{pred}}).\]

The FRD is then:

\[\begin{align} \mathrm{FRD}(\mathcal{D}_{\text{gt}}, \mathcal{D}_{\text{pred}}) = \left\| \mu_{\text{gt}} - \mu_{\text{pred}} \right\|_2^2 \\+ \mathrm{Tr}\!\left( \Sigma_{\text{gt}} + \Sigma_{\text{pred}} - 2\left( \Sigma_{\text{gt}} \Sigma_{\text{pred}} \right)^{1/2} \right), \end{align}\]

which corresponds to the 2-Wasserstein distance between Gaussian approximations of the radiomic feature distributions. In our approach, this epoch-level FRD calculation is used to determine which model checkpoint to save, with the checkpoint with the lowest FRD being the one saved and used for evaluation.

3.3 Direct Preference Optimization↩︎

DPO is an alternative to reinforcement learning from human feedback (RLHF) that eliminates the need for an explicit reward model by directly optimizing a policy from pairwise preference data [21]. Initially proposed for language modeling, recent work has demonstrated the efficacy of DPO when used as a training objective for segmentation models [28]. Here, we utilize DPO to help our proposed SARFA learn to prefer masks that have the lowest radiomic distance.

Following the calculation of the radiomic distances, we are left with \(k^+\) and \(k^-\), the masks with the lowest and highest radiomic distances to the ground truth, respectively. This defines a pairwise preference,

\[\hat{y}^{(k^{+})} \succ \hat{y}^{(k^{-})}.\]

For our DPO implementation, let \(\pi_\theta\) denote the current policy model and \(\pi_{\text{ref}}\) denote a frozen reference model initialized from the same weights. In our implementation, the “policy” corresponds to the IoU prediction head of SAM, which outputs scores over the \(K\) candidate masks. These scores are converted to probabilities using a softmax. So, for a preferred mask \(k^+\) and rejected mask \(k^-\) we define:

\[\log p_\theta(k | x) = \log\text{softmax}(s_\theta(x))_k,\]

where \(s_\theta(x)\) are the IoU logits.

The DPO loss for a single image is:

\[\begin{align} \mathcal{L}_{\text{DPO}} = - \log \sigma \Big( \beta \big[ &\big( \log p_{\theta}(k^{+} \mid x) - \log p_{\theta}(k^{-} \mid x) \big) \\ &- \big( \log p_{\text{ref}}(k^{+} \mid x) - \log p_{\text{ref}}(k^{-} \mid x) \big) \big] \Big). \end{align}\]

where \(\sigma(\cdot)\) is the logistic function and \(\beta\) is a temperature parameter controlling preference strength. The full training loss combines supervised segmentation losses with DPO,

\[\mathcal{L}_\text{total} = \mathcal{L}_\text{sup} + (\lambda \times \mathcal{L}_\text{DPO}),\]

where \(\mathcal{L}_\text{sup}\) includes cross-entropy, Dice, and Focal losses, and \(\lambda\) is a small weight applied to the DPO loss. To ensure stability in the loss calculation, we apply DPO intermittently during training every \(T\) steps.

Figure 2: Progression of loss values for P^2SAM and SARFA when trained for 100 epochs on the LIDC-IDRI. Both models exhibit rapid convergence within the first 10 epochs; hence, the selection of this value for the rest of our experiments.

4 Experimental Evaluation↩︎

4.1 Data↩︎

We validate our proposed SARFA on the Lung Image Database Consortium and Image Database (LIDC-IDRI) [42], which contains lesion annotations collected from four expert radiologists across 1,018 lung CT scans from 1,010 patients. To demonstrate efficacy across imaging modalities, we also evaluate our model on data from the 2017 Brain Tumor Segmentation (BraTS) Challenge [43]. This dataset contains annotations for GD-enhancing tumor, peritumoral edema, and necrotic/non-enhancing tumor across 285 3D MRI images, comprised of 155 slices in four modalities (T1, T1ce, T2, and FLAIR). Following Huang et al.’s method [8], we treat each of the three ground truth masks per slice as a unique ambiguous mask to be passed to SARFA. We follow the train/val/test splits in [8].

Table 1: Comparison of SOTA ambiguous segmentation models on LIDC. The best-performing method across the majority of metrics is highlighted in green.
Method GED(\(\downarrow\)) FRD(\(\downarrow\)) HM-IoU(\(\uparrow\)) \(D_{\max}(\uparrow)\)
Probabilistic U-Net [10] 0.324 0.423 0.370
HPU-Net [11] 0.270 0.530
PHiseg [6] 0.262 0.595
SSN [12] 0.259 0.555
CAR [9] 0.252 0.549 0.732
PixelSeg [14] 0.243 0.614 0.814
CIMD [4] 0.234 0.587
Mose [7] 0.234 0.623 0.702
SAMed [15] 0.380 0.357 0.703
P\(^2\)SAM [8] 0.353 3.648 0.654 0.772
SARFA (ours) 0.206 2.758 0.659 0.774

4.2 Implementation Details↩︎

Baselines: We compared our SARFA against ten state-of-the-art (SOTA) SAM-based and non-SAM baselines: Probabilistic U-Net [10], HPU-Net [11], PHiSeg [6], SSN [12], CAR [9], PixelSeg [14], CIMD [4], Mose [7], SAMed [15], and \(\text{P}^2\text{SAM}\) [8].

Training: We trained for 10 epochs using a batch size of 1, a learning rate of 0.001, and an image size of 128\(\times\)​128 for the generation of 16 candidate masks. Just 10 epochs were used for training after a convergence analysis (shown in Fig. 2) revealed that both SARFA and our strong baseline P\(^2\)SAM converged within the first 10 epochs, with no meaningful gains in performance after this point. For the DPO-specific hyperparameters, we used a \(\beta\) of 0.01, a \(\lambda\) of 0.05, and a \(T\) of 10 to apply DPO every 10 steps. All of these hyperparameters were empirically tuned for optimal model performance.

Machine Configuration: The models are trained on a Intel (R) Xeon (R) w7-2475X, 2600MHz machine with a dual NVIDIA A4000X2 GPUs (32GB).

a

b

c

d

e

f

g

h

i

j

k

l

m

Figure 3: Visual comparison of selected plausible masks generated by both P\(^2\)SAM and SARFA with the ground truth annotations for the LIDC-IDRI dataset. Note the failed segmentation of P\(^2\)SAM’s first mask compared to the successful segmentation of SARFA’s first mask. Additional visualizations are included in the Supplemental Material..

Evaluation: For evaluation, we used generalized energy distance (GED), FRD, Hungarian-matched IoU (HM-IoU), and maximum Dice matching (\(D_{max}\)).

4.3 Results and Discussion↩︎

Results on LIDC-IDRI: Table 1 reports the results of SARFA compared against ten SOTA ambiguous segmentation models. Comparatively, SARFA performs very well, outperforming each of the ten methods. Compared to the strongest baseline in P\(^2\)SAM, our SARFA outperforms it marginally on all metrics, with the biggest improvements being noted on the distance-based metrics GED and FRD, indicating that the plausible masks predicted by SARFA are more closely aligned with the ground truth distribution. Fig. 3 shows a qualitative comparison of the first four masks generated by both P\(^2\)SAM and SARFA. As evidenced by the first of P\(^2\)SAM’s predicted masks, SARFA can accurately which completely fails to properly segment the lung lesion, SARFA is very capable of outperforming P\(^2\)SAM in terms of visual output.

Results on BraTS2017: Table 2 shows the results of our proposed SARFA against both implementations of the strong baseline P\(^2\)SAM on the BraTS2017 dataset. Here, SARFA clearly outperforms P\(^2\)SAM across all metrics, demonstrating the efficacy of SARFA in achieving strong performance across imaging modalities and segmentation tasks. This is further validated by the visual comparison of the segmentation masks produced by SARFA and those produced by P\(^2\)SAM, as shown in Fig. 4.

Hyperparameter Tuning: In addition to investigating whether FRD is a valid training objective, the findings of which we have already discussed, we have also comprehensively evaluated different configurations of the loss functions used by SARFA. The outcomes of these experiments are shown in Table 3. To demonstrate that SARFA is capable of performing with varying values for the \(K\) candidate mask generation, all experiments in Table 3 were performed using \(K = 8\), whereas \(K = 16\) was used in the experiments reported in Tables 1 and 2. The components evaluated were the configuration of the supervised loss, \(\mathcal{L}_\text{sup}\), the temperature of \(\mathcal{L}_\text{DPO}\), \(\beta\), the weight of \(\mathcal{L}_\text{DPO}\), \(\lambda\), and the number of steps, \(T\), that pass before \(\mathcal{L}_\text{DPO}\) is calculated. We find that the configuration of \(\mathcal{L}_\text{sup}\) = \(\mathcal{L}_{\text{Dice}}\) + \(\mathcal{L}_{\text{CE}}\) + \(\mathcal{L}_{\text{Focal}}\), \(\beta = 0.1\), \(\lambda = 0.05\), and \(T = 10\) achieves the best balanced performance across the losses among the examined configurations, thus is the configuration used for our experiments reported in Tables 1 and 2.

Table 2: Comparison of SOTA ambiguous segmentation models on BraTS2017. \(\dagger\) indicates that the best model was saved using the lowest validation loss, while \(\ddagger\) indicates that the lowest FRD was used to select the best model. The best-performing method across the majority of metrics is highlighted in green.
Method GED(\(\downarrow\)) FRD(\(\downarrow\)) HM-IoU(\(\uparrow\)) \(D_{\max}(\uparrow)\)
P\(^2\)SAM (\(\dagger\)) 8.426 9.602 0.342 0.446
P\(^2\)SAM (\(\ddagger\)) 3.979 10.496 0.326 0.427
SARFA (ours) 2.644 6.198 0.358 0.458

a

b

c

d

e

f

g

h

i

j

Figure 4: Visual comparison of selected plausible masks generated by both P\(^2\)SAM and SARFA with the ground truth annotations for the BraTS2017 dataset. Additional visualizations are included in the Supplemental Material..

Ablation Experiments: To further assess the contribution of each component in SARFA, we performed two ablation studies. First, we evaluated whether the proposed radiomics-based ranking strategy is more effective for defining the DPO preference pairs than conventional overlap-based alternatives. Table 4 compares three strategies for selecting the positive and negative masks used in the DPO loss: highest/lowest IoU, highest/lowest Dice score, and lowest/highest FRD. Using FRD to construct the preference pairs achieves the strongest overall performance, with the lowest GED and FRD and the highest HM-IoU and \(D_{\max}\) among the evaluated ranking strategies. In particular, FRD-based ranking improves GED from 0.290 to 0.194 compared with Dice-based ranking and from 4.209 to 0.194 compared with IoU-based ranking. It also produces the lowest FRD score, reducing the distance from 3.410 with Dice ranking and 5.513 with IoU ranking to 2.818. These results indicate that selecting preferred masks based on radiomic similarity provides a more effective optimization signal than selecting them using only pixel-level overlap.

Figure 5: Correlation analysis between FRD and established segmentation quality metrics across all valid candidate masks generated during training (n = 80). Each point represents an individual candidate mask and solid lines denote least-squares regression fits. Lower FRD values are associated with lower GED and higher Dice and HM-IoU scores, indicating that radiomic similarity remains strongly aligned with conventional overlap- and distribution-based segmentation metrics.
Table 3: Segmentation performance under different configurations of the supervised and DPO loss components of SARFA by removing each of our proposed components. The best loss configuration is highlighted in green.
\(\mathcal{L}_{\text{sup}}\) Configuration \(\beta\) \(\lambda\) \(T\) GED(\(\downarrow\)) FRD(\(\downarrow\)) HM-IoU(\(\uparrow\)) \(D_{\max}(\uparrow)\)
\(\mathcal{L}_{\text{Dice}} + \mathcal{L}_{\text{CE}} + \mathcal{L}_{\text{Focal}}\) 0.1 0.05 5 0.321 3.087 0.578 0.704
0.217 2.780 0.559 0.690
10 0.194 2.818 0.602 0.728
15 6.502 5.632 0.225 0.331
0.323 3.714 0.493 0.629
\(\mathcal{L}_{\text{Dice}}\) + \(\mathcal{L}_{\text{CE}}\) + \(\mathcal{L}_{\text{Focal}}\) 0.1 0.10 10 0.252 2.903 0.600 0.722
0.01 0.341 3.275 0.579 0.706
\(\mathcal{L}_{\text{Dice}}\) + \(\mathcal{L}_{\text{CE}}\) + \(\mathcal{L}_{\text{Focal}}\) 0.2 0.05 10 0.064 2.543 0.591 0.716
\(\mathcal{L}_{\text{Dice}}\) + \(\mathcal{L}_{\text{CE}}\) 0.1 0.05 10 923.046 13.133 0.008 0.015
\(\mathcal{L}_{\text{Dice}}\) + \(\mathcal{L}_{\text{Focal}}\) 0.166 2.848 0.528 0.663
\(\mathcal{L}_{\text{CE}}\) + \(\mathcal{L}_{\text{Focal}}\) 1.316 4.307 0.423 0.563
\(\mathcal{L}_{\text{Focal}}\) 0.106 3.224 0.547 0.678
\(\mathcal{L}_{\text{CE}}\) 1.655 5.192 0.367 0.498
No \(\mathcal{L}_\text{sup}\) 38.274 8.853 0.035 0.056
Table 4: Segmentation performance when the positive and negative pairs for the DPO loss are selected based on the highest and lowest IoU, the highest and lowest Dice score, and the lowest and highest FRD. The best ranking strategy is highlighted in green.
Ranking Strategy GED(\(\downarrow\)) FRD(\(\downarrow\)) HM-IoU(\(\uparrow\)) \(D_{\max}(\uparrow)\)
IoU 4.209 5.513 0.301 0.402
Dice 0.290 3.410 0.543 0.673
FRD 0.194 2.818 0.602 0.728
Table 5: Ablation experiment on the architectural design of SARFA, where components are removed from the network in the order listed. The best-performing architecture is highlighted in green.
Architecture GED(\(\downarrow\)) FRD(\(\downarrow\)) HM-IoU(\(\uparrow\)) \(D_{\max}(\uparrow)\)
SARFA 0.194 2.818 0.602 0.728
– DPO 5.799 5.987 0.355 0.476
– FRD Ranking 2.358 5.472 0.470 0.590
– Mask Weighting 0.557 4.593 0.527 0.656

Second, we evaluated the architectural contributions of the main SARFA components, as shown in Table 5. Removing the DPO loss leads to the largest degradation in performance, increasing GED from 0.194 to 5.799 and FRD from 2.818 to 5.987, while reducing HM-IoU from 0.602 to 0.355 and \(D_{\max}\) from 0.728 to 0.476. This demonstrates that the DPO objective is essential for learning from ranked candidate masks rather than treating all multimask outputs equally. Removing FRD-based ranking also substantially harms performance, increasing GED to 2.358 and FRD to 5.472, which confirms that the radiomics-driven preference construction is a key component of the proposed framework. Finally, removing the mask weighting mechanism degrades performance across all metrics, though less severely than removing DPO or FRD ranking, with GED increasing to 0.557 and FRD increasing to 4.593. This suggests that mask weighting helps SARFA better combine SAM’s candidate outputs, but that its benefit is strongest when paired with radiomics-guided DPO optimization.

Radiomic Alignment as a Segmentation Objective: A central hypothesis of this work was that radiomic similarity between predicted and ground-truth regions could serve as a meaningful optimization target for medical image segmentation. While FRD was proposed as a dataset-level evaluation metric, its utility as a training objective had not been explored. Our results provide evidence that radiomic alignment is not only measurable but also strongly associated with conventional indicators of segmentation quality.

Fig. 5 demonstrates a strong relationship between FRD and three established segmentation metrics. Specifically, lower FRD values are associated with lower GED scores and higher HM-IoU and \(D_{\max}\) values, with both Pearson (\(r\)) and Spearman (\(\rho\)) correlations exceeding 0.90 in magnitude across all comparisons. All observed correlations were determined to be statistically significant (\(p < 0.001\)), providing strong evidence against the null hypothesis that FRD and the corresponding segmentation metrics are unrelated. This suggests that candidate masks that more closely preserve the radiomic characteristics of the ground-truth region also tend to exhibit stronger spatial agreement with expert annotations. Importantly, this relationship is monotonic as well as linear, indicating that improvements in radiomic similarity consistently correspond to improvements in segmentation quality over a wide range of masks.

Limitations: Despite the promising results achieved by SARFA, several limitations remain. First, radiomic feature extraction introduces additional computational overhead during training. Large-scale training on high-resolution datasets may require more efficient feature extraction pipelines or approximations of FRD. Second, the current implementation relies on radiomic descriptors extracted using PyRadiomics—alternative representations may capture complementary information and warrant investigation. Another limitation is that the evaluation was conducted on two datasets representing CT and MRI modalities. Furthermore, the current framework constructs preference pairs using only the most preferred and least preferred masks from each candidate set. More sophisticated ranking strategies that exploit the full ordering of candidate masks may provide a richer optimization signal and further improve performance.

5 Conclusions↩︎

In this paper, we introduced SARFA, a novel probabilistic segmentation model for medical image segmentation that learns during training to produce a distribution of \(K\) plausible masks that are radiomically similar to the ground truth distribution. We hypothesized that predicted masks that have similar radiomic features to the ground truth would be better matched anatomically, something that we demonstrated through comprehensive experimentation, demonstrating that the minimization of FRD is a valid training objective for medical segmentation models. Additionally, we demonstrate that incorporating a loss function based on DPO in combination with a standard supervised loss helps learn this training objective. Future work will consist of scaling this architecture to different imaging modalities and segmentation tasks, as well as gauging its performance when additional multimodal information, such as patient metadata, is provided to the model during training.

6 SARFA Training Algorithm↩︎

A formal algorithm showing the training procedure of SARFA is shown in Algorithm [alg:sarfa].

Dataset \(\mathcal{D}\), SAM checkpoint \(C\), epochs \(E\), learning rate \(\eta\), DPO interval \(K\), DPO weight \(\lambda\) Trained policy model \(\pi_\theta\)

\(\pi_\theta \gets\) initialize LoRA-SAM from checkpoint \(C\) \(\pi_{\mathrm{ref}} \gets\) frozen copy of \(\pi_\theta\) \(\mathbf{w} \gets\) initialize learnable mask weights \(\phi \gets\) initialize radiomics extractor \((\mu, \sigma) \gets\) radiomics statistics from ground-truth masks in \(\mathcal{D}\)

set \(\pi_\theta\) and \(\mathbf{w}\) to train mode

\((M, s_\theta) \gets \pi_\theta(x)\) \(\hat{m}_{w} \gets\) weighted combination of \(M\) using \(\mathbf{w}\) \(L_{\mathrm{sup}} \gets\) segmentation loss using \(M\), \(\hat{m}_{w}\), and \(y\) \(L_{\mathrm{DPO}} \gets 0\)

\(z_{\mathrm{gt}} \gets\) normalized radiomics features of \((x, y)\)

\(z_j \gets\) normalized radiomics features of \((x, \hat{m}_j)\) \(d_j \gets\) distance between \(z_j\) and \(z_{\mathrm{gt}}\)

\(c \gets\) candidate with smallest \(d_j\) \(r \gets\) candidate with largest \(d_j\)

\(s_{\mathrm{ref}} \gets \pi_{\mathrm{ref}}(x)\) \(\Delta_\theta \gets\) policy score margin between \(c\) and \(r\) using \(s_\theta\) \(\Delta_{\mathrm{ref}} \gets\) reference score margin between \(c\) and \(r\) using \(s_{\mathrm{ref}}\) \(L_{\mathrm{DPO}} \gets\) DPO loss from \(\Delta_\theta\), \(\Delta_{\mathrm{ref}}\)

\(L \gets L_{\mathrm{sup}} + \lambda L_{\mathrm{DPO}}\) update \(\pi_\theta\) and \(\mathbf{w}\) using \(L\)

\(\mathrm{FRD} \gets\) evaluate radiomics distance on epoch predictions save checkpoint

\(\pi_\theta\)

7 Additional Visualizations↩︎

Figs. 6 and 8 show extended visualizations of the samples used for Figs. 3 and 4 of the main text up to eight output masks. Note that while both P\(^2\)SAM and SARFA are probabilistic methods that can produce multiple mask variants, P\(^2\)SAM occasionally fails to segment samples (instead segmenting the background), something that is not observed in our SARFA visualizations. Visualizations for additional samples of the LIDC-IDRI and BraTS datasets are shown in Figs. 7 and 9.

Figure 6: image.

Figure 7: image.

Figure 8: image.

Figure 9: image.

References↩︎

[1]
Y. Gao, Y. Jiang, Y. Peng, F. Yuan, X. Zhang, and J. Wang, “Medical image segmentation: A comprehensive review of deep learning-based methods,” Tomography, vol. 11, no. 5, p. 52, 2025.
[2]
M. S. Fasihi and W. B. Mikhael, “Overview of current biomedical image segmentation methods,” in 2016 international conference on computational science and computational intelligence (CSCI), 2016, pp. 803–808.
[3]
S. Xi, S. Wang, M. Safari, M. Hu, Z. Tian, and X. Yang, “Uncertainty as a foundation for trustworthy medical imaging AI: A comprehensive review,” Authorea Preprints, 2025.
[4]
A. Rahman, J. M. J. Valanarasu, I. Hacihaliloglu, and V. M. Patel, “Ambiguous medical image segmentation using diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 11536–11546.
[5]
H. Zhou, L. Xu, C. Liu, and G. Li, UDNDSNet: A unified deterministic and non-deterministic segmentation network for multi-scene medical image analysis,” Knowledge-Based Systems, p. 114641, 2025.
[6]
C. F. Baumgartner et al., PHiSeg: Capturing uncertainty in medical image segmentation,” in International conference on medical image computing and computer-assisted intervention, 2019, pp. 119–127.
[7]
Z. Gao, Y. Chen, C. Zhang, and X. He, “Modeling multimodal aleatoric uncertainty in segmentation with mixture of stochastic experts,” in The eleventh international conference on learning representations, 2023.
[8]
Y. Huang et al., \(\text{P}^2\text{SAM}\): Probabilistically prompted SAMs are efficient segmentator for ambiguous medical images,” in Proceedings of the 32nd ACM international conference on multimedia, 2024, pp. 9779–9788.
[9]
E. Kassapis, G. Dikov, D. K. Gupta, and C. Nugteren, “Calibrated adversarial refinement for stochastic semantic segmentation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 7057–7067.
[10]
S. Kohl et al., “A probabilistic u-net for segmentation of ambiguous images,” Advances in Neural Information Processing Systems, vol. 31, 2018.
[11]
S. A. Kohl et al., “A hierarchical probabilistic u-net for modeling multi-scale ambiguities,” arXiv preprint arXiv:1905.13077, 2019.
[12]
M. Monteiro et al., “Stochastic segmentation networks: Modelling spatially correlated aleatoric uncertainty,” Advances in Neural Information Processing Systems, vol. 33, pp. 12756–12767, 2020.
[13]
T. Ward and A. Imran, “A probabilistic Segment Anything Model for ambiguity-aware medical image segmentation,” in Medical imaging 2026: Imaging informatics, 2026, vol. 13930, pp. 7–12.
[14]
W. Zhang, X. Zhang, S. Huang, Y. Lu, and K. Wang, PixelSeg: Pixel-by-pixel stochastic semantic segmentation for ambiguous medical images,” in Proceedings of the 30th ACM international conference on multimedia, 2022, pp. 4742–4750.
[15]
K. Zhang and D. Liu, “Customized Segment Anything Model for medical image segmentation,” arXiv preprint arXiv:2304.13785, 2023.
[16]
A. Kirillov et al., “Segment anything,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026.
[17]
J. Ma, Y. He, F. Li, L. Han, C. You, and B. Wang, “Segment anything in medical images,” Nature Communications, vol. 15, no. 1, p. 654, 2024.
[18]
M. E. Mayerhoefer et al., “Introduction to radiomics,” Journal of Nuclear Medicine, vol. 61, no. 4, pp. 488–495, 2020.
[19]
J. J. Van Griethuysen et al., “Computational radiomics system to decode the radiographic phenotype,” Cancer Research, vol. 77, no. 21, pp. e104–e107, 2017.
[20]
N. Konz et al., Fr\(\backslash\)’echet Radiomic Distance (FRD): A versatile metric for comparing medical imaging datasets,” arXiv preprint arXiv:2412.01496, 2024.
[21]
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems, vol. 36, pp. 53728–53741, 2023.
[22]
T. Shaharabany, A. Dahan, R. Giryes, and L. Wolf, AutoSAM: Adapting SAM to medical images by overloading the prompt encoder,” arXiv preprint arXiv:2306.06370, 2023.
[23]
B. Xie, H. Tang, B. Duan, D. Cai, Y. Yan, and G. Agam, MaskSAM: Auto-prompt SAM with mask classification for volumetric medical image segmentation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2025, pp. 24423–24433.
[24]
C. Zhou, K. Ning, Q. Shen, S. Zhou, Z. Yu, and H. Wang, SAM-SP: Self-prompting makes SAM great again,” arXiv preprint arXiv:2408.12364, 2024.
[25]
A. S. Wahd et al., Sam2Rad: A segmentation model for medical images with learnable prompts,” Computers in Biology and Medicine, vol. 187, p. 109725, 2025.
[26]
T. Ward and A. A. Z. Imran, “Annotation-efficient task guidance for medical Segment Anything,” in 2025 IEEE 22nd international symposium on biomedical imaging (ISBI), 2025, pp. 1–4.
[27]
T. Ward, M. K. Owen, O. Coleman, B. Noehren, and A.-A.-Z. Imran, “Autoadaptive medical Segment Anything Model,” arXiv preprint arXiv:2507.01828, 2025.
[28]
A. Konwer et al., “Enhancing SAM with efficient prompting and preference optimization for semi-supervised medical image segmentation,” in Proceedings of the computer vision and pattern recognition conference, 2025, pp. 20990–21000.
[29]
P. Hambarde, S. Talbar, A. Mahajan, S. Chavan, M. Thakur, and N. Sable, “Prostate lesion segmentation in MR images using radiomics based deeply supervised U-Net,” Biocybernetics and Biomedical Engineering, vol. 40, no. 4, pp. 1421–1435, 2020.
[30]
Y. Chen et al., “A radiomics-incorporated deep ensemble learning model for multi-parametric MRI-based glioma segmentation,” Physics in Medicine & Biology, vol. 68, no. 18, p. 185025, 2023.
[31]
X. Hao, H. Xu, N. Zhao, T. Yu, T. Hamalainen, and F. Cong, “Predicting pathological complete response based on weakly and semi-supervised joint learning from breast cancer MRI,” in 2023 45th annual international conference of the IEEE engineering in medicine & biology society (EMBC), 2023, pp. 1–4.
[32]
T.-W. Tang, W.-Y. Lin, J.-D. Liang, and K.-M. Li, “Artificial intelligence aided diagnosis of pulmonary nodules segmentation and feature extraction,” Clinical Radiology, vol. 78, no. 6, pp. 437–443, 2023.
[33]
G. Lekkas, E. Vrochidou, and G. A. Papakostas, “Deep-radiomic fusion for early detection of pancreatic ductal adenocarcinoma,” Applied Sciences, vol. 15, no. 24, p. 13024, 2025.
[34]
Z. Yang et al., “Deep-learning and radiomics ensemble classifier for false positive reduction in brain metastases segmentation,” Physics in Medicine & Biology, vol. 67, no. 2, p. 025004, 2022.
[35]
M. Bhattacharya, S. Jain, and P. Prasanna, GazeRadar: A gaze and radiomics-guided disease localization framework,” in International conference on medical image computing and computer-assisted intervention, 2022, pp. 686–696.
[36]
M. K. Faizi et al., “Graph neural network model using radiomics for lung CT image segmentation,” Scientific Reports, vol. 15, no. 1, p. 34148, 2025.
[37]
W. Xu and X. Shi, “Integrating radiomic texture analysis and deep learning for automated myocardial infarction detection in cine-MRI,” Scientific Reports, vol. 15, no. 1, p. 24365, 2025.
[38]
N. E. Zarch, H. Bagher-Ebadian, T. Alhanai, and M. M. Ghassemi, “GLoG-CSUnet: Enhancing vision transformers with adaptable radiomic features for medical image segmentation,” in ICASSP 2025-2025 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2025, pp. 1–5.
[39]
J. F. Barcroft et al., “Machine learning and radiomics for segmentation and classification of adnexal masses on ultrasound,” NPJ Precision Oncology, vol. 8, no. 1, p. 41, 2024.
[40]
Z. Deng et al., “From global radiomics to parametric maps: A unified workflow fusing radiomics and deep learning for PDAC detection,” in 2026 IEEE 23rd international symposium on biomedical imaging (ISBI), 2026, pp. 1–5.
[41]
J. Li, S. Pan, X. Zhang, C. T. Lin, J. W. Stayman, and G. J. Gang, “Generative adversarial networks with radiomics supervision for lung lesion generation,” IEEE Transactions on Biomedical Engineering, vol. 72, no. 1, pp. 286–296, 2024.
[42]
S. G. Armato III et al., “The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A completed reference database of lung nodules on CT scans,” Medical Physics, vol. 38, no. 2, pp. 915–931, 2011.
[43]
B. H. Menze et al., “The multimodal brain tumor image segmentation benchmark (BRATS),” IEEE Transactions on Medical Imaging, vol. 34, no. 10, pp. 1993–2024, 2014.