May 28, 2026
Large Language Models (LLMs) are increasingly used in clinical applications; however, their behavior remains highly sensitive to subtle linguistic variations, such as rephrasing or syntactic variation. This sensitivity poses risks in safety-critical healthcare settings, where semantically equivalent inputs should produce consistent predictions. However, a key challenge is to ensure that prompt variations truly preserve clinical meaning, as embedding-based similarity metrics often fail to capture distinctions involving negation, temporality, or severity. To address this limitation, we propose a semantic verification framework based on Natural Language Inference (NLI) to filter meaning-preserving prompt variations, which are further refined using an LLM-as-a-judge and audited by a clinical expert. In addition, we introduce three metrics to quantify model sensitivity: Meaning-Preserving Variation Sensitivity (MVS), confidence variation (\(\Delta\)C), and Worst-Case Instability (WCI). We evaluate 16 open-source general-purpose (GP) and medical LLMs within the same model families and parameter scales, using reformulated prompts derived from the DiagnosisQA and MedQA datasets. Our results demonstrate that robustness differences between domain-specific (DS) models are mixed and highly model-dependent, i.e., domain specialization does not consistently improve or reduce robustness to meaning-preserving prompt reformulations. Several DS models rank among the most robust (when compared with GP counterparts), and strong GP baselines remain competitive as well. We also find that confidence variation is only weakly associated with prediction instability, while overall confidence is only weakly aligned with model accuracy; many models remain overconfident despite reduced robustness or relatively low performance. This suggests that medical fine-tuning alone does not reliably guarantee stable behavior under natural prompt variations, thereby highlighting the need for explicit robustness evaluation prior to clinical deployment.
Large Language Models, Robustness Evaluation, Sensitivity Analysis, Semantic Stability, Healthcare AI.
Over the past few years, Large Language Models (LLMs) have emerged as transformative technologies that could streamline clinical workflows and enable flexible language-based interfaces to help clinicians navigate increasingly complex information landscapes. These models can support diverse tasks, including patient history summarization, discharge summary generation, medication reconciliation, and enhancing clinical diagnostics [1], [2]. As a result, LLMs are increasingly positioned not merely as auxiliary tools but as intermediaries between clinicians and the rapidly expanding ecosystem of structured records, unstructured clinical narratives, and evidence-based resources.
Despite their strong performance on standardized benchmarks [3], [4], recent studies reveal a critical limitation: medical LLMs exhibit significant sensitivity to minor linguistic variations in input prompts [2], [5]. Such variations include lexical substitutions (e.g., myocardial infarction vs. heart attack), syntactic reordering, abbreviation practices, and contextual rephrasing. Ideally, clinically reliable models should exhibit semantic invariance, producing consistent outputs across meaning-preserving formulations of the same clinical scenario. However, empirical evidence suggests that performance can vary significantly under such variations, with reported deviations ranging from 8–50% in decision-making tasks [6] and up to 45% across semantically equivalent prompt formulations [7]. This instability poses risks to patient safety, undermines clinician trust, and complicates the deployment of LLMs in safety-critical settings.
max width=
Evaluating semantic stability requires addressing a fundamental challenge, i.e., reliably determining whether a reformulated prompt preserves the original clinical meaning. Without reliable semantic verification, it becomes impossible to distinguish genuine model instability from legitimate variation in model outputs. Existing approaches based on embedding similarity, such as cosine similarity over sentence embeddings [8] and domain-adapted metrics such as MedSim [9], are demonstrably insufficient for clinical semantic verification [10], [11]. These methods often fail to capture clinically critical distinctions, such as negation, temporality, or severity. For example, patient has chest pain vs. patient denies chest pain may receive high similarity scores despite opposite clinical meanings.
This limitation was systematically demonstrated by Morris et al. [12], who reported that 10% \(\sim\) 15% of automated perturbations failed to preserve semantics even when constrained by high similarity thresholds, with 45% introducing grammatical errors. However, even such stricter thresholds can be insufficient for clinical contexts. This leads to a methodological dilemma: to evaluate whether a clinical LLM responds consistently to meaning-preserving variations, we must first reliably identify which variations actually preserve meaning, a task for which embedding-based metrics are demonstrably inadequate. Therefore, without principled semantic verification, sensitivity analysis becomes unreliable, potentially conflating genuine instability with legitimate responses to semantic differences.
To address this limitation, we propose a Natural Language Inference (NLI)-based semantic verification framework that formulates meaning preservation as a logical entailment problem. NLI models, particularly those adapted to clinical domains (e.g., MedNLI), can explicitly determine whether a reformulated prompt entails, contradicts, or is neutral with respect to the original. This framing provides a principled, interpretable mechanism for semantic verification: only variations that achieve bidirectional entailment (mutual implication) are retained for sensitivity evaluation. This ensures that the observed output variations reflect the model’s genuine instability rather than legitimate responses to altered clinical content. Specifically, we attempt to answer the following key questions:
RQ1: Do medical LLMs outperform general-purpose models on standard medical QA benchmarks under original prompt settings (i.e., without any natural variations)?
RQ2: To what extent does domain specialization improve prediction robustness to semantically equivalent prompt reformulations compared to general-purpose models?
RQ3: How stable are model confidence estimates under meaning-preserving prompt reformulations, and does confidence stability align with prediction stability?
To address aforementioned RQs, we made the following salient contributions in this paper.
We introduce an NLI-based semantic verification framework that formalizes meaning preservation as a bidirectional logical entailment, enabling sensitivity evaluation under rigorously validated semantic equivalence.
We propose three complementary metrics: Meaning-Preserving Variation Sensitivity (MVS), Confidence Variation (\(\Delta \text{C}\)), and worst-case instability, which disentangle prediction variability, confidence shifts, and per-sample correctness fragility.
We develop a semantically and human-verified benchmark of meaning-preserving prompt variations derived from two medical QA datasets and evaluate the sensitivity of both general-purpose and domain-specific LLMs.
Table [tab:related95work] summarizes existing related works across four related areas: (1) clinical LLM evaluation and prompt sensitivity; (2) semantic-preserving perturbations; (3) paraphrase-based consistency measurement; and (4) NLI-based semantic verification. Below, we highlight key gaps relevant to our approach.
Prompt sensitivity—variation in model outputs due to minor input changes has been widely studied in the literature. For instance, in clinical settings, Hager et al. [6] demonstrated that clinical LLMs exhibit 8–50% performance variance due to phrasing and 5–18% variance from information ordering. Similarly, Kim et al. [13] found that even the best performing LLM achieved only 52% accuracy versus 66% for physicians on adversarial clinical reasoning tasks, exhibiting hallucinations and overconfidence. On a similar note, studies focused on general-domain LLMs also reported similar instability. Cao et al. [7] demonstrated performance swings of up to 45% across semantically equivalent prompt formulations. Similarly, He et al. [14] showed that prompt formatting alone can induce performance variations of up to 40%. In a recent study, Yun et al. [15] demonstrated that variations in question framing (e.g., positive vs. negative) can lead to contradictory model responses even when grounded in identical clinical evidence in a controlled retrieval-augmented generation (RAG) setting.
To characterize and quantify this variability, Errica et al. [16] introduced sensitivity and consistency measures for prompt engineering, while Zhuo et al. [17] proposed the PromptSensiScore (PROSA) framework using decoding confidence to quantify response variability. In [18], the authors evaluated 10 large multimodal models across 61 prompt types, finding accuracy deviations up to 15% from prompt variations. Critically, they classified prompts by instructional intent (positive/neutral/negative framing) rather than by semantic equivalence, treating prompt variability as a feature of prompt engineering rather than as a requirement for consistency. This distinction is central to our work: we focus on unintended sensitivity to meaning-preserving variations—cases where models should remain invariant but do not.
Early work on semantic-preserving perturbations includes Semantically Equivalent Adversarial Rules (SEARs) [19], which use universal replacement rules (e.g., “What NOUN \(\rightarrow\) Which NOUN”) verified through back-translation and human evaluation. Jin et al. [20] introduced a black-box adversarial framework using synonym substitution with word importance ranking, verified by counter-fitted embeddings (cosine similarity \(>0.84\)) and Universal Sentence Encoder. However, human evaluation revealed that 97% preserved semantics, indicating that 3% still changed meaning, which highlights verification limitations. However, empirical analyses show that such constraints are insufficient; for instance, Morris et al. [12] found that 38% of automatically generated perturbations alter meaning despite high similarity scores. This finding directly motivates our choice of NLI-based verification over similarity metrics. Similarly, behavioral testing frameworks such as CheckList [21] and AdaTest [22] can systematically expose model limitations. However, they rely on templates or heuristic controls rather than formal semantic verification. As a result, they cannot guarantee that perturbations preserve meaning, limiting their applicability for consistency evaluation.
Several frameworks have been proposed to quantify consistency under meaning-preserving variations (i.e., paraphrases). Elazar et al. [23] evaluated invariance across manually curated paraphrases (with 95.5% inter-annotator agreement) and reported high variability in model predictions despite preserved meaning. In a similar study [24], the Paraphrastic Robustness in Textual Entailment (PaRTE) framework is presented that employs T5-based paraphrasers and measures changes in prediction rates. They found that bag-of-words and BiLSTM models showed 15–17% changes in predictions, while transformer models were more robust but still changed predictions. While PaRTE provides a consistency measurement similar to our MVS metric, it lacks a semantic verification step to ensure perturbations genuinely preserve meaning.
Beyond single-prompt sensitivity, recent work examines consistency across semantically equivalent prompt formulations. Cao et al. [7] introduced RobustAlpacaEval, a benchmark containing 10 paraphrases per query in AlpacaEval, generated with GPT-4 and manually reviewed for semantic integrity. They demonstrated performance swings up to 45% across semantically equivalent prompts, with worst-case performance often dramatically lower than best-case performance. This worst-case analysis reveals vulnerabilities obscured by single-prompt or average-case evaluation. Mizrahi et al. [25] advocated for multi-prompt evaluation, demonstrating that single-prompt assessments provide unstable estimates of model capabilities. Sclar et al. [26] introduced FormatSpread to quantify sensitivity to spurious features in prompt design and showed that formatting alone substantially affects performance. Sun et al. [27] evaluated the zero-shot robustness of instruction-tuned models and found that performance varies significantly across semantically equivalent prompt reformulations.
NLI provides a principled approach to semantic verification by framing similarity as logical entailment. Prior work has applied NLI to evaluate generation quality, detect factual inconsistencies, and enforce output-level consistency. Chen and Eger [28] developed MENLI, using bidirectional entailment checking for text generation quality assessment. They demonstrated that NLI metrics are 15–30% more robust to adversarial attacks than BERT-based metrics and better at detecting factual inconsistencies in machine translation and summarization. Laban et al. [29] introduced SummaC, a method for computing NLI scores between document-summary sentence pairs using a novel SummaCConv aggregation method. SummaC achieved 74.4% balanced accuracy (5% improvement over prior work) on inconsistency detection, demonstrating NLI’s effectiveness for document-level semantic preservation verification. Recent work [30] used DeBERTa-v3 NLI models for zero-shot question-answering evaluation, matching GPT-4o accuracy (89.9%) at orders of magnitude lower cost, demonstrating NLI’s viability for evaluating semantic equivalence in question answering relevant to clinical QA scenarios.
In addition to using NLI as an evaluation metric, it has also been used to enforce consistency across model outputs. For instance, Mitchell et al. [31] proposed ConCoRD, which uses pre-trained NLI models to check pairwise answer compatibility, builds factor graphs with NLI-based beliefs, and applies MaxSAT solvers. ConCoRD achieved 5% improvement on ConVQA and addressed transitivity failures without fine-tuning. This work complements ours: ConCoRD uses NLI to enforce consistency across model outputs, while we use NLI to verify the semantic equivalence of inputs before measuring sensitivity. Nighojkar and Licato [32] introduced the Adversarial Paraphrasing Task (APT), which generates mutually implicative sentences with minimal lexical overlap to verify mutual implication. They highlighted the difficulty of maintaining consistency even under NLI-verified conditions. In the clinical domain, Romanov and Shivade [33] created MedNLI, the first large-scale NLI dataset, comprising approximately 14,000 examples from MIMIC-III clinical notes. They demonstrated that transfer learning from SNLI is effective and that integrating medical knowledge is valuable for clinical NLI. Chowdhury et al. [34] improved MedNLI by incorporating UMLS and contextual knowledge, thereby demonstrating the value of domain-specific knowledge for clinical NLI. Our work leverages clinical NLI models, such as MedNLI, for semantic verification in clinical robustness testing, an application not explored in prior work.
While prior research has advanced perturbation-based testing, consistency measurement, and NLI-based evaluation, several critical gaps remain for the reliable deployment of clinical LLMs. First, existing approaches suffer from a semantic verification gap, as they primarily rely on similarity-based constraints (e.g., cosine similarity), which fail to capture fine-grained clinical distinctions involving negation, temporality, or severity. Second, there exists a clinical consistency gap: although MedNLI [33] enables clinical entailment reasoning, prior work has not systematically leveraged semantic verification to evaluate robustness under meaning-preserving variations, with most clinical evaluations focusing primarily on accuracy rather than consistency [35]. Third, current evaluation frameworks, such as CheckList [21] and AdaTest [22], along with consistency-focused methods like ParaRel [23] and PaRTE [24], generate or assess variations without enforcing semantic equivalence, limiting their ability to isolate true model instability. Finally, a measurement and interpretation gap persists, as existing metrics quantify performance under perturbations or prompt sensitivity but do not explicitly isolate variability under meaning-preserving transformations.
In this section, we present our proposed systematic framework to evaluate semantic stability in medical LLMs under meaning-preserving linguistic variations. Figure 1 illustrates the proposed pipeline, which consists of five stages: (1) base prompt selection, (2) variation generation, (3) NLI-based semantic verification, (4) hybrid judgment, and (5) sensitivity analysis.
We construct an evaluation set of 200 base prompts that are sampled from two public medical datasets, i.e., MedQA-USMLE and DiagnosisQA [5]. Specifically, we select 100 prompts from each dataset, representing complementary medical question answering and clinical interpretation tasks commonly used to evaluate LLMs in medical decision support settings. Each base prompt consists of a question stem, multiple-choice answer options, and a corresponding ground-truth answer. Following Cao et al. [7], for each base prompt \(p\), we generate exactly 10 candidate variations, resulting in a total of 2,000 candidate prompt variants across both datasets.
We used GPT-4o-mini with a fixed decoding temperature of 0.35 to generate ten predefined meaning-preserving variations, under strict semantic preservation constraints. These variations include abbreviation expansion, abbreviation compression, clinical-to-lay language conversion, lay-to-clinical language conversion, syntactic reordering, active-to-passive voice transformation, question reframing, documentation style modification, format shifts, and redundancy trimming. The generation process explicitly prohibits modification of clinical facts, including numeric values, units, negation, temporality, laterality, disease severity or stage, medications, and dosages. The model was further instructed not to add or remove information and to leave the stem unchanged when a requested variation could not be safely applied. To ensure diversity among generated variants, we apply a distinctness check within each prompt group using trigram Jaccard similarity, triggering regeneration when the similarity exceeded 0.92. Variants identified as identical to the original stem or insufficiently distinct from previously generated variants were regenerated up to 3 attempts per variation type. All generated variants, along with their associated metadata, were stored in a structured JSONL format for downstream processing. Table 1 summarizes the prompt template used for prompt variation generation.
| Task Context: Base prompts are medical multiple-choice questions sampled from the MedQA-USMLE and DiagnosisQA datasets. Each prompt consists of a single question stem accompanied by answer options and, where available, a reference answer. Only the question stem is subject to reformulation. |
|---|
| System Instruction: You are a clinical documentation expert. |
| Generation Instructions: Rewrite ONLY the question stem. Do NOT rewrite answer options or reference answers. You MUST preserve all clinical facts exactly as given. Do NOT change negation, temporality, laterality, disease severity or stage, medications, dosages, numeric values, or units. Do NOT add facts, remove facts, or introduce uncertainty. If a requested variation type cannot be safely applied, keep the stem unchanged for that type and explicitly record this in changes_made. Return ONLY valid JSON. |
Variation Types (exactly one per variant):
|
| Output Format (JSON per variant): |
To ensure that generated variations preserve meaning, we perform an NLI-based semantic verification procedure. This verification step operates exclusively on the question stems and is applied independently of the variation generation process. We employ three pretrained domain-specific NLI models: PubMedBERT-MNLI-MedNLI, FacebookAI’s roberta-large-mnli, and Microsoft’s deberta-large-mnli. Each NLI model outputs probabilities over three labels (i.e., entailment, neutral, and contradiction) for a given premise-hypothesis pair. Below, we discuss our proposed NLI-based semantic verification approach to determine whether generated prompt variations preserve the meaning of the corresponding base prompts.
Let \(p\) denote a base prompt and \(p'\) denote a candidate variation in the set of all variations \(\mathcal{V}(p)\). For each pair \((p, p')\), we compute entailment in both directions:
\[\mathbf{s}_{\text{forward}} = \text{NLI}(p \rightarrow p'), \quad \mathbf{s}_{\text{backward}} = \text{NLI}(p' \rightarrow p)\]
Here, \(\mathbf{s}_{\text{forward}} = (e_{\text{fwd}}, n_{\text{fwd}}, c_{\text{fwd}})\) and \(\mathbf{s}_{\text{backward}} = (e_{\text{bwd}}, n_{\text{bwd}}, c_{\text{bwd}})\) denote the predicted probabilities for entailment, neutral, and contradiction classes in the forward and backward directions, respectively. A candidate variation is considered to satisfy bidirectional entailment for a given NLI model if the following conditions are met: \[e_{\text{fwd}} \geq \tau_{\text{ent}}, \quad e_{\text{bwd}} \geq \tau_{\text{ent}}, \quad c_{\text{fwd}} < \tau_{\text{con}}, \quad c_{\text{bwd}} < \tau_{\text{con}},\]
where \(\tau_{\text{ent}}\) is the entailment threshold and \(\tau_{\text{con}}\) is the contradiction threshold. This bidirectional criterion enforces mutual implication between the base prompt and its variation, reducing false positives arising from asymmetric entailment.
Bidirectional entailment verification is performed independently for each of the three NLI models. For a given prompt pair \((p, p')\), a model is considered to satisfy semantic verification if it meets the predefined bidirectional threshold criteria. A multi-model consensus is then applied, requiring at least \(K = 2\) of the three models to verify the pair for it to be classified as meaning-preserving. Formally, let \(\mathbb{I}_m(p,p') \in {0,1}\) indicate whether model \(m\) satisfies the bidirectional entailment conditions for \((p,p')\). A variation is classified as meaning-preserving if, \(\sum_{m=1}^{3} \mathbb{I}_m(p,p') \geq K\).
After prompt generation and NLI-based semantic filtering, we apply a two-round hybrid verification process to ensure that retained prompt variations preserve the original clinical meaning and context. This process consists of automated review with expert adjudication and is summarized in Table 3. In Round 1, all NLI-verified samples are independently reviewed by both LLM-as-a-Judges and a human clinical expert. In Round 2, only samples rejected by the LLM judges are re-evaluated by the clinical expert. This adjudication step is designed to recover valid reformulations that may have been conservatively filtered by the LLM reviewers, while preserving expert oversight over ambiguous or borderline cases. The final benchmark therefore consists only of samples accepted after this sequential two-round process. Concretely, a prompt variation must (1) pass multi-model NLI semantic filtering, (2) survive independent Round 1 review, or be reinstated during Round 2 expert adjudication, and (3) satisfy final expert acceptance criteria.
To complement the NLI-based verification and introduce an additional semantic filtering layer, we employ an LLM-as-a-Judge mechanism. This stage operates as a secondary validation step to assess whether candidate variations \(P'\) preserve the same clinical intent as their corresponding base prompts \(P\). Specifically, we employ a committee of advanced models to evaluate semantic consistency between each base
prompt and its variation, utilizing a specific snapshot of GPT-5.2 (accessed April 2026) and the MedLLaMA3 model deployment via Ollama (version tag:
v20). Unlike the NLI models, which rely on probabilistic entailment scores, the LLM-as-judge provides a binary decision (YES/NO) on whether semantic equivalence is preserved. Table 2 illustrates the prompt template used for this evaluation.
| Role: You are a strict medical expert and clinical linguist. |
|---|
| Task Overview: Determine if the provided Candidate Question preserves the EXACT clinical meaning, intent, and factual granularity of the Original Question. |
Evaluation Criteria: The candidate question must be rejected (NO) if there is any change in:
|
| Input Variables: Original Question: {base_stem} |
| Candidate Question: {candidate_stem} |
| Constraint: Answer: YES or NO Reason: \(<\) brief explanation if NO, otherwise write None\(>\). |
To determine the optimal entailment threshold \(\tau_{\text{ent}}\), we evaluate LLM-as-a-judge decisions across a range of thresholds from \(0.70\) to \(0.90\). Figure 2 illustrates the number of variations classified as semantically equivalent (YES) and non-equivalent (NO) across different thresholds for both MedQA and DiagnosisQA datasets. We define the criteria for optimal threshold based on agreement between NLI verification and LLM judgments, with the objective of maximizing accepted meaning-preserving variations while minimizing clinically inconsistent cases. As shown in Figure 2, the number of accepted variations stabilizes between \(0.85\) and \(0.90\), indicating diminishing returns beyond this range. Therefore, we select \(\tau_{\text{ent}} = 0.85\) as the operating threshold for subsequent verification.
We conducted an independent clinical audit, in which one licensed clinician served as expert auditor and independently reviewed NLI-verified samples per dataset to assess whether each variation preserved the original clinical meaning and context. The evaluation was performed using the same criteria (as defined in Table 2), ensuring consistency between the two LLM-based judgments and the clinical expert audit. Expert evaluation was conducted under blinded conditions, without access to LLMs’ judgments, ensuring independence and minimizing potential bias. The clinical expert was conducted over three weeks to enable a careful and comprehensive assessment. The clinical expert confirmed semantic preservation with 98-99% agreement across all samples retained after NLI-based verification, indicating that these variations consistently preserved the intended clinical meaning. This high level of agreement validates the combined filtering process. To further assess potential false negatives, the clinical expert also reviewed the samples rejected by the LLM judges. A small number of LLM-rejected samples were identified by the expert as clinically equivalent and subsequently reinstated for further analysis. Table 3 summarizes the hybrid verification process to construct the meaning-preserving prompt variation benchmark.
| Dataset | Base Prompts | Variations | NLI Verified | Round 1 Independent Review | Round 2 Review | Final Dataset | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 5-8 (lr)9-10 (lr)11-12 | LLM-as-a-Judge | Human Expert | LLM-Rejected Samples | Accepted | Rejected | |||||||
| 5-6 (lr)7-8 (lr)9-10 | Accepted | Rejected | Accepted | Rejected | Reinstated | Confirmed Reject | ||||||
| MedQA | 100 | 1000 | 890 | 878 | 12 | 883 | 7 | 5 | 7 | 883 | 7 | |
| DiagnosisQA | 100 | 1000 | 886 | 875 | 11 | 878 | 8 | 3 | 8 | 878 | 8 | |
| Total | 200 | 2000 | 1776 | 1753 | 23 | 1761 | 15 | 8 | 15 | 1761 | 15 | |
We propose a set of quantitative metrics to measure the stability of model outputs under meaning-preserving prompt variations. These metrics quantify different aspects of stability, including prediction consistency, confidence variation, and worst-case robustness under semantically equivalent inputs. The metrics are defined independently of any specific model architecture or task and can be applied to both classification and generation settings. Let \(p\) denote a base prompt, and \(\mathcal{V}(p)\) denote the set of verified meaning-preserving variations associated with \(p\). Let \(\mathcal{M}\) denote a model under evaluation.
It quantifies the proportion of verified variations that lead to a change in the model output relative to the base prompt, defined as:
\[\text{MVS}_{\mathcal{M}}(p) = \frac{1}{|\mathcal{V}(p)|} \sum_{p' \in \mathcal{V}(p)} \mathbf{1}\left[\mathcal{M}(p) \neq \mathcal{M}(p')\right]\]
For classification tasks, \(\mathcal{M}(p) \neq \mathcal{M}(p')\) indicates that the predicted class differs between the base prompt and its variation. For generation tasks, this condition can be instantiated using task-specific equivalence criteria (e.g., exact match or semantic equivalence). Lower values of MVS indicate greater semantic stability, with \(\text{MVS} = 0\) corresponding to perfect invariance.
It captures changes in model confidence across meaning-preserving prompt variations, even when predicted outputs remain unchanged. Let \(C_{\mathcal{M}}(p)\) denote the maximum predicted probability associated with the model output for prompt \(p\), \(\Delta \text{C}\) is defined as:
\[\Delta C_{\mathcal{M}}(p) = \frac{1}{|\mathcal{V}(p)|} \sum_{p' \in \mathcal{V}(p)} \left| C_{\mathcal{M}}(p) - C_{\mathcal{M}}(p') \right|\]
Lower values of \(\Delta C_{\mathcal{M}}(p)\) indicate more stable confidence behavior across meaning-preserving variations.
While MVS quantifies how frequently model predictions change under meaning-preserving variations, it does not directly capture whether such changes affect answer correctness. To address this, we introduce a worst-case per-sample instability metric that measures whether the model’s correctness is invariant under all verified variations of a prompt. Intuitively, this metric checks whether any meaning-preserving rephrasing can flip the model from correct to incorrect (or vice versa). Let \(y(p)\) denote the ground-truth label associated with the prompt \(p\), \[s_{\mathcal{M}}(p) = \mathbf{1}\left[\mathcal{M}(p) = y(p)\right]\] denote the binary correctness of model \(\mathcal{M}\) on \(p\). We define the worst-case correctness over the set \(\{p\} \cup \mathcal{V}(p)\) as: \[\begin{align} s_{\mathcal{M}}^{\text{best}}(p) &= \max\Big( s_{\mathcal{M}}(p), \max_{p' \in \mathcal{V}(p)} s_{\mathcal{M}}(p') \Big), \\ s_{\mathcal{M}}^{\text{worst}}(p) &= \min\Big( s_{\mathcal{M}}(p), \min_{p' \in \mathcal{V}(p)} s_{\mathcal{M}}(p') \Big). \end{align}\]
The worst-case per-sample instability is then defined as: \[\text{Instability}_{\mathcal{M}}(p) = s_{\mathcal{M}}^{\text{best}}(p) - s_{\mathcal{M}}^{\text{worst}}(p)\]
This metric takes values in \(\{0,1\}\). A value of \(0\) indicates that the model’s correctness is invariant across all meaning-preserving variations of \(p\), while a value of \(1\) indicates that at least one variation causes a change in correctness. Lower values, therefore, correspond to stronger worst-case robustness.
All experiments are conducted on prompt variations that satisfy the multi-stage semantic verification pipeline, including NLI-based verification, LLM-as-a-judge filtering, and clinical expert audit. For each verified sample, we compute model predictions for both the original base prompt and all associated semantically-verified variations, facilitating direct comparison under controlled semantic conditions. We evaluate model sensitivity in a multiple-choice QA setting, where each input consists of a clinical question stem and a fixed set of answer options. The order of answer options remains unchanged across all prompt variations, ensuring that observed differences in model behavior are attributable solely to the linguistic reformulation of the question stem and isolating semantic sensitivity from confounding factors such as answer space variation or task ambiguity.
We evaluate a diverse set of open-weight large language models spanning both general-purpose (GP) and medically adapted domain-specific (DS) variants across similar families, with comparable parameter sizes ranging from 1.5B to 8B. The GP models include Qwen2.5-1.5B, Qwen2.5-3B, LLaMA-3.2-3B, Qwen2.5-7B, Mistral-7B-v0.2, LLaMA3.1-8B, and Qwen3-8B. These models are trained primarily on broad-domain corpora and serve as competitive baselines for reasoning and language understanding. The DS models include MedQwen2.5-1.5B, MedQwen2.5-3B, MedLLaMA-3.2-3B, BioMistral-7B, MedQwen2.5-7B, Meditron3-8B, MedLLaMA3-8B, MediChat-LLaMA3-8B, and Qwen3-8B-Biomedical. These models incorporate biomedical or clinical training data and are designed to better capture medical terminology and DS reasoning patterns. Including both GP and DS models enables us to assess whether domain specialization improves robustness to meaning-preserving linguistic reformulations.
All models are executed locally using the Hugging Face transformers library in inference-only mode, ensuring a consistent evaluation across different models. To enable efficient evaluation, we perform inference using 4-bit NF4 quantization
with half-precision computation, which substantially reduces memory requirements while maintaining practical inference fidelity for comparative analysis. Moreover, to ensure consistency across all models, we use a standardized prompt template that
instructs each model to act as a careful medical reasoning assistant, analyze the medical question, and select exactly one answer option from A–E. The prompt includes the question stem, all five answer choices, and concludes with the expected output cue
Answer:.
Rather than relying on unconstrained free-form text generation, we derive predictions directly from the model’s next-token output distribution following the Answer: cue. The candidate answer space is restricted to the five option tokens
(A–E), and the selected prediction corresponds to the option receiving the highest probability. Because predictions are obtained directly from token probabilities rather than stochastic decoding, the evaluation procedure is fully deterministic and
unaffected by sampling-based generation settings. We compute confidence by normalizing the probabilities assigned to the five answer-option tokens and selecting the probability of the predicted option (i.e., the highest value). Therefore, the reported
confidence score reflects the model’s relative preference among the available answer choices rather than a fully calibrated probability of correctness.
| DiagnosisQA | MedQA | |||||
|---|---|---|---|---|---|---|
| 4-5 (lr)6-7 Model | Type | Params | Acc | Conf | Acc | Conf |
| Qwen 2.5 | GP | 1.5B | \(0.46\) | \(0.46 \pm 0.13\) | \(0.29\) | \(0.42 \pm 0.10\) |
| LLaMA 3.2 | GP | 3B | \(0.77\) | \(0.64 \pm 0.19\) | \(0.54\) | \(0.51 \pm 0.17\) |
| Qwen 2.5 | GP | 3B | \(0.64\) | \(0.95 \pm 0.10\) | \(0.40\) | \(0.90 \pm 0.16\) |
| Mistral | GP | 7B | \(0.53\) | \(0.91 \pm 0.16\) | \(0.38\) | \(0.91 \pm 0.16\) |
| Qwen 2.5 | GP | 7B | \(0.73\) | \(0.90 \pm 0.16\) | \(0.49\) | \(0.89 \pm 0.15\) |
| LLaMA 3.1 | GP | 8B | \(0.80\) | \(0.80 \pm 0.18\) | \(0.66\) | \(0.62 \pm 0.23\) |
| Qwen 3 | GP | 8B | \(0.73\) | \(0.76 \pm 0.19\) | \(0.56\) | \(0.69 \pm 0.21\) |
| MedQwen 2.5 | DS | 1.5B | \(0.42\) | \(0.44 \pm 0.12\) | \(0.29\) | \(0.41 \pm 0.10\) |
| MedLLaMA 3.2 | DS | 3B | \(0.75\) | \(0.76 \pm 0.21\) | \(0.55\) | \(0.58 \pm 0.21\) |
| MedQwen 2.5 | DS | 3B | \(0.64\) | \(0.95 \pm 0.10\) | \(0.40\) | \(0.90 \pm 0.16\) |
| BioMistral | DS | 7B | \(0.51\) | \(0.62 \pm 0.20\) | \(0.36\) | \(0.58 \pm 0.20\) |
| MedQwen 2.5 | DS | 7B | \(0.79\) | \(0.78 \pm 0.20\) | \(0.52\) | \(0.60 \pm 0.22\) |
| MedLLaMA 3 | DS | 8B | \(0.69\) | \(0.80 \pm 0.21\) | \(0.51\) | \(0.70 \pm 0.22\) |
| MediChat-LLaMA 3 | DS | 8B | \(0.59\) | \(0.75 \pm 0.21\) | \(0.44\) | \(0.68 \pm 0.19\) |
| Meditron 3 | DS | 8B | \(0.80\) | \(0.68 \pm 0.19\) | \(0.59\) | \(0.50 \pm 0.19\) |
| Qwen 3 Biomedical | DS | 8B | \(0.77\) | \(0.84 \pm 0.17\) | \(0.72\) | \(0.73 \pm 0.21\) |
| Model | Params | DiagnosisQA | MedQA | ||||
|---|---|---|---|---|---|---|---|
| 3-5 (lr)6-8 | MVS (\(\downarrow\)) | \(\Delta\)C (\(\downarrow\)) | WCI (\(\downarrow\)) | MVS (\(\downarrow\)) | \(\Delta\)C (\(\downarrow\)) | WCI (\(\downarrow\)) | |
| General-Purpose Models (GP) | |||||||
| Qwen2.5-1.5B | 1.5B | 0.13 \(\pm\) 0.21 | 0.04 \(\pm\) 0.02 | 25.25% | 0.11 \(\pm\) 0.21 | 0.04 \(\pm\) 0.02 | 15.00% |
| 1-8 Qwen2.5-3B | 3B | 0.13 \(\pm\) 0.21 | 0.05 \(\pm\) 0.07 | 30.30% | 0.17 \(\pm\) 0.25 | 0.07 \(\pm\) 0.08 | 36.00% |
| LLaMA-3.2-3B | 3B | 0.08 \(\pm\) 0.17 | 0.05 \(\pm\) 0.02 | 25.25% | 0.13 \(\pm\) 0.21 | 0.06 \(\pm\) 0.03 | 33.00% |
| 1-8 Qwen2.5-7B | 7B | 0.11 \(\pm\) 0.21 | 0.07 \(\pm\) 0.09 | 30.30% | 0.07 \(\pm\) 0.15 | 0.07 \(\pm\) 0.07 | 18.00% |
| Mistral-7B-v0.2 | 7B | 0.16 \(\pm\) 0.26 | 0.06 \(\pm\) 0.08 | 28.28% | 0.15 \(\pm\) 0.24 | 0.08 \(\pm\) 0.09 | 27.00% |
| 1-8 LLaMA3.1-8B | 8B | 0.07 \(\pm\) 0.16 | 0.06 \(\pm\) 0.04 | 20.20% | 0.12 \(\pm\) 0.24 | 0.06 \(\pm\) 0.04 | 26.00% |
| Qwen3-8B | 8B | 0.10 \(\pm\) 0.19 | 0.07 \(\pm\) 0.05 | 28.28% | 0.12 \(\pm\) 0.20 | 0.07 \(\pm\) 0.05 | 32.00% |
| Domain-Specific Models (DS) | |||||||
| MedQwen2.5-1.5B | 1.5B | 0.14 \(\pm\) 0.23 | 0.04 \(\pm\) 0.02 | 26.26% | 0.14 \(\pm\) 0.23 | 0.04 \(\pm\) 0.02 | 20.00% |
1-8 MedQwen2.5-3B |
3B | 0.13 \(\pm\) 0.21 | 0.05 \(\pm\) 0.07 | 30.30% | 0.17 \(\pm\) 0.24 | 0.07 \(\pm\) 0.09 | 36.00% |
| MedLLaMA-3.2-3B | 3B | 0.10 \(\pm\) 0.19 | 0.06 \(\pm\) 0.04 | 30.30% | 0.15 \(\pm\) 0.24 | 0.07 \(\pm\) 0.05 | 32.00% |
1-8 BioMistral-7B |
7B | 0.14 \(\pm\) 0.22 | 0.07 \(\pm\) 0.05 | 34.34% | 0.14 \(\pm\) 0.22 | 0.06 \(\pm\) 0.03 | 28.00% |
| MedQwen2.5-7B | 7B | 0.10 \(\pm\) 0.20 | 0.05 \(\pm\) 0.05 | 30.30% | 0.13 \(\pm\) 0.19 | 0.06 \(\pm\) 0.04 | 32.00% |
1-8 Meditron3-8B |
8B | 0.09 \(\pm\) 0.20 | 0.05 \(\pm\) 0.03 | 21.21% | 0.16 \(\pm\) 0.27 | 0.05 \(\pm\) 0.03 | 30.00% |
| MedLLaMA3-8B | 8B | 0.09 \(\pm\) 0.17 | 0.06 \(\pm\) 0.05 | 24.24% | 0.16 \(\pm\) 0.26 | 0.08 \(\pm\) 0.06 | 29.00% |
| MediChat-LLaMA3-8B | 8B | 0.11 \(\pm\) 0.22 | 0.07 \(\pm\) 0.05 | 23.23% | 0.16 \(\pm\) 0.24 | 0.09 \(\pm\) 0.06 | 27.00% |
| Qwen3-8B-Biomedical | 8B | 0.09 \(\pm\) 0.19 | 0.05 \(\pm\) 0.05 | 18.18% | 0.15 \(\pm\) 0.24 | 0.07 \(\pm\) 0.05 | 37.00% |
Before evaluating sensitivity to meaning-preserving prompt variations, we establish a baseline using the original prompts from DiagnosisQA and MedQA. This baseline serves two purposes: it quantifies each model’s nominal capability under standard prompting conditions, and it provides a reference point for interpreting behavioral changes under prompt reformulations. Table 4 reports baseline accuracy and confidence scores for GP and DS LLMs across comparable parameter scales (1.5B–8B) and similar model families. It is evident from that table that DS LLMs do not consistently outperform GP models under original prompt settings. While several DS models achieve clear gains over their corresponding GP counterparts, these improvements are selective rather than universal. On DiagnosisQA, the strongest GP model (LLaMA 3.1 8B, 0.80) matches the best DS model (Meditron 3 8B, 0.80), and also exceeds several medically specialized alternatives. On MedQA, GP models remain competitive, with LLaMA 3.1 8B reaching 0.66 accuracy, surpassed only by Qwen 3 Biomedical (0.72). These results show that strong GP instruction-tuned models can perform competitively on medical QA benchmarks without explicit domain specialization.
Table 4 also shows that domain specialization can provide meaningful benefits when the adaptation process is effective. For instance, MedQwen 2.5 7B improves over Qwen 2.5 7B on both DiagnosisQA (0.79 vs. 0.73) and MedQA (0.52 vs. 0.49), while Qwen 3 Biomedical substantially exceeds standard Qwen 3 8B, especially on MedQA (0.72 vs. 0.56). These improvements suggest that targeted biomedical adaptation can enhance domain knowledge and reasoning. Notably, the strongest performance on MedQA is achieved by the Qwen 3 Biomedical (DS model), showing that well-executed domain adaptation can yield substantial gains. However, this effect is not seen uniformly across different DS models, such as BioMistral and MediChat-LLaMA 3, which underperform similarly sized GP baselines. In particular, the sizeable gap between MediChat-LLaMA 3 (0.59) and its stronger GP counterpart LLaMA 3.1 8B (0.80) on DiagnosisQA may reflect catastrophic forgetting, where domain-specific fine-tuning partially degrades broader reasoning capabilities learned during base pretraining. This suggests that the impact of specialization depends more on how adaptation is performed, including data quality, objective design, alignment strategy, and retention of general capabilities rather than on whether medical adaptation.
We also observe a persistent gap between predictive accuracy and model confidence across datasets. Nearly all models achieve lower accuracy on MedQA than on DiagnosisQA, suggesting that MedQA is the more challenging benchmark and likely requires deeper factual recall or more complex reasoning. Despite this drop in performance, confidence scores remain high for several models across both datasets, often ranging from 0.70 to 0.95, even when accuracy is modest. For example, Qwen 2.5 3B and MedQwen 2.5 3B achieve only 40% accuracy on MedQA while reporting confidence of 0.90, and Mistral 7B provides just 38% accuracy with confidence as high as 0.91. This misalignment indicates systematic overconfidence, where models express strong certainty without corresponding reliability. Importantly, this behavior occurs in both GP and DS models, suggesting that medical fine-tuning alone does not consistently resolve confidence misalignment.
To answer RQ2, we first provide an aggregate sensitivity analysis in Table 5, which analyzes whether domain specialization consistently improves robustness when prompts are reformulated without changing their meaning and context. These results demonstrate that domain specialization does not provide a universal robustness advantage over GP models, but neither does it systematically reduce robustness. Instead, its effects are selective, model-specific, and dataset-dependent. Across both DiagnosisQA and MedQA, several DS LLMs rank among the most stable models, while strong GP models remain highly competitive and frequently match DS counterparts. For example, on DiagnosisQA, LLaMA3.1-8B provides the lowest MVS among GP models (0.07) with a relatively low WCI (20.20%), while Qwen2.5-7B performs strongly on MedQA, with the lowest GP MVS (0.07) and a low WCI (18.00%). Among DS models, Qwen3-8B-Biomedical achieves the lowest WCI on DiagnosisQA (18.18%), and Meditron3-8B and MedLLaMA3-8B also demonstrate consistently strong robustness. On MedQA, MedQwen2.5-7B provides the lowest DS MVS (0.13), while MedQwen2.5-1.5B attains the lowest DS WCI (20.00%). Therefore, we argue that while domain specialization can improve robustness in specific cases, it is not a reliable predictor of robustness gains or losses relative to strong GP baselines.
We can also observe from Table 5 that robustness depends not only on model type, but also on model scale and adaptation quality. Smaller models generally show greater sensitivity and less consistent behaviour, suggesting that limited capacity may hinder stable decision-making under semantically equivalent reformulations. By contrast, several of the strongest robustness results are concentrated among 7B–8B models, where both GP and DS models more frequently achieve lower MVS and WCI values. However, scale alone does not guarantee robustness, as some smaller models still achieve competitive stability (e.g., Qwen2.5-1.5B with the best GP MedQA WCI of 15.00%), while some 8B models remain less stable than peers on specific metrics. This indicates that robustness emerges from the interaction of model capacity, architecture, and training quality rather than parameter count alone.
Finally, both GP and DS models become less robust on the more challenging MedQA datasets. Across most models, MedQA provides higher MVS and WCI than DiagnosisQA, indicating that semantic stability degrades as reasoning demands increase. For example, LLaMA3.1-8B rises from 20.20% WCI on DiagnosisQA to 26.00% on MedQA, while Qwen3-8B-Biomedical increases sharply from 18.18% to 37.00%. At the same time, confidence variation remains comparatively small for many models despite these correctness shifts, implying that models often maintain similar confidence even when their answers change under equivalent reformulations. These results complement our earlier argument that robustness is shaped more by architecture, scale, and training quality than by domain specialization alone.
Figure 3 provides a paired robustness comparison between GP and DS models under various meaning-preserving linguistic reformulations. Each GP model is aligned with a parameter-matched DS counterpart, allowing robustness differences to be interpreted with reduced confounding from model scale. Two high-level trends are immediately apparent from Figure 3. First, the dataset effect observed in Table 5 is again confirmed: MedQA consistently exhibits higher MVS than DiagnosisQA across nearly all model pairs and transformation types. This suggests that robustness degrades when input tasks demand more complex reasoning, denser factual recall, or greater sensitivity to contextual nuance. Second, robustness remains strongly model-dependent rather than domain-dependent (i.e., GP vs DS), where domain specialization does not uniformly improve stability. In some pairs, DS models outperform their GP counterparts—for example, BioMistral is more stable than Mistral under several MedQA transformations, and Qwen3-Bio generally improves over Qwen3-8B on DiagnosisQA. However, other DS variants show little benefit or mixed differences relative to their paired GP baselines, particularly in the Qwen small-scale pair, where MedQwen2.5-3B often mirrors or only marginally differs from Qwen2.5-3B. These results reinforce that domain specialization neither consistently helps nor harms robustness, with effects varying by architecture, scale, and dataset.
The figure also reveals that robustness is highly dependent on the type of linguistic reformulation. Across both datasets, Question Reframe is consistently among the most disruptive transformations and often produces the highest MVS scores across various models. This suggests that models remain sensitive to changes in discourse framing or interrogative structure, even when semantic intent is preserved. Clinical-to-Lay Language is similarly challenging, particularly on MedQA, where many models attain sharp increases in sensitivity. This implies that translating between professional and consumer-facing terminology remains a non-trivial robustness challenge, even for medically tuned LLMs. In contrast, variations focusing on localized edits such as Abbreviation Expansion, Abbreviation Compression, and Active-to-Passive voice changes generally lead to lower MVS values, indicating stronger invariance to lexical substitution and shallow syntactic alternation. Another notable pattern is that larger or stronger baseline models are not always the most stable: several 8B models remain vulnerable to question reframing and terminology shifts despite higher nominal accuracy. Overall, the paired comparisons suggest that semantic robustness is shaped by a combination of model architecture, adaptation strategy, and training diversity rather than by parameter scale or medical specialization alone. These findings further emphasize that evaluating medical LLMs solely on standard benchmark accuracy can obscure substantial fragility under realistic variations in clinical language.
We next investigate whether model confidence remains stable when prompts are reformulated without changing their meaning, and whether confidence instability is aligned with prediction instability. As shown in Figure 4, the association between MVS and \(\Delta\)C is weak overall, with almost no clear trend on DiagnosisQA (Figure 4a) and a modest positive trend on MedQA (Figure 4b). Models that are more prediction-sensitive under reformulation often exhibit larger confidence shifts, yet the relationship is far from deterministic. Several models with relatively modest MVS still show noticeable confidence volatility, whereas others with higher sensitivity maintain comparatively stable confidence. This dispersion indicates that prediction instability and confidence instability capture related but distinct failure modes. A model may preserve the same answer while substantially changing its certainty, or conversely flip predictions while expressing similar confidence. In practical terms, confidence scores alone do not provide a dependable proxy for semantic robustness. A second important finding is that MedQA tends to induce somewhat greater confidence instability than DiagnosisQA. The broader spread of points and steeper trend in Figure 4b suggest that more challenging clinical reasoning tasks can amplify fluctuations not only in predictions but also in certainty estimates. This implies that confidence becomes less stable precisely in settings where trustworthy model behavior is most critical. Thus, confidence estimates appear sensitive to task difficulty and prompt formulation, rather than reflecting a fixed internal measure of certainty.
To further assess whether model confidence reliably reflects actual performance, we analyzed the alignment between predictive confidence and empirical accuracy across all evaluated models. For each model and dataset, we computed mean variant predictive confidence and compared it against the observed accuracy. The results for this correlation are depicted in Figure 5, where the diagonal reference line (\(y=x\)) denotes perfect alignment between confidence and achieved performance, such that models positioned above the line exhibit confidence levels that exceed their observed accuracy. To quantify the overall association, we additionally report Pearson’s \(r\) and Spearman’s \(\rho\) for each dataset. As shown in Figure 5, confidence is not consistently aligned with performance across models. On DiagnosisQA, the relationship between confidence and accuracy is modest (\(r=0.41\)), suggesting that more accurate models tended to be somewhat more confident overall. In contrast, alignment is relatively weaker on MedQA (\(r=0.18\)), where higher confidence did not meaningfully correspond to higher accuracy. Several models on both benchmarks also appeared above the diagonal line, indicating confidence levels substantially higher than their achieved performance. This suggests that predictive confidence should be interpreted cautiously, as it does not always serve as a dependable indicator of correctness in medical QA tasks.
Our main sensitivity analysis evaluates all models under 4-bit NF4 quantization with half-precision compute. To quantify the effect of quantization on model sensitivity, we re-evaluated all eight 7–8B models on both DiagnosisQA and MedQA at full precision (16-bit) for paired GP and DS models (i.e., same size and family). The results of these experiments are summarized in Table 6, which are reported in terms of mean and standard deviation (to facilitate a direct comparison with Table 5). The table also reports the calculated change (\(\Delta\)) in the metric introduced by NF4 quantization with respect to full-precision settings, where positive values indicate degradation. The table shows that quantization has a small impact on model stability across both GP and DS models. For MVS, the observed changes are consistently small, with the difference (\(\Delta\)) for many models very close to zero. However, some models exhibit relatively large degradation in MVS, e.g., LLaMA3.1-8B and MedQwen2.5-7B both move by \(+0.04\) on DiagnosisQA, Mistral-7B-v0.2 by \(+0.04\) on MedQA, and MedLLaMA3-8B by \(+0.06\) on MedQA. This suggests that quantization introduces only negligible additional variance in model stability. A similar pattern is observed for \(\Delta\)C, where differences remain close to zero for nearly all models, indicating that this metric is largely insensitive to reduced precision. In contrast, WCI exhibits comparatively larger variations, with several models showing increases of up to approximately 8%, although these effects are not uniform and for Qwen2.5-7B, quantization leads to an improvement of 7% in WCI, and Qwen3-8B-Biomedical is the only model whose DiagnosisQA WCI improves under quantization (\(-1.01\)). The asymmetry between datasets is also noticeable for individual models: LLaMA3.1-8B shifts by \(+8.08\) on DiagnosisQA but \(0.00\) on MedQA, while Qwen2.5-7B moves in opposite directions on the two datasets (\(+6.06\) vs \(-7.00\)). These results highlight that NF4 quantization preserves models’ stability to a large extent, with only modest and metric- and dataset-dependent variations.
| MVS \((\downarrow)\) | \(\Delta C\) \((\downarrow)\) | WCI (%) \((\downarrow)\) | ||||
|---|---|---|---|---|---|---|
| 2-3 (lr)4-5 (lr)6-7 Model | FP | \(\Delta\) | FP | \(\Delta\) | FP | \(\Delta\) |
| DiagnosisQA | ||||||
| General-Purpose (GP) Models | ||||||
| Qwen3-8B | \(0.09 \pm 0.18\) | \(+0.01\) | \(0.06 \pm 0.05\) | \(+0.01\) | \(23.23\) | \(+5.05\) |
| Qwen2.5-7B | \(0.10 \pm 0.23\) | \(+0.01\) | \(0.05 \pm 0.07\) | \(\mathbf{+0.02}\) | \(24.24\) | \(\mathbf{+6.06}\) |
| LLaMA3.1-8B | \(0.03 \pm 0.10\) | \(\mathbf{+0.04}\) | \(0.05 \pm 0.04\) | \(+0.01\) | \(12.12\) | \(\mathbf{+8.08}\) |
| Mistral-7B-v0.2 | \(0.13 \pm 0.23\) | \(+0.03\) | \(0.06 \pm 0.08\) | \(\hphantom{+}0.00\) | \(23.23\) | \(+5.05\) |
| Qwen3-8B-Biomedical | \(0.08 \pm 0.18\) | \(+0.01\) | \(0.05 \pm 0.06\) | \(\hphantom{+}0.00\) | \(19.19\) | \(-1.01\) |
| MedQwen2.5-7B | \(0.06 \pm 0.16\) | \(\mathbf{+0.04}\) | \(0.05 \pm 0.04\) | \(\hphantom{+}0.00\) | \(22.22\) | \(\mathbf{+8.08}\) |
| MedLLaMA3-8B | \(0.10 \pm 0.21\) | \(-0.01\) | \(0.06 \pm 0.05\) | \(\hphantom{+}0.00\) | \(22.22\) | \(+2.02\) |
| BioMistral-7B | \(0.13 \pm 0.20\) | \(+0.01\) | \(0.06 \pm 0.04\) | \(+0.01\) | \(31.31\) | \(+3.03\) |
| MedQA | ||||||
| General-Purpose (GP) Models | ||||||
| Qwen3-8B | \(0.14 \pm 0.22\) | \(-0.02\) | \(0.07 \pm 0.05\) | \(\hphantom{+}0.00\) | \(31.00\) | \(+1.00\) |
| Qwen2.5-7B | \(0.07 \pm 0.14\) | \(\hphantom{+}0.00\) | \(0.06 \pm 0.07\) | \(+0.01\) | \(25.00\) | \(\mathbf{-7.00}\) |
| LLaMA3.1-8B | \(0.09 \pm 0.18\) | \(+0.03\) | \(0.06 \pm 0.04\) | \(\hphantom{+}0.00\) | \(26.00\) | \(\hphantom{+}0.00\) |
| Mistral-7B-v0.2 | \(0.11 \pm 0.19\) | \(\mathbf{+0.04}\) | \(0.06 \pm 0.07\) | \(\mathbf{+0.02}\) | \(25.00\) | \(+2.00\) |
| Qwen3-8B-Biomedical | \(0.12 \pm 0.21\) | \(+0.03\) | \(0.06 \pm 0.04\) | \(+0.01\) | \(32.00\) | \(+5.00\) |
| MedQwen2.5-7B | \(0.13 \pm 0.19\) | \(\hphantom{+}0.00\) | \(0.05 \pm 0.04\) | \(+0.01\) | \(29.00\) | \(+3.00\) |
| MedLLaMA3-8B | \(0.10 \pm 0.19\) | \(\mathbf{+0.06}\) | \(0.08 \pm 0.06\) | \(\hphantom{+}0.00\) | \(25.00\) | \(+4.00\) |
| BioMistral-7B | \(0.12 \pm 0.21\) | \(+0.02\) | \(0.07 \pm 0.04\) | \(-0.01\) | \(27.00\) | \(+1.00\) |
To determine whether domain specialization systematically alters robustness under meaning-preserving prompt reformulations, we perform pairwise statistical comparisons between strictly matched GP and DS models from the same architectural family and comparable parameter scale (e.g., Qwen vs. MedQwen, LLaMA vs. MedLLaMA, Mistral vs. BioMistral). For each matched pair, we compare aligned per-sample sensitivity metrics using two-sided paired permutation tests (100,000 permutations). The results of this analysis are summarized in Table 7 in terms of DS–GP differences in MVS and \(\Delta\)C, \(p\)-values, and paired effect sizes (\(d\)). The table reveals a highly heterogeneous pattern, providing little evidence that domain specialization consistently improves robustness. For instance, most matched comparisons are statistically non-significant across both datasets, and observed effect sizes are generally small. On DiagnosisQA, several DS models show numerically lower MVS than their GP counterparts, including BioMistral (7B), MedQwen2.5 (7B), and Qwen3-Biomedical (8B), but none of these differences are statistically significant. Conversely, MedLLaMA3 (8B) and MedLLaMA3.2 (3B) exhibit higher MVS than their matched GP baselines, again without reliable statistical significance evidence. This supports our finding that semantic robustness is largely model-specific rather than determined by specialization status.
A clearer trend emerges on MedQA, where task difficulty is higher and prompt sensitivity becomes more pronounced. Here, the only statistically significant MVS result is the 7B Qwen pair, where MedQwen2.5-7B is significantly less robust than Qwen2.5-7B (\(\mathrm{MVS}_{\mathrm{DS-GP}}\) = +0.052, \(p=0.0235\), \(d=0.230\)). Some other DS models also show positive MVS differences, including MedLLaMA3 (8B) and Qwen3-Biomedical (8B), indicating reduced stability relative to their GP counterparts, although these differences are not statistically significant. Importantly, none of the DS models demonstrates a significant MVS advantage on MedQA. These results support our argument that robustness differences are dataset-specific and that domain specialization does not reliably translate into stronger semantic invariance or weaker robustness in general, with effects varying across matched model pairs. Unlike MVS trends, confidence variation (\(\Delta\)C) presents a different picture, where significant differences are more frequent than for MVS, indicating that domain specialization more consistently changes how confident models respond to reformulated prompts than whether they preserve the same predictions. However, these shifts are not uniform across all models. Some DS models exhibit lower confidence instability (e.g., MedQwen2.5-1.5B on both datasets and Qwen3-Biomedical on DiagnosisQA), whereas others become significantly less stable (e.g., MedLLaMA3.2 on DiagnosisQA and MedLLaMA3 on MedQA). This asymmetry reinforces a key conclusion of our study: confidence behavior and prediction robustness are related but distinct properties, and gains in one do not guarantee gains in the other.
| Params | GP Model | DS Model | \(\mathrm{MVS}_{\mathrm{DS-GP}}\) | \(p\)(MVS) | \(d\)(MVS) | \(\Delta C_{\mathrm{DS-GP}}\) | \(p\)(\(\Delta\)C) | \(d\)(\(\Delta\)C) |
|---|---|---|---|---|---|---|---|---|
| DiagnosisQA | ||||||||
| 1.5B | Qwen2.5 | MedQwen2.5 | 0.009 | 0.5944 | 0.056 | -0.002 | 0.0073\(^{*}\) | -0.263 |
| 3B | Qwen2.5 | MedQwen2.5 | -0.001 | 0.8652 | -0.026 | -0.001 | 0.4471 | -0.078 |
| 7B | Qwen2.5 | MedQwen2.5 | -0.013 | 0.6542 | -0.047 | -0.013 | 0.1602 | -0.143 |
| 8B | Qwen3 | Qwen3-Biomedical | -0.010 | 0.6962 | -0.039 | -0.012 | 0.0234\(^{*}\) | -0.231 |
| 3B | LLaMA3.2 | MedLLaMA3.2 | 0.022 | 0.3142 | 0.102 | 0.014 | 0.0006\(^{*}\) | 0.351 |
| 8B | LLaMA3.1 | MedLLaMA3 | 0.027 | 0.2536 | 0.116 | -0.003 | 0.5639 | -0.058 |
| 7B | Mistral | BioMistral | -0.022 | 0.5185 | -0.065 | 0.005 | 0.5468 | 0.061 |
| MedQA | ||||||||
| 1.5B | Qwen2.5 | MedQwen2.5 | 0.026 | 0.0636 | 0.182 | -0.002 | 0.0003\(^{*}\) | -0.358 |
| 3B | Qwen2.5 | MedQwen2.5 | 0.003 | 0.6047 | 0.055 | 0.000 | 0.8015 | 0.026 |
| 7B | Qwen2.5 | MedQwen2.5 | 0.052 | 0.0235 | 0.230 | -0.006 | 0.4199 | -0.081 |
| 8B | Qwen3 | Qwen3-Biomedical | 0.039 | 0.1641 | 0.141 | -0.001 | 0.7694 | -0.029 |
| 3B | LLaMA3.2 | MedLLaMA3.2 | 0.021 | 0.4434 | 0.077 | 0.009 | 0.0341 | 0.213 |
| 8B | LLaMA3.1 | MedLLaMA3 | 0.038 | 0.2821 | 0.109 | 0.019 | 0.0010\(^{*}\) | 0.332 |
| 7B | Mistral | BioMistral | -0.004 | 0.8977 | -0.013 | -0.011 | 0.2665 | -0.112 |
Despite the promising findings, this study has several limitations that motivate future work. First, high-performing clinical models such as MedPaLM and MedPaLM 2 were not included due to access and deployment constraints, limiting comparisons to publicly available or deployable models; incorporating such systems would enable a more comprehensive evaluation across open-source, GP, and proprietary clinical LLMs. Second, our evaluation protocol intentionally employs standardized single-step decoding under constrained inference settings to isolate intrinsic model sensitivity under controlled conditions. Consequently, the study does not evaluate reasoning-oriented inference strategies, such as chain-of-thought prompting, self-consistency sampling, or adaptive test-time reasoning, which are increasingly adopted in modern LLM deployments. Incorporating such mechanisms would introduce additional inference-time adaptation and model-specific reasoning behaviors that could confound the attribution of instability to semantic-preserving prompt perturbations alone. Therefore, the reported results should be interpreted primarily as estimates of intrinsic decoding-level sensitivity rather than robustness after downstream stabilization or reasoning-based mitigation strategies have been applied. Third, our experiments are limited to multiple-choice benchmarks (MedQA-USMLE and DiagnosisQA). Although these datasets enable controlled and reproducible comparisons, they do not fully reflect the complexity of real-world clinical scenarios, which often involve free-text clinical notes, longitudinal patient histories, retrieval-augmented reasoning, and multi-turn interactions. Fourth, verification of meaning-preserving variations relies on a single clinical expert; although conducted under blinded conditions and showing high agreement with the automated pipeline, the absence of multiple annotators limits assessment of inter-rater reliability. In our future work, we will extend this framework to reasoning-mode and frontier models, apply standard prompt engineering and self-consistency techniques as additional baselines, report multiple-testing-adjusted p-values and inter-rater Cohen’s kappa, and validate variations against a small set of human-written paraphrases.
In this paper, we introduce a semantically grounded method for evaluating the robustness of medical Large Language Models (LLMs) to meaning-preserving prompt variations. Specifically, we put forward a benchmark of semantically equivalent prompts, developed using a multi-stage pipeline that establishes semantic equivalence via bidirectional entailment using multiple Natural Language Inference (NLI) models, further refines candidate variations through two LLM judges, and conducts final auditing by a clinical expert. In addition, we introduce three complementary robustness metrics: (1) Meaning-Preserving Variation Sensitivity (MVS); (2) confidence variation (\(\Delta \text{C}\)); and (3) Worst-Case Instability (WCI)—to capture prediction inconsistency, confidence shifts, and per-sample fragility under semantically equivalent reformulations. We perform a large-scale analysis of 16 open-source general-purpose (GP) and domain-specific (DS) medical LLMs across matched model families and parameter scales on the proposed benchmark, derived from the DiagnosisQA and MedQA datasets. Our results reveal that robustness to meaning-preserving variations is highly model-dependent, with no consistent advantage for domain-specialized models. In several experiments comparing GP and DS models, GP models were found equally or more robust than their medically adapted counterparts, while in some matched comparisons DS models improved upon their GP counterparts, indicating that specialization effects are architecture- and dataset-dependent rather than uniform. This suggests that medical fine-tuning alone does not necessarily guarantee stable behavior under natural prompt variability. We also observe that confidence estimates are only weakly aligned with prediction stability, with several models maintaining high confidence even when robustness is reduced or predictions are unstable. Our results highlight the importance of moving beyond standard benchmark accuracy when assessing the clinical readiness of LLMs, and demonstrate that robustness to semantically equivalent prompt reformulations should be treated as a core requirement before deployment in safety-critical healthcare environments.