January 01, 1970
In applied health research, Markov cohort models are built on a precisely specified transition probability matrix. However, in many applications the available evidence—transition counts, structural constraints, and treatment-effect data—identifies a set of admissible matrices rather than one uniquely justified matrix. This paper formulates an imprecise-probability extension in which inference yields lower and upper expectations over an evidence-compatible set of precise Markov cohort models. The contribution differs from existing imprecise Markov-chain work by focusing on finite-horizon cohort trajectories, additive accumulated outcomes, and transition matrices constructed from empirical transition counts. Under non-empty compact separately specified outgoing-row sets, the lower and upper accumulated outcomes are computed exactly by Bellman-style lower and upper transition operators. We prove the envelope theorem, reduction to the classical model, coherence properties of the lower transition operator, and algebraic conditions under which a single selected matrix yields a non-robust decision. We then show how multinomial transition counts induce admissible matrix sets through the Imprecise Dirichlet Model. A real-world cost-effectiveness example of patent foramen ovale closure after cryptogenic stroke illustrates the practical consequence: the empirical transition matrix slightly favors closure, whereas the imprecise analysis yields an incremental net monetary benefit interval crossing zero. The method provides both a rigorous lower-expectation formulation and a practical diagnostic for decisions that depend on transition probabilities not fully resolved by the evidence.
imprecise probability ,Markov cohort model ,lower expectation ,transition matrix ,Imprecise Dirichlet Model ,decision analysis ,cost-effectiveness analysis
In applied research in health and medicine, Markov cohort models are often used to simulate the prognosis of a patient or a cohort of patients following an intervention in the absence of long-term data [1]–[3]. The prognosis is represented as a stochastic transition model over a finite set of health states. These transitions are governed by transition probability matrices, which are the central numerical objects in such models. Once a matrix has been estimated, the cohort trajectories and all outcomes are precisely determined. Treating the matrix as fixed is appropriate when the transition probabilities are known or can be estimated with a high degree of certainty. However, it is less satisfactory when the matrix is estimated from sparse transition counts, heterogeneous evidence, and conflicting expert opinions. In those settings, several transition matrices may be compatible with the evidence. This paper treats that situation as an imprecise-probability problem. Instead of assigning a probability distribution over possible transition matrices, the analyst specifies the set of matrices that remain compatible with the evidence and structural assumptions. Each element of this set can be viewed as a standard transition probability matrix defining a precise Markov model. Given this set of matrices, we can compute the smallest and largest expected outcomes. Computing these bounds is the lower-envelope view used in imprecise probability [4]–[6] and in imprecise Markov models [7]–[9]. The lower-envelope view should not be confused with a Bayesian average over a prior distribution: all admissible precise models are retained on equal footing. Other related ideas include imprecise Markov chains and lower-transition-operator methods. Existing continuous-time work treats an imprecise model as a set of stochastic processes, computes lower expectations as tight lower bounds over that set, and develops lower transition operators as representational and computational tools for those lower expectations [7], [9], [10].
Discrete-time imprecise Markov chains have been developed in the same spirit, with global models and lower and upper expectations of functionals such as time averages and accumulated quantities [11], [12]. Another related literature is robust Markov decision processes, in which dynamic programming over a rectangular ambiguity set—one specified independently across states—yields an exact Bellman recursion [13]–[15].
In this paper, we extend these ideas to finite-horizon cohort trajectories with additive outcomes, using the same state-vector and transition-matrix objects as cohort models. In this discrete-time setting, the quantities of interest are accumulated outcomes, in addition to the probabilities of occupying a state at a future time point. Each such outcome is calculated by summing, over the time horizon, the outcome contribution of each state weighted by the fraction of the cohort occupying it, and may represent survival, cost, quality-adjusted survival, or net monetary benefit. This accumulated-outcome calculation is the formulation used in cohort models and cost-effectiveness analysis, where the cohort is repeatedly redistributed across states according to the transition matrix and outcomes are accumulated over the time horizon of the analyses [2], [3].
The contributions of this paper are as follows. First, we explicate an imprecise Markov cohort model as a set of admissible transition-matrix sequences and give its lower- and upper-expectation interpretation for additive cohort outcomes, including outcomes that depend on transitions as well as on occupied states. Second, under non-empty compact separately specified outgoing-row sets, we derive exact lower and upper Bellman recursions, prove the coherence of the associated lower transition operators, and show that the resulting accumulated-outcome functional is itself a coherent lower prevision over the admissible set. Third, we establish a performance-difference identity that exactly captures the gap between any two admissible models, and use it to obtain an algebraic diagnostic for determining when the result from a single selected matrix is invariant over the admissible set and when that matrix yields a non-robust decision. Fourth, we present a practical construction of admissible transition matrices from multinomial transition counts using the Imprecise Dirichlet Model (IDM), show that the induced row set coincides exactly with the IDM credal set, and demonstrate the method with a real-world example from a cost-effectiveness analysis.
Let \[\mathcal{S}=\{S_1,\ldots,S_s\}\] be a finite set of mutually exclusive states, where \(s\) is the number of states. We denote by \[\mathcal{L}(\mathcal{S})=\{f:\mathcal{S}\to\mathbb{R}\}\] the set of all real-valued functions on \(\mathcal{S}\). Because \(\mathcal{S}\) is finite, elements of \(\mathcal{L}(\mathcal{S})\) will be written as column vectors \(f=(f(1),\ldots,f(s))^\top\), where \(f(i)=f(S_i)\).
A transition probability matrix is an \(s\times s\) matrix \(P=(p_{ij})_{i,j=1}^s\) satisfying \(p_{ij}\ge0\) and \(\sum_{j=1}^s p_{ij}=1\) for every \(i=1,\ldots,s\). Its \(i\)th row is denoted by \(P_{i\cdot}\). If \(X_t\) is the state occupied by one individual at time step \(t\), then \(p_{ij,t}=\mathbb{P}(X_{t+1}=S_j\mid X_t=S_i)\).
A cohort vector is a row vector \[m_t=(m_{t,1},\ldots,m_{t,s}),\qquad m_{t,j}\ge0,\qquad \sum_{j=1}^s m_{t,j}=1.\] It represents the distribution of \(X_t\): \(m_{t,j}=\mathbb{P}(X_t=S_j)\). The deterministic cohort trajectory \(m_0,\ldots,m_T\) is therefore the sequence of marginal distributions of the stochastic process.
For \(t=0,\ldots,T\), let \[g_t\in\mathcal{L}(\mathcal{S})\] be an outcome-contribution function. The number \(g_t(i)\) is the amount of the chosen outcome assigned to a person in state \(S_i\) at time step \(t\). Depending on the application, \(g_t\) may represent survival, utility, cost, net monetary benefit, or an indicator for occupying a target state. Because \(g_t\) is allowed to depend on the time step, it also represents discounting and other time-varying valuations: a per-step discount factor \(\beta\in(0,1]\) is encoded by \(g_t(i)=\beta^{t}\tilde{g}_t(i)\) for an undiscounted contribution \(\tilde{g}_t\). Some applications additionally assign part of the outcome to the act of transitioning; this transition-based case is treated at the end of Section 4.
Given an initial distribution \(m_0\) and transition matrices \(P_0,\ldots,P_{T-1}\), the classical cohort recursion is \[m_{t+1}=m_tP_t,\qquad t=0,\ldots,T-1. \label{eq:forward}\tag{1}\] The accumulated cohort outcome is \[A(P_{0:T-1})=\sum_{t=0}^{T}m_tg_t, \label{eq:classical-accumulated}\tag{2}\] where \(P_{0:T-1}\) denotes the sequence \(P_0,\ldots,P_{T-1}\), and \(m_t\) is generated by 1 . Equation 2 is the usual cohort accounting: multiply the fraction of the cohort in each state by the outcome contribution assigned to that state, then add across states and time steps.
The same accumulated outcome can be written backward. Define \(V_t\in\mathcal{L}(\mathcal{S})\) by \[V_t(i)=\mathbb{E}\left[\sum_{\ell=t}^{T}g_\ell(X_\ell)\,\middle|\,X_t=S_i\right],\] where \(\ell\) indexes future time steps. Then \(V_T=g_T\). For \(t<T\), conditioning on the next state gives the Bellman recursion \[V_T=g_T,\qquad V_t=g_t+P_tV_{t+1},\qquad t=T-1,\ldots,0. \label{eq:classical-bellman}\tag{3}\] The forward and backward forms agree: \[m_0V_0=A(P_{0:T-1}).\] The proof is a direct unrolling of the recursion and is given in the appendix.
The key term in the Bellman recursion 3 is \(P_tV_{t+1}\). Its \(i\)th component is \[(P_tV_{t+1})(i)=\sum_{j=1}^s p_{ij,t}V_{t+1}(j).\] This is the expected outcome remaining after one transition from current state \(S_i\). If \(P_t\) is fixed, this weighted average is fixed. If several transition matrices remain compatible with the evidence, the same weighted average has a lower and an upper value. That observation is the computational link between Markov cohort models and lower transition operators.
For each time step \(t=0,\ldots,T-1\), let \(\mathcal{P}_t\) be a non-empty set of admissible transition probability matrices. A transition-matrix sequence is admissible if \[P_{0:T-1}\in\mathcal{A}:=\mathcal{P}_0\times\cdots\times\mathcal{P}_{T-1}.\] The imprecise model is the set of precise Markov cohort models generated by all sequences in \(\mathcal{A}\). For any additive outcome \(A(P_{0:T-1})\) defined in 2 , the lower and upper accumulated outcomes are \[\underline A=\inf_{P_{0:T-1}\in\mathcal{A}}A(P_{0:T-1}), \qquad \overline{A}=\sup_{P_{0:T-1}\in\mathcal{A}}A(P_{0:T-1}). \label{eq:envelope}\tag{4}\] These are lower and upper expectations over the admissible set of precise Markov models.
The tractable case used for the formal results is the row-separable, or separately specified, one, stated in the following assumption.
Assumption 1 (Compact separately specified outgoing rows). For every \(t=0,\ldots,T-1\) and every \(i=1,\ldots,s\), \(\mathcal{K}_{i,t}\) is a non-empty compact set of complete outgoing rows \[p_{i,t}=(p_{i1,t},\ldots,p_{is,t}),\qquad p_{ij,t}\ge0,\qquad \sum_{j=1}^s p_{ij,t}=1.\] The admissible matrix set at time step \(t\) is \[\mathcal{P}_t=\{P_t:P_{t,i\cdot}\in\mathcal{K}_{i,t},\;i=1,\ldots,s\}. \label{eq:row-separable-set}\qquad{(1)}\]
The interpretation is that one admissible outgoing row is chosen for each current state and the rows are stacked to form a transition matrix. Compactness guarantees that the row-wise infima and suprema below are attained. Separately specified rows are a structural assumption about the admissible set: the admissible outgoing distribution chosen for one state does not constrain the choice for any other state. This assumption can fail when the rows are linked, through any of three mechanisms that recur below. A shared parameter is a single quantity—such as an underlying event rate or a relative risk—that enters more than one row, so those rows cannot be moved independently. A calibration constraint requires the matrix to reproduce an observed aggregate, such as overall survival or disease prevalence, which couples entries across rows. A treatment-effect model derives one intervention’s transitions from another’s by applying an effect measure, such as a relative risk or hazard ratio, so the two arms’ rows are tied together. Whenever rows are linked in any of these ways, ?? does not hold and the linked matrix set must be optimized as a whole.
Remark 1 (Time-step-wise versus fixed-matrix admissibility). The set \(\mathcal{A}=\mathcal{P}_0\times\cdots\times\mathcal{P}_{T-1}\) allows an admissible matrix to be chosen at each time step. This product structure is the standard rectangular interpretation used by the lower and upper Bellman recursions. If the modeling assumption is instead that a single unknown matrix is fixed for all time steps (so that \(\mathcal{P}_0=\cdots=\mathcal{P}_{T-1}=:\mathcal{P}\)), the admissible set is the diagonal \(\{(P,\ldots,P):P\in\mathcal{P}\}\subseteq\mathcal{A}\). Two consequences follow. First, since the diagonal is contained in the rectangular set, the rectangular envelope is a valid outer* bound for the fixed-matrix envelope: the rectangular lower (upper) value is no larger (no smaller) than its fixed-matrix counterpart. Second, the two envelopes coincide whenever the row-wise minimizer and maximizer selected in ?? can be taken to be the same matrix at every time step. A sufficient condition is that the one-step objective is optimized at a time-invariant row; this holds in particular when each row set is governed by a scalar parameter and the accumulated outcome is monotone in that parameter, in which case the optimizing row is pinned to a parameter endpoint at every step. The example in Section 8 satisfies this condition, so its fixed-matrix endpoint evaluation equals the exact rectangular envelope; when no such time-invariant optimizer exists, the fixed-matrix envelope must be computed by global optimization over \(P\).*
For \(f\in\mathcal{L}(\mathcal{S})\), define the lower and upper transition operators \[(\underline{\mathcal{T}}_t f)(i)=\inf_{p_{i,t}\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}f(j), \qquad (\overline{\mathcal{T}}_t f)(i)=\sup_{p_{i,t}\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}f(j), \label{eq:lower-upper-operators-ijar}\tag{5}\] for \(i=1,\ldots,s\). These are lower and upper envelopes of the one-step conditional expectation of \(f(X_{t+1})\), conditional on \(X_t=S_i\).
Proposition 1 (Coherence properties). Under Assumption 1, \(\underline{\mathcal{T}}_t\) is a coherent lower transition operator in the following finite-state sense. For all \(f,h\in\mathcal{L}(\mathcal{S})\), all constants \(c\in\mathbb{R}\), all \(\mu\ge0\), and all states \(S_i\), \[(\underline{\mathcal{T}}_t f)(i)\ge \min_{1\le j\le s} f(j),\qquad \underline{\mathcal{T}}_t(f+h)\ge \underline{\mathcal{T}}_t f+\underline{\mathcal{T}}_t h,\] \[\underline{\mathcal{T}}_t(\mu f)=\mu\underline{\mathcal{T}}_t f,\qquad \underline{\mathcal{T}}_t(f+c)=\underline{\mathcal{T}}_t f+c.\] It is also monotone: if \(f\le h\) componentwise, then \(\underline{\mathcal{T}}_t f\le \underline{\mathcal{T}}_t h\). The upper operator is conjugate to the lower operator: \[\overline{\mathcal{T}}_t f=-\underline{\mathcal{T}}_t(-f).\] The lower bound, super-additivity, and non-negative homogeneity are the defining axioms of a coherent lower prevision in the sense of [4], [6]; constant additivity, monotonicity, and the conjugacy relation are consequences. In particular, combining the lower bound with conjugacy gives the two-sided envelope \[\min_{1\le j\le s} f(j)\le (\underline{\mathcal{T}}_t f)(i)\le (\overline{\mathcal{T}}_t f)(i)\le \max_{1\le j\le s} f(j).\]
Theorem 1 (Envelope Bellman recursion). Assume Assumption 1 holds for every \(t=0,\ldots,T-1\). Define \[\underline V_T=\overline{V}_T=g_T,\qquad \underline V_t=g_t+\underline{\mathcal{T}}_t\underline V_{t+1},\qquad \overline{V}_t=g_t+\overline{\mathcal{T}}_t\overline{V}_{t+1}, \label{eq:imprecise-bellman-ijar}\qquad{(2)}\] for \(t=T-1,\ldots,0\). Then, for every initial cohort vector \(m_0\), \[m_0\underline V_0=\inf_{P_{0:T-1}\in\mathcal{A}}A(P_{0:T-1}), \qquad m_0\overline{V}_0=\sup_{P_{0:T-1}\in\mathcal{A}}A(P_{0:T-1}).\]
The theorem states that the lower and upper Bellman recursions compute the exact lower and upper expectations in 4 . The proof is given in the appendix. The result is the discrete-time analogue of the lower-transition-operator view used in imprecise Markov models [7], [9]. The key simplifying condition is the row-separable structure: the worst outgoing row for one current state can be chosen without constraining the rows for the other current states.
Corollary 1 (Reduction to the classical model). If every \(\mathcal{P}_t\) is a singleton \(\{P_t\}\), then \(\underline{\mathcal{T}}_t f=\overline{\mathcal{T}}_t f=P_tf\) for all \(f\in\mathcal{L}(\mathcal{S})\). Hence ?? reduces to the classical Bellman recursion 3 .
Remark 2 (Coherent lower prevision on accumulated outcomes). Theorem 1 has a lower-prevision reading. Fix \(m_0\) and define the functional \(\underline E(g_{0:T})=m_0\underline V_0\) on outcome sequences \(g_{0:T}\). Because each one-step operator \(\underline{\mathcal{T}}_t\) is a coherent lower transition operator (Proposition 1) and \(m_0\) has non-negative entries summing to one, the composition in ?? preserves the lower bound, super-additivity, and non-negative homogeneity, together with constant additivity and monotonicity. Hence \(\underline E\) is a coherent lower prevision on \(\mathcal{L}(\mathcal{S})^{T+1}\), and \(\overline{E}=-\underline E(-\cdot)\) is its conjugate upper prevision. The envelope is therefore not a heuristic sensitivity range but the lower expectation of a coherent imprecise model, consistent with the operator view of imprecise Markov chains [7], [9].
Some outcomes are incurred when a transition occurs rather than while a state is occupied; an acute event cost charged on entry to an event state is the canonical example. Let \(r_t(i,j)\in\mathbb{R}\) be a one-step transition contribution, interpreted as the outcome assigned to a person who moves from \(S_i\) to \(S_j\) during step \(t\). The accumulated outcome 2 generalizes to \[A(P_{0:T-1})=\sum_{t=0}^{T}m_tg_t+\sum_{t=0}^{T-1}\sum_{i=1}^s\sum_{j=1}^s m_{t,i}\,p_{ij,t}\,r_t(i,j), \label{eq:accumulated-transition}\tag{6}\] and the backward recursion 3 becomes \[V_T=g_T,\qquad V_t(i)=g_t(i)+\sum_{j=1}^s p_{ij,t}\bigl(r_t(i,j)+V_{t+1}(j)\bigr),\qquad t=T-1,\ldots,0, \label{eq:bellman-transition}\tag{7}\] for which the forward–backward equivalence \(m_0V_0=A(P_{0:T-1})\) holds verbatim. For the imprecise model, the only change to the envelope recursion ?? is that the operators act on \(r_t(i,\cdot)+\underline V_{t+1}(\cdot)\) instead of \(\underline V_{t+1}(\cdot)\): \[\underline V_t(i)=g_t(i)+\inf_{p_{i,t}\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}\bigl(r_t(i,j)+\underline V_{t+1}(j)\bigr), \label{eq:imprecise-bellman-transition}\tag{8}\] and analogously for \(\overline{V}_t\) with the supremum. Because \(r_t(i,\cdot)\) is fixed and the objective in 8 is still linear in the row \(p_{i,t}\), Assumption 1 and the proof of Theorem 1 apply unchanged, so 8 computes the exact lower (upper) envelope of 6 . The state-based development of Sections 3–4 is the special case \(r_t\equiv0\). The performance-difference identity of Section 5 extends as well, with the row difference acting on the augmented one-step value \(r_t(i,j)+\widehat V_{t+1}(j)\). The example in Section 8 uses exactly such a contribution for the acute recurrent-stroke cost, so it is an instance of 6 .
For any transition-matrix sequence supplied to it, the classical cohort calculation returns the correct accumulated outcome; the difficulty is not the arithmetic but the input. When a single selected sequence is treated as uniquely justified, the outcome it produces—and the decision that outcome supports—can differ from those of other sequences equally compatible with the evidence. This section makes that dependence precise: a performance-difference identity gives the exact gap between any two admissible models, from which we obtain an algebraic condition for when the selected-matrix result is invariant over the admissible set, and a corresponding test for when it yields a non-robust decision.
Let \(\widehat P_{0:T-1}\in\mathcal{A}\) be the selected sequence used in a classical point analysis, and let \(\widehat V_t\) be the corresponding backward vector from 3 . For any admissible sequence \(Q_{0:T-1}\in\mathcal{A}\), let \(m_t^Q\) be the cohort trajectory generated by \(Q\).
Proposition 2 (Performance-difference identity). For any \(Q_{0:T-1}\in\mathcal{A}\), \[A(Q_{0:T-1})-A(\widehat P_{0:T-1}) = \sum_{t=0}^{T-1}m_t^Q(Q_t-\widehat P_t)\widehat V_{t+1}. \label{eq:performance-difference}\qquad{(3)}\]
Thus the point result is invariant over \(\mathcal{A}\) if and only if the right-hand side of ?? is zero for every admissible \(Q_{0:T-1}\). Row by row, \[m_t^Q(Q_t-\widehat P_t)\widehat V_{t+1} = \sum_{i=1}^s m_{t,i}^Q\sum_{j=1}^s (q_{ij,t}-\widehat p_{ij,t})\,\widehat V_{t+1}(j).\] Unresolved transition probabilities can affect the accumulated outcome only if the current state is reached, the admissible outgoing row differs from the selected row, and the row difference moves probability mass among destination states with different remaining outcomes. This algebraic condition is the precise version of the practical warning that an empirical zero should not be confused with a structural impossibility.
For decisions, let \(D(P^a,P^b;\lambda)\) be the incremental net monetary benefit (INMB) for intervention \(a\) versus \(b\) at willingness-to-pay threshold \(\lambda\). Define \[\underline D=\inf_{(P^a,P^b)\in\mathcal{A}_{ab}}D(P^a,P^b;\lambda), \qquad \overline{D}=\sup_{(P^a,P^b)\in\mathcal{A}_{ab}}D(P^a,P^b;\lambda),\] where \(\mathcal{A}_{ab}\) is the admissible joint set for the two interventions. When the interventions are informed by separate evidence, \(\mathcal{A}_{ab}=\mathcal{A}_a\times\mathcal{A}_b\) is a product set; writing \(\mathrm{NMB}_a\) and \(\mathrm{NMB}_b\) for the net monetary benefit of each intervention, so that \(D=\mathrm{NMB}_a-\mathrm{NMB}_b\), the lower endpoint then factorizes as \[\underline D=\inf_{P^a\in\mathcal{A}_a}\mathrm{NMB}_a-\sup_{P^b\in\mathcal{A}_b}\mathrm{NMB}_b,\] and analogously \(\overline{D}=\sup_{P^a}\mathrm{NMB}_a-\inf_{P^b}\mathrm{NMB}_b\), so the two arms are optimized separately by the recursion of Theorem 1. If the arms share parameters, \(\mathcal{A}_{ab}\) is not a product and \(D\) must be optimized jointly over the linked set.
Proposition 3 (Decision robustness). Suppose the selected matrices recommend intervention \(a\), so that \[D(\widehat P^a,\widehat P^b;\lambda)>0.\] Then the recommendation is robust over \(\mathcal{A}_{ab}\) if \(\underline D>0\). If \(\underline D\le0\), the admissible set does not support a strict recommendation for \(a\) against \(b\). If \(\underline D<0\), at least one admissible pair of transition-matrix sequences favors \(b\).
The previous sections assume that \(\mathcal{P}_t\) is specified. This section gives a data-based construction using multinomial transition counts and the IDM.
Fix current state \(S_i\). Let \(n_{ij}\) be the observed number of transitions from \(S_i\) to \(S_j\), let \(n_i=(n_{i1},\ldots,n_{is})\), and let \(N_i=\sum_{j=1}^s n_{ij}\). For the true outgoing row \(p_i=(p_{i1},\ldots,p_{is})\), the usual row likelihood is \[n_i\mid p_i\sim\mathrm{Multinomial}(N_i,p_i). \label{eq:multinomial-row}\tag{9}\] A classical empirical row uses \(\widehat p_{ij}=n_{ij}/N_i\) when \(N_i>0\). This estimator is convenient but treats observed zeros as zero transition probabilities.
The IDM uses a class of Dirichlet priors \[\{\mathrm{Dirichlet}(\alpha q_1,\ldots,\alpha q_s):q_j>0,\;\sum_{j=1}^s q_j=1\},\] where \(\alpha>0\) is a prior strength. After observing \(n_i\), the posterior class is \[\{\mathrm{Dirichlet}(n_{i1}+\alpha q_1,\ldots,n_{is}+\alpha q_s):q_j>0,\;\sum_{j=1}^s q_j=1\}.\] Optimizing the posterior mean of \(p_{ij}\) over this class gives \[\underline p_{ij}=\frac{n_{ij}}{N_i+\alpha}, \qquad \overline{p}_{ij}=\frac{n_{ij}+\alpha}{N_i+\alpha}, \qquad j=1,\ldots,s. \label{eq:idm-bounds-ijar}\tag{10}\] The induced IDM row set is \[\mathcal{K}_i^{IDM} = \left\{p_i: p_{ij}\ge0,\; \sum_{j=1}^s p_{ij}=1,\; \underline p_{ij}\le p_{ij}\le \overline{p}_{ij},\;j=1,\ldots,s \right\}. \label{eq:idm-row-set-ijar}\tag{11}\] Repeating this construction for all current states and applying structural constraints gives the admissible transition matrices used in ?? . The interval endpoints in 10 are not independent parameters; the simplex condition in 11 preserves the fact that each row must sum to one. The next lemma records that this set is exactly the IDM credal set, so it satisfies Assumption 1 and the upper bounds in 11 are in fact redundant.
Lemma 1 (Exactness of the IDM row set). With \(\underline p_{ij},\overline{p}_{ij}\) as in 10 , the row set is \[\mathcal{K}_i^{IDM} =\Bigl\{p_i:\;p_{ij}\ge \underline p_{ij}\;(j=1,\ldots,s),\;\textstyle\sum_{j=1}^s p_{ij}=1\Bigr\} =\operatorname{conv}\{v_1,\ldots,v_s\},\] where \[v_k=\frac{n_i+\alpha e_k}{N_i+\alpha},\qquad k=1,\ldots,s,\] and \(e_k\) is the \(k\)th standard basis vector. Hence \(\mathcal{K}_i^{IDM}\) is a non-empty compact polytope (the set of IDM posterior means), and the upper bounds \(p_{ij}\le\overline{p}_{ij}\) follow from the lower bounds and the simplex constraint.
Proof. If \(\sum_j p_{ij}=1\) and \(p_{ij}\ge\underline p_{ij}\) for all \(j\), then for each \(k\), \(p_{ik}=1-\sum_{j\ne k}p_{ij}\le 1-\sum_{j\ne k}\underline p_{ij}=1-\tfrac{N_i-n_{ik}}{N_i+\alpha}=\tfrac{n_{ik}+\alpha}{N_i+\alpha}=\overline{p}_{ik}\), so the upper bounds are redundant. Setting \(w_k=(p_{ik}-\underline p_{ik})(N_i+\alpha)/\alpha\) gives \(w_k\ge0\), \(\sum_k w_k=\bigl(1-N_i/(N_i+\alpha)\bigr)(N_i+\alpha)/\alpha=1\), and \(p_i=\sum_k w_k v_k\); conversely every such convex combination meets the constraints. Thus the three sets coincide. ◻
The framework does not depend on the IDM. Any non-empty compact row set \(\mathcal{K}_{i,t}\)—for instance one defined by a coherent lower prevision, a credal set from linear probability constraints, interval probabilities, or an \(\epsilon\)-contamination model—can be used in 5 and Theorem 1. We use the IDM in the example because it maps transition counts directly to row sets.
We provide a computational algorithm and the accompanying code as supplementary files. The computation follows directly from Theorem 1. For a selected outcome contribution \(g_t\), set \(\underline V_T=\overline{V}_T=g_T\). For \(t=T-1,\ldots,0\), compute the lower and upper one-step expectations in 5 and update \(\underline V_t,\overline{V}_t\) using ?? . The final interval for the initial cohort is \[[m_0\underline V_0,\;m_0\overline{V}_0].\]
For each state \(S_i\), the optimization in 5 is a linear program: \[\min_{p_{i,t}\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}f(j) \quad\text{or}\quad \max_{p_{i,t}\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}f(j).\] For IDM row sets, the constraints are linear: bounds, non-negativity, and summing to one. Hence the computation is transparent and scales linearly in the number of time steps once the row-level linear programs are solved.
The algorithm is as follows:
Define the states, initial distribution, time horizon, and outcome contributions \(g_t\).
Estimate an empirical transition matrix for comparison.
Apply the lower and upper Bellman recursion ?? .
Report the empirical point result together with the lower and upper outcomes.
For decisions, optimize INMB directly rather than optimizing costs and effects separately.
We use a cost-effectiveness example to demonstrate how the proposed method can change a decision. The clinical question is whether patent foramen ovale (PFO) closure is cost-effective compared with medical therapy alone for secondary prevention after cryptogenic stroke. Published economic evaluations have generally found PFO closure cost-effective or cost-saving over longer horizons [16]. The specific aim of this application is to examine the effect of sparse recurrent-stroke evidence near a decision boundary. The code for this example is included as supplementary material. We also conduct an independent verification: rather than evaluating the model at the interval endpoints, we use the full backward recursion of Theorem 1, applying the lower and upper transition operators at every time step in the transition-reward form 8 to assign a recurrent-stroke cost to each transition.
The example uses recurrent-stroke data from two high-risk PFO trials (Table 1). In CLOSE, no stroke occurred among 238 patients assigned to PFO closure, whereas 14 strokes occurred among 235 patients assigned to antiplatelet therapy [17]. In DEFENSE-PFO, 60 patients were enrolled in each arm; ischemic stroke occurred in 0% with closure and 10.5% with medical therapy [18]. We represent the latter as six events in the medical-therapy arm. Patient-years are approximated as the trial size times the follow-up period.
| Trial | Strategy | Patients | Follow-up | Events | PY |
|---|---|---|---|---|---|
| CLOSE | PFO closure | 238 | 5.3 | 0 | 1261.4 |
| CLOSE | Medical therapy | 235 | 5.3 | 14 | 1245.5 |
| DEFENSE-PFO | PFO closure | 60 | 2.8 | 0 | 168.0 |
| DEFENSE-PFO | Medical therapy | 60 | 2.8 | 6 | 168.0 |
| Pooled | PFO closure | 298 | – | 0 | 1429.4 |
| Pooled | Medical therapy | 295 | – | 20 | 1413.5 |
Let \(p_{\mathrm{stroke}}\) be the one-time-step probability of recurrent stroke. The empirical probability is 0 for PFO closure and 0.01415 for medical therapy. With IDM prior strength \(\alpha=2\), the IDM intervals are \[0\le p_{\mathrm{stroke}}^{\mathrm{closure}}\le 0.00140, \qquad 0.01413\le p_{\mathrm{stroke}}^{\mathrm{medical}}\le 0.01554.\] The closure-arm upper bound is small, but it is not zero. This nonzero upper bound is the distinction the empirical matrix hides. Here the patient-years at risk play the role of the multinomial denominator \(N_i\) in 10 , so \(p_{\mathrm{stroke}}\) is read as a per-time-step rate; for the small magnitudes involved, the difference between a rate and a one-step probability (\(p\approx 1-e^{-\text{rate}}\)) is negligible. Imprecision is placed only on the recurrent-stroke probability, with the background death probabilities held fixed, so the admissible “no recurrent stroke” row is the one-parameter subfamily \(\{(1-p-d_0,\,p,\,d_0):p\in[\underline p,\overline{p}]\}\) of the full IDM simplex of Lemma 1.
The model has three states: no recurrent stroke, after recurrent stroke, and dead. Everyone begins in the no-recurrent-stroke state. The horizon is seven years with 3% discounting. The willingness-to-pay threshold is \(\lambda=\$150{,}000\) per quality-adjusted life-year (QALY), matching the threshold used in the published PFO economic evaluation [16]. Table 2 gives the compact model inputs.
| Model parameters | Value |
|---|---|
| States | no recurrent stroke; after recurrent stroke; dead |
| Initial distribution | \(m_0=(1,0,0)\) |
| Horizon | 7 years |
| Discount rate | 3% per time step |
| Willingness-to-pay threshold | \(\$150{,}000\) per QALY |
| One-time closure cost | \(\$15{,}000\) |
| Annual cost without recurrent stroke | \(\$200\) |
| Acute recurrent-stroke cost | \(\$40{,}000\) |
| Annual post-stroke cost | \(\$8{,}000\) |
| Utility without recurrent stroke | 0.86 |
| Utility after recurrent stroke | 0.65 |
| Background death probability | 0.002 before recurrent stroke; 0.050 after recurrent stroke |
The empirical matrix sets \(p_{\mathrm{stroke}}^{\mathrm{closure}}=0\). A precise symmetric Dirichlet posterior mean with strength \(\alpha=2\) gives \(p_{\mathrm{stroke}}^{\mathrm{closure}}=0.00070\) and \(p_{\mathrm{stroke}}^{\mathrm{medical}}=0.01484\). The imprecise analysis evaluates all matrices with recurrent-stroke probabilities in the IDM intervals above.
Table 3 reports discounted incremental net monetary benefit results for PFO closure versus medical therapy. The empirical transition matrix slightly favors closure: INMB is \(\$167\). The precise Dirichlet posterior mean gives the same qualitative conclusion. The IDM analysis changes the interpretation. Its INMB interval crosses zero, from \(-\$1{,}387\) to \(\$1{,}618\) (Figure 1). Thus the empirical matrix recommends closure, but the evidence-compatible matrix set does not support a robust recommendation at this threshold. Because INMB is
monotone in each arm’s single recurrent-stroke probability, the optimizing transition row is pinned to an interval endpoint at every time step; by Remark 1 the
endpoint evaluation reported here coincides with the exact lower and upper envelopes of Theorem 1. The supplementary script bellman_envelope.R
confirms this by computing the same interval directly from the lower and upper transition operators 8 .
| Analysis | Incremental QALYs | Incremental cost | INMB |
|---|---|---|---|
| Empirical transition matrix | 0.0658 | \(\$9{,}705\) | \(\$167\) |
| Precise Dirichlet posterior mean | 0.0656 | \(\$9{,}727\) | \(\$115\) |
| IDM lower bound | 0.0591 | \(\$10{,}251\) | \(-\$1{,}387\) |
| IDM upper bound | 0.0721 | \(\$9{,}203\) | \(\$1{,}618\) |
The non-robust conclusion is not an artifact of the particular threshold or the IDM strength. Holding \(\alpha=2\) and sweeping the willingness-to-pay threshold \(\lambda\) (Table 4), the empirical point INMB changes sign at \(\lambda\approx\$147{,}500\) per QALY, whereas the lower and upper envelopes change sign at \(\lambda\approx\$173{,}500\) and \(\lambda\approx\$127{,}600\). The IDM interval therefore straddles zero throughout the band \(\$127{,}600<\lambda<\$173{,}500\): PFO closure is a robust choice only above this band and medical therapy only below it. At the \(\$150{,}000\) threshold used in the published evaluation [16] the decision is non-robust, and at the lower thresholds \(\$50{,}000\) and \(\$100{,}000\) the imprecise analysis robustly favors medical therapy, even though the empirical point INMB is positive at \(\$150{,}000\).
The interval width scales with the prior strength, as expected. Recomputing at the other conventional IDM strength \(\alpha=1\) gives the INMB interval \([-\$611,\;\$894]\) at \(\lambda=\$150{,}000\), against \([-\$1{,}387,\;\$1{,}618]\) at \(\alpha=2\); both straddle zero, so the non-robustness does not depend on the choice of \(\alpha\).
| Threshold \(\lambda\) ($/QALY) | Lower INMB | Upper INMB | Robustly preferred strategy |
|---|---|---|---|
| 50,000 | \(-\$7{,}297\) | \(-\$5{,}596\) | Medical therapy |
| 100,000 | \(-\$4{,}342\) | \(-\$1{,}989\) | Medical therapy |
| 150,000 | \(-\$1{,}387\) | \(\$1{,}618\) | None (non-robust) |
| 200,000 | \(\$1{,}568\) | \(\$5{,}225\) | PFO closure |
This example is deliberately compact: it does not claim that PFO closure is cost-ineffective in general, but shows that near a decision boundary a zero event count in a finite trial can be decisive in the empirical transition matrix. The imprecise analysis distinguishes an observed zero from a structural zero and makes the resulting non-robustness visible.
To show that the phenomenon is not tied to the pooled sample size, we study the three analyses as the amount of evidence varies. We fix a ground truth in which closure genuinely lowers the per-step recurrent-stroke probability, from \(0.014\) under medical therapy to \(0.010\) under closure, but not by enough to justify its one-time cost: the true INMB is \(-\$10{,}790\), so the correct decision is medical therapy. For each arm we draw recurrent-stroke counts from a Poisson model with mean \(N\times p\), where \(N\) is the patient-years at risk per arm, and we vary \(N\) from \(25\) to \(3200\). For each simulated data set we form the empirical matrix, the precise symmetric Dirichlet posterior mean (\(\alpha=2\)), and the IDM interval (\(\alpha=2\)), and record the decision each implies at \(\lambda=\$150{,}000\) per QALY. The two point analyses always commit to closure or to medical therapy; the IDM analysis additionally abstains when its INMB interval contains zero. Figure 2 reports, over \(16{,}000\) replications per sample size, how often each analysis is wrong and how often the IDM abstains.
Two patterns emerge. First, under sparsity the point analyses are frequently confident and wrong: the empirical matrix recommends the inferior arm in up to about a third of samples near \(N=50\), driven by closure-arm sampling zeros that mimic a highly effective treatment, and the precise Dirichlet posterior mean behaves almost identically, so its symmetric prior offers little protection in this regime. Second, the IDM analysis is almost never confidently wrong—its error rate stays below \(5\%\) throughout—because it abstains instead, reporting that the evidence does not resolve the decision. As the sample size grows the sampling zeros disappear, the abstention rate falls to zero, and all three analyses converge to the correct recommendation. The IDM thus trades the point analyses’ confident errors for transparent abstentions and commits to the correct decision once the evidence can support it.
This convergence has a simple analytical counterpart. From the IDM bounds 10 , each admissible row spans an interval of half-width \(\alpha/(N_i+\alpha)=O(\alpha/N_i)\), where \(N_i\) is the patient-years informing that row. Because the accumulated outcome is linear in each row, the INMB interval half-width—and hence the width of the non-robust threshold band of Section 8.4—is \(O(\alpha/N)\) in the per-arm patient-years \(N\). The imprecision therefore contracts at rate \(1/N\) as evidence accumulates, which is the formal counterpart of the convergence seen in Figure 2.
The paper develops a lower-expectation formulation of the Markov cohort model with imprecise transition probabilities. Rather than attach a single value to one selected matrix, the method reports the lower and upper expectation of an additive accumulated outcome over an evidence-compatible set of precise cohort models. The classical cohort model is recovered exactly when the admissible set is a singleton at every time step (Corollary 1); when the admissible set is larger, the lower and upper Bellman recursions of Theorem 1 return the exact range of the outcome over all admissible models.
The central result is the envelope theorem (Theorem 1), together with the coherence properties of Proposition 1. Under separately specified rows, the imprecise calculation is not a heuristic sensitivity analysis over hand-picked matrices; it is the exact lower and upper expectation of a coherent imprecise model in the sense of [4], [6], and it coincides with the lower- and upper-transition-operator view of imprecise Markov chains [7], [9]. Because the envelope is a coherent lower prevision, interval-valued cohort outcomes stand on the same footing as the point-valued outcomes of the classical model, and the bounds can be read as genuine expectations rather than as the spread of an ad hoc collection of scenarios.
The performance-difference identity of Proposition 2 provides an exact diagnostic that complements these bounds. A single selected matrix can mislead only when imprecision both reaches states that the cohort actually occupies and redistributes probability mass among states whose remaining outcomes differ. Imprecision confined to unreachable states, or that moves mass between states with the same continuation value, leaves the decision unchanged. The diagnostic thus identifies in advance where transition uncertainty can and cannot overturn a recommendation, and explains why its consequences are concentrated in a few influential transitions rather than spread uniformly across the matrix.
The method is a modest extension of the classical cohort model rather than a departure from it. An analyst can begin from the same transition counts used to build a classical cohort model, construct IDM row sets, run the lower and upper recursions at essentially the cost of a few point evaluations, and report whether the empirical recommendation is robust to the dependence and sampling uncertainty the data leave unresolved. The PFO closure example shows that this is not a formality: the empirical transition matrix and the precise Dirichlet posterior mean both favor closure, yet the evidence-compatible interval straddles the decision boundary, and it continues to do so across the conventional range of willingness-to-pay thresholds and across both conventional IDM strengths. The sparsity experiment of Section 8.5 shows that the interval behaves like a calibrated decision rule: it abstains when the evidence is too thin to support either arm, is almost never confidently wrong, and contracts at rate \(O(\alpha/N)\), so that it recovers the point decision as evidence accumulates.
There are two limitations which restrict the scope of the study results, though without the risk of undermining the rigor and utility of the approach in its intended setting. The first is the row-separable structure of Assumption 1, which is what makes the envelope recursion tractable: it lets the worst outgoing row for one state be chosen independently of the rows for the other states. Row separability is appropriate when outgoing rows are assessed separately by current state, which is the common case in applied cohort models, where each state’s transitions are estimated from its own data. When the rows are instead linked—by a shared parameter, a calibration constraint, or a treatment-effect model—the approach does not break down: the lower and upper expectations remain well defined over the linked set, and, because any linked set is contained in the rectangular product set, the row-by-row envelope is still a valid outer bound for the linked problem (cf.Remark 1) and can serve as a conservative screen before any costlier joint optimization. Only the row-wise decomposition, not the lower-expectation interpretation, is lost. The second limitation is that the IDM is only one way to convert evidence into an admissible set. Here too the core results are unaffected: the envelope theorem, the coherence of the lower and upper operators, and the performance-difference diagnostic hold for any non-empty compact row sets, so likelihood-based confidence regions, expert-elicited bounds, an imprecise treatment of Bayesian credible sets, or calibration-defined sets all plug into the same recursion. The IDM is simply the construction that matches multinomial transition counts; only its closed-form bounds and the \(O(\alpha/N)\) contraction rate are specific to that setting.
Several extensions follow naturally. Linked transition-matrix sets, in which rows co-vary through shared parameters or calibration constraints, would widen applicability at the cost of row-wise tractability and call for dedicated optimization algorithms, particularly for large state spaces where enumerating rows is impractical. A second direction is to combine imprecise transition matrices with parameter uncertainty analysis. The two answer different questions and are complementary rather than competing: imprecise transition matrices answer a robustness question—whether a recommendation survives every evidence-compatible dependence structure—whereas parameter uncertainty analysis answers a distributional question about the spread of outcomes under an assumed joint distribution of inputs. Reporting both a robustness interval and a distributional summary would give decision makers a fuller account of the uncertainty surrounding a transition-probability-driven decision.
Rowan Iskandar: Conceptualization, Methodology, Formal analysis, Software, Writing–original draft, Writing–review and editing.
The author declares that he has no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
No specific funding was received for this methodological work.
The case-study calculations are reproducible from the R scripts and Quarto guide supplied as supplementary files.
The author thanks colleagues and reviewers who provided feedback on earlier versions of this manuscript.
During the preparation of this work the author used Claude (a large language model developed by Anthropic) in order to improve the language, readability, and formatting of portions of the manuscript. After using this tool, the author reviewed and edited the content as needed and takes full responsibility for the content of the publication.
Unroll 3 . Since \(V_T=g_T\), \[V_0=g_0+P_0g_1+P_0P_1g_2+\cdots+P_0P_1\cdots P_{T-1}g_T.\] Multiplying by \(m_0\) gives \[m_0V_0 = m_0g_0+m_0P_0g_1+\cdots+m_0P_0P_1\cdots P_{T-1}g_T.\] Because \(m_t=m_0P_0\cdots P_{t-1}\), with the empty product interpreted as the identity for \(t=0\), the right-hand side is \(\sum_{t=0}^{T}m_tg_t=A(P_{0:T-1})\).
Fix \(t\) and \(i\). For any \(p_{i,t}\in\mathcal{K}_{i,t}\), \[\sum_{j=1}^s p_{ij,t}f(j)\ge \min_{1\le j\le s}f(j),\] which gives the lower bound after taking the infimum. For superadditivity, \[\inf_{p\in\mathcal{K}_{i,t}}p(f+h) \ge \inf_{p\in\mathcal{K}_{i,t}}pf+\inf_{p\in\mathcal{K}_{i,t}}ph,\] where \(pf\) denotes \(\sum_{j=1}^s p_jf(j)\). Non-negative homogeneity follows because, for \(\mu\ge0\), \[\inf_{p\in\mathcal{K}_{i,t}}p(\mu f)=\mu\inf_{p\in\mathcal{K}_{i,t}}pf.\] Constant additivity follows from \(p\mathbf{1}=1\) for every probability row \(p\). Monotonicity follows because \(f\le h\) implies \(pf\le ph\) for every admissible row. Finally, \[-\underline{\mathcal{T}}_t(-f)(i) = -\inf_{p\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}\{-f(j)\} = \sup_{p\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}f(j) = \overline{\mathcal{T}}_t f(i).\]
We prove the lower recursion; the upper recursion is identical with infima replaced by suprema. For an admissible sequence \(P_{0:T-1}\in\mathcal{A}\), let \(V_t^P\) denote the classical backward vector 3 , so that \(m_0V_0^P=A(P_{0:T-1})\) by the forward–backward equivalence. We establish two facts: (i) \(\underline V_t\le V_t^P\) componentwise for every admissible \(P\) and every \(t\); and (ii) there exists an admissible \(P^\star\) with \(V_t^{P^\star}=\underline V_t\). Both use only Assumption 1.
Lower bound. Backward induction on \(t\). At \(t=T\), \(\underline V_T=g_T=V_T^P\). Assume \(\underline V_{t+1}\le V_{t+1}^P\). Each row \(p_{i,t}\) of \(P_t\) has non-negative entries, so for every state \(S_i\), \[\begin{align} V_t^P(i)=g_t(i)+\sum_{j=1}^s p_{ij,t}V_{t+1}^P(j) &\ge g_t(i)+\sum_{j=1}^s p_{ij,t}\underline V_{t+1}(j)\\ &\ge g_t(i)+\inf_{p_{i,t}\in\mathcal{K}_{i,t}}\sum_{j=1}^s p_{ij,t}\underline V_{t+1}(j) =\underline V_t(i), \end{align}\] where the first inequality uses the induction hypothesis with \(p_{ij,t}\ge0\). Hence \(\underline V_t\le V_t^P\) for all \(t\). Since \(m_0\) has non-negative entries, \(A(P_{0:T-1})=m_0V_0^P\ge m_0\underline V_0\) for every admissible \(P\), and therefore \(m_0\underline V_0\le\inf_{P_{0:T-1}\in\mathcal{A}}A(P_{0:T-1})\).
Attainment. By Assumption 1 the map \(p_{i,t}\mapsto\sum_{j=1}^s p_{ij,t}\underline V_{t+1}(j)\) is linear, hence continuous, on the compact set \(\mathcal{K}_{i,t}\), so a minimizer \(p^\star_{i,t}\in\mathcal{K}_{i,t}\) exists. Because the rows are separately specified ?? , stacking \(p^\star_{1,t},\ldots,p^\star_{s,t}\) yields a matrix \(P^\star_t\in\mathcal{P}_t\); repeating this for every time step gives \(P^\star_{0:T-1}\in\mathcal{A}\). A second backward induction shows \(V_t^{P^\star}=\underline V_t\): it holds at \(t=T\), and if \(V_{t+1}^{P^\star}=\underline V_{t+1}\) then \[V_t^{P^\star}(i)=g_t(i)+\sum_{j=1}^s p^\star_{ij,t}\underline V_{t+1}(j)=g_t(i)+(\underline{\mathcal{T}}_t\underline V_{t+1})(i)=\underline V_t(i).\] Therefore \(A(P^\star_{0:T-1})=m_0V_0^{P^\star}=m_0\underline V_0\), the infimum is attained, and combined with the lower bound, \(m_0\underline V_0=\inf_{P_{0:T-1}\in\mathcal{A}}A(P_{0:T-1})\). The minimizer \(P^\star\) is constructed without reference to \(m_0\), so the same sequence attains the lower value for every initial distribution.
If \(\mathcal{P}_t=\{P_t\}\), each \(\mathcal{K}_{i,t}\) contains the single row \(P_{t,i\cdot}\). Therefore both the infimum and supremum in 5 equal \(\sum_{j=1}^s p_{ij,t}f(j)\), which is \((P_tf)(i)\). Substituting into ?? gives 3 .
Let \(V_t^{\widehat P}\) be the backward vector under \(\widehat P\). For any admissible sequence \(Q\), \[A(Q)-A(\widehat P)=m_0(V_0^Q-V_0^{\widehat P}).\] Using the backward recursions, \[V_t^Q-V_t^{\widehat P} = Q_t(V_{t+1}^Q-V_{t+1}^{\widehat P})+(Q_t-\widehat P_t)V_{t+1}^{\widehat P}.\] Multiplying by \(m_t^Q\) and using \(m_{t+1}^Q=m_t^QQ_t\) yields \[m_t^Q(V_t^Q-V_t^{\widehat P}) = m_{t+1}^Q(V_{t+1}^Q-V_{t+1}^{\widehat P}) + m_t^Q(Q_t-\widehat P_t)V_{t+1}^{\widehat P}.\] Summing over \(t=0,\ldots,T-1\) telescopes. The terminal difference is zero because \(V_T^Q=V_T^{\widehat P}=g_T\). This gives ?? .
Under Assumption 1 the joint admissible set \(\mathcal{A}_{ab}\) is compact and the map \((P^a,P^b)\mapsto D(P^a,P^b;\lambda)\) is continuous, so by the extreme value theorem the infimum \(\underline D\) is attained at some admissible pair. If \(\underline D>0\), then \(D(P^a,P^b;\lambda)\ge\underline D>0\) for every admissible pair, so intervention \(a\) has strictly greater net monetary benefit throughout \(\mathcal{A}_{ab}\) and the recommendation is robust. If \(\underline D\le0\), a minimizing pair satisfies \(D\le0\), so the strict guarantee fails. If moreover \(\underline D<0\), that minimizing pair has \(D(P^a,P^b;\lambda)<0\) and hence favors intervention \(b\).