A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning


1 Introduction↩︎

In risk-sensitive decision-making, it is important to specify which risk objective the agent aims to optimize. This identification problem is crucial in personalized applications such as robo-advising ([1], [2]) and autonomous driving ([3]), where agents may exhibit heterogeneous and complex risk preferences. Existing studies often focus on mean-variance objectives or coherent risk measures, which cover only part of the range of risk preferences observed in practice. Real-world preferences may involve quantiles, deviation-based objectives, asymmetric tail concerns, or dispersion control, which often fall outside the convex or coherent class. In quantitative risk management, [4] proposed distortion riskmetrics as a unified framework covering many risk measures, deviation measures, and other functionals for financial risk. [5] proposed a sensitivity estimator for the distortion risk measure. [6] further introduced conditional distortion risk measures and distortion risk contribution measures to quantify systemic risk. Despite the theoretical generality and unifying power of distortion riskmetrics, eliciting the appropriate risk objective remains challenging, since agents are often unable to precisely specify their own risk preferences. A common approach is to infer such preferences from agents’ observed behavior. Inverse reinforcement learning (IRL) estimates an agent’s objective from its policy ([7]) and has been applied to autonomous driving ([3], [8], [9]), legged robot locomotion ([10], [11]), and robo-advising ([12]). While these frameworks are promising, they assume that agents’ choices are always consistent with their risk preferences, which may not hold in practice.

Once the risk objective is identified, the next step is to optimize decisions under that objective. Risk-sensitive reinforcement learning (RSRL) has attracted growing attention for decision making under uncertainty. Unlike traditional risk-neutral RL, which optimizes expected cumulative rewards, RSRL explicitly accounts for the variability and tail behavior of returns, aiming to mitigate potentially catastrophic outcomes. Prior work has demonstrated its effectiveness in option hedging ([13], [14]), statistical arbitrage ([15], [16]), marketing ([17]), pandemic control ([18]), autonomous driving ([19], [20], [21]), robotics and control ([22]), and other domains. However, existing RSRL methods mainly focus on a limited set of risk measures or scoring functions, especially convex or coherent ones; see [16]. This leaves a gap between the broad class of risk objectives that may be elicited from agents and the narrower objectives currently supported by RL algorithms. Integrating general distortion riskmetrics into RL therefore provides a natural way to connect risk-objective identification with risk-sensitive policy optimization.

In this paper, we develop a noise-robust elicit-to-optimize framework that integrates IRL and RL for risk-preference elicitation and risk-sensitive decision making under a broad class of distortion riskmetrics. The framework first elicits the agents’ latent risk preferences from their noisy observed choices and then uses the elicited distortion riskmetrics as the objectives for subsequent policy optimization. To reduce the cognitive burden on agents, we collect agents’ choices through binary-choice questions. An adaptive framework ([23], [12]) is then applied, which dynamically adjusts the questions according to the previous responses. We establish the existence of a finite set of distinguishing questions, characterize their discrimination power, and derive an upper bound for the power. Since agents may act inconsistently with their true risk preferences, the observed choices are inherently noisy. To address this issue, we propose a Bayesian IRL method that identifies the agent’s preferred distortion riskmetric from stochastic choices. We prove that, under mild conditions, our algorithm converges to the agents’ preferred distortion riskmetrics at the rate \(O\left(\exp\left(-cm+O\left(\sqrt{m\log m}\right)\right)\right)\), where \(c>0\) is a constant and \(m\) denotes the number of adaptive questions. Compared with [12], whose IRL method achieves an inverse-square-root convergence rate up to logarithmic factors because it identifies more structural components, our approach is complementary: it covers a broad class of distortion riskmetrics and explicitly accounts for stochastic decision noise. Given the elicited risk objective, we develop a model-free RL algorithm that minimizes costs under the corresponding conditional distortion riskmetric. By representing the objective as a definite integral of the cost quantile function with respect to the distortion function on \([0,1]\), the RL method unifies all distortion-riskmetric objectives. We extend the Proximal Policy Optimization (PPO) algorithm ([24]) by introducing three neural networks into the RL framework: a policy network, a value network, and a quantile network. Unlike [14], which incorporates a VaR network to estimate a fixed tail quantile, our quantile network approximates the entire cost quantile function, allowing general distortion-riskmetric objectives to be evaluated by midpoint Riemann sums. By considering conditional distortion riskmetrics, the proposed RL method accounts for scenario-dependent losses and strengthens risk control across heterogeneous environments.

Our contribution can be summarized as follows: (i) We extend the conditional distortion risk measures studied in [6] to the broader class of distortion riskmetrics by allowing the distortion function to be non-monotone. (ii) We propose a Bayesian IRL framework for identifying a broad class of distortion riskmetrics through adaptive interactive questioning. The method accommodates stochastic choices and therefore applies to settings in which agents do not always act consistently with their latent risk preferences. Under mild conditions, we prove that the proposed IRL algorithm converges at the rate \(O\left(\exp\left(-cm+O\left(\sqrt{m\log m}\right)\right)\right)\), where \(c>0\) is a constant and \(m\) denotes the number of adaptive questions. Extensive numerical experiments verify the theoretical results and effectiveness of our IRL algorithm. (iii) We develop a model-free RL algorithm based on an extension of the Proximal Policy Optimization (PPO), for optimizing objectives defined by conditional distortion riskmetrics. Our algorithm accommodates a wide range of conditional distortion-riskmetric objectives, thereby capturing diverse attitudes toward downside risk, variability, and tail events within a unified framework. Moreover, by incorporating scenario dependence into the risk objective, the resulting policies perform effectively in complex and dynamic environments. Numerical experiments further validate the performance of the proposed RL algorithm.

The rest of the paper is organized as follows. Section 2 introduces the notion and properties of distortion riskmetrics and conditional distortion riskmetrics. Section 3 establishes the IRL algorithm for eliciting the true underlying risk preference, provides important quantities for the theoretical results, and proves the convergence rate of our algorithm. Section 4 develops an RL algorithm to solve the decision optimization problem under different distortion-riskmetric objectives. Section 5 illustrates the performances of our proposed algorithms by numerical experiments, and Section 6 concludes the paper.

2 Notation and Preliminaries↩︎

Let \((\Omega,\mathcal{F}, \mathbb{P})\) be an atomless probability space. Let \((\mathcal{F}_t)_{t\ge 0}\) denote the information available up to time \(t\). For any \(p\in[1,\infty)\), \(L^p\) is the space of random variables with finite \(p\)-th moment, and \(L^\infty\) is the space of essentially bounded random variables. Let \(\mathcal{X}\supset L^\infty\) be a convex cone and \(\mathcal{M}\) denote the set of distribution functions of random variables in \(\mathcal{X}\). For \(F\in\mathcal{M}\), \(X\sim F\) means that \(X\in\mathcal{X}\) has distribution \(F\). Denote by \(F_X\) the distribution function of the random variable \(X\). Equality in distribution is denoted by \(X \overset{\mathrm{d}}{=} Y\). Accordingly, \(X \overset{\mathrm{d}}{\ne} Y\) indicates that \(X\) and \(Y\) have different distributions. We define the left-continuous generalized inverse of \(F\) (left-quantile) as \[F^{-1}(t)=\inf\{x\in\mathbb{R}:F(x)\geq t\},\quad t\in (0,1],\] while the right-continuous generalized inverse of \(F\) (right-quantile) is defined as \[F^{-1+}(t)=\inf\{x\in\mathbb{R}:F(x)> t\},\quad t\in [0,1).\] Conventionally, we define \(F^{-1}(0)= F^{-1+}(0)\) and \(F^{-1+}(1)=F^{-1}(1)\).

The conditional distortion riskmetric is defined in terms of the distortion riskmetric studied in [4] on a general space denoted by \[\mathcal{H}= \{h: h \text{ maps }[0,1] \text{ to } \mathbb{R},~ h \text{ is of bounded variation, } h(0)=0 \}.\] For convenience, we include the definition of the distortion riskmetric below.

Definition 1. A functional \(\rho_h:\mathcal{X}\to\mathbb{R}\), whose domain \(\mathcal{X}\supset L^\infty\) is a law-invariant convex cone, is a distortion riskmetric* if there exists \(h\in \mathcal{H}\) such that \(\rho_h(X)=\int X\,\mathrm{d}h\circ\mathbb{P}\), where \(\int X\,\mathrm{d}h\circ\mathbb{P}\) is a signed Choquet integral defined by \[\label{eq:distortion-riskmetric-unconditional} \int X\,\mathrm{d}h\circ\mathbb{P}=\int_{-\infty}^{0}\left(h(\mathbb{P}(X\geq x))-h(1)\right)\,\mathrm{d}x+\int_0^\infty h(\mathbb{P}(X\geq x))\,\mathrm{d}x.\tag{1}\] The function \(h\) is called the distortion function of \(\rho_h\).*

For a given distortion riskmetric, the distortion function is unique. For example, if \(h(t)=\mathbf{1}_{\{t>1-\alpha\}}\), the distortion riskmetric is Value-at-Risk (VaR\(_\alpha\)): \(\rho_h(X)=F^{-1}_X(\alpha)\), where \(\alpha\in(0,1)\).

More examples can be found in [4]. In the following, we provide the definition of a conditional distortion riskmetric.

Definition 2. Let \(\mathcal{C}:=\{C\in\mathcal{F}:\mathbb{P}(C)>0\}.\) For \(C\in\mathcal{C}\), define \(\mathbb{P}_C(A):=\mathbb{P}(A\mid C),~ A\in\mathcal{F}.\) A conditional distortion riskmetric \(\rho_{h|\cdot}:\mathcal{X}\times\mathcal{C}\to\mathbb{R}\) is defined by \[\rho_{h|\cdot}(X,C)=\rho_{h|C}(X):=\int X\,\,\mathrm{d}h\circ\mathbb{P}_C.\] Equivalently, \[\begin{align} \rho_{h|C}(X) = & \int_{-\infty}^{0} \left( h\left( \mathbb{P}(X\ge x\mid C) \right) -h(1) \right)\,\mathrm{d}x + \int_{0}^{\infty} h\left( \mathbb{P}(X\ge x\mid C) \right)\,\mathrm{d}x . \end{align}\]

Below we obtain the quantile representation of conditional distortion riskmetrics, which can be seen as the parallel results of Lemma 1 in [4]. For \(X\in L^0\), denote by \(F_{X|C}\) the conditional distribution function of \(X\) under \(\mathbb{P}_C\).

Lemma 1. For \(h\in\mathcal{H}\), \(C\in\mathcal{C}\), and \(X\in L^0\) such that \(\rho_{h|C}(X)\) is well-defined, possibly taking values in \(\pm\infty\), the following statements hold.

(i) If \(h\) is right-continuous, then \[\rho_{h|C}(X) = \int_0^1 F_{X|C}^{-1+}(1-u)\,\,\mathrm{d}h(u).\]

(ii) If \(h\) is left-continuous, then \[\rho_{h|C}(X) = \int_0^1 F_{X|C}^{-1}(1-u)\, \,\mathrm{d}h(u).\]

(iii) If \(F_{X|C}^{-1}\) is continuous on \((0,1)\), then \[\rho_{h|C}(X) = \int_0^1 F_{X|C}^{-1}(1-u)\, \,\mathrm{d}h(u) = \int_0^1 F_{X|C}^{-1+}(1-u)\, \,\mathrm{d}h(u).\]

Proof. Note \(\mathbb{P}_C\) is a probability measure on \((\Omega,\mathcal{F})\). The result follows by applying the unconditional quantile representation (Lemma 1 in [4]) under the probability measure \(\mathbb{P}_C\). In particular, replace \(\mathbb{P}\) by \(\mathbb{P}_C\), \(F_X\) by \(F_{X|C}\), and \(\rho_h\) by \(\rho_{h|C}.\) This gives (i), (ii), and (iii). ◻

3 Noise-Robust IRL Approach for Elicitation↩︎

3.1 Setup↩︎

In this section, we investigate the IRL problem with the goal of recovering the agent’s preferred distortion function. We consider a finite candidate set of distortion functions: \[\hat{\mathcal{H}}:=\{h_1(t),h_2(t),\cdots,h_L(t) \}\subset \mathcal{H},\] where \(t\in [0,1]\) and \(|\hat{\mathcal{H}}|=L\). The reason for the restriction to a finite set of distortion functions is that different distortion functions may induce the same ordering over the available actions and are therefore indistinguishable from observed choices. Many applications naturally describe risk preferences through a finite set of representative types. For example, driving styles in autonomous driving are often classified as conservative, moderate, or aggressive. In robo-advising, clients are commonly assigned to a finite set of risk profiles or portfolio types. Thus it is sufficient to recover a representative distortion function from a finite set that best matches the observed behavior.

We consider binary-choice problems with action space \(\mathcal{A}'=\{a_1,a_2\}\).

Let \(\widehat{Q}_l^{(n)}\) denote the posterior probability at round \(n\), for \(l=1,2,\cdots, L\). We initialize the prior as \(\widehat{Q}_l^{(0)}=1/L\) for all \(l\), and update the probability vector \(\{\widehat{Q}_l^{(n)}:l=1,\ldots,L\}\) after each observed response. The binary-choice question at round \(n\) is defined as: \(G^{(n)}:=\left\{(X_n, Y_n): X_n\overset{\mathrm{d}}{\ne} Y_n, X_n,Y_n\in L^\infty\right\}\). Let \(C(a^{(n)})\in \{X_n,Y_n\}\) denote the choice under \(a^{(n)}\in\mathcal{A}'\). Hence, its risk value under the distortion risk measure \(\rho_{h_l}\) is given by \[\label{eq:value-function-Terminal-any-action} \begin{align} \rho_{h_l}(C(a^{(n)}))= \int_{-\infty}^{0}\left(h_l\left(\mathbb{P}\left( C(a^{(n)})\geq x \right)\right)-h_l(1)\right)\,\mathrm{d}x +\int_0^\infty h_l \left(\mathbb{P}\left(C(a^{(n)}) \geq x \right)\right) \,\mathrm{d}x, \end{align}\tag{2}\] Binary-choice questions arise naturally in many applications. For example, in the context of portfolio investment, a typical question with two choices to identify the risk preference of an investor takes the form in Example 1.

Example 1. Here are two portfolios A and B (see Tables ¿tbl:tab:sales95q1?-[tbl:tab:sales95q2]) in the current financial market. The first row of each table shows the cumulative probability, and the second row shows the corresponding loss quantile. Negative losses represent gains. Which of the two portfolios would you choose in the current market environment?

Cumulative probability 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100%
Loss (units) -10 -5 -3.7 -2.5 -1 0 3 5 7 9 13
Portfolio A
Cumulative probability 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100%
Loss (units) -12 -11 -6 -1.4 -0.3 0.5 4 8 9 9.5 10

If an agent is perfectly rational, he or she will always choose an action that minimizes the risk value induced by his or her preferred distortion riskmetric: \[a^{(n)}= \underset{a\in \mathcal{A}'}{\arg\min} \rho_{h^*}(C(a)),\] where \(\rho_{h^*}\) is the agent’s preferred distortion riskmetric. However, in practice, observed choices may deviate from this deterministic optimality rule. Such deviations may arise from fatigue during repeated questioning, limited awareness of latent risk preferences, or near indifference when the two choices have similar risk values under \(\rho_{h^*}\). Therefore, agents’ responses should be modeled as stochastic rather than perfectly rational choices. This motivates the development of noise-robust IRL algorithms for recovering the preferred distortion riskmetric from uncertain choice data.

3.2 Distinguishing Method↩︎

We measure the distance between two distortion functions by \(L^\infty\) distance: \[||h_i-h_j||_\infty=\sup_{t\in[0,1]}|h_i(t)-h_j(t)|.\] The following proposition establishes the quantitative relationship between the \(L^\infty\) distance of distortion functions and the \(L^\infty\) distance of distortion riskmetrics.

For all \(h_i, h_j\in \mathcal{H}\) and \(X\in L^\infty\), we have: \[|\rho_{h_i}(X)-\rho_{h_j}(X)|\le (|U|+2|M|)||h_i-h_j||_\infty,\] where \(U\ge\max\{0, \mathop{\mathrm{ess\,sup}}X\}\) and \(M\le\min\{0, \mathop{\mathrm{ess\,inf}}X\}\).

Proof. Proof of Proposition [prop:bound-of-risk-measure] Without loss of generality, we assume that \(M<0\). We set \(\Delta h_{ij}=h_i-h_j\). According to 1 , we have that \[\begin{align} \rho_{\Delta h_{ij}}(X)&=\int_{-\infty}^{0}\left(\Delta h_{ij}(\mathbb{P}(X\geq x))-\Delta h_{ij}(1)\right)\,\mathrm{d}x+\int_0^\infty \Delta h_{ij}(\mathbb{P}(X\geq x))\,\mathrm{d}x\\ &=\int_{M}^{0}\left(\Delta h_{ij}(\mathbb{P}(X\geq x))-\Delta h_{ij}(1)\right)\,\mathrm{d}x+\int_0^{U} \Delta h_{ij}(\mathbb{P}(X\geq x))\,\mathrm{d}x. \end{align}\] Note that \(||h_i-h_j||_\infty=\sup_{t\in[0,1]}|h_i(t)-h_j(t)| =\sup_{t\in[0,1]}|\Delta h_{ij}(t)|.\) So \[\begin{align} \int_0^{U} \Delta h_{ij}(\mathbb{P}(X\geq x))\,\mathrm{d}x&\le \bigg|\int_0^{U} \Delta h_{ij}(\mathbb{P}(X\geq x))\,\mathrm{d}x \bigg|\\ &\le \int_0^{U}||h_i-h_j||_\infty\,\mathrm{d}x =|U| ||h_i-h_j||_\infty. \end{align}\] Since \[\begin{align} \int_{M}^{0}\left(\Delta h_{ij}(\mathbb{P}(X\geq x))-\Delta h_{ij}(1)\right)\,\mathrm{d}x&\le \int_{M}^{0}\Big|\Delta h_{ij}(\mathbb{P}(X\geq x))-\Delta h_{ij}(1)\Big|\,\mathrm{d}x\\ &\le \int_{M}^{0}\left(\Big|\Delta h_{ij}(\mathbb{P}(X\geq x))\Big|+\Big|\Delta h_{ij}(1)\Big|\right)\,\mathrm{d}x\\ &\le \int_{M}^{0} 2||h_i-h_j||_\infty\,\mathrm{d}x =2|M|||h_i-h_j||_\infty, \end{align}\] we have \(|\rho_{h_i}(X)-\rho_{h_j}(X)|\le (|U|+2|M|)||h_i-h_j||_\infty\). ◻

Proposition [prop:existence-of-distinguishing-question] guarantees the existence of a distinguishing question for any pair of non-proportional distortion functions.

For all \(h_i\), \(h_j\in \mathcal{H}\), if \(h_i\ne c h_j\) for all \(c\ge 0\) and \(h_j\not\equiv 0\), there exist \(X,Y\in L^\infty\), such that \[\rho_{h_i}(X)<\rho_{h_i}(Y), \quad \rho_{h_j}(X)>\rho_{h_j}(Y).\]

Proof. Proof of Proposition [prop:existence-of-distinguishing-question] For any distortion function \(h\), we define the linear functional \(L_h\) on quantile functions \(q(u)\) by \[L_h(q)=\int_0^1 q(1-u)\,\,\mathrm{d}h(u).\] By the quantile representation, if a random variable \(Z\) has the quantile function \(q_Z\), then \(\rho_h(Z)=L_h(q_Z).\) We first show that there exists a bounded continuous function \(r\) such that \(L_{h_i}(r)<0\) and \(L_{h_j}(r)>0.\) If not, then every \(r\) satisfying \(L_{h_j}(r)>0\) must also satisfy \(L_{h_i}(r)\ge 0\). Since \(h_j\not\equiv 0\), the functional \(L_{h_j}\) is not identically zero. Hence there exists \(g\in C[0,1]\) such that \(L_{h_j}(g)>0.\) Now take any \(f\in \ker L_{h_j}\), so that \(L_{h_j}(f)=0\). We claim that \(L_{h_i}(f)=0\). If \(L_{h_i}(f)>0\), we set \(r_\lambda=g-\lambda f\). Then for sufficiently large \(\lambda>0\), we have \[L_{h_j}(r_\lambda)=L_{h_j}(g)-\lambda L_{h_j}(f)=L_{h_j}(g)>0, \quad L_{h_i}(r_\lambda)=L_{h_i}(g)-\lambda L_{h_i}(f)<0,\] which contradicts the assumption. Similarly, if \(L_{h_i}(f)<0\), then for sufficiently large \(\lambda>0\), by setting \(r_\lambda=g+\lambda f\), there is also a contradiction. Hence, we have \(\ker L_{h_j}\subseteq \ker L_{h_i},\) and there exists a constant \(c\ge 0\) such that \[L_{h_i}=cL_{h_j},\] which implies \(h_i=ch_j\). This contradicts the assumption that \(h_i\ne c h_j\) for all \(c\ge 0\). Therefore, there exists a bounded continuous function \(r\in C[0,1]\) such that \(L_{h_i}(r)<0\) and \(L_{h_j}(r)>0\). Since \(C^1[0,1]\) is dense in \(C[0,1]\) and the above inequalities are strict, we may choose such an \(r\) in \(C^1[0,1]\) such that \(L_{h_i}(r)<0\) and \(L_{h_j}(r)>0.\) Choose a constant \(M>\|r'\|_\infty\) and define \[q_Y(t)=Mt,\qquad q_X(t)=Mt+r(t), \qquad t\in[0,1].\] Then \(q_Y\) is strictly increasing. Moreover, we have \[q_X'(t)=M+r'(t)\ge M-\|r'\|_\infty>0,\] so \(q_X\) is also strictly increasing. Hence \(q_X\) and \(q_Y\) are valid bounded quantile functions. We define \(X=q_X(U)\) and \(Y=q_Y(U)\), where \(U\sim \mathcal{U}(0,1)\). Then \(X,Y\in L^\infty\), and their quantile functions are \(q_X\) and \(q_Y\). Using the quantile representation, we obtain \[\rho_{h_i}(X)-\rho_{h_i}(Y) = L_{h_i}(q_X)-L_{h_i}(q_Y) = L_{h_i}(q_X-q_Y) = L_{h_i}(r)<0.\] Therefore, we have \(\rho_{h_i}(X)<\rho_{h_i}(Y)\). Similarly, we can show that \(\rho_{h_j}(X)>\rho_{h_j}(Y)\). Hence, we complete the proof. ◻

Theorem 1 guarantees finite identifiability: for all \(h_i\in\hat{\mathcal{H}}\), a finite set of distinguishing questions is sufficient to recover \(h_i\) within the candidate class \(\hat{\mathcal{H}}\).

Theorem 1.

We define \[K_{\delta,i}=\left\{h\in \hat{\mathcal{H}}: \delta\le ||h-h_i||_{\infty}, h\ne ch_i\;\text{for every} \;c\ge 0\right\}\] for all \(\delta>0\) such that \(K_{\delta,i}\ne \emptyset\). For each distortion function \(h_{i}\in \hat{\mathcal{H}}\) and \(h_i\not\equiv 0\), there exists a constant \(M\in \mathbb{N}^*\), for all \(h\in K_{\delta,i}\), there exist \(k,l\in \{1,2,\cdots,M\}\) and \(k\ne l\), such that \[\rho_{h}(X_k)<\rho_{h}(X_l),\;\rho_{h_{i}}(X_k)>\rho_{h_{i}}(X_l).\]

Proof. Proof of Theorem 1 Since \(\hat{\mathcal{H}}\) is a finite set, \(K_{\delta,i}\) is finite and compact. Because \(h\in K_{\delta,i}\), we can construct \(X_h, Y_h\in L^\infty\), such that \(\rho_{h}(X_h)<\rho_{h}(Y_h)\) and \(\rho_{h_{i}}(X_h)>\rho_{h_{i}}(Y_h)\) by Proposition [prop:existence-of-distinguishing-question]. We fix \(X_h\) and \(Y_h\), and we set \(\Delta_h=\min \{|\rho_{h}(X_h)-\rho_{h}(Y_h)|,|\rho_{h_{i}}(X_h)-\rho_{h_{i}}(Y_h)| \}.\) By Proposition [prop:bound-of-risk-measure], there exists \(\eta_h>0\) such that \(||g-h||_\infty<\eta_h\), and \[|\rho_{g}(X_h)-\rho_{h}(X_h)|<\frac{\Delta_h}{4}, \;|\rho_{g}(Y_h)-\rho_{h}(Y_h)|<\frac{\Delta_h}{4}.\] Note that \[\big|\rho_{g}(X_h)-\rho_{g}(Y_h)-(\rho_{h}(X_h)-\rho_{h}(Y_h))\big|\le |\rho_{g}(X_h)-\rho_{h}(X_h)|+|\rho_{g}(Y_h)-\rho_{h}(Y_h)| <\frac{\Delta_h}{2}.\] We have \[\rho_{g}(X_h)-\rho_{g}(Y_h)<\rho_{h}(X_h)-\rho_{h}(Y_h)+\frac{\Delta_h}{2} \le -\Delta_h+\frac{\Delta_h}{2} <0.\] We construct an open set \(U_h=\{g\in K_{\delta,i}:||g-h||_\infty<\eta_h\}.\) Hence, for all \(g\in U_h\), there exists \(X_h, Y_h\in L^\infty\), such that \(\rho_{g}(X_h)<\rho_{g}(Y_h)\) and \(\rho_{h_{i}}(X_h)>\rho_{h_{i}}(Y_h)\). Since \(\{U_h\}_{h\in K_{\delta,i}}\) is an open cover of \(K_{\delta,i}\) and \(K_{\delta,i}\) is compact, there exist \(h_1\), \(h_2\), \(\cdots\), \(h_N\in K_{\delta,i}\), such that \[\label{eq:union-question} K_{\delta,i} \subset \bigcup_{j=1}^{N} U_{h_j}.\tag{3}\] We set \(M=2N\) and construct \[\{X_1,X_2,\cdots,X_M\}=\{X_{h_j},Y_{h_j}:\rho_{h_j}(X_{h_j})<\rho_{h_j}(Y_{h_j}),\;\rho_{h_{i}}(X_{h_j})>\rho_{h_{i}}(Y_{h_j}),\;j=1,2,\cdots,N\}.\] For all \(h\in K_{\delta,i}\), we can find a set \(U_{h_j}\), \(j\in \{1,2,\cdots,N\}\), such that \(h\in U_{h_j}\). And we can find \(X_k\), \(X_l\in \{X_1,X_2,\cdots,X_M\}\) where \(k\ne l\), such that \(\rho_{h}(X_k)<\rho_{h}(X_l)\) and \(\rho_{h_{i}}(X_k)>\rho_{h_{i}}(X_l)\). ◻

According to Theorem 1, we can construct a finite set of questions \(\mathcal{K}\) to identify the agent’s preferred distortion function: \[\mathcal{K}:=\left\{(X_i,Y_i): X_i\overset{\mathrm{d}}{\ne}Y_i, X_i,Y_i\in L^\infty, i=1,2,\cdots,N\right\}.\] We define the regret of the action \(a^{(n)}\) for each distortion function \(h_l\) under the binary-choice question \(G^{(n)}\in \mathcal{K}\) at round \(n\) as \[\label{eq:regret-of-action} \Phi(a^{(n)};G^{(n)},h_l)=\rho_{h_l}(C(a^{(n)}))-\min_{a\in\mathcal{A}'}\rho_{h_l}(C(a)).\tag{4}\] The regret \(\Phi\) quantifies how well the distortion function \(h_l\) aligns with the agent’s observed choices. To quantify the discriminative power of the question, we define the distinguishing power of \(G^{(n)}\) as: \[\begin{align}\label{eq:distinguish-power-of-the-environment} \Psi(G^{(n)},i,j)=\Phi(a^{(n), *,i};G^{(n)},h_j)\Phi(a^{(n), *,j};G^{(n)},h_i), \end{align}\tag{5}\] where \(a^{(n), *,i}\) and \(a^{(n), *,j}\) represent the optimal actions for the distortion functions \(h_i(t)\) and \(h_j(t)\) respectively. It is straightforward to verify that \[\Psi(G^{(n)},i,j)=\left(\rho_{h_i}(C(a^{(n), *,j}))-\rho_{h_i}(C(a^{(n), *,i}))\right)\left(\rho_{h_j}(C(a^{(n), *,i}))-\rho_{h_j}(C(a^{(n), *,j}))\right).\] Obviously, \(\Psi(G^{(n)},i,j)>0\) if and only if \(a^{(n), *,i}\ne a^{(n), *,j}\). If \(\Psi(G^{(n)}, i, j) = 0\) for some pair \(i \neq j\), then \(G^{(n)}\) cannot separate the distortion functions \(h_i\) and \(h_j\). To ensure that all distortion functions in \(\hat{\mathcal{H}}\) can be distinguished, we impose the following condition on every pair \((X,Y)\in \mathcal{K}\).

Assumption 1. For all \(h_i\in \hat{\mathcal{H}}\) and for all \((X,Y)\in \mathcal{K}\), we have \(\rho_{h_i}(X) \ne \rho_{h_i}(Y)\).

To improve the efficiency of the algorithm, we adopt an adaptive questioning method with exploration rate \(\epsilon>0\). At the end of round \(n-1\), we select the question for round \(n\). With probability \(1-\epsilon\), we select the question that best distinguishes the two distortion functions \(h_{i^*}\) and \(h_{j^*}\) with the highest and second-highest posterior probabilities until round \(n-1\): \[\label{eq:environment-selection} G^{(n)}= \underset{G\in \mathcal{K}}{\arg \max} \Psi(G,i^*,j^*).\tag{6}\] With exploration rate \(\epsilon\), we instead randomly choose two distinct distortion functions \(h_i\) and \(h_j\) from \(\hat{\mathcal{H}}\) and select the question that best distinguishes them. To quantify how effectively the question \(G\) can discriminate between different distortion riskmetrics, we derive an upper bound on the distinguishing power \(\Psi(G,i,j)\) given two different distortion functions in Proposition [prop:upper-bound-of-distingushing-power].

Under Assumption 1, let \((X,Y)\in \mathcal{K}\) and define \(\ell=\max\{||X||_\infty, ||Y||_\infty\}\). For all \(h_i, h_j\in \hat{\mathcal{H}}\), if there exists a constant \(m\in \mathbb{N}^*\) and a partition: \(0=t_0<t_1<\cdots<t_m=1\), such that \(h_i(t)-h_j(t)\) is monotone on each interval \([t_{r-1},t_r]\) for \(r=1,\cdots, m\), we have \[\Psi(G,i,j)\le 4m^2\ell^2||h_i-h_j||_\infty^2.\]

Proof. Proof of Proposition [prop:upper-bound-of-distingushing-power] Let \(A=\rho_{h_i}(X) - \rho_{h_i}(Y)\), \(B=\rho_{h_j}(Y) - \rho_{h_j}(X)\), and \(g=h_i-h_j\). If \(AB<0\), then we have \(\Psi(G,i,j)=0\), and the proposition holds. If \(AB>0\), without loss of generality, we let \(A>0\) and \(B>0\). By [4], we have \[A+B =\rho_g(X)-\rho_g(Y) \le\mathop{\mathrm{ess\,sup}}|X-Y|\cdot \text{TV}_g[0,1] \le 2\ell \cdot \text{TV}_g[0,1],\] where \(\text{TV}_g[0,1]\) is the total variation of \(g\) on [0,1]. Then we have \[AB\le \frac{(A+B)^2}{4} \le\frac{4\ell^2 (\text{TV}_g[0,1])^2}{4} =\ell^2 (\text{TV}_g[0,1])^2.\] Since \(g\) is of bounded variation on [0,1] and monotone on each interval \([t_{r-1},t_r]\) for \(r=1,\cdots, m\), by additive property of total variation, \[\text{TV}_g[0,1]=\sum^m_{r=1}\text{TV}_g[t_{r-1},t_r] =\sum^m_{r=1}|g(t_r)-g(t_{r-1})| \le \sum^m_{r=1}2||g||_\infty =2m ||h_i-h_j||_\infty,\] where \(\text{TV}_g[t_{r-1},t_r]\) is the total variation of \(g\) on \([t_{r-1},t_r]\). We have \[\Psi(G,i,j)=(\rho_{h_i}(X) - \rho_{h_i}(Y))(\rho_{h_j}(Y) - \rho_{h_j}(X)) \le 4m^2\ell^2 ||h_i-h_j||_\infty^2.\] ◻

Proposition [prop:upper-bound-of-distingushing-power] suggests that the upper bound of the distinguishing power decays quadratically with respect to the distance between the distortion functions. Hence, candidates with similar induced risk preferences are intrinsically more difficult to separate and we therefore construct \(\hat{\mathcal{H}}\) as a set of representative and well-separated distortion functions.

3.3 Probabilistic Models for the Agent’s Decision-Making Process↩︎

We define \(h_{\hat{l}}\) as the distortion function that best approximates the agent’s risk preference. Since the agent may not always choose the option that is optimal under his or her preference, we propose the following decision-making models of an agent.

3.3.1 The Specific Decision-Making Model↩︎

We introduce a parameter \(\beta > 0\) reflecting the degree of rationality in the agent’s decision-making process, which is assumed to be unknown. We model the probability of choosing action \(a_1\) under question \(G^{(n)}\) and the preferred distortion function \(h_{\hat{l}}\) as: \[\label{eq:probability-agent-a1} \widehat{P}(a_1|G^{(n)},h_{\hat{l}})=\frac{\exp{(-\beta\Phi(a_1;G^{(n)},h_{\hat{l}}))}}{\exp{(-\beta\Phi(a_1;G^{(n)},h_{\hat{l}}))}+\exp{(-\beta\Phi(a_2;G^{(n)},h_{\hat{l}}))}}.\tag{7}\] The probability of choosing the action \(a_2\) can be computed as: \[\label{eq:probability-agent-a2} \widehat{P}(a_2|G^{(n)},h_{\hat{l}})=1-\widehat{P}(a_1|G^{(n)},h_{\hat{l}}).\tag{8}\]

The above model is reasonable for the following reasons. First, the model assigns a probability greater than \(50\%\) to the choice preferred under the agent’s distortion riskmetric. Second, the model captures the behavioral pattern that a larger risk-value gap between the two choices leads to a higher probability of selecting the option with lower risk under the agent’s preferred distortion riskmetric. Moreover, the inclusion of the parameter \(\beta\) enables the model to flexibly reflect real-world behaviors. If \(\beta\) is large, the agent behaves with more confidence and is more likely to choose the optimal choice. If \(\beta\) is small, the agent’s choice becomes more uncertain, increasing the likelihood of selecting a suboptimal choice. If \(\beta=0\), the agent exhibits no risk preferences and is indifferent between the two choices. When \(\beta\rightarrow{\infty}\), the agent consistently selects the optimal choice.

The value of \(\beta\) is influenced by various factors including the user’s fatigue and the extent to which they have insight into their own risk preferences.

3.3.2 A General Decision-Making Model↩︎

As for the general model, we assume that there exists a constant \(\bar{P}\in (\frac{1}{2},1]\), such that the agent selects the optimal choice \(a^*\in \mathcal{A}'\) given his or her preferred distortion riskmetric at round \(n\) with a stochastic probability \[\label{eq:fixed-probability-agent-a1} \widehat{P}(a^*|G^{(n)},h_{\hat{l}})=\widetilde{P}_n,\tag{9}\] where \(\widetilde{P}_n\in [\bar{P},1]\), while the other choice is selected with probability \[\label{eq:fixed-probability-agent-a2} 1-\widehat{P}(a^*|G^{(n)},h_{\hat{l}})=1-\widetilde{P}_n.\tag{10}\] The above model represents a more general case, since it does not impose a specific functional form. The model in Section 3.3.1 can be viewed as a special case of this general decision-making model.

3.4 Estimation Method↩︎

We establish a Bayesian-style framework to update the probability vector of the distortion functions. The Bayesian paradigm suggests that \[\text{Posterior Probability}\propto \text{Likelihood} \times \text{Prior Probability}.\] For any distortion function \(h_i\in\hat{\mathcal{H}}\), we introduce the normalized regrets of selecting different actions \(a_1\) and \(a_2\) under the question \(G^{(n)}\) as: \[\begin{align} &\widetilde{\Phi}(a_1;G^{(n)},h_i)=\frac{\Phi(a_1;G^{(n)},h_i)}{\max\{\Phi(a_1;G^{(n)},h_i), \Phi(a_2;G^{(n)},h_i)\}}, \\ &\widetilde{\Phi}(a_2;G^{(n)},h_i)=\frac{\Phi(a_2;G^{(n)},h_i)}{\max\{\Phi(a_1;G^{(n)},h_i), \Phi(a_2;G^{(n)},h_i)\}}. \end{align}\] If \(a_1\) is the optimal action, we have \(\widetilde{\Phi}(a_1;G^{(n)},h_i)=0\) and \(\widetilde{\Phi}(a_2;G^{(n)},h_i)=1\). Conversely, if \(a_2\) is optimal, then \(\widetilde{\Phi}(a_2;G^{(n)},h_i)=0\) and \(\widetilde{\Phi}(a_1;G^{(n)},h_i)=1\). The pseudo likelihood of selecting action \(a_1\) under \(G^{(n)}\) is computed as: \[\label{eq:likelihood-agent-a1} P(a_1|G^{(n)},h_i)=\frac{\exp{(-\kappa\widetilde{\Phi}(a_1;G^{(n)},h_i))}}{\exp{(-\kappa\widetilde{\Phi}(a_1;G^{(n)},h_i))}+\exp{(-\kappa\widetilde{\Phi}(a_2;G^{(n)},h_i))}},\tag{11}\] where \(\kappa>0\) is the learning rate. Meanwhile, the pseudo likelihood of action \(a_2\) is \[\label{eq:likelihood-agent-a2} P(a_2|G^{(n)},h_i)=1-P(a_1|G^{(n)},h_i).\tag{12}\] Based on the agent’s choice \(a^{(n)} \in \{a_1, a_2\}\) in round \(n\), we update the cumulative likelihood of each distortion function \(h_i\) in \(\hat{\mathcal{H}}\) by \[\label{eq:update-process} Q^{(n)}_i=P(a^{(n)}|G^{(n)},h_i)Q^{(n-1)}_{i}, \quad Q^{(0)}_i=\widehat{Q}_i^{(0)}, \quad n=1,\cdots, N.\tag{13}\] We normalize the cumulative likelihood across different distortion functions to obtain the set of posterior probabilities for the distortion functions: \(\{\widehat{Q}_i^{(n)}: i=1,2,\cdots,L\}\), where \[\widehat{Q}_i^{(n)}=\frac{Q_i^{(n)}}{\sum^L_{j=1}Q_j^{(n)}}.\] Once the update process has stabilized or the final round is reached, we select the distortion function with the highest posterior probability as the estimate of the agent’s preferred distortion function. The details for the estimation process are provided in Algorithm [alg:estimate-distortion-function].

Early stopping parameter \(q\), total number of rounds \(N\), number of questions \(S\), learning rate \(\kappa\), exploration rate \(\epsilon\), question pool \(\mathcal{K}\), initial cumulative likelihood set \(\{Q^{(0)}_i: i=1,2,\cdots,L \}\). Generate a random number \(u \sim \mathcal{U}(0,1)\). Select two distortion functions \(h_{i}\) and \(h_{j}\) with the largest and the second largest probability. Select two distortion functions \(h_{i}\) and \(h_{j}\) randomly. Choose the question \(G^{(n)}_{s^*}\) that best distinguishes \(h_i\) and \(h_j\) and collect the action \(a^{(n)}\) chosen by the agent. Obtain \(a'\in \mathcal{A}'\setminus \{a^{(n)}\}\). Compute \[P(a^{(n)}|G^{(n)},h_l)=\frac{\exp{(-\kappa\widetilde{\Phi}(a^{(n)};G^{(n)}_{s^*},h_l))}}{\exp{(-\kappa\widetilde{\Phi}(a^{(n)};G^{(n)}_{s^*},h_l))}+\exp{(-\kappa\widetilde{\Phi}(a';G^{(n)}_{s^*},h_l))}}.\] Update the cumulative likelihood: \[Q^{(n)}_l=P(a^{(n)}|G^{(n)}_{s^*},h_l)Q^{(n-1)}_{l}.\] Obtain \(l^{*,(n)}=\arg \max_{l\in \{1,2,\cdots,L\}}{Q}_l^{(n)}\). Compute \(\mu=\frac{1}{q}\sum^{q-1}_{j=0}{Q}_{l^{*,(n-j)}}^{(n-j)}\) and \(\sigma=\sqrt{\frac{1}{q}\sum^{q-1}_{j=0}({Q}_{l^{*,(n-j)}}^{(n-j)}-\mu)^2}\). break Obtain the distortion function \(h_{l^*}\) where \(l^{*,(n)}=\arg \max_{l\in \{1,2,\cdots,L\}}{Q}_l^{(n)}\).

3.5 Convergence Rate Analysis↩︎

We establish the convergence rate of the algorithm under the case \(h_{\hat{l}}\in \hat{\mathcal{H}}\) and the general setting 9 and 10 . We introduce a quantity: \[Z_{i,n}(a^{(n)})=\log \frac{P(a^{(n)}|G^{(n)},h_i)}{P(a^{(n)}|G^{(n)},h_{\hat{l}})},\] where \(a^{(n)}\) is the action chosen by the agent at round \(n\). Here \(Z_{i,n}(a^{(n)})\) represents the advantage of the likelihood of the distortion function \(h_i\) over the likelihood of the agent’s preferred distortion function \(h_{\hat{l}}\), given that the agent selects action \(a^{(n)}\).

To simplify the notation, we set \(d_j(n)=\rho_{h_j}(X_n)-\rho_{h_j}(Y_n)\). We define \(\sigma(u)=\frac{1}{1+e^{-u}}\) and the set of distinguishing rounds of the distortion function pair \(\{h_i, h_{\hat{l}}\}\) by \(\mathcal{N}_i=\left\{t\ge 1:\;d_{\hat{l}}(t)d_i(t)<0\right\}\). We also define \(\mathcal{N}_i^e\) as the set of distinguishing rounds when distinguishing questions between \(h_i\) and \(h_{\hat{l}}\) are generated by the exploration step, and \(\mathcal{N}_i^e\subset \mathcal{N}_i\).

The following lemma shows the upper bound of \(Z_{i,n}\).

Lemma 2. Under Assumption 1, for all \(i\in \{1,2,\cdots, L\}\setminus \{\hat{l}\}\) and every round \(n\), we have \[|Z_{i,n}(a^{(n)})|\le \kappa.\]

Proof. Proof of Lemma 2 Without loss of generality, let \(a_1^{(n)}\) denote the action of choosing the option \(X_n\), and \(a_2^{(n)}\) the action of selecting option \(Y_n\). According to the likelihood 11 and 12 , if \(d_i(n)<0\) and \(a_1^{(n)}\) is the optimal choice, we have \[P(a_1^{(n)}|G^{(n)},h)=\sigma(\kappa), \;P(a_2^{(n)}|G^{(n)},h)=\sigma(-\kappa).\] Hence, for \(a^{(n)}\in\{a_1^{(n)},a_2^{(n)}\}\) and \(i\ne \hat{l}\), we have \[|Z_{i,n}(a^{(n)})|=\left|\log \frac{P(a^{(n)}|G^{(n)},h_i)}{P(a^{(n)}|G^{(n)},h_{\hat{l}})}\right|\le \log \frac{\sigma(\kappa)}{\sigma(-\kappa)}=\kappa.\] ◻

The following proposition demonstrates that distinguishing questions can increase the expected log-likelihood ratio of the distortion function preferred by the agent more than that of other distortion functions. By contrast, for non-distinguishing questions, a wrong risk preference will not gain any advantage over the true one from these questions.

Under Assumption 1, for all \(\kappa>0\) and \(i\in \{1,\cdots, L\}\setminus \{\hat{l}\}\), given the question \((X_n, Y_n)\), if \((\rho_{h_{\hat{l}}}(X_n)-\rho_{h_{\hat{l}}}(Y_n))(\rho_{h_{i}}(X_n)-\rho_{h_{i}}(Y_n))<0\), under the general decision-making model 9 and 10 , the expectation of \(Z_{i,n}(a^{(n)})\) satisfies \[\widehat E[Z_{i,n}(a^{(n)})| \mathcal{F}_{n-1}]\le -\gamma,\] where \(\gamma=(2\bar{P}-1)\kappa>0\). If \((\rho_{h_{\hat{l}}}(X_n)-\rho_{h_{\hat{l}}}(Y_n))(\rho_{h_{i}}(X_n)-\rho_{h_{i}}(Y_n))>0\), we have \[\widehat E[Z_{i,n}(a^{(n)})| \mathcal{F}_{n-1}]=0.\]

Proof. Proof of Proposition [prop:negative-upper-bound-of-the-ratio] Without loss of generality, let \(\rho_{h_{\hat{l}}}(X_n)<\rho_{h_{\hat{l}}}(Y_n)\) and \(\rho_{h_{i}}(X_n)>\rho_{h_{i}}(Y_n)\). Let \(a_1^{(n)}\) denote the action of choosing \(X_n\), and \(a_2^{(n)}\) denote the action of choosing \(Y_n\). To simplify the notation, we set \(p_n:=\widetilde{P}_n\in[\bar P,1]\) according to 9 . Under the general decision-making model, we have \(\widehat P(a^{(n)}_1\mid G^{(n)},h_{\hat{l}})=p_n\) and \(\widehat P(a^{(n)}_2\mid G^{(n)},h_{\hat{l}})=1-p_n.\) By the likelihood 11 and 12 , we have \[\begin{align} \widehat E[Z_{i,n}(a^{(n)})| \mathcal{F}_{n-1}] &= p_n\log\frac{1-\sigma(\kappa )}{\sigma(\kappa )} +(1-p_n)\log\frac{\sigma(\kappa )}{1-\sigma(\kappa )} \\ &=p_n\log\frac{\sigma(-\kappa )}{\sigma(\kappa )} +(1-p_n)\log\frac{\sigma(\kappa )}{\sigma(-\kappa )}\\ &=-p_n\kappa+(1-p_n)\kappa\\ &=-(2p_n-1)\kappa \le -(2\bar P-1)\kappa. \end{align}\] If \(\rho_{h_{\hat{l}}}(X_n)<\rho_{h_{\hat{l}}}(Y_n)\) and \(\rho_{h_{i}}(X_n)<\rho_{h_{i}}(Y_n)\), we have \[\widehat E[Z_{i,n}(a^{(n)})| \mathcal{F}_{n-1}]=p_n\log\frac{\sigma(\kappa )}{\sigma(\kappa )} +(1-p_n)\log\frac{1-\sigma(\kappa )}{1-\sigma(\kappa )}=0.\] ◻

Corollary 1. Under Assumption 1, we have \[\sum_{t=1}^n \widehat {E}[Z_{i,t}(a^{(t)})| \mathcal{F}_{t-1}] \le -(2\bar P-1)\kappa \pi_0n,\] where \(\pi_0=\frac{2\epsilon}{L(L-1)}\).

Proof. Proof of Corollary 1 We define \(\mathcal{G}_t=\mathcal{F}_{t-1}\vee \sigma(G^{(t)})\) as the \(\sigma-\)algebra containing all information available up to round \(t-1\), together with the question selected for round \(t\). Let \(V_{i,t}\) be the event that the selected question for round \(t\) is to distinguish the pair of distortion functions \(\{h_{\hat{l}},h_i\}\). According to Proposition [prop:negative-upper-bound-of-the-ratio], we have \[\widehat {E}[Z_{i,t}(a^{(t)})| \mathcal{G}_{t}]\le -(2\bar P-1)\kappa\mathbf{1}_{V_{i,t}}.\] Since \(\widehat{P}(V_{i,t}|\mathcal{F}_{t-1})\ge \frac{2\epsilon}{L(L-1)}\), we have \[\begin{align} \widehat {E}[Z_{i,t}(a^{(t)})| \mathcal{F}_{t-1}]&=\widehat {E}[\widehat {E}[Z_{i,t}(a^{(t)})| \mathcal{G}_{t}]| \mathcal{F}_{t-1}]\\ &\le -(2\bar P-1)\kappa \widehat{P}(V_{i,t}|\mathcal{F}_{t-1})\le -(2\bar P-1)\kappa\cdot \frac{2\epsilon}{L(L-1)}. \end{align}\] Hence, we complete the proof. ◻

Theorem 2 establishes an upper bound on the convergence rate of our algorithm under the general setting 9 and 10 . The convergence rate is of order \(O\left(\exp\left(-cm+O\left(\sqrt{m\log m}\right)\right)\right)\) for some constant \(c>0\), where \(m\) is the number of questions.

Theorem 2. Under Assumption 1, if \(\kappa>0\), under the general model setting 9 and 10 , for all \(\delta\in(0,1)\) and all \(m\ge 1\), we have \[\widehat P\left(\left|\widehat{Q}_{\hat{l}}^{(m)}-1\right|\le r(m)\right)\ge 1-\delta,\] where \(r(m)=(L-1)\exp\left(-g_0\pi_0m+ 2\kappa\sqrt{2ma_{m}}\right)\), \(g_0=(2\bar P-1)\kappa\), \(\pi_0=\frac{2\epsilon}{L(L-1)}\), and \(a_{m}=\log\left(\frac{(L-1)\pi^2m^2}{6\delta}\right)\).

Proof. Proof of Theorem 2

Let \(S_i(m):=\log \frac{Q_{\hat{l}}^{(m)}}{Q_i^{(m)}}\). We have \[\widehat{Q}_{\hat{l}}^{(m)}=\frac{Q_{\hat{l}}^{(m)}}{\sum_{j=1}^LQ_{j}^{(m)} } =\frac{1}{1+\sum_{i\ne \hat{l}}\frac{Q_{i}^{(m)}}{Q_{\hat{l}}^{(m)}}} =\frac{1}{1+\sum_{i\ne \hat{l}}e^{-S_i(m)}},\] and \(1-\widehat{Q}_{\hat{l}}^{(m)}=\frac{\sum_{i\ne \hat{l}}e^{-S_i(m)}}{1+\sum_{i\ne \hat{l}}e^{-S_i(m)}}\le \sum_{i\ne \hat{l}}e^{-S_i(m)}.\) We set \(\widehat{Z}_{i,t}(a^{(t)})=-{Z}_{i,t}(a^{(t)})\). According to the update process 13 , we have \(S_i(t)=S_i(t-1)+\widehat{Z}_{i,t}(a^{(t)})\). We define \(\widehat{\eta}_{i,t}:=\widehat{Z}_{i,t}(a^{(t)})-\widehat E[\widehat{Z}_{i,t}(a^{(t)})|\mathcal{F}_{t-1}]\) and have \(\widehat E[\widehat{\eta}_{i,t}|\mathcal{F}_{t-1}]=0\). We define \[\widehat{M}_{i,m}:=\sum_{t=1}^m \widehat{\eta}_{i,t}, \quad \widehat{M}_{i,0}:=0.\] Hence, \(\widehat{M}_{i,m}\) is a martingale with respect to \(\mathcal{F}_{m}\). According to Lemma 2, we have \(\left|\widehat E[\widehat Z_{i,t}(a^{(t)})|\mathcal{F}_{t-1}]\right|\le \kappa\). Hence, \[|\widehat \eta_{i,t}|\le \left|\widehat Z_{i,t}(a^{(t)})\right|+\left|\widehat E[\widehat Z_{i,t}(a^{(t)})|\mathcal{F}_{t-1}]\right|\le 2\kappa.\] Noting that \(S_i(0)=0\), we have \[S_i(m)=\sum_{t=1}^{m}\widehat{Z}_{i,t}(a^{(t)})=\sum_{t=1}^{m}\widehat{\eta}_{i,t}+\sum_{t=1}^{m}\widehat E[\widehat{Z}_{i,t}(a^{(t)})|\mathcal{F}_{t-1}] = -\sum_{t=1}^m\Delta_{i,t}+\widehat{M}_{i,m},\] where \(\Delta_{i,t}=-\widehat E[\widehat{Z}_{i,t}(a^{(t)})|\mathcal{F}_{t-1}]\). According to Corollary 1, we have \(-\sum_{t=1}^m\Delta_{i,t}\ge g_0\pi_0m>0,\) where \(g_0=(2\bar{P}-1)\kappa\) and \(\pi_0=\frac{2\epsilon}{L(L-1)}\). Since \(\widehat M_{i,m}\) is a martingale and \(|\widehat \eta_{i,t}|\le 2\kappa\), by Azuma-Hoeffding inequality, for any \(x>0\), \[\widehat P(\widehat{M}_{i,m}\le -x)\le \exp\left(-\frac{x^2}{2\sum_{t=1}^{m}4\kappa^2}\right)=\exp\left( -\frac{x^2}{8\kappa^2 m}\right).\] Let \(\delta_{m}=\frac{6\delta}{\pi^2(m)^2}\), \(a_{m}=\log\left(\frac{(L-1)\pi^2(m)^2}{6\delta}\right)\), \(x_{m}=2\kappa \sqrt{2ma_{m}}\). Then \[\widehat P(\widehat{M}_{i,m}\le -x_m)\le \exp(-a_{m})=\frac{\delta_{m}}{L-1}.\] By Boole’s inequality, \[\widehat P\left(\exists i\ne\hat{l}: \widehat{M}_{i,m}\le -x_{m}\right)\le \sum_{i\ne \hat{l}}\mathbb{P}\left(\widehat{M}_{i,m}\le -x_{m} \right)\le (L-1)*\frac{\delta_{m}}{L-1}=\delta_{m}.\] Since \(\sum^\infty_{m=1}\delta_{m}=\frac{6\delta}{\pi^2}\sum^\infty_{m=1}\frac{1}{m^2}=\frac{6\delta}{\pi^2}\cdot \frac{\pi^2}{6}=\delta,\) by Boole’s inequality, \[\widehat P\left(\exists m\ge 1, \exists i\ne\hat{l}: \widehat{M}_{i,m}\le -x_{m} \right)\le \sum_{m=1}^\infty \widehat P\left(\exists i\ne\hat{l}: \widehat{M}_{i,m}\le -x_{m}\right)=\delta.\] Hence, \(\widehat P\left(\forall m\ge 1, \forall i\ne\hat{l}: \widehat{M}_{i,m}\ge -x_{m} \right)\ge 1-\delta.\) In the event \(\Theta :=\{\forall m\ge 1, \forall i\ne\hat{l}: \widehat{M}_{i,m}\ge -x_{m}\}\), we have \(\sum_{t=1}^{m}\widehat{Z}_{i,t}(a^{(t)})\ge g_0\pi_0m-x_{m}.\) Since \(S_i(m)=\sum^{m}_{t=1}\widehat{Z}_{i,t}(a^{(t)}),\) in the event \(\Theta\), we have \[\begin{align} S_i(m)&\ge g_0\pi_0m-2\kappa \sqrt{2ma_{m}},\\ \sum_{i\ne \hat{l}}e^{-S_i(m)}&\le \sum_{i\ne \hat{l}}\exp\left(-g_0\pi_0m+2\kappa \sqrt{2ma_{m}}\right),\\ \sum_{i\ne \hat{l}}e^{-S_i(m)}&\le (L-1)\exp\left(-g_0\pi_0m+2\kappa \sqrt{2ma_{m}}\right). \end{align}\] Since \(1-\widehat{Q}_{\hat{l}}^{(m)}\le \sum_{i\ne \hat{l}}e^{-S_i(m)}\), we have \[\widehat P\left(\left|\widehat{Q}_{\hat{l}}^{(m)}-1\right|\le (L-1)\exp\left(-g_0\pi_0m+2\kappa \sqrt{2ma_{m}}\right)\right)\ge 1-\delta.\] ◻

Theorem 2 shows that regardless of the specific pattern the agent follows during the selection process, as long as the agent does not choose options completely at random, we can identify the distortion riskmetric that reflects the agent’s underlying preferences by setting an appropriate learning rate \(\kappa\).

Although any learning rate \(\kappa>0\) is sufficient for the convergence analysis, it should be chosen carefully in practice when the agent’s level of uncertainty is unknown: an overly large value of \(\kappa\) may lead to overconfident updates of the posterior probabilities, whereas an excessively small value may slow the update of the posterior probabilities. An appropriate learning rate should align the likelihood 11 and 12 with the agent’s true probability. Since the general decision-making model assumes \(\widetilde{P}_n\in[\bar P,1]\), we choose \(\kappa\) conservatively by minimizing the worst-case expected negative log-likelihood: \[\min_{s\in(1/2,1)} \sup_{p\in[\bar P,1]} \left[-p\log s-(1-p)\log(1-s)\right],\] where \(s=\sigma(\kappa)\). Since \(f(s,p):=-p\log s-(1-p)\log(1-s)\) is decreasing in \(p\), we have \[\min_{s\in(1/2,1)} \sup_{p\in[\bar P,1]} \left[-p\log s-(1-p)\log(1-s)\right]\iff \min_{s\in(1/2,1)} \left[-\bar P\log s-(1-\bar P)\log(1-s)\right].\] If \(\bar P<1\), the minimizer of \(f(s,\bar P)\) is that \(s=\bar P\), and we have \(\kappa=\log\frac{\bar P}{1-\bar P}\). If \(\bar P=1\), then \(\kappa\) can be chosen arbitrarily large. In practice, we can set \(\kappa=\log\frac{0.55}{0.45}\approx 0.2\), which corresponds to assuming that the minimum probability with which agents choose their preferred choices is 0.55.

4 Elicit-to-Optimize Reinforcement Learning↩︎

4.1 Formulation↩︎

In this section, we solve the risk-sensitive reinforcement learning problems after identifying the agent’s preferred distortion function. Let \(T \in \mathbb{N}\) be a finite time horizon, and \(\mathcal{L} := \{0, 1, 2, \ldots, T\}\). Let \(\mathcal{S}\) and \(\mathcal{A}\) denote the state and the action spaces, respectively.

The gain of the agent at time \(t+1\) satisfies the following dynamics: \[W_{t+1}=\sum_{\tau=0}^tR(S_\tau, A_\tau, S_{\tau+1}; W_0), \quad 0<t\le T-1,\] where \(S_\tau\in \mathcal{S}\), \(A_\tau\in \mathcal{A}\), \(W_0\) denotes the initial gain, and \(R:\mathcal{S}\times \mathcal{A}\times \mathcal{S}\times \mathbb{R}\rightarrow{\mathbb{R}}\) is the deterministic measurable function.

We are interested in the RL problem where a conditional distortion riskmetric is applied to the accumulated costs. Throughout the RL section, we use the conditional distortion riskmetric in Definition 2. Since the condition variable \(Y\) is discrete, for a fixed scenario \(y_0\) with \(\mathbb{P}(Y=y_0)>0\), we take \(C=\{Y=y_0\}\). For simplicity, we write \(\rho_{h|y_0}(X):=\rho_{h|C}(X)\). We use \(\mathbf{s}_t\) to denote the collection of variables at time \(t\), which includes the state variable \(S_t\), time-to-maturity \(\tau_t=T-t\), condition variable \(y_t\), and the gain \(W_t^\pi\): \(\mathbf{s}_t=(S_t, \tau_t, y_t, W_t^\pi)\). We denote by \(\pi(a|\mathbf{s}_t)\) a stochastic feedback policy, which is a probability density function on the action space \(\mathcal{A}\) for any fixed \(\mathbf{s}_t\in \mathcal{S}\times \mathcal{L}\times \mathcal{X}\times \mathbb{R}\). Specifically, we aim to solve the following optimization problem: \[\begin{align}\label{prob:SRP} & \inf_{\pi \in \Pi} \int \left(W_0-W_T^\pi \right) \,\mathrm{d}h \circ \mathbb{P}_{Y = y_0} \\ =&\inf_{\pi\in\Pi} \int_{-\infty}^{0}\left(h\left(\mathbb{P}\left(W_0-W_T^\pi\geq x \bigg|Y= y_0\right)\right)-h(1)\right)\,\mathrm{d}x +\int_0^\infty h \left(\mathbb{P}\left(W_0-W_T^\pi \geq x \bigg| Y = y_0 \right)\right) \,\mathrm{d}x \\ =&\inf_{\pi\in\Pi} \rho_{h|y_0}\left(W_0-W_T^\pi \right), \end{align}\tag{14}\] where \(\Pi\) is the set of admissible feedback control distributions and \(W_t^\pi\) is the gain at time \(t\) under policy \(\pi\). Across different fields, \(W_t^\pi\) and the condition variable \(Y\) may take a wide range of meanings. For example, in investment, \(W_t^\pi\) may represent the total wealth of an investor and \(Y\) may represent macroeconomic conditions or market regimes. In operations management, \(W_t^\pi\) may describe order revenues and \(Y\) may describe order cancellation rate or port congestion.

According to Lemma 1, the objective 14 can be represented by \[\sup_{\pi \in \Pi} -\int^1_0F^{-1+,\pi}_{X|Y= y}(1-z)\,\mathrm{d}h(z),\] where \(F^{-1+,\pi}_{X|Y=y}(1-z)=\inf\left\{x\in\mathbb{R}:\mathbb{P}\left(W_0-W_T^\pi\le x|Y=y_0\right)>1- z\right\},\) or \[\sup_{\pi \in \Pi}-\int^1_0F^{-1,\pi}_{X|Y=y}(1-z)\,\mathrm{d}h(z),\] where \(F^{-1,\pi}_{X|Y=y}(1-z)=\inf\left\{x\in\mathbb{R}: \mathbb{P}\left(W_0-W_T^\pi\le x|Y=y_0\right)\ge 1-z\right\}.\) Hence, the reward received by the agent at time \(T\) can be defined as \(R_T=-\int^1_0F^{-1+,\pi}_{X|Y=y}(1-z)\,\mathrm{d}h(z)\) or \(R_T=-\int^1_0F^{-1,\pi}_{X|Y=y}(1-z)\,\mathrm{d}h(z)\). If \(t+1<T\), the reward can be defined as \[\label{eq:reward} R_{t+1}=-\left|W_{t+1}^\pi-W_0\right|\mathbf{1}_{\left\{{W_{t+1}^\pi-W_0<0}\right\}}.\tag{15}\] The formulation of the reward in Equation 15 indicates that only positive costs are penalized, which is used as a loss-oriented shaping term to guide the policy toward optimizing the final conditional distortion-riskmetric objective. Starting from the initial state \(\mathbf{s}_0\), the agent takes actions following a policy \(\pi\) until maturity, which results in a trajectory \(\mathcal{T}=(\mathbf{s}_0,a_0,R_1, \mathbf{s}_1,a_1,R_2,\cdots, \mathbf{s}_{T-1}, a_{T-1},R_T, \mathbf{s}_{T}).\) Under the policy \(\pi\), the value function at the state \(\mathbf{s}_t\) is \[V^\pi(\mathbf{s}_t)=\mathbb{E}^\pi\left[\sum_{k=0}^{T-t-1}\gamma^k R_{t+1+k}\Bigg| \mathbf{s}_t \right],\] where \(\gamma\) is the discount rate.

4.2 Policy Network and Value Network↩︎

Both the value function and the policy are represented by fully connected, multilayer perceptrons (MLP), with their parameters denoted by \(\psi\in \Psi\) and \(\theta\in \Theta\), where \(\Psi\) and \(\Theta\) are two compact domains. The two networks are denoted by \(V^{\pi,\psi}(\mathbf{s}_t)\) and \(\pi^{\theta}(\mathbf{s}_t)\), respectively. The policy network \(\pi^\theta\) parameterizes a Gaussian distribution over the continuous action space: \[\mu_\theta(\mathbf{s}_t) = \mathrm{MLP}_\theta(\mathbf{s}_t)\in\mathbb{R}, \; \sigma_\theta = \exp(\hat{\sigma}_\theta), \; \pi^\theta(a\mid\mathbf{s}_t) = \mathcal{N}\!\left(a;\,\mu_\theta(\mathbf{s}_t),\,\sigma_\theta^2\right),\] where \(\hat{\sigma}_\theta\) is a learnable scalar parameter. The value network \(V^{\pi,\psi}\) estimates the expected discounted return of the policy at the current state: \[V^{\pi,\psi}(\mathbf{s}_t) = \mathrm{MLP}_\psi(\mathbf{s}_t)\in\mathbb{R}. \label{eq:value95net}\tag{16}\] The 1-step TD residual of the value function is defined as \[\delta_t^{\pi^\theta}= R_{t+1} + \gamma\,V^{\pi,\psi}(\mathbf{s}_{t+1}) - V^{\pi,\psi}(\mathbf{s}_t),\] where \(\gamma=1\) is the discount factor. Then the advantage function and target returns are computed using Generalized Advantage Estimation (GAE): \[\label{eq:GAE} \hat{A}_t^{\pi^\theta} = \sum_{l=0}^{T-t-1}(\gamma\lambda_{\mathrm{GAE}})^l\,\delta_{t+l}^{\pi^\theta}, \; \hat{G}_t = \hat{A}_t^{\pi^\theta} + V^{\pi,\psi}(\mathbf{s}_t),\tag{17}\] where \(\lambda_{\mathrm{GAE}}\) is the GAE decay coefficient. Hence, the policy network updates are carried out using the clipped surrogate objective of the PPO algorithm: \[\mathcal{L}_{\mathrm{PPO}}(\theta) = - \min\!\left( r_t(\theta)\,\hat{A}_t^{\pi^{\theta_{\mathrm{old}}}},\; g\left(\epsilon,\hat{A}_t^{\pi^{\theta_{\mathrm{old}}}}\right) \right) , \label{eq:ppo}\tag{18}\] where the policy probability ratio \(r_t(\theta)\) and the function \(g(\epsilon, A)\) that prevents the policy from changing too aggressively in one update are defined as \[r_t(\theta) = \frac{\pi^\theta(a_t\mid\mathbf{s}_t)}{\pi^{\theta_{\mathrm{old}}}(a_t\mid\mathbf{s}_t)} = \exp\!\left( \log\pi^\theta(a_t\mid\mathbf{s}_t) - \log\pi^{\theta_{\mathrm{old}}}(a_t\mid\mathbf{s}_t) \right), \label{eq:ratio}\tag{19}\] and \[\begin{align} g(\epsilon, A)=\begin{cases} (1 + \epsilon)A, & \text{if } A \geq 0, \\ (1 - \epsilon)A, & \text{otherwise}. \end{cases} \end{align}\] with clipping threshold \(\epsilon=0.2\) and \(\theta_{\text{old}}\) representing the parameters of the policy network obtained from the previous update. The value network is trained using a clipped mean squared error loss: \[\mathcal{L}_V(\psi) = \max\!\left( \left(V^{\pi,\psi}(\mathbf{s}_t) - \hat{G}_t\right)^2,\; \left(V^{\pi,\psi}_{\mathrm{clip}}(\mathbf{s}_t) - \hat{G}_t\right)^2 \right) , \label{eq:value95loss}\tag{20}\] where the clipped value estimate is given by \[V^{\pi,\psi}_{\mathrm{clip}}(\mathbf{s}_t) = \max\left(\min \left(V^{\pi,\psi}(\mathbf{s}_t), V^{\pi,\psi_{\mathrm{old}}}(\mathbf{s}_t) + \epsilon_V\right), V^{\pi,\psi_{\mathrm{old}}}(\mathbf{s}_t) - \epsilon_V\right),\] and \(\epsilon_V=0.2\) is the value clipping threshold and \(\psi_{\mathrm{old}}\) represents the parameters of the value network obtained from the previous update. To encourage sufficient exploration, an analytical entropy regularization term for the Gaussian policy is included: \[\mathcal{L}_{\mathrm{ent}}(\theta) = -\mathcal{H}[\pi_\theta(\cdot\mid\mathbf{s}_t)] = -\tfrac{1}{2}\bigl(1 + \log(2\pi) + 2\log\sigma_\theta\bigr). \label{eq:entropy95loss}\tag{21}\] Overall, we minimize the following optimization objective for the policy and value networks: \[\mathcal{L}_{\mathrm{total}}(\theta,\psi) = \mathbb{E}\left[\mathcal{L}_{\mathrm{PPO}}(\theta) + c_0\,\mathcal{L}_{\mathrm{ent}}(\theta) + c_1\,\mathcal{L}_V(\psi)\right]. \label{eq:total95loss}\tag{22}\]

4.3 Quantile Network↩︎

In order to calculate the objective 14 , we approximate the \(\alpha\)-quantile \(F^{-1+,\pi}_{X|Y= y}(\alpha)\) and \(F^{-1,\pi}_{X|Y= y}(\alpha)\) by a quantile network \(\omega^\phi\), which is also represented by a multilayer feedforward network. The quantile network takes the initial state \(\mathbf{s}_0=(S_0,\tau_0,Y_0,W_0^\pi)\) and a confidence level \(\alpha\) as inputs. According to [14], we can set the loss function of the quantile network \(\omega^\phi(\boldsymbol{s}_0,\alpha)\) as \[\mathcal{L}_{\mathrm{QR}}(\phi) = \begin{cases} \alpha \left| \omega^\phi(\boldsymbol{s}_0,\alpha) - \left(W_0-W_T^\pi \right) \right|, & W_0-W_T^\pi \geq \omega^\phi(\boldsymbol{s}_0,\alpha), \\ (1 - \alpha) \left| \omega^\phi(\boldsymbol{s}_0,\alpha) -\left(W_0-W_T^\pi \right) \right|, & W_0-W_T^\pi \leq \omega^\phi(\boldsymbol{s}_0,\alpha). \end{cases}\] To enforce the structural constraint that higher confidence levels correspond to larger loss quantiles, we introduce a monotonicity regularization term: \[\mathcal{L}_{\mathrm{mono}}(\phi) = l\bigl( \omega_\phi(\mathbf{s}_0,\alpha) - \omega_\phi(\mathbf{s}_0,\alpha+\delta) \bigr), \label{eq:mono95loss}\tag{23}\] where \(\delta=0.05\) is the finite-difference step size and \(l(x)=\max (x,0)\) is the ReLU term. For \(\alpha_2=\alpha+\delta>\alpha\), the ReLU term produces a positive penalty whenever the network output violates \(\omega^\phi(\cdot,\alpha_2)\geq\omega^\phi(\cdot,\alpha)\). Hence, the overall training objective for the quantile network is \[\mathcal{L}_{\mathrm{quantile}}(\phi) = \mathbb{E}\left[c_2\,\mathcal{L}_{\mathrm{QR}}(\phi) + c_3\,\mathcal{L}_{\mathrm{mono}}(\phi)\right]. \label{eq:var95total95loss}\tag{24}\]

With the quantile network, the objective can be approximated numerically using the midpoint Riemann sum. Suppose that the distortion function \(h\) has finitely many right jumps at \(\mathcal{J}_h=\{\xi_1,\ldots,\xi_J\}\). Partitioning \([0,1]\) into \(n\) equal subintervals \(\{[t_{i-1},t_i]\}_{i=1}^n\) with midpoints \(t_i^{\mathrm{mid}}=(t_{i-1}+t_i)/2\), the estimate for the objective is \[\hat{\mathcal{C}}_h = -\sum_{i=1}^n \omega^\phi\left(\mathbf{s}_0,\;1-t_i^{\mathrm{mid}}\right)\, \left(h^c(t_i) - h^c(t_{i-1}) \right)- \sum_{j=1}^{J} \Delta h(\xi_j) \omega^\phi \left(\mathbf{s}_0,1-\xi_j \right), \label{eq:midpoint}\tag{25}\] where \(h^{c}\) denotes the continuous component of \(h\). For a left-continuous distortion function, we have \(\Delta h(\xi_j)=h(\xi_j+)-h(\xi_j)\). For a right-continuous distortion function, we have \(\Delta h(\xi_j)=h(\xi_j)-h(\xi_j-)\).

The details of the RL algorithm are provided in Algorithm [alg:RL].

Initialize policy network \(\pi^\theta\), value network \(V^{\pi, \psi}\), and quantile network \(\omega^\phi(s_0,\alpha)\) with parameters \(\theta_0\), \(\psi_0\), \(\phi_0\), experience replay buffer \(\mathcal{R}\leftarrow\emptyset\) with capacity \(C_{\mathcal{R}}\), initial gain \(W_0\), distortion function \(h\), discount factor \(\gamma\), GAE decay coefficient \(\lambda_{\mathrm{GAE}}\) Initialize buffers \(\mathcal{D}_k\leftarrow\emptyset\) and \(\mathcal{H}_k\leftarrow\emptyset\). Randomly sample \(N\) trajectories from the training set. Draw \(N\) confidence levels \(\alpha^{(n)}\) with \(n=1,\ldots,N.\) Observe current states \(\{\mathbf{s}_t^{(n)}\}_{n=1}^{N}\) from the vectorized environment. Sample actions \(a_t^{(n)}\!\sim\!\pi^{\theta_k}(\cdot\mid \mathbf{s}_t^{(n)})\); record \(\log\pi^{\theta_k}(a_t^{(n)}\mid \mathbf{s}_t^{(n)})\) and \(V^{\pi, \psi}(\mathbf{s}_t^{(n)})\). Execute \(a_t^{(n)}\); receive next states \(\mathbf{s}_{t+1}^{(n)}\), intermediate rewards \(R_{t+1}^{(n)}\), and gain \(W_{t+1}^{(n)}\). Record initial states \(\mathbf{s}_0^{(n)}\) and normalized terminal gain \(\tilde{W}_T^{(n)}=W_T^{(n)}/W_0\) for \(n=1,\ldots,N\). Approximate the objective value by 25 : \[\hat{\mathcal{C}}_h=-\sum_{i=1}^n \omega^\phi\left(\mathbf{s}_0^{(n)},\;1-t_i^{\mathrm{mid}}\right)\, \left(h^c(t_i) - h^c(t_{i-1}) \right)- \sum_{j=1}^{J} \Delta h(\xi_j) \omega^\phi \left(\mathbf{s}_0^{(n)},1-\xi_j \right).\] Set the terminal reward: \(R_T^{(n)}\leftarrow \hat{\mathcal{C}}_h\). Compute GAE advantages and target returns: \[\begin{align} \hat{A}_{T-1}^{(n)} &= R_T^{(n)}-V^{\pi, \psi}(\mathbf{s}_{T-1}^{(n)}), \qquad G_{T-1}^{(n)} = R_T^{(n)};\\ \delta_t^{(n)} &= R_{t+1}^{(n)}+\gamma\,V^{\pi, \psi}(\mathbf{s}_{t+1}^{(n)})-V^{\pi, \psi}(\mathbf{s}_t^{(n)}),\\ \hat{A}_t^{(n)} &= \delta_t^{(n)}+\gamma\lambda_{\mathrm{GAE}}\,\hat{A}_{t+1}^{(n)}, \qquad G_t^{(n)} = \hat{A}_t^{(n)}+V^{\pi, \psi}(\mathbf{s}_t^{(n)}). \end{align}\] Add tuples \(\bigl\{\bigl(\mathbf{s}_t^{(n)},\,a_t^{(n)},\,\pi^{\theta_k}(a_t^{(n)}\mid s_t^{(n)}),\, G_t^{(n)},\,\hat{A}_t^{(n)}\bigr)\bigr\}_{t=0}^{T-1}\) to \(\mathcal{D}_k\), for each \(n\). Append \(\bigl(\mathbf{s}_0^{(n)},\,\tilde{W}_T^{(n)},\,\alpha^{(n)}\bigr)\) to \(\mathcal{H}_k\) and to replay buffer \(\mathcal{R}\), for each \(n\). Set \(\theta\leftarrow\theta_k\), \(\psi\leftarrow\psi_k\), \(\phi\leftarrow\phi_k\). Shuffle \(\mathcal{D}_k\) and partition into mini-batches of size \(n_{\mathcal{B}}\). Normalize advantages \(\hat{A}_j\). Update policy parameter: \[\theta\leftarrow\theta -\eta_k\!\left[\frac{1}{n_{\mathcal{B}}}\sum_{j=1}^{n_{\mathcal{B}}}\left( \nabla_{\theta}\mathcal{L}_{\mathrm{PPO}}(\theta) +c_0\nabla_{\theta}\mathcal{L}_{\mathrm{ent}}(\theta)\right)\right] ;\] Update value parameter: \[\psi\leftarrow\psi -c_1\eta_k\!\left[\frac{1}{n_{\mathcal{B}}}\sum_{j=1}^{n_{\mathcal{B}}} \nabla_{\psi}\mathcal{L}_V(\psi)\right] ;\] Sample mini-batch \(\mathcal{H}'=\bigl\{\bigl(\mathbf{s}_0^{(j)},1-\tilde{W}_T^{(j)},\alpha^{(j)}\bigr) \bigr\}_{j=1}^{n_\phi}\) uniformly from \(\mathcal{R}\). Update quantile network parameter: \[\phi\leftarrow\phi -\eta_k^{\phi}\,\nabla_{\phi}\!\left[\frac{1}{n^\phi} \sum_{j=1}^{n^\phi}\left( c_2L_{\mathrm{QR}}(\phi) +c_3 L_{\mathrm{mono}}(\phi)\right)\right] ;\] \(\theta_{k+1}\leftarrow\theta\); \(\psi_{k+1}\leftarrow\psi\); \(\phi_{k+1}\leftarrow\phi\).

5 Numerical Experiment↩︎

In our numerical experiments, we conduct the relevant studies in a financial investment setting. We evaluate the IRL and RL components of our framework separately in order to provide a detailed and transparent assessment of the effectiveness of each module.

5.1 Estimation of the Distortion Function↩︎

We set \[\begin{align} \hat{\mathcal{H}}=&\bigg\{\mathbf{1}_{\{t>1-\alpha_1\}}, \frac{t}{1-\alpha_1}\wedge 1, \mathbf{1}_{\{\alpha_1\ge t\ge 1-\alpha_1\}}, \frac{t}{1-\alpha_1} \wedge 1 + \frac{\alpha_1 - t}{1-\alpha_1} \wedge 0, \mathbf{1}_{\{0<t<1\}}, \\&\quad \omega_1 \left( \frac{t}{1-\alpha_2} \wedge 1 \right) + \omega_2 \left( \frac{t}{1-\beta_1} \wedge 1 \right) + \omega_3 \mathbf{1}_{\{t > 1 - \alpha_2\}}, \left( \frac{\max(t-1+\beta_2, 0)}{\beta_2 - \alpha_3} \right) \wedge 1\bigg\}, \end{align}\] where \(\alpha_1\in \{0.7, 0.8, 0.9\}\), \(\alpha_2=0.7\), \(\beta_1=0.9\), \((\alpha_3,\beta_2)\in \{(0.7, 0.9), (0.7, 0.8), (0.8, 0.9)\}\), and \((\omega_1, \omega_2, \omega_3) = \bigl\{(0.4, 0.4, 0.2), (0.3, 0.5, 0.2), (0.5, 0.3, 0.2)\bigr\}\). Each distortion function in \(\hat{ \mathcal{H}}\) corresponds to a specific riskmetric: Value-at-Risk (\(\text{VaR}_\alpha\)), Expected Shortfall (\(\text{ES}_\alpha\), \(\text{CVaR}_\alpha\)), Inter-Quantile Range (\(\text{IQR}_\alpha\)), Inter-ES Range (\(\text{IER}_\alpha\)), Range, Range Value-at-Risk (\(\text{RVaR}_{\alpha,\beta}\)) and GlueVaR (see [4]). We select 100 stocks from the S\(\&\)P 500 constituents to form the basis of our choice set \(G\), utilizing their historical data to compute the risk thresholds at the 20 quantile points. We test our estimation algorithm under the two decision-making models: (1) The agent’s probabilities of choosing actions are given by 7 and 8 ; (2) The agent selects the choice consistent with his or her risk preference at a fixed probability \(\hat{p} > 0.5\) and the other choice is selected with probability \(1 - \hat{p}\). We set the learning rate \(\kappa=0.2\) and exploration rate \(\epsilon=0.01\). Figure 1 shows that our algorithm accurately identifies the agent’s risk preference under both adaptive questioning strategy (\(\epsilon=0.01\)) and random questioning strategy (\(\epsilon=1\)). When the adaptive questioning method is used, the algorithm converges faster and exhibits narrower uncertainty bands.

Figure 1: Posterior probabilities of the agent’s true risk preference (VaR_{0.9}) at each round of the IRL algorithm when adopting an adaptive questioning method with exploration (\epsilon=0.01) or selecting the two distortion functions to be distinguished completely at random (\epsilon=1). The probability that the agent chooses the option consistent with his or her risk preference is fixed at 0.9.

5.1.1 Specific Decision-Making Model with Different \(\beta\)↩︎

Figures 2-3 show the posterior probabilities of distortion functions in \(\hat{\mathcal{H}}\) at each round under different \(\beta\). We can observe that the posterior probability of the true risk preference always converges to one upon algorithm termination. When \(\beta=8\), the algorithm shows fast convergence and exhibits few fluctuations, because the agent behaves in a highly rational manner in the decision-making process. When \(\beta=0.7\), the probability of the true risk preference converges relatively slowly, and its convergence curve displays noticeable oscillations. This slow and volatile convergence is due to the high stochasticity in the agent’s decision-making process. Even under this high level of randomness, the algorithm still converges to the true risk preference within a relatively small number of rounds. Our algorithm not only identifies the agent’s risk preference effectively but also provides visual evidence of the agent’s degree of decision-making rationality.

Figure 2: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{VaR}_{0.8} that belongs to \hat{\mathcal{H}}. The posterior probability of \text{VaR}_{0.8} converges to 1.
Figure 3: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{RVaR}_{0.7,0.8} that belongs to \hat{\mathcal{H}}. The posterior probability of \text{RVaR}_{0.7,0.8} converges to 1.

We also investigate scenarios in which the model is misspecified: the agent’s true distortion function is not in the defined set \(\hat{\mathcal{H}}\). To illustrate this, we consider the following distortion functions that are not in \(\hat{\mathcal{H}}\): \[\left( \frac{\max(t-0.1, 0)}{0.3} \right) \wedge 1, \quad 0.3 \left( \frac{t}{0.2} \wedge 1 \right) + 0.5 \left( \frac{t}{0.1} \wedge 1 \right) + 0.2 \mathbf{1}_{\{t > 0.2\}},\] which are \(\text{RVaR}_{0.6,0.9}\) and \(0.3\text{ES}_{0.8}+0.5\text{ES}_{0.9}+0.2\text{VaR}_{0.8}\), respectively. Figures 4-5 present the evolution of the probabilities of the distortion functions in \(\hat{\mathcal{H}}\) when the true distortion function is not in the set \(\hat{\mathcal{H}}\). We can find that our algorithm will finally converge to a risk preference in \(\hat{\mathcal{H}}\) that is close to the true risk preference of the agent for both small and large values of \(\beta\) within a relatively small number of rounds. In the case of the true risk preferences \(\text{RVaR}_{0.6,0.9}\) and \(0.3\text{ES}_{0.8}+0.5\text{ES}_{0.9}+0.2\text{VaR}_{0.8}\), the algorithm converges to \(\text{RVaR}_{0.7,0.9}\) and \(0.3\text{ES}_{0.7}+0.5\text{ES}_{0.9}+0.2\text{VaR}_{0.7}\) respectively, which are close to the true risk preferences. Our algorithm can effectively address the challenges arising from both the model misspecification and the uncertainty of the agent in making decisions.

In all the above examples, we can preliminarily determine the agent’s preferred distortion function within 100 rounds, even when its structure is complex or the agent behaves randomly.

Figure 4: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{RVaR}_{0.6,0.9} that is not included in \hat{\mathcal{H}}. The posterior probability of \text{RVaR}_{0.7,0.9} converges to 1.
Figure 5: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is 0.3\text{ES}_{0.8}+0.5\text{ES}_{0.9}+0.2\text{VaR}_{0.8} that is not included in \hat{\mathcal{H}}. The posterior probability of 0.3\text{ES}_{0.7}+0.5\text{ES}_{0.9}+0.2\text{VaR}_{0.7} converges to 1.

5.1.2 General Decision-Making Model with Different \(\hat{p}\)↩︎

Under the general decision-making model 9 and 10 , we examine a setting where the agent, driven by his or her risk preference, makes an optimal choice with a fixed probability \(\hat{p}>0.5\) and a suboptimal choice with probability \(1-\hat{p}\) at each round. Figures 6-7 show the performances of our algorithm under this setting for \(\hat{p}=0.6\) and \(\hat{p}=0.9\) respectively. We find that our algorithm still performs well in this setting. When \(\hat{p}=0.6\), the agent’s choices are already close to being completely random, but our algorithm can still preliminarily identify the agent’s true risk preference after approximately 100 rounds. When \(\hat{p}=0.9\), the algorithm converges quickly with fewer oscillations. In the case of model misspecification, we consider two distortion functions that are not in the set \(\hat{\mathcal{H}}\): \(\frac{t}{0.05}\wedge 1\) (\(\text{ES}_{0.95}\)) and \(\mathbf{1}_{\{0.6\ge t\ge 0.4\}}\) (\(\text{IQR}_{0.6}\)). Figures 8-9 show that the algorithm converges to \(\text{ES}_{0.9}\) and \(\text{IQR}_{0.7}\) respectively, which are close to the true risk preferences that are not in the set \(\hat{\mathcal{H}}\). The results of this experiment demonstrate the robustness, flexibility, and stability of our algorithm under different settings.

Figure 6: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{ES}_{0.9} that belongs to \hat{\mathcal{H}}. The posterior probability of \text{ES}_{0.9} converges to 1.
Figure 7: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{IQR}_{0.9} that belongs to \hat{\mathcal{H}}. The posterior probability of \text{IQR}_{0.9} converges to 1.
Figure 8: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{ES}_{0.95} that is not included in \hat{\mathcal{H}}. The posterior probability of \text{ES}_{0.9} converges to 1.
Figure 9: Posterior probabilities of the distortion riskmetrics at each round. The agent’s true risk preference is \text{IQR}_{0.6} that is not included in \hat{\mathcal{H}}. The posterior probability of \text{IQR}_{0.7} converges to 1.

5.2 Decision-Making under Distortion Riskmetrics Optimization Objectives↩︎

After identifying the agent’s risk preference, we construct a portfolio that trades the S\(\&\)P 500 index and our objective is to solve problem 14 . We assume that all trading actions occur at market close and therefore use the S\(\&\)P 500 daily closing prices for our tests. We collect 21-year daily data of the S\(\&\)P 500 index from 2005/01/03 to 2026/03/27 from Wharton Research Data Services (WRDS). The dataset consists of 5342 daily observations in total. We use the first 4333 observations as the training set and the subsequent 1009 observations as the test set. Rolling-window tests are conducted on the test set, resulting in a total of 1000 test paths. The input variables of the policy network at time \(t\) include the index value \(S_t\), the time-to-maturity \(\tau_t\), the current wealth \(W_t\) and the regime \(Y_t\), while the output \(\delta_t\) is the quantity of shares to buy or sell. The wealth process follows the dynamic \[W_{t+1}=W_{t}+\delta_{t}(S_{t+1}-S_t).\] The initial wealth is \(W_0=300000\) and the largest time to maturity \(\tau_0\) is \(\frac{10}{252}\). The loss is defined as \(W_0-W_T\). We identify bull and bear market regimes from the historical data using a method similar to that of [25] and [26]. The market regime is classified as a bear market if the current index price falls by 19% from its most recent peak. Conversely, it is classified as a bull market if the current index price rises by 24% from its most recent trough. We choose the following distortion-riskmetric objectives: \(\mathrm{RVaR}_{0.6,0.9}\), \(\mathrm{RVaR}_{0.5,0.9}\), \(\mathrm{IQR}_{0.95}\), \(\mathrm{IQR}_{0.9}\), \(\mathrm{CVaR}_{0.9}\), \(\mathrm{CVaR}_{0.95}\), \(\mathrm{VaR}_{0.85}\), \(\mathrm{VaR}_{0.95}\), \(\mathrm{IER}_{0.9}\), and GlueVaR (\(0.6 \mathrm{CVaR}_{0.8}(X) + 0.25 \mathrm{CVaR}_{0.9}(X) + 0.15 \mathrm{VaR}_{0.8}(X)\)). In the algorithm implementation, the coefficients in 17 are \(\gamma=1\) and \(\lambda_{\text{GAE}}=0.95\); the coefficients in 22 are \(c_0=0\) and \(c_1=0.04\); the coefficients in 24 are \(c_2=0.08\) and \(c_3=0.05\). As for the neural networks, there are 3 hidden layers, with each hidden layer comprising 64 neurons.

Figures 10 and 11 compare the learned policies under different distortion-riskmetric objectives in bull and bear markets. Overall, for a fixed wealth level, the agent tends to sell at high prices and buy at low prices. As wealth increases, the agent can tolerate higher asset prices, so the buy–sell boundary generally slopes upward with diminishing marginal effects. The boundary often exhibits a hump-shaped pattern or an arch-shaped pattern: low-wealth agents buy at lower prices to recover capital, whereas high-wealth agents trade less to satisfy the risk-optimization objective. In the bear market, the buy region expands markedly relative to the bull market, and the buy–sell boundary shifts upward and rightward. This is plausible because depressed prices imply greater upside potential, while short positions are exposed to large losses from price rebounds. Since the policy patterns are broadly similar across regimes aside from the sizes of the buy and sell regions, we focus on differences among distortion-riskmetric objectives in the bull market.

  • Under \(\mathrm{CVaR}_{0.90}\), the buy region is mainly concentrated in low-price and low-to-intermediate-wealth states. Relative to \(\mathrm{CVaR}_{0.90}\), \(\mathrm{CVaR}_{0.95}\) yields a smaller buy region and a larger short region, reflecting greater conservatism. A similar pattern holds for \(\mathrm{VaR}_{0.85}\) and \(\mathrm{VaR}_{0.95}\), with the buy-sell boundary under \(\mathrm{VaR}_{0.95}\) shifting downward.

  • The GlueVaR policy, combining \(\mathrm{VaR}_{0.8}\), \(\mathrm{CVaR}_{0.8}\), and \(\mathrm{CVaR}_{0.9}\), is similar to the \(\mathrm{CVaR}_{0.9}\) policy but has a wider buy region. This is because GlueVaR also incorporates the relatively less conservative risk measures \(\mathrm{VaR}_{0.8}\) and \(\mathrm{CVaR}_{0.8}\).

  • \(\mathrm{RVaR}_{0.5,0.9}\) produces the broadest buy region, because it averages quantiles only over an intermediate range and is therefore more tolerant of downside risk while emphasizing wealth recovery. Compared with \(\mathrm{RVaR}_{0.5,0.9}\), \(\mathrm{RVaR}_{0.6,0.9}\) shifts the buy–sell boundary uniformly downward.

  • The transition from \(\mathrm{IQR}_{0.9}\) to \(\mathrm{IQR}_{0.95}\) is non-monotone. Under \(\mathrm{IQR}_{0.95}\), the boundary is slightly higher at very low wealth levels but clearly lower at medium and high wealth levels. This reflects the fact that \(\mathrm{IQR}_{\alpha}\) compresses the gap between favorable and unfavorable outcomes: loss compression dominates at very low wealth, making buying more attractive, whereas profit compression becomes more important as wealth increases, leading to more conservative policies.

  • The \(\mathrm{IER}_{0.90}\) policy has a geometry similar to that of \(\mathrm{CVaR}_{0.90}\), but it penalizes both upper- and lower-tail averages of the loss distribution. Consequently, when the wealth is high and the price is low, the policy tends to hold zero positions to avoid extreme profits; when the wealth is low and the price is low, it tends to buy more to reduce extreme losses.

Figure 10: Learned policy in the bull market (regime =1). Colors indicate normalized actions in [-1,1], with positive and negative values representing long and short positions, respectively.
Figure 11: Learned policy in the bear market (regime =2). Colors indicate normalized actions in [-1,1], with positive and negative values representing long and short positions, respectively.

6 Conclusion↩︎

In this paper, we develop an IRL-RL framework for eliciting agents’ risk preferences and optimizing decisions under the elicited preferences. Our framework applies to a broad class of distortion riskmetrics, including non-convex and non-coherent cases. We collect agents’ choices through iterative questioning and employ a Bayesian IRL framework to identify their risk preferences, which remains effective despite agents’ stochastic choices. We establish the existence of a finite set of distinguishing questions and prove the convergence rate of our IRL algorithm. In the decision-making stage, we extend the PPO algorithm to the conditional distortion riskmetric objective by representing the objective as an integral of the cost quantile function over \([0,1]\) with respect to the distortion function, and approximate it using a quantile neural network and the midpoint Riemann sum. Numerical experiments demonstrate the effectiveness of our IRL-RL framework.

References↩︎

[1]
H. Alsabah, A. Capponi, O. Ruiz Lacedelli, and M. Stern, “Robo-advising: Learning investors’ risk preferences via portfolio choices,” Journal of Financial Econometrics, vol. 19, no. 2, pp. 369–392, 2021.
[2]
A. Capponi, S. Olafsson, and T. Zariphopoulou, “Personalized robo-advising: Enhancing investment through client interaction,” Management Science, vol. 68, no. 4, pp. 2485–2512, 2022.
[3]
A. Majumdar, S. Singh, A. Mandlekar, and M. Pavone, “Risk-sensitive inverse reinforcement learning via coherent risk models.” in Robotics: Science and systems, 2017, vol. 16, p. 117.
[4]
Q. Wang, R. Wang, and Y. Wei, “Distortion riskmetrics on general spaces,” ASTIN Bulletin: The Journal of the IAA, vol. 50, no. 3, pp. 827–851, 2020.
[5]
P. W. Glynn, Y. Peng, M. C. Fu, and J.-Q. Hu, “Computing sensitivities for distortion risk measures,” INFORMS Journal on Computing, vol. 33, no. 4, pp. 1520–1532, 2021.
[6]
J. Dhaene, R. J. Laeven, and Y. Zhang, “Systemic risk: Conditional distortion risk measures,” Insurance: Mathematics and Economics, vol. 102, pp. 126–145, 2022.
[7]
S. Arora and P. Doshi, “A survey of inverse reinforcement learning: Challenges, methods and progress,” Artificial Intelligence, vol. 297, p. 103500, 2021.
[8]
M. Kuderer, S. Gulati, and W. Burgard, “Learning driving styles for autonomous vehicles from demonstration,” in 2015 IEEE international conference on robotics and automation (ICRA), 2015, pp. 2641–2646.
[9]
D. Sadigh, S. Sastry, S. A. Seshia, and A. D. Dragan, “Planning for autonomous cars that leverage effects on human actions.” in Robotics: Science and systems, 2016, vol. 2, pp. 1–9.
[10]
M. Zucker, J. A. Bagnell, C. G. Atkeson, and J. Kuffner, “An optimization approach to rough terrain locomotion,” in 2010 IEEE international conference on robotics and automation, 2010, pp. 3589–3595.
[11]
T. Park and S. Levine, “Inverse optimal control for humanoid locomotion,” in Robotics science and systems workshop on inverse optimal control and robotic learning from demonstration, 2013, pp. 4887–4892.
[12]
Z. Cheng, A. Coache, and S. Jaimungal, “Eliciting risk aversion with inverse reinforcement learning via interactive questioning,” arXiv preprint arXiv:2308.08427, 2023.
[13]
H. Buehler, L. Gonon, J. Teichmann, and B. Wood, “Deep hedging,” Quantitative Finance, vol. 19, no. 8, pp. 1271–1291, 2019.
[14]
X. Peng, X. Zhou, B. Xiao, and Y. Wu, “A risk sensitive contract-unified reinforcement learning approach for option hedging,” arXiv preprint arXiv:2411.09659, 2024.
[15]
A. Coache and S. Jaimungal, “Reinforcement learning with dynamic convex risk measures,” Mathematical Finance, vol. 34, no. 2, pp. 557–587, 2024.
[16]
S. Han, Y. Liu, and X. Yu, “Risk-sensitive reinforcement learning based on convex scoring functions,” arXiv preprint arXiv:2505.04553, 2025.
[17]
Y. Chow, M. Ghavamzadeh, L. Janson, and M. Pavone, “Risk-constrained reinforcement learning with percentile risk criteria,” Journal of Machine Learning Research, vol. 18, no. 167, pp. 1–51, 2018.
[18]
M. Du, H. Yu, and N. Kong, “Transfer reinforcement learning for mixed observability markov decision processes with time-varying interval-valued parameters and its application in pandemic control,” INFORMS Journal on Computing, vol. 37, no. 2, pp. 315–337, 2025.
[19]
K. Stachowicz and S. Levine, “Racer: Epistemic risk-sensitive rl enables fast driving with fewer crashes,” arXiv preprint arXiv:2405.04714, 2024.
[20]
J. Bernhard, S. Pollok, and A. Knoll, “Addressing inherent uncertainty: Risk-sensitive behavior generation for automated driving using distributional reinforcement learning,” in 2019 IEEE intelligent vehicles symposium (IV), 2019, pp. 2148–2155.
[21]
D. Kamran, C. F. Lopez, M. Lauer, and C. Stiller, “Risk-aware high-level decisions for automated driving at occluded intersections with reinforcement learning,” in 2020 IEEE intelligent vehicles symposium (IV), 2020, pp. 1205–1212.
[22]
S. Zhang, B. Liu, and S. Whiteson, “Mean-variance policy iteration for risk-averse reinforcement learning,” in Proceedings of the AAAI conference on artificial intelligence, 2021, vol. 35, pp. 10905–10913.
[23]
T. K. Büning, A.-M. George, and C. Dimitrakakis, “Interactive inverse reinforcement learning for cooperative games,” in International conference on machine learning, 2022, pp. 2393–2413.
[24]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
[25]
M. Dai, Q. Zhang, and Q. J. Zhu, “Trend following trading under a regime switching model,” SIAM Journal on Financial Mathematics, vol. 1, no. 1, pp. 780–810, 2010.
[26]
M. Dai, Z. Yang, Q. Zhang, and Q. J. Zhu, “Optimal trend following trading rules,” Mathematics of Operations Research, vol. 41, no. 2, pp. 626–642, 2016.