June 25, 2026
We study the squared price-of-risk premium of a portfolio—an integrated conditional squared Sharpe-ratio functional, not an expected excess return—and its attribution to causal drivers. Relative to a declared admissible benchmark it decomposes into intervention-stable premium, a signed causal distortion (the confounding wedge), and a nonnegative information loss; the loss is an \(L^2\) projection residual, the wedge is not. The decomposition is well posed exactly when the driver filtration is immersed in the price filtration. It need not aggregate across portfolios pooling drivers: we identify an order-three obstruction that is invisible to every singleton and pairwise admissibility screen—each one- and two-driver sub-book is immersed while the pooled triple reveals a future innovation—the analogue of Bernstein’s pairwise-but-not-mutually-independent triple, and minimal relative to such pairwise diagnostics. We separate its two ingredients, combinatorial masking and anticipative coupling. The failure is one of immersion, not of no-arbitrage. Experiments on synthetic single- and multi-driver panels show the decomposition and its causal correction are estimable, and that a permutation-calibrated screen detects planted order-three leakage with controlled false positives.
Keywords: price of risk , causal inference , filtration enlargement , immersion , portfolio attribution
MSC 2020: 91G10 , 60G44 , 62H22
The use of conditioning information in portfolio choice has a long history. Hansen and Richard [1] characterized the mean–variance frontier attainable by an investor who conditions a dynamic strategy on an information set, and Ferson and Siegel [2] derived the closed-form optimal weights and the associated efficiency tests; Abhyankar, Basu, and Stremme [3] showed that the increment in the maximal squared Sharpe ratio afforded by additional conditioning information equals the \(R^2\) of a predictive regression. In parallel, the theory of initial and progressive enlargement of filtrations [4] and the quantification of the information drift it induces [5] provide the probabilistic apparatus for comparing pricing under nested information sets. This paper joins the two literatures. We treat the conditioning information of a portfolio as the filtration generated by its causal drivers, decompose the attainable conditional squared Sharpe ratio into causal and non-causal components, and identify the precise sense in which the admissibility of the driver filtration—its immersion in the price filtration—governs whether that decomposition is well posed and whether it aggregates across a book of portfolios.
Let \((\Omega,\mathcal{F},(\mathcal{F}_t),\mathbb{P})\) carry the price filtration \(\mathcal{F}\). For a portfolio with excess-return dynamics \(dr_t=\mu_t\,dt+\sigma_t\,dW_t\), the price of risk is \(\lambda_t=\sigma_t^{-1}\mu_t\), the integrand of the Girsanov change to an equivalent martingale measure. Conditioning on a driver filtration \(\mathcal{G}\) and integrating the squared projection defines \(\pi(\mathcal{G})=\mathbb{E}\int_0^T\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\,dt\). By [1], \(\pi(\mathcal{G})\) is the maximal conditional squared Sharpe ratio attainable from \(\mathcal{G}\), and we establish (Proposition 2) that it is the value of a conditional mean–variance problem; \(\pi\) is therefore not a descriptive attribution statistic but the value realized by an admissible \(\mathcal{G}\)-measurable trading strategy. The object lives in the pricing regime: the price of risk mediates the change of measure, so \(\pi\) is a Sharpe-type functional rather than an expected excess return. We work in \(L^2(\mathbb{P}\otimes dt)\) because its inner-product structure supplies the orthogonal projection on which the decomposition rests. The distinction between the existence of a deflator and of an equivalent local martingale measure, invoked when we separate immersion failure from arbitrage, follows [6], [7].
The paper makes four contributions. First (Section 3), it gives a reference-relative decomposition of the conditional squared price-of-risk premium \(\mathrm{SPR}\) into intervention-stable premium, a signed causal distortion (the confounding wedge), and a nonnegative admissible information loss, the loss measured relative to a declared admissible benchmark. Second, it proves the causal distortion is not a projection residual—it is signed and non-Pythagorean (Proposition 7)—thereby separating causal bias from informational incompleteness, and it isolates the wedge identity (Theorem 4, requiring only the identification assumption) from the generic-nonzero converse (Theorem 5, requiring faithfulness). Third (Sections 4–5), it identifies a minimal order-three obstruction to admissible aggregation that is invisible to every singleton and pairwise screen: each one- and two-driver sub-book is immersed while the pooled book reveals a future innovation, the filtration-theoretic analogue of Bernstein’s pairwise-but-not-mutually-independent triple, minimal relative to pairwise diagnostics (Definition 2, Theorem 9). Fourth, it separates masking from anticipation (Remark 10): adapted crowding reproduces the algebra of the obstruction without breaking immersion, so the obstruction requires both combinatorial masking and anticipative coupling. We further link projection to realized premium, distinguishing admissible realized premium from inadmissible diagnostic premium (Proposition 15), and characterize the maximal jointly admissible sub-books as the independent sets of a masking hypergraph (Section 5.3).
The immersion framework and the masking construction are developed in the companion [8], from which we import the independent-enlargement lemma and which we restate with explicit revelation timing so the present results are self-contained. The driver-selection and causal-allocation methodology that an optimizer would employ is developed in [9]; we treat such an optimizer generically. Causal discovery—the estimation of the driver graph itself [10]–[12]—is orthogonal and upstream to the attribution problem studied here.
Fix a filtered probability space \((\Omega,\mathcal{F},(\mathcal{F}_t)_{0\le t\le T},\mathbb{P})\) carrying the price filtration \(\mathcal{F}\), the usual augmentation of a Brownian filtration [13]. For a fixed portfolio, \(\lambda_t=\mu_t/\sigma_t\) with \(\sigma_t>0\) is its price-of-risk process, treated as an element of the Hilbert space \(L^2(\mathbb{P}\otimes dt)\) of square-integrable \(dt\)-progressive processes, with inner product \(\langle a,b\rangle=\mathbb{E}\int_0^T a_t b_t\,dt\) and norm \(\|\cdot\|\).
Throughout, premium means a squared price-of-risk functional, equivalently an integrated conditional squared Sharpe-ratio value; it is not an expected excess-return level. Expected realized gain appears only after a trading strategy is specified (Section 5.2). We name the central object the squared price-of-risk premium (or Sharpe premium) of a driver filtration \(\mathcal{G}\), \[\mathrm{SPR}(\mathcal{G}):=\mathbb{E}\int_0^T\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\,dt=\|P_\mathcal{G}\lambda\|^2,\] the squared norm of the \(L^2\) projection \(P_\mathcal{G}\lambda\) of the price of risk onto \(\mathcal{G}\). We retain the lighter notation \(\pi(\mathcal{G})=\mathrm{SPR}(\mathcal{G})\) in informal passages but use \(\mathrm{SPR}\) in the technical statements where the distinction from an expected return matters. The choice of \(L^2\) is not a matter of convenience: the Sharpe premium is a second-moment object, so the relevant geometry is that of variance, and \(L^2\) is the unique \(L^p\) space whose inner product supplies the orthogonal projection and Pythagorean identity on which the loss term, and its distinction from the signed wedge, depend.
A driver information structure is a \(\sigma\)-field \(H\) (the information carried by drivers \(Y\), valued in \(M\subseteq\mathbb{R}^k\)) together with the enlarged filtration \(\mathcal{G}\), the usual augmentation of \(\mathcal{F}\vee\sigma(H_t)\). The structure is admissible if \(\mathcal{F}\hookrightarrow\mathcal{G}\), i.e.every \((\mathbb{P},\mathcal{F})\)-martingale remains a \((\mathbb{P},\mathcal{G})\)-martingale (immersion). Admissibility is the condition under which conditioning on the drivers does not anticipate the price innovation, so that the projected price of risk is a legitimate pricing object. We take from [9] the common-driver setting in which each portfolio conditional mean is measurable with respect to a shared driver manifold; we adopt it as setting rather than develop it.
Lemma 1 (Independence preserves immersion, initial and delayed). If \(H\) is a \(\sigma\)-field independent of \(\mathcal{F}_T\), the initial enlargement of \(\mathcal{F}\) by \(H\) preserves immersion. The same holds for the delayed enlargement \(\mathcal{G}_t=\mathcal{F}_t\vee\sigma(H)\mathbf{1}_{\{t\ge \tau\}}\) at a deterministic time \(\tau\). Hence drivers jointly independent of the terminal reference \(\sigma\)-field admit a globally immersed union filtration, whether revealed initially or at \(\tau\).
Proof. For the initial case it suffices to verify the \((\mathcal{H}')\) immersion criterion: every \((\mathbb{P},\mathcal{F})\)-martingale is a \((\mathbb{P},\mathcal{F}\vee\sigma(H))\)-martingale. Let \(M\) be a bounded \((\mathbb{P},\mathcal{F})\)-martingale and \(s<t\). For bounded \(\mathcal{F}_s\)-measurable \(G\) and bounded measurable \(h\), independence of \(H\) from \(\mathcal{F}_T\supseteq\sigma(M_t,M_s,G)\) gives \[\mathbb{E}[h(H)\,G\,(M_t-M_s)]=\mathbb{E}[h(H)]\,\mathbb{E}[G\,(M_t-M_s)]=0,\] the last equality by the \((\mathbb{P},\mathcal{F})\)-martingale property. Since such \(h(H)G\) generate \(\mathcal{F}_s\vee\sigma(H)\), we get \(\mathbb{E}[M_t-M_s\mid\mathcal{F}_s\vee\sigma(H)]=0\), i.e.\(M\) is a \((\mathbb{P},\mathcal{F}\vee\sigma(H))\)-martingale. For the delayed case, nothing is adjoined on \([0,\tau)\), so immersion is trivial there; on \([\tau,T]\) the added field \(\sigma(H)\) is independent of \(\mathcal{F}_T\), so the initial-enlargement argument applies verbatim to \((\mathcal{F}_t)_{t\ge\tau}\), and the two regimes glue at \(\tau\) because \(M_\tau\) is \(\mathcal{F}_\tau\)-measurable. The union-filtration statement follows by taking \(H\) the joint field of the independent drivers. See [8] for the general theory. ◻
Let \(\mathcal{G}^\star\) be the declared admissible benchmark: the largest driver filtration the analyst retains after admissibility screening, against which losses are measured. It is not an omniscient or “true” filtration; \(\lambda\) is taken \(\mathcal{G}^\star\)-measurable by the modeling convention that \(\mathcal{G}^\star\) is the reference universe, so the loss term below is not absolute unexplained premium but unexplained premium relative to this declared universe. In applications \(\mathcal{G}^\star\) is the most complete admissible driver set available.
Remark 1 (Reference-relative decomposition). If \(\mathcal{G}^\star_1\subseteq\mathcal{G}^\star_2\) are two admissible benchmarks, then for a fixed driver set \(A\) the intervention-stable component \(\pi^{\mathrm{do}}(A)\) and the confounding wedge are unchanged—both are defined intrinsically from \(Y_A\) and the causal kernels, without reference to \(\mathcal{G}^\star\)—while the information loss increases by exactly the additional projection residual \(\|\,P_{\mathcal{G}^\star_2}\lambda-P_{\mathcal{G}^\star_1}\lambda\,\|^2\ge0\) between the two benchmarks. The decomposition is thus reference-relative in its loss term only, and the causal terms are benchmark-invariant.
We first record that the premium is the value of a portfolio problem, then decompose it.
Proposition 2 (Investor problem behind \(\pi\)). Fix information \(\mathcal{G}_t\) and a portfolio with excess-return process \(dr_t=\mu_t\,dt+\sigma_t\,dW_t\), \(\sigma_t>0\) and \(\mathcal{G}_t\)-measurable (the drivers carry the conditional volatility), price of risk \(\lambda_t=\mu_t/\sigma_t\). For a \(\mathcal{G}\)-predictable position \(\phi\), the discounted wealth \(V^\phi\) with \(dV^\phi_t=\phi_t\,dr_t\) has instantaneous conditional drift \(\phi_t\mathbb{E}[\mu_t\mid\mathcal{G}_t]\) and instantaneous conditional quadratic variation rate \(\phi_t^2\sigma_t^2\) (since \(\sigma_t\) is \(\mathcal{G}_t\)-measurable). The pointwise mean–variance rate functional \[J_t(\phi):=\phi_t\,\mathbb{E}[\mu_t\mid\mathcal{G}_t]-\tfrac12\,\phi_t^2\,\sigma_t^2\] is maximized over \(\mathcal{G}_t\)-measurable \(\phi_t\) at \(\phi_t^\star=\mathbb{E}[\mu_t\mid\mathcal{G}_t]/\sigma_t^2\), with maximal rate \(\tfrac12\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\). Hence \[\pi(\mathcal{G})=\mathbb{E}\int_0^T\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\,dt =2\,\mathbb{E}\int_0^T \max_{\phi_t\;\mathcal{G}_t\text{-meas.}} J_t(\phi)\,dt,\] so \(\pi(\mathcal{G})\) is twice the integrated maximal conditional mean–variance rate attainable from \(\mathcal{G}\), equivalently the integrated maximal conditional squared Sharpe ratio of [1]. The benchmark \(\pi(\mathcal{G}^\star)\) is the value under the richest admissible information, and the loss \(\pi(\mathcal{G}^\star)-\pi(\mathcal{G})\) is the mean–variance tradeoff forgone by conditioning on \(\mathcal{G}\) rather than \(\mathcal{G}^\star\).
Proof. The drift and quadratic-variation rates of \(V^\phi\) follow from Itô’s isometry: over \([t,t+h]\), \(\mathbb{E}[V^\phi_{t+h}-V^\phi_t\mid\mathcal{G}_t]=\mathbb{E}[\int_t^{t+h}\phi_s\mu_s\,ds \mid\mathcal{G}_t]\) has rate \(\phi_t\mathbb{E}[\mu_t\mid\mathcal{G}_t]\), and \(\operatorname{Var}(V^\phi_{t+h}-V^\phi_t\mid\mathcal{G}_t)=\mathbb{E}[\int_t^{t+h}\phi_s^2\sigma_s^2\,ds\mid\mathcal{G}_t]\) has rate \(\phi_t^2\sigma_t^2\), the cross terms vanishing as \(o(h)\). Maximizing the scalar quadratic \(\phi\,m-\tfrac12\phi^2 v\) with \(m=\mathbb{E}[\mu_t\mid\mathcal{G}_t]\) and \(v=\sigma_t^2\) gives \(\phi^\star=m/v\) and value \(m^2/(2v) =\mathbb{E}[\mu_t\mid\mathcal{G}_t]^2/(2\sigma_t^2)\). Because \(\sigma_t\) is \(\mathcal{G}_t\)-measurable, \(\mathbb{E}[\mu_t\mid\mathcal{G}_t]/\sigma_t=\mathbb{E}[\lambda_t\mid\mathcal{G}_t]\), so the maximal rate is \(\tfrac12\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\). Integrating over \([0,T]\) and taking expectations, and using the Hansen–Richard identification of the maximal conditional squared Sharpe ratio with \(\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\), gives the stated value; monotonicity of the optimum in the conditioning \(\sigma\)-field gives \(\pi(\mathcal{G})\le\pi(\mathcal{G}^\star)\). ◻
The decomposition rests on two classical ingredients, stated as recalls. The first is the Hansen–Richard projection: conditioning on coarser information can only lower the attainable squared Sharpe ratio, the shortfall being an integrated conditional variance.
Lemma 2 (Projection optimality; Hansen–Richard [1]). Let \(\lambda_t(\mathcal{G})=\mathbb{E}[\lambda_t\mid\mathcal{G}_t]\) and \(\pi(\mathcal{G})=\mathbb{E}\int_0^T\lambda_t(\mathcal{G})^2\,dt=\|\mathbb{E}[\lambda\mid\mathcal{G}]\|^2\), and let \(\mathcal{G}^\star\) be the reference information set against which losses are measured: the filtration of the richest admissible driver structure entertained (of type (D1)), so that \(\lambda\) is \(\mathcal{G}^\star\)-measurable by construction. It is a benchmark, not an object that must be known—the decomposition below measures every shortfall relative to it, and in applications it is the most complete admissible driver set available rather than a claim about a true model. Then for any subfiltration \(\mathcal{G}\subseteq\mathcal{G}^\star\), by the orthogonality of nested \(L^2\) projections, \[\pi(\mathcal{G}^\star)-\pi(\mathcal{G}) =\big\|\mathbb{E}[\lambda\mid\mathcal{G}^\star]-\mathbb{E}[\lambda\mid\mathcal{G}]\big\|^2 ,\] and since \(\lambda\) is \(\mathcal{G}^\star\)-measurable so that \(\mathbb{E}[\lambda\mid\mathcal{G}^\star]=\lambda\), this equals the integrated conditional variance, \[\pi(\mathcal{G})=\pi(\mathcal{G}^\star)-\mathbb{E}\int_0^T\operatorname{Var}(\lambda_t\mid\mathcal{G}_t)\,dt\le\pi(\mathcal{G}^\star),\] with equality iff \(\lambda\) is \(\mathcal{G}\)-measurable. We record this only to fix notation for the loss term; it is the conditional-expectation projection of [1].
The premium is the \(L^2\) norm of the projected price of risk, projected as one object. The information-loss term for a driver set \(A\) is \(\pi(\mathcal{G}^\star)-\pi(\mathcal{G}^{Y_A})=\mathbb{E}\int_0^T\operatorname{Var}(\lambda_t\mid\mathcal{G}^{Y_A}_t)\,dt\), the squared norm of the projection residual, hence nonnegative.
Remark 3 (Benchmark invariance). The choice of \(\mathcal{G}^\star\) affects only the loss term: enlarging the benchmark leaves the intervention-stable component \(\pi^{\mathrm{do}}(A)\) and the wedge—both defined intrinsically from \(Y_A\) and the causal kernels, without reference to \(\mathcal{G}^\star\)—unchanged. The two causal terms are benchmark-free; only the bookkeeping of unexplained premium is measured relative to the reference set.
The second ingredient is the causal correction. We adopt a standard identification setup.
Assumption 1 (Causal identification). For the driver set \(A\) under study: (I1) the DAG \(D\) is known and a set \(U\) satisfying the back-door criterion for \(Y_A\to\lambda\) is available; (I2) the conditional densities entering the adjustment are strictly positive on the support of \(Y_A\) used; (I3) the intervention \(\mathrm{do}(Y_A)\) leaves the remaining structural mechanisms intact; (I4) the structural relations are stable over \([0,T]\) to the extent needed for the adjustment to be defined; (I5) the processes \((t,\omega)\mapsto Y_t,U_t,\lambda_t\) are jointly \(\mathcal{B}([0,T])\otimes\mathcal{F}\)-measurable, \(U\) is a (possibly path-valued) latent adjustment variable measurable at each \(t\), \(\lambda\in L^2(\mathbb{P}\otimes dt)\), and the conditional kernels entering the adjustment are regular.
Lemma 3 (Well-posedness of the interventional premium). Under Assumption 1, the back-door \(g\)-formula above is defined for a.e. \((t,\omega)\), the map \((t,\omega)\mapsto m^{\mathrm{do}}_A(t,Y_A(\omega))\) is \(\mathcal{B}([0,T])\otimes\mathcal{F}\)-measurable, and \(m^{\mathrm{do}}_A(\cdot,Y_A)\in L^2(\mathbb{P}\otimes dt)\); consequently \(\pi^{\mathrm{do}}(A)\) is finite.
Proof. Joint measurability of \((t,\omega)\mapsto Y_t,U_t,\lambda_t\) (I5) and regularity of the conditional kernels make \(\mathbb{E}[\lambda_t\mid Y_A=y,U=u]\) jointly measurable, so the integral over \(\mathbb{P}_U\) defines \(m^{\mathrm{do}}_A(t,y)\) for a.e.\((t,\omega)\) and the composition with \(Y_A\) is measurable. By conditional Jensen, \(\|m^{\mathrm{do}}_A(\cdot,Y_A)\|_{L^2(\mathbb{P}\otimes dt)}\le\|\lambda\|_{L^2(\mathbb{P}\otimes dt)} <\infty\) since \(\lambda\in L^2(\mathbb{P}\otimes dt)\) by (I5); Fubini applies, and \(\pi^{\mathrm{do}}(A)=\|m^{\mathrm{do}}_A(\cdot,Y_A)\|_{L^2(\mathbb{P}\otimes dt)}^2\) is finite. ◻
Assumption 2 (Faithfulness / no cancellation). Fix a structural causal model in which, for each \(t\), the conditional means \(\mathbb{E}[\lambda_t\mid Y_A]\) and \(m^{\mathrm{do}}_A(t,Y_A)\) are real-analytic functions of a finite parameter vector \(\vartheta\in\Theta\subseteq\mathbb{R}^d\) (the structural coefficients), the open set \(\Theta\) carrying Lebesgue measure; the linear–Gaussian family of Example 1 is the running instance. Open back-door paths from \(Y_A\) to \(\lambda\) then induce a nonzero observational–interventional mean distortion \(\delta_t:=\mathbb{E}[\lambda_t\mid Y_A]-m^{\mathrm{do}}_A(t,Y_A)\) for all \(\vartheta\) outside a Lebesgue-null (real-analytic) subvariety of \(\Theta\). This uses the standard fact that the zero set of a real-analytic function that is not identically zero has Lebesgue measure zero [14]: \(\delta_t\) is real-analytic in \(\vartheta\) and not identically zero precisely when a back-door path is open, so its zero set is null. Equivalently, whenever the distortion is nonzero the wedge is nonzero; the exact cancellation \(\|\mathbb{E}[\lambda\mid Y_A]\|=\|m^{\mathrm{do}}_A\|\) with \(\delta\neq0\) defines a lower-dimensional algebraic set, hence is nongeneric, as the Gaussian example below illustrates.
The interventional premium \(\pi^{\mathrm{do}}(A)\) is the value realized by a strategy that responds to the intervention-stable driver surface rather than the observational one. The gap between the two is the confounding wedge.
Theorem 4 (Wedge identity). Under Assumption 1 alone, for a driver set \(A\), \[\pi(\mathcal{G}^{Y_A})-\pi^{\mathrm{do}}(A) =\mathbb{E}\!\int_0^T\Big(\mathbb{E}[\lambda_t\mid Y_A]^2-m^{\mathrm{do}}_A(t,Y_A)^2\Big)\,dt .\] If the observational conditional mean equals the back-door-adjusted interventional mean, \(\mathbb{E}[\lambda_t\mid Y_A]=m^{\mathrm{do}}_A(t,Y_A)\) for a.e.\(t\), the wedge is zero; graphical blocking of all open back-door paths from \(Y_A\) to \(\lambda\) (so that the adjustment set \(U\) satisfies the back-door criterion) is a sufficient condition. The wedge can take either sign, unlike the nonnegative information loss. This identity requires no faithfulness or analyticity hypothesis.
Proof. Both terms are squared \(L^2\) norms; their difference is the stated integral by expanding each norm. The zero-wedge sufficiency is the back-door criterion [15], [16]; the sign is unconstrained because the integrand is a difference of squares. ◻
Theorem 5 (Generic nonzero wedge). Under Assumptions 1 and 2, if a back-door path from \(Y_A\) to \(\lambda\) is open (the adjustment \(U\) omitted), then the wedge of Theorem 4 is nonzero for all structural parameters outside a Lebesgue-null real-analytic subvariety of \(\Theta\); that is, the observational attribution is generically distorted away from the intervention-stable value.
Proof. By Theorem 4 the wedge vanishes iff \(\mathbb{E}[\lambda_t\mid Y_A]= m^{\mathrm{do}}_A(t,Y_A)\) a.e. Under Assumption 2 the distortion \(\delta_t=\mathbb{E}[\lambda_t\mid Y_A]-m^{\mathrm{do}}_A(t,Y_A)\) is real-analytic in the structural parameter \(\vartheta\) and not identically zero when a path is open, so by [14] its zero set is Lebesgue-null; off that set the wedge is nonzero. ◻
These assemble into the three-way split.
Theorem 6 (Anatomy of realizable premium). Under Assumption 1, fix a portfolio with \(\lambda\) measurable w.r.t. \(\mathcal{G}^\star\) and an admissible driver set \(A\) with \(\mathcal{G}^{Y_A}\subseteq\mathcal{G}^\star\). Then \[\underbrace{\pi(\mathcal{G}^\star)}_{\text{attainable}} =\underbrace{\pi^{\mathrm{do}}(A)}_{\text{captured}} +\underbrace{\big[\pi(\mathcal{G}^{Y_A})-\pi^{\mathrm{do}}(A)\big]}_{\text{confounding wedge (signed)}} +\underbrace{\big[\pi(\mathcal{G}^\star)-\pi(\mathcal{G}^{Y_A})\big]}_{\text{information loss }(\ge 0)} .\]
The three terms answer three distinct questions. The captured term asks how much premium is intervention-stable; the wedge—a causal distortion of the conditional Sharpe surface—asks how much of the observed conditional premium is not intervention-stable; and the loss asks how much premium is forgone because the driver set is informationally incomplete relative to \(\mathcal{G}^\star\). The taxonomy is summarized in Table 1.
| Term | Object | Sign / interpretation |
|---|---|---|
| \(\pi^{\dop}(A)\) | norm of interventional surface | \(\ge0\): stable causal contribution |
| \(\pi(\G^{Y_A})-\pi^{\dop}(A)\) | difference of squared norms | signed: observational (causal) distortion |
| \(\pi(\G^\star)-\pi(\G^{Y_A})\) | projection residual | \(\ge0\): missing admissible information |
The geometric structure of the three terms is central to the single-portfolio decomposition: the loss is a Pythagorean projection residual, the wedge is not.
Proposition 7 (Residual versus wedge). Let \(P_A\) be the \(L^2\) projection onto \(\mathcal{G}^{Y_A}\)-measurable processes. The information loss is a squared projection residual, \(\pi(\mathcal{G}^\star)-\pi(\mathcal{G}^{Y_A})=\|\lambda-P_A\lambda\|^2\) with \(\langle P_A\lambda,\lambda-P_A\lambda\rangle=0\), hence nonnegative. The confounding wedge is not of this form: it is a signed inner product, \[\pi(\mathcal{G}^{Y_A})-\pi^{\mathrm{do}}(A) =\big\langle P_A\lambda-\lambda^{\mathrm{do}}(Y_A),\,P_A\lambda+\lambda^{\mathrm{do}}(Y_A)\big\rangle,\] the inner product of the sum and difference of two \(\mathcal{G}^{Y_A}\)-measurable vectors, neither the orthogonal projection of the other. It therefore carries the sign of that inner product and is not a squared norm; in the Gaussian Example 1 it takes the values \(+1.12\), \(-0.48\), and \(0\) as the confounding coefficient varies, with fixed nonnegative loss—a sign change no projection residual can exhibit. Consequently the decomposition of Theorem 6 is additive but not Pythagorean.
Proof. The loss identity is the projection theorem, giving orthogonality and nonnegativity. For the wedge, both \(P_A\lambda=\mathbb{E}[\lambda\mid\mathcal{G}^{Y_A}]\) and the back-door-adjusted mean \(\lambda^{\mathrm{do}}(Y_A)=\int\mathbb{E}[\lambda\mid Y_A,U=u]\,d\mathbb{P}_U(u)\) are \(\mathcal{G}^{Y_A}\)-measurable, and \(\pi(\mathcal{G}^{Y_A})-\pi^{\mathrm{do}}(A)=\|P_A\lambda\|^2-\|\lambda^{\mathrm{do}}(Y_A)\|^2 =\langle P_A\lambda-\lambda^{\mathrm{do}}(Y_A),P_A\lambda+\lambda^{\mathrm{do}}(Y_A)\rangle\) by the polarization of a difference of squared norms. When every back-door path is blocked, \(\mathbb{E}[\lambda\mid Y_A,U]\) does not depend on \(U\), so \(\lambda^{\mathrm{do}}(Y_A)=P_A\lambda\) and the wedge vanishes. When a back-door path is open the two differ, and the Gaussian Example 1 computes the wedge explicitly as \((\kappa^2-b^2)v_Y\), whose sign is that of \((\kappa-b)(\kappa+b)\) and which is positive, negative, or zero as the confounding coefficient \(c\) varies while the loss stays fixed and positive. A signed quantity that changes sign under a parameter holding the loss fixed cannot be a squared projection residual, which proves the wedge is not of that form. ◻
Example 1 (Gaussian). Suppress \(t\). Let the chosen driver be scalar \(Y\), with confounder \(U\) and omitted driver \(Z\): \[U\sim N(0,1),\;Z\sim N(0,1),\;Y=aU+\varepsilon_Y,\;\lambda=bY+cU+dZ,\] \(\varepsilon_Y\sim N(0,\sigma_Y^2)\), all independent; \(A=\{Y\}\) with adjustment \(U\). Since \(\mathbb{E}[U\mid Y]=\tfrac{a}{a^2+\sigma_Y^2}Y\), \(\mathbb{E}[\lambda\mid Y]=\big(b+\tfrac{ca}{a^2+\sigma_Y^2}\big)Y=:\kappa Y\), while severing \(U\to Y\) gives \(m^{\mathrm{do}}_{\{Y\}}(y)=by\). With \(v_Y=a^2+\sigma_Y^2\), \[\pi(\mathcal{G}^Y)=\kappa^2 v_Y,\quad \pi^{\mathrm{do}}(\{Y\})=b^2 v_Y,\quad \text{wedge}=(\kappa^2-b^2)v_Y,\] and the loss from omitting \(Z\) is \(\mathbb{E}[\operatorname{Var}(\lambda\mid Y)]=c^2(1-a^2/v_Y)+d^2\). The wedge sign is that of \((\kappa-b)(\kappa+b)\) with \(\kappa-b=ca/v_Y\): \(a{=}1,b{=}0.5,c{=}0.8,\sigma_Y^2{=}1\) gives wedge \(+1.12\); \(c{=}-0.8\) gives \(-0.48\); and \(c=-2b\,v_Y/a\) gives \(\kappa=-b\), wedge \(0\) despite an open back-door path, which is why Assumption 2 is needed. (Monte Carlo with \(8\times10^6\) draws reproduces \(1.120,-0.480,0.000\).) In a multi-driver panel the same calculation holds with scalars replaced by covariance matrices: the wedge is the difference of two quadratic forms in the observational and back-door-adjusted conditional-mean coefficients.
A book pools the drivers of many portfolios. Write \(\lambda^B=\sum_j\alpha_j\lambda_j\) for the book’s price-of-risk process, the \(\alpha_j\) being attribution weights, and \(P_\cup\) for the \(L^2\) projection onto the pooled filtration \(\mathcal{G}^\cup=\bigvee_j\mathcal{G}^j\). Because premia are squared norms, pooling generates pairwise bilinear cross terms, so book wedges and losses do not add linearly across desks.
Proposition 8 (Book-level decomposition). The aggregation below is an attribution aggregation in the Hilbert space of portfolio-level price-of-risk processes, with \(\lambda^B=\sum_j\alpha_j\lambda_j\) the book attribution process of the preceding paragraph. For a jointly admissible book with union residuals \(r_j:=\lambda_j-P_\cup\lambda_j\), the information loss is \[\ell_B=\|\lambda^B-P_\cup\lambda^B\|^2 =\sum_j\alpha_j^2\|r_j\|^2+2\sum_{j<k}\alpha_j\alpha_k\langle r_j,r_k\rangle,\] so losses add only under residual orthogonality \(\langle r_j,r_k\rangle=0\); the residual cross-covariance is a distinct aggregation correction. The book wedge is \[w_B=\sum_j\alpha_j^2 w_j^{\cup} +2\sum_{j<k}\alpha_j\alpha_k\big(\langle P_\cup\lambda_j,\delta_k\rangle +\langle P_\cup\lambda_k,\delta_j\rangle-\langle\delta_j,\delta_k\rangle\big),\] with no higher-order terms, since a squared norm generates only pairwise bilinear terms. Both cross terms can be read off the network: they are generically nonzero when the shared driver lies on a back-door path to more than one incident price of risk, and may still vanish by orthogonality or exact cancellation in nongeneric configurations.
Proof. Expand both squared norms by bilinearity; the diagonals give the squared-weight marginal terms and the off-diagonals the stated cross terms. ◻
This algebraic non-additivity is the first, quantitative layer of aggregation. The second, qualitative layer—the subject of the next section—is structural: the pooled filtration \(\mathcal{G}^\cup\) may cease to be admissible, and then no decomposition applies to the book at all, since the projection is no longer onto an immersed filtration.
This is the paper’s central result. We do not classify all failures of immersion under driver pooling; we identify a minimal, well-defined obstruction class and prove existence and minimality within it. The class is the filtration-theoretic counterpart of Bernstein’s phenomenon, in which three events are pairwise but not mutually independent.
Definition 2 (Masking obstruction; pairwise-invisible obstruction). A family \(\{H_1,\dots,H_m\}\) of driver fields has a masking obstruction of order \(m\) (relative to the terminal price innovation generating \(\mathcal{F}_T\)) if every proper subfamily \(\{H_i:i\in I\}\), \(I\subsetneq\{1,\dots,m\}\), is independent of \(\mathcal{F}_T\), while the full \(m\)-tuple \(\bigvee_{i=1}^m H_i\) reveals a nontrivial function of \(\mathcal{F}_T\). Masking can occur at order two: with \(\varepsilon_2=\varepsilon_1 Z\) for \(Z\) a function of \(\mathcal{F}_T\) and \(\varepsilon_1\) independent of \(\mathcal{F}_T\), each singleton is independent of \(\mathcal{F}_T\) while the pair reveals \(Z\). We therefore study the stronger pairwise-invisible obstruction: the family is such that every singleton and every pair is independent of \(\mathcal{F}_T\) (so every one- and two-driver sub-book is immersed by Lemma 1 and passes every singleton and pairwise admissibility screen), while some larger subfamily reveals a function of \(\mathcal{F}_T\). “Order three” below is minimal relative to pairwise diagnostic regimes: it is the smallest order at which an obstruction can hide from all singleton and pairwise screens, not a claim that masking cannot occur algebraically at order two.
The next theorem establishes that a pairwise-invisible obstruction (every singleton and every pair admissible, the triple not) exists at order three, and that three is the minimal order at which an obstruction can hide from all singleton and pairwise screens. As noted in Definition 2, masking can occur algebraically at order two (the XOR relation \(\varepsilon_2=\varepsilon_1 Z\)); what cannot occur at order two is a pairwise-invisible obstruction, since the revealing pair is itself a two-driver sub-book that a pairwise screen detects.
Theorem 9 (Existence and minimality of a pairwise-invisible obstruction). Let \(Z=\operatorname{sgn}(B^1_T-B^1_{T/2})\in\mathcal{F}_T\), let \(\varepsilon_1,\varepsilon_2\) be independent Rademacher variables independent of \(\mathcal{F}_T\), and set \(\varepsilon_3=\varepsilon_1\varepsilon_2 Z\). Fix the revelation timing: each \(\varepsilon_i\) is the externally enlarged driver information of portfolio \(i\) of type (D2), revealed at the common time \(T/2\); that is, portfolio \(i\) has enlarged filtration \(\mathcal{G}^i\), the usual augmentation of \(\mathcal{F}\vee\sigma(\varepsilon_i \mathbf{1}_{\{t\ge T/2\}})\), and the pooled book has union filtration \(\mathcal{G}^{\cup}\), the augmentation of \(\mathcal{F}\vee\sigma(\varepsilon_1,\varepsilon_2, \varepsilon_3)\) activated at \(T/2\). Then \(\{\varepsilon_1,\varepsilon_2, \varepsilon_3\}\) is a pairwise-invisible obstruction in the sense of Definition 2:
for every proper \(I\subsetneq\{1,2,3\}\), \(\sigma(\varepsilon_i:i\in I)\) is independent of \(\mathcal{F}_T\), so each portfolio and each pair is jointly admissible (immersion holds by Lemma 1) and its decomposition is well posed;
\(\varepsilon_1\varepsilon_2\varepsilon_3=Z\) is a nontrivial function of a future increment, revealed at \(T/2\), so \(\mathcal{G}^{\cup}_{T/2}\supseteq\sigma(Z)\) and the triple is not jointly admissible.
Moreover three is minimal relative to pairwise diagnostic regimes: no pairwise-invisible obstruction exists at order one or two. At order one a single revealing field is itself detected by a singleton screen; at order two a revealing pair is itself a two-driver sub-book detected by a pairwise screen (as in the XOR relation \(\varepsilon_2=\varepsilon_1 Z\), where the pair \(\{\varepsilon_1, \varepsilon_2\}\) reveals \(Z\) and so fails pairwise invisibility). Three is therefore the smallest order at which the obstruction can hide from all singleton and pairwise admissibility screens. We make no claim that masking is impossible at order two, nor any classification claim about non-masking obstructions.
Proof. For (i): for \(I=\{3\}\), \(\mathbb{P}(\varepsilon_3=b\mid\mathcal{F}_T)=\mathbb{P}(\varepsilon_1\varepsilon_2=bZ\mid\mathcal{F}_T)=\tfrac12\) since \(\varepsilon_1\varepsilon_2\) is Rademacher independent of \(\mathcal{F}_T\); for \(I=\{1,3\}\), \(\mathbb{P}(\varepsilon_1=a,\varepsilon_3=b\mid\mathcal{F}_T) =\mathbb{P}(\varepsilon_1=a,\varepsilon_2=abZ\mid\mathcal{F}_T)=\tfrac14\), symmetrically for \(\{2,3\}\); \(\{1\},\{2\},\{1,2\}\) are immediate, since \(\sigma(\varepsilon_1, \varepsilon_2)\) is generated by two Rademacher variables independent of \(\mathcal{F}_T\) by construction—this is the relevant input to Lemma 1 for the \(\{1,2\}\) case, and it is precisely here that the dependence of \(\varepsilon_3\) on \(Z\) is quarantined: \(Z\) enters only through the third coordinate, so no pair recovers it. For each proper subcollection \(S\), independence of \(\sigma(S)\) from \(\mathcal{F}_T\) gives immersion of the delayed enlargement \(\mathcal{G}^S_t=\mathcal{F}_t\vee\sigma(S)\mathbf{1}_{t\ge T/2}\) by Lemma 1: nothing is added before \(T/2\), and from \(T/2\) on the added field is independent of \(\mathcal{F}_T\), so the initial-enlargement criterion applies to \((\mathcal{F}_t)_{t\ge T/2}\) and martingales are preserved throughout. For (ii): \(\varepsilon_1\varepsilon_2\varepsilon_3=Z\), so \(\mathcal{G}^{\cup}_{T/2}\supseteq\sigma(Z)\). Consider the \((\mathbb{P},\mathcal{F})\)-martingale \(t\mapsto B^1_t-B^1_{T/2}\) on \([T/2,T]\), with terminal increment \(X=B^1_T-B^1_{T/2}\sim N(0,T/2)\) and \(\mathbb{E}[X\mid\mathcal{F}_{T/2}]=0\). Under the enlargement, \(Z=\operatorname{sgn}(X)\) is \(\mathcal{G}^{\cup}_{T/2}\)-measurable, so \[\mathbb{E}[X\mid\mathcal{G}^{\cup}_{T/2}]=\mathbb{E}[X\mid Z]=Z\,\mathbb{E}|X|=Z\sqrt{T/\pi}\neq0 ,\] where \(\pi=3.14159\ldots\) is the constant (the only place it denotes the constant rather than the premium functional \(\pi(\cdot)\); the meaning is clear from the argument), using \(\mathbb{E}|X|=\sqrt{2\sigma^2/\pi}\) for \(X\sim N(0,\sigma^2)\) with \(\sigma^2=T/2\). Thus the \(\mathcal{F}\)-martingale \(B^1\) acquires a nonzero drift on \([T/2,T]\) under \(\mathcal{G}^{\cup}\) and is no longer a \(\mathcal{G}^{\cup}\)-martingale; immersion fails. ◻
The construction reveals information externally, at \(T/2\). One may ask whether a purely adapted economic mechanism—desks reacting to present information—can supply the same obstruction. Such a mechanism supplies the third-order masking structure but not the anticipation; we make the two ingredients precise, since the obstruction requires both and an adapted mechanism delivers only one.
Remark 10 (Two ingredients: masking and anticipation). Theorem 9 and the crowding model of Definition 3 below are best read as separating two ingredients of a pairwise-invisible obstruction. The first is combinatorial masking: no singleton or pair reveals the aggregate signal \(Z\), so the obstruction is invisible to all order-\(\le2\) screens; this is a combinatorial property of the driver fields. The second is anticipative coupling: the aggregate signal \(Z\) is not measurable at its revelation time but is a nontrivial function of a future innovation, so adjoining it breaks immersion; this is a filtration-theoretic property. Neither ingredient alone gives the obstruction studied here. Masking without anticipation—as in the crowding model, where \(Z\in\mathcal{F}_0\)—yields lower-order invisibility while immersion is preserved (Proposition 11); anticipation without masking—a single field or pair that reveals a future innovation—breaks immersion but is detected by a lower-order screen. Table 2 summarizes the two models.
| Model | Masking | Anticipation | Immersion failure |
|---|---|---|---|
| Theorem [thm:order3] (external revelation) | yes | yes | yes |
| Crowding (Definition [def:crowding]) | yes | no | no |
Definition 3 (Discrete crowding model). On \((\Omega,\mathcal{F},\mathbb{P})\) fix a two-date filtration \(\mathcal{F}_0\subseteq\mathcal{F}_1\). Let \(s_1,s_2,c\) be independent Rademacher (\(\pm1\), fair) random variables, all \(\mathcal{F}_0\)-measurable, and set the third desk’s position by the \(\mathcal{F}_0\)-measurable hedging rule \(s_3:=s_1 s_2 c\). Let \(R>0\) be a strictly positive random variable with \(\mathbb{E}R^2<\infty\), \(\mathcal{F}_1\)-measurable and independent of \((s_1,s_2,c)\). At date \(1\) a margin rule liquidates the aggregate position and realizes the return \[X_1:=R\,(s_1 s_2 s_3),\qquad Z:=\operatorname{sgn}X_1=s_1 s_2 s_3 ,\] so the realized direction is the crowding configuration and the magnitude \(R\) is exogenous. Define the pooled filtration \(\mathcal{G}_t=\mathcal{F}_t\vee\sigma(s_1,s_2,s_3)\).
Proposition 11 (What the crowding model does and does not give). In the model of Definition 3: (i) each \(s_i\) is \(\mathcal{F}_0\)-measurable, so no desk uses information beyond date \(0\); (ii) the reference increment \(X_1\), taken net of its \(\mathcal{F}_0\)-conditional mean, is a martingale increment generated by the positions through the liquidation map, not an exogenous innovation; (iii) each singleton \(s_i\) and each pair \(s_i s_j\) (\(i\neq j\)) is independent of \(Z\), while \(s_1 s_2 s_3=Z\) determines it, so the masking algebra \(s_3=s_1 s_2 Z\) of Theorem 9 holds. (iv) However, because \(Z=c\) is here \(\mathcal{F}_0\)-measurable, \(\sigma(Z)\subseteq\mathcal{F}_0\) and the pooled filtration does not* fail immersion: the model reproduces the masking algebra and the lower-order invisibility, but not the filtration obstruction itself.*
Proof. (i) holds by construction: \(s_1,s_2,c\) are \(\mathcal{F}_0\)-measurable and \(s_3=s_1 s_2 c\) is a measurable function of them, so \(Z=s_1 s_2 s_3=c\) is \(\mathcal{F}_0\)-measurable. (ii) Since \(X_1=R\,Z\) with \(Z\) \(\mathcal{F}_0\)-measurable and \(R>0\) independent of \(\mathcal{F}_0\), \(\mathbb{E}[X_1\mid\mathcal{F}_0]=Z\,\mathbb{E}R\); the centered increment \(\widetilde{X}_1:=X_1-\mathbb{E}[X_1\mid\mathcal{F}_0]=Z(R-\mathbb{E}R)\) satisfies \(\mathbb{E}[\widetilde{X}_1\mid\mathcal{F}_0]=0\). (iii) \(Z=c\) is a fair Rademacher drawn independently of \(s_1,s_2\), and \(s_3=s_1 s_2 c\); hence each \(s_i\) is independent of \(Z\), and each pair is independent of \(Z\): \(\mathbb{P}(s_1=a,s_2=b,Z=z)=\tfrac18=\mathbb{P}(s_1=a,s_2=b)\mathbb{P}(Z=z)\), and for \((s_1,s_3)\), since \((s_2,c)\perp s_1\), the pair is independent of \(Z=c\) by the same count; the triple gives \(s_1 s_2 s_3=c=Z\). (iv) \(\sigma(Z)=\sigma(c)\subseteq\mathcal{F}_0\), so adjoining \(\sigma(Z)\) adds nothing to \(\mathcal{F}_0\) and every \(\mathcal{F}\)-martingale remains a \(\mathcal{G}\)-martingale; immersion holds, and no drift is acquired. ◻
Remark 12 (Crowding versus revealed information: the irreducible distinction). Proposition 11 isolates exactly what an adapted economic mechanism supplies and what it does not. Read constructively, the crowding model demonstrates the harder half of the phenomenon: it shows how an adapted positioning rule can manufacture a third-order aggregate signal \(Z=s_1s_2s_3\) that is invisible to every pairwise test, with all the masking algebra of Theorem 9 present. What it does not supply, on its own, is anticipation: because \(Z=c\in\mathcal{F}_0\), conditioning on the triple reveals no future innovation, and the pooled filtration remains immersed. The crowding model produces the appearance of a third-order effect but not anticipation; the two are separate ingredients.
The distinction is therefore not that crowding fails, but that admissibility failure needs both ingredients, and crowding provides only the first. To obtain a genuine immersion failure the third-order aggregate signal must in addition be coupled to a future innovation—equivalently, the pooled object must be a price/order-flow statistic that reveals information not in \(\mathcal{F}_0\), so that \(\sigma(Z)\not\subseteq\mathcal{F}_t\) at the relevant \(t\). In the construction of Theorem 9 this coupling is supplied externally, with \(Z\in\mathcal{F}_T\setminus\mathcal{F}_{T/2}\); then immersion genuinely fails, but the revealing information is external to the adapted drivers. The two regimes—adapted-but-not- anticipative (crowding) and anticipative-but-externally-revealed (Theorem 9)—supply complementary halves, and cannot be merged within either model alone. Building a single dynamic mechanism in which adapted positioning endogenously generates a price filtration that reveals a future innovation—thereby supplying both halves at once—is the open problem we do not resolve here.
We separate immersion failure from the stronger notions it does not entail.
Remark 13 (Admissibility, not arbitrage). We separate four notions kept distinct throughout. Immersion (\(\mathcal{F}\hookrightarrow\mathcal{G}\), the \(\mathcal{H}\)-hypothesis) is that every \((\mathbb{P},\mathcal{F})\)-martingale remains a \((\mathbb{P},\mathcal{G})\)-martingale; what we prove is its failure. This is weaker than failure of the \(\mathcal{H}'\)-hypothesis (semimartingale preservation), of NUPBR (existence of a strictly positive local-martingale deflator), or of an equivalent local martingale measure: a filtration can fail immersion while still preserving semimartingales and admitting a deflator. Theorem 9 proves only the immersion failure. The sign \(Z\in\mathcal{F}_T\) becomes known at the deterministic time \(T/2\); this is an enlargement that adds nothing before \(T/2\) and adjoins \(\sigma(Z)\) from \(T/2\) on, i.e.\(\mathcal{G}_t=\mathcal{F}_t\) for \(t<T/2\) and \(\mathcal{G}_t=\mathcal{F}_t\vee\sigma(Z)\) for \(t\ge T/2\), an initial enlargement of \((\mathcal{F}_t)_{t\ge T/2}\) by \(\sigma(Z)\), not a progressive enlargement by a random time. Across \(T/2\) the reference martingale \(B^1\) acquires the drift \(\mathbb{E}[B^1_T-B^1_{T/2}\mid Z]=Z\sqrt{T/\pi}\), so it is no longer a \(\mathcal{G}\)-martingale. We do not use, and do not require, any failure of no-arbitrage: under standard initial-enlargement criteria (Jacod [4]) one may still obtain semimartingale preservation and a deflator, but this is not part of our contribution. The point we use is only that the pooled wedge and loss, though well defined as economic quantities, are computed on an information set that anticipates the future and so do not constitute an admissible projection.
Exact masking is the zero-noise limit of a continuous family: a noisy higher-order signal degrades the obstruction smoothly.
Proposition 14 (Approximate masking and the single-increment immersion defect). Let \(X\) be a future reference innovation with \(\mathbb{E}[X]=0\) (a martingale increment), \(Z=\operatorname{sgn}X\), and let three driver signs carry incremental third-order predictive information \[\varepsilon:=R^2\big(X\mid\varepsilon_1,\varepsilon_2,\varepsilon_3\big) -\max_{\{i,j\}}R^2\big(X\mid\varepsilon_i,\varepsilon_j\big)\;\in[0,1],\] the triple’s predictive \(R^2\) for \(X\) in excess of its best pair, with \(R^2\) the explained-variance fraction (well defined since \(X\) is centered). By the nested-projection inequality \(R^2(X\mid\varepsilon_1,\varepsilon_2,\varepsilon_3) \ge\max_{\{i,j\}}R^2(X\mid\varepsilon_i,\varepsilon_j)\), so \(\varepsilon\ge0\), with maximum \(2/\pi\) at exact sign masking. Let \(P_3\) and \(P_{ij}\) denote the \(L^2\) projections of \(X\) onto \(\sigma(\varepsilon_1,\varepsilon_2,\varepsilon_3)\) and onto \(\sigma(\varepsilon_i,\varepsilon_j)\); since \(\sigma(\varepsilon_i,\varepsilon_j) \subseteq\sigma(\varepsilon_1,\varepsilon_2,\varepsilon_3)\), the projections are nested and \(P_3 X-P_{ij}X\) is orthogonal to the range of \(P_{ij}\). Define the single-increment immersion defect* as the squared \(L^2\) norm of the incremental predictable drift, \[\Delta:=\big\|P_3 X-P_{i^\star j^\star}X\big\|_2^2, \qquad \{i^\star,j^\star\}=\mathop{\mathrm{arg\,max}}_{\{i,j\}}\|P_{ij}X\|_2^2,\] the orthogonal increment of the triple projection over the best pair projection, not a difference of norms. (This is a finite-horizon, single-innovation quantity, not a statement about immersion at all martingales and times.) Then \[\Delta=\operatorname{Var}(X)\,\varepsilon .\] Hence \(\Delta\to0\) as \(\varepsilon\to0\), and the exact obstruction of Theorem 9 is recovered as \(\varepsilon\uparrow 2/\pi\). For \(\varepsilon>0\) the pooled filtration acquires a forecasting advantage of size \(\operatorname{Var}(X)\varepsilon\) but, by Remark 13, no arbitrage; approximate masking degrades admissible projection continuously.*
Proof. Since \(X\) is centered, for any conditioning \(\sigma\)-field \(\mathcal{C}\) the projection satisfies \(\|P_{\mathcal{C}}X\|_2^2=\|\mathbb{E}[X\mid\mathcal{C}]\|_2^2 =\operatorname{Var}(X)\,R^2(X\mid\mathcal{C})\), the explained-variance identity. By nesting, \(P_{i^\star j^\star}X=P_{i^\star j^\star}P_3 X\), so \(P_3 X-P_{i^\star j^\star}X\) is the component of \(P_3 X\) orthogonal to \(\sigma(\varepsilon_{i^\star},\varepsilon_{j^\star})\), and by the Pythagorean identity \(\|P_3X-P_{i^\star j^\star}X\|_2^2=\|P_3X\|_2^2-\|P_{i^\star j^\star}X\|_2^2\). Substituting the explained-variance identity for each term gives \(\Delta=\operatorname{Var}(X)\big(R^2(X\mid\varepsilon_1,\varepsilon_2,\varepsilon_3) -R^2(X\mid\varepsilon_{i^\star},\varepsilon_{j^\star})\big) =\operatorname{Var}(X)\,\varepsilon\). At exact masking every pair has \(R^2=0\) while the triple has \(R^2=\mathbb{E}[X\mid Z]^2/\operatorname{Var}(X)=(T/\pi)/(T/2)=2/\pi\), recovering Theorem 9. ◻
The premium is a projection; realizing it requires a position that responds to the projected price of risk. The order-three obstruction is precisely where the projection algebra assigns value to information that is inadmissible.
Proposition 15 (Admissible realized premium versus inadmissible diagnostic premium). Let a portfolio have excess-return increment \(dr_t=\sigma_t\lambda_t\,dt+\sigma_t\,dW_t\) with \(\sigma_t>0\) and \(\mathcal{G}_t\)-measurable. (i) The \(\mathcal{G}\)-measurable position \(\phi^\star_t=\mathbb{E}[\lambda_t\mid\mathcal{G}_t]/\sigma_t\) realizes expected premium \(\mathbb{E}\int_0^T\phi^\star_t\,dr_t=\pi(\mathcal{G})\); the attainable projection is exactly what an information-responsive strategy realizes. (ii) If \(\mathcal{G}\) is admissible (\(\mathcal{F}\hookrightarrow\mathcal{G}\)), this is admissible realized premium: intervention-stable up to the confounding wedge of Theorem 4, and no strategy realizes premium from future information. (iii) If a pooled book filtration fails admissibility through the order-three masking relation of Theorem 9, then the same algebra, applied to the pooled signal \(\varepsilon_1\varepsilon_2\varepsilon_3=Z\), computes a value \(\tfrac12\mathbb{E}|X|>0\) that depends on the future increment \(X\), while every singleton- or pair-measurable position computes zero. This is inadmissible diagnostic premium: it quantifies the anticipative content of the inadmissible projection, and is not* a claim that the signal is available to a non-anticipative trader. It is a diagnostic statement about the economic content of the projection, invisible to all lower-order diagnostics and nonzero only at order three.*
Proof. (i) With \(\sigma_t\) \(\mathcal{G}_t\)-measurable, \(\mathbb{E}[\phi^\star_t\,dr_t\mid\mathcal{G}_t]=\mathbb{E}[\lambda_t\mid\mathcal{G}_t]\sigma_t^{-1}\cdot\sigma_t\mathbb{E}[\lambda_t\mid\mathcal{G}_t]\,dt =\mathbb{E}[\lambda_t\mid\mathcal{G}_t]^2\,dt\) (the noise term is mean-zero given \(\mathcal{G}_t\)); integrate and take expectations. (ii) Under immersion the projected price of risk is an admissible object and its causal content is \(\pi^{\mathrm{do}}\) plus the wedge by Theorem 4; the \(\mathcal{F}\)-martingale part contributes no predictable premium, so the realized premium is implementable by a non-anticipative trader. (iii) Under the masking relation \(Z=\varepsilon_1\varepsilon_2\varepsilon_3\in\mathcal{F}_T\) is measurable with respect to the pooled field, so the value \(\mathbb{E}[\,Z\cdot X\,]=\mathbb{E}[\operatorname{sgn}(X)X]=\mathbb{E}|X|>0\) is computed by the projection algebra; because \(Z\) is a function of the future increment, this value is diagnostic, not implementable. Any singleton or pair is independent of \(Z\) and hence of \(\operatorname{sgn}(X)\), giving zero by independence. 0◻ ◻
This is the link to the decision layer. The projections of this paper say what is attributable; an optimizer mapping admissible driver information to positions, as in the causal allocation framework of [9], realizes them. Operating on a book whose pooled filtration is inadmissible, such a strategy realizes the anticipative premium of Proposition 15(iii) while passing every pairwise diagnostic: the order-three obstruction is the exact point at which the projection algebra assigns value to inadmissible information. Restricting the optimizer to the maximal admissible sub-books eliminates the anticipative component.
Proposition 16 (Attribution domains are independent sets). Encode in the \(3\)-uniform hypergraph \(\mathcal{H}\) one hyperedge per driver triple realizing the masking relation of Theorem 9. Within this masking-obstruction class—that is, when the only obstructions present are order-three masking relations—the domains on which the book-level decomposition of Proposition 8 is admissibly defined are exactly the maximal independent sets of \(\mathcal{H}\). A portfolio common to several masking triples is the natural vertex to drop, since its removal breaks all of them simultaneously. We do not claim \(\mathcal{H}\) captures every possible obstruction; higher-order or non-masking obstructions, if present, would add further excluded sets.
Proof. By Theorem 9, a vertex set carrying a masking triple fails joint admissibility, and by construction \(\mathcal{H}\) records exactly those triples; so within the masking-obstruction class, joint admissibility fails iff the set contains a hyperedge of \(\mathcal{H}\), the admissible sub-books are the independent sets, and the maximal ones the maximal independent sets. On each, \(P_\cup\) projects onto an admissible filtration so Proposition 8 applies; on any set containing a hyperedge it does not. ◻
We demonstrate finite-sample behavior in controlled synthetic settings: the decomposition is estimable from data, the causal correction is necessary and recoverable, and the order-three screen detects planted leakage while producing no false positives among the sampled clean triples. The settings are synthetic and controlled, with the structural model known, so estimates can be checked against ground truth; no proprietary data is used and every figure is reproducible from the supplementary scripts. We do not claim a particular empirical frequency of the obstruction in real books—that is an empirical question beyond this paper—but the experiments show the tools work where the truth is known and degrade gracefully away from the exact construction.
In a linear–Gaussian single-driver model with a known confounder \(U\) entering both \(Y_A\) and \(\lambda\), the back-door (g-formula) estimator recovers the intervention-stable premium \(\pi^{\mathrm{do}}\) and the wedge, with mean absolute error decaying as \(T^{-1/2}\), while the confounding-agnostic estimator—ordinary regression of \(\lambda\) on \(Y_A\), the computation a standard factor-attribution method performs—returns the observational premium and so misattributes the wedge as captured. The gap is exactly the confounding wedge (Figure 1).
The scalar picture extends to a realistic panel. We take \(K=5\) observed drivers loaded on \(L=3\) latent confounders that also drive the price of risk, so the observational attribution mixes the direct causal channel with confounding through the shared latent factors. Here the population wedge is \(+197\%\) of the intervention-stable premium: a naive method would credit the drivers with roughly three times the premium that survives intervention. Figure 2 (a) shows the per-driver attribution under the naive and back-door methods—some drivers are over-credited, and one even flips sign—and panel (b) confirms the matrix g-formula estimator is consistent at the \(T^{-1/2}\) rate. The decomposition is therefore not a scalar curiosity: it operates on a covariance-matrix panel and the causal correction matters most precisely when confounding is strong.
The order-three masking obstruction suggests a concrete screen for candidate driver triples: (1)–(3) estimate the singleton, pairwise, and triple predictive \(R^2\) of the reference innovation; (4) form the incremental third-order signal \(\varepsilon=R^2(X\mid\varepsilon_1,\varepsilon_2,\varepsilon_3) -\max_{i,j}R^2(X\mid\varepsilon_i,\varepsilon_j)\); (5) calibrate \(\varepsilon\) against a permutation null that breaks the driver–future link, flagging the triple when the increment exceeds the upper-\(\alpha\) null quantile. All \(R^2\) are estimated out of sample (cell means fit on one half, scored on the other).
Figure 3 instantiates the population construction with \(X=B^1_T-B^1_{T/2}\), \(Z=\operatorname{sgn}(X)\), independent Rademacher \(\varepsilon_1,\varepsilon_2\), \(\varepsilon_3=\varepsilon_1\varepsilon_2 Z\): every singleton and pair has out-of-sample \(R^2\) indistinguishable from zero while the triple attains \(\approx2/\pi\) (A); the observed incremental order-three \(R^2\) lies far outside the permutation null (\(200\) permutations, \(p\approx0.005\)) while the best pair sits inside it (B); and the noise path, replacing the exact relation by \(X=\beta\,\operatorname{sgn}(\varepsilon_1\varepsilon_2\varepsilon_3)+\sigma\eta\), degrades the incremental \(R^2\) continuously while \(\Delta=\operatorname{Var}(X)\,\varepsilon\) stays constant (C), a finite-sample confirmation of Proposition 14.
Finally we run the screen as a screen: one masking triple is planted among clean drivers in a synthetic book, and we measure detection power and false positives. In the synthetic design considered here, Figure 4 (a) shows the screen separates the planted triple from clean triples and produces no false positives among the sampled clean triples at the chosen permutation threshold; panel (b) traces detection power against masking corruption \(q\) (the fraction of the triple’s relation replaced by noise) at two sample sizes, exhibiting the expected sample-size/signal-strength trade-off; panel (c) confirms the transition at an intermediate sample size. The screen is reliable when the masking relation is close to exact and the sample is adequate, and degrades gracefully otherwise.
The price of risk, and through it the premium a causal strategy can realize, is a function of the admissibility of the driver filtration. For a single portfolio the conditional price-of-risk attribution decomposes exactly into an intervention-stable component, a signed confounding wedge, and a nonnegative information loss, with the loss an \(L^2\) projection residual and the wedge a causal object that is not. This decomposition need not aggregate across portfolios pooling their drivers: we exhibit an obstruction at order three that is invisible to every singleton and pairwise admissibility screen, at which the pooled filtration ceases to be admissible and the projection algebra assigns value to information that anticipates future innovations. The maximal jointly admissible sub-books are the independent sets of the masking hypergraph.
Three limitations bound the claims. The wedge requires a specified causal graph and a valid adjustment set; we do not solve causal discovery [10]–[12], which is orthogonal and upstream to the attribution studied here. The experiments demonstrate that the decomposition is estimable, that the causal correction is necessary and recoverable, and that the order-three screen detects planted leakage with controlled false positives; we do not, however, claim a particular empirical frequency of the obstruction in real books—establishing that is a separate empirical undertaking. And the obstruction requires two ingredients: a third-order aggregate signal invisible to pairwise tests, and the coupling of that signal to a future innovation. A purely adapted crowding mechanism supplies the first but not the second, so it reproduces the masking algebra without the immersion failure (Proposition 11); whether an adapted book can endogenously supply both—generating a price filtration that reveals a future innovation—is left open. The contribution is thus both structural and operational: it identifies the object to look for, the order at which lower-order diagnostics are guaranteed to miss it, and a concrete screen for it, while leaving its empirical prevalence to future work.
All experiments are synthetic and seeded with numpy.random.default_rng; the parameters below fully specify each design, and the archived code (https://doi.org/10.5281/zenodo.20843643) regenerates every figure.
Single-driver linear–Gaussian model \(U,Z\sim N(0,1)\) independent, \(Y=aU+\varepsilon_Y\) with \(\varepsilon_Y\sim N(0,\sigma_Y^2)\), and \(\lambda=bY+cU+dZ\), with \((a,b,c,d)=(1.0,0.5,0.8,0.6)\) and \(\sigma_Y^2=1\). The adjustment set is \(U\). The back-door estimator regresses \(\lambda\) on \((Y,U)\) and averages over the empirical marginal of \(U\); the confounding-agnostic estimator regresses \(\lambda\) on \(Y\) alone. Consistency is assessed over sample sizes \(T\in\{250,\dots,2\times10^5\}\) with \(40\) replications per size.
\(K=5\) drivers and \(L=3\) latent confounders, with \(Y=AU+E\), \(E\sim N(0,S)\) diagonal, and \(\lambda=b^\top Y+c^\top U\); the loading matrix \(A\), the idiosyncratic variances \(S\), and the coefficient vectors \(b,c\) are fixed in the script. Per-driver credit is \(m_i(\Sigma_Y m)_i\) so the credits sum to the respective premium. Consistency uses the same sample-size grid and \(40\) replications.
\(X=B^1_T-B^1_{T/2}\sim N(0,T/2)\) with \(T=2\), \(Z=\operatorname{sgn}X\), independent Rademacher \(\varepsilon_1,\varepsilon_2\), and \(\varepsilon_3=\varepsilon_1\varepsilon_2 Z\), with \(N=2\times10^6\) samples. Predictive \(R^2\) is estimated out of sample on a \(50/50\) split. The permutation null uses \(200\) permutations of \(X\) against the fixed driver signs. The noise path replaces the exact relation by \(X=\beta\,\operatorname{sgn}(\varepsilon_1\varepsilon_2\varepsilon_3)+\sigma\eta\), \(\eta\sim N(0,1)\), over \(\sigma\in[0,3]\) on a \(13\)-point grid.
A synthetic book of clean Rademacher drivers with one planted triple \(\varepsilon_3=\varepsilon_1\varepsilon_2 Z\), optionally corrupted by replacing a fraction \(q\) of the relation with independent signs. The screen flags a triple when its out-of-sample incremental \(R^2\) exceeds the upper-\(5\%\) quantile of a permutation null. Detection power is averaged over \(40\) replications and traced against \(q\in[0.3,0.9]\) at sample sizes \(T\in\{400,800,3000\}\).
Funding The author received no specific funding for this work.
Conflict of interest The author declares no conflict of interest.
Data availability All experiments are synthetic and use no proprietary data. The code reproducing every figure is openly available at https://doi.org/10.5281/zenodo.20843643
(archived) and https://github.com/AlejandroRodriguezDominguez/order-three-attribution (development); each script is seeded, so every figure is exactly
reproducible.
Miralta Finance Bank S.A., Madrid, Spain, and Department of Computer Science, University of Reading, UK. arodriguez@miraltabank.com↩︎