The ‘Right’ Extension of Type-I Error to Data-Dependent Levels


Abstract

The literature on hypothesis testing with data-dependent and post-hoc significance levels relies on a particular extension of the Type-I error to data-dependent levels. Existing arguments for this extension are heuristic, and primarily motivated by a resulting connection to the E-value. Our main contribution is to argue that the extension is ‘right’, by showing that it emerges from three axioms: it is the only extension that nests classical Type-I error validity for data-independent levels, preserves classical validity for data-dependent levels and is monotone in the strength of the rejection claim. We subsequently apply this result to support the common definition of the E-value, by showing that it arises as the ‘right’ notion of validity for the numerical representation of a generalized hypothesis test that may reject at different data-driven significance levels.

Keywords: Data-dependent significance level, Post-hoc hypothesis testing, E-values

1 Introduction↩︎

The hypothesis testing framework of Neyman & Pearson has developed into one of the methodological pillars of modern empirical research. Perhaps the central idea within this framework is to choose a test that is valid: with a Type-I error bounded by some level of significance \(\alpha\): \[\begin{align} \textrm{Type-I error} \equiv P(\textrm{test falsely rejects the hypothesis}) \leq \alpha. \end{align}\]

In this classical framework, it is mandatory to choose the significance level in advance, or at least independently from the data. The standard argument is that for a data-driven choice of the level \(\widetilde{\alpha}\), there is generally no hope that the (conditional) Type-I error given chosen level \(\widetilde{\alpha} = a\) is always smaller than \(a\) \[\begin{align} \label{ineq:distortion} P(\textrm{test falsely rejects the hypothesis} \mid \widetilde{\alpha} = a) \leq a. \end{align}\tag{1}\]

This classical Neyman & Pearson framework was recently shaken by a series of papers on testing with data-dependent and even fully post-hoc selected significance levels. Showing that 1 is indeed hopeless, the key idea in this literature is to settle for a weaker target: controlling the expected distortion-ratio between the conditional Type-I error and the reported level \[\begin{align} \label{ineq:expected95distortion95ratio} \Ex^P\left[\frac{P(\textrm{test falsely rejects hypothesis} \mid \widetilde{\alpha})}{\widetilde{\alpha}}\right] \leq 1. \end{align}\tag{2}\]

Unfortunately, the arguments for this extension of the Type-I error to data-dependent levels are mostly heuristic, and perhaps primarily driven by a resulting connection to the E-value (see Section 5).

This leaves open the question:

Is this the ‘right’ extension of the classical Type-I error to data-dependent levels?

In particular: why should 2 involve a ratio? And why should it be bounded in expectation? We believe that answering these questions is critical in order to convince anyone to go beyond a framework as well established as that of Neyman and Pearson.

1.1 Contributions↩︎

The main contribution of this paper is to show that 2 is indeed the natural extension of the Type-I error to data-dependent levels. To support this claim, we show that, within a broad framework, it is the only extension of the Type-I error to data-dependent levels that satisfies three natural axioms:

  1. it should nest classical validity when using data-independent levels,

  2. it should preserve classical validity when using data-dependent levels: the unconditional probability that a test rejects at a data-dependent level smaller than a prespecified threshold \(a\) is bounded by \(a\),

  3. it does not ban obviously conservative choices of the data-dependent level: if a test is valid for data-dependent level \(\widetilde{\alpha}\) and \(\widetilde{\alpha}'\) yields a uniformly weaker decision, then the test is also valid for \(\widetilde{\alpha}'\).

In particular, we find that the first axiom pins the ratio within the expectation of 2 , for any reasonable notion of aggregation. The second and third axioms together subsequently pin down the expectation among such aggregators.

We apply this result to support the fundamental interpretation of the E-value as a generalization of a hypothesis test that may reject at different data-driven levels. The common definition of the E-value then emerges as the numerical representation that nests classical hypothesis testing.

1.2 Literature review↩︎

To the best of our knowledge, the first connection between E-values and testing with data-dependent levels was made by [1], [2] and [3], who find that E-values enable the post-hoc selection of the level in a multiple testing context.

This was developed into a generalization of the classical Neyman–Pearson framework to data-dependent levels in an expected loss setting by [4]. [5] subsequently goes beyond expected loss by proposing a framework that conditions on the data-dependent level \(\widetilde{\alpha}\) to capture 1 , introducing the expected distortion ratio 2 , connecting to the p-value and studying optimality. [6] introduces the idea to view the E-value as a generalization of a hypothesis test that rejects at some data-driven level, deriving an extension of the Neyman–Pearson lemma to E-values. [7] study admissibility within the framework of [4].

Post-hoc inference has already been applied in several settings, including conformal prediction [8][10], knockoffs [11], and equivalence testing [12]. Moreover, [13], [14] and [15] consider multiple testing with a data-driven level, and [16] study post-hoc inference in an asymptotic framework.

Our axioms are inspired by examples studied in [5]. In particular, the second axiom is used as an argument in [5] to dismiss another extension of the Type-I error to data-dependent levels: \[\begin{align} P(\textrm{test falsely rejects hypothesis}) \leq \Ex^P[\widetilde{\alpha}]. \end{align}\] This extension nevertheless appears in [8] and [10], showing how the literature struggles with whether 2 is truly the ‘right’ notion of validity to consider. A violation of our third axiom underlies an example in [5] that dismisses 1 as a reasonable extension. The current work goes far beyond these examples, showing that the combination of these axioms can in fact be used to dismiss any notion of validity besides 2 within a large class of options.

1.3 Notation and latent assumptions↩︎

We use \(\X\) to denote our sample space. Our results are all easily extended to the composite setting, so we focus on a simple (null) hypothesis \(\{P\}\), where \(P\) is a probability distribution on \(\X\).

Throughout, we assume that we have access to an external source of randomization. Formally, whenever needed we assume we may enrich our probability space with a uniform random variable \(U \sim \textrm{Unif}[0, 1]\) that is independent from any other random variables being considered. We suppress this in our notation, for the sake of brevity.

2 Testing with data-dependent levels↩︎

2.1 Classical testing↩︎

Throughout, we compare tests across different significance levels. For this reason, it is important to treat rejections at different levels as different decisions. We capture these decisions in what we call an evidence space, which is a totally ordered decision space \(\D\) with a least element that we denote by \(0\). We use this to define a level \(\alpha \in (0, 1)\) test as a \(\{0, d_\alpha\}\)-valued random variable, where \(0, d_\alpha \in \D\). Here, \(d_\alpha\) represents the decision to reject at significance level \(\alpha\) and \(0\) represents a non-rejection. We further assume that \(d_{\alpha^+} \leq d_{\alpha^-}\), for every \(\alpha^- \leq \alpha^+\), expressing that a rejection at a smaller level is a ‘stronger decision’.

To prepare for testing at data-dependent levels, we assume we have access to a family \(\phi\) of level \(\alpha\) tests \(\phi(\alpha) : \X \to \{0, d_\alpha\}\) across different levels \(\alpha \in (0, 1)\). Here, we say that \(\phi\) is valid for \(\alpha\) if the test \(\phi(\alpha)\) is valid, in the sense that the Type-I error is below \(\alpha\).

Definition 1 (Classical validity). \(\phi\) is valid at level \(\alpha\) if \(P(\phi(\alpha) = d_\alpha) \leq \alpha\).

2.2 Testing with data-dependent levels↩︎

We use \(\widetilde{\alpha} : \X \to (0, 1)\) to denote an arbitrary data-dependent level. To test with a data-dependent level, we plug such a data-dependent level into \(\phi\): \(\phi(\widetilde{\alpha})\). This should be viewed as using the decision produced by the random variable \(x \mapsto \phi(\widetilde{\alpha}(x))(x)\).

While it is easy to define a test with a data-dependent level, it is not obvious how to extend validity of tests from data-independent to data-dependent levels. The key problem is that the object \(\phi(\widetilde{\alpha})\) is no longer a test in the classical sense: its outcome is no longer binary: \(\{0, d\}\) for some \(d \in \D\). Indeed, \(\phi(\widetilde{\alpha})\) may produce a decision to reject at various different levels.

To work towards a notion of validity for data-dependent levels, we formulate three properties that we believe a reasonable extension of validity to data-dependent levels \(\widetilde{\alpha}\) should satisfy. These properties are the key to pinning a notion of validity for data-dependent levels.

A first natural property is that an extension of validity should nest classical validity: using a data-dependent level that happens to be independent of the data should coincide with classical validity. We formalize this in Property 1.

Property 1 (Nesting classical validity). \(\phi\) is valid for the data-dependent level \(\widetilde{\alpha} \equiv \alpha\) if and only if \(\phi\) is valid for \(\alpha\).

A second property is that it should preserve classical validity when using data-dependent levels. In particular, the probability that \(\phi(\widetilde{\alpha})\) yields a decision at least as strong as a rejection at level \(\alpha\) should be at most \(\alpha\). In other words, thresholding at \(d_\alpha\) yields a binary test that is classically valid at level \(\alpha\): \(\phi'(\alpha) = d_\alpha\ind{\phi(\widetilde{\alpha}) \geq d_\alpha}\).

Property 2 (Preserving classical validity). For every \(\widetilde{\alpha}\), if \(\phi\) is valid for \(\widetilde{\alpha}\) then \(P(\phi(\widetilde{\alpha}) \geq d_\alpha) \leq \alpha\), for every \(\alpha \in (0, 1)\).

A third property is a monotonicity property, which demands that the notion of validity should not ban obviously conservative choices. In words, it states that if \(\phi\) is valid for \(\widetilde{\alpha}\) and \(\widetilde{\alpha}'\) yields a weaker decision than \(\widetilde{\alpha}\), then \(\phi\) is also valid for \(\widetilde{\alpha}'\).

Property 3 (Monotonicity). If \(\phi\) is valid for \(\widetilde{\alpha}\) and \(\phi(\widetilde{\alpha}') \leq \phi(\widetilde{\alpha})\) pointwise, then \(\phi\) is valid for \(\widetilde{\alpha}'\).

3 Starting point: expected loss↩︎

As a starting point, we consider a class of notions of validity that can be expressed as an expected-loss bound. In particular, let \(L : \D \to \mathbb{R}\) be an increasing ‘loss function’, and let \(C \in \mathbb{R}\). We then say that \(\phi\) is valid for a data-dependent level \(\widetilde{\alpha}\) if \[\begin{align} \phi \textrm{ is valid for } \widetilde{\alpha} \iff \Ex^P[L(\phi(\widetilde{\alpha}))] \leq C. \end{align}\] Here, we assume \(C > L(0)\), which ensures there exist families of tests \(\phi\) that are valid, beyond \(\phi \equiv 0\).

It is convenient to normalize the loss function and threshold \(C\), because different loss function-threshold pairs \((L, C)\) yield the same notion of validity. Indeed, any positive affine transformation \((L, C) \to (aL + b, aC + b)\), with \(a > 0\) and \(b \in \mathbb{R}\), leaves the validity criterion unchanged. We therefore represent such a class of loss function-threshold pairs by a single canonical choice. In particular, a pair \((L, C)\) can be represented by the pair \((\overline{L}, 1)\), where \(\overline{L}(x) = (L(x) - L(0)) / (C - L(0))\). We then have that \(\overline{L}(0) = 0\) and \[\begin{align} \Ex^P[\overline{L}(\phi(\widetilde{\alpha}))] \leq 1 \iff \Ex^P[L(\phi(\widetilde{\alpha}))] \leq C. \end{align}\]

The expected distortion ratio 2 corresponds to the choice \(\overline{L}(d_\alpha) = 1/\alpha\) (see Remark 2). Theorem 4 shows that this is the only notion of validity for data-dependent levels that can be expressed as an expected loss bound and nests classical validity. The proof may be found in the appendix, alongside all other omitted proofs.

Theorem 4. Consider a notion of validity that can be expressed as an expected loss bound. Then this notion nests classical validity if and only if \(\overline{L}(d_\alpha) = 1/\alpha\) for every \(\alpha \in (0, 1)\).

Remark 1 (Relationship to [4]). The expected loss framework presented here is a notationally simplified version of the framework of [4]. Within this framework, [4] suggests choosing \(\overline{L}(d_\alpha) = 1 / \alpha\). Theorem 4 formalizes this choice, though the underlying reasoning is not highly involved.

Remark 2 (Link expected loss and distortion ratio). The choice \(\overline{L}(d_\alpha) = 1/\alpha\) yields control of the expected distortion ratio 2 . Indeed, we then have \[\begin{align} \label{eq:conditional95link95loss95distortion} \Ex^P[\overline{L}(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = \frac{P(\phi(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \mid \widetilde{\alpha})}{\widetilde{\alpha}}, \end{align}\tag{3}\] since \(\overline{L}(0) = 0\). Hence, by the law of iterated expectations, we have \[\begin{align} \Ex^P[\overline{L}(\phi(\widetilde{\alpha}))] = \Ex^P\left[\frac{P(\phi(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \mid \widetilde{\alpha})}{\widetilde{\alpha}}\right]. \end{align}\]

4 Beyond expected loss↩︎

Unfortunately, the expected loss framework is not rich enough to cover the natural notion of validity \[\begin{align} \label{ineq:dd95validity95strong} P(\phi(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \mid \widetilde{\alpha}) \leq \widetilde{\alpha}, \end{align}\tag{4}\] for (almost) every realization of \(\widetilde{\alpha}\), featured in 1 in the introduction. This means that Theorem 4 does not cover the notion of validity for data-dependent levels that forms the classical counterargument against data-dependent levels.

The gap between 4 and the expected loss framework from Section 3 is bridged by the more general framework introduced in [5], which is based on conditioning on the data-dependent level \(\widetilde{\alpha}\). This generalization comes with substantive consequences for our results: we find that Property 1 by itself does not suffice to pin a single notion of validity for data-dependent levels within this conditioning framework.

To introduce the framework, we use 3 to rewrite 4 as \[\begin{align} \label{ineq:dd95validity95strong95essup} \esssup_P \frac{P(\phi(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \mid \widetilde{\alpha} )}{\widetilde{\alpha}} \equiv \esssup_P \Ex^P\left[L(\phi(\widetilde{\alpha}))\;\middle|\;\widetilde{\alpha} \right] \leq 1, \end{align}\tag{5}\] for \(L(d_\alpha) = 1/\alpha\) and \(L(0) = 0\). Observing that 5 yields an overly conservative notion of post-hoc validity, [5] considers replacing the essential supremum by something weaker. We can capture this by introducing a certainty equivalent \(\rho\), which ranges from \([L(0), \infty]\)-valued random variables to \([L(0), \infty]\) in a way that fixes constant random variables: \[\begin{align} X \equiv c \implies \rho(X) = c. \end{align}\] This leads to a more general notion of validity for data-dependent levels.

Definition 2 (General validity). For a certainty equivalent \(\rho\) and increasing loss function \(L : \D \to \mathbb{R}\) with \(C > L(0)\), we say \[\begin{align} \label{ineq:validity95general} \phi \textrm{ is valid for } \widetilde{\alpha} \iff \rho(\Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}]) \leq C. \end{align}\tag{6}\]

For \(\rho = \Ex^P\), this notion of validity recovers the expected loss setting by the tower property. For \(\rho = \esssup_P\) with \(L(d_\alpha) = 1/\alpha\), \(L(0) = 0\) and \(C = 1\) this recovers the notion of validity considered in 4 (see Remark 3).

For this more general notion of validity, Theorem 5 shows that nesting classical validity is sufficient to force \(\overline{L}(d_\alpha) = 1/\alpha\) on \(\alpha \in (0, 1)\) for any certainty equivalent \(\rho\). At the same time, this condition places no restriction on the certainty equivalent \(\rho\) itself, as illustrated in Example 1.

Theorem 5. Consider a notion of validity of the form in 6 . Then, classical validity is nested if and only if \[\begin{align} \overline{L}(d_\alpha) := \frac{L(d_\alpha) - L(0)}{C - L(0)} = \frac{1}{\alpha}. \end{align}\]

Remark 3. Within this conditioning framework, [5] suggests to also use \(\overline{L}(d_\alpha) = 1 / \alpha\), based on heuristic arguments. Theorem 5 formalizes this choice.

Remark 4. For \(\rho = \esssup_P\), with \(L(d_\alpha) = 1 / \alpha\), \(L(0) = 0\), and \(C = 1\) we obtain 4 . Indeed, following 3 , we have \[\begin{align} \esssup_P \Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = \esssup_P \frac{P(\phi(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \mid \widetilde{\alpha})}{\widetilde{\alpha}}. \end{align}\]

Example 1 (\(\esssup_P\) nests classical validity). The choice \(\rho = \esssup_P\) with \(L(d_\alpha) = 1/\alpha\) and \(L(0) = 0\) nests classical validity since \(\widetilde{\alpha} \equiv \alpha\) yields \[\begin{align} \esssup_P \frac{P(\phi(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \mid \widetilde{\alpha} )}{\widetilde{\alpha}} = \frac{P(\phi(\alpha) = d_{\alpha})}{\alpha}. \end{align}\]

4.1 Main result: pinning a notion of validity↩︎

As nesting classical validity does not restrict the certainty equivalent \(\rho\), we must impose additional conditions to pin a notion of validity based on a specific choice of \(\rho\).

This leads to our main result: Theorem 6, which shows that Property 2 (Preserving classical validity) and Property 3 (Monotonicity) are sufficient to pin the notion of validity \[\begin{align} \Ex^P[\overline{L}(\phi(\widetilde{\alpha}))] \leq 1, \end{align}\] with \(\overline{L}(d_\alpha) = 1/\alpha\) and \(\overline{L}(0) = 0\), under the mild technical condition that \(\rho\) is continuous from below: \(Y_n \uparrow Y\) implies \(\rho(Y_n) \uparrow \rho(Y)\).

Theorem 6. Assume \(\rho\) is continuous from below. Assume that our notion of validity 6 both nests and preserves classical validity, and is monotone. Then the normalized loss equals \(\overline{L}(d_\alpha) = 1/\alpha\) for every \(\alpha \in (0, 1)\), and for every \(\phi\) and \(\widetilde{\alpha}\) we have \[\begin{align} \phi \textrm{ is valid for } \widetilde{\alpha} \iff \Ex^P[\overline{L}(\phi(\widetilde{\alpha}))] \leq 1. \end{align}\]

The proof of Theorem 6 relies on constructing data-dependent levels that violate the extension of classical validity or monotonicity. Example 2 illustrates such a construction to show that \(\rho = \esssup_P\) violates monotonicity.

Example 2 (\(\rho = \esssup_P\) violates monotonicity). Let \(p \sim \textrm{Unif}[0,1]\) and let \(\phi(\alpha)\) reject at level \(\alpha\) whenever \(p\leq \alpha\). The fixed data-dependent level \(\widetilde{\alpha}_0 \equiv 0.01\) is valid under the \(\esssup_P\) criterion, since \(P(p\leq0.01)/0.01=1\).

Now consider the more conservative data-dependent level \[\begin{align} \widetilde{\alpha}_1 = \begin{cases} 0.02, & p \leq 0.01, \\ 0.01, & p > 0.01. \end{cases} \end{align}\] Then \(\widetilde{\alpha}_1\geq \widetilde{\alpha}_0\), so \(\phi(\widetilde{\alpha}_1)\leq\phi(\widetilde{\alpha}_0)\) pointwise. However, conditional on \(\widetilde{\alpha}_1=0.02\), rejection occurs with probability one. Hence \[\begin{align} \esssup_P \frac{ P(\phi(\widetilde{\alpha}_1)=d_{\widetilde{\alpha}_1} \mid \widetilde{\alpha}_1) }{ \widetilde{\alpha}_1 } \geq \frac{1}{0.02} > 1. \end{align}\] As a result, \(\rho = \esssup_P\) violates monotonicity.

Remark 5 (Relationship to [5]). Example 2 is taken from [5], where it is used to dismiss \(\rho = \esssup_P\) as a candidate notion of validity. This example formed the source of inspiration for the monotonicity property, which we use in Theorem 6 to dismiss a much larger class of options for \(\rho\).

The extension property is inspired by Proposition 5 of [5], who uses this property to dismiss a particular subset of generalized means as options for \(\rho\). Theorem 6 shows that these ideas can be taken much further, pinning a single notion of validity.

Remark 6 (Post-hoc validity). Our results characterize validity for a single data-dependent level \(\widetilde{\alpha}\). The main focus of the literature has been post-hoc validity: validity for every data-dependent level. Theorem 6 immediately extends to the post-hoc setting, and simply forces \(\Ex^P[\overline{L}(\phi(\widetilde{\alpha}))] \leq 1\) for every \(\widetilde{\alpha}\).

5 E-values as a generalization of a test↩︎

This work can also be used to support the proposal to view the E-value as a multi-significance level generalization of a hypothesis test [5], [6].

An E-value is commonly defined as ‘some \([0, \infty]\)-valued random variable’ \(\e\) that satisfies \[\begin{align} \label{dfn:e-value95old} \Ex^P[\e] \leq 1. \end{align}\tag{7}\] We believe this definition does not do justice to the true nature of the E-value.

To support this claim, recall that we defined a level \(\alpha\) test \(\phi(\alpha)\) as a map \(\phi(\alpha) : \X \to \{0, d_\alpha\}\), where \(0, d_\alpha \in \D\) represent the decision to not reject and the decision to reject at level \(\alpha\). Adding additional rejection decisions into its codomain, we can generalize the notion of a hypothesis test: \(\{0, d_{\alpha_1}, d_{\alpha_2}, d_{\alpha_3}, \dots\}\). We believe this is how the E-value should fundamentally be defined: as a multi-decision generalization of a hypothesis test: \(\e : \X \to \D\). That is, an E-value is a generalization of a hypothesis test that produces a decision to reject at a data-dependent significance level.

Definition 3 (E-value: fundamental). We say that \(\e : \X \to \D\) is an E-value.

In relation to the current manuscript, the key implication is that this implies our main object of study \(\phi(\widetilde{\alpha})\) is an E-value! Our results can therefore be interpreted as defining a notion of validity for the E-value which appropriately extends the classical Type-I error. Viewing our results in this light suggests that the right notion of validity for the E-value is \[\begin{align} \Ex^P[\overline{L}(\e)] \leq 1, \end{align}\] with \(\overline{L}(d_\alpha) = 1/\alpha\) and \(\overline{L}(0) = 0\).

The common definition of an E-value as a \([0, \infty]\)-valued map can then be recovered as a convenient numerical representation of the underlying decision space \(\D\), by associating realizations of \(\e\) to the numerical value \(\overline{L}(\e)\). This aligns with the proposal in [6] to numerically represent a level \(\alpha\) test as a \(\{0, 1/\alpha\}\)-valued map, instead of the classical \(\{0, 1\}\)-valued representation.

If we restrict ourselves to \(\alpha \in (0, 1)\), the numerical representation coming from tests technically only permits an E-value to take value in \(\{0\} \cup (1, \infty)\), where the \(\{0\}\) comes from \(\overline{L}(0) = 0\). While there is nothing inherently wrong with this, we can obtain the full codomain \([0, \infty]\) by associating the decision to not reject to \(\alpha = \infty\) and also add the decisions \(d_\alpha\) to reject at levels \(\alpha \in \{0\} \cup [1, \infty)\). This simply means that the evidence space \(\D\) of an E-value is richer than that of classical hypothesis tests.

5.1 Abstract post-hoc validity↩︎

This view on the E-value also gives a clean interpretation of the connection between E-values and post-hoc validity. We briefly explain this here for interested readers; see also Section 7 in [5] for a slightly more abstract framework that drops \(\alpha\) altogether.

For the result, we need a slightly more abstract notion of monotonicity than presented in Property 3: if \(\e\) is valid and \(\e' \leq \e\) pointwise, then \(\e'\) is valid. Moreover, we also need to define continuity from below for a notion of validity: if \((\e_n)_{n \geq 1}\) is an increasing sequence of valid evidence-valued decisions and \(\e = \sup_{n \geq 1} \e_n\) is a \(\D\)-valued random variable, then \(\e\) is valid. In addition, we assume that \(\phi\) can be approximated post-hoc, in the sense that there exists a sequence of data-dependent levels \((\widetilde{\alpha}_n)_{n \geq 1}\) such that the sequence \(\phi(\widetilde{\alpha}_n)\) is increasing and \[\begin{align} \phi(\widetilde{\alpha}_n) \uparrow \sup_{\alpha \in (0, 1)} \phi(\alpha). \end{align}\]

To present the result, we define the E-value of a family of tests \(\phi\) as \[\begin{align} \e_\phi := \sup_{\alpha\in(0,1)} \phi(\alpha), \end{align}\] which we assume always exists and is a \(\D\)-valued random variable. We can use this E-value to define the closure \(\overline{\phi}\) of \(\phi\) as the threshold family generated by \(\e_\phi\): \[\begin{align} \overline{\phi}(\alpha) = d_\alpha \ind{\e_\phi\geq d_\alpha}. \end{align}\] By construction, \(\phi\) is dominated by its closure: \(\phi(\alpha) \leq \overline{\phi}(\alpha)\).

We are now ready to present the result that links E-values and post-hoc validity.

Theorem 7 (Post-hoc validity and E-values). Consider a notion of validity that satisfies the abstract monotonicity and continuity from below conditions. Assume that \(\phi\) can be approximated post-hoc. If \(\phi\) is post-hoc valid, then its closure \(\overline{\phi}\) is post-hoc valid. Moreover, \(\overline{\phi}\) is post-hoc valid if and only if its associated E-value \(\e_{\overline{\phi}}\) is valid.

6 Proof of Theorem 4↩︎

Proof. Fix a level \(\alpha\) and define the shorthands \(p_\alpha(\phi) := P(\phi(\alpha) = d_\alpha)\) and \(\ell_\alpha := \overline{L}(d_\alpha)\). Since \(\phi(\alpha)\) is \(\{0, d_\alpha\}\)-valued and \(\overline{L}(0) = 0\), we have \[\begin{align} \Ex^P[\overline{L}(\phi(\alpha))] = p_\alpha(\phi) \ell_\alpha. \end{align}\] In this notation, classical fixed-\(\alpha\) validity corresponds to \[\begin{align} p_\alpha(\phi) \leq \alpha, \end{align}\] and expected-loss validity corresponds to \[\begin{align} p_\alpha(\phi) \ell_\alpha \leq 1. \end{align}\]

By the external randomization assumption, for every \(p \in [0,1]\) there exists an event \(B\) with \(P(B) = p\). Hence the test \[\begin{align} \phi(\alpha) = d_\alpha \ind{B} \end{align}\] has \(p_\alpha(\phi) = p\). This means nesting classical validity at level \(\alpha\) is equivalent to \[\begin{align} p \leq \alpha \iff p\ell_\alpha \leq 1, \textrm{ for every } p \in [0,1]. \end{align}\] As this needs to hold for every \(p\), it forces \(\ell_\alpha = 1/\alpha\). Indeed, for \(p = \alpha\) we have \(\alpha\ell_\alpha\leq1\), so that \(\ell_\alpha \leq 1/\alpha\). On the other hand, for every \(p>\alpha\), the equivalence implies \(p \ell_\alpha > 1\), and so \(\ell_\alpha > 1/p\). Letting \(p \downarrow \alpha\) gives \(\ell_\alpha \geq 1/\alpha\). Hence \(\ell_\alpha=1/\alpha\). ◻

7 Proof of Theorem 5↩︎

Proof. Fix a constant level \(\widetilde{\alpha} \equiv \alpha\) and again write \[\begin{align} p_\alpha(\phi) := P(\phi(\alpha) = d_\alpha). \end{align}\] Since \(\phi(\alpha) \in \{0, d_\alpha\}\), \[\begin{align} \Ex^P[L(\phi(\alpha))] = L(0) + p_\alpha(\phi) (L(d_\alpha) - L(0)). \end{align}\] Moreover, for a constant level, the conditional expectation itself is constant \[\begin{align} \Ex^P[L(\phi(\widetilde{\alpha}))\mid\widetilde{\alpha}] = \Ex^P[L(\phi(\alpha))]. \end{align}\] Since \(\rho\) fixes constants, \[\begin{align} \rho\left(\Ex^P[L(\phi(\widetilde{\alpha}))\mid\widetilde{\alpha}]\right) = L(0) + p_\alpha(\phi) (L(d_\alpha) - L(0)). \end{align}\] Hence, at fixed levels, the \(\rho\)-notion of validity is equivalent to \[\begin{align} p_\alpha(\phi)(L(d_\alpha) - L(0)) \leq C - L(0). \end{align}\] This is exactly the condition studied in the expected-loss case, so that the proof follows from the same reasoning as in Appendix 6. ◻

8 Proof of Theorem 6↩︎

8.1 Technical lemma↩︎

To prove Theorem 6, we first prove a technical lemma. The lemma shows that any bounded random variable \(Y \geq L(0)\) satisfying a bound on its expectation can be replicated as the conditional expectation of a test with a data-dependent level \(\phi(\widetilde{\alpha})\) that has certain properties.

As the conditions of Theorem 6 require certain properties to hold for every test and data-dependent level, we can use this lemma construct particular examples that lead to the claim of the theorem.

Lemma 1. Assume that \(\overline{L}(d_\alpha) = 1/\alpha\). Let \(Y \geq L(0)\) be an arbitrary random variable that is bounded from above.

  1. If \(\Ex^P[Y] < C\), then there exists a pair \((\phi, \widetilde{\alpha})\) and a fixed level \(a \in (0, 1)\) such that \(Y = \Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}]\), \(P(\phi(a) = d_a) \leq a\) and \(\phi(\widetilde{\alpha}) \leq \phi(a)\), pointwise.

  2. If \(\Ex^P[Y] > C\), then there exists a pair \((\phi, \widetilde{\alpha})\) and a fixed level \(a \in (0, 1)\) such that \(\Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = Y\) and \(P(\phi(\widetilde{\alpha}) \geq d_a) > a\).

Proof. Define \(\overline{Y} := (Y - L(0)) / (C - L(0))\). Since \(Y \geq L(0)\) and \(Y\) is bounded from above, there exists an \(M > 0\) such that \(0 \leq \overline{Y} \leq M\).

We start by describing a randomized construction that is used to prove both claims. Let \(\widetilde{\alpha}\) be a data-dependent level such that \(\overline{Y}\) is measurable with respect to \(\widetilde{\alpha}\) and \(\overline{Y} \leq 1/\widetilde{\alpha}\). Recall the externally randomized random variable \(U \sim \textrm{Unif}[0, 1]\), independent from the \((\overline{Y}, \widetilde{\alpha})\), and define the event \[\begin{align} A := \{\overline{Y} \geq U / \widetilde{\alpha}\} \end{align}\] Now, define the family of tests which rejects at every level on the event \(A\), \[\begin{align} \phi(\alpha) := d_\alpha 1_A, \quad \textrm{ for every } \alpha \in (0, 1). \end{align}\] As \(U\) is independent, we have \[\begin{align} P(A \mid \widetilde{\alpha}) = \widetilde{\alpha} \overline{Y}. \end{align}\] Hence, by \(\overline{L}(d_\alpha) = 1 / \alpha\), we have \[\begin{align} \Ex^P[\overline{L}(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = P(A \mid \widetilde{\alpha}) / \widetilde{\alpha} = \overline{Y}. \end{align}\] Equivalently, on the unnormalized scale, \[\begin{align} \Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = L(0) + (C - L(0)) \overline{Y} = Y. \end{align}\] It therefore remains to choose \(\widetilde{\alpha}\) in such a way that the desired properties in the claims hold.

For the first claim, suppose that \(\Ex^P[Y] < C\), which is equivalent to \(\Ex^P[\overline{Y}] < 1\). Choose \(\delta > 0\) such that \((1 + \delta) \Ex^P[\overline{Y}] \leq 1\), and select \(a \in (0, 1)\) to be sufficiently small so that both \(a (1 + \delta) < 1\) and \(a (1 + \delta) M \leq 1\). We then define the data-dependent level \[\begin{align} \widetilde{\alpha} := a (1 + \delta \overline{Y} / M). \end{align}\]

This level satisfies \(\widetilde{\alpha} \in [a, a (1 + \delta)]\), so that \(\widetilde{\alpha} \geq a\). In addition, \(\overline{Y}\) is measurable with respect to \(\widetilde{\alpha}\), since \[\begin{align} \overline{Y} = \frac{M}{\delta} \left(\frac{\widetilde{\alpha}}{a} - 1\right). \end{align}\] Moreover, we have \(\overline{Y} \leq 1/\widetilde{\alpha}\), since \(\overline{Y} \leq M\) so that \[\begin{align} \widetilde{\alpha}\overline{Y} = a (1 + \delta \overline{Y} / M) \overline{Y} \leq a (1 + \delta) M \leq 1. \end{align}\] This means that the construction above gives \(\Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = Y\). Furthermore, \[\begin{align} P(\phi(a) = d_a) = P(A) = \Ex^P[\widetilde{\alpha}\overline{Y}] \leq a (1 + \delta) \Ex^P[\overline{Y}] \leq a. \end{align}\] Finally, because \(\widetilde{\alpha} \geq a\), we have that a rejection at level \(\widetilde{\alpha}\) is weaker than a rejection at level \(a\). As a consequence, \(\phi(\widetilde{\alpha}) \leq \phi(a)\), pointwise. This proves the first claim.

For the second claim, suppose that \(\Ex^P[Y] > C\), which is equivalent to \(\Ex^P[\overline{Y}] > 1\). We now choose \(\delta \in (0, 1)\) such that \((1 - \delta) \Ex^P[\overline{Y}] > 1\). Then choose \(a \in (0, 1)\) to be sufficiently small such that \(a M \leq 1\). We then define the data-dependent level \[\begin{align} \widetilde{\alpha} := a (1 - \delta \overline{Y} / M), \end{align}\]

This level satisfies \(\widetilde{\alpha} \in [a (1 - \delta), a]\) so that \(\widetilde{\alpha} \leq a\). Moreover, \(\widetilde{\alpha} \overline{Y} \leq a M \leq 1\), and \(\overline{Y}\) is measurable with respect to \(\widetilde{\alpha}\), since \[\begin{align} \overline{Y} = \frac{M}{\delta} \left(1 - \frac{\widetilde{\alpha}}{a}\right), \end{align}\] so that the construction above gives \(\Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = Y\). Since \(\widetilde{\alpha} \leq a\), a rejection at the data-dependent level is at least as strong as a rejection at level \(a\). Hence, \[\begin{align} \{\phi(\widetilde{\alpha}) \geq d_a\} = A. \end{align}\] As a consequence, \[\begin{align} P(\phi(\widetilde{\alpha}) \geq d_a) = P(A) = \Ex^P[\widetilde{\alpha}\overline{Y}] = a \Ex^P\left[(1 - \delta \frac{\overline{Y}}{M})\overline{Y}\right] \geq a (1 - \delta) \Ex^P[\overline{Y}] > a. \end{align}\] This proves the second claim. ◻

8.2 Proof of the theorem↩︎

Proof of Theorem 6. To start, by nesting classical validity, Theorem 5 pins the normalized loss to \(\overline{L}(d_\alpha) = 1/\alpha\), \(\overline{L}(0) = 0\).

We now prove the core of the result. Let \(Y\) be an arbitrary \([L(0), \infty]\)-valued random variable \(Y\). We will show that \[\begin{align} \label{iff:core} \rho(Y) \leq C \iff \Ex^P[Y] \leq C. \end{align}\tag{8}\]

We start by assuming \(Y\) is bounded, and later lift this using continuity from below. We split the proof into three cases: \(\Ex^P[Y] < C\), \(\Ex^P[Y] > C\) and \(\Ex^P[Y] = C\).

Suppose that \(\Ex^P[Y] < C\). By Lemma 1, we can manufacture a test \(\phi(\widetilde{\alpha})\) at some data-dependent level \(\widetilde{\alpha}\) with \(\Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}] = Y\), whose probability to reject at level \(a\) is bounded by \(a\), and \(\widetilde{\alpha} \geq a\). By nesting classical validity, \(\phi\) is therefore valid for the constant data-dependent level \(\widetilde{\alpha}' \equiv a\). Since \(\widetilde{\alpha} \geq a\), we have that \(\phi(\widetilde{\alpha}) \leq \phi(\widetilde{\alpha}')\) for \(\widetilde{\alpha}' \equiv a\). By the monotonicity property, \(\phi\) must therefore be valid at level \(\widetilde{\alpha}\), and so \(\rho(Y) \leq C\).

Next, suppose that \(\Ex^P[Y] > C\). For the sake of contradiction, suppose \(\rho(Y) \leq C\). Lemma 1 then allows us to manufacture a test \(\phi(\widetilde{\alpha})\) at some data-dependent level \(\widetilde{\alpha}\), whose probability to reject at level at least \(a\) exceeds \(a\). Since \(\rho(Y) \leq C\), this selected test would be valid, contradicting the preservation of classical validity property. As a consequence, \(\rho(Y) \leq C \implies \Ex^P[Y] \leq C\).

Finally, suppose that \(\Ex^P[Y] = C\). We approximate \(Y\) from below by \[\begin{align} Y_n := L(0) + (1 - 1/n) (Y - L(0)) \leq Y, \end{align}\] for \(n \geq 1\). Then \(Y_n \uparrow Y\) and \(\Ex^P[Y_n] < C\) for every \(n\). Hence, the \(\Ex^P[Y] < C\)-case above gives \(\rho(Y_n) \leq C\). By continuity from below, \(\rho(Y_n) \uparrow \rho(Y)\) so that \(\rho(Y) \leq C\).

Combining these three cases proves 8 for bounded \(Y\). We now use continuity from below to lift boundedness to the general case. Let \(Y\) be \([L(0), \infty]\)-valued, write \(H = (Y - L(0)) / (C - L(0))\), and for \(n \geq 1\), define \[\begin{align} Y_n' := L(0) + (C - L(0)) (H \wedge n) \leq Y. \end{align}\]

Suppose first that \(\Ex^P[Y] \leq C\). Since \(Y_n'\) is bounded and \(\Ex^P[Y_n'] \leq C\), the bounded comparison gives \(\rho(Y_n') \leq C\) for every \(n \geq 1\). Since \(Y_n' \uparrow Y\), continuity from below gives \(\rho(Y) \leq C\).

Conversely, suppose \(\rho(Y) \leq C\). For the sake of contradiction, assume that \(\Ex^P[Y] > C\). By monotone convergence, there exists an \(n'\) such that \(\Ex^P[Y_{n'}'] > C\). Since \(Y_n' \uparrow Y\), continuity from below gives \(\rho(Y_n') \uparrow \rho(Y)\). In particular, because \(\rho(Y) \leq C\), we must have \(\rho(Y_{n'}') \leq C\). But \(Y_{n'}'\) is bounded and \(\Ex^P[Y_{n'}'] > C\), contradicting the bounded comparison above. Hence, \(\Ex^P[Y] \leq C\).

This completes the proof of the core comparison 8 . We now apply this to an arbitrary test \(\phi\) and data-dependent level \(\widetilde{\alpha}\) by taking \[\begin{align} Y = \Ex^P[L(\phi(\widetilde{\alpha})) \mid \widetilde{\alpha}]. \end{align}\] By the defined notion of validity 6 and 8 , \[\begin{align} \phi \textrm{ is valid for } \widetilde{\alpha} \iff \Ex^P[Y] \leq C. \end{align}\] By the tower property, \(\Ex^P[Y] = \Ex^P[L(\phi(\widetilde{\alpha}))]\), so that validity is equivalent to \[\begin{align} \Ex^P[L(\phi(\widetilde{\alpha}))] \leq C. \end{align}\] Normalizing to \(\overline{L}\) and \(1\) yields the claim. ◻

9 Proof of Theorem 7↩︎

Proof. We start by observing some simple consequences of the definition of closure. For every data-dependent level \(\widetilde{\alpha}\), \[\begin{align} \label{ineq:proof95closure} \overline{\phi}(\widetilde{\alpha}) \leq \e_\phi. \end{align}\tag{9}\] Indeed, on the event \(\{\e_\phi\geq d_{\widetilde{\alpha}}\}\) we have \(\overline{\phi}(\widetilde{\alpha}) = d_{\widetilde{\alpha}} \leq \e_\phi\), while on its complement \(\overline{\phi}(\widetilde{\alpha}) = 0 \leq \e_\phi\). Moreover, by a squeezing-argument we have \[\begin{align} \label{eq:proof95squeeze} \e_{\overline{\phi}} \equiv \sup_{\alpha \in (0, 1)} \overline{\phi}(\alpha) = \sup_{\alpha \in (0, 1)} \phi(\alpha) \equiv \e_\phi, \end{align}\tag{10}\] as \(\phi(\alpha) \leq \overline{\phi}(\alpha) \leq \e_\phi\) for every \(\alpha\).

We now first prove an intermediate result that post-hoc validity of \(\phi\) implies validity of \(\e_\phi\). By post-hoc approximability, there exists a sequence of data-dependent levels \((\widetilde{\alpha}_n)_{n \geq 1}\) such that \(\phi(\widetilde{\alpha}_n)\) is increasing and \[\begin{align} \phi(\widetilde{\alpha}_n) \uparrow \sup_{\alpha \in (0, 1)} \phi(\alpha) = \e_\phi. \end{align}\] Since \(\phi\) is post-hoc valid, \(\phi(\widetilde{\alpha}_n)\) is valid for every \(n \geq 1\). By continuity from below, \(\e_\phi\) is valid.

We now use this to prove that \(\overline{\phi}\) is post-hoc valid. As \(\e_\phi\) is valid and \(\overline{\phi}(\widetilde{\alpha}) \leq \e_\phi\) pointwise by 9 for a data-dependent level \(\widetilde{\alpha}\), abstract monotonicity implies that \(\overline{\phi}\) is valid for \(\widetilde{\alpha}\). As this holds for every \(\widetilde{\alpha}\), \(\overline{\phi}\) is post-hoc valid.

It remains to prove the equivalence between post-hoc validity of \(\overline{\phi}\) and the validity of its associated E-value. If \(\e_{\overline{\phi}}\) is valid, then \(\overline{\phi}(\widetilde{\alpha}) \leq \e_{\overline{\phi}}\) for every \(\widetilde{\alpha}\), so that \(\overline{\phi}\) is post-hoc valid by monotonicity.

Conversely, suppose that \(\overline{\phi}\) is post-hoc valid. Since \(\phi(\widetilde{\alpha}_n) \leq \overline{\phi}(\widetilde{\alpha}_n)\), monotonicity implies that every \(\phi(\widetilde{\alpha}_n)\) is valid. As \(\phi(\widetilde{\alpha}_n) \uparrow \e_\phi\), continuity from below gives validity of \(\e_\phi\). Since \(\e_{\overline{\phi}} = \e_\phi\) by 10 , this proves the validity of \(\e_{\overline{\phi}}\). ◻

References↩︎

[1]
R. Wang and A. Ramdas, “False discovery rate control with e-values,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 84, no. 3, pp. 822–852, 2022.
[2]
E. Katsevich and A. Ramdas, “Simultaneous high-probability bounds on the false discovery proportion in structured, regression and online settings,” The Annals of Statistics, vol. 48, no. 6, pp. 3465–3487, 2020.
[3]
Z. Xu, R. Wang, and A. Ramdas, “Post-selection inference for e-value based confidence intervals,” Electronic Journal of Statistics, vol. 18, no. 1, pp. 2292–2338, 2024.
[4]
P. D. Grünwald, “Beyond neyman–pearson: E-values enable hypothesis testing with a data-driven alpha,” Proceedings of the National Academy of Sciences, vol. 121, no. 39, p. e2302098121, 2024.
[5]
N. W. Koning, “Post-hoc \(\alpha\) hypothesis testing and the post-hoc \(p\)-value,” arXiv preprint arXiv:2312.08040, 2025.
[6]
N. W. Koning, “Continuous testing: Unifying tests and e-values,” arXiv preprint arXiv:2409.05654, 2024.
[7]
B. Chugg, T. Lardy, A. Ramdas, and P. Grünwald, “On admissibility in post-hoc hypothesis testing,” International Journal of Approximate Reasoning, vol. 191, p. 109634, 2026, doi: https://doi.org/10.1016/j.ijar.2026.109634.
[8]
E. Gauthier, F. Bach, and M. I. Jordan, “E-values expand the scope of conformal prediction,” arXiv preprint arXiv:2503.13050, 2025.
[9]
N. W. Koning and S. van Meer, “Fuzzy prediction sets: Conformal prediction with e-values,” arXiv preprint arXiv:2509.13130, 2025.
[10]
M. Zhu and O. Simeone, “Beyond fixed false discovery rates: Post-hoc conformal selection with e-variables,” arXiv preprint arXiv:2604.11305, 2026.
[11]
L. Fischer and K. Sechidis, “Knockoffs for low dimensions: Changing the nominal level post-hoc to gain power while controlling the FDR,” arXiv preprint arXiv:2511.11166, 2025.
[12]
S. Koobs and N. W. Koning, “Equivalence testing with data-dependent and post-hoc equivalence margins,” arXiv preprint arXiv:2603.16213, 2026.
[13]
W. Hartog and L. Lei, “Family-wise error rate control with e-values,” arXiv preprint arXiv:2501.09015, 2025.
[14]
Z. Xu, A. Solari, L. Fischer, R. de Heide, A. Ramdas, and J. Goeman, “Bringing closure to false discovery rate control: A general principle for multiple testing,” arXiv preprint arXiv:2509.02517, 2025.
[15]
N. W. Koning, “The e-measure,” arXiv preprint arXiv:2604.20788, 2026.
[16]
B. Chugg, E. Gauthier, M. I. Jordan, A. Ramdas, and I. Waudby-Smith, “Post-hoc large-sample statistical inference,” arXiv preprint arXiv:2603.08002, 2026.

  1. n.w.koning@ese.eur.nl, Econometric Institute, Erasmus University Rotterdam.↩︎