April 21, 2026
We study economies where consumers interact independently with many monopolists. When consumer valuations over goods are correlated, correlation can distort the induced distribution of consumer surplus (information rents). We identify which shifts in the correlation structure over valuations make the induced distribution more or less fair, in the sense of second order stochastic dominance. We then investigate the role taxation can have on shaping the distribution of information rents, and show the tax authority never benefits from randomizing the allocation of goods. We characterize the set of mechanisms that are on the fairness-efficiency frontier under regularity conditions on the distribution of types. Furthermore, under these conditions all allocations on the fairness-efficiency frontier ration the good more than an unregulated monopolist. Finally, we discuss implications of our model for luxury commodity taxation.
Keywords: Excise Taxes, Information Rents, Second Order Stochastic Dominance.
Commodity taxes are often a staple of redistributive taxation. For example, luxury taxes are frequently implemented by U.S. state governments—such as Connecticut’s tax on jewelry and handbags ([1]) and Washington’s tax on luxury motor vehicles ([2])—and were utilized by the US federal government as recently as the 1990s. Moreover, commodity taxes are often used in lieu of income taxes by developing countries as their primary redistributive instrument ([3]).
The redistributive impact of these taxes is often at the center of public debate. For example, opponents of California’s excise tax on gas criticize its supposedly regressive nature, while proponents argue that the burden falls primarily on more affluent households, who already benefit disproportionately from market activity. Despite their prevalence, a formal understanding of the redistributive effects of commodity taxes remains elusive.
In this paper, we propose a redistributive theory of commodity taxes. We adopt the perspective of a regulator—–such as a state lawmaker barred from income taxation2 or a developing country with limited state capacity3—–who lacks access to income taxes but can regulate the market for goods. We ask whether our redistributive-minded regulator can ameliorate inequality using only these a-priori regressive tools. Our main contribution shows not only that the answer is yes, but we also characterize (under mild regularity conditions) the frontier way to attain redistributive targets with minimal efficiency losses.
Our model is as follows. There is a single consumer with unit demand for \(n\) distinct goods, each supplied by a different monopolist. The consumer’s valuation is represented by a multidimensional parameter \(\theta \in \Theta \subset \mathbb{R}^n\). A regulator implements quantity-dependent taxes or subsidies to influence the distribution of consumer surplus, with the goal of making the market “more fair.”
We define fairness through the lens of second order stochastic dominance (SOSD). A policy \(\sigma\) is more fair than another, \(\sigma'\), if and only if the induced distribution of consumer surplus by \(\sigma\) dominates in the SOSD sense the distribution induced by \(\sigma'\). We seek to characterize the frontier of fair and efficient policies–those where consumer surplus cannot be raised without making the market less fair.
Standard economic theory often dismisses excise taxes as an inefficient redistributive tool for two primary reasons. First is the welfare theorems. These suggest redistribution and allocation should be separated, so that only lump sum taxes should be considered because commodity taxes unnecessarily distort production when chasing allocative redistribution. Second is the seminal result by [4], which states that, given an optimal nonlinear income tax, differentiated commodity taxes are redundant as a redistributive instrument when preferences are separable between consumption and labor.
Our model departs from standard economic theory in two fundamental respects. First, we assume that a consumer’s place in the distribution of surplus is both private information and endogenous to the tax policy, rendering direct lump-sum targeting impossible4. Second, we abstract away from income to isolate heterogeneity arising from consumer tastes and purchasing patterns. This shift departs from the [4] framework, where inequality is traditionally driven by heterogeneous ability and addressed via income taxation. Collectively, these deviations from canonical models allow us to model the specific constraints faced by redistribution-minded regulators whose primary available policy lever is commodity taxation.
Our departures afford several advantages. First, they allow us to rationalize and evaluate the redistributive impact of commodity taxes as they are implemented in practice. Second, they allow us to to analyze the impact that correlated values for goods have on the distribution of consumer surplus, and answer the question of when correlated values for goods a-priori lead to more or less fair distributions of consumer surplus. Finally, by explicitly excluding income taxation and direct monopoly regulation from the model, we provide rigorous second-best guidance to policymakers who lack those instruments but maintain redistributive objectives5. To the best of our knowledge, our model is the first to be able to provide insights in this vein.
Our key results come in two flavors. First, we fix an excise tax structure and ask what changes in the correlation of values between goods \((F \in \Delta(\Theta))\) lead to more or less fair distributions of consumer surplus. Our first main result (Theorem 1) shows that one correlation structure \(F\) is more fair than another \(G\) if and only if \(G\) dominates \(F\) in the supermodular order (supposing they have the same marginal distributions). We interpret this result as stating that positive affiliation between valuations for goods—a high value for good \(i\) being predictive of high valuations for good \(j\)—engenders inequity in the distribution of consumer surplus. To the best of our knowledge, ours is the first result that points to correlation between valuations for goods as being a determinant of inequality in the market.
Our second set of results considers the design of excise tax structures and the impact of excise taxation on market behavior and the resulting distribution of consumer information rents. To shut down the impact of correlation on information rents, we consider the case where \(F\) is independent across markets. Our second pair of main results (Theorem 2 and Theorem 3) shows only threshold mechanisms—excise tax and subsidy policies which induce posted prices in each market—can be on the fairness-efficiency frontier, and furthermore (under some mild regularity conditions on the distribution of valuations \(F\)) characterize exactly which mechanisms are on the frontier. Theorems 2 and 3 have the striking consequence that every frontier mechanism is deterministic: It is never strictly optimal to randomize allocations of the good. Furthermore, it is sufficient for a regulatory agency to consider the fairness-efficiency tradeoff in each market separately. Such an optimization program market by market will yield a frontier mechanism for the joint distribution given that values are independent. Moreover, Theorem 3 gives an easy way to compute the thresholds as a function of the marginal response of the monopolist to the tax policy. We leverage this simple characterization in an example with two goods in order to give a complete computational characterization of the set of all excise taxes that induce a distribution of consumer surplus on the frontier.
Finally, we show that any mechanism on the fairness-efficiency frontier must induce strictly more rationing than the monopolist’s unrestricted choice, i.e. only excise taxes (and not subsidies) can be on the fairness-efficiency frontier. In particular, it is always more efficient to attain a redistributive goal by taxing the good and giving the proceeds to those who did not buy than to subsidize the good to encourage more people to purchase. We interpret this result as identifying conditions under which commodity taxes are exactly the right mechanism to attain our regulator’s redistributive goals.
The rest of the paper is organized as follows. The remainder of this section briefly discusses the related literature. Section 2 introduces the formal model. Section 3 analyzes correlation structures and ranks them by fairness. Section 4 characterizes fair excise tax policies. Section 5 works through a luxury taxation example that applies our results. Section 6 concludes. All omitted proofs, as well as supplementary technical material, can be found in the appendix.
We relate to a few strands of literature. First, our exercise is inspired by the literature on redistributive market design ([5], [6], and [7]), which studies the impact that non-market interventions such as quotas, a public option ([8]), or ordeals ([9]) have on allocative and redistributive efficiency. While we share a common motivation—seeking to understand the redistributive impact of market interventions on the distribution of consumer surplus—our tools and messages are quite distinct. In particular, in contrast to these papers, which often assume the designer’s objective includes more variables than they have tools to screen on, we show that there is a role for government intervention even when consumers’ valuation for money is not their private information. Moreover, their policy conclusions suggest a range of possible interventions are constrained optimal (including subsidization and rationing); in contrast, we show commodity taxes can be used to attain our regulator’s redistributive goals. Our redistributive criteria—which looks at the distribution of information rents in a monopoly screening problem—is similar in spirit to the criteria of [10], though our goals and methods are quite different; for example, they study an information design problem. Closer to our paper, [11] compare screening devices by how they distribute information rents when welfare weights are unobservable; we similarly study the distribution of information rents, but through tax-and-subsidy policies without unobservable welfare weights, in many-good screening problems.
Further back, there is a large and developed literature studying the incidence of excise taxes and their impact on consumer surplus in public finances (see [12] for a survey). We will not discuss our relation to every paper in this field, but a few are particularly related. First, in the absence of screening, several papers study the impact of excise taxes on welfare in environments with a single good ([13], [14], [15], and [16]). Because our environment is separable, we are able to study the impact of excise taxes on welfare in the presence of consumer private information, and provide a full characterization of the monopolist’s optimal behavior. Second, there is work studying the impact of excise taxes in multi-good markets, including on prices and welfare ([17] and [18]). We contribute to this literature by studying taxation of economies with many goods but separate monopolies. Finally, our supra-pricing result is similar in spirit to the result that monopolists’ may oversupply quality relative to first best in two-sided markets ([19]), though we obtain it in a very different environment.
Relatedly, we comment on a smaller literature that tries to ground redistributive motives for commodity taxes under different distortions in the market. [20] showed commodity taxes can be optimal in an [4] world when tastes are positively correlated with earning ability conditional on income; further afield, [21] and [22] give a salience-based microfoundation for optimal commodity taxes, and explicitly allow the designer to have a redistributive motive. We depart from this literature by (1) removing income taxes and (2) explicitly modeling a monopolistic seller. We contribute to this literature by showing positive affiliation between goods amplifies the need for redistribution, and by computing the set of all tax policies on the fairness-efficiency frontier when good are independently distributed and the designer has a redistributive motive for consumer surplus.
Finally, we relate to the literature on stochastic orders. Our first main result gives an equivalence between three stochastic orders on the distribution of information rents, drawing heavily on past work on the supermodular order ([23], [24], and [25] Chapter 10).
We study economies with three types of players: a single consumer6, a single regulator, and \(n\) many firms. Each firm \(i \in [n]\) supplies good \(i\) at zero marginal cost to the consumer, who collects payoff \(\theta_iv_i(q_i)\) from consuming \(q_i\) of each good, and has private value \(\theta_i \in \Theta_i\) for the good. Suppose \(\Theta_i \subset \mathbb{R}\) is a compact interval. Denote by \(\Theta = \prod_{i = 1}^n \Theta_i\) the space of all possible valuations of goods, and suppose the consumer’s utility is additively separable across goods (hence, if they purchase a bundle \(\{q_i(\theta_i)\}_{i = 1}^n\), their total consumption utility is given by \(\sum_{i \in I} \theta_i v_i(q_i(\theta_i))\)). Moreover suppose \(\theta \sim F \in \Delta(\Theta)\) is drawn from a commonly known fully supported joint distribution admitting an everywhere positive continuous density \(f\), and let \(F_i\) denote the marginal distribution of \(F\) on \(\Theta_i\).
Next, the game. Fix a Borel message space \(M\). A firm-\(i\) game is a function \(\Gamma_i: M \to [0, 1] \times \mathbb{R}_+\), specifying an allocation quantity and a transfer for each message sent to firm \(i\); let \(G_i\) be the set of all firm-\(i\) games and \(G = \prod_i G_i\) be the set of all games for all firms. When there is no loss of ambiguity, we represent a firm-\(i\) game as \((q_i, t_i) = \Gamma_i\), representing the quantity and transfer rules, respectively.
The timing (and strategies) of the game are as follows.
The government chooses a (separable) tax policy \(\sigma = \{\sigma_i: [0,1] \to \mathbb{R}\}\) mapping firm quantities to subsidies or taxes given to the consumer.
Given \(\sigma\), each firm \(i\) simultaneously chooses \(\Gamma_i \in G_i\).
Given \(\{\Gamma_i\}_{i = 1}^n\) and \(\sigma\), consumers choose a messaging strategy \(m_i: \Theta \to \Delta(M)\) for each \(i \in [n]\).
Given \(m\), and Nature’s draw of \(\theta\), \(\Gamma_i(m_i)\) is realized in each dimension, with consumers getting payoff7 \[U(\theta) = \sum_{i = 1}^n \theta_i v_i(q_i(m_i(\theta))) - t_i(m_i(\theta)) + \sigma_i(q_i(m_i(\theta))).\] while firm \(i\) gets a payoff of \[t_i(m_i(\theta)).\]
A mechanism is a tuple \((\sigma, \Gamma, m)\). Two mechanisms are equivalent if they lead to the same payoffs for all agents. A mechanism is direct if \(M = \Theta\), and incentive compatible if \(m_i(\theta) = \theta\) for each type \(\theta\). More formally, let the interim utility of type \(\theta\) reporting \(m^i\) to firm \(i\) to be \[\mathcal{U}(\{m^i\}; \theta) = \sum_{i = 1}^n \theta_i v_i\left( q_i(m^i) \right) - t_i(m^i) + \sigma_i(q_i(m^i)).\] Incentive compatibility requires that (1) the mechanism is direct (so that \(M = \Theta\)), and (2) \[\{\theta\}_{i = 1}^n \in \mathop{\mathrm{\arg\!\max}}_{\{\tilde{\theta}^i\} \in \Theta^n} \mathcal{U}(\{\tilde{\theta}^i\}; \theta)\] for all \(\theta\).
A profile of firm-\(i\) games \(\{\Gamma_i\}_{i \in [n]}\) is optimal if it is played on path in a subgame perfect equilibrium of the induced game (that is, \(\Gamma_i\) best responds to \(\sigma\) and \(\Gamma_{-i}\) for fixed \(\sigma\)).
Throughout the exposition, we will restrict to separable mechanisms, those where the government’s tax policy for firm \(i\), \(\sigma_i(\cdot)\), depends only on \(q_i(m_i(\theta))\). We interpret this assumption as a restriction that commodity taxes (or subsidies) are levied good-by-good, and not based on the entire bundle, which we think is reasonable given our motivation8. Given a separable mechanism, it will sometimes be helpful to write \(\sigma(q(\theta)) = \sum_i \sigma_i(q_i(\theta))\) as the total tax/subsidy levied per type profile \(\theta\) given \((q, t)\). With the assumption of separable mechanisms in hand, we can state our first result—a version of the revelation principle, adapted to our setting with multiple firms.
Proposition 1 (Revelation Principle). Let \((\sigma, \Gamma, m)\) be some separable mechanism. There exists an equivalent incentive compatible direct mechanism \((\sigma, \{q_i, t_i\}, \text{Id})\) where the message space is \(\Theta\). If \(\Gamma, m\) were optimal then so is \((\{q_i, t_i\}, \text{Id})\).
The first part of Proposition 1—that it is without loss of generality to restrict to truthful incentive compatible messaging strategies for the consumers—follows from standard arguments. The novel part of the argument (and the one that requires separability) is that such direct mechanisms are optimal given \(\sigma\) so long as \((\Gamma, m)\) was optimal given \(\sigma\) in the original mechanism. The core intuition here is that so long as the mechanism is separable, strategic interactions between firms cannot happen off the path of play, which allows the direct mechanism to inherit optimality. [Appendix32II] gives an example of how optimality might not be inherited in the absence of separability.
We can, in fact, say a little bit more. In particular, note that while a-priori each consumer is reporting their entire multidimensional type \(\theta\) to firm \(i\), firm \(i\) cares only about dimension \(\theta_i\), the consumer’s particular valuation for the good that they are providing. This is because valuations are additively separable, different firms control each good, and the tax policy is additively separable—hence, it is without loss of generality for each consumer to report only \(\theta_i\) to firm \(i\). Formally,
Proposition 2. Suppose \((q, t, \sigma)\) is incentive compatible and separable. Then for all \(i\), \(q_i(\theta_i, \theta_{-i})\) is constant in \(\theta_{-i}\) for fixed \(\theta_i\).
Note that Proposition 2 implies transfers \(t_i(\theta_i)\) are independent of \(\theta_{-i}\) as well, as otherwise the partial information rent (which must be constant in \(\theta_{-i}\)), would have variation in a payoff-irrelevant part of the report, which is impossible. From here on out, we will adopt this restriction and write \((q_i, t_i)\) as functions of \(\theta_i\) only. Given this formulation, we define the information rent given to a consumer of type \(\theta\) as their interim utility from truthful reporting; \[I(\theta; q, t, \sigma) = \sum_{i = 1}^n \Big[ \theta_i v_i(q_i(\theta_i)) - t_i(\theta_i) + \sigma_i(q_i(\theta_i)) \Big].\] If \(I(\theta; q, t, \sigma)\) is a function induced by an incentive compatible, direct \((q, t, \sigma)\) where \((q, t)\) is optimal given \(\sigma\)9, we say that \(I\) is implementable under \(\sigma\). When there is no loss of ambiguity we will omit the terms after the semicolon and write \(I(\theta)\). Given such an information rent \(I: \Theta \to \mathbb{R}\), define the distribution of information rents induced by \(F\) to be \[(F \circ I)(x) = \mathbb{P}_F(I(\theta) \leq x),\] noting this is a distribution on \(\mathbb{R}\).
Our first main results compare the effect that variation over correlation structures has on the induced distribution of information rents. Throughout this section, fix a tax policy \(\sigma\) and suppose \((q, t)\) is the resulting induced mechanism given \(\sigma\) and \(F\), i.e. the incentive compatible direct mechanism \((q, t)\) that maximizes firm revenue at \((\sigma, F)\).
Note by Proposition 2, firm \(i\) behavior depends only on \(F_i\) and \(\sigma_i\), so in particular \((q, t)\) will remain optimal for any other distribution \(G\) whose marginal distributions are the same as \(F\), i.e. \(F_i = G_i\) for all \(i\). This observation allows us to characterize when one distribution \(F\) leads to a distribution of information rents which is more or less “fair” (in a manner to be made precise) than another distribution \(G\), even as firm pricing remains unchanged.
Our main result of this section is that correlation between valuations for goods is exactly the right predictor that shifts distributions of surplus to be more or less fair. We need two definitions to state the result. First, suppose that \(\Theta_i = [a_i, b_i]\) for each \(i \in [n]\). Say that \(F \in \Delta(\Theta)\) dominates \(G \in \Delta(\Theta)\) in the supermodular (stochastic) order, \(F \succsim_{SM} G\) if
\[\int h dF \geq \int h dG \text{ for all } h \text{ supermodular in \Theta}.\]
This defines a partial order over lotteries, and following [23], is a quantitative measure of correlation. Second, we need a definition which formalizes our notion of fairness for distributions.
Definition 1. A distribution \(F\) is more fair than a distribution \(G\) given \(\sigma\) if
\(F\) and \(G\) have the same marginal distributions.
For all \(I\) implementable under \(\sigma\), \((F \circ I) \succsim_{SOSD} (G \circ I)\).
We discuss here some motivation for our use of second order stochastic dominance as the main notion of inequality aversion. Note that given Definition 1 and using the fact that \(F\) and \(G\) have the same marginal distributions, \(F \circ I\) and \(G \circ I\) will have the same average consumer surplus. Consequently, Definition 1 is equivalent to dominance in the Lorenz order, which is well-known to be a canonical measurement of inequality when looking at distributions of welfare (see [26] and consequent citations for a detailed discussion).
We have the following core result.
Theorem 1. Fix policy \(\sigma\). The following are equivalent.
\(F \succsim_{SM} G\).
\(G\) is more fair than \(F\) given \(\sigma\).
When \(n = 2\), \(F(x, y) \geq G(x, y)\) for all \(x, y\).
Theorem 1 is the first main takeaway of the paper, and has two implications. First, it says whenever \(F\) is more correlated than \(G\), then the corresponding distribution of information rents under \(F\) will be less fair than \(G\). We thus have a sense in which, even though the problem is separable in each dimension, correlation can influence the agglomeration of information rents to consumers of different types, even under the same optimal pricing structure.10
Second, it gives a general test in the two-dimensional case of when a distribution is more or less fair than the independent coupling. Directly reading off the third equivalence gives the following, where \(F_1 \otimes F_2\) is the independent coupling of the marginals \(F_1\) and \(F_2\).
Corollary 1. Let \(n = 2\). \(F_1 \otimes F_2\) is more fair than \(F\) given \(\sigma\) if and only if for all \(x, y\), \(F(x, y) \geq F_1(x) F_2(y)\).
When there are two goods, our characterization via the supermodular order allows us to show that there exist maximally and minimally fair correlational structures. Here, “maximality” means that among all correlational structures which attain some level of total surplus, there are no other ways to correlate types to induce fairer distributions of consumer surplus. For fixed marginal distributions, maximality should be seen as the analogue to unimprovability for information rents (see Definition 7).
Definition 2. \(F \in \Delta(\Theta)\) is \(r\)-maximally fair if
There exists an implementable information rent \(I\) such that \(\mathbb{E}_F[I] \geq r\).
Among all \(G\) such that \(F_i = G_i\) for all \(i\), and all \(I\) such that \(\mathbb{E}_F[I] \geq r\), \((F \circ I) \succsim_{SOSD} (G \circ I)\).
Corollary 2. Suppose \(n = 2\) and fix marginal distributions \(\bar F_1, \bar F_2\). For any attainable \(r \geq 0\), the antitone coupling is \(r\)-maximally fair.
The key observation is that, because the test functions are submodular, the maximal-fairness problem can be rewritten as a family of optimal transport problems indexed by concave utility functions (using the equivalence between SOSD and the concave order). The antitone coupling simultaneously maximizes all of these transport problems, by the classical extremal property of antitone couplings for submodular objectives. Combined with Theorem 1, this implies that the antitone coupling is maximally fair. By symmetry, the monotone coupling is \(r\)-minimally fair for every \(r\).
Our second set of results fixes a correlation structure \(F\) and asks instead when and how tax policies affect the distribution of consumer surplus. Throughout, suppose \(F\) is the independent coupling (so that \(F_i | \theta_{-j}\) is constant in \(\theta_{-j}\)). Moreover, suppose \(v_i(q_i) = q_i\) for each \(i\); that is, consumers have constant marginal value for quality. Next, by a normalization, suppose \(\Theta_i = [0, 1]\) for each \(i\). Finally, suppose that \(F\) is such that each \(F_i\) is Myersonian regular, i.e. \(\theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)}\) is nondecreasing.
This section proceeds as follows. First, we characterize the set of all allocations a regulator can implement subject to firm optimality. Second, we leverage this characterization and extreme point techniques to show only threshold mechanisms are on the frontier (Theorem 2). Third, we impose a regularity condition on \(F\) to characterize which thresholds are on the frontier (Theorem 3). Finally, we use our regularity condition to show that luxury taxes are optimal (Proposition 5).
How do firm’s price optimally given policy \(\sigma\)? Fix one such \(\sigma\). The envelope theorem implies for each dimension \(i\) that \[\int_{0}^{\theta_i} q_i(s) ds = \theta_i q_i(\theta_i) - t_i(\theta_i) + \sigma_i(q_i(\theta_i)) - \left( 0 q_i(0) - t_i(0) + \sigma_i(q_i(0)) \right).\] This formulation implies that the information rent generated to the lowest type is given by \(\sigma_i(q_i(0))\); that is, the subsidy at the lowest type modifies the individual rationality constraint. Standard arguments imply that at an optimal mechanism, the individual rationality constraint must bind for the lowest type. Because not participating in the good \(i\) mechanism gives the agent an allocation \(q_i = 0\) and transfer \(t_i = 0\), this implies \[0 q_i(0) - t_i(0) + \sigma_i(q_i(0)) = \sigma_i(0).\] Rearranging the expression above yields an expression for revenue, \(\int_{\Theta_i} t_i(s) dF(s)\), given a mechanism \(q\) as \[\int_{\Theta_i} \left(\theta_i q_i(\theta_i) - \int_{0}^{\theta_i} q_i(s) ds + \sigma_i(q_i(\theta_i)) - \sigma_i(0) \right) dF_i(\theta_i)\]
Integrating by parts and rearranging thus gives that the firms’ revenue for some fixed \(q\) is exactly \[R_i(q_i; \sigma) = \underbrace{\int_{\Theta_i} q_i(\theta_i) \left(\theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)} \right) f_i(\theta_i) d\theta_i}_{\text{Myersonian Revenue}} + \underbrace{\int_{\Theta_i} (\sigma_i(q_i(\theta_i)) - \sigma_i(0)) f_i(\theta_i) d\theta_i}_{\text{Policy Distortion}}.\] Standard arguments imply that \(q_i\) is implementable by some transfer rule \(t\) if and only if it is increasing; we thus have that the firm’s optimal pricing problem solves \[R_i(\sigma) = \max_{q_i \text{ increasing}} R_i(q_i; \sigma).\]
Classical Myersonian reasoning tells us that all and only increasing allocations \(q_i\) are implementable by some transfer rule for a fixed tax. The main result of this subsection is a stronger, complementary result: all (and only) increasing allocations are part of an optimal choice for the firm for some tax policy. In particular, for any increasing \(q_i\) we construct some tax policy \(\sigma\) such that \(q_i\) maximizes firm \(i\)’s revenue. Formally,
Definition 3. An allocation \(q_i\) is optimal given \(\sigma\) if \[q_i \in \mathop{\mathrm{\arg\!\max}}_{\tilde{q}_i \text{ implementable}} R_i(\tilde{q}_i; \sigma).\]
Proposition 3. \(q_i\) is nondecreasing if and only if there exists \(\sigma\) where \(q_i\) is optimal given \(\sigma\).
The intuition behind this argument is as follows. Clearly only increasing allocations are optimal because they are the only ones that can be incentive compatible. To construct some \(\sigma\) that renders \(q_i\) optimal, first restrict to \(q_i\) which are strictly increasing (and thus invertible). On every interior value of \(q_i\), the first order condition is sufficient. \(q_i\) is then optimally pinned down by a differential equation \(\sigma'(Q)\), written in terms of the quantity and the inverse hazard rate of the distribution. Explicitly solving for this equation gives \(\sigma(q)\) as the following integral equation \[\sigma_i(q_i(\theta_i)) = \sigma_i(0) + \int_0^{q_i(\theta_i)} \frac{1-F_i(q_i^{-1}(s))}{f_i(q_i^{-1}(s))} - q_i^{-1}(s) ds.\] Setting \(\sigma(Q)\) as above then implies \(q_i\) is optimal for each type \(\theta_i\), as desired. A compactness argument extends this construction to all \(q\), even when it is only weakly increasing. The details are spelled out in the appendix.
Proposition 3 implies that there are no a-priori restrictions that can be made on the set of allocation rules that can be induced by a tax policy; any allocation rule which is incentive compatible for consumers can be made optimal for some specification of taxes. Thus, there might a-priori exist points on the fairness-efficiency frontier which are quite complicated, and which require complex taxation schemes to attain. We turn to this problem next.
We start by introducing a particularly simple class of mechanisms; those which sell the the good to all (and only) consumers above a certain cutoff type, \(k_i\).
Definition 4. \((q, t)\) is a threshold mechanism if for all \(i\), \(q_i(\theta_i) = \mathbf{1}_{[k_i, 1]}(\theta_i)\) for some \(k_i \in [0, 1]\).
Definition 5. An excise tax is a policy of the form \(\sigma_i(q_i) = C_i - \tau_i q_i\) for constants \(C_i, \tau_i \in \mathbb{R}\).
Graphically, an excise tax subsidizes some amount \(C_i\) for not buying \(q_i = 0\), and taxes some amount (\(C_i - \tau_i\)) for buying \(q_i = 1\). Excise taxes are useful because they exactly characterize the set of all threshold mechanisms.
Observation 1. \((q, t)\) is a threshold mechanism if and only if \((q, t)\) is optimal given an excise tax \(\sigma\).
Observation 1 implies that, whenever the entire fairness-efficiency frontier is traced out by threshold mechanisms, they can be implemented by exactly the set of excise taxes. Thus economies where this is the case are relatively easy to analyze. As we show in Theorem 2, the property that the frontier is traced out by threshold mechanisms is quite natural. Towards stating this result, though, we need a formal definition of the fairness-efficiency frontier.
Definition 6. An information rent \(I: \Theta \to \mathbb{R}\) is feasible if it is induced by an \(F\)-budget balanced (\(\mathbb{E}_F[\sigma(q(\theta))] \leq 0\)) optimal mechanism \((q, t)\) given some \(\sigma\).
Feasible information rents are those that can “naturally” be induced by some tax policy and optimal pricing by the firm, supposing that the government must be budget-balanced—that is, they cannot artificially add surplus into the economy11. A bookkeeping argument implies that requiring budget balance overall is equivalent to requiring budget balance in each market.
Observation 2. Suppose \(\mathbb{E}_F[\sigma(q(\theta))] \leq 0\). Then there exists \(\sigma'\) such that \(\mathbb{E}_{F_i}[\sigma_i'(q_i(\theta_i))] \leq 0\) for all \(i\) and \(\sigma(q(\theta)) = \sum_i \sigma_i'(q_i(\theta_i))\).
We have the following definition of the fairness-efficiency frontier for information rents.
Definition 7. Let \(I, I'\) be feasible.
\(I\) is more efficient than \(I'\) if \(\mathbb{E}_F[I] \geq \mathbb{E}_F[I']\).
\(I\) is more fair than \(I'\) if \((F \circ I) \succsim_{SOSD} (F \circ I')\).
\(I\) is \(F\)-unimprovable if there does not exist a feasible \(I'\) which is more fair than \(I\) and more efficient than \(I\), with at least one strict.12
Unimprovability trades off on two dimensions: efficiency, which asks about the total surplus in the economy, and fairness, which tracks how that surplus is distributed among people with different types. Together, \(F\)-unimprovable information rents are exactly those which trace out the efficiency-fairness frontier for a given correlational structure \(F\).
In principle, the set of \(F\)-unimprovable information rents could be induced by potentially complicated mechanisms; Proposition 3 implies that any information rent which is induced by an arbitrary increasing \(q\) can be implemented by some tax policy. Moreover, the second order stochastic dominance order is not particularly complete; a natural conjecture then is that the frontier of \(F\)-unimprovable allocations is large. Theorem 2, however, shows that this is not the case. In particular, we show, perhaps surprisingly, that the frontier consists of only threshold mechanisms. Formally,
Theorem 2. \(I\) is \(F\)-unimprovable only if \(I\) is induced by a threshold mechanism.
Theorem 2 is the second main takeaway of the paper. It implies that excise taxes or subsidies are enough to trace out the fairness-efficiency frontier, even though the tax authority could in principle implement a complicated tax schedule to induce any increasing allocation and consequently potentially rich distributions of information rents(Proposition 3). Consequently, simple luxury taxes or subsidies—those that impose a per-unit tax/subsidy on each unit of the item sold—outperform any other way to redistribute surplus via commodity taxation. This includes possibly randomizing the allocation of the good through some type of quota, and also taxing goods in a nonlinear way.
We conclude this subsection with an outline of the proof of Theorem 2.
Proof. The proof proceeds in three steps.
First, we show it is without loss of generality to show threshold mechanisms are on the frontier in each dimension. To do this, we show that dominance coordinate-by-coordinate is enough to imply dominance of the entire multidimensional policy. The following lemma is instrumental.
Lemma 1. Let \(H = (F, G)\) be the independent coupling of \(F\) and \(G\). Fix information rents \(I_F, I_F': \Theta_1 \to \mathbb{R}\) and \(I_G, I_G': \Theta_2 \to \mathbb{R}\). If \(F \circ I_F \succsim_{SOSD} F \circ I_F'\) and \(G \circ I_G \succsim_{SOSD} G \circ I_G'\), then \(H \circ (I_F + I_G) \succsim H \circ (I_F' + I_G')\).
This lemma implies that if we can show threshold mechanisms are on the frontier in each dimension, then only threshold mechanisms can be on the frontier in general. To see this, fix \(F \in \Delta(\Theta) = (F_1, F_2, F_3, \dots F_n)\) coupled independently. Inductively applying Lemma 1 shows that if threshold information rents \(I_{F_i'}\) satisfy \(F_i \circ I_{F_i'} \succsim F_i \circ I_{F_i}\) for all \(i\), this must also be the case on the entire joint distribution: \(F \circ I_{F'} \succsim F \circ I_{F}\).
The second step requires some notation. Throughout the remainder of this proof, given Lemma 1, we suppose \(n = 1\) (and hence drop subscripts on \(i\)). First, let \[W(s) = (1 - F(s))\psi'(s) - f(s) \psi(s)\] be the monopolist’s marginal response to government spending (see Section 4.3 for discussion). Define \[R_q(u) = I(F^{-1}(u); q) = \int_0^{F^{-1}(u)}q(s) ds - \int_\Theta W(s) q(s) ds\] to be the information rents an individual of percentile \(u\) receives. Define \(H_q(p) = \int_0^p R_q(u) du\) be cumulative information rents up until the \(p\)-th percentile. Let \(G_q(x) = F(I^{-1}(x; q))\) be the distribution of information rents under allocation rule \(q\) where \(I^{-1}\) is the generalized inverse. When there is no loss of ambiguity, we will drop subscripts for all statements that hold for all \(q\). Finally, note \[G(R(u)) = F(I^{-1}(I(F^{-1}(u)))) = u\] so \(G = R^{-1}\) and \(G, R\) are inverse mappings between distributions over the space of information rents and the percentiles of information rents. Finally, define \(\Psi(k; p) = H_{q_k}(p)\) to be the cumulative information rents to percentile \(p\) for a threshold mechanism at threshold \(k\).
With this notation, we can state the second lemma, which characterizes second order stochastic dominance in terms of cumulative information rents.
Lemma 2. \(F \circ I(\cdot; q) \succsim_{SOSD} F \circ I(\cdot; q')\) if and only if \(H_q(p) \geq H_{q'}(p)\) for all \(p \in [0, 1]\).
We note this result is known,13 and is recorded (albeit with significantly distinct notation) in [25], Theorem 4.A.3 (albeit with an incomplete proof). For completeness of the exposition, we give a formal proof in the appendix. The key idea is to rewrite the SOSD order as an inequality based on the convex conjugates of the cumulative information rents, which follow by an integration by parts argument.
We can now proceed to the final step. Let \(\tilde{q}(k)\) be the distribution over thresholds induced by an allocation rule \(q\). Note the mapping \(q \to \tilde{q}\) is affine. Moreover note \(\tilde{q}\) is indeed a distribution over threshold mechanisms, since all (and only) increasing allocations can possibly be implemented by some subsidy policy, so we can decompose \(\tilde{q}\) into a mixture of its extreme points. We can thus write cumulative information rents up until percentile \(p\) as \[H_q(p) = \int_\Theta H_{q_k} \tilde{q}(k) dk = \int_\Theta \Psi(k; p) \tilde{q}(k) dk.\] By Lemma 2, we know that \(q\) is undominated if and only if \(H_q(p)\) is pointwise undominated by any other \(H_{q'}(p)\). This is equivalent to requiring that \[q \in \mathop{\mathrm{\arg\!\max}}_{q' \text{ increasing}} \int_0^1 \Lambda(p)H_{q'}(p)dp\] for some Pareto weights \(\Lambda(p): [0, 1] \to \mathbb{R}_+\). However, note that \[\int_0^1 \Lambda(p) H_q(p) dp = \int_0^1 \int_\Theta \Lambda(p) \Psi(k; p) \tilde{q}(k) dk dp\] is linear in \(q\). Bauer’s theorem of the maximum then implies that one solution to the maximization problem will be an extreme point of the set of increasing allocation rules. But these are precisely the set of threshold mechanisms (see [28], Chapter 2). ◻
Theorem 2 says that the fairness-efficiency frontier is exactly given by threshold mechanisms, but leaves open the question of which thresholds are exactly on the frontier. The goal of this section is to answer this question. For a sharp result characterizing the set of \(F\)-unimprovable information rents, we will need the following regularity condition.
Definition 8. \(F_i \in \Delta([0, 1])\) is strongly regular if \(f_i\) is nondecreasing and log-concave.
Strong regularity is thus named because it is a strengthening of Myersonian regularity, which is implied by the assumption \(f_i\) is log-concave. Hence it is stronger by exactly requiring that the pdf of each \(F_i\), \(f_i\) is increasing.
Strong regularity is a property of only the marginal distributions \(\{F_i\}\). It is satisfied by some of the standard distributions used in auction theory, including the uniform and power distributions (\(\beta(\alpha, 1)\) for some \(\alpha\)), but is not satisfied for (truncated) exponential distributions or other concave distributions which are sometimes convenient to assume.
What does strong regularity buy us in our model? Consider the “total cost” levied in some market \(\sigma_i\), \(\mathbb{E}_{F_i}[\sigma_i(q_i(\theta_i))]\). By Proposition 3, one has that \[\mathbb{E}_{F_i}[\sigma_i(q_i(\theta_i))] = \int_0^1 \int_0^{q_i(\theta_i)} \frac{1-F_i(q_i^{-1}(s))}{f_i(q_i^{-1}(s))} - q_i^{-1}(s) ds d F_i(\theta_i).\] Interchanging the order of integration and integrating by parts allows us to write this as \[-\int_0^1 \underbrace{\left[(1 - F_i(\theta_i))\psi_i'(\theta_i) - f_i(\theta_i) \psi_i(\theta_i) \right]}_{W_i(\theta_i)} q_i(\theta_i) d\theta_i\] where \(\psi_i(\theta_i) = \theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)}\) is the Myersonian virtual value of type \(\theta_i\). By Observation 2, we can use individual budget balance (assuming it binds) to pin down the constant in the integral equation defining \(\sigma_i(q_i(\theta_i))\); in particular, one has that \(\sigma_i(0) = \mathbb{E}_{F_i}[\sigma_i(q_i(\theta_i))]\). We thus have that the “lump sum” paid out to the lowest type is exactly \(-\int_0^1 W_i(\theta_i) q_i(\theta_i) d\theta_i\).
What does \(W_i(\theta_i)\) represent economically? Differentiating under the integral in \(\theta_i\) gives that by spending a little bit more in expectation at \(\theta_i\), the monopolist can incentivize \(W_i(\theta_i)\) more of the good. \(W_i\) is thus the marginal response of the monopolist to government spending when it comes to allocating marginally more of the good to type \(\theta_i\).
Proposition 4. Let \(F \in \Delta([0, 1])\) be strongly regular. Then \(W\) and \(\frac{W}{1 - F}\) are strictly decreasing.
The proof follows from some straightforward computations and can be found in the appendix. These monotonicity properties, however, afford us the following theorem.
Theorem 3. Suppose \(F\) is such that each \(F_i\) is strongly regular. Then \(I\) is \(F\)-unimprovable if and only if \(I\) is induced by a threshold mechanism in each dimension with thresholds \[k_i \in \{\theta_i: W_i(\theta_i) \leq 1 - F_i(\theta_i), W_i(\theta_i) \geq 0\}.\]
Together, we see Theorem 2 and 3 as being analogous to the standard Myersonian framework, but applied to our characterization of the fairness-efficiency frontier. First, Theorem 2 uses an ordinal argument to extremize information rents, and shows that posted prices and excise taxes are undominated; Theorem 3 then uses the appropriate analogue of regularity of the virtual value function—strong regularity—to guarantee monotonicity of \(W_i\) instead of \(\psi_i\), and thus characterize exactly which allocations live on the frontier.
Theorem 3 has three implications. First, it gives a complete characterization of the set of all possible thresholds that can arise on the fairness-equity frontier in terms of the marginal cost of moving the monopolist (\(W_i(\theta_i)\)). Second, it shows that there is no need to consider the impact of other goods on the optimal subsidy/tax of good \(j\). Even if one good contributes in a particularly spectacular way to inequality in the realized distribution of information rents, the tax authority does not change their redistributive target in other goods markets in order to ameliorate this concern. Using our program, a redistribution-minded policymaker need not do anything fancy to hit a redistributive target subject to some minimal efficiency loss—they need only tune a one-dimensional parameter (the frontier threshold) in each dimension separately, and then implement an excise tax as desired to realize this threshold.
Third, the characterization requires only strong regularity of the marginal distributions, so long as the coupling between distributions is independent. Note however, that when combined with Theorem 1, these two main results give a complete characterization of the fairness efficiency frontier—varying first the correlation structure and second the feasible taxes that can be levied—for strongly regular environments.
Finally, our analysis affords us the following comparative result. Let \((\bar q, \bar t, 0)\) be the threshold mechanism induced by the unrestricted monopoly which is subject to no taxes or subsidies. When are luxury taxes optimal? Note that a luxury tax must ration the good at a rate higher than the monopolist already rations, since taxes will cause the monopolist to further decrease demand. A luxury tax, then, is optimal only when the redistributive effect of giving money to those who do not have the good outweighs the value of providing them with the good. It turns out that this is satisfied exactly when \(F\) is strongly regular. Proposition 5 is stated in the language of a single goods market.
Proposition 5 (Supra-Pricing). Suppose \(F \in \Delta([0, 1])\) is strongly regular. For any \(F\)-unimprovable \(I\) induced by \((q, t, \sigma)\), \(\bar q(\theta) = 0 \implies q(\theta) = 0\). Moreover, the monopoly allocation is not \(F\)-unimprovable.
Proposition 5 implies that any \(F\)-unimprovable allocation will ration the good more than the monopolist, i.e. sell to fewer types than an unconstrained monopolist would prefer to sell to. This has two implications. First, it gives conditions under which a luxury tax is strictly optimal from a redistributive standpoint. To the best of our knowledge, we are the first paper to give guidance to policymakers on this point. Second, it gives new settings where a monopolist strictly oversupplies the good from a welfare perspective. In particular, it must be that any frontier allocation (for any target level of surplus) rations the good at a rate strictly higher than the monopolist. Consequently, it can never be fairness-maximizing to redistribute the good to those with lower valuations for the good; it is always better to tax the good, raise revenue, and redistribute the resulting revenue to those who would not have bought the good in the first place.
The intuition behind Proposition 5 is as follows. There are two effects that arise from increasing the marginal tax rate by a little bit. First, some people are excluded from buying the good, because it has become more expensive. Second, however, the tax raises revenue, \(\mathbb{E}_F[\sigma]\), and thus increases (by the boundary condition above) the subsidy given to everyone who does not purchase the good, \(\sigma(0)\). While the first force makes the equilibrium distribution less fair (and efficient) in an SOSD sense—the surplus of all people weakly decreases with a tax, not accounting for the redistributive effect—the second force compensates those who no longer buy and can redistribute surplus from those with high valuations to those with lower valuations. Because the distribution is convex (by strong regularity), so that those with high valuations have “much higher” values than those with low valuations, the second force (raising revenue by taxing those who still buy and redistributing to those who do not buy) outweighs the first force. Thus taxation (as opposed to subsidization) is strictly optimal.
Consider the case when there are two goods (\(n = 2\)), and suppose each marginal distribution is uniform, i.e. \(F_i = \text{Unif}[0, 1]\). We use this simple case to illustrate the main results of the paper. First, suppose there is no tax or subsidy, so the monopolist optimally posts a price of \(\frac{1}{2}\) in each market. By Corollary 2, the antitone coupling is maximally fair, while the monotone coupling is minimally fair. What are their distributions of information rents? Consider the following graph, which looks at the percentile of information rents under the maximally and minimally fair couplings over information rents. We can use these graphs to compute the cumulative distributions of information rents, giving that under the antitone coupling cumulative information rents are \(\frac{1}{4}x^2\) while under the monotone coupling cumulative information rents are \(0\) if \(x \leq \frac{1}{2}\), and otherwise is \((x - \frac{1}{2})^2\). Comparing these cumulative distributions of information rents gives that \(H_{\text{antitone}}(p) \geq H_{\text{monotone}}(p)\) for all percentiles \(p\), consistent with Lemma 2.
Figure 1:
.
Second, the \(F\)-unimprovability frontier. Suppose that the joint distribution is independently coupled against the two uniform marginals. Note the uniform distribution is strongly regular, since \(f_i(\theta_i) = 1\) is nondecreasing and log-concave. Moreover, we have that \[W_i(\theta_i) = (1 - \theta_i) \cdot 2 - 1(2\theta_i - 1) = 3 - 4\theta_i\] and so we know by Theorem 3 that a threshold mechanism is on the frontier if and only if it has a threshold such that \(W_i(\theta_i) \leq 1 - \theta_i\) and \(W_i(\theta_i) \geq 0\). These two inequalities together imply that \(k \in [\frac{2}{3}, \frac{3}{4}]\) is the set of admissible thresholds which are on the frontier, by Theorem 3. Moreover, by Theorem 3, for any pair of thresholds \((k_1, k_2) \in [\frac{2}{3}, \frac{3}{4}]^2\) in each dimension, those thresholds are on the frontier.
Note that \(\frac{2}{3} > \frac{1}{2}\), i.e. the set of admissible threshold mechanisms will always ration away at least \(\frac{1}{6}\) more of mass of types. This is a consequence of the supra-pricing property. Hence the example is also consistent with Proposition 5. Note to implement these prices, we need to set a tax \(\tau_i = \phi_i(k) = 2k - 1\) for \(k \in [\frac{2}{3}, \frac{3}{4}]\), so the excise tax varies between \([\frac{1}{3}, \frac{1}{2}]\).
This paper develops a model of redistributive commodity taxation in monopolistic markets where consumer valuations across product bundles are correlated. We introduce a fairness metric based on the second-order stochastic dominance (SOSD) of information rents. Our analysis demonstrates a direct link between how correlated valuations for goods are and market inequality: As valuations for goods becomes more correlated (in the supermodular order), the resulting distribution of consumer surplus becomes less fair.
When taxes are separable in each dimension, we use (suitably adapted) Myersonian arguments to analyze the set of all implementable and optimal allocations. We identify specific conditions under which frontier-optimal policies take the form of excise taxes—either taxing or subsidizing units to induce specific posted prices. Furthermore, we establish the boundaries of the fairness-equity frontier, identifying when luxury taxes are strictly optimal and subsidies are excluded from the optimal policy mix.
Our model has several limitations and opens up promising directions for future research. First is the restriction to separable mechanisms, which simplifies the space of tax-and-subsidy policies afforded to the government. Appendix II shows that this assumption cannot be dropped without losing the revelation principle, via a counterexample; thus, understanding more complex tax-and-subsidy policies remains an open question that would likely require new tools to analyze. Second is the behavior of threshold mechanisms outside of the strongly regular case. While we conjecture we can relax this assumption to the case that \(W(\theta)\) is nondecreasing in each dimension whenever it is nonnegative, we do not know what the weakest possible condition to put on the distributions \(F\) are to ensure that we can pin down exactly which threshold mechanisms are on the frontier. Third, understanding when subsidies of goods are optimal is an interesting direction of future work. Finally, understanding how luxury taxation interactions with income taxation seems like a particularly promising area for future research. While for this project we abstracted away from income taxation, focusing on policymakers who only have access to commodity taxes (such as Washington state), a full, more general analysis would likely endow the policymaker with both tools, and we see this as a fruitful area for future research.
Proof. Fix \((\Gamma, m)\) and let \(q_i(\theta), t_i(\theta)\) be the associated direct mechanism. Any deviation to sending the vector \(\{\theta'_i\}_{i \in [n]}\) could have been replicated in the original mechanism by sending \(\{m_i(\theta'_i)\}_{i\in [n]}\) instead of \(\{m_i(\theta)\}_{i\in [n]}\). Thus, optimal reporting of the original mechanism implies incentive compatibility of the direct mechanism. This argument is isomorphic to the standard revelation principle.
Now we wish to show that the direct mechanism defined above is optimal given fixed \(\sigma\). Suppose the original mechanism \((\Gamma, m)\) was optimal and suppose the statement was false, so that the direct mechanism \((\{q_i, t_i\}, \text{Id})\) was not optimal. Then there exists some firm \(j\), game \(\hat{\Gamma}_j\), and induced direct messaging strategy \(\hat{m}_j: \Theta \to \Theta\) which yields firm \(j\) higher profit. Then, firm \(j\) could have replicated this deviation in the original setting by choosing some bijection \(m_j': \Theta \to M\) and \(\Gamma'_j(m_j'(\theta)) = \hat{\Gamma}_j(\hat{m}_j(m^{-1}(m(\theta)))) = \hat{\Gamma}_j(\hat{m}_j(\theta))\) so \(\Gamma_j', m_j'\) induces the same allocation rule as \(\hat{\Gamma}_j, \hat{m}_j\). We thus have a contradiction so long as \(m_j'\) is induced by \(\Gamma_j'\), i.e. is the consumer’s best response to \(\Gamma_j'\).
To show this, suppose this was not the case and the consumer had a profitable deviation. Since consumer utility (from their allocation, transfers per market, and tax per market) is separable, there must also exist a profitable deviation where only the consumer’s report to firm \(j\) is changed.14 By the same revelation principle logic as before, this deviation could have been replicated under \(\hat{\Gamma}_j, \hat{m}_j\), a contradiction. ◻
Proof. We first show the information rents the consumer obtains from firm \(i\) must be constant in \(\theta_{-i}\). Suppose not. Then there exists two types \((\theta_i, \theta_{-i})\) and \((\theta_i, \theta_{-i}')\) whom obtain different information rents when they report the truth; without loss of generality, suppose \[\begin{align} \theta_i v_i(q_i(\theta_i, \theta_{-i})) - t_i(\theta_i, \theta_{-i}) + \sigma_i(q_i(\theta_i, \theta_{-i})) \\ > \theta_i v_i(q_i(\theta_i, \theta_{-i}')) - t_i(\theta_i, \theta_{-i}') + \sigma_i(q_i(\theta_i, \theta_{-i}')). \end{align}\] Here, we abuse notation and set \(q_{-i}(\theta) = (q_j(\theta))_{j \neq i}\), i.e. it is the vector of quantities that arise supposing that \(\theta\) is reported to all of the other firms. Consider the deviation for type \(\theta'\) where they report to firm \(i\) they are type \(\theta\), and keep their messages the same to all other firms. This deviation must secure type \(\theta'\) a strictly higher interim payoff, and is thus not incentive compatible under \((q, t, \sigma)\), a contradiction.
Recall incentive compatibility along all reports must imply incentive compatibility in dimension \(i\) when the consumer reports truthfully to all other firms. Define \(U_i(\theta_i, \theta_{-i})\) to be the information rent from a firm when they report truthfully to all other firms, i.e. \[U_i(\theta_i, \theta_{-i}) = \max_{\theta'} \left\{ \theta_i v_i(q_i(\theta')) - t_i(\theta') + \sigma_i(q_i(\theta'))\right\}.\] This is a convex function in \(\theta_i\) (it is the upper envelope of affine functions in \(\theta_i\)) and constant in all other coordinates \(\theta_{-i}\). In particular, this implies that \(U_i\) is absolutely continuous over \(\Theta_i\), and hence the standard envelope formula applies: for any fixed \(\theta_{-i}\),
\[\frac{\partial }{\partial \theta_i}U_i(\theta_i, \theta_{-i}) = v_i(q_i(\theta_i, \theta_{-i})).\] As \(U_i(\theta_i, \theta_{-i})\) is constant in \(\theta_{-i}\) and \(v_i\) is strictly increasing and hence injective, it must be that \(q_i\) also be constant in \(\theta_{-i}\). ◻
Proof. We show \((1) \iff (2)\) and \((1) \iff (3)\).
First, if \(F \succsim_{SM} G\) then \(F\) and \(G\) have the same marginals. Consider the test function \(h(x) = h(x_1)\) for some differentiable \(h(x_1)\); this is supermodular. But then \[\int h dF \geq \int h dG \text{ and } \int -h dF \geq \int -h dG \implies \int h dF = \int h dG\] for every differentiable test function \(h\). Moreover, because they have the same marginal distributions, \(\mathbb{E}_F[x] = \mathbb{E}_G[x]\).
Now note \(G\) is more fair than \(F\) if and only if for all concave nondecreasing functions \(h: \mathbb{R} \to \mathbb{R}\), \[\int h(I(x)) dG \geq \int h(I(x)) dF.\] \(h \circ I\) is a multivariate function of \(h\). Let \(\{h_n\} \subset \mathcal{C}^2(\mathbb{R})\) be a sequence of twice-differentiable concave functions such that \(||h_n - h||_{\bar \theta} \to 0\) as \(n \to \infty\). Recall \[I(x) = \sum_{i = 1}^n U_i(\theta) = \sum_{i = 1}^n \max_{\theta_i' \in \Theta_i} \left\{ \theta_i v_i(q_i(\theta_i')) - t_i(\theta_i') + \sigma_i(q_i(\theta_i')) \right\}\] Differentiating \(h_n\) twice gives that \[\frac{\partial^2 (h_n \circ I)}{\partial \theta_i \partial \theta_j} = \frac{\partial^2 h_n}{\partial I^2} \frac{\partial I}{\partial \theta_j} \frac{\partial I}{\partial \theta_i} + \frac{\partial^2 I}{\partial \theta_j \partial \theta_i} \frac{\partial h_n}{\partial I}.\] This derivative is nonpositive. To see why, note that \(\frac{\partial I}{\partial \theta_k} = v_k(q_k(\theta_k)) \geq 0\) by the envelope theorem, and so \(\frac{\partial^2 I}{\partial \theta_j \partial \theta_i} = 0\). Finally, since \(h_n\) is concave, \(\frac{\partial^2 h_n}{\partial I^2} < 0\). Hence each \(h_n \circ I\) is submodular. But since \(h_n \to h\) in the sup-norm, this must imply \(h \circ I\) is submodular as well for arbitrary concave nondecreasing \(h\). Hence since \(F \succsim_{SM} G\), (multiplying by \(-1\) as necessary), \[\int h(I(x)) dG \geq \int h(I(x)) dF\] and so we have \(G \circ I \succsim_{SOSD} F \circ I\).
We are then done if \(I\) is implementable under both \(F\) and \(G\). Fix the additively separable tax \(\sigma\) and suppose that \(I\) is implementable under \(F\). This implies that, for each \(F_i\), \(q\) maximizes the firm’s revenue function given \(\sigma\). Since \(F_i = G_i\) as the marginal distributions are the same, \(q\) is also implementable under \(G\). This completes the proof.
The converse. First, suppose \(G\) is not more fair than \(F\). If (1) fails, our construction shows it cannot be that \(F \succsim_{SM} G\). If (1) is satisfied by (2) fails, then there exists a concave function \(h\) such that \[\int h(I(x)) dF > \int h(I(x)) dG \implies \int -h(I(x)) dF < \int -h(I(x)) dG\] but \(-h(I(x))\) is supermodular by the argument in the proof of the forward direction. Hence it cannot be that \(F \succsim_{SM} G\).
Now suppose there are only two goods, so \(n = 2\). First, suppose that \(F \succsim_{SM} G\). Consider the supermodular functions \(h(x, y) = \mathbf{1}\{x \geq a, y \geq b\}\). Integrating against this yields \[\mathbb{P}_F((x, y) \geq (a, b)) = \bar F(a, b) \geq \bar G(a, b) = \mathbb{P}_G((x, y) \geq (a, b))\] From here, recall the inclusion-exclusion identity for bivariate random variables: \[\bar F(a, b) = F(a, b) - F_1(a) - F_2(b) + 1.\] Because \(F\) and \(G\) have the same marginals, substituting these in gives that \(F(a, b) \geq G(a, b)\). Since this is true for all \(a, b\), we have the forward direction.
Now suppose \(F(x, y) \geq G(x, y)\) for all \(x, y\). Fix some \((a, b) \in \mathbb{R}^2\) and supermodular, \(\mathcal{C}^2\) function \(h\). Note the integral representation of Taylor’s formula implies \[h(x, y) = h(a, b) + \int_a^x h_1(s, b) ds + \int_b^y h_2(a, t) dt + \int_a^x\int_b^y h_{12}(s,t) dsdt\] whenever \(h \in \mathcal{C}^2\) as well. This implies that \[\begin{align} \mathbb{E}_F[h] = \int h(x, y) dF(x, y) = \int\left(h(a, b) + \int_a^x h_1(s, b) ds + \int_b^y h_2(a, t) dt + \int_a^x\int_b^y h_{12}(s,t) dsdt \right)dF \\ = \int\left(h(a, b) + \int_a^x h_1(s, b) ds + \int_b^y h_2(a, t) dt \right)dF + \int_a^{\bar \theta} \int_b^{\bar \theta} h_{12}(s, t) \left( \int_s^{\bar \theta} \int_t^{\bar \theta} dF(x, y) \right) ds dt \end{align}\] by using Fubini’s theorem. Since \(F\) and \(G\) have the same marginals, we know moreover that \[\int \int_a^x h_1(s, b) ds dF = \int \int_a^x h_1(s, b) ds dG\] with a similar equality holding for the \(h_2(a, t)\) term. Thus, we get that \[\begin{align} \mathbb{E}_F[f] - \mathbb{E}_G[f] = & \int_a^{\bar \theta} \int_b^{\bar \theta} h_{12}(s, t) \left( \int_s^{\bar \theta} \int_t^{\bar \theta} dF(x, y) - \int_s^{\bar \theta} \int_t^{\bar \theta} dG(x, y) \right) ds dt \\ = & \int_a^{\bar \theta} \int_b^{\bar \theta} h_{12}(s,t) [\bar F(s, t) - \bar G(s, t)] ds dt = \int_a^{\bar \theta} \int_b^{\bar \theta} h_{12}(s, t) [F(s, t) - G(s, t)] \geq 0 \end{align}\] The last equality uses the fact \(\bar F - \bar G = F - G\) because \(F\) and \(G\) have the same marginals, again applying the inclusion-exclusion identity for bivariate random variables. But then the inequality follows by noting \(h_{12}(s, t) \geq 0\) by supermodularity, and \(F(s, t) \geq G(s, t)\) by assumption. A density argument extends this observation from all \(\mathcal{C}^2\) supermodular functions to all supermodular functions, and we are done. ◻
Proof. Let \(r \geq 0\) be chosen so that there is some feasible information rent where \(\mathbb{E}_F[I] \geq r\). For any \(G\) with \(G_1 = \bar F_1\), \(G_2 = \bar F_2\), we have that (1) \(I\) is also feasible under \(G\) and (2) \(\mathbb{E}_G[I] \geq r\). Thus, the problem of finding an \(r\)-maximally fair coupling simplifies to asking whether the solution the program \[\mathop{\mathrm{\arg\!\max}}_{F \in \Delta(\Theta_1 \times \Theta_2)} \int_{\theta\in\Theta} v(I(\theta_1,\theta_2)) dF(\theta) \text{ s.t. } F_1 = \bar F_1, F_2 = \bar F_2\] has a constant selection as we vary \(v\) over all non-decreasing concave functions \(v\). This is because the SOSD order is equivalent to dominance in the (increasing) concave order. Using arguments from Theorem 1 gives that the objective is submodular in \((\theta_1, \theta_2)\), and hence this is now a standard optimal transport problem for a submodular objective, and so Theorem 2.9 of [29] implies perfect anti-correlation solves the problem (that is, \(F\) is the antitone coupling). ◻
Proof. The converse. Clearly if \(q_i\) is optimal then it is implementable and thus nondecreasing.
The forward direction. Fix an allocation \(q_i\) and let \(\Theta^o = q_i^{-1}((0, 1))\). We first prove the result under the assumption \(q_i\) is strictly increasing on \(\Theta^o\).
Fix a type \(\theta_i\); to implement \(q_i(\theta_i)\), we need \[q_i(\theta_i) \in \mathop{\mathrm{\arg\!\max}}_{Q \in [0, 1]} \left( \theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)} \right) Q + (\sigma_i(Q) - \sigma_i(0))\] Note the objective is supermodular in \((\theta, Q)\) and so there is a selection from the maximum correspondence such that \(q_i(\theta_i)\) is nondecreasing regardless of \(\sigma\).
There are now two cases. First, suppose \(\theta_i \in \Theta^o\). In that case the first order condition is necessary, and hence we have that \(q_i(\theta_i)\) must solve the first order condition \[\theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)} + \sigma_i'(Q) = 0.\] We will construct \(\sigma\) to satisfy the FOC . Since \(q\) is strictly monotone on \(\Theta^o\), it is locally invertible, and hence setting \(Q = q_i(\theta_i)\), \(\theta_i = q_i^{-1}(Q)\) implies that the first order condition is solved by \[\sigma'(Q) = \frac{1 - F(q_i^{-1}(Q))}{f(q_i^{-1}(Q))} - q_i^{-1}(Q)\] where \(Q = q_i(\theta_i)\) is the target allocation. This formulation also allows us to verify the second order condition for optimality, which requires \(\sigma''(Q) \leq 0\); in particular, note \[\begin{align} \sigma''(Q) = - \left( \frac{[f_i(q_i^{-1}(Q))^2 + (1 - F_i(q_i^{-1}(Q))) f_i'(q_i^{-1}(Q))] \frac{d}{dQ} q_i^{-1}(Q)}{f_i(q_i^{-1}(Q))^2} \right) - \frac{d}{dQ} q_i^{-1}(Q) \\ = - \frac{d}{d\theta}\left(\theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)} \right) \cdot \frac{d}{dQ} q_i^{-1}(Q) \leq 0 \end{align}\] where the last inequality follows from regularity of \(F\) and the fact \(q^{-1}\) is strictly increasing on \((0, 1)\) and hence \(\frac{d}{dQ} q^{-1}(Q) > 0\). Note this formula holds for all \(Q \in (0, 1)\) chosen; in particular, this implies that one solution to the differential equation for \(\sigma'\) is \[\sigma(Q) = \int_0^Q \frac{1-F_i(q_i^{-1}(s))}{f_i(q_i^{-1}(s))} - q_i^{-1}(s) ds.\] Now suppose \(\theta_i \not\in \Theta_i^o\), so that the target allocation is either \(1\) or \(0\). Define \(\sigma_i(1)\) to be \[\int_0^1 \frac{1-F_i(q_i^{-1}(s))}{f(q_i^{-1}(s))} - q_i^{-1}(s) ds\] and let \(\sigma_i(0)\) be chosen by continuity so that \(\sigma_i(0) = \lim\limits_{Q \to 0} \sigma_i(Q)\).
Having defined \(\sigma\) as above, we now need to show that \(q\) is chosen optimally in response to \(\sigma\). Let the firm’s optimal allocation rule be \(\hat{q}(\theta).\) If \(\hat{q}(\theta) \in (0, 1)\) then by the FOC analysis above, \(\hat{q}(\theta) = q(\theta)\). By the monotone comparative statics from earlier, \(\hat{q}\) is (weakly) increasing in \(\theta\). Thus, \(\hat{q}^{-1}((0, 1))\) is connected; furthermore, \(\hat{q} = 0\) for all \(\theta\) below \(\hat{q}^{-1}((0, 1))\) and \(\hat{q} = 1\) for all \(\theta\) above \(\hat{q}^{-1}((0, 1))\). This exactly coincides with \(q\) since \(q\) is also weakly increasing.
We can now generalize the result to arbitrary \(q_i(\theta_i)\), even if it is not strictly increasing on \(\Theta_i^o\) via a standard upper hemicontinuity argument. In particular, note that \[q_i \in \mathop{\mathrm{\arg\!\max}}_{\tilde{q}_i \text{ increasing}} R_i(\tilde{q}_i; \sigma)\] is a compact maximization problem over a continuous domain (see [28] for the argument), and hence \(q\) is upper hemi-continuous in \(\sigma\).
Let \(\{q_n\}_n\) be some strictly increasing sequence of functions approximating pointwise a nondecreasing \(q\). Next, given \(\{q_n\}_n\) define \(\{\sigma_n\}_n\) as above to ensure \[q_n \in \mathop{\mathrm{\arg\!\max}}_{\tilde{q}_i \text{ increasing}} R_i(\tilde{q}_i; \sigma_n).\] Note by definition that \(\{\sigma_n\}_n\) are uniformly bounded and of uniformly bounded variation; Helly’s generalized selection theorem (for functions of bounded variation) provides the existence of a pointwise convergence subsequence \(\{\sigma_{n_k}\}\). We then have that \({q_{n_k}} \to q\) an \(\sigma_{n_k} \to \sigma\) for some tax \(\sigma\); upper hemi-continuity guarantees \(q \in \mathop{\mathrm{\arg\!\max}}_{\tilde{q}_i \text{ increasing}} R(\tilde{q}_i; \sigma)\). ◻
Proof. Let \(\sigma(q_i) = C_i - \tau_i q_i\) be an excise tax. Then for \(q_i(\theta_i)\) to be optimal, it must be that \[q_i(\theta_i) \in \mathop{\mathrm{\arg\!\max}}_{Q \in [0, 1]} \left( \theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)} - \tau_i\right) Q + C_i.\] Yet it is clear the solution to the above is given by \[q_i(\theta_i) = \mathbf{1}\left\{ \theta_i - \frac{1 - F_i(\theta_i)}{f_i(\theta_i)} \geq \tau_i \right\}\] which is clearly a threshold mechanism by Myersonian regularity. This gives the converse. For the forward direction, note that we can recover every threshold mechanism by varying \(\tau_i\) as the virtual value function is continuous. ◻
Proof. For \(i = 1,...,n-1\), let \[\sigma'_i(q_i(\theta_i)) = \sigma_i(q_i(\theta_i)) - \mathbb{E}_{F_i}[\sigma_i(q_i(\theta_i))]\] so by construction, \(\mathbb{E}_{F_i}[\sigma_i'(q_i(\theta_i))] = 0\). For \(i = n\), let \[\sigma_n'(q_n(\theta_n)) = \sigma_n(q_n(\theta_n)) + \sum_{j \neq n} \mathbb{E}_{F_j}[\sigma_j(q_j(\theta_j)))].\] But then \[\mathbb{E}_{F_n}[\sigma_n'(q_n(\theta_n))] = \mathbb{E}_{F_n}[\sigma_n(q_n(\theta_n))] + \sum_{j \neq n} \mathbb{E}_{F_j}[\sigma_j(q_j(\theta_j))] = \sum_j \mathbb{E}_{F_j}[\sigma_j(q_j(\theta_j))] = \mathbb{E}_F[\sigma(q(\theta))] \leq 0\] using budget balance of \(\sigma\). Similarly, \[\begin{align} \sigma'(q(\theta)) &= \sum_j \sigma_j'(q_j(\theta_j)) \\ &= \sum_{j \neq n} \sigma_j(q_j(\theta_j)) - \mathbb{E}_{F_j}[\sigma_j(q_j(\theta_j))] + \left(\sigma_n(q_n(\theta_n)) + \sum_{j \neq n} \mathbb{E}_{F_j}[\sigma_j(q_j(\theta_j))]\right) \\ &= \sum_j \sigma_j(q_j(\theta_j)) = \sigma(q(\theta)). \end{align}\] This finishes the proof. ◻
Proof. Let \(X_F \sim F \circ I_F\), \(X_G \sim G \circ I_G\), \(Y_F \sim F \circ I_F'\), and \(Y_G \sim G \circ I_G'\). Note that \(X_F + X_G \sim H \circ (I_F + I_G)\); the same is true of the other possible combinations.
We first show \(X_F + X_G \succsim_{SOSD} Y_F + X_G\) (independently coupled). This follows by noting that \[\begin{align} \int_{-\infty}^t \mathbb{P}_{F, G}(I_F(\theta_1) + I_G(\theta_2) \leq x) dx = \int_{-\infty}^t \mathbb{E}[\mathbf{1}\left\{I_F(\theta_1) \leq x - I_G(\theta_2)\right\} dx \\ = \int_{-\infty}^t \mathbb{E}_G [\mathbb{E}_{F | G}[\mathbf{1}\left\{I_F(\theta_1) \leq x - I_G(\theta_2)\right\}]] dx = \mathbb{E}_G \left[\int_{-\infty}^t \mathbb{E}_F[\mathbf{1}\left\{I_F(\theta_1) \leq x - I_G(\theta_2)\right\}] dx \right] \\ \leq \mathbb{E}_G \left[\int_{-\infty}^t \mathbb{E}_{F}[\mathbf{1}\left\{I_F'(\theta) \leq x - I_G(\theta_2)\right\}] dx \right] = \int_{-\infty}^t \mathbb{P}_{F, G}(I_F'(\theta_1) + I_G(\theta_2) \leq x) dx \end{align}\] where the second equality is the law of iterated expectations, followed by Fubini’s theorem and independence of \(F\) and \(G\), followed by the fact \(X_F \succsim_{SOSD} Y_F\), followed by closing the interval using the same steps.
Applying the same methodology again then gives that \(Y_F + X_G \succsim_{SOSD} Y_F + Y_G\), since \(X_G \succsim_{SOSD} Y_G\). Finally, transitivity implies \(X_F + X_G \succsim_{SOSD} Y_F + Y_G\), as desired. Thus we have that \(H \circ (I_F + I_G) \succsim_{SOSD} H \circ (I_F' + I_G')\). This finishes the argument. ◻
Proof. First, some computations for an arbitrary CDF \(G\). Applying integration by parts gives \[\int_{-\infty}^{z} G(x) dx = [x G(x)]_{-\infty}^{z} - \int_{-\infty}^{z} x dG(x).\] Since all possible information rents are bounded, \([x G(x)]_{-\infty}^{z} = zG(z)\). Next, using change of variables with \(x = R(u)\) and letting \(p^* = G(z)\) gives \(dG(x) = dG(R(u)) = du\) and \(R^{-1}(-\infty) = G(-\infty) = 0, R^{-1}(x) = G(x) = p^*\) so \[\int_{-\infty}^{z} x dG(x) = \int_{0}^{p^*} R(u) du = H(p^*)\] so \[\int_{-\infty}^{z} G(x) dx = z G(z) - H(p^*) =zp^* - H(p^*).\] This value of \(p^*\) satisfies \[\frac{d}{dp} \left[zp - H(p)\right] = z - \frac{d}{dp} \int_{0}^{p} R(u) du = z - R(p^*) = 0\] and \[\frac{d^2}{dp^2} \left[zp - H(p)\right] = -R'(p^*) \leq 0\] so \(p^*\) maximizes \(zp - H(p)\). As such, \[\int_{-\infty}^{z} G(x) dx = zp^* - H(p^*) = \max_p \{zp - H(p)\}.\] Then, by the definition of SOSD, \(F \circ I(\cdot; q) \succsim_{SOSD} F \circ I(\cdot; q')\) if and only if \[\int_{-\infty}^{z} G_q(x) dx \leq \int_{-\infty}^{z} G_{q'}(x) dx \text{ for all } z.\] Plugging in expressions from before, this is equivalent to \[\max_p \{zp - H_q(p)\} \leq \max_p \{zp - H_{q'}(p)\} \text{ for all } z.\] If \(H_q(p) \geq H_{q'}(p)\) for all \(p\) then \(zp - H_q(p) \leq zp -H_{q'}(p)\) for all \(p\) so taking maximums over \(p\) on both sides preserves the inequality. Conversely, if there were any \(p^*\) for which \(H_q(p^*) < H_{q'}(p^*)\) we can choose \(z^*\) to have \(G(z^*) = p^*\), in which case \[H_q(p^*) < H_{q'}(p^*) \implies z^*p^* - H_q(p^*) > z^* p^* H_{q'}(p^*) \implies \int_{-\infty}^{z^*} G_q(x) dx > \int_{-\infty}^{z^*} G_{q'}(x) dx\] as desired. This finishes the proof. ◻
Proof. First, we remark strong regularity implies convexity of \(F\) and concavity of the virtual value function. To see this, note that \[\frac{d}{ds} \log(f) = \frac{f'}{f} \geq 0 \iff f' \geq 0\] which gives convexity of \(F\). Second, let \(H(s) = \frac{1 - F(s)}{f(s)}\) denote the inverse hazard rate. Note that \[\psi(s) = s - H(s) \implies \psi''(s) = -H''(s)\] so we need that \(H''(s) \geq 0\) for \(\psi\) to be concave. A computation gives that \[H'(s) = -1 - \frac{(1 - F(s))f'(s)}{f(s)^2} = -1 - H(s) \frac{f'(s)}{f(s)}\] Differentiating again gives \[H''(s) = -\left( H'(s)\frac{f'(s)}{f(s)} + H(s)\left( \frac{f(s) f''(s) - f'(s)^2}{f^2} \right) \right)\] Note moreover that \(\frac{d}{ds} \log(f(s)) \geq 0\), and \[\frac{d^2}{ds^2}\log(f(s)) = \frac{d}{ds} \frac{f'(s)}{f(s)} = \frac{f(s) f''(s) - f'(s)^2}{f(s)^2} \leq 0\] by log-concavity of \(f\). Hence \(H''(s)\) is positive, since
\(H'(s) \leq 0\) by Myersonian regularity (the inverse hazard rate is decreasing).
\(\frac{f'(s)}{f(s)} \geq 0\) since it is increasing.
\(H(s) \geq 0\) by assumption.
\(\frac{f(s)f''(s) - f'(s)^2}{f(s)^2} \leq 0\) by log-concavity of \(f\).
Hence \(H''(s) \geq 0\), so \(\psi''(s) \leq 0\) and hence \(\psi\) is concave.
Now nonincreasingness of \(W\). A basic computation implies we can rewrite \(W(\theta)\) as \[W(\theta) = 3(1-F(\theta)) - \theta f(\theta) + \frac{(1-F(\theta))^2 f'(\theta)}{f(\theta)^2}\] and \[W'(\theta) = -4f(\theta) - \theta f'(\theta) + \frac{d}{d\theta} \left[ \frac{f'(\theta)}{f(\theta)} \cdot \frac{(1-F(\theta))^2}{f(\theta)} \right]\] by a computation.
Since strong regularity implies \(f' \geq 0\), we know that \[\frac{d}{d\theta}\frac{(1-F(\theta))^2}{f(\theta)}= -\frac{1-F(\theta)}{f(\theta)^2} [2f(\theta)^2 + (1-F(\theta))f'(\theta)] < 0 \text{ and } -\theta f'(\theta) \leq 0.\] Moreover, strong regularity implies \(\frac{f'(\theta)}{f(\theta)}\) is positive and decreasing, and hence we know that \[\frac{d}{d\theta} \left[ \frac{f'(\theta)}{f(\theta)} \cdot \frac{(1-F(\theta))^2}{f(\theta)} \right] \leq 0.\] Together, these two observations allow us to bound \(W'(\theta)\) by \(-4 f(\theta)\). Since \(f\) is fully supported, \(W'(\theta) < 0\), as desired.
The second statement. Note \[\frac{W(s)}{1 - F(s)} = \psi'(s) - \frac{f(s)}{1 - F(s)} \psi(s) = \psi'(s) - \frac{\psi(s)}{H(s)}\] where \(H(s) = \frac{1 - F(s)}{f(s)}\).
Further simplifying gives that this can be written as \[\psi'(s) - \frac{\psi(s)}{H(s)} = \psi'(s) - \frac{s - H(s)}{H(s)} = \psi'(s) - \frac{s}{H(s)} + 1\] and so its derivative can be written as \[\psi''(s) - \frac{H(s) - H'(s) s}{H(s)^2}\] Recall that \(\psi''(s) < 0\). Moreover, note \(H(s) \geq 0\) always, while by Myersonian regularity \(H'(s) \leq 0\), i.e. the inverse hazard rate is decreasing. Thus the second term is the difference of a negative and a positive term and hence is negative. Because the sum of (strictly) negative terms is (strictly) negative, \(\frac{W(s)}{1 - F(s)}\) is strictly decreasing. ◻
Proof. The proof follows in three steps.
The first step is to show that all threshold mechanisms which induce thresholds in each dimension as stated by the theorem are on the frontier. Formally,
Lemma 3. Fix a tax policy \(\sigma\) inducing a threshold mechanism in each dimension with thresholds \(k_i \in \{\theta_i : W_i(\theta_i) \leq 1 - F_i(\theta_i), W_i(\theta_i) \geq 0\}\), and let \(I\) be its information rent. Then \(I\) is \(F\)-unimprovable.
Proof. We start with some notation. Note that the consumer’s individual information rent in market \(i\) given a threshold of \(k_i\) is given by \[I_i(\theta_i | k_i) = \max\left\{\theta_i - k_i, 0\right\} - \int W_i(s) \mathbf{1}_{\{s \geq k_i\}} ds\] Total information rents are then \(I(\theta | \mathbf{k}) = \sum_i I_i(\theta_i | k_i)\) for some \(\mathbf{k} \in [0, 1]^n\). Let \(I_p(\mathbf{k})\) denote the \(p\)-th percentile of aggregate information rents. Cumulative information rents up to percentile \(p\) induced by thresholds \(\mathbf{k}\) are then \[H(p | \mathbf{k}) = \mathbb{E}_F[I(\theta | \mathbf{k})\mathbf{1}_{\{I(\theta | \mathbf{k}) \leq I_p(\mathbf{k})\}}].\] We have the following computational lemma.
Lemma 4. For all \(i\), \[\frac{\partial H(p | \mathbf{k})}{\partial k_i} = \mathbb{E}_F\left[ \frac{\partial I(p | \mathbf{k})}{\partial k_i}\mathbf{1}_{\{I(\theta | \mathbf{k}) \leq I_p(\mathbf{k})\}} \right]\]
Proof. Define \(\Omega_p(\mathbf{k}) = \{\theta \in \Theta : I(\theta | \mathbf{k}) \leq I_p(\mathbf{k})\}\) to be the region of types where information rents are below the \(p\)-th percentile. We then have that \[H(p | \mathbf{k}) = \int_{\Omega_p(\mathbf{k})} I(\theta | \mathbf{k}) f(\theta) d\theta.\] Differentiating with respecting to \(k_i\) and applying the Leibniz integral rule gives \[\frac{\partial H(p | \mathbf{k})}{\partial k_i} = \underbrace{\int_{\Omega_p(\mathbf{k})} \frac{\partial I(\theta|\mathbf{k})}{\partial k_i} f(\theta) d\theta}_{\text{Integrand Term}} + \underbrace{\int_{\partial \Omega_p(\mathbf{k})} [I(\theta| \mathbf{k}) f(\theta)] \left( v \cdot n \right) dS}_{\text{Boundary Term}}\] where \(\partial \Omega_p(\mathbf{k})\) is the boundary of \(\Omega_p(\mathbf{k})\), \(v\) is the velocity of the boundary changing, and \(n\) is the outward unit normal vector. We are done if we can show the boundary term vanishes. On the boundary, we know that \(I(\theta | \mathbf{k}) = I_p(\mathbf{k})\), and hence the boundary term can be rewritten as \[I_p(\mathbf{k}) \int_{\partial \Omega_p(\mathbf{k})} f(\theta)(v \cdot n) dS\] But we know that the \(p\)-th percentile is equivalent to \[p = \int_{\Omega_p(\mathbf{k})} f(\theta) d\theta.\] Differentiating with respect to \(k_i\) (again applying the Leibniz integral rule to the right hand side) gives that \[0 = \int_{\Omega_p(\mathbf{k})} \frac{\partial f(\theta)}{\partial k_i} d\theta + \int_{\partial \Omega_p(\mathbf{k})} f(\theta)(v \cdot n) dS\] but of course \(\frac{\partial f(\theta)}{\partial k_i} = 0\), and hence we get that \[\frac{\partial H(p | \mathbf{k})}{\partial k_i} = \int_{\Omega_p(\mathbf{k})} \frac{\partial I(\theta | \mathbf{k})}{\partial k_i} f(\theta) d\theta + I_p(\mathbf{k}) \cdot 0 = \mathbb{E}\left[ \frac{\partial I(\theta | \mathbf{k})}{\partial k_i} \mathbf{1}_{I(\theta | \mathbf{k}) \leq I_p(\mathbf{k})} \right]\] as desired. ◻
A further computation gives \[\frac{\partial I(\theta|\mathbf{k})}{\partial k_i} = -\mathbf{1}_{\{\theta_i \geq k_i\}} + W_i(k_i).\] Combining this with Lemma 4 implies \[\frac{\partial H(p | \mathbf{k})}{\partial k_i} = \mathbb{E}\left[ \left( - \mathbf{1}_{\{\theta_i \geq k_i\}} + W_i(k_i) \right) \mathbf{1}_{I(\theta | \mathbf{k}) \leq I_p(\mathbf{k})} \right] = p W_i(k_i) - \mathbb{E}\left[\mathbf{1}_{\{\theta_i \geq k_i\}}\mathbf{1}_{I(\theta | \mathbf{k}) \leq I_p(\mathbf{k})} \right].\] Consider now the set \[k_i \in \{\theta_i: W_i(\theta_i) \leq 1 - F_i(\theta_i), W_i(\theta_i) \geq 0\}.\] Define \[\underline k_i = \min \{\theta_i : W_i(\theta_i) \leq 1 - F_i(\theta_i), W_i(\theta_i) \geq 0\}\] to be the lowest potential threshold for each \(i\). Since we know that \[W_i(0) = \psi_i'(0) - f_i(0)\left(-\frac{1}{f_i(0))} \right) > 1,\] we have that \(\underline k_i > 0\) for all \(i\). This implies that no agent in the set \(S = \prod_i [0, \underline k_i]\) consumes any good at any potential threshold. Since \(F\) is fully supported, \(F(S) > 0\). Taking \(p = \frac{F(S)}{2}\) gives that \[\frac{\partial H(p | \mathbf{k})}{\partial k_i} = \frac{F(S)}{2} W_i(k_i) - 0\] since no agent below the \(p\)-th percentile receives any good. Since \(W_i(k_i) \geq 0\) for all potential thresholds, this gives that \(\frac{\partial H(p | \mathbf{k})}{\partial k_i}\) is positive for all \(k_i\).
Finally, taking \(p = 1\) instead gives that \[\frac{\partial H(1 | \mathbf{k})}{\partial k_i} = W_i(k_i) - \mathbb{E}[\mathbf{1}_{\left\{\theta_i \geq k_i\right\}}] = W_i(k_i) - (1 - F_i(k_i)).\] Since \(W_i(k_i) \leq 1 - F_i(k_i)\) for every potential threshold, this gives that \(\frac{\partial H(1|\mathbf{k})}{\partial k_i} \leq 0\).
Together, these two inequalities imply for any potential threshold \(k_i\) on the frontier, raising \(k_i\) will always increase cumulative information rents for low \(p\) and decrease them for high \(p\), and so no threshold within the set of potential thresholds can dominate another. This finishes the proof. ◻
The next steps show that only these distributions of information rents are \(F\)-unimprovable. For the remainder of this proof, given given Lemmas 1 and 3, we drop subscripts and focus on the case \(n = 1\).
Lemma 5. A threshold mechanism uniquely maximizes the surplus of the lowest type subject to producing any required level of surplus \(S\) if and only if the threshold is in \(\{\theta: W(\theta) \leq 1-F(\theta), W(\theta) \geq 0\}\).
Proof. The utility of the lowest type is \[\int_{\underline{\theta}}^{\underline{\theta}} q(s) ds - \int_\Theta W(s)q(s)ds = - \int_\Theta W(s)q(s)ds\] while total consumer surplus is (after using the usual mechanism design approach of interchanging the order of integration) \[\int_\Theta \left( \int_{\underline{\theta}}^{\theta} q(s)ds - \int_\Theta W(s)q(s) ds \right) d\theta = \int_\Theta (1-F(\theta)-W(\theta))q(\theta) d\theta.\] As such, maximizing the surplus of the lowest type subject to a total consumer surplus constraint is equivalent to the program \[\begin{align} \max_q & \int_\Theta -W(\theta)q(\theta)d\theta \\ \text{s.t.} & \int_\Theta (1-F(\theta)-W(\theta))q(\theta)d\theta \geq S; \\ & q(\theta) \text{ increasing}, q(\theta) \in [0, 1]. \end{align}\] The Lagrangian of the relaxed problem ignoring the monotonicity constraint is \[\begin{align} \mathcal{L}(q, \lambda) &= \int_\Theta -W(\theta)q(\theta)d\theta - \lambda \left( S - \int_\Theta (1-F(\theta)-W(\theta))q(\theta)d\theta \right) \\ &= \int_\Theta [-W(\theta) + \lambda (1-F(\theta)-W(\theta)]q(\theta)d\theta - \lambda S. \end{align}\] with \(\lambda \geq 0\). The solution for any \(\lambda\) is \[q = \mathbb{1}\left\{-W(\theta) + \lambda (1-F(\theta)-W(\theta) \geq 0\right\} = \mathbb{1}\left\{\frac{W(\theta)}{1-F(\theta)-W(\theta)} \leq \lambda \right\}.\] Moreover, by strong regularity, \[\frac{W(\theta)}{1-F(\theta)-W(\theta)} = \frac{1}{\frac{1-F(\theta)}{W(\theta)}-1}\] is strictly decreasing if \(\frac{W(t)}{1-F(t)}\) is strictly decreasing, which we know by Proposition 4. Thus, a threshold mechanism uniquely maximizes information rents for the lowest type.
Next, which thresholds are admissible? If \(W(\theta) \leq 0\) then \(-W(\theta) + \lambda (1-F(\theta)-W(\theta) \geq 0\) for all \(\lambda\) so the good must be allocated in this case. If \(W(\theta) \geq 1-F(\theta)\) then \[-W(\theta) + \lambda (1-F(\theta)-W(\theta) \leq -(1-F(\theta)) +\lambda (0) = 0\] and the good must not be allocated for any \(\lambda\). As such, for there to exist some \(\lambda\) for which a threshold \(k\) is optimal, it must be that \[k \in \{\theta: W(\theta) \leq 1-F(\theta), W(\theta) \geq 0\}.\] Such a \(\lambda\) exists since \(-W(\theta) + \lambda (1-F(\theta)-W(\theta)\) is continuous in \(\lambda\) and is both positive and negative as \(\lambda\) ranges from \(0\) to \(\infty\). ◻
Finally, we combine the above lemmata to show no allocation rule can dominate a threshold mechanism at a threshold \(k \in \{\theta: W(\theta) \leq 1-F(\theta), W(\theta) \geq 0\}\).
Let \(q\) be a threshold mechanism in the above class. Towards a contradiction, suppose the distribution of information rents generated by \(q'\) SOSD dominates the distribution of information rents generated by \(q\). By Lemma 2, it must be that \(H_{q'}(1) \geq H_q(1)\); i.e. allocation rule \(q'\) generates weakly more surplus than the threshold mechanism \(q\). By Lemma 5, \(q\) uniquely maximizes the rents of the lowest type subject to total rents being at least \(H_q(1)\) so \(q'\) must give strictly lower information rents to the agent of the lowest type; that is \(R_q(0) > R_{q'}(0)\). Next, \(H_q(0) = H_{q'}(0) = 0\) since integrals over empty regions are zero. Finally, \[H'_q(0) = R_q(0) \text{ and } H'_{q'}(0) = R_{q'}(0)\] by the fundamental theorem of calculus. Thus, for sufficiently small \(\varepsilon\), it must be that \[H_q(\varepsilon) > H_{q'}(\varepsilon).\] As such, by Lemma 2, the distribution of information rents generated by \(q'\) cannot SOSD dominate the distribution of information rents generated by \(q\).
Conversely, for any threshold mechanism \(q\) outside of that class, by Lemma 5 there exists some threshold mechanism \(q'\) in that class which leads to to the same total information rents while giving the lowest type higher information rents. Thus, \(q'\) dominates \(q\). ◻
Proof. We start with a computation. \[\begin{align} W(s) = (1 - F(s))\psi'(s) - f(s)\psi(s) = 3(1 - F(s)) - s f(s) + \frac{(1 - F(s))^2}{f(s)^2} f'(s) \\ = (1 - F(s)) - s f(s) + 2 (1 - F(s)) + \frac{(1 - F(s))^2}{f(s)^2} f'(s) \end{align}\] At the monopoly price, we know that \(\theta^* = \frac{1 - F(\theta^*)}{f(\theta^*)}\) and so \[1 - F(\theta^*) - \theta^* f(\theta^*) = f(\theta^*) \left(\frac{1 - F(\theta^*)}{f(\theta^*)} - \theta^* \right) = 0\] and so we get that \[W(\theta^*) = 2(1 - F(\theta^*)) + \frac{(1 - F(\theta^*))^2}{f(\theta^*)^2} f'(\theta^*) > 1 - F(\theta^*).\] In particular, this implies that at the monopoly threshold, \(W(\theta^*) > 1 - F(\theta^*)\). By strong regularity and Theorem 3, \(\theta^*\) cannot be a threshold which is on the fairness-efficiency frontier.
Finally, consider \(1 - \frac{W(\theta)}{1 - F(\theta)}\). A necessary condition by Theorem 3 for this to be on the frontier is that this expression is positive. However, we know that \(1 - \frac{W(\theta^*)}{1 - F(\theta^*)} < 0\); since \(\frac{W(\theta)}{1 - F(\theta)}\) is nonincreasing by Proposition 4, for this to be positive it must be that any admissible threshold has \(k > \theta^*\). This finishes the proof. ◻
In this appendix we elaborate on the importance of separability in proving the second half of the Proposition 1. In particular, we give an example of a nonseparable mechanism whose direct implementation is not firm optimal.
Suppose there are two firms and consumer valuations are degenerate at \(1\) for both goods. Consider a government policy of the form \[\sigma(0,0) = \sigma(1,1) = 0; \sigma(0, 1) = \sigma(1, 0) = 1.\] It is then an equilibrium for both firms to price the good at \(0\) and offer the good; and consumers to buy from firm \(1\) only. Note that this is indeed an equilibrium; if any firm deviates and offers a different price or quantity other than prescribed, then consumers buy only from the other firm and get a payoff of \(2\) (their highest possible payoff). The implied (constant) direct mechanism is then for firms to set \[(q_1, t_1)(\theta) = (1, 0) \text{ and } (q_2, t_2)(\theta) = (0, 0) \text{ for all } \theta.\]
Yet firm \(1\) can profitably deviate by posting a price of \(1 - \varepsilon\) (instead of posting a price of zero as before); since firm \(2\) is not offering the good in direct mechanisms, each consumer will still buy the good from firm \(1\). The problem here is that the direct mechanism contracts only on the realized allocation; thus firm \(2\) offers \((0, 0)\) instead of the “off-path” proposed allocation \((1, 0)\). However, because the tax is not separable, there are strategic interactions between firms, and thus the off-path allocation \((1, 0)\) was necessary to discipline firm \(1\)’s incentives in offering \((1, 0)\) as well. Thus direct mechanisms are with loss of generality in the class of all mechanisms when \(\sigma\) is nonseparable.
In this appendix we give a microfoundation of where the correlation for valuations might come from between goods. Suppose that we are solving a traditional two good utility maximization problem, namely \[\max_{x_1, x_2} u(x_1, x_2) \text{ s.t. } p_1 x_1 + p_2 x_2 \leq w\] for a fixed level of wealth \(w\). Under standard differentiability conditions, this will induce indirect demand functions \(x_1(p_1, p_2)\) and \(x_2(p_1, p_2)\). Taking the continuum approach as in [30] and assuming a maximal price \(\bar p\) for both goods, we can interpret \(x_i(p_i, p_j)\) as the mass of consumers whose types are above \(p_i\) conditional on their type being \(p_j\) in the second good. Formally, we have that there is a candidate joint distribution \[F(p_i | p_j) = 1 - x_1(p_i, p_j) \text{ and } F(p_j | p_i) = 1 - x_2(p_i, p_j).\] This induces a correlational structure over types so long as \(F\) is well-defined. Thus, we can view our exercise—characterizing how correlational structures affect the distribution of information rents—also as an exercise that characterize consumer surplus at difference prices as the demand structure or utility function changes. The only remaining question is whether \(u\) and \((x_1, x_2)\) give rise of a valid joint distribution, given that we have only specified the family of conditional distributions.
The sufficient condition, which we borrow directly from the literature on Gibbs sampling, is log-separability of the indirect demand functions. Formally,
Proposition 6. Suppose demand is derivative log-separable, i.e. \(\frac{\frac{\partial x_1(p_1, p_2)}{\partial p_1}}{\frac{\partial x_2(p_1, p_2)}{\partial p_2}} = \frac{f(p_1)}{g(p_2)}\) for some \(f, g: \mathbb{R} \to \mathbb{R}\). Then there is a well-defined joint distribution \(F\) induced by the indirect demand functions.
Proof. See [31], Theorem 4.1. ◻
Massachusetts Institute of Technology, ericgao@mit.edu, daniel57@mit.edu. We are indebted to Drew Fudenberg and Alex Wolitzky for invaluable guidance and Eric Yan for conversations early in the development of this project. We also thank Piotr Dworczak, Marina Halac, Michelle Hyun-Kim, Navin Kartik, Andrew Komo, Elliot Lipnowski, Stephen Morris, Ellen Muir, Bryant Xia, Ivan Werning, Kai Hao Yang, and seminar participants at MIT and Yale for helpful discussion. Coarse.ink and Refine.ink were used to check the paper for clarity and correctness. Luo acknowledges financial support from the NSF Graduate Research Fellowship. Part of this paper was written while Luo was visiting Yale University. All errors are, of course, ours alone.↩︎
For example, Washington state is constitutionally barred from implementing income taxes, and thus lawmakers must resort to property or commodity taxes to raise revenue and achieve redistributive goals.↩︎
[3] suggest that in states with low capacity, commodity taxes are much easier to implement than income taxes.↩︎
Formally, we suppose that taxes and subsidies can only depend on observed consumer purchasing decisions.↩︎
Famously, Washington state was able to impose a capital gains tax by arguing that capital gains were not property—and therefore not income—and instead an excise tax on the sale of a financial commodity. Our analysis can shed light on what types of excise taxes Washington policymakers should use to attain progressive redistributive targets.↩︎
In our model, we will abuse notation and switch between the ex-ante single consumer interpretation and the continuum consumer interpretation of uncertainty. Note these interpretations are mathematically equivalent.↩︎
Here we abuse notation, since \(m_i(\theta)\) is a lottery over messages; we are extending the value of \(q_i, t_i\) to lotteries in the usual way.↩︎
Absent such an assumption, the government can consider pathological mechanisms that coordinate payments bundle-by-bundle so that direct mechanisms on the part of the firms are not without loss of generality, which would render further analysis of the model intractable. While relaxing this assumption is a potentially interesting direction of research, it is out of the scope of our current analysis and we defer it to future work.↩︎
We remark here that the equivalence between (1) and (3) is known to the literature, though from what we can tell a formal proof of the statement is missing. Specifically, [23] and [24] reference [25], who give equivalences of the supermodular order to the concordance order but defer a formal proof to the original manuscript by [27]. However, we are unable to find the explicit proof provided by [27] of this statement. Thus, our proof of this equivalence is, as far as we can tell, a new proof of a known result, and an independent contribution to the literature in that regard.↩︎
If this assumption were violated, then the government could always improve everyone’s welfare by introducing an arbitrarily high lump-sum subsidy.↩︎
Note that it is sufficient to focus on the second order stochastic dominance part of the definition, since \(I\) being more fair than \(I'\) requires that \(I\) is weakly more efficient than \(I'\). We include the more efficient part of this definition to illuminate our frontier characterization.↩︎
We thank Kai Hao Yang for pointing us to this reference and result therein.↩︎
Without separability, it may be possible that all profitable deviations are joint across multiple markets. If this is the case, applying the revelation principle to other markets changes the problem firm \(j\) faces. Thus, the equivalent mechanism may no longer be optimal, even if it is implementable.↩︎