Should Demand Models Incorporate Competitor Prices?
Oblivious Learning and Algorithmic Collusion


Abstract

On a platform with many sellers, should a pricing algorithm explicitly model competitors’ prices when learning demand? Classical learning arguments suggest an affirmative answer: ignoring competitors induces model misspecification and inefficiency. In contrast, recent work on algorithmic collusion suggests that strategic obliviousness—deliberately ignoring competitor prices—may facilitate collusive outcomes and improve profits. We study this modeling choice in a stylized competitive market with unknown noisy demand, in which multiple sellers repeatedly set prices and estimate demand via iterated least squares, and either incorporate competitors’ prices into their demand models (informed) or ignore them (oblivious). We first show that, relative to a monopolist, an oblivious seller in a competitive market must explore more aggressively to compensate for the loss of dynamic competitor information. Building on this insight, we characterize market dynamics when all sellers are oblivious and show that prices converge to the competitive outcome under sufficient exploration, while a continuum of pseudo-equilibria arises when exploration decays. Analyzing the resulting price trajectories, we uncover an excursion phenomenon that gives rise to transient collusive patterns that dissipate as learning progresses. In markets with both oblivious and informed sellers, the informed strictly out-earn the oblivious. Read as a strategy game, the modeling choice has a unique Nash equilibrium: the all-informed market, in which prices converge to the competitive outcome efficiently. Overall, our results indicate that collusive patterns are not robust and are not sustained by oblivious modeling; therefore, incorporating competitor information, together with sufficient price exploration, remains a reliable strategy for sellers in competitive markets.

Algorithmic collusion, Dynamic pricing, Multi-agent learning, Demand learning, Model misspecification

1 Introduction↩︎

1.1 Background↩︎

The rising popularity of learning algorithms and the advent of AI have significantly expanded algorithmic decision-making, allowing algorithms to make autonomous or semi-autonomous decisions that were once under the full purview of humans. The resulting algorithmic economy is hence one where interactions among human decision makers are increasingly replaced by interactions among algorithms. This is particularly true in the context of pricing analytics, where decades of research have given rise to an abundance of dynamic pricing algorithms [1].

One consequence of these trends is a growing need to better understand how pricing algorithms interact and influence market outcomes, both for regulatory oversight and revenue management practices. One of the most recent concerns is algorithmic collusion, where independent pricing algorithms operating in competitive markets may learn to exhibit or even sustain supracompetitive prices (i.e., prices that are higher than the competitive level or even near-collusive). The concept of algorithmic collusion emerged around 2015 [2], [3] and became a topic of heated discussion with the seminal work of [4]. Since then, related issues have made headlines in the popular press and have received regulatory attention [5][8]. Despite extensive numerical studies and theoretical exposition, many problems remain open in this area [9], [10]. It is nevertheless important to recognize that algorithmic collusion has introduced new perspectives on the interaction of learning and pricing algorithms, especially when competition is involved.

1.2 Motivation↩︎

One such perspective, fundamental to pricing practices in competitive markets, is whether sellers should explicitly incorporate competitors’ prices into their demand models when learning to price. Consider a market with multiple sellers offering imperfectly substitutable products, a setting representative of many online retail platforms. In such environments, it is natural to posit that a seller’s demand depends on all sellers’ prices, so explicitly modeling competitors’ prices appears both economically sound and methodologically prudent.

1.2.0.1 Oblivious learning and collusive outcomes.

The algorithmic collusion literature, however, has challenged this view. While explicit collusion is prohibited by antitrust laws, tacit collusion through a seemingly benign pricing algorithm is attractive to sellers, as opacity of algorithmic decision-making affords plausible deniability, and there is little consensus on what constitutes a legally “correct” learning algorithm in competitive settings. For example, several works suggest that sellers may benefit from strategic obliviousness—deliberately ignoring competitors’ prices when modeling demand—as this may facilitate collusive outcomes [2], [11], [12]. In practice, there are also many sensible reasons for sellers to consider oblivious modeling: collecting competitors’ prices may be costly or infeasible, and learning from one’s own price-demand data is simpler. This means that, if ignoring competitors’ prices can lead to collusive outcomes, sellers may have both economic and operational incentives to do so.

1.2.0.2 Impact of demand noise and price exploration.

Although oblivious sellers effectively behave as monopolists in their own demand models, the interaction of “reasonable” monopolistic learning algorithms yields puzzling market behavior. In monopolistic dynamic pricing, iterated least squares with cumulative exploration of order \(\Theta(\sqrt{n})\), where \(n\) denotes the pricing horizon (or number of interactions with buyers), is widely regarded as a sensible and near-optimal choice [13]. Yet, when all sellers in a competitive market are oblivious and adopt this exploration rate, numerical simulations reveal a wide range of possible outcomes, spanning near-competitive prices, near-collusive prices, and indeterminate regimes. Figure 1 illustrates three representative sample paths. This phenomenon echoes earlier observations in settings without demand noise or price exploration: [2] documents a continuum of possible price limit points under oblivious learning. That similar patterns persist even in the presence of exploration and noise raises several fundamental questions. Do price exploration and demand noise affect the nature of the limit points? Is the widely accepted \(\Theta(\sqrt{n})\) exploration rate still appropriate in competitive markets? How do model misspecification and exploration interact to shape market dynamics?

a
b
c

Figure 1: Three sample paths of the prices for two oblivious sellers under cumulative exploration of order \(\Theta(\sqrt n)\) over \(10{,}000\) periods. The blue and orange lines are the per-period prices of sellers 1 and 2; the red and green dotted lines mark the competitive and collusive prices \(p^{NE}\) and \(p^C\). Implementation details are in Appendix 8.2.. a — A near-competitive outcome., b — An indeterminate outcome., c — A near-collusive outcome.

1.3 Contributions↩︎

In this work, we study this modeling choice in a dynamic pricing environment with unknown and noisy demand. We consider a stylized competitive market in which multiple sellers repeatedly set prices and estimate demand via iterated least squares, and sellers may either incorporate competitors’ prices into their demand models (informed) or ignore them (oblivious). A key challenge is to analyze the market dynamics when oblivious sellers are present, as the asymmetry in information evolves and feeds back into the dynamics. Our contributions include the following:

  1. We show that oblivious sellers must explore more aggressively than a monopolist because a “spiral-up” effect compels them to escalate exploration to offset model misspecification, and we identify linear exploration as an optimal strategy that nevertheless incurs a persistent exploration tax (Section 4).

  2. We characterize market outcomes when all sellers are oblivious: under persistent (linear-rate) exploration, prices converge to the competitive (Nash) outcome, while decaying exploration produces a continuum of pseudo-equilibria, mirroring [2] and explains the puzzling outcomes of Figure 1 (Section 5).

  3. We uncover an excursion phenomenon in the price trajectories that explains observed collusive patterns, but these patterns are fragile and dissipate as learning progresses (Section 5.3.1).

  4. When informed sellers are present, they consistently learn the true demand function and the market converges to the competitive outcome without the persistent exploration tax that oblivious sellers bear. In mixed markets, the informed strictly out-earn the oblivious (Section 6).

  5. Consolidating these three market compositions, we show that the choice between oblivious and informed modeling can be read as a strategy game (Table 1), in which informed strictly dominates oblivious and the unique strict pure-strategy Nash equilibrium is the all-informed market. The same conclusion carries over to a general \(N\)-seller setting (Section 7).

12pt

Table 1: The strategy game. Rows index seller 1, columns index seller 2. Each seller chooses whether to model competitor prices (informed) or not (oblivious). Cell entries are the revenues (payoffs) for the corresponding strategy profile, where “Nash” denotes the payoff at the competitive outcome. The unique Nash equilibrium of this game is (informed, informed); see Section 7 and Table 3 for more details.
oblivious informed
oblivious \(<\) Nash, \(<\) Nash \(<\) Nash, \(\ge\) Nash
informed \(\ge\) Nash, \(<\) Nash Nash, Nash

Overall, oblivious demand modeling does not robustly sustain collusive pricing, and explicitly modeling competitors together with sufficient price exploration remains the reliable design principle for learning-based pricing in competitive markets.

2 Literature Review↩︎

2.0.0.1 Algorithmic collusion.

The seminal work of [4] drew significant attention by demonstrating through numerical experiments that Q-learning agents can learn to charge supracompetitive prices and even exhibit collusive reward-punishment schemes. Numerous works followed up and added further numerical evidence; see, e.g., [14] and references therein. Most theoretical work on algorithmic collusion proceeds through the lens of a repeated Prisoner’s Dilemma with a deterministic payoff matrix, primarily for reinforcement learning (RL) algorithms [11], [12], [15][20]. Despite numerous insights, the above works are limited in that they do not consider demand noise. Another line of work on algorithmic collusion stems from an algorithmic game-theoretic view and largely studies repeated (Bertrand) pricing games [21][24] but under deterministic demands as well. Some works also design algorithms intended to collude [25][27]. Empirical evidence on algorithmic collusion has been sparse overall, with [28] and [29] considering simplified settings. For more comprehensive reviews, we refer readers to [9] and [10].

2.0.0.2 Pricing and learning in competition.

A common result in multi-seller revenue management is that the system will achieve optimal regret with respect to a Nash equilibrium if every seller deploys a specific (often informed) algorithm (see, e.g., [30]). However, the often-used dynamic benchmark is not necessarily appropriate in the context of algorithmic collusion (Appendix 8.3). An earlier line of work studies multi-seller pricing with a fixed inventory and a finite time horizon; see [31] for a comprehensive review. Learning in competition is also related to learning in non-stationary environments, but the commonly assumed sublinear “variation budget” that constrains how much the environment can change over time (e.g., [32]) is equivalent to assuming a priori that the competitor’s price will converge in our case. Our convergence-to-Nash conclusion also has a long classical antecedent in game theory: when payoffs are known and players adaptively best respond to historical averages of opponents’ actions, the dynamics converge to the Nash set (see, e.g., [33], [34] and the subsequent literature on adaptive learning in games). Our setting departs from this benchmark in that the demand parameters are unknown, so the question is no longer one of purely adaptive decision-making but also of joint estimation under (possibly misspecified) demand models.

2.0.0.3 Price exploration.

Most closely related to our work is [2], which considers two sellers with linear demand and studies the limit prices when sellers do not incorporate competitor prices in their modeling. Crucially, however, they do not consider price exploration (versus exploitation or greedy strategies). In the dynamic pricing literature, it has long been known that, without proper price exploration, the presence of demand noise can cause suboptimal performance for the sellers [13], [35]. As a result, the algorithms analyzed in [2]—as well as the works mentioned above that consider deterministic demand and/or no price exploration—may not reflect practical settings accurately.

2.0.0.4 Oblivious learning.

Besides [2], multiple works related to algorithmic collusion consider oblivious settings [11], [12] without formally discussing this modeling decision. There is also a fruitful line of work whose focus is to design a distributed algorithm such that, if it is deployed by all sellers, the market converges to a Nash outcome efficiently; e.g., [36] and [37] consider multinomial logit demand models and design online gradient descent algorithms. In contrast, our work lets market outcomes steer the agents’ modeling choices and establishes that oblivious modeling is suboptimal. In addition, we consider linear demand models, which provide less information to oblivious sellers than logit demand models do (see Section 3.2).

2.0.0.5 Learning under misspecified models.

Our analysis connects to the literature on statistical learning under a misspecified model. Classical results show that least-squares and maximum-likelihood estimators converge to a “pseudo-true” parameter [38], [39]. [40] extends this idea to a game-theoretic setting and proposes a relevant solution concept, whereas we show explicit convergence. Closer to our pricing context, [41] and [42] study single-seller misspecification, whereas we explicitly consider a multi-seller setting where the misspecification comes from ignoring competitors with dynamic and interacting prices.

3 Model↩︎

3.1 Demand Model↩︎

We consider a market with \(N \ge 2\) sellers indexed by \([N] = \{1,\ldots,N\}\), where each seller offers a single product. Time is discrete and indexed by \(n \in \mathbb{Z}_+\). Let \(\mathbf{p}_n = (p_{n,1},\ldots,p_{n,N}) \in \mathbb{R}^N\) denote the vector of prices at time \(n\). The demand faced by seller \(i\) at time \(n\) follows a linear model: \[\label{eq:demand} d_{n,i} = \alpha_i + \beta_i p_{n,i} + \sum_{j \neq i} \gamma_{i,j} p_{n,j} + \varepsilon_{n,i},\tag{1}\] where the following assumptions hold:

  • Prices are bounded: \(p_{n,i} \in [l,u]\) for all \(n\) and \(i\), with \(u>l>0\).

  • The noise terms \(\varepsilon_{n,i}\) are i.i.d.across \(n\) and \(i\), with zero mean and bounded support.

  • The demand parameters satisfy \(\alpha_i>0\), \(\beta_i<0\), \(\gamma_{i,j}>0\) for all \(i\neq j\), and \(-\beta_i>\gamma_i\) for all \(i\), where \(\gamma_i\triangleq \sum_{j\neq i}\gamma_{i,j}\).

These assumptions are standard in the dynamic pricing literature [2], [13]. They ensure that expected demand is non-negative, products are imperfect substitutes, own-price effects dominate cross-price effects, and the competitive benchmark is well-defined. The price bounds reflect practical considerations, such as marginal costs and reasonable upper limits on prices. Under the true demand model, we distinguish sellers by whether they explicitly model competitive effects.

Definition 1 (Oblivious vs. informed). Seller \(i\) is said to be oblivious* if they model demand using a misspecified monopolistic form, \[d_{n,i} = a_i + b_i p_{n,i} + \varepsilon_{n,i}.\] Seller \(i\) is informed if they correctly specify demand according to 1 .*

3.1.0.1 Remark.

Oblivious modeling may arise for several reasons. A seller may lack reliable access to competitor prices, face substantial missing data, or prefer an operationally simpler model. Even when competitor prices are observable, it may be unclear how to incorporate them effectively. See [2] for more discussion of this modeling choice.

3.2 Learning and Pricing Dynamics↩︎

All sellers, oblivious or informed, follow the same learning-and-pricing protocol, differing only in the demand model they estimate. Each seller repeatedly estimates demand via least squares, computes a myopic revenue-maximizing price based on the estimate, and adds random perturbations to ensure exploration. At time \(n\), an oblivious seller \(i\) observes their own price-demand history \(\mathcal{H}_{n,i}^{ob} = \{(p_{m,i}, d_{m,i})\}_{m=1}^n\) only, whereas an informed seller \(i\) observes the full price vector, \(\mathcal{H}_{n,i}^{in} = \{(\mathbf{p}_m, d_{m,i})\}_{m=1}^n\). No seller observes other sellers’ historical demands. Define the corresponding regressors \[x_{n,i}^{ob} = (1, p_{n,i})^\top \in \mathbb{R}^2, \qquad x_{n,i}^{in} = (1, p_{n,1},\ldots,p_{n,N})^\top \in \mathbb{R}^{N+1}.\] An oblivious seller thus estimates two parameters, while an informed seller estimates \(N+1\) parameters. For each seller \(i\), let \(\Theta_i^{ob}\subset\mathbb{R}^2\) and \(\Theta_i^{in}\subset\mathbb{R}^{N+1}\) denote known compact and convex parameter sets. At time \(n\), seller \(i\) computes the least-squares estimator and projects it onto the feasible set: \[\tilde{\theta}_{n,i}^{ob} = \argmin_{\theta\in\mathbb{R}^2}\sum_{m=1}^n [d_{m,i} - (x_{m,i}^{ob})^\top\theta]^2, \qquad \hat{\theta}_{n,i}^{ob} = \pi_{\Theta_i^{ob}}(\tilde{\theta}_{n,i}^{ob}),\] \[\tilde{\theta}_{n,i}^{in} = \argmin_{\theta\in\mathbb{R}^{N+1}}\sum_{m=1}^n [d_{m,i} - (x_{m,i}^{in})^\top\theta]^2, \qquad \hat{\theta}_{n,i}^{in} = \pi_{\Theta_i^{in}}(\tilde{\theta}_{n,i}^{in}),\] where \(\pi_{\mathcal{C}}\) denotes Euclidean projection onto the set \(\mathcal{C}\). The oblivious seller’s expected single-period revenue under price \(p\) is then estimated as \[\hat{r}_{n,i}^{ob}(p, \hat{\theta}_{n,i}^{ob}) = p \cdot \left(\hat{a}_{n,i} + \hat{b}_{n,i} p\right).\] For the informed seller, the expected single-period revenue under price \(p\) given predicted competitors’ prices \(\mathbf{p}_{-i} = (p_{j})_{j \neq i}\) is estimated as \[\hat{r}_{n,i}^{in}(p, \mathbf{p}_{-i}, \hat{\theta}_{n,i}^{in}) = p \cdot \left(\hat{\alpha}_{n,i} + \hat{\beta}_{n,i} p + \sum_{j \neq i} \hat{\gamma}_{n,i,j} p_{j}\right).\] Define the myopic revenue-maximizing prices as \[\label{eq:greedy-prices} \phi^{ob}(\theta^{ob}) = \argmax_{p \in [l,u]} \hat{r}_{n,i}^{ob}(p, \theta^{ob}), \quad \phi_i^{in}(\theta^{in}, \mathbf{p}_{-i}) = \argmax_{p \in [l,u]} \hat{r}_{n,i}^{in}(p, \mathbf{p}_{-i}, \theta^{in}).\tag{2}\] Since both revenue functions are strictly concave quadratics in own price (with \(\hat{b}_{n,i}<0\) and \(\hat{\beta}_{n,i}<0\), respectively), the unconstrained greedy prices are \[\tilde{p}_{n+1,i}^{ob} = -\frac{\hat{a}_{n,i}}{2 \hat{b}_{n,i}}, \quad \tilde{p}_{n+1,i}^{in} = -\frac{\hat{\alpha}_{n,i} + \sum_{j \neq i} \hat{\gamma}_{n,i,j} p_{n,j}}{2 \hat{\beta}_{n,i}}.\] For each oblivious seller \(i\), we assume that \(\phi^{ob}(\Theta_i^{ob}) \subset (l,u)\), so that the oblivious greedy price always lies in the feasible range; this is a standard interior assumption in dynamic pricing [13], [41], [43]. For informed sellers, we assume that the true parameters are in the interior of the feasible set, i.e., \((\alpha_i, \beta_i, \gamma_{i,1}, \ldots, \gamma_{i,i-1}, \gamma_{i,i+1}, \ldots, \gamma_{i,N})^\top \in \mathrm{int}(\Theta_i^{in})\) for all \(i\). To explore, all sellers then add a random perturbation \(z_{n+1,i}\) to the greedy price, where \(z_{n,i}\) are independent across \(n\) and \(i\) with mean zero, variance \(\nu_{n,i}^2\), and bounded support. The final price is \[p_{n+1,i} = \begin{cases} \tilde{p}_{n+1,i}^{ob} + z_{n+1,i}, & \text{if seller } i \text{ is oblivious}, \\ \tilde{p}_{n+1,i}^{in} + z_{n+1,i}, & \text{if seller } i \text{ is informed}. \end{cases}\] Price feasibility can be ensured by truncation or by enlarging the bounds \([l,u]\).

3.2.0.1 Remark.

Our model follows a standard estimate–exploit–explore paradigm that isolates the role of information and misspecification while remaining analytically tractable. We have deliberately left the dithering terms \(z_{n,i}\) unspecified, treating the exploration schedule as a strategic choice to be studied in subsequent sections. Table 5 in Appendix 8.1 consolidates the notation used throughout the paper.

Two alternative modeling choices are common in the algorithmic-collusion literature. Finite price grids paired with bandit-style or RL algorithms (Section 2) admit no-regret guarantees but discard the parametric demand structure, and existing theoretical analyses are often confined to two-price grids. The (multinomial) logit demand model [25], [27], [36] differs informationally: every successful sale by one seller is a sale a competitor did not make, so sellers observe a partial signal about rival demand from their own data, whereas our linear-demand model with unknown intercept offers no such handle and is therefore strictly harder.

3.3 Solution Concepts↩︎

We benchmark learning dynamics against two full-information outcomes: the competitive outcome (Nash equilibrium) and the collusive outcome (cartel pricing). In a pure-strategy Nash equilibrium (NE), each seller best responds to competitors’ prices. The unconstrained equilibrium price vector \(\mathbf{p}^{NE}\) satisfies \[p_i^{NE} = \argmax_{p_i\in\mathbb{R}} p_i\!\left(\alpha_i + \beta_i p_i + \sum_{j\neq i}\gamma_{i,j}p_j^{NE}\right), \qquad i\in[N].\] Let \(\Gamma\) denote the \(N\times N\) matrix with \(\Gamma_{ii}=2\beta_i\) and \(\Gamma_{ij}=\gamma_{i,j}\) for \(i\neq j\), and let \(\boldsymbol{\alpha}=(\alpha_1,\ldots,\alpha_N)^\top\). Since \(-\beta_i>\gamma_i\) for all \(i\), \(\Gamma\) is strictly diagonally dominant with negative diagonal, hence invertible, yielding the unique equilibrium \[\Gamma \mathbf{p}^{NE} = -\boldsymbol{\alpha}, \qquad \mathbf{p}^{NE} = -\Gamma^{-1}\boldsymbol{\alpha}.\] On the other hand, the unconstrained collusive outcome maximizes total expected revenue: \[\mathbf{p}^{C} = \argmax_{\mathbf{p}\in\mathbb{R}^N} \sum_{i=1}^N p_i\!\left(\alpha_i + \beta_i p_i + \sum_{j\neq i}\gamma_{i,j}p_j\right).\] The first-order conditions yield the \(N\times N\) symmetric matrix \(H\) with \(H_{ii}=2\beta_i\) and \(H_{ij}=\gamma_{i,j}+\gamma_{j,i}\) for \(i\neq j\). A unique collusive outcome \[\mathbf{p}^{C} = -H^{-1}\boldsymbol{\alpha}\] exists whenever \(H\) is negative definite, a standard requirement in differentiated-demand oligopoly models [44][46]. Define \(\gamma_i^{\mathrm{col}}\triangleq \sum_{j\neq i}\gamma_{j,i}\), the total cross-price effect of seller \(i\)’s price on all other sellers’ demands. For ease of exposition, we assume the slightly stronger diagonal-dominance condition \(-2\beta_i>\gamma_i+\gamma_i^{\mathrm{col}}\) for all \(i\), which is sufficient for \(H\prec 0\) and additionally renders both \(-\Gamma\) and \(-H\) into \(M\)-matrices, so that the clean component-wise ordering \(\mathbf{p}^{C}\ge\mathbf{p}^{NE}\) holds and the discussion of competition vs. collusion is cleaner.1 To make the discussion of competitive and collusive outcomes meaningful, we also assume that \(\mathbf{p}^{NE},\mathbf{p}^{C}\in(l,u)^N\).

3.3.0.1 Remark.

The competitive outcome can be interpreted as an indication of market efficiency. Under this outcome, no seller can unilaterally improve their revenue by deviating from the price vector and consumer welfare is preserved. The collusive outcome, on the other hand, maximizes the total revenue of the sellers. It represents the case that sellers act as if they have formed a cartel and charge the monopoly price, which negatively impacts consumers because \(\mathbf{p}^{C} \ge \mathbf{p}^{NE}\) component-wise.2

3.4 Performance Metric↩︎

3.4.0.1 Price convergence.

Much of the competitive dynamic pricing literature evaluates algorithms by their regret—the cumulative revenue gap relative to an oracle that best-responds to the realized competitor profile at every step [30], [37], [49], [50]. Formally, write \(R_i(p_i, \mathbf{p}_{-i}) \triangleq p_i(\alpha_i + \beta_i p_i + \sum_{j\neq i}\gamma_{i,j} p_j)\) for seller \(i\)’s expected per-period revenue and \(\phi_i^{in}(\theta_i, \mathbf{p}_{-i}) \triangleq (\alpha_i + \sum_{j\neq i}\gamma_{i,j} p_j)/(-2\beta_i)\) for the full-information best response. The regret of seller \(i\) over horizon \(T\) is \[\label{eq:dynamic-regret-main} \Delta_i(\theta_i, T) \;\triangleq\; \sum_{n=1}^{T} \EE_{n-1}\!\left[R_i\!\left(\phi_i^{in}(\theta_i, \mathbf{p}_{n,-i}), \mathbf{p}_{n,-i}\right) - R_i\!\left(p_{n,i}, \mathbf{p}_{n,-i}\right)\right].\tag{3}\] A standard workflow in monopolistic dynamic pricing often rewrites \(\Delta_i\) as the cumulative squared distance from the per-period price to a fixed reference price—the monopoly optimum [13]. The same reformulation applied in our competitive setting (cf. Appendix 8.3) uncovers the reference to be the Nash equilibrium: \[\label{eq:regret-equiv-main} \Delta_i(\theta_i, T) \;\lesssim\; \sum_{n=1}^{T} \EE_{n-1}\!\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2, \qquad \sum_{i=1}^{N} \Delta_i(\theta_i, T) \;\asymp\; \sum_{n=1}^{T} \EE\,\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2.\tag{4}\] Minimizing regret is therefore equivalent to driving prices toward the Nash equilibrium. In particular, sellers who sustain a collusive outcome and earn higher revenues than at \(\mathbf{p}^{NE}\) nevertheless incur linear regret, revealing an inherent tension between the benchmark and the phenomenon of interest. Accordingly, our analysis foregrounds price convergence—characterizing where prices converge to (\(\mathbf{p}^{NE}\), \(\mathbf{p}^C\), or elsewhere) and at what rate—and discusses implications for regret as a complement.

3.4.0.2 Surplus-capture ratio.

To normalize realized time-averaged revenue across markets with different demand parameters, we follow the algorithmic-collusion literature [4] and define seller \(i\)’s surplus-capture ratio along a price path \(\{\mathbf{p}_n\}_{n\ge 1}\) by \[\label{eq:surplus-capture} S_i \;\triangleq\; \liminf_{n\to\infty}\; \frac{\frac{1}{n}\sum_{m=1}^{n} p_{m,i}\, d_{m,i} \;-\; \Pi_i^{NE}}{\Pi_i^{C} \;-\; \Pi_i^{NE}},\tag{5}\] where \(d_{m,i}\) is the realized demand at time \(m\), and the seller-specific per-period revenues at the competitive and collusive outcomes are \[\Pi_i^{NE} \;\triangleq\; p_i^{NE}\!\left(\alpha_i + \beta_i p_i^{NE} + \sum_{j\neq i}\gamma_{i,j} p_j^{NE}\right), \qquad \Pi_i^{C} \;\triangleq\; p_i^{C}\!\left(\alpha_i + \beta_i p_i^{C} + \sum_{j\neq i}\gamma_{i,j} p_j^{C}\right).\] By construction, \(S_i = 0\) means that seller \(i\) exactly matches their Nash benchmark on average, \(S_i = 1\) means that they match their collusive benchmark, and the super-collusive case \(S_i > 1\) and the sub-Nash case \(S_i < 0\) are both possible. The ratio is unit-free and handles asymmetric markets transparently, since \(\Pi_i^{NE}\) and \(\Pi_i^{C}\) are both well-defined seller-by-seller from \(\mathbf{p}^{NE}\) and \(\mathbf{p}^{C}\). We will use \(S_i\) as a metric for revenue in both theoretical and numerical analyses.

4 Price Exploration in Oblivious Modeling↩︎

In monopolistic dynamic pricing with unknown demand, the role of price exploration is relatively well understood; see, e.g., [35] and [13]. In contrast, under oblivious demand modeling, estimation is misspecified because competitor prices are omitted, affecting both what can be learned and what exploration is needed for stable performance. Rather than imposing an ad hoc exploration rule, in this section we adopt the perspective of an oblivious seller and ask: what exploration behavior is defensible from first principles? The answer will serve as a baseline behavioral prediction for oblivious sellers and as the foundation for our market-level analysis. To set the stage, we first introduce some key quantities.

Fix an oblivious seller \(i\). Define their cumulative exploration up to time \(n\) by \[J_{n,i} \;\triangleq\; \sum_{m=1}^n \left(p_{m,i} - \bar p_{n,i}\right)^2 \;=\; \sum_{m=2}^{n} \left(1 - \frac{1}{m}\right)\left(p_{m,i} - \bar p_{m-1,i}\right)^2,\] where \(\bar p_{n,i}\triangleq \frac{1}{n}\sum_{m=1}^n p_{m,i}\). The growth rate of \(J_{n,i}\) captures how much variation seller \(i\) injects into their own prices, and we use exploration rate to refer to the asymptotic growth of \(J_{n,i}\) (e.g., linear vs.sublinear) as a function of \(n\). We assume throughout that the first two prices are not equal so that \(J_{n,i}>0\) for all \(n\ge 2\). Two more auxiliary quantities will be used repeatedly in what follows: \[w_{n,i} \;\triangleq\; \frac{1}{J_{n,i}} \sum_{m=1}^n (p_{m,i} - \bar p_{n,i})\,\varepsilon_{m,i}, \qquad r_{n,i\leftarrow j} \;\triangleq\; \frac{1}{J_{n,i}} \sum_{m=1}^{n} (p_{m,i} - \bar p_{n,i})(p_{m,j} - \bar p_{n,j}).\] The term \(w_{n,i}\) is a self-noise term capturing the correlation between seller \(i\)’s price deviations and idiosyncratic demand shocks. The term \(r_{n,i\leftarrow j}\) captures the omitted-variable channel: it is a normalized empirical cross-covariance between seller \(i\)’s prices and seller \(j\)’s prices, and it determines how competitor price variation contaminates seller \(i\)’s misspecified regression.

To make the dependence on \((w_{n,i}, r_{n,i\leftarrow j})\) explicit, we write each oblivious parameter set as a rectangle \(\Theta_i^{ob}=[\underline{a}_i,\bar a_i]\times[\underline{b}_i,\bar b_i]\) with \(0<\underline a_i<\bar a_i\) and \(\underline b_i<\bar b_i<0\). A direct least-squares calculation and projection yields, for all \(n\ge 2\), \[\hat{a}_{n,i} = \pi_{[\underline a_i,\bar a_i]}\!\left( \alpha_i + \sum_{j\neq i}\gamma_{i,j}\left(\bar p_{n,j} - r_{n,i\leftarrow j}\bar p_{n,i}\right) + \bar\varepsilon_{n,i} - w_{n,i}\bar p_{n,i} \right),\] \[\hat{b}_{n,i} = \pi_{[\underline b_i,\bar b_i]}\!\left( \beta_i + \sum_{j\neq i}\gamma_{i,j}\,r_{n,i\leftarrow j} + w_{n,i} \right),\] where \(\bar\varepsilon_{n,i}\triangleq \frac{1}{n}\sum_{m=1}^n \varepsilon_{m,i}\). Consequently, the oblivious seller’s greedy (conditional-mean) price at time \(n+1\) is \[\label{eq:obliviousgreedynextprice} \tilde{p}_{n+1,i} = \frac{ \pi_{[\underline a_i,\bar a_i]}\!\left( \alpha_i + \sum_{j\neq i}\gamma_{i,j}\left(\bar p_{n,j} - r_{n,i\leftarrow j}\bar p_{n,i}\right) + \bar\varepsilon_{n,i} - w_{n,i}\bar p_{n,i} \right) }{ -2\, \pi_{[\underline b_i,\bar b_i]}\!\left( \beta_i + \sum_{j\neq i}\gamma_{i,j}\,r_{n,i\leftarrow j} + w_{n,i} \right) }.\tag{6}\] This expression isolates two distinct sources of instability for oblivious learning: idiosyncratic demand noise via \(w_{n,i}\) and \(\bar\varepsilon_{n,i}\), and strategic interaction via the cross terms \(r_{n,i\leftarrow j}\).

4.1 Baseline: Divergent Exploration Rate↩︎

We first examine the role of demand noise. By the strong law of large numbers, \(\bar\varepsilon_{n,i}\to 0\) a.s., so the relevant term is \(w_{n,i}\). Intuitively, \(w_{n,i}\) should vanish under effective learning, otherwise idiosyncratic shocks persistently distort demand estimates and pricing decisions. The next theorem formalizes this link between exploration and noise attenuation; the proof is in Appendix 9.1.

Theorem 1 (Impact of demand noise). As the horizon \(n \to \infty\), the following holds:

  1. On the event that \(\frac{J_{n,i}}{\log n \log \log n} \rightarrow \infty\), we have \(w_{n,i} \rightarrow 0\).

  2. On the event that \(\sum_{m=2}^{\infty} \left(p_{m,i} - \bar p_{m-1,i}\right)^2 (\log m)^{1+\delta} < \infty\) for some \(\delta>0\), we have \(w_{n,i} \rightarrow w_{\infty,i}\) for some a.s.finite random variable \(w_{\infty,i}\).

4.1.0.1 Discussion.

Since \(J_{n,i}\) is nondecreasing, each sample path satisfies either \(J_{n,i}\to\infty\) or \(J_{n,i}\to J_{\infty,i}<\infty\), and the two cases of Theorem 1 correspond to these scenarios up to log factors: idiosyncratic demand noise becomes asymptotically negligible only if exploration is unbounded, otherwise the long-run price (and revenue) depends on the random limit \(w_{\infty,i}\) and is sample-path dependent. This mirrors the well-known failure of certainty-equivalent control in monopolistic dynamic pricing [13], [51]: competition does not automatically save an oblivious seller, as \(w_{n,i}\) vanishes only if the seller’s own prices exhibit sufficient dispersion. We henceforth impose the mild condition \[\frac{J_{n,i}}{\log n \log \log n} \longrightarrow \infty\] for all oblivious sellers, so that \(w_{n,i}\) is asymptotically negligible. By Lemma 8, this can be achieved by adding perturbations with cumulative variance diverging faster than \(\log n \log\log n\) (e.g., \(\mathrm{Var}(z_{n,i})\asymp n^{-c}\) for any \(c\in[0,1)\)).

4.2 The “Spiral-up” Phenomenon: Linear Exploration Rate↩︎

We now turn to the cross-seller terms \(r_{n,i\leftarrow j}\), which mediate misspecification across sellers and, unlike \(w_{n,i}\), need not vanish under divergent exploration. The weighted aggregate \(\sum_{j\neq i}\gamma_{i,j}\,r_{n,i\leftarrow j}\) enters the greedy update directly, and it is straightforward to verify that setting all \(r_{n,i\leftarrow j}\) to zero recovers the Nash prices. Therefore, \(r_{n,i\leftarrow j}\) captures the distortion of seller \(i\)’s price due to misspecification. See Appendix 8.4 for more interpretations.

In this section, we ask what controls \(r_{n,i\leftarrow j}\), and what an oblivious seller’s rational response to that control looks like. We begin with a sufficient condition under which seller \(j\) is asymptotically immune to seller \(i\)’s variation: seller \(j\) must explore at a larger scale. The proof is in Appendix 9.2.

Lemma 1 (Variance dominance). Fix sellers \(i \neq j\). On the event that \(J_{n,i}/J_{n,j} \longrightarrow 0\), we have \(r_{n,\, j \leftarrow i} \longrightarrow 0\).

When \(J_{n,i}/J_{n,j}\to 0\), seller \(j\)’s price dispersion dwarfs seller \(i\)’s, so in \(j\)’s misspecified regression seller \(i\)’s price behaves asymptotically like a constant and the omitted-variable impact vanishes. This creates an informational asymmetry: \(j\) is effectively well-specified with respect to \(i\), while \(i\) may still suffer persistent misspecification induced by \(j\)’s variation. To sharpen the revenue implication, let \(N=2\) with sellers indexed by \(\{1,2\}\) and, for each \(i\neq j\), define the full-information best-response function \(\phi_i(p)\triangleq \tfrac{\alpha_i + \gamma_{i,j} p}{-2\beta_i}\). The Nash equilibrium prices \((p_1^{NE},p_2^{NE})\) satisfy \(p_i^{NE}=\phi_i(p_j^{NE})\) for \(i\neq j\). The next proposition characterizes the long-run behavior under variance dominance; the proof is in Appendix 9.3.

Let \(i,j\in\{1,2\}\) with \(i\neq j\), and write the two corner prices \(\ell_{-,i}\triangleq \underline{a}_i/(-2\bar{b}_i)\) and \(\ell_{+,i}\triangleq \bar{a}_i/(-2\underline{b}_i)\) defined by the parameter projection box \(\Theta_i^{ob}\). On the event \(J_{n,i}/J_{n,j} \rightarrow 0\):

  1. Every subsequential limit \(\tilde{p}_{\infty,i}\) of \(\tilde{p}_{n,i}\) belongs to \(\{\ell_{-,i},\,\ell_{+,i},\, p_i^{NE}\}\), and along that subsequence \(\tilde{p}_{n,j}\) converges to the corresponding best response \(\phi_j(\tilde{p}_{\infty,i})\).

  2. Suppose in addition that the parameters are symmetric at \((\alpha, \beta, \gamma)\) and that \(J_{n,k} = o(n)\) a.s.for \(k\in\{i,j\}\). Define \[q_- \;\triangleq\; \frac{\alpha}{2|\beta|+\gamma}\;<\; p^{NE}\;=\;\frac{\alpha}{2|\beta|-\gamma}.\] If the dominated seller’s projection box \(\Theta_i^{ob}\) satisfies \[\label{eq:box-strict-dominance} \ell_{-,i}\;<\;q_- \qquad\text{and}\qquad \ell_{+,i}\;>\;p^{NE},\tag{7}\] then the surplus-capture ratios satisfy \(S_j \ge S_i\) a.s. Moreover, on any sample path along which \(\tilde{p}_{n,i}\) converges to a limit \(\tilde{p}_{\infty,i}\in\{\ell_{-,i},\ell_{+,i},p_i^{NE}\}\): \[S_i \;=\; S_j \;=\; 0 \quad\text{if} \quad\tilde{p}_{\infty,i} = p_i^{NE}, \qquad S_i \;<\; S_j \quad\text{if} \quad\tilde{p}_{\infty,i} \in \{\ell_{-,i},\ell_{+,i}\}.\]

4.2.0.1 Discussion.

Part (a) says the variance-dominated seller locks at a boundary-induced price or at \(p_i^{NE}\) in a sample-path-dependent way while the dominant seller asymptotically best-responds. Part (b) sharpens this into a surplus-capture ordering: under condition 7 , \(S_j \ge S_i\) a.s., with \(S_j = S_i = 0\) on every Nash convergent sample path and \(S_j > S_i\) on every non-Nash one.

4.2.0.2 Numerical experiments.

Condition 7 is a non-trivial sufficient assumption: the proof in Appendix 9.3 shows that a dominated seller locked inside \((q_-, p^{NE})\) could out-capture the dominant. Table 2 stress-tests the conclusion by placing \(\ell_{-,i}\) inside the reverse regime—violating 7 —yet the dominant still out-captures the dominated by \(S_j - S_i \in [0.79, 0.96]\) across every dominance cell, suggesting that the ordering \(S_j > S_i\) is robust.3

Table 2: Variance dominance in a symmetric oblivious–oblivious duopoly under a stress-test projection box that violates the sufficient condition 7 of Proposition [prop:variancedominancetwosellercase](b). Both sellers use polynomial exploration schedules \(\nu_{n,k}^2 = 0.05\,(n+1)^{-\eta_k}\); the dominant seller’s exponent is fixed at \(\eta_2 = 0.01\), and the dominated’s exponent \(\eta_1 \in \{0.5, 0.7, 1.0, 1.5\}\). Columns report the cross-regression ratios \(r_{T,j\leftarrow i}\), running-mean prices \(\bar p_{T,i}\), and surplus-capture ratios \(S_i\). Lemma 1’s prediction that \(r_{T,2\leftarrow 1}\) should be smaller than \(r_{T,1\leftarrow 2}\) in magnitude holds throughout; the surplus-capture ordering \(S_2 > S_1\) is preserved with margin \(S_2 - S_1 \in[0.79, 0.96]\), confirming Proposition [prop:variancedominancetwosellercase](b) empirically even when its sufficient condition fails. Implementation details are in Appendix 8.2.
\(\eta_1\) \(r_{T,2\leftarrow 1}\) \(r_{T,1\leftarrow 2}\) \(\bar p_{T,1}\) \(\bar p_{T,2}\) \(S_{1}\) \(S_{2}\)
\(0.5\) \(0.14\) \(0.95\) \(2.45\) \(2.09\) \(+0.00\) \(+0.96\)
\(0.7\) \(0.18\) \(0.92\) \(2.46\) \(2.12\) \(+0.04\) \(+0.96\)
\(1.0\) \(0.24\) \(0.92\) \(2.47\) \(2.15\) \(+0.10\) \(+0.96\)
\(1.5\) \(0.26\) \(0.85\) \(2.44\) \(2.17\) \(+0.11\) \(+0.90\)

4.2.0.3 “Spiral-up” phenomenon and a behavioral prediction.

Taken together, these results suggest that empirical price variance acts as a strategic resource in oblivious markets: Lemma 1, Proposition [prop:variancedominancetwosellercase], and Table 2 jointly argue that sellers with higher price variance can asymptotically out-capture their competitors in revenue terms. Rational sellers should not accept this structural disadvantage, so each oblivious seller has an incentive to ensure they are not the variance-dominated one, and the only sample-path-stable way to do so is to keep their own price variance growing at least at the same order as the competitor’s. This creates a natural spiral-up logic: if sellers attempt to out-explore one another, rates escalate to the maximal sustainable order—linear, \[J_{n,i} = \Theta(n),\] since any \(\omega(n)\) growth would require unbounded fluctuations and persistent clipping. By Lemma 8, a linear rate is implementable with controllably small non-diminishing perturbations (e.g., \(z_{n,i}\sim\mathrm{Unif}(-c,c)\) for \(c>0\)), at a persistent exploration tax of \(|\beta|\nu^2\) per period (Proposition [prop:costlinearexploration], Appendix 8.5). With the small \(\nu^2\) values used in our experiments, this tax is small in \(S\) units relative to the surplus-capture gap of \(\approx 1\) in Table 2, so linear exploration functions as cheap “insurance” for an oblivious seller. The rate is undesirable in monopolistic pricing but emerges as a rational choice in competition; whether this structural inefficiency is compensated by collusion gains is the focus of the next section.

5 Outcomes in a Market of Oblivious Sellers↩︎

We now turn to market outcomes when all sellers are oblivious. The central question is whether oblivious learning (misspecified) can produce long-term supracompetitive, collusive-like outcomes. Our answer is no: under reasonable exploration, the competitive outcome is the dominant attractor.

5.1 Global Convergence to the Competitive Outcome↩︎

Even though an oblivious seller estimates a misspecified two-parameter model, at the competitive benchmark it admits a natural “pseudo-true” target. Define \[\theta_i^{*, \,ob} \triangleq (a_i^*,b_i^*) \;=\; \left(\alpha_i + \sum_{j\neq i}\gamma_{i,j}p_j^{NE},\;\beta_i\right),\] and assume \(\theta_i^{*, \,ob} \in \mathrm{int}(\Theta_i^{ob})\) for all \(i\). Note that \(\phi^{ob}(\theta_i^{*,\,ob}) = p_i^{NE}\) by the Nash first-order condition. Let \(L_\phi^{ob}\) denote the Lipschitz constant of the oblivious pricing map \(\phi^{ob}(\cdot)\) over \(\bigcup_i\Theta_i^{ob}\), and let \(C_x\) be a uniform bound on \(\|x_{n,i}^{ob}\|_2\). Let \(\mathcal{F}_n \triangleq \sigma\{(p_{m,i},d_{m,i}) : m\le n,\;i\in[N]\}\) be the filtration generated by the entire price-demand history up to time \(n\). Recall that \(\gamma_i = \sum_{j\neq i}\gamma_{i,j}\) and \(\gamma_i^{\mathrm{col}} = \sum_{j\neq i}\gamma_{j,i}\) (cf.Sections 3.1 and 3.3). Define \[\bar\gamma \triangleq \frac{1}{2} \left(\max_i \gamma_i + \max_i \gamma_i^{\mathrm{col}}\right).\] The next theorem establishes global convergence to the competitive outcome when exploration is sufficiently strong relative to strategic misspecification; its proof is deferred to Appendix 9.4.

Theorem 2 (Global convergence to the competitive outcome). Suppose there exists \(C_M>0\) such that for all \(n\) and \(i\), \[\label{eq:globalconvergencespectrallowerboundcondition} \mathbb{E}\!\left[x_{n,i}^{ob}(x_{n,i}^{ob})^\top\mid\mathcal{F}_{n-1}\right] \succeq C_M I_{2}.\qquad{(1)}\] If, in addition, \[\label{eq:globalconvergencecondition} \bar\gamma\, L_\phi^{ob}\, C_x \;<\; C_M,\qquad{(2)}\] then \(\hat{\theta}_{n,i}^{ob}\to \theta_i^{*,ob}\) and \(\tilde{p}_{n,i}\to p_i^{NE}\) a.s.for every \(i\in[N]\). Moreover, the aggregate mean-squared error satisfies \[\sum_{i=1}^{N}\left[ \mathbb{E}\|\hat{\theta}_{n,i}^{ob}-\theta_i^{*, \,ob} \|_2^2 + \mathbb{E}|\tilde{p}_{n,i}-p_i^{NE}|^2 \right] = \begin{cases} O\!\left(1 / n\right), & \text{if } 2\bar\gamma L_\phi^{ob} C_x < C_M,\\[6pt] O\!\left(\log n / n\right), & \text{if } 2\bar\gamma L_\phi^{ob} C_x = C_M,\\[6pt] O\!\left(n^{-2(1-\rho)}\right), & \text{if } \bar\gamma L_\phi^{ob} C_x < C_M < 2\bar\gamma L_\phi^{ob} C_x, \end{cases}\] where \(\rho \triangleq \frac{\bar\gamma L_\phi^{ob} C_x}{C_M}\in(1/2,1)\) in the last case.

5.1.0.1 Discussion.

Theorem 2 uncovers a misspecification–exploration tug-of-war. The quantity \(\bar\gamma L_\phi^{ob} C_x\) in ?? upper-bounds the strategic feedback loop created by omitted competitor prices (\(\bar\gamma\) scales the worst-case cross-price interaction, while \(L_\phi^{ob}\) and \(C_x\) measure the oblivious best-response sensitivity to estimation errors), and \(C_M\) measures exploration strength via the smallest eigenvalue of the conditional regressor covariance. When exploration dominates misspecification, the induced bias behaves like a controlled perturbation and the dynamics contract toward \(\mathbf{p}^{NE}\); the rates accelerate from \(n^{-2(1-\rho)}\) to \(1/n\) as \(C_M / (\bar\gamma L_\phi^{ob} C_x)\) grows. The persistent-excitation condition ?? is standard [52], [53], and by Lemmas 8 and 9 it is implied by any uniformly bounded-below exploration variance \(\Var(z_{n,i})\), i.e., a linear exploration rate \(J_{n,i} = \Theta(n)\). Since this is exactly the spiral-up prediction of Section 4.2, ?? is best read as a behavioral prediction rather than a technical assumption.4

5.1.0.2 Connection to regret.

Combining Theorem 2 with the dynamic-benchmark equivalence \(\sum_i \Delta_i(\theta_i, T) \asymp \sum_n \EE\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2\) (Corollary 1 in Appendix 8.3) and the decomposition \(\mathbf{p}_n = \tilde{\mathbf{p}}_n + \mathbf{z}_n\), the misspecification-induced portion of regret is at most \(O(\log T)\) in the fastest regime while the persistent-exploration term contributes a \(\Theta(T)\) tax. Though undesirable, this \(\Theta(T)\) rate is qualitatively different from the linear regret of an algorithm that fails to learn: the tax is controlled by the seller (Proposition [prop:costlinearexploration]), whereas linear regret from learning failure is pinned by an irreducible bias.

a
b

Figure 2: Empirical stress test of Theorem 2: seed-averaged \(\mathrm{MSE}(\tilde{\mathbf{p}}_n)\) on log–log axes. Left: asymmetric markets at \(N \in \{3, 5, 10\}\); right: symmetric \(N = 5\) markets sweeping the cross-price coefficient \(\gamma\); both panels sweep exploration variance \(\nu^2 \in \{0.05, 0.10, 0.20\}\). Across all 18 configurations the sufficient condition ?? is violated by roughly two orders of magnitude, yet every \(\mathrm{MSE}\) trajectory decays at the asymptotic \(n^{-1}\) rate. Implementation details are in Appendix 8.2.. a — Asymmetric markets, \(N\in\{3,5,10\}\)., b — Symmetric \(N=5\), shrinking \(\gamma\).

5.1.0.3 Numerical experiments.

Figure 2 stress-tests Theorem 2 over an asymmetric \(N \in \{3, 5, 10\}\) sweep (left panel) and a symmetric \(N=5\) sweep over the cross-price coefficient \(\gamma\) (right panel), each crossed with three exploration variances \(\nu^2 \in \{0.05, 0.10, 0.20\}\). Across all \(18\) configurations, \(\bar\gamma\, L_\phi^{ob}\, C_x\) is of order \(1\)\(10\) while \(C_M(\nu^2)\) stays below \(10^{-1}\), so condition ?? is violated by roughly two orders of magnitude; yet every \(\mathrm{MSE}(\tilde{\mathbf{p}}_n)\) trajectory decays at the asymptotic \(n^{-1}\) rate predicted in the strongest regime. The contraction condition is therefore conservative but not necessary in practice. Appendix 8.6 adds a 16-point \(\nu^2\) sweep that pins down the empirical convergence threshold and a Gaussian-clipped dithering robustness check (as opposed to the uniform perturbations used throughout the numerical experiments), both of which reinforce this conclusion.

5.2 A Mean-Dynamics ODE Perspective and Local Convergence↩︎

When ?? fails, the discrete-time learning dynamics need not be globally contractive. Yet numerical experiments suggest that convergence to the competitive outcome still occurs in practice. To understand this phenomenon, we next analyze the learning dynamics from a complementary mean-dynamics perspective that isolates the local deterministic drift induced by the estimate–exploit–explore protocol. We track the running averages of prices and pairwise price products—the mean-price vector \(m(t)\in\RR^N\) and the second-moment matrix \(Q(t)\in\RR^{N\times N}\)—under i.i.d.exploration \(z_{n,i}\) with \(\Var(z_{n,i})=\nu^2>0\). A standard stochastic-approximation derivation [54], detailed in Appendix 8.7, yields a \((2N+N(N-1)/2)\)-dimensional mean-dynamics ODE: \[\label{eq:ode-N-mQ} \begin{align} \dot{m}_i(t) &= p_i^g(m(t),Q(t)) - m_i(t), \qquad i\in[N],\\ \dot{Q}_{ij}(t) &= p_i^g(m(t),Q(t))\,p_j^g(m(t),Q(t)) - Q_{ij}(t),\qquad i\neq j,\\ \dot{Q}_{ii}(t) &= (p_i^g(m(t),Q(t)))^2 + \nu^2 - Q_{ii}(t),\qquad i\in[N], \end{align}\tag{8}\] where \(p_i^g(m,Q)\) is the misspecified greedy price for seller \(i\) given the moments \((m,Q)\) (explicit form in Appendix 8.7). Heuristically, 8 is the deterministic continuous-time skeleton of the discrete moment recursion, analogous in spirit to the fluid limit in queueing applications. The result below shows that 8 admits a unique equilibrium, whose mean component coincides with the full-information Nash equilibrium \(\mathbf{p}^{NE}\), and that this equilibrium is locally asymptotically stable. The proof is in Appendix 9.5.

Theorem 3 (Equilibrium point and local convergence). The ODE 8 admits a unique equilibrium \((m^*,Q^*)\) that satisfies \[m^*=\mathbf{p}^{NE}, \qquad Q^* = m^*(m^*)^\top + \nu^2 I_N.\] Moreover, \((m^*,Q^*)\) is locally asymptotically stable.

5.2.0.1 Discussion.

Theorem 3 says that, within the regime where the mean-dynamics approximation is accurate, the competitive outcome is the only plausible limit and is locally stable. Together with Theorem 2, this confirms \(\mathbf{p}^{NE}\) as a robust local attractor even when global small-gain conditions fail.

5.3 Collusive Excursions and the Continuum of Pseudo-Equilibria↩︎

Section 1.2 highlighted two puzzling features of an all-oblivious market: possible collusive behaviors [2], [11], [12], and the wide spread of long-run sample-path prices ranging from near-competitive to near-collusive (Figure 1). The mean-dynamics ODE 8 explains both: within the persistent-exploration regime that drives convergence to \(\mathbf{p}^{NE}\), the system admits striking finite-time excursions (upward and collusive-looking, or downward); and when persistent excitation fails, sample paths trace out a continuum of pseudo-equilibria, qualitatively matching the noiseless and exploration-free analysis of [2].

5.3.1 Short-Run Excursions↩︎

The same ODE 8 that locks the long-run mean to \(\mathbf{p}^{NE}\) admits striking transients along the way. Specializing to the symmetric duopoly, Figure 3 displays two types of trajectories of 8 and their discrete-time analogues: an upward excursion (panels (a),(b)) in which the mean price overshoots \(\mathbf{p}^{NE}\) toward the collusive benchmark \(\mathbf{p}^C\) before relaxing back, and a downward excursion (panels (c),(d)) in which the mean price first dips below \(\mathbf{p}^{NE}\) before recovering. Appendix 8.8 formalizes the governing forces behind these phenomena and shows that the direction of any such excursion is sample-path dependent and not controllable by the sellers. Consequently, observable collusive behaviors are random finite-time transients that oblivious learning cannot reliably induce to earn supracompetitive profits.

Figure 3: Positive (panels (a), (b)) and negative (panels (c), (d)) excursions of the mean price in the symmetric duopoly, each shown for the five-dimensional ODE 15 and the discrete-time dynamics 6 . The y-axis is average price; red and green dotted lines mark p^{NE} and p^C. ODE panels use a linear time scale; discrete-time panels plot the rolling empirical mean of each seller on a logarithmic n-axis. Implementation details are in Appendix 8.2.

5.3.2 Diminishing Exploration↩︎

When exploration decays over time (\(\Var(z_{n,i})\downarrow 0\)), persistent excitation can fail (cf.Lemmas 8 and 9) and the ODE drift can become ill-defined, and so neither Theorem 2 nor Theorem 3 forces convergence to \(\mathbf{p}^{NE}\). In the duopoly \(N=2\), this connects directly to the deterministic/no-exploration analysis of [2]. Define the empirical regression ratios \[r_{n,1}\triangleq \frac{Q_{n,12}-m_{n,1}m_{n,2}}{Q_{n,11}-m_{n,1}^2}, \qquad r_{n,2}\triangleq \frac{Q_{n,12}-m_{n,1}m_{n,2}}{Q_{n,22}-m_{n,2}^2}.\] Conditional on a fixed pair \((r_1,r_2)\), the duopoly’s misspecified greedy conditions \(m_i=-a_i/(2b_i)\) reduce to the linear system \[\label{eq:cooper-pseudo-equilibrium} (2\beta_1+\gamma_{1,2} r_1)m_1+\gamma_{1,2} m_2=-\alpha_1, \qquad \gamma_{2,1} m_1+(2\beta_2+\gamma_{2,1} r_2)m_2=-\alpha_2,\tag{9}\] whose solution yields a continuum of candidate pseudo-equilibria indexed by \((r_1,r_2)\), mirroring [2].

Figure 4: Running-mean prices (panels (a), (c)) and per-period revenues (panels (b), (d)) of 1{,}800 sample paths in the symmetric duopoly under two decaying-exploration schedules \nu_n^2 = 0.3(n+1)^{-\eta}: \eta = 0.5 (panels (a), (b)) and \eta = 0.85 (panels (c), (d)). Shading in the price panels is the theoretical region induced by the admissible regression-ratio set, and its image under the revenue map in the revenue panels. Red and green markers are \mathbf{p}^{NE}, \mathbf{p}^{C} (and their revenue images). A non-trivial fraction of seeds earns less than \Pi^{NE} (33\% at \eta = 0.5, 19\% at \eta = 0.85). Implementation details are in Appendix 8.2.

Figure 4 probes this continuum in a symmetric duopoly with long-run sample paths from diverse warm-up prices, under two decaying-exploration schedules \(\nu_n^2 = 0.3(n+1)^{-\eta}\) with \(\eta\in\{0.5, 0.85\}\), overlaid on the theoretical admissible regression-ratio region (Appendix 8.2). The empirical clouds fill two-dimensional regions reaching \(\mathbf{p}^{NE}\), \(\mathbf{p}^{C}\), and points outside the bounding rectangle, reproducing the insight of [2] that the limit can be competitive, collusive, intermediate, or random.

In contrast to Theorems 2 and 3, this continuum is the manifestation of incomplete learning [51] in a competitive setting: persistent exploration facilitates identification of the competitive equilibrium, and without it learning fails and long-run outcomes scatter across an entire continuum of random prices. In line with this picture, the smaller-variance schedule, though injecting less exploration noise, produces the wider cross-seed cloud, as smaller exploration variance lets early drifts freeze, whereas larger exploration variance keeps the persistent-excitation regime active for more iterations and allows the Nash attractor to undo collusive drifts. This closes the loop with Figure 1: under the canonical \(\Theta(\sqrt n)\) rate, the design becomes nearly singular as exploration fades and different sample paths freeze at different points of the pseudo-equilibrium continuum.

6 Markets with Informed Sellers↩︎

So far, we have analyzed the market when all sellers are oblivious. We now turn to the case when at least one seller is informed. We first analyze the all-informed market as an ideal baseline, and then we study mixed markets in which a strict subset of sellers is oblivious and the rest informed.

6.0.0.1 Heterogeneous exploration rates.

To allow for heterogeneous exploration behavior, let \(\nu_{n,i}^2=\Var(z_{n,i})\) denote seller \(i\)’s exploration variance at time \(n\), and throughout this section suppose that \[\nu_{n,i}^2=\Theta\!\left(n^{-\eta_i}\right), \qquad 0\le \eta_i<1, \qquad \eta_{\min}=\min_{i\in[N]}\eta_i, \qquad \eta_{\max}=\max_{i\in[N]}\eta_i.\] This polynomial rate assumption is not necessary for the results, but it allows for a clean ordering of exploration rates among sellers and is thus adopted for expositional clarity. Since informed sellers do not suffer from model misspecification, we further assume that their exploration decays, i.e., \(\eta_j > 0\) for every informed seller \(j \in \cI^{in}\). Without loss of generality, assume that the first \(N+1\) prices induce a full-rank empirical Fisher information matrix for each informed seller, allowing least squares thereafter.

6.0.0.2 Forecast rules.

As explained in Section 3.2, an informed seller needs a forecast rule to map competitors’ past prices into a prediction of their next prices, which is then used to compute the greedy response 2 . There are many possible forecast rules, but a comprehensive survey is beyond the scope of this paper, as our focus is on the strategic-modeling choice between oblivious and informed sellers. Therefore, we adopt the running-mean forecast rule \(\hat{p}_{n+1, j} = m_{n, j}\), which is operationally simple and robust to competitors’ price-exploration noise. This is also the canonical choice for the classical adaptive-best-response results in the game theory literature (Section 2). As ablations, we will consider two alternative forecast patterns in Section 6.2 and show that the qualitative insights carry over.

6.1 All-Informed Markets↩︎

We begin with the all-informed market, \(\mathcal{I}^{in} = [N]\). Throughout this subsection, let \(\Gamma\) be the Nash matrix from Section 3.3 and define the linearized best-response map around \(\mathbf{p}^{NE}\), \[B \;\triangleq\; I - \tfrac{1}{2}\,\mathrm{diag}(1/\beta_i)\,\Gamma.\] Under the standing assumption \(\gamma_i < |\beta_i|\), \(I - B\) is positive stable, so the matrix Lyapunov equation \(P(I-B) + (I-B)^\top P = I\) admits a unique symmetric positive-definite solution \(P\); we set \(\mu \triangleq 1/\sigma_{\max}(P) > 0\). The following result characterizes the convergence of realized prices to the competitive equilibrium; the proof is in Appendix 9.7.

Theorem 4 (All-informed market). Suppose that \[\label{eq:conditioninformedsellersmixed} \eta_{\min} + 1 \;>\; 2\, \eta_{\max}.\qquad{(3)}\] Then, \(\hat{\theta}_{n,k}\to\theta_k^*\) and \(\tilde{p}_{n,k}\to p_k^{NE}\) a.s.for every \(k\in[N]\), and the realized price satisfies \[\label{eq:cor:allinformed-rate} \mathbb{E}\|\mathbf{p}_n - \mathbf{p}^{NE}\|_2^2 \;=\; O\!\left(n^{-\mu/2}(\log n)^2 \;+\; n^{\eta_{\max} - 1} \;+\; n^{-\eta_{\min}}\right).\qquad{(4)}\] Moreover, in the special case where \(\eta_i = 1/2\) for every \(i\in[N]\) and, additionally, \[\label{eq:cor:symmpart-regularity} \lambda_{\max}(B + B^\top) \;<\; 1,\qquad{(5)}\] we have \(\mathbb{E}\|\mathbf{p}_n - \mathbf{p}^{NE}\|_2^2 = O(n^{-1/2})\). Summing over \(n\le T\) in 4 gives individual dynamic regret \(\Delta_i(\theta_i,T) = O(\sqrt T)\) for every \(i\in[N]\) and the same rate for the aggregate regret.

6.1.0.1 Discussion.

The general rate ?? decomposes into three terms: a market-side rate independent of the exploration schedule, the parameter-estimation error (Theorem 6), and the residual exploration variance controlled by the slowest-decaying schedule. In the special case, every seller achieves the optimal \(\sqrt T\) regret rate, matching the optimal rate for dynamic pricing with demand learning in the monopolistic case [13]. Thus, the all-informed market can be interpreted as an example of a well-behaved market in which the sellers reach a stable outcome efficiently without paying a persistent exploration tax.

6.1.0.2 Interpreting the conditions.

Condition ?? is a mild regularity requirement on relative exploration rates to ensure that the learning signal (mean growth) dominates the noise (variance). Condition ?? requires that the symmetric part of the linearized best-response map \(B\) has spectral radius strictly less than \(1\), which holds whenever \(B\) is symmetric. When all sellers share a common own-price slope \(\beta_i \equiv \beta\), a direct calculation gives \(B + B^\top = I - H / (2\beta)\), so ?? is equivalent to \(H \prec 0\), the condition from Section 3.3 that guarantees a unique unconstrained collusive outcome.

6.1.0.3 Numerical experiments.

Figure 5 corroborates Theorem 4 in the special case with exploration rate \(\eta_i = 1/2\). Panel () is a symmetric market at \(N \in \{2, 5\}\); panel () is an asymmetric market at \(N \in \{3, 5\}\) with heterogeneous demand primitives. In both panels the regularity condition ?? holds and the seed-averaged price-MSE trajectory decays at the slope \(\approx -1/2\) predicted by Theorem 4. Panel () is an asymmetric \(N = 3\) market where condition ?? is violated. In this regime, the MSE still tends to zero, but at a markedly slower empirical slope \(\approx -0.3\).

a
b
c

Figure 5: Empirical corroboration of Theorem 4 in the special case \(\eta_i = 1/2\). All three panels overlay seed-averaged price-MSE \(\EE\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2\) on log–log axes. Panel (): symmetric all-informed markets at \(N \in \{2, 5\}\). Panel (): asymmetric all-informed markets at \(N \in \{3, 5\}\). In both, regularity ?? holds and the empirical slope matches the predicted \(-1/2\). Panel (): asymmetric \(N = 3\) market designed so every cell violates regularity ?? . The MSE still decays in every cell but at a slower slope (\(\approx -0.3\) vs.the predicted \(-1/2\)), confirming that condition ?? is sufficient but not strictly necessary in practice. Implementation details are in Appendix 8.2.. a — Symmetric, regularity holds., b — Asymmetric, regularity holds., c — Every cell violates regularity.

6.2 Mixed Markets↩︎

We now consider mixed markets in which a strict subset of sellers is oblivious and the remainder is informed, i.e., \(\cI^{ob} \cup \cI^{in} = [N]\) with both subsets non-empty. Recall from Section 3.2 that both oblivious and informed sellers follow an estimate–exploit–explore paradigm, and recall the heterogeneous-exploration setup from the opener of Section 6: oblivious sellers explore persistently (\(\eta_i = 0\)), while informed sellers’ dithering decays at rate \(\nu_{n,j}^2 = \Theta(n^{-\eta_j})\) with \(\eta_j > 0\).

Theorem 5 below characterizes the market outcome in mixed markets. To set the stage, we first introduce some constants. From Theorem 2 we inherit the oblivious greedy-map Lipschitz constant \(L_\phi^{ob}\), the regressor envelope \(C_x\) (a uniform a.s. bound on \(\|x_{n,i}^{ob}\|_2\)), and the persistent-excitation lower bound \(C_M\)—all restricted to the set of oblivious sellers \(\cI^{ob}\). We also define the oblivious-on-oblivious cross-coupling weight \[\bar\gamma^{ob} \;\triangleq\; \frac{1}{2}\Bigl(\max_i \sum_{j \in \cI^{ob}\setminus\{i\}} \gamma_{i,j} + \max_i \sum_{j \in \cI^{ob}\setminus\{i\}}\gamma_{j,i}\Bigr),\] the analogue of \(\bar\gamma\) from Theorem 2 restricted to oblivious–oblivious links. We further introduce \(L_\phi^{in,\theta}\), the maximum Lipschitz constant of \(\phi_j^{in}(\theta;\,\mathbf{m})\) in \(\theta\) over the projection boxes of the informed sellers and over possible competitor price means \(\mathbf{m} \in [l,u]^{N-1}\).

Since informed sellers and oblivious sellers interact through different modeling dynamics, three additional cross-coupling constants are needed to capture the feedback effects. The first measures the aggregate effect of informed sellers on oblivious sellers’ regression error: \[\bar\Lambda \;\triangleq\; \max_{i \in \cI^{ob}} \sum_{j \in \cI^{in}}\frac{\gamma_{i,j}\,\gamma_j}{2|\beta_j|} \;+\; L_\phi^{in,\theta}\,\max_{i \in \cI^{ob}}\sum_{j \in \cI^{in}}\gamma_{i,j}.\] The second represents how a perturbation in the running mean \(m_{n,k}\) propagates through informed sellers’ best responses and lands on oblivious sellers’ regression: \[\bar\Psi \;\triangleq\; \max_{k \in [N]} \sum_{i \in \cI^{ob}}\sum_{j \in \cI^{in}\setminus\{k\}}\frac{\gamma_{i,j}\,\gamma_{j,k}}{2|\beta_j|}.\] The third is the small-gain margin on the informed side: the diagonal self-stabilizing rate—combining the effects of oblivious sellers and running-mean forecasts on informed sellers’ regression—minus the informed–informed cross-coupling drag, \[\bar\kappa \;\triangleq\; \min\,\Bigl\{2 - L_\phi^{ob},\;\; 2 - L_\phi^{in,\theta} - \max_{j \in \cI^{in}}\frac{\gamma_j}{2|\beta_j|}\Bigr\} \;-\; \max_{k \in [N]}\sum_{j \in \cI^{in}\setminus\{k\}}\frac{\gamma_{j,k}}{2|\beta_j|}.\] We are now ready to state the convergence result for mixed markets. The proof is in Appendix 9.8.

Theorem 5 (Mixed-market convergence). Suppose:

  • (Exploration regularity, informed side) \(\eta_{\max} < 1/2\);

  • (Small-gain, informed side) \(\bar\kappa \;>\; 0\);

  • (Persistent excitation, oblivious side) There exists \(C_M > 0\) such that \(\EE[\,x_{n,i}^{ob}(x_{n,i}^{ob})^\top \mid \cF_{n-1}] \succeq C_M\, I_2\) for every \(n\) and every \(i \in \cI^{ob}\);

  • (Small-gain, oblivious side) \(C_M > C_x\left(\, 2\bar\gamma^{ob}\, L_\phi^{ob} + \bar\Lambda + L_\phi^{ob}\, \bar\Psi/\bar\kappa \,\right)\).

Then, \(\hat{\theta}_{n,i}^{ob} \to \theta_i^{*,ob}\), \(\hat{\theta}_{n,j}^{in} \to \theta_j^{*,in}\), and \(\tilde{p}_{n,k} \to p_k^{NE}\) a.s.for every \(i \in \cI^{ob}\), \(j \in \cI^{in}\), and \(k \in [N]\). Moreover, the aggregate mean-squared error satisfies \[\sum_{i \in \cI^{ob}} \EE\norm{\hat{\theta}_{n,i}^{ob} - \theta_i^{*,ob}}_2^2 \;+\; \sum_{k \in [N]} \EE\!\left[(\tilde{p}_{n,k} - p_k^{NE})^2\right] \;=\; \begin{cases} O\!\bigl(n^{\eta_{\max} - 1}\bigr), & c^* > 1 - \eta_{\max},\\[3pt] O\!\bigl(n^{\eta_{\max} - 1}\,\log n\bigr), & c^* = 1 - \eta_{\max},\\[3pt] O\!\bigl(n^{-c^*}\bigr), & c^* < 1 - \eta_{\max}; \end{cases}\] and every informed seller’s realized-price mean-squared error satisfies \(\EE[(p_{n,j} - p_j^{NE})^2] = O \left(n^{-\min\{c^*,\, \eta_j\}} \right)\) for \(j \in \cI^{in}\), where \(c^* \triangleq \sup_{\lambda > 0} \min\bigl\{1 + \bigl[C_M - C_x(2\bar\gamma^{ob} L_\phi^{ob} + \bar\Lambda) - \lambda L_\phi^{ob}\bigr]/C_x^2,\;\; \bar\kappa - C_x \bar\Psi/\lambda\bigr\} > 0\).

6.2.0.1 Discussion.

Despite the presence of misspecified oblivious competitors, informed sellers learn the true demand model and the market converges jointly to \(\mathbf{p}^{NE}\). Oblivious sellers reproduce the Theorem 2 outcome—near-Nash asymptotic prices, persistent exploration tax, linear regret—while informed sellers earn sublinear regret (Proposition [prop:dynamicbenchmark]). The informed realized-price MSE decays at rate \(\min\{c^*,\,\eta_j\}\), where the stability bottleneck \(c^*\) pins the fastest joint convergence rate consistent with both subsystems’ small-gain margins and \(\eta_j\) is the informed seller’s exploration exponent. The two bottlenecks decouple, since \(\eta_{\max}\) enters neither \(c^*\) nor conditions (ii)–(iv). Any larger-than-\(\sqrt T\) regret thus reflects the price informed sellers pay for having oblivious competitors who inject persistent misspecification into the market.

6.2.0.2 Interpreting the conditions.

The four conditions pair an exploration-regularity requirement with a small-gain inequality on each side of the market. Condition (i) is ?? from Theorem 4 specialized to \(\eta_{\min} = 0\) for the persistently exploring oblivious sellers, while condition (ii) is the small-gain inequality on the informed side, requiring that the diagonal self-decay rate exceeds the informed-on-informed cross-coupling drag. On the oblivious side, condition (iii) is the persistent-excitation requirement ?? from Theorem 2—the rational behavior predicted in Section 4.2—and condition (iv) is the small-gain inequality requiring that persistent excitation dominate every cross-coupling channel into the oblivious regression, with its leading oblivious–oblivious term recovering the (strongest) all-oblivious condition ?? .

a
b
c

Figure 6: Empirical verification of Theorem 5: seed-averaged price MSE \(\EE\norm{\tilde{\mathbf{p}}_n - \mathbf{p}^{NE}}_2^2\) on log–log axes. Panel (): symmetric \(N = 5\) with \(|\cI^{ob}| = 2\), \(|\cI^{in}| = 3\). Panel (): asymmetric markets at \(N \in \{3, 5, 10\}\) with one oblivious and \(N-1\) informed sellers. In both panels \(\bar\kappa < 0\), so conditions (ii) and (iv) fail jointly. Panel (): symmetric \(N = 5\) with \(|\cI^{ob}| = 4\), \(|\cI^{in}| = 1\), primitives chosen so \(\bar\kappa>0\) (condition (ii) holds) but condition (iv) fails. Across all 36 cells the price-MSE trajectory decays log-linearly to zero, with no qualitative difference between cells in which (ii) and (iv) jointly fail and cells in which (iv) alone fails. Implementation details are in Appendix 8.2.. a — Symm., (ii) and (iv) both fail., b — Asymm., (ii) and (iv) both fail., c — Symm., (ii) holds, (iv) fails.

6.2.0.3 Numerical experiments.

Figure 6 stress-tests Theorem 5 in three settings: a symmetric \(N = 5\) mixed market sweeping \((\gamma, \nu^2)\) on a \(5 \times 3\) grid (panel ()), an asymmetric mixed market at \(N \in \{3, 5, 10\}\) with one oblivious and \(N-1\) informed sellers (panel ()), and a symmetric \(N = 5\) market with primitives tuned so condition (ii) holds but (iv) fails (panel ()). In the first two panels the projection-box-induced Lipschitz envelopes are large enough that \(\bar\kappa < 0\), so conditions (ii) and (iv) jointly fail.5 Across all 36 cells the seed-averaged price-MSE decays cleanly to zero on log–log axes, indicating that the small-gain conditions are sufficient but not necessary in practice.

6.2.0.4 Alternative forecast rules.

Appendix 8.10 ablates the running-mean rule against an implementable lag-1 rule \(\hat{p}_{n+1,j} = p_{n,j}\) and a clairvoyant greedy-component rule \(\hat{p}_{n+1,j} = \tilde{p}_{n+1,j}\). All three rules share the practical constraint that the competitor’s next-period exploration noise \(z_{n+1,1}\) is unobservable, and they all leave \(\mathbf{p}^{NE}\) as the price attractor, showing that the qualitative insight of Theorem 5 carries over. A fourth, clairvoyant forecast that does peek at \(z_{n+1,1}\) is also ablated in the appendix and is discussed in the following section.

7 Conclusion: The Strategy Game↩︎

Sections 5 (obob), 6.1 (inin), and 6.2 (obin) characterize the three distinct market compositions via their convergence theorems. In this section we lift those three theorems into a strategy game on \(\{\textsf{oblivious},\textsf{informed}\}\) and identify its unique strict pure-strategy Nash equilibrium. The key upshot is that informed strictly dominates oblivious in every market composition. For clarity, we first present the duopoly case, then extend to the \(N\)-seller case.

For the asymptotic revenue argument we impose an additional requirement on the bounded-support exploration shocks of Section 3.2: for every \(i\in[N]\) there exists a deterministic constant \(\delta_i < \min\,\left\{p_i^{NE} - l,\; u - p_i^{NE}\right\}\) such that \[\label{eq:explore-bound} |z_{n,i}| \;\le\; \delta_i \qquad \text{for all sufficiently large } n.\tag{10}\] Condition 10 allows clipping to be asymptotically inactive once \(\tilde{p}_{n,i}\) has converged into a neighborhood of \(p_i^{NE}\).6

7.0.0.1 Strict dominance.

Consider the duopoly strategy game in which each seller \(i\in\{1,2\}\) chooses \(s_i\in\{\textsf{oblivious},\textsf{informed}\}\) and earns the asymptotic surplus-capture \(S_i(s_1, s_2)\), defined as in 5 along the price path induced by \((s_1, s_2)\). The strategy profile \((s_1, s_2)\) determines the market composition and thus which of the three convergence theorems applies. The following proposition shows that the informed strategy strictly dominates the oblivious strategy for each seller, and that the unique strict pure-strategy Nash equilibrium of the strategy game is the all-informed market. The proof is in Appendix 9.9.

Suppose sellers’ exploration schedules satisfy 10 , and the assumptions of Theorems 2, 4, and 5 hold under the corresponding strategy profiles (with constant oblivious variance \(\nu_i^2 > 0\) on the oblivious side and a decaying schedule on the informed side). Then, for each seller \(i\in\{1,2\}\) and any opponent choice \(s_{-i}\), almost surely, \[S_i\bigl(\textsf{informed},\, s_{-i}\bigr) \;=\; 0 \;>\; -\,\frac{|\beta_i|\,\nu_i^2}{\Pi_i^{C}-\Pi_i^{NE}} \;=\; S_i\bigl(\textsf{oblivious},\, s_{-i}\bigr).\] Hence, informed strictly dominates oblivious for each seller, and \((\textsf{informed},\textsf{informed})\) is the unique strict pure-strategy Nash equilibrium of the strategy game.

7.0.0.2 Numerical experiments.

Table 3 reports per-seller realized average revenue \(\bar R_{T,i} \triangleq R_{T,i}/T\) and the surplus-capture ratio \(S_i\) for each composition in a symmetric duopoly. In the obob composition, both sellers pay the persistent exploration tax, with realized revenues sitting below \(\Pi^{NE}\) in proportion to \(\nu^2\); in the obin composition, the oblivious seller pays the tax while the informed seller essentially captures \(\Pi^{NE}\) under the running-mean forecast; and in the inin composition, both sellers’ realized revenues are essentially equal to \(\Pi^{NE}\).

Table 3: Revenue comparison across the three duopoly compositions of oblivious (ob) and informed (in) sellers. \(\bar R_{T,i}\) is the per-period running-mean revenue (brackets are cross-seed \(5\%\)\(95\%\) ranges over \(S = 200\) seeds); \(S_i\) is the surplus-capture ratio. Implementation details are in Appendix 8.2.
composition exploration / forecast \(\bar R_{T,1}\) [5–95%] \(\bar R_{T,2}\) [5–95%] \(S_1\) \(S_2\)
obob linear, \(\nu^2 = 0.20\) \(2.99\;[2.95, 3.03]\) \(2.99\;[2.96, 3.03]\) \(-0.27\) \(-0.27\)
obob linear, \(\nu^2 = 0.05\) \(3.16\;[3.12, 3.23]\) \(3.16\;[3.12, 3.23]\) \(-0.04\) \(-0.04\)
obin running mean, \(\eta = 0.5\) \(3.09\;[3.05, 3.16]\) \(3.19\;[3.16, 3.22]\) \(-0.13\) \(\phantom{-}0.00\)
obin running mean, \(\eta = 0.25\) \(3.09\;[3.06, 3.13]\) \(3.18\;[3.15, 3.21]\) \(-0.14\) \(-0.01\)
obin perfect prediction \(3.14\;[3.09, 3.21]\) \(3.41\;[3.37, 3.45]\) \(-0.07\) \(\phantom{-}0.31\)
inin decaying, \(\eta = 0.5\) \(3.19\;[3.15, 3.23]\) \(3.19\;[3.15, 3.23]\) \(\phantom{-}0.00\) \(\phantom{-}0.00\)
inin decaying, \(\eta = 0.7\) \(3.19\;[3.12, 3.25]\) \(3.19\;[3.12, 3.26]\) \(\phantom{-}0.00\) \(\phantom{-}0.00\)

The last obin row probes how much an informed seller could in principle gain from sharper competitor-price information: under a clairvoyant perfect-prediction forecast \(\hat{p}_{n+1,1} = p_{n+1,1}\), the informed-side revenue rises to the Stackelberg revenue \(\Pi^{S}\)—the follower payoff in a Stackelberg duopoly with the oblivious seller as leader—and the surplus-capture ratio climbs from \(0\) to \(0.31\). Because the competitor’s next-period exploration noise is unobservable in practice, \(\Pi^{S}\) acts as an information upper bracket on the informed-side revenue against an oblivious competitor; any sharper-but-implementable forecast would yield a realized revenue somewhere in \([\Pi^{NE}, \Pi^{S}]\). See Appendix 8.10 for more details.

7.0.0.3 The strategy game.

Treating each seller’s choice of oblivious vs.informed as a strategic move yields a two-player normal-form game whose payoff is the realized \(S_i\). Table 4 summarizes the four cells, with the payoffs taken from Proposition [prop:strict-dom] and confirmed numerically by Table 3. The off-diagonal “\(\ge 0\)” entries acknowledge a range of possibilities for an informed seller facing an oblivious competitor: under the practical forecasts of Proposition [prop:strict-dom] the informed-side payoff is the Nash surplus capture \(0\), but should richer side information about a competitor’s next-period price—say from market-microstructure data, public algorithmic disclosures, or stronger forecasting models—become available, the payoff could be strictly higher, capped by the Stackelberg upper bracket (Appendix 8.10). In every case, informed strictly dominates oblivious for each seller, and (informed, informed) is the unique strict pure-strategy Nash equilibrium of the strategy game.

12pt

Table 4: The strategy game. Rows index seller 1’s strategy, columns index seller 2’s strategy; each cell reports the asymptotic surplus-capture payoffs \((S_1, S_2)\). The entry \(-\nu^2\) is shorthand for the oblivious payoff under a constant exploration variance \(\nu^2\), and \(0\) is the informed payoff at Nash. The off-diagonal “\(\ge 0\)” entries cover both the Nash value (attained under the practical forecasts of Proposition [prop:strict-dom]) and the Stackelberg upper bracket (Appendix 8.10). The unique strict pure-strategy Nash equilibrium is (informed, informed), shown in the bottom-right cell.
oblivious informed
oblivious \(-\nu^2\) \(-\nu^2\) \(-\nu^2\) \(\ge 0\)
informed \(\ge 0\) \(-\nu^2\) \(\boldsymbol{0}\) \(\boldsymbol{0}\)

7.0.0.4 Beyond duopoly.

The two-seller strategy game extends cleanly to \(N\)-seller markets at every composition \(|\cI^{ob}| \in \{0, 1, \ldots, N\}\): the on-diagonal cells generalize via Theorems 2 and 4, and the off-diagonal cells via Theorem 5. Appendix 8.11 substantiates this numerically at \(N \in \{3, 5\}\) across all \(|\cI^{ob}| \in \{0, \ldots, N\}\), and the four-cell qualitative ordering of Table 4 carries over: oblivious sellers strictly below Nash, informed sellers operationally at Nash, with a strict positive gap \(\bar S^{in} - \bar S^{ob}\) in every mixed cell. As a result, any unilateral deviation from the all-informed equilibrium moves the deviator from the operationally-zero \(\bar S^{in}\) cell to the strictly-negative \(\bar S^{ob}\) cell, confirming the all-informed market as the unique strict pure-strategy Nash equilibrium of the \(N\)-seller strategy game as well.

7.0.0.5 Future directions.

Several directions warrant further investigation. First, our framework distinguishes only between fully oblivious and fully informed sellers, while richer modeling strategies that interpolate between these two extremes would be valuable to study. Second, our results rely on the running-mean forecast for every informed seller, and a fuller understanding of the forecast-rule design space remains open. Third, we consider only mean-zero dithering on the greedy price, so alternative exploration designs are worth investigating. Fourth, more elaborate ways of incorporating competitor information into both the demand model and the pricing decision deserve attention. The interplay between model specification, learning dynamics, and market outcomes in competitive settings where algorithmic collusion is possible remains a rich and largely open area, and we hope that our work provides a useful starting point for future research in this direction.

Acknowledgments↩︎

We thank Santiago R. Balseiro, Nicola Rosaia, Hongyao Ma, Izzy Grosof, Vineet Goyal, and Julius Durmann for helpful comments and suggestions. We thank the audience at NYC Operations Day 2025, Market Innovation Workshop 2025, INFORMS Revenue Management and Pricing Section Conference 2025, and INFORMS Annual Meeting 2025 for their comments on preliminary versions of this work. We also appreciate the feedback received from the anonymous reviewers at EC 2026, which helped us improve the presentation and clarify the results.

8 Additional Details↩︎

8.1 Notation Summary↩︎

Table 5 consolidates the main notation used throughout the paper. Symbols introduced only locally (within a single proof or appendix subsection) are not listed.

Table 5: Notation used throughout the paper.
Symbol Description
(continued from previous page)
Symbol Description
continued on next page
\(N\) Number of sellers, \(N\ge 2\)
\([N]\) Seller index set \(\{1,\ldots,N\}\)
\(i,j,k\) Seller indices
\(n,m\) Discrete time indices
\(T\) Pricing horizon
\(\cI^{ob},\cI^{in}\) Sets of oblivious and informed sellers; \(\cI^{ob}\cup\cI^{in}=[N]\)
\(\cF_n\) Filtration generated by all price–demand history up to time \(n\)
\([l,u]\) Price box with \(u>l>0\); all prices lie in \([l,u]\)
\(d_{n,i}\) Demand realized by seller \(i\) at time \(n\), eq. 1
\(\alpha_i,\beta_i,\gamma_{i,j}\) True demand intercept, own-price slope, and cross-price coefficient (\(i\neq j\))
\(\varepsilon_{n,i}\) Demand noise; i.i.d.across \(n\) and \(i\), mean zero, bounded support
\(\gamma_i\) \(\sum_{j\neq i}\gamma_{i,j}\), total cross-price effect of competitors on seller \(i\)
\(\gamma_i^{\mathrm{col}}\) \(\sum_{j\neq i}\gamma_{j,i}\), total effect of seller \(i\)’s price on others
\(\bar\gamma\) \(\frac{1}{2}\left(\max_i \gamma_i + \max_i \gamma_i^{\mathrm{col}}\right)\), worst-case cross-price coupling
\(\boldsymbol{\alpha}\) \((\alpha_1,\ldots,\alpha_N)^\top\)
\(a_i,b_i\) Misspecified intercept and slope used by an oblivious seller
\(\theta_i^{*,ob},\theta_i^{*,in}\) (Pseudo-)true oblivious and informed parameters for seller \(i\)
\(p_{n,i}\) Seller \(i\)’s posted price at time \(n\)
\(\mathbf{p}_n\) Price vector \((p_{n,1},\ldots,p_{n,N})\in\RR^N\)
\(\mathbf{p}_{n,-i}\) Competitor price vector \((p_{n,j})_{j\neq i}\)
\(\tilde{p}_{n+1,i}\) Unconstrained myopic greedy price (oblivious or informed) for seller \(i\)
\(z_{n,i}\) Exploration perturbation; mean zero, variance \(\nu_{n,i}^2\), bounded support
\(\nu_{n,i}^2\) Exploration variance for seller \(i\) at time \(n\); abbreviated \(\nu^2\) (constant, homogeneous) or \(\nu_n^2\) (time-varying, homogeneous) when no confusion arises
\(\eta_i\) Polynomial exploration exponent: \(\nu_{n,i}^2 = \Theta(n^{-\eta_i})\)
\(\eta_{\min},\eta_{\max}\) \(\min_i\eta_i\) and \(\max_i\eta_i\)
\(x_{n,i}^{ob},x_{n,i}^{in}\) Oblivious and informed regressor vectors
\(\Theta_i^{ob},\Theta_i^{in}\) Compact convex projection sets for parameter estimates
\(\tilde{\theta}_{n,i}^{ob},\tilde{\theta}_{n,i}^{in}\) Unprojected least-squares estimates
\(\hat{\theta}_{n,i}^{ob},\hat{\theta}_{n,i}^{in}\) Projected least-squares estimates
\(\hat{r}_{n,i}^{ob},\hat{r}_{n,i}^{in}\) Estimated per-period revenue functions for seller \(i\)
\(\phi^{ob}(\theta)\) Oblivious greedy map; identical functional form across sellers
\(\phi_i^{in}(\theta,\mathbf{p}_{-i})\) Informed greedy map for seller \(i\) given parameters \(\theta\) and (forecasted) competitor prices \(\mathbf{p}_{-i}\)
\(L_\phi^{ob}\) Lipschitz constant of \(\phi^{ob}(\cdot)\) over \(\bigcup_i \Theta_i^{ob}\)
\(L_\phi^{in,\theta}\) Lipschitz constant of \(\phi_i^{in}(\theta;\mathbf{m})\) in \(\theta\) over the informed projection boxes and over possible competitor price means \(\mathbf{m} \in [l,u]^{N-1}\)
\(\hat{p}_{n+1,j}\) Informed seller’s forecast of competitor \(j\)’s next price (default: running mean \(m_{n,j}\))
\(R_i(p_i,\mathbf{p}_{-i})\) Seller \(i\)’s expected per-period revenue under the true demand model
\(\mathbf{p}^{NE},\,p_i^{NE}\) Nash equilibrium price vector and seller-\(i\) component
\(\mathbf{p}^{C},\,p_i^{C}\) Collusive (joint-revenue-maximizing) price vector and seller-\(i\) component
\(\Pi_i^{NE},\,\Pi_i^{C}\) Seller \(i\)’s per-period revenue at \(\mathbf{p}^{NE}\) and \(\mathbf{p}^{C}\)
\(\Gamma\) \(N\times N\) Nash matrix: \(\Gamma_{ii}=2\beta_i\), \(\Gamma_{ij}=\gamma_{i,j}\); \(\Gamma\mathbf{p}^{NE}=-\boldsymbol{\alpha}\)
\(H\) \(N\times N\) symmetric collusion matrix: \(H_{ii}=2\beta_i\), \(H_{ij}=\gamma_{i,j}+\gamma_{j,i}\)
\(B\) Linearized informed best-response map at \(\mathbf{p}^{NE}\): \(B=I-\tfrac12\,\mathrm{diag}(1/\beta_i)\,\Gamma\)
\(\Delta_i(\theta_i,T)\) Dynamic regret of seller \(i\) over horizon \(T\), eq. 3
\(S_i\) Surplus-capture ratio of seller \(i\), eq. 5
\(m_{n,i},\,m_i(t)\) Running mean of seller \(i\)’s prices (discrete time \(n\) / continuous time \(t\))
\(Q_{n,ij},\,Q_{ij}(t)\) Running cross second-moment of prices (discrete / continuous)
\(p_i^g(m,Q)\) Misspecified greedy price as a function of empirical moments \((m,Q)\)
\(C_x\) Uniform almost-sure bound on \(\|x_{n,i}^{ob}\|_2\)
\(C_M\) Persistent-excitation lower bound on the smallest eigenvalue of the oblivious conditional regressor covariance
\(f\lesssim g\) \(f=O(g)\) up to constants depending only on demand primitives
\(f\asymp g\) \(f=\Theta(g)\) in the same sense

8.2 Experiment Details↩︎

This subsection consolidates the demand primitives, exploration schedules, initial conditions, sample sizes, and other replication details for the figures and tables in the main body. Within each entry, \(\alpha\), \(\beta\), and \(\gamma\) refer to the linear-demand primitives of 1 , \([l, u]\) to the price box, \(\nu^2\) to per-period exploration variance, \(T\) to the horizon, and \(S\) to the number of independent seeds. Unless otherwise stated, demand noise is uniform on \([-0.2, 0.2]\) (so \(\sigma_\varepsilon = 0.2/\sqrt{3}\)), the dithering distribution is uniform on \([-\sqrt{3\nu^2}, +\sqrt{3\nu^2}]\), and warm-up prices are seeded near a fixed pair within \([l, u]^N\) for each draw.

8.2.0.1 Figure 1 (intro sample paths).

Symmetric duopoly with \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.4\), \([l, u] = [0.5, 2.5]\) (so \(p^{NE} \approx 1.56\), \(p^C \approx 2.08\)); cumulative exploration \(J_n = \Theta(\sqrt n)\) implemented via \(\nu_n^2 = 0.025/\sqrt{n+1}\). Horizon \(n = 10^4\); one seed per panel selected for illustration of near-competitive, indeterminate, and near-collusive paths.

8.2.0.2 Table 2 (variance-dominance stress test).

Symmetric oblivious–oblivious duopoly with \(\alpha_i = 2.5\), \(\beta_i = -1\), \(\gamma_{i,j} = 0.6\) for \(i \neq j\), \([l, u] = [0.5, 3.5]\) (yielding \(\Pi^{NE} = 3.189\), \(\Pi^{C} = 3.906\), \(p^{NE} \approx 1.786\), and \(q_- \approx 0.962\)). The oblivious projection box is \((a, b) \in [1.2, 8.0] \times [-2.0, -0.5]\), chosen as a stress test of Proposition [prop:variancedominancetwosellercase](b): the non-Nash lock candidates are \(\ell_{-,i} = 1.2\) and \(\ell_{+,i} = 2.0\), so \(\ell_{+,i} > p^{NE}\) but \(\ell_{-,i} \in (q_-, p^{NE})\) falls inside the reverse-regime interval where the dominated would out-capture the dominant if it locked there; condition 7 therefore fails for this box. Both sellers use polynomial exploration schedules \(\nu_{n,k}^2 = 0.05\,(n+1)^{-\eta_k}\); the dominant seller’s exponent is fixed at \(\eta_2 = 0.01\), and the dominated’s exponent \(\eta_1 \in \{0.5, 0.7, 1.0, 1.5\}\) varies the strength of the variance-dominance event. Horizon \(T = 2\times 10^5\), \(S = 120\) seeds per cell.

8.2.0.3 Figure 2 (global-convergence stress test).

Both panels use \([l, u] = [0.5, 2.5]\) and uniform demand noise with standard deviation \(0.2\), \(S = 60\) seeds per configuration. Left (asymmetric): \(N \in \{3, 5, 10\}\) and \(\nu^2 \in \{0.05, 0.10, 0.20\}\) (line style); demand primitives drawn once per \(N\) from \(\alpha_i \sim \mathcal{N}(2.5, 0.4^2)\) clipped to \([1.5, 3.5]\), \(\beta_i \sim \mathcal{N}(-1, 0.2^2)\) clipped to \([-1.5, -0.7]\), and \(\gamma_{i,j} = (0.4/(N-1)) \cdot U[0.7, 1.3]\) for each \(i \neq j\); horizon \(T = 5\times 10^4\). Right (symmetric): \(N = 5\), \(\alpha_i = 2.5\), \(\beta_i = -1\), cross-price \(\gamma \in \{0.05, 0.10, 0.15\}\) (color); same \(\nu^2\) grid; \(T = 3\times 10^4\). Across all 18 configurations \(\bar\gamma\, L_\phi^{ob}\, C_x\) is of order \(1\)\(10\), while \[C_M(\nu^2) \;\le\; \tfrac{1}{2}\Big(1 + p_*^2 + \nu^2 - \sqrt{(1 + p_*^2 + \nu^2)^2 - 4\nu^2}\Big)\] for any limit price \(p_* \in [l, u]\), which stays below \(10^{-1}\) for every \(\nu^2 \le 0.20\); the sufficient condition ?? is therefore violated by roughly two orders of magnitude across the entire sweep.

8.2.0.4 Figure 3 (excursions in the symmetric duopoly).

Common parameters: \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.4\), \([l, u] = [0.5, 2.5]\) (so \(p^{NE} = 1.56\), \(p^C = 2.08\)); exploration variance \(\nu^2 = 0.04\) with uniform dithering. Upward (panels (a), (b)): ODE initial condition \(m_1(0) = 0.7\), \(m_2(0) = 0.8\), \(Q_{11}(0) = 0.8\), \(Q_{22}(0) = 0.9\), \(Q_{12}(0) = 1.2\); discrete-time warm-up prices \((0.5, 1.0)\) and \((0.8, 1.2)\); demand noise \(\mathcal{N}(0, 0.2)\); discrete horizon \(n = 10^7\). Downward (panels (c), (d)): ODE initial condition \(m_1(0) = 1.8\), \(m_2(0) = 1.9\), \(Q_{11}(0) = 3.5\), \(Q_{22}(0) = 4.0\), \(Q_{12}(0) = 2.1\); discrete-time warm-up prices \((1.5, 1.7)\) and \((1.7, 1.6)\); demand noise \(\mathcal{N}(0, 0.05)\); discrete horizon \(n = 10^4\).

8.2.0.5 Figure 4 (pseudo-equilibrium continuum, prices and revenues).

Symmetric duopoly with \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.4\), \([l, u] = [0.5, 2.5]\). Decaying-exploration schedules \(\nu_n^2 = 0.3(n+1)^{-\eta}\) with \(\eta \in \{0.5, 0.85\}\). Each sample path is started from one of \(60\) warm-up price pairs jittered around six anchors that span the price box: the four near-corner anchors \((l + 0.2, l + 0.2)\), \((u - 0.2, u - 0.2)\), \((l + 0.2, u - 0.2)\), \((u - 0.2, l + 0.2)\), plus two on-diagonal anchors \((1.5, 1.5)\) near \(\mathbf{p}^{NE}\) and \((2.0, 2.0)\) near \(\mathbf{p}^{C}\), with \(S = 30\) seeds per anchor (total \(S = 1{,}800\)) and horizon \(T = 8\times 10^4\). Following [2], the admissible regression-ratio pairs \((r_1, r_2)\) satisfy (i) the non-singularity condition \(r_i < -\beta_i / \gamma_{i,j}\) for each \(i \neq j\); and (ii) either the Nash branch \(r_1 = r_2 = 0\) or the co-moving branch \(0 < r_1 r_2 \le 1\). For our market \(-\beta_i/\gamma_{i,j} = 1/0.4 = 2.5\), so we sweep \(r \in [-2.5, 2.5]^2\) on a \(1{,}201\)-point interior grid augmented by a \(4{,}000\)-point boundary sweep along the ridge \(r_1 r_2 = 1\). Shading in the price panels is the theoretical region induced by this admissible set \(\{(r_1, r_2)\}\); shading in the revenue panels is its image under the revenue map \(R_i(\mathbf{p}) = p_i(\alpha_i + \beta_i p_i + \sum_{j\neq i}\gamma_{i,j} p_j)\).

8.2.0.6 Figure 5 (all-informed empirical corroboration).

All three panels: \(\eta_i = 1/2\), \(S = 100\) seeds per cell, \(T = 5\times 10^4\). Panel (): symmetric markets at \(N \in \{2, 5\}\) with \(\alpha_i = 2.5\), \(\beta_i = -1\), \(\gamma_{i,j} = 0.4/(N-1)\), sweeping \(\nu^2 \in \{0.05, 0.10, 0.20\}\). Panel (): asymmetric markets at \(N \in \{3, 5\}\) with heterogeneous demand primitives drawn from \(\alpha_i \sim \cN(2.5, 0.4^2)\) clipped to \([1.5, 3.5]\), \(\beta_i \sim \cN(-1, 0.2^2)\) clipped to \([-1.5, -0.7]\), and \(\gamma_{i,j} = (0.4/(N-1)) \cdot U[0.7, 1.3]\) (drawn once per \(N\) and reused in Section 6.2), sweeping the same \(\nu^2\) grid. Panel (): asymmetric \(N = 3\) market with \(\alpha = (1.5, 1.5, 0.15)\), \(\beta = (-1, -1, -0.05)\), and \([l, u] = [0.5, 2.5]\). The only swept coefficient is the coupling between the two strong sellers, \(\gamma_{12} = \gamma_{21} \in \{0.88, 0.91, 0.94, 0.96\}\); the cross-coefficients involving the third (small-\(\beta\)) seller are fixed at \(\gamma_{13} = \gamma_{23} = 0.025\) and \(\gamma_{31} = \gamma_{32} = 0.024\). The primitives are chosen so that every cell violates regularity ?? (\(\lambda_{\max}(B + B^\top) \in \{1.01, 1.03, 1.06, 1.08\}\), shown in the legend) while the spectral radius \(\rho(B) \in \{0.45, 0.47, 0.48, 0.49\}\) stays well below \(1\).

8.2.0.7 Figure 6 (mixed-market empirical verification).

All three panels: \(\eta_j = 0.25\) for every informed seller, \(T = 5\times 10^4\), \(S = 100\) seeds per cell. Panel (): \(5 \times 3\) grid of \((\gamma, \nu^2)\) in a symmetric \(N = 5\) market with \(|\cI^{ob}| = 2\), \(|\cI^{in}| = 3\), \((\alpha, \beta) = (2.5, -1)\). Panel (): \(3 \times 3\) grid of \((N, \nu^2)\) at \(N \in \{3, 5, 10\}\) with one oblivious and \(N - 1\) informed sellers, demand primitives drawn from the heterogeneous distribution introduced in Figure 5 (one draw per \(N\)). In both panels the projection-box-induced \(L_\phi^{ob}\) alone exceeds \(2\), so \(\bar\kappa < 0\) and conditions (ii) and (iv) fail jointly. Panel (): \(4 \times 3\) grid of \((\gamma, \nu^2)\) in a symmetric \(N = 5\) market with \(|\cI^{ob}| = 4\), \(|\cI^{in}| = 1\), \((\alpha, \beta, u) = (3.0, -2.0, 2.0)\); here \(\bar\kappa \in [0.73, 0.82]\) across cells (condition (ii) holds) while the small-gain margin of (iv) is in \([-2.27, -0.76]\) (condition (iv) fails).

8.2.0.8 Table 3 (strategy-game revenue summary).

Symmetric revenue duopoly with \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.6\), \([l, u] = [0.5, 3.5]\) (so \(\Pi^{NE} = 3.189\), \(\Pi^{C} = 3.906\); the wider Nash–collusive gap makes the composition effects more visible than the baseline \(\gamma = 0.4\) duopoly); \(S = 200\) seeds per cell; \(T = 60{,}000\). The obob cells use constant exploration \(\nu^2 \in \{0.20, 0.05\}\). The obin cells use the running-mean forecast with \(\nu_{n,1}^2 \equiv 0.10\) for the oblivious seller and \(\nu_{n,2}^2 = 0.10\,(n+1)^{-\eta}\) for \(\eta \in \{0.5, 0.25\}\) for the informed seller; the final obin row replaces the running-mean forecast with the clairvoyant perfect-prediction forecast \(\hat{p}_{n+1,1} = p_{n+1,1}\) at \(\eta = 0.5\). The inin cells use running-mean forecasting on both sellers with \(\nu_{n,i}^2 = 0.30\,(n+1)^{-\eta}\) and \(\eta \in \{0.5, 0.7\}\).

8.3 A Dynamic Benchmark↩︎

How should one evaluate a learning algorithm in a competitive market? A standard answer in the dynamic-pricing literature is a dynamic benchmark: the cumulative gap between the algorithm’s revenue and that of an oracle which knows the seller’s own demand function and best-responds to the realized competitor profile at every step [30], [37], [49], [50]. In this appendix we show that, in our linear-demand setting, the dynamic benchmark coincides up to a positive-definite weighting with the cumulative mean-square distance from the realized price path to the Nash equilibrium. Minimizing it is therefore equivalent to driving prices to the competitive outcome, which makes the benchmark ill-suited to studying tacit collusion, a setting in which prices are sustained above the competitive level by design.

8.3.0.1 The benchmark.

Consider the asymmetric \(N\)-seller market of Section 3.1 with linear demand 1 . Seller \(i\)’s expected revenue at price profile \(\mathbf{p}=(p_1,\ldots,p_N)\) is \[R_i(p_i, \mathbf{p}_{-i}) \coloneq p_i \!\left(\alpha_i + \beta_i p_i + \sum_{j\neq i}\gamma_{i,j}p_j\right),\] and seller \(i\)’s informed best-response oracle, viewed as a function of competitor prices \(\mathbf{q}_{-i}\), is \[\phi_i^{in}(\theta_i, \mathbf{q}_{-i}) \coloneq \frac{\alpha_i + \sum_{j\neq i}\gamma_{i,j}q_j}{-2\beta_i},\] where \(\theta_i = (\alpha_i, \beta_i, \{\gamma_{i,j}\}_{j\neq i})\) collects seller \(i\)’s demand parameters. The dynamic benchmark for seller \(i\) over horizon \(T\) is the cumulative expected revenue lost against this oracle: \[\Delta_i(\theta_i, T) = \sum_{n=1}^{T} \mathbb{E}_{n-1}\!\left[R_i\!\left(\phi_i^{in}(\theta_i, \mathbf{p}_{n,-i}), \mathbf{p}_{n,-i}\right) - R_i\!\left(p_{n,i}, \mathbf{p}_{n,-i}\right)\right].\] Equivalently, \(\Delta_i\) is the regret against hindsight-optimal own-prices when competitors’ prices are treated as exogenous. The next proposition expresses \(\Delta_i\) as a quadratic in the price-deviation vector \(\mathbf{p}_n - \mathbf{p}^{NE}\).

For every seller \(i\in[N]\) and every \(T\ge 1\), \[\Delta_i(\theta_i, T) = -\beta_i \sum_{n=1}^{T} \mathbb{E}_{n-1}\!\left\{\left[\sum_{j\neq i}\frac{\gamma_{i,j}}{-2\beta_i}\!\left(p_{n,j} - p_j^{NE}\right) - \!\left(p_{n,i} - p_i^{NE}\right)\right]^2\right\}.\] In particular, \(\Delta_i(\theta_i, T) \lesssim \sum_{n=1}^{T} \mathbb{E}_{n-1}\,\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2\).

Proof of Proposition [prop:dynamicbenchmark]. Denote \(p_{n,i}^* = \phi_i^{in}(\theta_i, \mathbf{p}_{n,-i})\). Expanding the per-period revenue gap, \[R_i\!\left(p_{n,i}^*, \mathbf{p}_{n,-i}\right) - R_i\!\left(p_{n,i}, \mathbf{p}_{n,-i}\right) = \alpha_i \left(p_{n,i}^* - p_{n,i}\right) + \beta_i \left[\left(p_{n,i}^*\right)^2 - p_{n,i}^2\right] + \sum_{j\neq i}\gamma_{i,j}p_{n,j}\!\left(p_{n,i}^* - p_{n,i}\right).\] The first-order condition \(\alpha_i + 2\beta_i p_{n,i}^* + \sum_{j\neq i}\gamma_{i,j} p_{n,j} = 0\) that defines \(p_{n,i}^*\) yields, after multiplying by \(p_{n,i}^*\), \[\alpha_i p_{n,i}^* + 2 \beta_i \left(p_{n,i}^*\right)^2 + \sum_{j\neq i}\gamma_{i,j} p_{n,j}\, p_{n,i}^* = 0.\] Substituting this identity and completing the square, \[\begin{align} &\quad \; R_i\!\left(p_{n,i}^*, \mathbf{p}_{n,-i}\right) - R_i\!\left(p_{n,i}, \mathbf{p}_{n,-i}\right) \\ &= \underbrace{\alpha_i p_{n,i}^* + 2 \beta_i \left(p_{n,i}^*\right)^2 + \sum_{j\neq i}\gamma_{i,j} p_{n,j}\, p_{n,i}^*}_{=0} - \alpha_i p_{n,i} - \beta_i \left[p_{n,i}^2 + \left(p_{n,i}^*\right)^2\right] - \sum_{j\neq i}\gamma_{i,j} p_{n,j}\, p_{n,i} \\ &= - \alpha_i p_{n,i} - \beta_i \left[p_{n,i}^2 + \left(p_{n,i}^*\right)^2\right] - \sum_{j\neq i}\gamma_{i,j} p_{n,j}\, p_{n,i} \\ &= - \beta_i \left(p_{n,i}^* - p_{n,i}\right)^2 -2 \beta_i p_{n,i}^* p_{n,i} - \alpha_i p_{n,i} - \sum_{j\neq i}\gamma_{i,j}p_{n,j} p_{n,i} \\ &= - \beta_i \left(p_{n,i}^* - p_{n,i}\right)^2 + 2 \beta_i p_{n,i} \underbrace{\left(\frac{\alpha_i + \sum_{j\neq i}\gamma_{i,j}p_{n,j}}{-2 \beta_i} - p_{n,i}^*\right)}_{=0} \\ &= - \beta_i \left[\phi_i^{in}\!\left(\theta_i, \mathbf{p}_{n,-i}\right) - \phi_i^{in}\!\left(\theta_i, \mathbf{p}^{NE}_{-i}\right) + p_i^{NE} - p_{n,i}\right]^2 \\ &= - \beta_i \left[\sum_{j\neq i}\frac{\gamma_{i,j}}{-2\beta_i}\left(p_{n,j} - p_j^{NE}\right) - \left(p_{n,i} - p_i^{NE}\right)\right]^2, \end{align}\] where the penultimate equality uses \(p_i^{NE} = \phi_i^{in}\!\left(\theta_i, \mathbf{p}^{NE}_{-i}\right)\) and the last equality uses the linearity of \(\phi_i^{in}\) in \(\mathbf{q}_{-i}\). Taking \(\mathbb{E}_{n-1}[\cdot]\) and summing over \(n\) gives the displayed identity. ◻

Proposition [prop:dynamicbenchmark] admits a compact expression that exposes the link to the Nash equilibrium. It turns out that this relationship can be made more precise by aggregating over all sellers. Define the deviation \(\mathbf{e}_n \triangleq \mathbf{p}_n - \mathbf{p}^{NE}\) and the row vector \(v_i \in \RR^{N}\) with \(v_{i,i} = -1\) and \(v_{i,j} = \gamma_{i,j}/(-2\beta_i)\) for \(j\neq i\). Then \[\label{eq:Delta-vi} \Delta_i(\theta_i, T) \;=\; -\beta_i \sum_{n=1}^{T} \mathbb{E}_{n-1}\!\left[\left(v_i^\top \mathbf{e}_n\right)^2\right].\tag{11}\] Stacking the rows yields the matrix \(V \triangleq D_\beta^{-1}\Gamma\), where \(D_\beta \triangleq \mathrm{diag}(-2\beta_i)\) and \(\Gamma\) is the matrix from Section 3.3 (\(\Gamma_{ii}=2\beta_i\), \(\Gamma_{ij}=\gamma_{i,j}\) for \(i\neq j\)). Each summand on the right-hand side of 11 measures only a one-dimensional projection of \(\mathbf{e}_n\), but aggregating over sellers recovers the full norm.

Corollary 1 (Aggregate dynamic regret). Let \(D \triangleq \mathrm{diag}(-\beta_i) \succ 0\). The matrix \(V^\top D V\) is positive definite, and for every \(T\ge 1\), \[\begin{align} \lambda_{\min}\!\left(V^\top D V\right) \sum_{n=1}^{T} \mathbb{E}_{n-1}\,\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2 &\;\le\; \sum_{i=1}^{N} \Delta_i(\theta_i, T) \\ &\;\le\; \lambda_{\max}\!\left(V^\top D V\right) \sum_{n=1}^{T} \mathbb{E}_{n-1}\,\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2. \end{align}\] In particular, \(\sum_{i=1}^{N}\Delta_i(\theta_i, T) \asymp \sum_{n=1}^{T}\EE\,\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2^2\).

Proof of Corollary 1. Summing 11 over \(i\), \[\sum_{i=1}^{N} \Delta_i(\theta_i, T) \;=\; \sum_{n=1}^{T} \mathbb{E}_{n-1}\!\left[\mathbf{e}_n^\top V^\top D V\, \mathbf{e}_n\right].\] Since \(-\beta_i>\gamma_i\) for all \(i\), \(\Gamma\) is strictly diagonally dominant and hence invertible; consequently \(V = D_\beta^{-1}\Gamma\) is invertible, and combined with \(D\succ 0\) this implies \(V^\top D V \succ 0\). The two-sided bound now follows from the elementary inequalities \[\lambda_{\min}(V^\top D V)\,\norm{\mathbf{e}_n}_2^2 \;\le\; \mathbf{e}_n^\top V^\top D V \,\mathbf{e}_n \;\le\; \lambda_{\max}(V^\top D V)\,\norm{\mathbf{e}_n}_2^2.\] ◻

The two-sided bound holds only for the aggregate \(\sum_i \Delta_i\), not for any individual \(\Delta_i\). To see why a per-seller lower bound fails, suppose seller \(i\) best-responds to the realized competitor profile, i.e., \(p_{n,i} = \phi_i^{in}(\theta_i, \mathbf{p}_{n,-i})\). Then \(v_i^\top \mathbf{e}_n = 0\) and \(\Delta_i(\theta_i, T) = 0\), yet \(\norm{\mathbf{p}_n - \mathbf{p}^{NE}}_2\) can be arbitrarily large because the competitors may sit far from \(\mathbf{p}_{-i}^{NE}\). Each \(v_i^\top \mathbf{e}_n\) probes one direction of \(\RR^N\); only the joint span of \(\{v_i\}_{i=1}^N\)—which is all of \(\RR^N\) exactly when \(\Gamma\) is invertible—recovers the full deviation vector.

8.3.0.2 Tension between the benchmark and collusion.

Two consequences of Corollary 1 clarify why the dynamic benchmark is poorly suited to questions of algorithmic collusion.

  • Sublinear aggregate regret pins down the Nash equilibrium. If \(\sum_{i=1}^{N} \Delta_i(\theta_i, T) = o(T)\), then \(\frac{1}{T}\sum_{n=1}^{T}\EE\,\norm{\mathbf{p}_n-\mathbf{p}^{NE}}_2^2 \to 0\), and hence the price path converges in mean square to \(\mathbf{p}^{NE}\).

  • Sustained supracompetitive prices imply linear aggregate regret. If the price path is on average bounded away from the Nash equilibrium—in particular, whenever the system sustains a collusive outcome that strictly Pareto-dominates \(\mathbf{p}^{NE}\) in terms of revenue—then \(\sum_{i=1}^{N}\Delta_i(\theta_i, T) = \Theta(T)\).

The dynamic benchmark therefore identifies the Nash equilibrium as the algorithm’s target by construction. Any market dynamic that earns supracompetitive revenues—the kind of dynamic that motivates the algorithmic-collusion literature—is, by the same construction, labeled as linearly regretful. Using this benchmark to evaluate learning in markets where supracompetitive pricing is a possibility therefore confounds the question being asked. This observation motivates our use of price convergence as the primary metric throughout the paper (see the discussion in Section 3.4). At the same time, Proposition [prop:dynamicbenchmark] shows that the mean-square convergence rates we derive in Sections 5 and 6 bound the learning part of the regret, so the two perspectives remain quantitatively connected.

8.4 Cross-Seller Propagation: Price Correlation and Modeling Error↩︎

This appendix unpacks the cross-seller residual terms \(r_{n,i\leftarrow j}\) summarized at the start of Section 4.2. Recall that, for sellers \(i \neq j\), \[r_{n,i\leftarrow j} \;=\; \frac{1}{J_{n,i}} \sum_{m=1}^{n} \left(p_{m,i}-\bar p_{n,i}\right)\left(p_{m,j}-\bar p_{n,j}\right), \qquad J_{n,i}=\sum_{m=1}^{n}\left(p_{m,i}-\bar p_{n,i}\right)^2,\] and the interaction-weighted aggregate of substitution effects is \[s_{n,i} \;\triangleq\; \sum_{j\neq i} \gamma_{i,j}\, r_{n,i\leftarrow j}.\] Ignoring noise terms and projections, recall from 6 that the next-period greedy price satisfies \[\tilde{p}_{n+1,i} \approx \frac{\alpha_i + \sum_{j\neq i}\gamma_{i,j}\bar p_{n,j} - \bar p_{n,i}s_{n,i}}{-2\left(\beta_i + s_{n,i}\right)}.\] It is then straightforward to verify that, when average demand is positive, larger values of \(s_{n,i}\) lead to higher next-period prices. Thus, large positive values of \(r_{n,i\leftarrow j}\) across important substitution links \(\gamma_{i,j}\) generate upward pricing pressure and may lead to supracompetitive outcomes. Conversely, negative values of \(r_{n,i\leftarrow j}\) generate downward pricing pressure and may lead to sub-Nash outcomes. In this sense, \(r_{n,i\leftarrow j}\) captures the degree to which seller \(i\)’s response strategy deviates from the full-information best response to seller \(j\)’s price path, and the sign and magnitude of this deviation directly influence the direction and strength of price adjustments.

8.4.0.1 Price correlation and tacit coordination.

Define the de-trended price vectors \[\mathbf{p}'_{n,i} = (p_{1,i},\ldots,p_{n,i})^\top - \bar p_{n,i}\mathbf{1}_n.\] Then, \[r_{n,i\leftarrow j} = \frac{\langle \mathbf{p}'_{n,i}, \mathbf{p}'_{n,j} \rangle}{\|\mathbf{p}'_{n,i}\|^2_2} = \frac{\|\mathbf{p}'_{n,j}\|_2}{\|\mathbf{p}'_{n,i}\|_2} \cos(\theta_{n, ij}),\] where \(\theta_{n,ij}\) is the angle between the two de-trended price paths. In particular, \[r_{n,i\leftarrow j}\, r_{n,j\leftarrow i} = \frac{\langle \mathbf{p}'_{n,i}, \mathbf{p}'_{n,j} \rangle^2}{\|\mathbf{p}'_{n,i}\|^2_2 \|\mathbf{p}'_{n,j}\|^2_2} = \mathrm{Corr}(\mathbf{p}_{n,i},\mathbf{p}_{n,j})^2 \in [0,1].\] Thus, pairwise products \(r_{n,i\leftarrow j} r_{n,j\leftarrow i}\) directly measure the degree of tacit coordination between sellers’ price movements. At the aggregate level, let \(\Sigma_n\) denote the empirical covariance matrix of de-trended prices across sellers up to time \(n\). Large values of \(r_{n,i\leftarrow j}\) across many pairs correspond to \(\Sigma_n\) developing a dominant low-rank structure. Economically, this reflects common-factor price experimentation: sellers’ exploratory price movements become aligned along a shared direction, whether due to similar algorithms, shared signals, or coincident responses to demand feedback. This low-rank structure is precisely what amplifies misspecification effects under oblivious learning: correlated experimentation increases the projection of omitted competitors’ prices onto a seller’s own price path, thereby inflating the magnitude of \(r_{n,i\leftarrow j}\) and feeding back into higher greedy prices. This mechanism closely parallels recent discussions of correlated price experiments in the algorithmic collusion literature [15], [19].

8.4.0.2 Modeling error allocation.

From a learning perspective, \(r_{n,i\leftarrow j}\) also governs how the omitted-variable bias induced by competitor \(j\) is allocated between the intercept and slope of seller \(i\)’s misspecified demand model. When \(\mathbf{p}'_{n,j}\) is nearly orthogonal to \(\mathbf{p}'_{n,i}\), the corresponding bias primarily shifts the intercept. When \(\mathbf{p}'_{n,j}\) aligns with \(\mathbf{p}'_{n,i}\), the bias loads onto the slope estimate, increasing the perceived own-price effect and hence the seller’s inferred market power. For illustration, consider \(N=2\) and \(\gamma_{1,2}=\gamma_{2,1}=\gamma\). As discussed in [2], if the slope parameter \(\beta\) is known, the intercept absorbs all modeling error and prices converge to the competitive outcome; if the intercept \(\alpha\) is known, the slope absorbs all modeling error and prices converge to the collusive outcome. In our notation, these correspond respectively to \(r_{n,i}\approx 0\) and \(r_{n,i}\approx 1\). Correlated price movements thus determine how omitted-variable bias is distributed across parameters, and high correlation—especially along economically important substitution links—induces supracompetitive pricing.

8.4.0.3 Regulatory insights.

These observations suggest two diagnostic tools for monitoring algorithmic collusion without the need to access proprietary learning algorithms or elicit information from sellers. First, regulators or platforms may track the correlation network of sellers’ prices, paying particular attention to the emergence of strong low-rank structure or large leading eigenvalues. Second, comparing sellers’ estimated own-price effects to benchmark estimates can reveal systematic inflation of perceived market power, especially when such inflation coincides with heightened price correlation. Together, these signals provide actionable indicators of tacit coordination arising from correlated price movements.

8.5 The Cost of Linear Exploration↩︎

In Section 4.2, we argued that rational oblivious sellers should avoid being variance-dominated by others. One simple way to achieve this is to add non-diminishing perturbations to prices, which guarantees linear growth of \(J_{n,i}\) by Lemma 10. However, a natural question arises: does such linear exploration come at a significant cost?

Let the index be \(\left\{1, 2\right\}\). For own price \(p\) and opponent price \(p_{j}\), denote the expected revenue function of seller \(i\) as \[R_i(p; p_{j}) = p (\alpha_i + \beta_i p + \gamma_{i,j} p_{j}).\] The following result quantifies, per-period and uniformly in the history, the expected-revenue cost of injecting a mean-zero perturbation into a seller’s pricing decision. The long-run analogue, applied to a learning protocol whose greedy price converges to \(\mathbf{p}^{NE}\) a.s., is sharpened to an exact asymptotic by Lemma 4.

Consider any pricing protocol for seller \(i\) of the form \[p_{n,i} \;=\; \tilde{p}_{n,i} \;+\; z_{n,i}, \qquad n \ge 1,\] where the underlying decision \(\tilde{p}_{n,i}\) and the opponent’s realized price \(p_{n,j}\) are \(\mathcal{F}_{n-1}\)-measurable and the perturbation \(z_{n,i}\) satisfies \[\mathbb{E}[z_{n,i} \mid \mathcal{F}_{n-1}] \;=\; 0, \qquad |z_{n,i}| \;\le\; \delta \quad \text{a.s.}\] Then, for every \(n\ge 1\), \[\mathbb{E}[R_i(\tilde{p}_{n,i} + z_{n,i};\, p_{n,j}) \mid \mathcal{F}_{n-1}] \;=\; R_i(\tilde{p}_{n,i};\, p_{n,j}) \;-\; |\beta_i|\,\mathbb{E}[z_{n,i}^2 \mid \mathcal{F}_{n-1}].\] In particular, the per-period expected revenue loss due to the perturbation is at most \(|\beta_i|\,\delta^2\), uniformly in \(n\), and vanishes as \(\delta\downarrow 0\).

Proof of Proposition [prop:costlinearexploration]. Because \(R_i(\cdot;\,p_{n,j})\) is a quadratic in its first argument with \(\partial_p^2 R_i \equiv 2\beta_i\), Taylor’s theorem gives the exact expansion \[R_i(\tilde{p}_{n,i} + z_{n,i};\, p_{n,j}) \;=\; R_i(\tilde{p}_{n,i};\, p_{n,j}) \;+\; z_{n,i}\,\frac{\partial R_i}{\partial p}(\tilde{p}_{n,i};\, p_{n,j}) \;+\; \beta_i\, z_{n,i}^2.\] Taking conditional expectation given \(\mathcal{F}_{n-1}\) and using that \(\tilde{p}_{n,i}, p_{n,j}\in\mathcal{F}_{n-1}\) together with \(\mathbb{E}[z_{n,i}\mid\mathcal{F}_{n-1}]=0\), the linear term vanishes. Substituting \(\beta_i = -|\beta_i|\) yields the claimed identity. The per-period magnitude bound follows from \(\mathbb{E}[z_{n,i}^2\mid\mathcal{F}_{n-1}]\le \delta^2\). ◻

8.6 Empirical Threshold and Robustness Checks for Theorem 2↩︎

This appendix complements the stress test in Figure 2 with two additional analyses in the symmetric duopoly with \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.4\), \([l, u] = [0.5, 2.5]\): (i) a 16-point logarithmically spaced sweep in the exploration variance \(\nu^2 \in [10^{-3}, 0.3]\), used to pin down the empirical convergence threshold; and (ii) a robustness check that replaces the uniform dithering \(z_{n,i} \sim \mathrm{Unif}(-c, c)\) with a Gaussian-clip scheme \(z_{n,i} \sim \mathcal{N}(0, \nu^2)\), truncated at \(\pm 4\sigma\). In both cases we report the tail log–log slope of \(\mathrm{MSE}(\tilde{\mathbf{p}}_n)\) (fitted over the last \(25\%\) of a horizon \(T = 6\times 10^4\) run with \(S = 80\) seeds per cell), and—for the \(\nu^2\) sweep—the closed-form upper bound on the persistent-excitation constant \(C_M(\nu^2) \le \tfrac{1}{2}(1 + p_*^2 + \nu^2 - \sqrt{(1 + p_*^2 + \nu^2)^2 - 4\nu^2})\) evaluated at \(p_* = p^{NE}\).

8.6.0.1 Empirical threshold for convergence.

Table 6 shows that the tail \(\log\mathrm{MSE}(\tilde{\mathbf{p}}_n)\) vs.\(\log n\) slope crosses \(-1\) at \(\nu^2_{\mathrm{emp}} \approx 0.014\) and stays at the asymptotic \(n^{-1}\) rate thereafter; below \(\nu^2_{\mathrm{emp}}\) the system is trapped on the pseudo-equilibrium continuum and the slope flattens. At the empirical threshold the closed-form bound gives \(C_M(\nu^2_{\mathrm{emp}}) \approx 4 \times 10^{-3}\), while Theorem 2’s formal RHS \(\bar\gamma\, L_\phi^{ob}\, C_x \approx 10\) for this market: the formal sufficient threshold is roughly \(2{,}500\times\) larger than the empirically allowable one.

Table 6: Empirical regime transition along a 16-point log-spaced \(\nu^2\) grid in the symmetric duopoly with \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.4\), \([l, u] = [0.5, 2.5]\). The tail log–log slope of \(\mathrm{MSE}(\tilde{\mathbf{p}}_n)\) (column 3) crosses \(-1\) at \(\nu^2_{\mathrm{emp}} \approx 0.014\) (bolded row) and stays at the asymptotic \(n^{-1}\) rate thereafter. Theorem 2’s formal threshold, by contrast, requires \(\bar\gamma\, L_\phi^{ob}\, C_x \approx 10\) to be smaller than \(C_M(\nu^2)\) (column 4), which is at most \(\approx 0.08\) for any \(\nu^2\) that keeps prices in the box \([l, u]\): the formal sufficient threshold is roughly \(2{,}500\times\) larger than the empirically allowable threshold at the transition. Horizon \(T = 6\times 10^4\), \(S = 80\) seeds per cell.
\(\nu^2\) slope \(\mathrm{MSE}(\hat{\theta})\) slope \(\mathrm{MSE}(\tilde{\mathbf{p}})\) \(C_M(\nu^2)\) (upper bound)
\(0.0010\) \(-0.02\) \(-0.06\) \(2.91\times 10^{-4}\)
\(0.0015\) \(-0.09\) \(-0.12\) \(4.25\times 10^{-4}\)
\(0.0021\) \(-0.06\) \(-0.06\) \(6.21\times 10^{-4}\)
\(0.0031\) \(-0.07\) \(-0.07\) \(9.09\times 10^{-4}\)
\(0.0046\) \(-0.09\) \(-0.09\) \(1.33\times 10^{-3}\)
\(0.0067\) \(-0.38\) \(-0.33\) \(1.94\times 10^{-3}\)
\(0.0098\) \(-0.69\) \(-0.62\) \(2.84\times 10^{-3}\)
\(\boldsymbol{0.0143}\) \(\boldsymbol{-1.01}\) \(\boldsymbol{-0.87}\) \(\boldsymbol{4.15\times 10^{-3}}\)
\(0.0210\) \(-1.21\) \(-1.18\) \(6.06\times 10^{-3}\)
\(0.0306\) \(-1.10\) \(-1.22\) \(8.85\times 10^{-3}\)
\(0.0448\) \(-1.02\) \(-1.19\) \(1.29\times 10^{-2}\)
\(0.0656\) \(-0.99\) \(-1.15\) \(1.88\times 10^{-2}\)
\(0.0959\) \(-0.99\) \(-1.11\) \(2.73\times 10^{-2}\)
\(0.1402\) \(-1.01\) \(-1.08\) \(3.96\times 10^{-2}\)
\(0.2051\) \(-1.04\) \(-1.05\) \(5.71\times 10^{-2}\)
\(0.3000\) \(-1.07\) \(-1.03\) \(8.20\times 10^{-2}\)

8.6.0.2 Gaussian-clip dithering.

Table 7 repeats the same constant-\(\nu^2\) sweep with Gaussian-clip dithering (\(z_{n,i} \sim \mathcal{N}(0, \nu^2)\), truncated at \(\pm 4\sigma\)). The transition pattern is qualitatively unchanged: the tail log–log slope is still well below \(-1\) for \(\nu^2 \le 0.005\) and stabilizes at \(\approx -1\) for \(\nu^2 \ge 0.025\), reproducing the uniform-dithering sweep of Table 6. The conclusions of Theorem 2 and the spiral-up discussion in §4.2 therefore do not depend on the particular shape of the dithering distribution.

Table 7: Gaussian-clip dithering with \(z_{n,i} \sim \mathcal{N}(0, \nu^2)\) truncated at \(\pm 4\sigma\), symmetric duopoly with \(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.4\), \([l, u] = [0.5, 2.5]\). The same constant-\(\nu^2\) transition as Table 6 is reproduced: tail \(\log\mathrm{MSE}\) slope flat for \(\nu^2 \le 0.005\), then settles at \(\approx -1\) for \(\nu^2 \ge 0.025\). Horizon \(T = 6\times 10^4\), \(S = 80\) seeds per cell.
\(\nu^2\) slope \(\mathrm{MSE}(\hat{\theta})\) slope \(\mathrm{MSE}(\tilde{\mathbf{p}})\)
\(0.005\) \(-0.45\) \(-0.44\)
\(0.010\) \(-0.76\) \(-0.67\)
\(0.025\) \(-1.14\) \(-1.06\)
\(0.050\) \(-1.09\) \(-1.12\)
\(0.100\) \(-1.03\) \(-1.10\)
\(0.200\) \(-0.94\) \(-1.06\)
\(0.300\) \(-0.79\) \(-0.93\)

8.7 Mean-Dynamics ODE: Empirical Moments and Derivation↩︎

This appendix derives the mean-dynamics ODE 8 used in Section 5.2. Throughout we assume the dithering terms \(z_{n,i}\) are i.i.d.across \(n\) and \(i\), independent of \(\{\varepsilon_{n,i}\}\), with \(\EE[z_{n,i}]=0\) and \(\Var(z_{n,i})=\nu^2>0\).

We describe the price process through its empirical first and second moments: \[m_{n,i}\triangleq \frac{1}{n}\sum_{s=1}^n p_{s,i},\qquad i\in[N], \qquad Q_{n,ij}\triangleq \frac{1}{n}\sum_{s=1}^n p_{s,i}p_{s,j},\qquad i,j\in[N].\] Let \(m_n=(m_{n,1},\ldots,m_{n,N})\in\RR^N\) and \(Q_n=(Q_{n,ij})_{i,j}\in\RR^{N\times N}\). These moments evolve by averaging: \[\label{eq:moment-recursion-N} m_{n+1,i}=m_{n,i}+\frac{1}{n+1}\left(p_{n+1,i}-m_{n,i}\right),\qquad Q_{n+1,ij}=Q_{n,ij}+\frac{1}{n+1}\left(p_{n+1,i}p_{n+1,j}-Q_{n,ij}\right).\tag{12}\] Define the centered covariance and variance (as functions of \((m,Q)\)): \[S_{ij}(m,Q)\triangleq Q_{ij}-m_i m_j\quad(i\neq j),\qquad V_i(m,Q)\triangleq Q_{ii}-m_i^2.\] Under the true demand model 1 with \(\EE[\varepsilon_{n,i}]=0\), the population regression coefficients of the misspecified regression \(d_i\sim a_i+b_i p_i\) satisfy \[b_i(m,Q)=\frac{\Cov(p_i,d_i)}{\Var(p_i)},\qquad a_i(m,Q)=\EE[d_i]-b_i(m,Q)\EE[p_i].\] A direct calculation yields the closed-form maps \[\label{eq:ab-map-N-ode} b_i(m,Q)=\beta_i+\sum_{j\neq i}\gamma_{i,j}\frac{S_{ij}(m,Q)}{V_i(m,Q)},\qquad a_i(m,Q)=\alpha_i+\sum_{j\neq i}\gamma_{i,j}m_j \;-\;m_i\sum_{j\neq i}\gamma_{i,j}\frac{S_{ij}(m,Q)}{V_i(m,Q)}.\tag{13}\] Note that \(V_i(m,Q)\) is bounded away from \(0\) over long horizons by persistent excitation and Lemma 8. We then define the mean-dynamics greedy price \[\label{eq:pg-map-N-ode} p_i^g(m,Q)\triangleq -\frac{a_i(m,Q)}{2b_i(m,Q)},\qquad i\in[N],\tag{14}\] and let \(p^g(m,Q)\in\RR^N\) denote the resulting price vector.

Let \(\mathcal{F}_n\) denote the filtration generated by the entire price–demand history up to time \(n\). Under the pricing rule \[p_{n+1,i}=p_i^g(m_n,Q_n)+z_{n+1,i},\] and the independence assumptions on \(z_{n+1,i}\), we have \[\EE[p_{n+1,i}\mid \mathcal{F}_n]=p_i^g(m_n,Q_n).\] Moreover, using independence across sellers and \(\Var(z_{n+1,i})=\nu^2\), \[\EE[p_{n+1,i}p_{n+1,j}\mid \mathcal{F}_n]= \begin{cases} p_i^g(m_n,Q_n)p_j^g(m_n,Q_n), & i\neq j,\\ (p_i^g(m_n,Q_n))^2+\nu^2, & i=j. \end{cases}\] Taking conditional expectations in 12 motivates viewing \((m_n,Q_n)\) as a stochastic approximation recursion with step size \(1/(n+1)\) and drift determined by \(p^g(m,Q)\). The associated mean-dynamics ODE is 8 . Under the standard conditions of the ODE method for stochastic approximation (bounded iterates, diminishing step sizes, and locally Lipschitz drift), the piecewise-linear interpolation of \((m_n,Q_n)\) tracks the flow of 8 on long time scales [54].7

8.8 Short-Run Excursions in the Oblivious Duopoly↩︎

This appendix formalizes the short-run excursion phenomenon discussed in Section 5.3.1 and illustrated in Figure 3. Throughout, we focus on the duopoly \(N=2\) with symmetric demand parameters.

8.8.0.1 Five-dimensional duopoly ODE.

Specializing the mean-dynamics ODE 8 to \(N=2\) yields the five-dimensional moment system \[\label{eq:duopoly-mQ-ode} \begin{align} \dot{m}_i(t) &= p_i^g(m(t),Q(t)) - m_i(t), \qquad i=1,2,\\ \dot{Q}_{i,j}(t) &= p_i^g(m(t),Q(t))\,p_j^g(m(t),Q(t)) - Q_{i,j}(t),\qquad i \ne j,\\ \dot{Q}_{i,i}(t) &= (p_i^g(m(t),Q(t)))^2 + \nu^2 - Q_{i,i}(t),\qquad i=1,2. \end{align}\tag{15}\] Recall the (symmetric) competitive and collusive benchmarks in the duopoly case: \[p^{NE}=\frac{\alpha}{-2\beta-\gamma},\qquad p^C=\frac{\alpha}{-2\beta-2\gamma}, \qquad l<p^{NE}<p^C<u.\] By Theorem 3, 15 admits a unique (locally asymptotically stable) equilibrium \[m^* = (p^{NE}, p^{NE}), \qquad Q^*_{i,j} = (p^{NE})^2, \quad i\neq j, \qquad Q^*_{i,i} = (p^{NE})^2 + \nu^2, \quad i=1,2.\]

8.8.0.2 Reduced ODE on the symmetric moment manifold.

Finite-time dynamics of nonlinear ODEs can be complicated to delineate even for the simplified ODE 15 . We therefore take a two-fold approach: we first theoretically characterize the structure of trajectories on a further reduced symmetric moment manifold, and then we empirically verify that the patterns predicted by the reduced system are indeed present in the five-dimensional system.

The convergent patterns established in Section 5 imply that, over long horizons, the two sellers’ prices tend to align, i.e., the empirical moments should approximately satisfy \(m_1\approx m_2\) and \(Q_{11}\approx Q_{22}\). Thus, the symmetric moment manifold \[m_1=m_2=:y_1,\qquad Q_{11}=Q_{22}=:y_2,\qquad Q_{12}=:y_3,\] should capture key aspects of finite-time dynamics without overly complicated characterization of the dissipating inter-seller differences. Restricting to this manifold reduces the five-dimensional duopoly moment system to a three-dimensional ODE for \(\mathbf{y}=(y_1,y_2,y_3)\): \[\label{eq:reducedsystem} \dot{\mathbf{y}}_t = \begin{bmatrix} B(y_{t,1},r(\mathbf{y}_t)) - y_{t,1}\\[2pt] B(y_{t,1},r(\mathbf{y}_t))^2+\nu^2-y_{t,2}\\[2pt] B(y_{t,1},r(\mathbf{y}_t))^2-y_{t,3} \end{bmatrix},\qquad t\ge 0,\tag{16}\] where the greedy price map is \[\label{eq:reduced-greedy} B(y_1,r)\;\triangleq\;-\frac{a(\mathbf{y})}{2b(\mathbf{y})} \;=\; \frac{\alpha+\gamma(1-r)y_1}{-2\beta-2\gamma r}\tag{17}\] and the variance and covariance proxies are \[V(\mathbf{y})\triangleq y_2-y_1^2,\qquad C(\mathbf{y})\triangleq y_3-y_1^2,\qquad r(\mathbf{y})\triangleq \frac{C(\mathbf{y})}{V(\mathbf{y})}.\] We assume \(V(\mathbf{y}_0)>0\) so that \(r(\mathbf{y}_t)\) is well-defined initially.8 We also assume \(\alpha+(\beta+\gamma)u>0\), a technical condition meaning that the demand remains positive when both sellers set the upper price bound \(u\). Following Theorem 3, the reduced system 16 admits a unique (locally asymptotically stable) equilibrium \[\label{eq:3dode-equilibrium} \mathbf{y}^*=\left(p^{NE},(p^{NE})^2+\nu^2,(p^{NE})^2\right).\tag{18}\]

8.8.0.3 Upward excursions: a single overshoot.

Let \(y_i(t)\) denote the coordinates of \(\mathbf{y}_t\). The next proposition provides a finite-time characterization of trajectories: the mean price either converges directly to \(p^{NE}\), or it makes a single supracompetitive excursion before converging back. The proof is in Appendix 9.6.

The following hold for any solution \(\mathbf{y}_t\) of 16 :

  1. \(V(t)>0\) for all \(t\ge 0\), hence \(r(t)\) is well-defined for all \(t\ge 0\).

  2. There exists \(T_{r, C}<\infty\) such that \(r(t)<1\) and \(y_1(t)<p^C\) for all \(t\ge T_{r, C}\).

  3. Either (a) \(\mathbf{y}(t)\) converges directly to \(\mathbf{y}^*\); or (b) there exists \(T_{NE}<\infty\) such that \(\dot{y}_1(T_{NE}) > 0\) and \(r(t) \in (0,1)\) and \(y_1(t)\in(p^{NE},p^C)\) for all \(t\ge T_{NE}\).

  4. If \(\dot{y}_1(t_0)=0\) and \(r(t_0)>0\) for some time \(t_0\), then \(\ddot y_1(t_0)<0\). Consequently, in case (iii.b), the function \(t\mapsto y_1(t)\) has exactly one local maximum on \([T_{NE},\infty)\) and no local minima. After attaining this maximum, \(y_1(t)\) decreases monotonically to \(p^{NE}\).

8.8.0.4 Discussion.

Proposition [prop:reduced-excursion] shows that the reduced ODE admits two modes of convergence to the competitive equilibrium: either the mean price converges directly to \(p^{NE}\) (case (iii.a)), or it exhibits a single supracompetitive excursion—it increases above \(p^{NE}\), attains exactly one local maximum in \((p^{NE},p^C)\), and then decreases monotonically and converges to \(p^{NE}\) (case (iii.b)+(iv)). The latter captures a collude-in-the-short-run, compete-in-the-long-run pattern: prices can be transiently supracompetitive even though the unique long-run equilibrium is competitive. The same overshoot–then–convergence pattern appears in the full five-dimensional duopoly ODE 15 and in the original discrete-time dynamics 6 (Figure 3, panels (a)–(b)). The mean-dynamics ODE captures the drift of the empirical moments under the stochastic-approximation scaling with step size \(1/(n+1)\). Thus, ODE time \(t\) corresponds to the cumulative step size \(\sum_{k\le n}\frac{1}{k}\approx \log n\) in discrete time. Consequently, an \(O(1)\)-length excursion in ODE time can translate into a long window of discrete periods, making transient supracompetitive pricing economically relevant even when the eventual limit is \(p^{NE}\).

8.8.0.5 Downward excursions and excursion direction.

The existence of an upward excursion alone does not imply that oblivious learning is a reliable collusion device. In fact, if we no longer restrict the five-dimensional duopoly ODE 15 to the symmetric manifold, the same dynamics can generate downward excursions in which mean prices temporarily fall below \(p^{NE}\) before recovering (Figure 3, panels (c)–(d)). Such episodes are economically harmful: when both sellers price below \(p^{NE}\), both could improve profits by switching to the competitive benchmark. The coexistence of upward and downward excursions raises a core strategic question: can oblivious learning reliably control the direction of the excursion? The answer is no.

To understand why both directions are possible when the system is not restricted to the symmetric manifold, consider the centered covariance, centered variances, and the associated covariance–variance ratios \[C(t)\triangleq Q_{12}(t)-m_1(t)m_2(t), \quad V_i(t)\triangleq Q_{ii}(t)-m_i(t)^2, \quad r_i(t)\triangleq \frac{C(t)}{V_i(t)},\quad i=1,2.\] The terms \(V_i(t)\) play the same role as the discrete-time information variables \(J_{n,i}\). The identity \(\dot{V}_i(t)=\dot{m}_i(t)^2+\nu^2 - V_i(t)\) implies that, after an initial transient, \(V_i(t)\) is bounded away from zero because \(\nu^2>0\) provides persistent excitation. Thus, \(r_i(t)\) is well-defined. Moreover, \(r_1(t)\) and \(r_2(t)\) share the same numerator \(C(t)\), so their signs coincide.

The key object is \(C(t)\). A direct calculation gives \[\dot{C}(t) \;=\; \dot{m}_1(t)\,\dot{m}_2(t) - C(t),\] i.e., \(C(t)\) is a stable first-order filter driven by the signed “input” \(\dot{m}_1(t)\dot{m}_2(t)\). Under a mild positive-demand condition \(\alpha+\beta u+\gamma l>0\), the misspecified greedy map is locally increasing in \(r_i\): larger (positive) \(r_i\) biases the oblivious best response upward, while smaller (negative) \(r_i\) biases it downward. In the symmetric reduction, the two sellers’ mean deviations coincide, so \(\dot{m}_1(t) \dot{m}_2(t)=\dot{y}_1(t)^2\ge 0\) and \(C(t)\) cannot be persistently negative. In the full 5-D dynamics, however, the two sellers’ mean deviations can differ in sign, yielding a switching logic:

  • If \(\dot{m}_1(t)\) and \(\dot{m}_2(t)\) move in the same direction early on, then \(\dot{m}_1(t)\dot{m}_2(t)\) is predominantly positive, pushing \(C(t)\) upward and making \(r_i(t)>0\), which can generate an upward excursion.

  • If \(\dot{m}_1(t)\) and \(\dot{m}_2(t)\) move in opposite directions early on, then \(\dot{m}_1(t)\dot{m}_2(t)\) is primarily negative, pushing \(C(t)\) downward and making \(r_i(t)<0\), which can generate a downward excursion.

The same co-movement mechanism is present in discrete time: the empirical ratios \(r_{n,i\leftarrow j}\) summarize whether seller \(i\)’s centered price variations align with seller \(j\)’s, and this alignment is driven by early transients and realized exploration shocks. The excursion direction is therefore governed by the sign of an endogenous co-movement statistic that is shaped by early-time transients and is sample-path dependent in the underlying stochastic dynamics due to both demand noise and price exploration. Consequently, oblivious learning cannot reliably induce supracompetitive behavior to deliver extra revenue.

8.9 Informed OLS Rate in Mixed Markets↩︎

Consider a market in which at least one seller is informed, and let \(\mathcal{I}^{in}\subseteq\{1,\ldots,N\}\) denote the set of informed sellers. Without loss of generality, assume that the first \(N+1\) prices induce a full-rank empirical Fisher information matrix for each informed seller, allowing least squares thereafter. We work under the heterogeneous-exploration setup introduced in the opener of Section 6—in particular, \(\nu_{n,i}^2 = \Theta(n^{-\eta_i})\) with \(\eta_i \in [0,1)\), and \(\eta_{\min}\), \(\eta_{\max}\) as defined there—and denote the true (correctly specified) demand parameter by \[\theta^*_i=(\alpha_i,\beta_i,\gamma_{i,1},\ldots,\gamma_{i,N})^\top \in\mathbb{R}^{N+1}, \qquad i\in\mathcal{I}^{in}.\]

Theorem 6 (Informed OLS rate in mixed markets). Suppose condition ?? holds. Then, \[\sum_{i \in \mathcal{I}^{in}} \mathbb{E}\,\|\hat{\theta}_{n,i}^{in}-\theta^*_i\|_2^2 = O\!\left(n^{\eta_{\max}-1}\right).\]

The proof is in Appendix 9.11. Theorem 6 shows that, despite the presence of oblivious competitors whose pricing behavior may be misspecified and potentially erratic, informed sellers achieve complete learning of the true demand model: by correctly specifying the demand model, informed sellers treat competitors’ price movements as informative covariates rather than as unmodeled shocks, and thus convert market variation into identification.

8.10 Forecast-Rule Ablations↩︎

In Section 6.2, every informed seller adopts the running-mean forecast \(\hat{p}_{n+1,k} = m_{n,k}\) for each competitor \(k\). This appendix compares it against three additional one-step forecast patterns. For tractability of the closed-form analysis below, we work in a mixed duopoly (\(N = 2\)) used as a tractable test bed: seller 1 is oblivious with persistent \(\nu_1^2 \equiv 0.10\), seller 2 is informed with decaying \(\nu_{n,2}^2 = 0.10\, n^{-1/2}\), \(S = 200\) seeds, horizon \(T = 60{,}000\). The four forecast rules that seller 2 could use for seller 1’s next-period price are:

  • running mean: \(\hat{p}_{n+1,1} = m_{n,1}\), the running mean of seller 1’s realized prices (Section 6.2).

  • lag-1: \(\hat{p}_{n+1,1} = p_{n,1}\), the most recent realized price, which inherits seller 1’s previous exploration realization \(z_{n,1}\).

  • greedy component: \(\hat{p}_{n+1,1} = \tilde{p}_{n+1,1}\), a clairvoyant rule that reads seller 1’s deterministic best-response component before exploration noise is added; not implementable in practice but useful as a noise-free reference point.

  • perfect prediction: \(\hat{p}_{n+1,1} = p_{n+1,1} = \tilde{p}_{n+1,1} + z_{n+1,1}\), a clairvoyant rule that reads seller 1’s full realization, including its exploration noise, before committing seller 2’s price; not implementable in practice but serves as an informational upper bound.

a
b
c
d

Figure 7: Cross-seed scatter of the seed-by-seed running-mean prices \((\bar p_{T,1}, \bar p_{T,2})\) at horizon \(T = 60{,}000\) for each forecast rule, \(S = 200\) seeds. Demand primitives match the mixed-duopoly cell of Table 3 (\(\alpha = 2.5\), \(\beta = -1\), \(\gamma = 0.6\), \([l, u] = [0.5, 3.5]\)); all four panels share the same axis limits and mark the Nash price \(\mathbf{p}^{NE}\), the collusive price \(\mathbf{p}^{C}\), and the Stackelberg price \(\mathbf{p}^*\) (Equation 19 ) computed in closed form from the same primitives. The running-mean, lag-1, and greedy-component rules (panels –) all cluster tightly around \(\mathbf{p}^{NE}\), while the clairvoyant perfect-prediction rule (panel ) clusters around the Stackelberg point with the informed seller as follower.. a — running mean., b — lag-1., c — greedy component., d — perfect prediction.

Table 8: Seed-averaged running-mean prices, per-period revenues, and surplus-capture ratios at horizon \(T = 60{,}000\) for the four forecast rules in the mixed duopoly. Seller 1 is oblivious, seller 2 is informed. Benchmarks: \(\Pi^{NE} = (2.441, 2.441)\), \(\Pi^{C} = (2.604, 2.604)\); the Stackelberg revenues are \(\Pi^* = (2.446, 2.484)\). Across \(S = 200\) seeds, the first three rules sit at the Nash limit price; lag-1 incurs a small per-period revenue penalty for the informed seller (the Jensen tax of 20 below); perfect-prediction shifts the limit to the Stackelberg point and lifts the informed seller’s surplus above \(\Pi^{NE}\).
rule \(\bar p_1\) \(\bar p_2\) \(\bar R_1\) \(\bar R_2\) \(S_1\) \(S_2\)
running mean \(1.564\) \(1.563\) \(2.341\) \(2.441\) \(-0.615\) \(-0.004\)
lag-1 \(1.564\) \(1.563\) \(2.341\) \(2.437\) \(-0.614\) \(-0.025\)
greedy component \(1.564\) \(1.563\) \(2.341\) \(2.441\) \(-0.615\) \(-0.004\)
perfect prediction \(1.632\) \(1.576\) \(2.353\) \(2.487\) \(-0.542\) \(+0.281\)

8.10.0.1 Limit prices: three near-Nash rules.

The common feature of the running-mean, lag-1, and greedy-component rules is that none of them uses the competitor’s next-period exploration noise \(z_{n+1,1}\). The price limit in all three cases is thus the Nash equilibrium \(\mathbf{p}^{NE}\), consistent with Theorem 5, and the empirical clusters in Figure 7 confirm this: the seed-by-seed final-period running means concentrate tightly around \(\mathbf{p}^{NE}\) in all three panels.

8.10.0.2 Perfect-prediction limit: Stackelberg.

The clairvoyant perfect-prediction rule, by contrast, reads the competitor’s exploration noise realization \(z_{n+1,1}\) and uses it in seller 2’s best response. A standard martingale argument shows that the empirical co-movement statistic \[r_n \triangleq \frac{\sum_{m=1}^n (p_{m,1}-\bar p_{n,1})(p_{m,2}-\bar p_{n,2})}{\sum_{m=1}^n (p_{m,1}-\bar p_{n,1})^2}\] converges to \(r^\ast = \gamma_{2,1}/(-2\beta_2)\) a.s., exactly the slope of seller 2’s best-response map. Setting up the linearized recursion for \((\bar p_{n,1}, \bar p_{n,2})\) around the fixed point and solving the limiting affine system yields the Stackelberg price profile with seller 1 (oblivious) as leader and seller 2 (informed) as follower: \[\label{eq:stackelberg-prices-app} p^*_1 \;=\; \frac{-2\beta_2\alpha_1+\gamma_{1,2}\alpha_2}{4\beta_1\beta_2-2\gamma_{1,2}\gamma_{2,1}}, \qquad p^*_2 \;=\; \phi_2^{in}(p^*_1) \;=\; \frac{\alpha_2+\gamma_{2,1}p^*_1}{-2\beta_2}.\tag{19}\] A direct check of the \(2\times 2\) Hurwitz conditions on the linearization at \((p^*_1, p^*_2)\) gives \(\mathrm{tr} = -2\) and \(\det = 1\), so the fixed point is locally asymptotically stable and the running means converge to \((p^*_1, p^*_2)\) a.s. In a symmetric duopoly (\(\alpha_i \equiv \alpha\), \(\beta_i \equiv \beta\), \(\gamma_{i,j} \equiv \gamma\)), one obtains the ordering \(p^{NE} < p^*_2 < p^*_1 < p^C\) and the deterministic limit-revenue ordering \(\Pi^{NE} < \Pi^*_1 < \Pi^*_2 < \Pi^C\): joint profits at the limit prices exceed Nash, the informed follower’s limit revenue is strictly above the oblivious leader’s, and even the leader’s limit revenue is nominally above Nash. The follower’s gain is large enough to register cleanly in the realized table entry \(S_2 = +0.28\), the empirical signature of Stackelberg lock-in. The leader, by contrast, posts \(S_1 = -0.54\) in the same table even though \(\Pi^*_1 > \Pi^{NE}\), because the deterministic Stackelberg gain \(\Pi^*_1 - \Pi^{NE}\) is tiny in this symmetric duopoly cell (\(2.446 - 2.441 = 0.005\)) while the oblivious leader continues to pay a persistent exploration tax of order \(|\beta_1|\,\nu_1^2 = 0.10\) on its realized revenue (Proposition [prop:costlinearexploration]); the tax dwarfs the gain by an order of magnitude, and the leader’s realized revenue therefore sits below \(\Pi^{NE}\) rather than slightly above \(\Pi^*_1\). The perfect-prediction rule is, however, not implementable in practice—it requires real-time observation of the competitor’s i.i.d.exploration realization—so we record this limit only as a conceptual upper bracket on what an informed seller could earn against an oblivious competitor.

8.10.0.3 Lag-1 Jensen tax.

Among the three near-Nash rules, the lag-1 rule pays a small per-period revenue penalty visible in Table 8 (\(\bar R_2 = 2.437\) vs. \(2.441\) for the running mean rule; \(S_2 = -0.025\) vs.\(-0.004\)). Heuristically, the informed best response \(\phi_2^{in}(p_{n,1}) = (\alpha_2 + \gamma_{2,1}\, p_{n,1})/(-2\beta_2)\) is linear in \(p_{n,1}\), so plugging the noisy realized price \(p_{n,1} = \tilde{p}_{n,1} + z_{n,1}\) inherits the exploration variance \(\nu_{n,1}^2 = \Var(z_{n,1})\) in the next-period price \(p_{n+1,2}\). Revenue \(R_2 = p_{n+1,2}\,(\alpha_2 + \beta_2 p_{n+1,2} + \gamma_{2,1} p_{n+1,1})\) is concave-parabolic in \(p_{n+1,2}\), so by Jensen’s inequality the variance translates into a persistent revenue penalty. At Nash, expanding the three moments \(\EE[p_{n+1,2}], \EE[p_{n+1,2}^2], \EE[p_{n+1,2}\, p_{n+1,1}]\) around \((p_1^{NE}, p_2^{NE})\) and using the Nash identity \(\alpha_2 + 2\beta_2 p_2^{NE} + \gamma_{2,1} p_1^{NE} = 0\) gives the asymptotic decomposition \[\label{eq:lag1-jensen-app} \EE\!\left[R_{2,n+1}\right] \;=\; \Pi_2^{NE} \;-\; \frac{\gamma_{2,1}^2}{4\,|\beta_2|}\, \nu_{n,1}^2 \;-\; |\beta_2|\, \nu_{n+1,2}^2 \;+\; o(1).\tag{20}\] The second penalty term \(|\beta_2|\, \nu_{n+1,2}^2\) is the informed seller’s own-exploration tax and vanishes whenever \(\nu_{n,2}^2 \downarrow 0\), which holds under the decaying schedule. The first penalty term, however, is paid on the competitor’s persistent exploration variance and does not vanish in the spiral-up regime where \(\nu_{n,1}^2 \equiv \nu^2 > 0\). In contrast, the running-mean rule de-noises \(z_{n,1}\) before computing seller 2’s best response, so it does not inherit this variance term and the per-period informed revenue collapses to \(\Pi_2^{NE}\) in expectation. This is the structural Jensen cost of forecasting noisy realized prices rather than a de-noised average, and it is the source of the small \(S_2\) gap between the lag-1 rule and the running-mean rule in Table 8.

8.10.0.4 Takeaway.

The four rules bracket the canonical running mean from both sides. On the implementable side, the lag-1 rule sits below: it preserves the Nash limit price but inherits the competitor’s persistent exploration variance as a Jensen tax. On the clairvoyant side, the greedy-component rule matches the running-mean’s Nash limit (and its zero Jensen tax) but requires reading the competitor’s deterministic best-response component before noise. The qualitative ordering \(S^{in} > S^{ob}\) from Section 7 survives every rule—including lag-1, where the gap is only narrowed, not closed—further suggesting that the strategic conclusion is robust to forecast-rule choice.

8.11 Multi-Seller Strategy Game↩︎

The two-player game of Table 4 extends cleanly to general \(N\)-seller markets at every composition. The two on-diagonal cells generalize directly from the pure-type results: obob markets via Theorem 2, and inin markets via Theorem 4. The off-diagonal cells—mixed markets with arbitrary informed/oblivious compositions \(|\cI^{ob}| \in \{1, \ldots, N - 1\}\)—are characterized by Theorem 5: under the running-mean forecast, prices converge to \(\mathbf{p}^{NE}\) jointly, informed sellers learn the true demand model, and oblivious sellers continue to pay the persistent exploration tax.

8.11.0.1 Numerical experiments.

An asymmetric multi-seller experiment substantiates this generalization numerically at \(N \in \{3, 5\}\), sweeping the full composition \(|\cI^{ob}| \in \{0, 1, \ldots, N\}\). Table 9 reports the group-mean surplus-capture ratios \(\bar S^{ob}\) (over \(\cI^{ob}\)) and \(\bar S^{in}\) (over \(\cI^{in}\)) for every cell. The two pure-type endpoints generalize the on-diagonal duopoly cells: every all-oblivious cell shows the group-mean oblivious surplus \(\bar S^{ob}\) strongly negative, with \(\bar S^{ob} = -1.38\) at \(N = 3\) and \(\bar S^{ob} = -1.10\) at \(N = 5\), while every all-informed cell shows the group-mean informed surplus \(\bar S^{in}\) tightly bracketing zero (\(-0.010\) at \(N = 3\), \(-0.007\) at \(N = 5\)). The six off-diagonal cells generalize the duopoly obin cell: in every cell with \(|\cI^{ob}| \in \{1, \ldots, N - 1\}\), the gap \(\bar S^{in} - \bar S^{ob}\) is robustly positive (range \([+0.61, +1.25]\)), the oblivious leg \(\bar S^{ob}\) is strictly negative, and the informed leg \(\bar S^{in}\) is operationally at zero.

Table 9: Surplus-capture ratios across the full strategy-game grid in asymmetric multi-seller markets at \(N \in \{3, 5\}\), sweeping the composition \(|\cI^{ob}| \in \{0, 1, \ldots, N\}\) under the running-mean forecast (Theorem 5). \(\nu^2 = 0.10\) oblivious dithering, \(\nu_n^2 = 0.01\,(n+1)^{-0.25}\) informed dithering, horizon \(T = 10^6\), \(S = 80\) seeds per cell, with demand primitives drawn from the heterogeneous distribution introduced in Figure 5. Changing \(\nu_n^2\) to a faster or slower decay pattern does not alter the qualitative ordering of cells. Brackets are cross-seed \(5\%\)\(95\%\) ranges; “gap” is \(\bar S^{in} - \bar S^{ob}\), reported only for mixed cells.
\(N\) \(|\cI^{ob}|/|\cI^{in}|\) \(\bar S^{ob}\) 5%–95% \(\bar S^{in}\) 5%–95% gap
3 0/3 \(-0.010\) \([-0.134, +0.124]\)
3 1/2 \(-1.262\) \([-1.475, -1.077]\) \(-0.010\) \([-0.168, +0.140]\) \(+1.252\)
3 2/1 \(-0.911\) \([-1.035, -0.812]\) \(-0.014\) \([-0.249, +0.245]\) \(+0.897\)
3 3/0 \(-1.377\) \([-1.493, -1.260]\)
5 0/5 \(-0.007\) \([-0.079, +0.066]\)
5 1/4 \(-0.620\) \([-0.724, -0.517]\) \(-0.010\) \([-0.086, +0.069]\) \(+0.609\)
5 2/3 \(-0.721\) \([-0.816, -0.643]\) \(-0.015\) \([-0.115, +0.087]\) \(+0.706\)
5 3/2 \(-1.033\) \([-1.125, -0.945]\) \(-0.016\) \([-0.132, +0.090]\) \(+1.018\)
5 4/1 \(-0.970\) \([-1.047, -0.907]\) \(-0.006\) \([-0.228, +0.157]\) \(+0.964\)
5 5/0 \(-1.098\) \([-1.163, -1.011]\)

8.11.0.2 Multi-seller strategy game.

The qualitative four-cell ordering of Table 4 therefore extends to asymmetric multi-seller markets at every composition \(|\cI^{ob}| \in \{0, \ldots, N\}\) and \(N \in \{3, 5\}\): oblivious sellers strictly below Nash, informed sellers operationally at Nash, with a strict positive gap \(\bar S^{in} > \bar S^{ob}\) in every mixed cell. A unilateral deviation from the all-informed equilibrium to oblivious modeling moves the deviator from the operationally-zero \(\bar S^{in}\) cell to the strictly-negative \(\bar S^{ob}\) cell of Table 9, while all other sellers’ performances are unaffected.

9 Proofs↩︎

9.1 Proof of Theorem 1↩︎

Let \((\mathcal{F}_n)_{n \ge 2}\) be the filtration generated by \(\left\{z_{m+1,i}, \varepsilon_{m,i}: 1 \le m \le n, i \in [N]\right\}\), meaning that \(\mathcal{F}_n\) captures the information right after the \((n+1)\)th prices are set and before the \((n+1)\)th demand is observed. Concretely, this means that \(p_{m,i} \in \mathcal{F}_{m-1}\) and \(\varepsilon_{m,i} \in \mathcal{F}_{m}\) for all \(m \ge 2\). We drop the subscript \(i\) for notational simplicity.

We write \[J_n = \sum_{m=1}^{n} \left(p_{m} - \lbar{p}_{n}\right)^2 = \sum_{m=2}^{n} \left(1 - \frac{1}{m}\right) (p_{m} - \lbar{p}_{m-1})^2.\] Similarly, denote the numerator of \(w_{n}\) by \[C_n = \sum_{m=1}^{n} \left(p_{m} - \lbar{p}_{n}\right) \varepsilon_{m} = \sum_{m=2}^{n} \left(1 - \frac{1}{m}\right) (p_m - \lbar{p}_{m-1}) (\varepsilon_m - \bar{\varepsilon}_{m-1}),\] where \(\bar{\varepsilon}_{m-1} = \frac{1}{m-1} \sum_{k=1}^{m-1} \varepsilon_{k}\). Denote \[a_m = \left(1 - \frac{1}{m}\right) (p_{m} - \lbar{p}_{m-1})\] and note that \(a_m\) is \(\mathcal{F}_{m-1}\)-measurable. Write \[J_n = \sum_{m=2}^{n} \frac{m}{m-1} a_m^2, \quad C_n = D_n - R_n,\] \[D_n = \sum_{m=2}^{n} a_m \varepsilon_{m}, \quad R_n = \sum_{m=2}^{n} a_m \bar{\varepsilon}_{m-1}.\] Since \(\frac{m}{m-1} a_m^2 \in [1,2]\), for all \(n \ge 2\), \[\sum_{m=2}^{n} a_m^2 \le J_n \le 2 \sum_{m=2}^{n} a_m^2.\] This means that \(J_n\) and \(\sum_{m=2}^{n} a_m^2\) are of the same order.

We first analyze \(D_n\). Since prices and demand noises are bounded, \((D_n, \mathcal{F}_n)_{n \ge 2}\) is a square-integrable martingale with uniformly bounded increments. Also, \[V_n \triangleq \sum_{m=2}^{n} \mathbb{E}\left[(D_m - D_{m-1})^2 \middle| \mathcal{F}_{m-1}\right] = \sigma^2 \sum_{m=2}^{n} a_m^2.\] By the strong law of large numbers for martingales (e.g., [55], pp.122-124):

  1. If \(V_n \rightarrow \infty\), then \(D_n / V_n \rightarrow 0\) a.s.

  2. If \(V_n \rightarrow V_\infty\) for some \(V_\infty < \infty\), then \(D_n \rightarrow D_\infty\) a.s. for some finite \(D_\infty\).

Since \(J_n\) is squeezed by \(\sum_{m=2}^{n} a_m^2\) up to constant factors, the same conditions can be expressed in terms of \(J_n\).

Next, we analyze \(R_n\). Again because \(J_n\) is squeezed by \(\sum_{m=2}^{n} a_m^2\), it suffices to analyze \[\frac{\sum_{m=2}^{n} a_m \bar{\varepsilon}_{m-1}}{\sum_{m=2}^{n} a_m^2}.\] Note that \(\bar{\varepsilon}_{m}\) is a normalized sum of bounded martingale differences. By the law of the iterated logarithm for martingales (e.g., [56]), we have \[\bar{\varepsilon}_{m} = O\left(\sqrt{\frac{\log \log m}{m}}\right) \quad \text{a.s.}\] By Cauchy-Schwarz, \[\left|\frac{\sum_{m=2}^{n} a_m \bar{\varepsilon}_{m-1}}{\sum_{m=2}^{n} a_m^2}\right| \le \sqrt{\frac{\sum_{m=2}^{n} \bar{\varepsilon}_{m-1}^2}{\sum_{m=2}^{n} a_m^2}} \preceq \sqrt{\frac{\sum_{m=3}^{n} \frac{\log \log (m-1)}{m-1}}{\sum_{m=2}^{n} a_m^2}} \preceq \sqrt{\frac{\log n \log \log n}{\sum_{m=2}^{n} a_m^2}} \quad \text{a.s.}\] If \(\frac{J_n}{\log n \log \log n} \rightarrow \infty\), then \(\frac{\sum_{m=2}^{n} a_m \bar{\varepsilon}_{m-1}}{\sum_{m=2}^{n} a_m^2} \rightarrow 0\) a.s. Again by Cauchy-Schwarz, \[\begin{align} \left|\frac{\sum_{m=2}^{n} a_m \bar{\varepsilon}_{m-1}}{\sum_{m=2}^{n} a_m^2}\right| &\le \frac{\sqrt{\sum_{m=2}^{n} a_m^2 (\log m)^{1+\delta}} \cdot \sqrt{\sum_{m=2}^{n} \frac{\bar{\varepsilon}_{m-1}^2}{(\log m)^{1+\delta}}}}{\sum_{m=2}^{n} a_m^2} \\ &\preceq \frac{\sqrt{\sum_{m=2}^{n} a_m^2 (\log m)^{1+\delta}} \cdot \sqrt{\sum_{m=3}^{n} \frac{\log \log (m-1)}{(m-1)(\log m)^{1+\delta}}}}{\sum_{m=2}^{n} a_m^2} \quad \text{a.s.} \end{align}\] Since \(\sum_{m=3}^{n} \frac{\log \log (m-1)}{(m-1)(\log m)^{1+\delta}} < \infty\), if \(\sum_{m=2}^{n} a_m^2 (\log m)^{1+\delta} < \infty\), then \(\frac{\sum_{m=2}^{n} a_m \bar{\varepsilon}_{m-1}}{\sum_{m=2}^{n} a_m^2}\) converges a.s. to some finite limit. Putting the analyses of \(D_n\) and \(R_n\) together, we have:

  • If \(\frac{J_n}{\log n \log \log n} \rightarrow \infty\), then \[w_n = \frac{C_n}{J_n} = \frac{D_n}{J_n} - \frac{R_n}{J_n} \rightarrow 0 \quad \text{a.s.}\]

  • If \(\sum_{m=2}^{n} \left(p_{m} - \lbar{p}_{m-1}\right)^2 (\log m)^{1+\delta} < \infty\) for some \(\delta > 0\), then \[w_n = \frac{C_n}{J_n} = \frac{D_n}{J_n} - \frac{R_n}{J_n} \rightarrow w_{\infty} \quad \text{a.s.}\] for some finite \(w_{\infty}\).

The proof is complete.

9.2 Proof of Lemma 1↩︎

By the Cauchy–Schwarz inequality, \[|r_{n,\, j \leftarrow i}| = \frac{\left|\sum_{m=1}^{n} (p_{m,j} - \bar{p}_{n,j})(p_{m,i} - \bar{p}_{n,i})\right|}{J_{n,j}} \le \frac{\sqrt{J_{n,j} J_{n,i}}}{J_{n,j}} = \sqrt{\frac{J_{n,i}}{J_{n,j}}}.\] Since \(\tfrac{J_{n,i}}{J_{n,j}} \to 0\) almost surely, the result follows.

9.3 Proof of Proposition [prop:variancedominancetwosellercase]↩︎

9.3.0.1 Proof of (a).

By Lemma 1, \(\tfrac{J_{n,i}}{J_{n,j}} \to 0\) implies \[r_{n,j \leftarrow i} \to 0 \quad \text{a.s.}\] Hence, seller \(j\)’s misspecified regression becomes asymptotically correct, and their greedy price converges to the best response to seller \(i\)’s limiting price whenever that limit exists. For seller \(i\), write \[r_{n,i \leftarrow j} = \sqrt{\frac{J_{n,j}}{J_{n,i}}} \cos(\theta_n),\] where \(\theta_n\) is the angle between the centered price vectors. Since \(\tfrac{J_{n,j}}{J_{n,i}} \to \infty\), the behavior of \(r_{n,i \leftarrow j}\) depends on the cosine term. If \(r_{n,i \leftarrow j} \to 0\) due to \(\cos(\theta_n) \to 0\) fast, then the effective misspecification vanishes and both greedy prices converge to the NE. Otherwise, \(|r_{n,i \leftarrow j}| \to \infty\) along a subsequence, which forces the projected coefficients \((\hat{a}_{n,i}, \hat{b}_{n,i})\) to hit the boundary of \(\Theta_i^{ob}\), and the greedy price converges to one of the boundary values \(\ell_{-,i} = \frac{\underline{a}_i}{-2 \bar{b}_i}\) or \(\ell_{+,i} = \frac{\bar{a}_i}{-2 \underline{b}_i}\) along that subsequence. In all cases, seller \(j\) best responds to seller \(i\)’s limiting price along the same subsequence.

9.3.0.2 Proof of (b).

Specialize to the symmetric duopoly and write \(b\triangleq|\beta|>0\) and \(\phi(q) \triangleq (\alpha + \gamma q)/(2b)\) for the common best response. For a candidate dominated price \(q\) and the corresponding best response \(p_j^* = \phi(q)\), the first-order condition \(\alpha + 2\beta p_j^* + \gamma q = 0\) gives \(\alpha + \beta p_j^* + \gamma q = -\beta p_j^* = b\,p_j^*\), so \[R_j(p_j^*,q) \;=\; p_j^*\bigl(\alpha + \beta p_j^* + \gamma q\bigr) \;=\; b\,(p_j^*)^2.\] The asymptotic revenue gap as a function of \(q\) is then \[\begin{align} \Delta(q) &\;\triangleq\; R_j(p_j^*,q) - R_i(q,p_j^*) \;=\; b\,(p_j^*)^2 - q\bigl(\alpha + \beta q + \gamma p_j^*\bigr) \\ &= b\,(p_j^*)^2 + b\,q^2 - \alpha q - \gamma q\, p_j^*. \end{align}\] Substituting \(p_j^* = (\alpha+\gamma q)/(2b)\) and simplifying, \[\label{eq:delta-q-formula} \Delta(q) \;=\; \frac{\bigl[(2b-\gamma)q - \alpha\bigr]\,\bigl[(2b+\gamma)q - \alpha\bigr]}{4b}.\tag{21}\] The two roots are \(q = \alpha/(2b-\gamma) = p^{NE}\) and \(q = q_- = \alpha/(2b+\gamma) < p^{NE}\). With the leading coefficient \((2b-\gamma)(2b+\gamma)/(4b)>0\), the sign of \(\Delta\) is positive for \(q < q_-\) and for \(q > p^{NE}\), negative for \(q \in (q_-, p^{NE})\), and zero at the two roots. Under assumption 7 , \(\ell_{-,i} < q_-\) and \(\ell_{+,i} > p^{NE}\) so \(\Delta(\ell_{\pm,i}) > 0\), while \(\Delta(p^{NE}) = 0\) identically. Hence at every lock candidate, \[\label{eq:pointwise-delta} \Delta(q)\;\ge\;0,\quad q\in\{\ell_{-,i},\ell_{+,i},p^{NE}\}, \qquad \Delta(q) > 0\;\text{iff } q\ne p^{NE}.\tag{22}\]

Write \(p_{n,k} = \tilde{p}_{n,k} + z_{n,k}\) where \(z_{n,k}\) is the (mean-zero, \(\mathcal{F}_{n-1}\)-conditional-variance-\(\nu_{n,k}^2\)) exploration shock and \(z_{n,i}\perp z_{n,j}\). Expanding \(R_{n,k} = p_{n,k}(\alpha + \beta p_{n,k} + \gamma p_{n,-k})\) around \(\tilde{p}_n\) and taking \(\mathbb{E}[\cdot\mid \mathcal{F}_{n-1}]\), \[\label{eq:Rnk-decomp} \mathbb{E}[R_{n,k}\mid\mathcal{F}_{n-1}] \;=\; R_k(\tilde{p}_n) + \beta\,\nu_{n,k}^2.\tag{23}\] The martingale residual \(R_{n,k} - \mathbb{E}[R_{n,k}\mid\mathcal{F}_{n-1}]\) has uniformly bounded conditional second moments (since prices and \(z_{n,k}\) are bounded), so its Cesàro average vanishes a.s.by the strong law for martingale differences. Hence, recalling \(\bar\nu_{T,k}^2 = J_{T,k}/T\), \[\bar R_{T,k}\;=\;\frac{1}{T}\sum_{m\le T} R_k(\tilde{p}_m) \;+\; \beta\,\bar\nu_{T,k}^2 \;+\; o(1) \quad\text{a.s.}\] Under the hypothesis \(J_{T,k} = o(T)\), \(\bar\nu_{T,k}^2\to 0\), so the exploration-tax term vanishes: \[\label{eq:cesaro-no-tax} \bar R_{T,k} \;-\; \frac{1}{T}\sum_{m\le T} R_k(\tilde{p}_m) \;\xrightarrow[T\to\infty]{a.s.}\; 0, \qquad k\in\{i,j\}.\tag{24}\] By Part (a), \(\tilde{p}_{n,j} - \phi(\tilde{p}_{n,i})\to 0\) a.s., so \(R_j(\tilde{p}_n) - R_i(\tilde{p}_n) \to \Delta(\tilde{p}_{n,i})\) a.s.along any subsequential lock of \(\tilde{p}_{n,i}\). Continuity of \(\Delta\) together with 22 gives that, for every \(\varepsilon>0\), \(\Delta(\tilde{p}_{n,i}) \ge -\varepsilon\) eventually a.s. Cesàro-averaging and combining with 24 , \[\liminf_{T\to\infty}\bigl(\bar R_{T,j} - \bar R_{T,i}\bigr) \;\ge\; 0 \quad\text{a.s.},\] and dividing by \(\Pi^{C} - \Pi^{NE}>0\) yields \(S_j \ge S_i\) a.s. On any sample path with \(\tilde{p}_{n,i}\to \tilde{p}_{\infty,i} \in\{\ell_{-,i},\ell_{+,i}\}\), continuity gives \(\Delta(\tilde{p}_{n,i})\to \Delta(\tilde{p}_{\infty,i}) > 0\), hence by 24 , \(\bar R_{T,j} - \bar R_{T,i}\to \Delta(\tilde{p}_{\infty,i}) > 0\) a.s., and \(S_j - S_i = \Delta(\tilde{p}_{\infty,i})/(\Pi^{C} - \Pi^{NE}) > 0\). On any sample path with \(\tilde{p}_{n,i}\to p^{NE}\), \(R_k(\tilde{p}_n)\to\Pi^{NE}\) for both \(k\), so \(\bar R_{T,k}\to\Pi^{NE}\) and \(S_j = S_i = 0\) a.s.\(\square\)

9.4 Proof of Theorem 2↩︎

For notational simplicity, we denote \(x_{n,i}^{ob}\) by \(x_{n,i}\), \(\tilde{\theta}_{n,i}^{ob}\) by \(\tilde{\theta}_{n,i}\), \(\theta_{n,i}^{ob}\) by \(\hat{\theta}_{n,i}\), and \(\theta_i^{*, \,ob}\) by \(\theta_i^*\). Recall that each oblivious seller \(i\) models the demand function as \[d_{n,i}=a_i + b_i p_{n,i} + \varepsilon_{n,i} = x_{n,i}^\top \theta_i^* + \eta_{n,i},\] where \[\eta_{n,i} = \varepsilon_{n,i} + \sum_{j\ne i}\gamma_{ij}\left(p_{n,j}-p_j^{NE}\right).\] For each seller \(i\), define the adaptive covariance matrix \[S_{n,i}=\sum_{m=1}^n x_{m,i}x_{m,i}^\top, \quad M_{n,i}=S_{n,i}/n,\] and note that \(S_{n,i}\) is invertible for all \(n \ge 2\) since the first two prices are not the same. Then, the unprojected least squares estimator can be written as \[\tilde{\theta}_{n,i} = S_{n,i}^{-1} \sum_{m=1}^n x_{m,i} d_{m,i} = \theta_i^* + S_{n,i}^{-1} \sum_{m=1}^n x_{m,i} \eta_{m,i}.\] Define the estimation error \[\tilde{e}_{n,i} = \tilde{\theta}_{n,i} - \theta_i^*, \quad e_{n,i} = \hat{\theta}_{n,i} - \theta_i^*.\] Then, we have \[\tilde{\theta}_{n+1, i} - \tilde{\theta}_{n,i} = S_{n+1,i}^{-1} x_{n+1,i} (d_{n+1,i} - x_{n+1,i}^\top \tilde{\theta}_{n,i})\] and \[\begin{align} \tilde{e}_{n+1,i} - \tilde{e}_{n,i} &= S_{n+1,i}^{-1} x_{n+1,i} (x_{n+1,i}^\top \theta_i^* + \eta_{n+1,i} - x_{n+1,i}^\top \tilde{\theta}_{n,i}) \\ &= - S_{n+1,i}^{-1} x_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i} + S_{n+1,i}^{-1} x_{n+1,i} \eta_{n+1,i}. \end{align}\] Define Lyapunov functions \[V_{n,i} = \tilde{e}_{n,i}^\top S_{n,i} \tilde{e}_{n,i}, \quad W_{n,i} = \tilde{e}_{n,i}^\top M_{n,i} \tilde{e}_{n,i} = \frac{1}{n} V_{n,i},\] and \[V_n = \sum_{i=1}^{N} V_{n,i}, \quad W_n = \sum_{i=1}^{N} W_{n,i} = \frac{1}{n} V_n.\] Then, \[\begin{align} V_{n+1, i} &= (\tilde{e}_{n,i} - S_{n+1,i}^{-1} x_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i} + S_{n+1,i}^{-1} x_{n+1,i} \eta_{n+1,i})^\top S_{n+1,i} \\ &\quad \quad (\tilde{e}_{n,i} - S_{n+1,i}^{-1} x_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i} + S_{n+1,i}^{-1} x_{n+1,i} \eta_{n+1,i}) \\ &= V_{n,i} - (x_{n+1, i}^\top \tilde{e}_{n,i})^2 (1 - x_{n+1, i}^\top S_{n+1,i}^{-1} x_{n+1, i}) \\ &\quad \quad + 2 \eta_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i} (1 - x_{n+1, i}^\top S_{n+1,i}^{-1} x_{n+1, i}) \\ &\quad \quad + x_{n+1,i}^\top S_{n+1,i}^{-1} x_{n+1,i} \eta_{n+1,i}^2. \end{align}\] The Sherman-Morrison formula gives \[S_{n+1,i}^{-1} = S_{n,i}^{-1} - \frac{S_{n,i}^{-1} x_{n+1,i} x_{n+1,i}^\top S_{n,i}^{-1}}{1 + x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i}}.\] Then, \[S_{n+1,i}^{-1} x_{n+1,i} = S_{n,i}^{-1} x_{n+1,i} - \frac{S_{n,i}^{-1} x_{n+1,i} (x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i})}{1 + x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i}} = \frac{S_{n,i}^{-1} x_{n+1,i}}{1 + x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i}},\] and \[1 - x_{n+1, i}^\top S_{n+1,i}^{-1} x_{n+1, i} = \frac{1}{1 + x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i}}.\] Denoting \(h_{n+1,i} = x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i}\), we have \[V_{n+1, i} = V_{n,i} - \frac{(x_{n+1, i}^\top \tilde{e}_{n,i})^2}{1 + h_{n+1,i}} + \frac{2 \eta_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i}}{1 + h_{n+1,i}} + \frac{h_{n+1,i}}{1 + h_{n+1,i}} \eta_{n+1,i}^2,\] \[W_{n+1, i} = \frac{n}{n+1} W_{n,i} - \frac{(x_{n+1, i}^\top \tilde{e}_{n,i})^2}{(n+1)(1 + h_{n+1,i})} + \frac{2 \eta_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i}}{(n+1)(1 + h_{n+1,i})} + \frac{h_{n+1,i}}{(n+1)(1 + h_{n+1,i})} \eta_{n+1,i}^2.\] Since \[p_{n+1,j}=\phi^{ob}(\hat{\theta}_{n,j})+z_{n+1,j}=\phi^{ob}(\theta_j^*+e_{n,j})+z_{n+1,j}\] and \(\mathbb{E}[z_{n+1,j}| \mathcal{F}_n]=0\), we have \[\left|\mathbb{E}[p_{n+1,j}-p_j^{NE}\mid\mathcal{F}_n]\right| = \left|\phi^{ob}(\theta_j^*+e_{n,j})-\phi^{ob}(\theta_j^*)\right| \le L_\phi^{ob} \norm{e_{n,j}}_2 \le L_\phi^{ob} \norm{\tilde{e}_{n,j}}_2,\] where the last inequality holds because projection onto a convex set is non-expansive. Given \(\mathcal{F}_n\), \(x_{n+1, i}\) depends only on \(z_{n+1,i}\), and \(p_{n+1,j}-p_j^{NE}\) depends only on \(z_{n+1,j}\). Since \(z_{n+1,i}\) are independent across \(i\), we have, for \(j \neq i\), \[(p_{n+1,j}-p_j^{NE})\, \perp\, x_{n+1,i}\quad\text{conditional on }\mathcal{F}_n.\] Recall that \[\eta_{n+1,i} =\varepsilon_{n+1,i}+\sum_{j\ne i}\gamma_{ij}(p_{n+1,j}-p_j^{NE}).\] Since \(\varepsilon_{n+1,i}\) is independent of everything else with zero mean, we have \[\mathbb{E}[\eta_{n+1,i}x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n] =\sum_{j\ne i}\gamma_{ij}\mathbb{E}[(p_{n+1,j}-p_j^{NE})x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n].\] By the conditional independence noted above, \[\mathbb{E}[(p_{n+1,j}-p_j^{NE})x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n] =\mathbb{E}[p_{n+1,j}-p_j^{NE}\mid\mathcal{F}_n]\cdot \mathbb{E}[x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n].\] Thus, \[\begin{align} \left|\mathbb{E}[\eta_{n+1,i}x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n]\right| &\le \sum_{j\ne i}\gamma_{ij} \left|\mathbb{E}[p_{n+1,j}-p_j^{NE}\mid\mathcal{F}_n]\right| \cdot \left|\mathbb{E}[x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n]\right| \\ &\le \sum_{j\ne i}\gamma_{ij} L_\phi^{ob} \norm{\tilde{e}_{n,j}}_2 \cdot C_x \norm{\tilde{e}_{n,i}}_2, \end{align}\] where the first inequality is due to Cauchy-Schwarz. Using Young’s inequality, \[\left|\mathbb{E}[\eta_{n+1,i}x_{n+1,i}^\top \tilde{e}_{n,i}\mid\mathcal{F}_n]\right| \le \frac{1}{2} L_\phi^{ob} C_x \sum_{j\ne i}\gamma_{ij} \left(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}}_2^2\right).\] Since \[\mathbb{E}\left[ x_{n+1, i} x_{n+1, i}^\top \middle| \mathcal{F}_n\right] \succeq C_M I_{2} \quad \text{for all } n \text{ and } i\] for some constant \(C_M > 0\), we have \[\mathbb{E}\left[ (x_{n+1, i}^\top \tilde{e}_{n,i})^2 \middle| \mathcal{F}_n\right] = \tilde{e}_{n,i}^\top \mathbb{E}\left[ x_{n+1, i} x_{n+1, i}^\top \middle| \mathcal{F}_n\right] \tilde{e}_{n,i} \ge C_{M} \norm{\tilde{e}_{n,i}}_2^2.\] By Lemma 10, we have that, for any \(0 < \delta < C_M\), eventually a.s., \[h_{n+1,i} \le C_x^2 \lambda_{\max}(S_{n,i}^{-1})= \frac{C_x^2}{C_{M} - \delta} \frac{1}{n} \le \delta_h\] for arbitrarily small \(\delta_h > 0\). Since \(\frac{1}{1+h_{n+1,i}}\ge \frac{1}{1+\delta_h}\) eventually, \(\frac{h_{n+1,i}}{1+h_{n+1,i}}\le h_{n+1,i}\), and \(C_\eta^2 \triangleq \sup_{n,i}\mathbb{E}[\eta_{n+1,i}^2\mid\mathcal{F}_n]<\infty\) (finite because demand noises and prices are bounded), taking conditional expectation of the \(W_{n+1,i}\) recursion gives that, for all large \(n\) almost surely, \[\begin{align} \mathbb{E}[W_{n+1,i}\mid\mathcal{F}_n] &\le \frac{n}{n+1}W_{n,i} -\frac{C_M}{(n+1)(1+\delta_h)}\norm{\tilde{e}_{n,i}}_2^2 \\ &\quad \quad + \frac{L_\phi^{ob} C_x }{n+1} \sum_{j\ne i}\gamma_{ij} \left(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}}_2^2\right) +\frac{C_x^2 C_\eta^2}{(C_M-\delta)}\frac{1}{n(n+1)}. \end{align}\] Summing over \(i\) from \(1\) to \(N\), we have \[\begin{align} \mathbb{E}[W_{n+1}\mid\mathcal{F}_n] &\le \frac{n}{n+1} W_n - \frac{C_M }{(n+1)(1+\delta_h)} \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 \\ &\quad \quad + \frac{L_\phi^{ob} C_x }{n+1} \sum_{i=1}^N \sum_{j\ne i}\gamma_{ij} \left(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}}_2^2\right) + \frac{N C_x^2 C_\eta^2}{(C_M-\delta)}\frac{1}{n(n+1)}. \end{align}\] Note that each \(\norm{\tilde{e}_{n,i}}_2^2\) appears with weight \(\gamma_i\) from the first term and weight \(\gamma_i^{\mathrm{col}}\) from the second (by relabeling summation indices). Hence, \[\sum_{i=1}^N \sum_{j\ne i}\gamma_{ij} \left(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}}_2^2\right) = \sum_{i=1}^N (\gamma_i + \gamma_i^{\mathrm{col}}) \norm{\tilde{e}_{n,i}}_2^2 \le 2\bar\gamma \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2.\] So, \[\mathbb{E}[W_{n+1}\mid\mathcal{F}_n] \le \frac{n}{n+1} W_n + \frac{\kappa}{n+1} \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 + \frac{C_\delta}{n(n+1)},\] where \[\kappa = 2 L_\phi^{ob} C_x \bar\gamma - \frac{C_M}{1+\delta_h}, \quad \text{and} \quad C_\delta = \frac{N C_x^2 C_\eta^2}{C_M-\delta}.\] By Lemma 10 again, for all \(0 < \delta < C_M\), eventually a.s., \[W_n \ge (C_{M} - \delta) \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 \; \Longrightarrow \; \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 \le \frac{1}{C_{M} - \delta} W_n.\] Similarly, a trivial spectral upper bound gives \[W_n \le C_x^2 \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 \; \Longrightarrow \; \sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 \ge \frac{1}{C_x^2} W_n.\]

From the condition of the theorem, we have \[\bar\gamma L_\phi^{ob} C_x < C_{M}.\] We consider three cases. Suppose \[2 \bar\gamma L_\phi^{ob} C_x < C_{M}.\] Then, since \(\delta\) and \(\delta_h\) can be arbitrarily small, \(\kappa < 0\) eventually, and eventually a.s., \[\begin{align} \mathbb{E}\left[W_{n+1} \middle| \mathcal{F}_n\right] &\le \frac{n}{n+1} W_{n} + \frac{\kappa}{C_x^2(n+1)} W_n + \frac{C_\delta}{n(n+1)} \\ &= W_n - \left(1 - \frac{\kappa}{C_x^2}\right)\frac{1}{n+1} W_n + \frac{C_\delta}{n(n+1)}. \end{align}\] Because \(\kappa < 0\), the Robbins–Siegmund conditions [57] are satisfied, and we have \[W_n \text{ converges a.s.} \quad \text{and} \quad \sum_{n=1}^{\infty} \frac{1 - \kappa / C_x^2}{n+1} W_n < \infty \quad \text{a.s.}\] This implies that \(W_n \rightarrow 0\) a.s. because \(\sum_{n=1}^{\infty} \frac{1}{n+1} = \infty\). If \[2 \bar\gamma L_\phi^{ob} C_x = C_{M},\] then \(\kappa\) goes to 0 as \(\delta_h \rightarrow 0\), which is at the rate or \(O(1/n)\). Then, the second term in the \(W_{n+1}\) recursion can be absorbed into the last term, and we still have \(W_n \rightarrow 0\) a.s. If \[\bar\gamma L_\phi^{ob} C_x < C_{M} < 2 \bar\gamma L_\phi^{ob} C_x,\] then \(\kappa > 0\) infinitely often, and for all \(n\) large enough, \[\Delta_\delta \triangleq 1 - \frac{\kappa}{C_{M} - \delta} \in (0,1).\] We have \[\begin{align} \mathbb{E}\left[W_{n+1} \middle| \mathcal{F}_n\right] &\le \frac{n}{n+1} W_{n} + \frac{\kappa}{(n+1)(C_{M} - \delta)} W_n + \frac{C_\delta}{n(n+1)} \\ &= \left(1 - \frac{1}{n+1}\left(1 - \frac{\kappa}{C_{M} - \delta}\right)\right) W_n + \frac{C_\delta}{n(n+1)} \\ &= W_n - \frac{\Delta_\delta}{n+1} W_n + \frac{C_\delta}{n(n+1)}. \end{align}\] Then, the Robbins–Siegmund conditions are again satisfied and we still have \(W_n \rightarrow 0\) a.s. By the spectral lower bound \(\sum_{i=1}^N \norm{\tilde{e}_{n,i}}_2^2 \le \frac{1}{C_{M} - \delta} W_n\), we have that \(\tilde{\theta}_{n,i} \rightarrow \theta_i^*\) a.s. for all \(i\). Since the projections do not take effect eventually, we also have \(\hat{\theta}_{n,i} \rightarrow \theta_i^*\) a.s. for all \(i\). By the continuity of \(\phi^{ob}\), we have \(\tilde{p}_{n,i} \rightarrow p_i^{NE}\) a.s. for all \(i\). This completes the almost sure convergence part of the proof.

To obtain the rate of convergence, take the unconditional expectation on both sides of the \(W_{n+1}\) recursion. We have \[\begin{cases} \mathbb{E}\left[W_{n+1}\right] \le \mathbb{E}\left[W_{n}\right] - \left(1 - \frac{\kappa}{C_x^2}\right)\frac{1}{n+1} \mathbb{E}\left[W_{n}\right] + \frac{C_\delta}{n(n+1)}, & \kappa < 0 \text{ eventually}, \\ \mathbb{E}\left[W_{n+1}\right] \le \mathbb{E}\left[W_{n}\right] -\frac{1}{n+1} \mathbb{E}\left[W_{n}\right] + \frac{C_\delta}{n(n+1)}, & \kappa \rightarrow 0, \\ \mathbb{E}\left[W_{n+1}\right] \le \mathbb{E}\left[W_{n}\right] - \frac{\Delta_\delta}{n+1} \mathbb{E}\left[W_{n}\right] + \frac{C_\delta}{n(n+1)}, & \kappa > 0 \text{ infinitely often}. \end{cases}\] In the first case, applying Lemma 11 with \(a = 1 - \frac{\kappa}{C_x^2} > 1\) and \(b = C_\delta\), we have \[\mathbb{E}\left[W_{n}\right] = O\left(\frac{1}{n}\right).\] In the second case, applying Lemma 11 with \(a = 1\) and \(b = C_\delta\), we have \[\mathbb{E}\left[W_{n}\right] = O\left(\frac{\log n}{n}\right).\] In the third case, applying Lemma 11 with \(a = \Delta_\delta < 1\) and \(b = C_\delta\), since \(\delta\) and \(\delta_h\) can be arbitrarily small, we have \[\mathbb{E}\left[W_{n}\right] = O\left(n^{- 2\left(1 - \rho\right)}\right),\] where \[\rho \triangleq \frac{\bar\gamma L_\phi^{ob} C_x}{C_{M}} \in (0,1).\] So, \[\mathbb{E}\left[W_{n}\right] = \begin{cases} O\left(\frac{1}{n}\right), & \text{if } 2 \bar\gamma L_\phi^{ob} C_x < C_{M}, \\ O\left(\frac{\log n}{n}\right), & \text{if } 2 \bar\gamma L_\phi^{ob} C_x = C_{M}, \\ O\left(n^{- 2\left(1 - \rho\right)}\right), & \text{if } \bar\gamma L_\phi^{ob} C_x < C_{M} < 2 \bar\gamma L_\phi^{ob} C_x. \end{cases}\] Again, by the spectral lower bound \(\sum_{i=1}^{N} \norm{\tilde{e}_{n,i}}_2^2 \le \frac{1}{C_{M} - \delta} W_n\), we can similarly write the mean squared error bounds as \[\sum_{i=1}^{N} \mathbb{E}\|\tilde{\theta}_{n,i} - \theta_i^*\|_2^2 = \begin{cases} O\left(\frac{1}{n}\right), & \text{if } 2 \bar\gamma L_\phi^{ob} C_x < C_{M}, \\ O\left(\frac{\log n}{n}\right), & \text{if } 2 \bar\gamma L_\phi^{ob} C_x = C_{M}, \\ O\left(n^{- 2\left(1 - \rho\right)}\right), & \text{if } \bar\gamma L_\phi^{ob} C_x < C_{M} < 2 \bar\gamma L_\phi^{ob} C_x, \end{cases}\] where \(\rho\in(1/2, 1)\). The projected estimators have the same bounds because, as the unprojected estimators converge to \(\theta_i^*\) which is in the interior of \(\Theta_i^{ob}\), eventually the projections do not take effect. Since \(\phi^{ob}\) is uniformly Lipschitz continuous on the compact set \(\bigcup_i\Theta_i^{ob}\), we have the same bound for the greedy prices: \[\mathbb{E}\|\tilde{\mathbf{p}}_{n} - \mathbf{p}^{NE}\|_2^2 = \begin{cases} O\left(\frac{1}{n}\right), & \text{if } 2 \bar\gamma L_\phi^{ob} C_x < C_{M}, \\ O\left(\frac{\log n}{n}\right), & \text{if } 2 \bar\gamma L_\phi^{ob} C_x = C_{M}, \\ O\left(n^{- 2\left(1 - \rho\right)}\right), & \text{if } \bar\gamma L_\phi^{ob} C_x < C_{M} < 2 \bar\gamma L_\phi^{ob} C_x, \end{cases}\] where \(\rho\in(1/2, 1)\).

9.5 Proof of Theorem 3↩︎

Let \((m^*,Q^*)\) be any equilibrium of 8 . From \(\dot{m}=0\) we have \[m_i^* = p_i^g(m^*,Q^*),\qquad i\in[N].\] From \(\dot{Q}=0\) we obtain, for \(i\neq j\), \[Q_{ij}^* = p_i^g(m^*,Q^*)p_j^g(m^*,Q^*) = m_i^* m_j^*,\] and for \(i=j\), \[Q_{ii}^* = (p_i^g(m^*,Q^*))^2+\nu^2 = (m_i^*)^2+\nu^2.\] Therefore, \(S_{ij}(m^*,Q^*)=Q_{ij}^*-m_i^* m_j^*=0\) for all \(i\neq j\), and \(V_i(m^*,Q^*)=Q_{ii}^*-(m_i^*)^2=\nu^2\) for all \(i\). Plugging \(S_{ij}=0\) into 13 yields \[b_i(m^*,Q^*)=\beta_i, \qquad a_i(m^*,Q^*)=\alpha_i+\sum_{j\neq i}\gamma_{i,j}m_j^*.\] Using 14 and \(m_i^*=p_i^g(m^*,Q^*)\) then gives, for each \(i\), \[m_i^* = -\frac{\alpha_i+\sum_{j\neq i}\gamma_{i,j}m_j^*}{2\beta_i}.\] Equivalently, \[2\beta_i\,m_i^*+\sum_{j\neq i}\gamma_{i,j}m_j^*=-\alpha_i,\qquad i\in[N].\] This coincides with the first-order condition for the full-information Nash equilibrium. By the same development as in Section 3.3, \(m^*=\mathbf{p}^{NE}\) is the unique solution. The result for \(Q^*\) then follows.

For local asymptotic stability, we linearize 8 around \((m^*,Q^*)\). Define \[y_i(m,Q)\triangleq p_i^g(m,Q)-m_i,\qquad i\in[N],\] and the centered second-moment deviations \[v_i\triangleq (Q_{ii}-m_i^2)-\nu^2,\qquad i\in[N], \qquad s_{ij}\triangleq Q_{ij}-m_im_j,\qquad i\neq j.\] In a neighborhood of \((m^*,Q^*)\) this defines a smooth change of coordinates from \((m,Q)\) to \((m,v,s)\). By 8 , we have \[\label{eq:ode-vs-exact-rep} \dot{v}_i = -v_i + y_i(m,Q)^2,\qquad \dot{s}_{ij} = -s_{ij} + y_i(m,Q)\,y_j(m,Q)\quad (i\neq j),\tag{25}\] together with \(\dot{m}_i=y_i(m,Q)\).

At equilibrium, \(y(m^*,Q^*)=0\), \(v^*= 0\), and \(s^*= 0\). Let \(J\) denote the Jacobian of the vector field of 8 expressed in \((m,v,s)\) coordinates, evaluated at \((m^*,0,0)\). Since the maps \(y\mapsto y_i^2\) and \((y_i,y_j)\mapsto y_i y_j\) have derivative zero at \(y=0\), linearizing 25 yields \[\delta\dot{v} = -\,\delta v,\qquad \delta\dot{s} = -\,\delta s.\] Therefore, \(J\) has the block upper-triangular form \[J= \begin{pmatrix} J_{mm} & *\\ 0 & -I \end{pmatrix},\] where the \(-I\) block corresponds to \((v,s)\) and contributes eigenvalues \(-1\) with multiplicity \(N+\frac{N(N-1)}{2}\). Hence, the spectrum of \(J\) is the union of \(\mathrm{spec}(J_{mm})\) and \(\{-1\}\), and it remains to show that \(J_{mm}\) is Hurwitz.

The mean block is the Jacobian of \(m\mapsto \dot{m}\) at equilibrium with \((v,s)\) held fixed. Since \(s^*\equiv 0\), the greedy map, locally continuously differentiable, simplifies to the linear best response \[p_i^g(m,Q^*)= -\frac{\alpha_i+\sum_{j\neq i}\gamma_{i,j}m_j}{2\beta_i},\] so \[\frac{\partial p_i^g}{\partial m_j}(m^*,Q^*)= -\frac{\gamma_{i,j}}{2\beta_i}\;(i\neq j), \qquad \frac{\partial p_i^g}{\partial m_i}(m^*,Q^*)=0.\] Since \(\dot{m}_i=p_i^g-m_i\), it follows that \[(J_{mm})_{ii}=-1,\qquad (J_{mm})_{ij}=-\frac{\gamma_{i,j}}{2\beta_i}\;\;(i\neq j).\] Define \(A\) by \(A_{ii}=0\) and \(A_{ij}=\frac{\gamma_{i,j}}{-2\beta_i}\) for \(i\neq j\). We therefore have \[J_{mm} = A - I.\] Because \(\beta_i<0\) for all \(i\), \(A\) is entrywise non-negative and its row sums satisfy \[\sum_{j}A_{ij}=\sum_{j\neq i}\frac{\gamma_{i,j}}{-2\beta_i}=\frac{\gamma_i}{-2\beta_i}<\frac{1}{2},\] using \(-\beta_i>\gamma_i\). Hence, \(\|A\|_\infty<1/2\) and every eigenvalue \(\mu\) of \(A\) satisfies \(|\Re(\mu)|\le \|A\|_\infty<1/2\). Thus, every eigenvalue \(\lambda\) of \(J_{mm}\) therefore satisfies \[\Re(\lambda)=\Re(\mu-1)\le |\Re(\mu)|-1 < -\frac{1}{2}.\] Therefore, \(J_{mm}\) is Hurwitz, and consequently the full Jacobian \(J\) is Hurwitz. Standard linearization theory implies local asymptotic stability of \((m^*,Q^*)\) for 8 .

9.6 Proof of Proposition [prop:reduced-excursion]↩︎

We first prove two auxiliary lemmas. Define \(V(t) = V(\mathbf{y}_t)\), \(C(t)=C(\mathbf{y}_t)\), and \(r(t)=r(\mathbf{y}_t)\). Since \(\dot{y}_1(t)=B(y_1(t),r(t))-y_1(t)\), we have \((B(y_1(t),r(t))-y_1(t))^2=\dot{y}_1(t)^2\). A direct calculation from 16 yields the following lemma.

Lemma 2 (Variance–covariance identities). Along any solution of 16 , \[\label{eq:reduced-V-C} \dot{V}(t)= \dot{y}_1(t)^2+\nu^2 - V(t), \qquad \dot{C}(t)= \dot{y}_1(t)^2 - C(t).\qquad{(6)}\] Consequently, whenever \(V(t)>0\), the ratio \(r(t)=C(t)/V(t)\) is differentiable and satisfies \[\label{eq:reduced-r} \dot{r}(t) = \frac{\dot{y}_1(t)^2+\nu^2}{V(t)} \left(q(t)-r(t)\right), \qquad q(t)\triangleq \frac{\dot{y}_1(t)^2}{\dot{y}_1(t)^2+\nu^2}\in[0,1).\qquad{(7)}\]

Proof of Lemma 2. Differentiate \(V(t)=y_2(t)-y_1(t)^2\) and use 16 : \[\dot{V}=\dot{y}_2-2y_1\dot{y}_1 =\left(B^2+\nu^2-y_2\right)-2y_1(B-y_1) =(B-y_1)^2+\nu^2-(y_2-y_1^2),\] which gives the identity for \(V(t)\). The derivation for \(C(t)\) is identical using \(C(t)=y_3(t)-y_1(t)^2\). For ?? , apply the quotient rule to \(r=C/V\) and substitute ?? . ◻

Lemma 3 (Monotonicity of the greedy map). Assume \(\alpha+(\beta+ \gamma) u>0\). Then for all \((y_1,r)\in[l,u]\times[0,1)\), \[\frac{\partial}{\partial r}B(y_1,r)>0.\] In particular, for \(r\in[0,1)\), \[\label{eq:reduced-envelope} B(y_1,0)\le B(y_1,r)<B(y_1,1)=p^C.\qquad{(8)}\]

Proof. A direct differentiation of 17 gives \[\frac{\partial}{\partial r}B(y_1,r) = \frac{\gamma\left(\alpha+(\beta+\gamma)y_1\right)}{(-2\beta-2\gamma r)^2}.\] The denominator is positive. Since \(y_1\in[l,u]\), we have \(\alpha+(\beta+\gamma)y_1\ge \alpha+(\beta+ \gamma)u>0\), hence \(\partial_r B>0\). The envelope ?? follows by monotonicity and the identity \(B(y_1,1)=\alpha/(-2\beta-2\gamma)=p^C\). ◻

We now prove the claims in Proposition [prop:reduced-excursion].

(i) This is immediate from Lemma 2: solving ?? gives \[V(t)=e^{-t}V(0)+\int_0^t e^{-(t-s)}\left(\dot{y}_1(s)^2+\nu^2\right)\,ds \ge e^{-t}V(0)+\nu^2(1-e^{-t})>0.\]

(ii) Define \(W(t)\triangleq V(t)-C(t)=y_2(t)-y_3(t)\). From 16 , \[\dot{W}(t)=\dot{y}_2(t)-\dot{y}_3(t) =\left(B^2+\nu^2-y_2\right)-\left(B^2-y_3\right) =\nu^2-W(t),\] so \(W(t)=\nu^2+(W(0)-\nu^2)e^{-t}\). Hence, there exists \(T_r<\infty\) such that \(W(t)>0\) for all \(t\ge T_r\). Since \(r(t)=C(t)/V(t)\) and \(V(t)>0\) by (i), it follows that \(r(t)<1\) for all \(t\ge T_r\). By Lemma 3, \(\partial_r B>0\) on \([l,u]\times[0,1)\), and \(B(y_1,1)=p^C\). Under the standing assumption that \(l < p^{NE}<p^C<u\) and since \(r(t)<1\) for all \(t\ge T_r\) by (ii), we have \(B(y_1(t),r(t))<p^C\) and therefore \[\dot{y}_1(t)=B(y_1(t),r(t))-y_1(t)< p^C-y_1(t), \qquad t\ge T_r.\] This implies \(y_1(t)\) falls below \(p^C\) in finite time and cannot cross above thereafter, yielding a finite \(T_C\) such that \(y_1(t)<p^C\) for all \(t\ge T_C\). Defining \(T_{r,C} = \max\{T_r,T_C\}\) gives the claim.

(iii) From Lemma 2, \(C(t)\) satisfies \(\dot{C}(t)=\dot{y}_1(t)^2-C(t)\), hence \(C(t)\) is pulled toward the nonnegative input \(\dot{y}_1(t)^2\), and if \(C(t)\) is positive at some time, it remains so thereafter. Suppose \(C(t)> 0\) for all \(t \ge t_0\) for some \(t_0\). Then, \(r(t)=C(t)/V(t)> 0\) for all \(t \ge 0\), and Lemma 3 implies \[B(y_1(t),r(t))> B(y_1(t),0).\] Since \(B(y,0)-y\) is strictly positive for \(y<p^{NE}\) and vanishes at \(y=p^{NE}\), a standard one-dimensional comparison argument implies that \(y_1(t)\) crosses above \(p^{NE}\) in finite time (with \(\dot{y}_1(T_{NE}) > 0\)) and cannot cross below thereafter, giving (b).

On the other hand, suppose \(C(t)\le 0\) for all \(t\ge 0\). Since \(\dot{C}(t)=\dot{y}_1(t)^2 - C(t)\ge -C(t)\ge 0\), \(C(t)\) is non-decreasing and therefore converges to a limit \(\bar C\in(-\infty, 0]\), with \(\dot{C}(t)\to 0\). The identity \(\dot{C}(t)=\dot{y}_1(t)^2 - C(t)\) then forces \(\dot{y}_1(t)^2\to \bar C\); since \(\dot{y}_1(t)^2\ge 0\), \(\bar C\ge 0\), so \(\bar C=0\) and \(\dot{y}_1(t)\to 0\). The compact level set \(\{(y_1,y_2,y_3): y_1\in[l,u],\;y_2,y_3\;\text{bounded}\}\) together with \(\dot{y}_1(t)\to 0\) implies \(y_1(t)\) converges to a limit \(y_1^*\) (by LaSalle’s invariance principle), and continuity of \(B\) in 16 gives \(B(y_1^*,0)=y_1^*\), i.e., \(y_1^*=p^{NE}\). Finally, \(V(t)\to\nu^2\) from ?? (since \(\dot{y}_1(t)^2\to 0\) and \(V(t)\) satisfies a stable linear ODE with input \(\dot{y}_1^2+\nu^2\)), and \(C(t)\to 0\) implies \(y_3(t)\to y_1^{*2}=(p^{NE})^2\), while \(V(t)\to\nu^2\) implies \(y_2(t)\to y_1^{*2}+\nu^2=(p^{NE})^2+\nu^2\). Hence \(\mathbf{y}(t)\to \mathbf{y}^*\), yielding (a).

(iv) Differentiate \(\dot{y}_1=B(y_1,r)-y_1\) to obtain \[\ddot y_1 = (\partial_{y_1}B-1)\dot{y}_1 + (\partial_r B)\dot{r}.\] At any \(t_0\) with \(\dot{y}_1(t_0)=0\), we have \(q(t_0)=0\) in ?? , hence \(\dot{r}(t_0)= -\frac{\dot{y}_1(t_0)^2+\nu^2}{V(t_0)}\,r(t_0)<0\) whenever \(r(t_0)>0\). By Lemma 3, \(\partial_r B>0\) on \([l,u]\times[0,1)\), so \[\ddot y_1(t_0)=(\partial_r B)(y_1(t_0),r(t_0))\,\dot{r}(t_0)<0.\] Thus, every critical point with \(r>0\) is a strict local maximum, implying no local minima and at most one critical point on any interval where \(r\in(0,1)\). In case (b), we have \(r(t)\in(0,1)\) for all \(t\ge T_{NE}\), and there must be at least one critical point since otherwise \(y_1(t)\) would be monotonically increasing and either diverge or converge to a limit above \(p^{NE}\), contradicting (iii) or the equilibrium characterization 18 . After this maximum, \(\dot{y}_1<0\) and \(y_1(t)\) decreases monotonically toward \(p^{NE}\).

9.7 Proof of Theorem 4↩︎

Applying Theorem 6 with \(\mathcal{I}^{in} = [N]\) gives the parameter rate \[\label{eq:cor:param-rate} \sum_{i \in [N]} \mathbb{E}\|\hat{\theta}_{n,i} - \theta^*_i\|_2^2 \;=\; O\!\left(n^{\eta_{\max} - 1}\right).\tag{26}\]

Step 1: Linearizing the greedy best response around \(\mathbf{p}^{NE}\).

Stack the seller-\(i\) greedy best-response functions into the vector-valued map \(\boldsymbol{\Phi}(\theta;\mathbf{q}) \in \mathbb{R}^N\) with \(i\)th coordinate \[\label{eq:cor:phi-def} \boldsymbol{\Phi}_i(\theta_i; \mathbf{q}) \;\triangleq\; -\,\frac{\alpha_i + \sum_{j \neq i}\gamma_{i,j} q_j}{2\beta_i}, \qquad i \in [N].\tag{27}\] Recall from Section 3.3 the Nash matrix \(\Gamma \in \mathbb{R}^{N \times N}\) with \(\Gamma_{ii} = 2\beta_i\), \(\Gamma_{ij} = \gamma_{i,j}\) for \(j \neq i\), and \(\boldsymbol{\alpha} = (\alpha_1, \dots, \alpha_N)^\top\). The first-order condition \(\Gamma\mathbf{p}^{NE} = -\boldsymbol{\alpha}\) from Section 3.3 reads coordinate-wise as \[\label{eq:cor:NE-foc} 2\beta_i\,p_i^{NE} \;+\; \sum_{j \neq i}\gamma_{i,j}\,p_j^{NE} \;=\; -\alpha_i, \qquad i \in [N].\tag{28}\] Subtracting \(p_i^{NE}\) from 27 at the true parameters \(\theta^* = (\alpha, \beta, \gamma)\) and using 28 to substitute \(\alpha_i = -2\beta_i p_i^{NE} - \sum_{j \neq i}\gamma_{i,j} p_j^{NE}\): \[\begin{align} \boldsymbol{\Phi}_i(\theta^*; \mathbf{q}) - p_i^{NE} &\;=\; -\frac{\alpha_i + \sum_{j\neq i}\gamma_{i,j} q_j}{2\beta_i} - p_i^{NE} \;=\; -\frac{\alpha_i + 2\beta_i p_i^{NE} + \sum_{j\neq i}\gamma_{i,j} q_j}{2\beta_i} \notag\\ &\;=\; -\frac{-\sum_{j \neq i}\gamma_{i,j} p_j^{NE} + \sum_{j\neq i}\gamma_{i,j} q_j}{2\beta_i} \;=\; \sum_{j \neq i}\left(-\frac{\gamma_{i,j}}{2\beta_i}\right)(q_j - p_j^{NE}). \label{eq:cor:phi-shift-coord} \end{align}\tag{29}\] Define \(B \in \mathbb{R}^{N \times N}\) by \[B_{ii} \;=\; 0, \qquad B_{ij} \;=\; -\frac{\gamma_{i,j}}{2\beta_i} \quad \text{for } j \neq i, \qquad i \in [N].\] Stacking 29 into vector form yields the linearization identity \[\label{eq:cor:linear-BR} \boldsymbol{\Phi}(\theta^*;\mathbf{q}) \;-\; \mathbf{p}^{NE} \;=\; B\,(\mathbf{q} - \mathbf{p}^{NE}).\tag{30}\] Recall that \(B = I - \tfrac{1}{2}\mathrm{diag}(1/\beta_i)\,\Gamma\). The standing condition \(-\beta_i > \gamma_i \triangleq \sum_{j \neq i}\gamma_{i,j}\) from Section 3.1, which there ensured the strict diagonal dominance of \(\Gamma\) (i.e., \(|\Gamma_{ii}| = 2|\beta_i| > \gamma_i = \sum_{j \neq i}|\Gamma_{ij}|\)) and hence its invertibility, gives directly \[\sum_{j \neq i}\,|B_{ij}| \;=\; \sum_{j \neq i}\frac{\gamma_{i,j}}{2|\beta_i|} \;=\; \frac{\gamma_i}{2|\beta_i|} \;<\; \frac{1}{2}, \qquad i \in [N],\] since \(\gamma_{i,j} \ge 0\) and \(\gamma_i < |\beta_i|\). Hence \[\label{eq:cor:B-infty-norm} \|B\|_\infty \;\triangleq\; \max_{i \in [N]} \sum_{j \in [N]}|B_{ij}| \;=\; \max_{i \in [N]} \sum_{j \neq i}|B_{ij}| \;<\; \frac{1}{2}.\tag{31}\] Set \[\lambda \;\triangleq\; 1 - \|B\|_\infty \;\in\; \bigl(\tfrac{1}{2},\, 1\bigr].\] Applying Gershgorin’s circle theorem to \(I - B\): the \(i\)th diagonal entry is \((I-B)_{ii} = 1\) and the \(i\)th off-diagonal row-sum is \(\sum_{j \neq i}|(I-B)_{ij}| = \sum_{j \neq i}|B_{ij}| \le \|B\|_\infty = 1 - \lambda\). Thus every eigenvalue \(\zeta\) of \(I-B\) lies in the disc \(\{z \in \mathbb{C} : |z - 1| \le 1 - \lambda\} \subset \{z : \mathrm{Re}(z) \ge \lambda > 1/2\}\). In particular, \(I - B\) is positive stable: all its eigenvalues have real part bounded below by \(\lambda\).

By the interiority assumption \(\theta^*_i \in \operatorname{int}(\Theta^{in}_i)\) from Section 3.2, the LS-projection set \(\Theta^{in}_i\) has positive distance from \(\{\beta_i = 0\}\), so there exists \(\beta_{\min} > 0\) such that \(|\hat{\beta}_{n,i}| \ge \beta_{\min}\) for all \(i, n\) a.s. Differentiating 27 with respect to the components of \(\theta_i = (\alpha_i, \beta_i, (\gamma_{i,j})_{j \neq i})\): \[\begin{align} \frac{\partial \boldsymbol{\Phi}_i}{\partial \alpha_i} &\;=\; -\frac{1}{2\beta_i}, & \frac{\partial \boldsymbol{\Phi}_i}{\partial \beta_i} &\;=\; \frac{\alpha_i + \sum_{j \neq i}\gamma_{i,j} q_j}{2\beta_i^2}, & \frac{\partial \boldsymbol{\Phi}_i}{\partial \gamma_{i,j}} &\;=\; -\frac{q_j}{2\beta_i} \quad (j \neq i). \end{align}\] On \(\Theta^{in} \times [l, u]^N\) the parameters and \(q_j\) are bounded, so \(|\hat{\alpha}|, |\hat{\gamma}_{i,j}| \le C_\theta\) and \(|q_j| \le u\), hence each partial derivative is bounded in magnitude by some constant \(K_0 = K_0(\beta_{\min}, C_\theta, u)\). By the mean-value theorem along the segment from \(\theta^*_i\) to \(\hat{\theta}_{n,i}\), \[|\boldsymbol{\Phi}_i(\hat{\theta}_{n,i}; \mathbf{m}_n) - \boldsymbol{\Phi}_i(\theta^*_i; \mathbf{m}_n)| \;\le\; K_0\,\|\hat{\theta}_{n,i} - \theta^*_i\|_2.\] Stacking and using \(\|\mathbf{v}\|_2 \le \sqrt{N}\|\mathbf{v}\|_\infty\), \[\label{eq:cor:phi-lipschitz} \|\boldsymbol{\Phi}(\hat{\theta}_n; \mathbf{m}_n) - \boldsymbol{\Phi}(\theta^*; \mathbf{m}_n)\|_2 \;\le\; K\,\sum_{i \in [N]}\|\hat{\theta}_{n,i} - \theta^*_i\|_2, \qquad K \triangleq K_0\sqrt{N}.\tag{32}\] Write the realized price as \[\mathbf{p}_{n+1} \;=\; \boldsymbol{\Phi}(\hat{\theta}_n;\mathbf{m}_n) \;+\; \mathbf{z}_{n+1},\] where \(\mathbf{z}_{n+1}\) is the exploration noise vector with \(\mathbb{E}[\mathbf{z}_{n+1} \mid \mathcal{F}_n] = \mathbf{0}\), independent coordinates, and second moments \(\mathbb{E}[z_{n+1,i}^2] = \nu_{n+1,i}^2\). Adding and subtracting \(\boldsymbol{\Phi}(\theta^*; \mathbf{m}_n)\) and applying 30 with \(\mathbf{q} = \mathbf{m}_n\): \[\begin{align} \mathbf{p}_{n+1} - \mathbf{p}^{NE} &\;=\; \bigl[\boldsymbol{\Phi}(\theta^*; \mathbf{m}_n) - \mathbf{p}^{NE}\bigr] \;+\; \bigl[\boldsymbol{\Phi}(\hat{\theta}_n; \mathbf{m}_n) - \boldsymbol{\Phi}(\theta^*; \mathbf{m}_n)\bigr] \;+\; \mathbf{z}_{n+1} \notag\\ &\;=\; B\,(\mathbf{m}_n - \mathbf{p}^{NE}) \;+\; \boldsymbol{\delta}_n \;+\; \mathbf{z}_{n+1}, \label{eq:cor:price-decomp} \end{align}\tag{33}\] where \[\boldsymbol{\delta}_n \;\triangleq\; \boldsymbol{\Phi}(\hat{\theta}_n; \mathbf{m}_n) - \boldsymbol{\Phi}(\theta^*; \mathbf{m}_n).\] By 32 , \(\boldsymbol{\delta}_n\) is \(\mathcal{F}_n\)-measurable and obeys \[\label{eq:cor:delta-bound} \mathbb{E}\|\boldsymbol{\delta}_n\|_2^2 \;\le\; K^2\,\mathbb{E}\!\left(\sum_{i \in [N]}\|\hat{\theta}_{n,i} - \theta^*_i\|_2\right)^{\!2} \;\le\; K^2 N\,\sum_{i \in [N]}\mathbb{E}\|\hat{\theta}_{n,i} - \theta^*_i\|_2^2 \;=\; O\!\left(n^{\eta_{\max} - 1}\right),\tag{34}\] where the second inequality is Cauchy–Schwarz applied to the unit-weight inner product on \(\mathbb{R}^N\), and the last equality is 26 .

Step 2: Stochastic-approximation recursion for the running mean.

Let \[\mathbf{e}_n \;\triangleq\; \mathbf{m}_n - \mathbf{p}^{NE}.\] The running-mean update \(\mathbf{m}_{n+1} = \mathbf{m}_n + (n+1)^{-1}(\mathbf{p}_{n+1} - \mathbf{m}_n)\) subtracted on both sides by \(\mathbf{p}^{NE}\) becomes \[\begin{align} \mathbf{e}_{n+1} &\;=\; \mathbf{e}_n + \frac{1}{n+1}\bigl(\mathbf{p}_{n+1} - \mathbf{m}_n\bigr) \;=\; \mathbf{e}_n + \frac{1}{n+1}\bigl[(\mathbf{p}_{n+1} - \mathbf{p}^{NE}) - (\mathbf{m}_n - \mathbf{p}^{NE})\bigr] \\ &\;=\; \left(1 - \frac{1}{n+1}\right)\mathbf{e}_n + \frac{1}{n+1}\bigl(\mathbf{p}_{n+1} - \mathbf{p}^{NE}\bigr). \end{align}\] Substituting 33 for \(\mathbf{p}_{n+1} - \mathbf{p}^{NE}\): \[\begin{align} \mathbf{e}_{n+1} &\;=\; \left(1 - \frac{1}{n+1}\right)\mathbf{e}_n + \frac{1}{n+1}\bigl[B\,\mathbf{e}_n + \boldsymbol{\delta}_n + \mathbf{z}_{n+1}\bigr] \notag\\ &\;=\; \left[I - \frac{I - B}{n+1}\right]\mathbf{e}_n + \frac{\boldsymbol{\delta}_n + \mathbf{z}_{n+1}}{n+1}. \end{align}\] Define \[\label{eq:cor:An-def} A_n \;\triangleq\; I - \frac{I - B}{n + 1},\tag{35}\] so that \(\mathbf{e}_{n+1} = A_n \mathbf{e}_n + (n+1)^{-1}(\boldsymbol{\delta}_n + \mathbf{z}_{n+1})\). Since \(I-B\) is positive stable, the Lyapunov theorem produces a unique symmetric \(P \succ 0\) such that \[\label{eq:cor:Lyap-eq} P(I-B) + (I-B)^\top P \;=\; I.\tag{36}\] Since \(P \succ 0\) implies \(P \preceq \sigma_{\max}(P)\,I\), we have \(I \succeq \sigma_{\max}(P)^{-1}\,P\), so 36 gives the spectral inequality \[\label{eq:cor:Lyap-bound} P(I-B) + (I-B)^\top P \;\succeq\; \mu\,P, \qquad \mu \;\triangleq\; \frac{1}{\sigma_{\max}(P)} \;>\; 0.\tag{37}\] Compute \(A_n^\top P A_n\) from 35 , using 36 to collapse the first-order term: \[\begin{align} A_n^\top P A_n &\;=\; \left[I - \frac{(I-B)^\top}{n+1}\right] P \left[I - \frac{I-B}{n+1}\right] \notag\\ &\;=\; P \;-\; \frac{P(I-B) + (I-B)^\top P}{n+1} \;+\; \frac{(I-B)^\top P (I-B)}{(n+1)^2} \notag\\ &\;=\; P \;-\; \frac{I}{n+1} \;+\; \frac{(I-B)^\top P (I-B)}{(n+1)^2}. \end{align}\] Using \(-I \preceq -\mu P\) from 37 and bounding the second-order term \((I-B)^\top P (I-B) \preceq \|I-B\|_2^2\,\sigma_{\max}(P)\,I \preceq C_3\,P\) with \(C_3 \triangleq \|I-B\|_2^2\,\sigma_{\max}(P)/\sigma_{\min}(P)\), \[A_n^\top P A_n \;\preceq\; \left(1 - \frac{\mu}{n+1} + \frac{C_3}{(n+1)^2}\right) P.\] For all \(n \ge n_0 \triangleq 2C_3/\mu\), \(C_3/(n+1)^2 \le \mu/(2(n+1))\), hence \[\label{eq:cor:AnPAn-tight} A_n^\top P A_n \;\preceq\; \left(1 - \frac{\mu}{2(n+1)}\right) P, \qquad n \ge n_0.\tag{38}\]

Define the weighted Lyapunov \(V_n\), the associated \(P\)-norm, and the root-mean-square \(r_n\) by \[V_n \;\triangleq\; \mathbf{e}_n^\top P\,\mathbf{e}_n, \qquad \|u\|_P \;\triangleq\; \sqrt{u^\top P\,u}, \qquad r_n \;\triangleq\; \sqrt{\mathbb{E} V_n} \;=\; \sqrt{\mathbb{E}\|\mathbf{e}_n\|_P^2}.\] By norm equivalence, \[\label{eq:cor:V-norm-equiv} \sigma_{\min}(P)\,\|\mathbf{e}_n\|_2^2 \;\le\; V_n \;\le\; \sigma_{\max}(P)\,\|\mathbf{e}_n\|_2^2.\tag{39}\] Applying the \(P\)-norm triangle inequality to the recursion \(\mathbf{e}_{n+1} = A_n\mathbf{e}_n + (n+1)^{-1}(\boldsymbol{\delta}_n + \mathbf{z}_{n+1})\), \[\label{eq:cor:Pnorm-tri} \|\mathbf{e}_{n+1}\|_P \;\le\; \|A_n \mathbf{e}_n\|_P \;+\; \frac{\|\boldsymbol{\delta}_n\|_P}{n+1} \;+\; \frac{\|\mathbf{z}_{n+1}\|_P}{n+1}.\tag{40}\] View the three terms on the right as random variables \(X_n \triangleq \|A_n\mathbf{e}_n\|_P\), \(Y_n \triangleq \|\boldsymbol{\delta}_n\|_P/(n+1)\), \(Z_n \triangleq \|\mathbf{z}_{n+1}\|_P/(n+1)\), and write \(\|U\|_{L^2(\mathbb{P})} \triangleq \sqrt{\mathbb{E}\,U^2}\). Squaring 40 and taking expectations preserves the inequality (since all terms are nonnegative), giving \[r_{n+1}^2 \;=\; \mathbb{E}\|\mathbf{e}_{n+1}\|_P^2 \;\le\; \mathbb{E}\bigl[(X_n + Y_n + Z_n)^2\bigr] \;=\; \|X_n + Y_n + Z_n\|_{L^2(\mathbb{P})}^2.\] Taking square roots yields \(r_{n+1} \le \|X_n + Y_n + Z_n\|_{L^2(\mathbb{P})}\). Minkowski’s inequality applied twice gives \(\|X_n + Y_n + Z_n\|_{L^2(\mathbb{P})} \le \|X_n\|_{L^2(\mathbb{P})} + \|Y_n\|_{L^2(\mathbb{P})} + \|Z_n\|_{L^2(\mathbb{P})}\), hence \[\label{eq:cor:rn-Mink} r_{n+1} \;\le\; \sqrt{\mathbb{E}\|A_n\mathbf{e}_n\|_P^2} \;+\; \frac{\sqrt{\mathbb{E}\|\boldsymbol{\delta}_n\|_P^2}}{n+1} \;+\; \frac{\sqrt{\mathbb{E}\|\mathbf{z}_{n+1}\|_P^2}}{n+1}.\tag{41}\] We bound each of the three \(L^2\) norms. For the first, 38 gives \(\mathbb{E}\|A_n\mathbf{e}_n\|_P^2 \le (1 - \mu/(2(n+1)))\,r_n^2\) for \(n \ge n_0\); the elementary inequality \(\sqrt{1-x} \le 1 - x/2\) for \(x \in [0, 1]\) then yields \[\label{eq:cor:rn-contract} \sqrt{\mathbb{E}\|A_n\mathbf{e}_n\|_P^2} \;\le\; \sqrt{1 - \mu/(2(n+1))}\,r_n \;\le\; \left(1 - \frac{\mu}{4(n+1)}\right) r_n.\tag{42}\] For the second, \(\mathbb{E}\|\boldsymbol{\delta}_n\|_P^2 \le \sigma_{\max}(P)\,\mathbb{E}\|\boldsymbol{\delta}_n\|_2^2 \le \sigma_{\max}(P)\,D_\delta\,n^{\eta_{\max}-1}\), where \(D_\delta > 0\) is the implied constant in 34 (concretely, \(D_\delta \le K^2 N\,D_\theta\) with \(D_\theta\) the constant in \(\sum_i \mathbb{E}\|\hat{\theta}_{n,i} - \theta^*_i\|_2^2 \le D_\theta\,n^{\eta_{\max}-1}\) from Theorem 6), so \[\sqrt{\mathbb{E}\|\boldsymbol{\delta}_n\|_P^2}/(n+1) \;\le\; K_\delta\,n^{(\eta_{\max}-3)/2}, \qquad K_\delta \triangleq \sqrt{\sigma_{\max}(P)\,D_\delta}.\] For the third, \(\mathbb{E}\|\mathbf{z}_{n+1}\|_P^2 = \mathrm{tr}(P\,\Sigma_{n+1}) \le N\sigma_{\max}(P)\,\nu_{\max,n+1}^2 \le N\sigma_{\max}(P)\,c_\nu\,n^{-\eta_{\min}}\), where \(c_\nu\) is the constant in \(\nu_{\max,n}^2 \le c_\nu\,n^{-\eta_{\min}}\), so \[\label{eq:cor:rn-noise} \sqrt{\mathbb{E}\|\mathbf{z}_{n+1}\|_P^2}/(n+1) \;\le\; K_z\,n^{-(2+\eta_{\min})/2}, \qquad K_z \triangleq \sqrt{N\sigma_{\max}(P)\,c_\nu}.\tag{43}\] Substituting 4243 into 41 : for \(n \ge n_0\), abbreviating \(\rho_n \triangleq 1 - \mu/(4(n+1))\) and \(g_n \triangleq K_\delta\,n^{(\eta_{\max}-3)/2} + K_z\,n^{-(2+\eta_{\min})/2}\), \[\label{eq:cor:rn-recursion} r_{n+1} \;\le\; \rho_n\,r_n \;+\; g_n.\tag{44}\] Iterating 44 from \(n_0\) to \(n - 1\) gives \[r_n \le \left(\prod_{m=n_0}^{n-1}\rho_m\right) r_{n_0} \;+\; \sum_{k=n_0}^{n-1} g_k\,\prod_{m=k+1}^{n-1}\rho_m,\] with the convention that the empty product \(\prod_{m=n}^{n-1}\rho_m = 1\). Using \(\log(1 - x) \le -x\) for \(x \in [0, 1)\) and the integral lower bound \(\sum_{m=k+1}^{n-1}1/(m+1) \ge \log((n+1)/(k+2))\) on the contraction products, \[\begin{align} \prod_{m=k+1}^{n-1}\rho_m &\;=\; \exp\!\left(\sum_{m=k+1}^{n-1}\!\log\!\left(1 - \frac{\mu}{4(m+1)}\right)\right) \;\le\; \exp\!\left(-\frac{\mu}{4}\!\sum_{m=k+1}^{n-1}\!\frac{1}{m+1}\right) \\ &\;\le\; \left(\frac{k+2}{n+1}\right)^{\!\mu/4} \;\le\; C_0\,\frac{(k+1)^{\mu/4}}{n^{\mu/4}}, \end{align}\] where \(C_0\) absorbs the index shifts. Applying this bound to the recursion gives, \[\label{eq:cor:rn-telescope} r_n \;\le\; C_1\,n^{-\mu/4} \;+\; C_0\,K_\delta\,n^{-\mu/4}\!\!\sum_{k=n_0}^{n-1}\!\!(k+1)^{\mu/4}\,k^{(\eta_{\max}-3)/2} \;+\; C_0\,K_z\,n^{-\mu/4}\!\!\sum_{k=n_0}^{n-1}\!\!(k+1)^{\mu/4}\,k^{-(2+\eta_{\min})/2},\tag{45}\] where \(C_1 \;\triangleq\; C_0\,r_{n_0}\,(n_0+1)^{\mu/4}\) is the constant carrying the initial condition \(r_{n_0}\) through the homogeneous decay. It remains to evaluate the two sums in 45 . We reduce each to a pure power sum via \((k+1)^{\mu/4} \le 2^{\mu/4}\,k^{\mu/4}\) (valid for \(k \ge 1\)): for any \(b \in \mathbb{R}\), \[\label{eq:cor:sum-reduction} \sum_{k=n_0}^{n-1}(k+1)^{\mu/4}\,k^{b} \;\le\; 2^{\mu/4}\!\sum_{k=n_0}^{n-1}\!k^{\,\mu/4 + b} \;\le\; 2^{\mu/4}\!\sum_{k=1}^{n}\!k^{\,\mu/4 + b}.\tag{46}\] The standard integral comparison gives, for any \(a \in \mathbb{R}\) and \(n \ge 2\), \[\label{eq:cor:power-sum} \sum_{k=1}^{n}k^{a} \;\le\; C(a)\, \begin{cases} n^{a+1} & \text{if } a > -1, \\ 1 + \log n & \text{if } a = -1, \\ 1 & \text{if } a < -1, \end{cases}\tag{47}\] with \(C(a) > 0\) depending only on \(a\).

First sum. Set \(a_1 \triangleq \mu/4 + (\eta_{\max} - 3)/2\). By 46 with \(b = (\eta_{\max}-3)/2\) and 47 with \(a = a_1\), then multiplying by \(n^{-\mu/4}\) and using \(a_1 + 1 - \mu/4 = (\eta_{\max}-1)/2\): \[\begin{align} n^{-\mu/4}\!\!\sum_{k=n_0}^{n-1}\!\!(k+1)^{\mu/4}\,k^{(\eta_{\max}-3)/2} &\;\le\; 2^{\mu/4}\,C(a_1)\,n^{-\mu/4}\,\times\! \begin{cases} n^{a_1+1} & \text{if } a_1 > -1, \\ 1 + \log n & \text{if } a_1 = -1, \\ 1 & \text{if } a_1 < -1 \end{cases}\notag \\ &\;=\; \begin{cases} O\!\left(n^{(\eta_{\max}-1)/2}\right) & \text{if } a_1 > -1, \\[2pt] O\!\left(n^{-\mu/4}\,\log n\right) & \text{if } a_1 = -1, \\[2pt] O\!\left(n^{-\mu/4}\right) & \text{if } a_1 < -1. \end{cases} \end{align}\] Combining the three regimes into a single uniform bound by attaching the boundary \(\log n\) factor to the \(n^{-\mu/4}\) term, \[\label{eq:cor:S1-bound} n^{-\mu/4}\!\!\sum_{k=n_0}^{n-1}\!\!(k+1)^{\mu/4}\,k^{(\eta_{\max}-3)/2} \;=\; O\!\left(n^{-\mu/4}\,\log n \;+\; n^{(\eta_{\max}-1)/2}\right).\tag{48}\]

Second sum. The same argument with \(a_2 \triangleq \mu/4 - (2+\eta_{\min})/2\), observing \(a_2 + 1 - \mu/4 = -\eta_{\min}/2\), yields \[\label{eq:cor:S2-bound} n^{-\mu/4}\!\!\sum_{k=n_0}^{n-1}\!\!(k+1)^{\mu/4}\,k^{-(2+\eta_{\min})/2} \;=\; O\!\left(n^{-\mu/4}\,\log n \;+\; n^{-\eta_{\min}/2}\right).\tag{49}\]

Substituting 4849 into 45 and collapsing constants into \[C_2 \;\triangleq\; C_0\,K_\delta\,C(\tfrac{\mu}{4} + \tfrac{\eta_{\max}-3}{2}), \qquad C_3 \;\triangleq\; C_0\,K_z\,C(\tfrac{\mu}{4} - \tfrac{2+\eta_{\min}}{2}),\] we obtain \[\label{eq:cor:rn-rate} r_n \;\le\; \widetilde{C}_1\,n^{-\mu/4}\,\log n \;+\; C_2\,n^{(\eta_{\max}-1)/2} \;+\; C_3\,n^{-\eta_{\min}/2},\tag{50}\] where \(\widetilde{C}_1\) absorbs \(C_1\) and the multiplicative constants from 4849 (the bare \(C_1 n^{-\mu/4}\) term is dominated by \(\widetilde{C}_1 n^{-\mu/4}\log n\)). Squaring 50 and using \((a+b+c)^2 \le 3(a^2+b^2+c^2)\): \[\begin{align} W_n \;=\; r_n^2 &\;\le\; 3\,\widetilde{C}_1^2\,n^{-\mu/2}\,(\log n)^2 \;+\; 3\,C_2^2\,n^{\eta_{\max}-1} \;+\; 3\,C_3^2\,n^{-\eta_{\min}} \notag\\ &\;=\; O\!\left(n^{-\mu/2}(\log n)^2 + n^{\eta_{\max} - 1} + n^{-\eta_{\min}}\right). \end{align}\] Combining with 39 gives the running-mean \(L^2\) rate \[\label{eq:cor:mn-rate} \mathbb{E}\|\mathbf{m}_n - \mathbf{p}^{NE}\|_2^2 \;\le\; \frac{W_n}{\sigma_{\min}(P)} \;=\; O\!\left(n^{-\mu/2}(\log n)^2 + n^{\eta_{\max} - 1} + n^{-\eta_{\min}}\right).\tag{51}\]

Step 3: \(L^2\) rate for the realized price.

Squaring 33 in the \(\ell_2\) norm: \[\|\mathbf{p}_{n+1} - \mathbf{p}^{NE}\|_2^2 \;=\; \|B\mathbf{e}_n + \boldsymbol{\delta}_n + \mathbf{z}_{n+1}\|_2^2.\] Expanding and taking conditional expectation, using \(\mathbb{E}[\mathbf{z}_{n+1}\mid \mathcal{F}_n] = 0\) to kill the cross-terms involving \(\mathbf{z}_{n+1}\): \[\begin{align} \mathbb{E}\!\left[\|\mathbf{p}_{n+1} - \mathbf{p}^{NE}\|_2^2 \,\middle|\, \mathcal{F}_n\right] &\;=\; \|B\mathbf{e}_n + \boldsymbol{\delta}_n\|_2^2 \;+\; \mathbb{E}\|\mathbf{z}_{n+1}\|_2^2 \\ &\;\le\; 2\|B\|_2^2 \|\mathbf{e}_n\|_2^2 + 2\|\boldsymbol{\delta}_n\|_2^2 + \sum_{i \in [N]}\nu_{n+1,i}^2, \end{align}\] where the inequality uses \(\|u+v\|^2 \le 2\|u\|^2 + 2\|v\|^2\). Taking unconditional expectations, \[\label{eq:cor:pn-L2-pre} \mathbb{E}\|\mathbf{p}_{n+1} - \mathbf{p}^{NE}\|_2^2 \;\le\; 2\|B\|_2^2\,\mathbb{E}\|\mathbf{e}_n\|_2^2 + 2\,\mathbb{E}\|\boldsymbol{\delta}_n\|_2^2 + N\,\nu_{\max,n+1}^2.\tag{52}\] The matrix norm \(\|B\|_2\) is finite with \(\|B\|_2 < \infty\). Substituting 51 , 34 , and \(\nu_{\max,n+1}^2 = O(n^{-\eta_{\min}})\) into 52 : \[\label{eq:cor:pn-rate-3terms} \mathbb{E}\|\mathbf{p}_n - \mathbf{p}^{NE}\|_2^2 \;=\; O\!\left(n^{-\mu/2}(\log n)^2 \;+\; n^{\eta_{\max} - 1} \;+\; n^{-\eta_{\min}}\right).\tag{53}\]

Step 4: Almost-sure convergence.

We now upgrade the \(L^2\) statement to a.s.convergence. Lemma 6(b) applies with \(\cI^{in} = [N]\) and gives \[\label{eq:cor:theta-as} \hat{\theta}_{n,k} \;\to\; \theta_k^* \quad\text{a.s., for every } k\in[N],\tag{54}\] which by 32 also yields \(\boldsymbol{\delta}_n \to 0\) a.s.and \(\sum_n \|\boldsymbol{\delta}_n\|_2^2/(n+1) < \infty\) a.s.(Lemma 6(c)). Take the recursion \(\mathbf{e}_{n+1} = A_n\mathbf{e}_n + (n+1)^{-1}(\boldsymbol{\delta}_n + \mathbf{z}_{n+1})\) from Step 2 and compute \[\begin{align} \EE[V_{n+1}\mid\cF_n] &\;=\; \mathbf{e}_n^\top A_n^\top P A_n\,\mathbf{e}_n \;+\; \frac{2}{n+1}\,\mathbf{e}_n^\top A_n^\top P\,\boldsymbol{\delta}_n \\ &\quad +\; \frac{1}{(n+1)^2}\,\|\boldsymbol{\delta}_n\|_P^2 \;+\; \frac{1}{(n+1)^2}\,\mathrm{tr}(P\,\Sigma_{n+1}), \end{align}\] where the cross-term involving \(\mathbf{z}_{n+1}\) vanishes because \(\EE[\mathbf{z}_{n+1}\mid\cF_n] = 0\). Applying 38 to the first piece and Young’s inequality with \(\epsilon = \mu/(4(n+1))\) to the second, \[\label{eq:cor:RS-ineq} \EE[V_{n+1}\mid\cF_n] \;\le\; \left(1 - \frac{\mu}{4(n+1)}\right)V_n \;+\; h_n\tag{55}\] holds eventually a.s., with \[h_n \;\triangleq\; \frac{4\sigma_{\max}(P)\,\|\boldsymbol{\delta}_n\|_2^2}{\mu(n+1)} \;+\; \frac{\sigma_{\max}(P)\,\|\boldsymbol{\delta}_n\|_2^2}{(n+1)^2} \;+\; \frac{N\sigma_{\max}(P)\,\nu_{\max,n+1}^2}{(n+1)^2}.\] The first two pieces of \(h_n\) are bounded by a constant multiple of \(\|\boldsymbol{\delta}_n\|_2^2/(n+1)\), which is a.s.-summable by Lemma 6(c) (applied with \(\cI^{in} = [N]\) together with 32 ). The third piece is \(O(n^{-(2+\eta_{\min})})\), deterministically summable. Hence \(\sum_n h_n < \infty\) a.s.

Robbins–Siegmund (in the form \(\EE[V_{n+1}\mid\cF_n] \le (1+a_n)V_n - b_n + c_n\) with \(a_n \equiv 0\), \(b_n \triangleq (\mu/(4(n+1)))\,V_n \ge 0\), \(c_n \triangleq h_n\), \(\sum_n c_n < \infty\) a.s.) applied to 55 gives that \(V_n\) converges to a finite limit \(V_\infty\) a.s.and \(\sum_n b_n = (\mu/4)\sum_n V_n/(n+1) < \infty\) a.s. Since \(\sum_n 1/(n+1) = \infty\), the latter forces \(V_\infty = 0\) a.s. Combining with the \(P\)-norm equivalence 39 , \(\mathbf{m}_n \to \mathbf{p}^{NE}\) a.s. Finally, from the identity \[\tilde{\mathbf{p}}_{n+1} - \mathbf{p}^{NE} \;=\; \boldsymbol{\Phi}(\hat{\theta}_n;\mathbf{m}_n) - \mathbf{p}^{NE} \;=\; B\,(\mathbf{m}_n - \mathbf{p}^{NE}) \;+\; \boldsymbol{\delta}_n\] (from 30 and the definition of \(\boldsymbol{\delta}_n\)), both summands vanish a.s., so \(\tilde{p}_{n,k} \to p_k^{NE}\) a.s.for every \(k\in[N]\).

Step 5: Special case.

It remains to prove the regret bound under \(\eta_i = 1/2\) for every \(i \in [N]\) and ?? . Write \[A_s \;\triangleq\; \tfrac12\bigl((I - B) + (I - B)^\top\bigr) \;=\; I - \tfrac12 (B + B^\top)\] for the symmetric part of \(I - B\). Let \(v \in \mathbb{R}^N\) be a unit eigenvector of \(P\) corresponding to its largest eigenvalue, so that \(P v = \sigma_{\max}(P)\, v\) and \(\|v\|_2 = 1\). Multiplying the Lyapunov equation \(P(I-B) + (I-B)^\top P = I\) on the left by \(v^\top\) and on the right by \(v\), \[\begin{align} 1 \;=\; v^\top v \;&=\; v^\top\!\bigl[P(I-B) + (I-B)^\top P\bigr] v \\ \;&=\; v^\top P (I-B) v + v^\top (I-B)^\top P v \\ \;&=\; 2\,\sigma_{\max}(P)\, v^\top (I-B) v \\ \;&=\; 2\,\sigma_{\max}(P)\, v^\top A_s v, \end{align}\] where the fourth equality uses \(v^\top P = \sigma_{\max}(P)\, v^\top\) together with \(v^\top (I-B)^\top P v = (v^\top P (I-B) v)^\top = v^\top P (I-B) v\) (the transpose of a scalar), and the last equality uses that \(v^\top A_a v = 0\) for the antisymmetric part \(A_a \triangleq \tfrac12\bigl((I-B) - (I-B)^\top\bigr)\) (which equals its own transpose as a scalar and also satisfies \(v^\top A_a v = -v^\top A_a v\) since \(A_a^\top = -A_a\)). Therefore, by the standard Rayleigh quotient bound, \[\sigma_{\max}(P) \;=\; \frac{1}{2\, v^\top A_s v} \;\le\; \frac{1}{2\,\lambda_{\min}(A_s)}.\] By definition, \(\mu \triangleq 1/\sigma_{\max}(P)\) and \(\lambda_{\min}(A_s) = 1 - \tfrac12\,\lambda_{\max}(B+B^\top)\), so \(\mu \ge 2 - \lambda_{\max}(B+B^\top)\). Condition ?? then yields \(\mu > 1\).

At \(\eta_i = 1/2\) for every \(i\) we have \(\eta_{\max} = \eta_{\min} = 1/2\), so condition ?? holds (\(\eta_{\min} + 1 - 2\eta_{\max} = 1/2 > 0\)) and the general rate ?? specializes to \[\mathbb{E}\|\mathbf{p}_n - \mathbf{p}^{NE}\|_2^2 \;=\; O\!\left(n^{-\mu/2}(\log n)^2 \;+\; n^{-1/2}\right).\] Since \(\mu > 1\), the first term is dominated, giving \(\mathbb{E}\|\mathbf{p}_n - \mathbf{p}^{NE}\|_2^2 = O(n^{-1/2})\). Summing over \(n\) from \(1\) to \(T\) and writing \(\mathbf{e}_n \triangleq \mathbf{p}_n - \mathbf{p}^{NE}\), \[\sum_{n=1}^T \mathbb{E}\|\mathbf{e}_n\|_2^2 \;=\; O(\sqrt T).\] Corollary 1 then gives \(\sum_{i=1}^N \Delta_i(\theta_i, T) \asymp \sum_{n=1}^T \mathbb{E}\|\mathbf{e}_n\|_2^2 = O(\sqrt T)\), and Proposition [prop:dynamicbenchmark] gives \(\Delta_i(\theta_i, T) = O(\sqrt T)\) for every \(i \in [N]\). This completes the proof.

9.8 Proof of Theorem 5↩︎

We follow the structure of the proof of Theorem 2 (multi-seller Lyapunov on the projected oblivious OLS error) combined with an auxiliary Lyapunov for the running-mean deviations that drive the informed-side dynamics. Throughout the proof, \(\cF_n\) is the natural filtration of the price–demand history up to time \(n\), and \(h_{n+1,i} \triangleq x_{n+1,i}^\top S_{n,i}^{-1} x_{n+1,i}\) satisfies \(h_{n+1,i} \le \delta_h\) eventually a.s.for any \(\delta_h > 0\) by Lemma 10. We work with the quantities \[\begin{align} \tilde{e}_{n,i} &\;\triangleq\; \tilde{\theta}_{n,i}^{ob} - \theta_i^{*,ob} & &(i\in\cI^{ob}), & \tilde{e}_{n,j}^{in} &\;\triangleq\; \hat{\theta}_{n,j}^{in} - \theta_j^{*,in} & &(j\in\cI^{in}),\\ u_{n,k} &\;\triangleq\; m_{n,k} - p_k^{NE} & &(k\in[N]), & U_{n,k} &\;\triangleq\; u_{n,k}^2, & U_n^{tot} &\;\triangleq\; \sum_{k=1}^N U_{n,k},\\ W_{n,i} &\;\triangleq\; \frac{1}{n}\, \tilde{e}_{n,i}^\top S_{n,i}\, \tilde{e}_{n,i}, & & & W_n^{ob} &\;\triangleq\; \sum_{i \in \cI^{ob}} W_{n,i}. \end{align}\] Within this proof we further split \[\bar\Lambda \;=\; \bar\Delta + \bar\Theta \qquad\text{and}\qquad \bar\kappa \;=\; \bar c_{\mathrm{diag}} - \bar D\] into their four constituents: \[\begin{align} \bar\Delta &\;\triangleq\; \max_{i \in \cI^{ob}}\sum_{j \in \cI^{in}}\frac{\gamma_{i,j}\,\gamma_j}{2|\beta_j|},\\ \bar\Theta &\;\triangleq\; L_\phi^{in,\theta}\,\max_{i \in \cI^{ob}}\sum_{j \in \cI^{in}}\gamma_{i,j},\\ \bar c_{\mathrm{diag}} &\;\triangleq\; \min\Bigl\{2 - L_\phi^{ob},\;\; 2 - L_\phi^{in,\theta} - \max_{j \in \cI^{in}}\frac{\gamma_j}{2|\beta_j|}\Bigr\},\\ \bar D &\;\triangleq\; \max_{k \in [N]}\sum_{j \in \cI^{in}\setminus\{k\}}\frac{\gamma_{j,k}}{2|\beta_j|}, \end{align}\] which arise as natural coefficients in the recurrences below. We also abbreviate the oblivious-side small-gain slack from condition (iv) and the informed-side small-gain margin from condition (ii) as \[\label{eq:K1-K2-proof} K_1 \;\triangleq\; C_M - C_x\,(2\bar\gamma^{ob}\, L_\phi^{ob} + \bar\Lambda) \;>\; 0, \qquad K_2 \;\triangleq\; \bar\kappa \;>\; 0,\tag{56}\] two scalar constants that drive the per-\(\lambda\) rate analysis below. With these, the bottleneck \(c^*\) in the statement of Theorem 5 reads \(c^* = \sup_{\lambda > 0}\min\bigl\{1 + (K_1 - \lambda L_\phi^{ob})/C_x^2,\; K_2 - C_x\bar\Psi/\lambda\bigr\}\).

9.8.0.1 Step 1. Oblivious OLS recursion.

For each \(i \in \cI^{ob}\), the regression residual against the pseudo-true parameter is \[\label{eq:proof-mm-eta-def} \eta_{n,i} \;=\; \varepsilon_{n,i} \;+\; \sum_{j \ne i}\gamma_{i,j}\bigl(p_{n,j} - p_j^{NE}\bigr).\tag{57}\] The Sherman–Morrison derivation of Step 1 in the proof of Theorem 2 applies verbatim to give \[\begin{align} W_{n+1,i} &\;=\; \frac{n}{n+1}\, W_{n,i} \;-\; \frac{(x_{n+1,i}^\top \tilde{e}_{n,i})^2}{(n+1)(1 + h_{n+1,i})}\notag\\ &\quad +\;\frac{2\,\eta_{n+1,i}\, x_{n+1,i}^\top \tilde{e}_{n,i}}{(n+1)(1 + h_{n+1,i})} \;+\; \frac{h_{n+1,i}\,\eta_{n+1,i}^2}{(n+1)(1 + h_{n+1,i})}. \label{eq:proof-mm-Wrec-pre} \end{align}\tag{58}\] Take conditional expectation given \(\cF_n\). The first term keeps \(\tfrac{n}{n+1}\,W_{n,i}\) as is. For the second, condition (iii) gives the persistent-excitation lower bound \(\EE[(x_{n+1,i}^\top \tilde{e}_{n,i})^2\mid\cF_n] \ge C_M\,\norm{\tilde{e}_{n,i}}_2^2\), which, combined with \(h_{n+1,i} \le \delta_h\) eventually a.s., yields \[\label{eq:proof-mm-pe-bd} \frac{\EE[(x_{n+1,i}^\top \tilde{e}_{n,i})^2\mid\cF_n]}{(n+1)(1 + h_{n+1,i})} \;\ge\; \frac{C_M\,\norm{\tilde{e}_{n,i}}_2^2}{(n+1)(1+\delta_h)} \quad\text{eventually a.s.}\tag{59}\] The third term is the cross term, retained for separate treatment in Step 2. For the fourth term, \(h/(1+h)\le h\) for \(h\ge 0\), and Lemma 10 gives \(h_{n+1,i} \le C_x^2/[(C_M - \delta)\,n]\) a.s.eventually, so together with \(C_\eta^2 \triangleq \sup_{n,i}\EE[\eta_{n+1,i}^2\mid\cF_n] < \infty\) (bounded prices and demand noise), \[\label{eq:proof-mm-h-eta-bd} \frac{\EE[h_{n+1,i}\,\eta_{n+1,i}^2\mid\cF_n]}{(n+1)(1 + h_{n+1,i})} \;\le\; \frac{C_x^2\, C_\eta^2}{(C_M - \delta)}\cdot \frac{1}{n(n+1)} \quad\text{eventually a.s.}\tag{60}\] Combining 5960 in 58 , eventually a.s., \[\begin{align} \EE[W_{n+1,i}\mid\cF_n] &\;\le\; \frac{n}{n+1}\,W_{n,i} \;-\; \frac{C_M\,\norm{\tilde{e}_{n,i}}_2^2}{(n+1)(1+\delta_h)}\notag\\ &\quad +\;\frac{2\,\bigl|\EE[\eta_{n+1,i}\, x_{n+1,i}^\top \tilde{e}_{n,i}\mid\cF_n]\bigr|}{n+1} \;+\; \frac{C_1}{n(n+1)},\label{eq:proof-mm-Wrec-init} \end{align}\tag{61}\] where \(C_1 \triangleq C_x^2\,C_\eta^2/(C_M - \delta)\). The cross-term bound is the work of Step 2.

9.8.0.2 Step 2. Bounding the cross term.

We expand \(\EE[\eta_{n+1,i}\, x_{n+1,i}^\top \tilde{e}_{n,i}\mid\cF_n]\) using the definition 57 of \(\eta_{n+1,i}\). Since \(\varepsilon_{n+1,i}\) has zero conditional mean and is independent of everything else, that term drops. For each \(j\ne i\), the price \(p_{n+1,j} = \phi_j(\theta_{n,j};\,\cdot) + z_{n+1,j}\) is the sum of an \(\cF_n\)-measurable greedy price and a mean-zero exploration shock \(z_{n+1,j}\) that is independent of \(z_{n+1,i}\), while \(x_{n+1,i}\) depends only on \(z_{n+1,i}\) (and on \(\cF_n\)-measurable quantities). Hence \(z_{n+1,j}\) and \(x_{n+1,i}^\top \tilde{e}_{n,i}\) are conditionally independent for \(j\ne i\), and the conditional expectation factorizes: \[\begin{align} \EE[\eta_{n+1,i}\, x_{n+1,i}^\top \tilde{e}_{n,i}\mid\cF_n] &\;=\; \sum_{j\ne i}\gamma_{i,j}\,\EE[p_{n+1,j} - p_j^{NE}\mid\cF_n]\,\cdot\,\EE[x_{n+1,i}^\top \tilde{e}_{n,i}\mid\cF_n].\label{eq:proof-mm-cross-factor} \end{align}\tag{62}\] The second factor is bounded by Cauchy–Schwarz and the regressor envelope: \[\label{eq:proof-mm-cs} \bigl|\EE[x_{n+1,i}^\top \tilde{e}_{n,i}\mid\cF_n]\bigr| \;\le\; C_x\,\norm{\tilde{e}_{n,i}}_2.\tag{63}\] The first factor depends on whether \(j\) is oblivious or informed. For \(j\in\cI^{ob}\setminus\{i\}\), the greedy price is \(\phi^{ob}(\hat{\theta}_{n,j}^{ob})\), and since \(\phi^{ob}(\theta_j^{*,ob}) = p_j^{NE}\), \[\label{eq:proof-mm-ob-comp} \bigl|\EE[p_{n+1,j} - p_j^{NE}\mid\cF_n]\bigr| \;=\; \bigl|\phi^{ob}(\hat{\theta}_{n,j}^{ob}) - \phi^{ob}(\theta_j^{*,ob})\bigr| \;\le\; L_\phi^{ob}\,\norm{\tilde{e}_{n,j}}_2,\tag{64}\] using that projection onto the convex set \(\Theta_j^{ob}\) is non-expansive. For \(j\in\cI^{in}\), the greedy price is \(\phi_j^{in}(\hat{\theta}_{n,j}^{in};\, \mathbf{m}_{n,-j})\), with \(\phi_j^{in}(\theta_j^{*,in};\, \mathbf{p}_{-j}^{NE}) = p_j^{NE}\) by the Nash first-order condition. Decompose into parameter and forecast variations: \[\begin{align} \bigl|\phi_j^{in}(\hat{\theta}_{n,j}^{in};\, \mathbf{m}_{n,-j}) - \phi_j^{in}(\theta_j^{*,in};\, \mathbf{p}_{-j}^{NE})\bigr| &\;\le\; \bigl|\phi_j^{in}(\hat{\theta}_{n,j}^{in};\, \mathbf{m}_{n,-j}) - \phi_j^{in}(\theta_j^{*,in};\, \mathbf{m}_{n,-j})\bigr|\notag\\ &\quad +\;\bigl|\phi_j^{in}(\theta_j^{*,in};\, \mathbf{m}_{n,-j}) - \phi_j^{in}(\theta_j^{*,in};\, \mathbf{p}_{-j}^{NE})\bigr|.\notag \end{align}\] The first piece is bounded by the projection-box Lipschitz constant \(L_\phi^{in,\theta}\) in \(\theta\). The second piece uses the linearity of \(\phi_j^{in}(\theta_j^{*,in};\cdot)\) in \(\mathbf{m}\) with \(\partial_{m_k}\phi_j^{in} = \gamma_{j,k}/(2|\beta_j|)\). Hence \[\label{eq:proof-mm-in-comp} \bigl|\EE[p_{n+1,j} - p_j^{NE}\mid\cF_n]\bigr| \;\le\; L_\phi^{in,\theta}\,\norm{\tilde{e}_{n,j}^{in}}_2 \;+\; \sum_{k\ne j}\frac{\gamma_{j,k}}{2|\beta_j|}\,|u_{n,k}|.\tag{65}\] Inserting 63 , 64 , and 65 into 62 , \[\begin{align} \bigl|\EE[\eta_{n+1,i} x_{n+1,i}^\top \tilde{e}_{n,i}\mid\cF_n]\bigr| &\;\le\; C_x\,\norm{\tilde{e}_{n,i}}_2\Biggl[\,\sum_{j\in\cI^{ob}\setminus\{i\}} \gamma_{i,j}\, L_\phi^{ob}\, \norm{\tilde{e}_{n,j}}_2\notag\\ &\qquad +\; \sum_{j\in\cI^{in}}\gamma_{i,j}\,\Bigl(L_\phi^{in,\theta}\,\norm{\tilde{e}_{n,j}^{in}}_2 + \sum_{k\ne j}\frac{\gamma_{j,k}}{2|\beta_j|}\,|u_{n,k}|\Bigr)\,\Biggr]. \label{eq:proof-mm-cross-unyoung} \end{align}\tag{66}\] We next apply Young’s inequality to each product on the right-hand side of 66 , separately on the ob–ob terms, the ob–in parameter terms, and the ob–in running-mean terms. We collect each family in turn.

ob–ob. Young’s gives \[\begin{align} \sum_{j\in\cI^{ob}\setminus\{i\}} \gamma_{i,j}\, L_\phi^{ob}\,\norm{\tilde{e}_{n,i}}_2\,\norm{\tilde{e}_{n,j}}_2 &\;\le\; \frac{L_\phi^{ob}}{2}\sum_{j\in\cI^{ob}\setminus\{i\}}\gamma_{i,j}\,\bigl(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}}_2^2\bigr). \end{align}\] Summing over \(i\in\cI^{ob}\) and re-indexing the double sum into row and column sums of \(\gamma\) on \(\cI^{ob}\times\cI^{ob}\), \[\sum_{i\in\cI^{ob}}\sum_{j\in\cI^{ob}\setminus\{i\}}\gamma_{i,j}\,\bigl(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}}_2^2\bigr) \;\le\; 2\bar\gamma^{ob}\,\sum_{i\in\cI^{ob}}\norm{\tilde{e}_{n,i}}_2^2\] by the definition of \(\bar\gamma^{ob}\), analogously to the step for Theorem 2.

ob–in (parameter). Young’s gives \[\sum_{j\in\cI^{in}}\gamma_{i,j}\, L_\phi^{in,\theta}\,\norm{\tilde{e}_{n,i}}_2\,\norm{\tilde{e}_{n,j}^{in}}_2 \;\le\; \frac{L_\phi^{in,\theta}}{2}\sum_{j\in\cI^{in}}\gamma_{i,j}\,\bigl(\norm{\tilde{e}_{n,i}}_2^2 + \norm{\tilde{e}_{n,j}^{in}}_2^2\bigr).\] Summing over \(i\in\cI^{ob}\) yields a row-sum coefficient \(\bar\Theta = L_\phi^{in,\theta}\,\max_{i\in\cI^{ob}}\sum_{j\in\cI^{in}}\gamma_{i,j}\) on \(\sum_i\norm{\tilde{e}_{n,i}}_2^2\) and a column-sum coefficient \[\bar\Theta^{\mathrm{col}} \;\triangleq\; L_\phi^{in,\theta}\,\max_{j\in\cI^{in}}\sum_{i\in\cI^{ob}}\gamma_{i,j}\] on \(\sum_j\norm{\tilde{e}_{n,j}^{in}}_2^2\).

ob–in (running mean). Young’s gives \[\begin{align} \sum_{j\in\cI^{in}}\gamma_{i,j}\sum_{k\ne j}\frac{\gamma_{j,k}}{2|\beta_j|}\,\norm{\tilde{e}_{n,i}}_2\,|u_{n,k}| &\;\le\; \frac{1}{2}\sum_{j\in\cI^{in}}\sum_{k\ne j}\gamma_{i,j}\,\frac{\gamma_{j,k}}{2|\beta_j|}\,\bigl(\norm{\tilde{e}_{n,i}}_2^2 + U_{n,k}\bigr). \end{align}\] Summing over \(i\in\cI^{ob}\), the diagonal piece gives coefficient \(\bar\Delta\) on \(\sum_i\norm{\tilde{e}_{n,i}}_2^2\) by the definition of \(\bar\Delta\), while re-indexing the off-diagonal piece over \(k\) first, \[\sum_{i\in\cI^{ob}}\sum_{j\in\cI^{in}}\sum_{k\ne j} \gamma_{i,j}\,\frac{\gamma_{j,k}}{2|\beta_j|}\,U_{n,k} \;=\; \sum_{k\in[N]}\,U_{n,k}\sum_{i\in\cI^{ob}}\sum_{j\in\cI^{in}\setminus\{k\}}\frac{\gamma_{i,j}\,\gamma_{j,k}}{2|\beta_j|} \;\le\; \bar\Psi\,U_n^{tot}\] by the definition of \(\bar\Psi\).

Aggregation. Plugging the three families back into 61 and summing over \(i\in\cI^{ob}\), \[\begin{align} \sum_{i\in\cI^{ob}} \EE[W_{n+1,i}\mid\cF_n] &\;\le\; \frac{n}{n+1}\,W_n^{ob}\notag\\ &\quad -\;\frac{1}{n+1}\Bigl[\frac{C_M}{1+\delta_h} - C_x\,\bigl(2\bar\gamma^{ob}\,L_\phi^{ob} + \bar\Delta + \bar\Theta\bigr)\Bigr]\sum_{i\in\cI^{ob}}\norm{\tilde{e}_{n,i}}_2^2\notag\\ &\quad +\;\frac{C_x\,\bar\Psi}{n+1}\,U_n^{tot} \;+\; \frac{C_x\,\bar\Theta^{\mathrm{col}}}{n+1}\sum_{j\in\cI^{in}}\norm{\tilde{e}_{n,j}^{in}}_2^2 \;+\; \frac{C_1\,|\cI^{ob}|}{n(n+1)}.\label{eq:proof-mm-Wrec-agg} \end{align}\tag{67}\]

9.8.0.3 Step 3. Running-mean recursion.

For every \(k\in[N]\), expand \(m_{n+1,k} = \tfrac{n}{n+1}\,m_{n,k} + \tfrac{1}{n+1}\,p_{n+1,k}\), so subtracting \(p_k^{NE}\), \[u_{n+1,k} \;=\; \frac{n}{n+1}\,u_{n,k} + \frac{1}{n+1}\,(p_{n+1,k} - p_k^{NE}).\] Squaring and taking conditional expectation, \[\begin{align} \EE[U_{n+1,k}\mid\cF_n] &\;=\; \Bigl(\frac{n}{n+1}\Bigr)^{\!2}\,U_{n,k} \;+\; \frac{2n}{(n+1)^2}\,u_{n,k}\,\EE[p_{n+1,k} - p_k^{NE}\mid\cF_n]\notag\\ &\quad +\; \frac{1}{(n+1)^2}\,\EE[(p_{n+1,k} - p_k^{NE})^2\mid\cF_n].\label{eq:proof-mm-Urec-raw} \end{align}\tag{68}\] Using \((\tfrac{n}{n+1})^2 \le 1 - \tfrac{2}{n+1} + \tfrac{1}{(n+1)^2}\) and bounding the last term by a price envelope \(C_p^2 \triangleq \sup_{n,k}\EE[(p_{n+1,k} - p_k^{NE})^2\mid\cF_n] < \infty\) (bounded prices), 68 simplifies to \[\begin{align} \EE[U_{n+1,k}\mid\cF_n] &\;\le\; U_{n,k} \;-\; \frac{2}{n+1}\,U_{n,k}\notag\\ &\quad +\;\frac{2}{n+1}\,|u_{n,k}|\,\bigl|\EE[p_{n+1,k} - p_k^{NE}\mid\cF_n]\bigr| \;+\;\frac{U_{n,k} + C_p^2}{(n+1)^2}.\label{eq:proof-mm-Urec-mid} \end{align}\tag{69}\] The cross-term bound now distinguishes \(k\in\cI^{ob}\) and \(k\in\cI^{in}\). For \(k\in\cI^{ob}\), 64 gives \(|\EE[p_{n+1,k}-p_k^{NE}\mid\cF_n]|\le L_\phi^{ob}\,\norm{\tilde{e}_{n,k}}_2\), and Young’s inequality \[|u_{n,k}|\cdot L_\phi^{ob}\,\norm{\tilde{e}_{n,k}}_2 \;\le\; \frac{L_\phi^{ob}}{2}\,\bigl(U_{n,k} + \norm{\tilde{e}_{n,k}}_2^2\bigr)\] yields the per-row bound \[\label{eq:proof-mm-Urec-ob} \EE[U_{n+1,k}\mid\cF_n] \;\le\; U_{n,k} \;-\; \frac{2 - L_\phi^{ob}}{n+1}\,U_{n,k} \;+\; \frac{L_\phi^{ob}}{n+1}\,\norm{\tilde{e}_{n,k}}_2^2 \;+\; \frac{U_{n,k} + C_p^2}{(n+1)^2}.\tag{70}\] For \(k\in\cI^{in}\), 65 gives the two-piece bound. Young’s on the parameter piece gives \[|u_{n,k}|\cdot L_\phi^{in,\theta}\,\norm{\tilde{e}_{n,k}^{in}}_2 \;\le\; \frac{L_\phi^{in,\theta}}{2}\,\bigl(U_{n,k} + \norm{\tilde{e}_{n,k}^{in}}_2^2\bigr),\] and on each running-mean piece, \[|u_{n,k}|\cdot \frac{\gamma_{k,\ell}}{2|\beta_k|}\,|u_{n,\ell}| \;\le\; \frac{\gamma_{k,\ell}}{2|\beta_k|}\,\frac{U_{n,k} + U_{n,\ell}}{2} \qquad (\ell\ne k).\] Summing the latter over \(\ell\ne k\) and combining, \[|u_{n,k}|\cdot\bigl|\EE[p_{n+1,k}-p_k^{NE}\mid\cF_n]\bigr| \;\le\; \frac{1}{2}\Bigl[L_\phi^{in,\theta} + \frac{\gamma_k}{2|\beta_k|}\Bigr]U_{n,k} + \frac{L_\phi^{in,\theta}}{2}\,\norm{\tilde{e}_{n,k}^{in}}_2^2 + \frac{1}{2}\sum_{\ell\ne k}\frac{\gamma_{k,\ell}}{2|\beta_k|}\,U_{n,\ell}.\] Plugging into 69 gives, for \(k\in\cI^{in}\), \[\begin{align} \EE[U_{n+1,k}\mid\cF_n] &\;\le\; U_{n,k} \;-\; \frac{2 - L_\phi^{in,\theta} - \gamma_k/(2|\beta_k|)}{n+1}\,U_{n,k}\notag\\ &\quad +\;\frac{1}{n+1}\sum_{\ell\ne k}\frac{\gamma_{k,\ell}}{2|\beta_k|}\,U_{n,\ell} \;+\; \frac{L_\phi^{in,\theta}}{n+1}\,\norm{\tilde{e}_{n,k}^{in}}_2^2 \;+\; \frac{U_{n,k}+C_p^2}{(n+1)^2}.\label{eq:proof-mm-Urec-in} \end{align}\tag{71}\] Summing the per-row bounds 7071 over \(k\in[N]\), applying \[\min\,\left\{2 - L_\phi^{ob},\;\, 2 - L_\phi^{in,\theta} - \max_{j\in\cI^{in}}\gamma_j/(2|\beta_j|)\right\} \;=\; \bar c_{\mathrm{diag}}\] as a uniform lower bound on the diagonal coefficient, and re-indexing the informed–informed cross term \[\sum_{k\in\cI^{in}}\sum_{\ell\ne k}\frac{\gamma_{k,\ell}}{2|\beta_k|}\,U_{n,\ell} \;=\; \sum_{\ell\in[N]}\,U_{n,\ell}\sum_{k\in\cI^{in}\setminus\{\ell\}}\frac{\gamma_{k,\ell}}{2|\beta_k|} \;\le\; \bar D\,U_n^{tot}\] by the definition of \(\bar D\), we arrive at \[\begin{align} \sum_{k=1}^N \EE[U_{n+1,k}\mid\cF_n] &\;\le\; U_n^{tot} \;-\; \frac{\bar c_{\mathrm{diag}} - \bar D}{n+1}\,U_n^{tot} \;+\; \frac{L_\phi^{ob}}{n+1}\sum_{k\in\cI^{ob}}\norm{\tilde{e}_{n,k}}_2^2\notag\\ &\quad +\;\frac{L_\phi^{in,\theta}}{n+1}\sum_{k\in\cI^{in}}\norm{\tilde{e}_{n,k}^{in}}_2^2 \;+\; \frac{N\,C_2}{(n+1)^2},\label{eq:proof-mm-Urec-agg} \end{align}\tag{72}\] where \(C_2 \triangleq \sup_{n,k} U_{n,k} + C_p^2 < \infty\). Condition (ii) guarantees \(\bar c_{\mathrm{diag}} - \bar D > 0\), so the diagonal contraction is non-trivial.

9.8.0.4 Step 4. Combined Lyapunov contraction and the per-\(\lambda\) rate.

For \(\lambda > 0\) to be chosen, define the joint Lyapunov function \[\cL_n \;\triangleq\; W_n^{ob} \;+\; \lambda\,U_n^{tot}.\] Multiplying 72 by \(\lambda\), adding to 67 , and using the identity \(\tfrac{n}{n+1}\,W_n^{ob} = W_n^{ob} - \tfrac{1}{n+1}\,W_n^{ob}\), \[\label{eq:proof-mm-Lstep4-bracket} \EE[\cL_{n+1}\mid\cF_n] \;\le\; \cL_n \;-\; \frac{1}{n+1}\,\Bigl[\,W_n^{ob} + A(\lambda)\sum_{i\in\cI^{ob}}\norm{\tilde{e}_{n,i}}_2^2 + B(\lambda)\,U_n^{tot}\,\Bigr] \;+\; R_n,\tag{73}\] where \[\begin{align} A(\lambda) &\;\triangleq\; \frac{C_M}{1+\delta_h} - C_x\,\bigl(2\bar\gamma^{ob}\,L_\phi^{ob} + \bar\Delta + \bar\Theta\bigr) - \lambda\,L_\phi^{ob},\tag{74}\\ B(\lambda) &\;\triangleq\; \lambda\,(\bar c_{\mathrm{diag}} - \bar D) \;-\; C_x\,\bar\Psi,\tag{75} \end{align}\] and the residual is \[\label{eq:proof-mm-R} R_n \;\triangleq\; \frac{C_1\,|\cI^{ob}| + \lambda\,N\,C_2}{n(n+1)} \;+\; \frac{C_x\,\bar\Theta^{\mathrm{col}} + \lambda\,L_\phi^{in,\theta}}{n+1}\sum_{j\in\cI^{in}}\norm{\tilde{e}_{n,j}^{in}}_2^2.\tag{76}\] Both \(A(\lambda) > 0\) and \(B(\lambda) > 0\) hold iff \(\lambda\) lies in the open interval \[\label{eq:proof-mm-lambda-range} \frac{C_x\,\bar\Psi}{\bar c_{\mathrm{diag}} - \bar D} \;<\; \lambda \;<\; \frac{C_M/(1+\delta_h) - C_x\,(2\bar\gamma^{ob}\,L_\phi^{ob} + \bar\Delta + \bar\Theta)}{L_\phi^{ob}}.\tag{77}\] A feasible \(\lambda\) exists iff the upper bound exceeds the lower bound, which after letting \(\delta_h \downarrow 0\) is exactly condition (iv).

Fix any such \(\lambda\). By Lemma 10, the spectral upper bound \(W_n^{ob} \le C_x^2\sum_{i\in\cI^{ob}}\norm{\tilde{e}_{n,i}}_2^2\) holds, so \(\sum_i\norm{\tilde{e}_{n,i}}_2^2 \ge W_n^{ob}/C_x^2\). Substituting into the bracket of 73 , \[\begin{align} W_n^{ob} + A(\lambda)\sum_i\norm{\tilde{e}_{n,i}}_2^2 + B(\lambda)\,U_n^{tot} &\;\ge\; \Bigl(1 + \frac{A(\lambda)}{C_x^2}\Bigr)\,W_n^{ob} \;+\; \frac{B(\lambda)}{\lambda}\,\lambda\,U_n^{tot}\\ &\;\ge\; c^*(\lambda)\,\cL_n, \end{align}\] where \[\label{eq:proof-mm-cstar-lambda} c^*(\lambda) \;\triangleq\; \min\,\left\{1 + \frac{A(\lambda)}{C_x^2},\;\, \frac{B(\lambda)}{\lambda}\right\} \;=\; \min\,\left\{1 + \frac{K_1 - \lambda\,L_\phi^{ob}}{C_x^2},\;\, K_2 - \frac{C_x\,\bar\Psi}{\lambda}\right\} \;>\; 0,\tag{78}\] using \(K_1\) from 56 , after letting \(\delta_h \downarrow 0\) in \(A(\lambda)\). Substituting back into 73 , \[\label{eq:proof-mm-Lcontract} \EE[\cL_{n+1}\mid\cF_n] \;\le\; \cL_n \;-\; \frac{c^*(\lambda)}{n+1}\,\cL_n \;+\; R_n.\tag{79}\]

9.8.0.5 Step 5. Almost-sure convergence.

By Lemma 6(c), \(\sum_n \tfrac{1}{n+1}\sum_j\norm{\tilde{e}_{n,j}^{in}}_2^2 < \infty\) a.s., which together with the \(O(1/n^2)\) term in 76 gives \(\sum_n R_n < \infty\) a.s. Applying the Robbins–Siegmund theorem [57] to 79 yields \[\cL_n \to 0 \quad\text{a.s.}, \qquad \sum_{n}\frac{c^*(\lambda)}{n+1}\,\cL_n < \infty \quad\text{a.s.,}\] so \(W_n^{ob}\to 0\) and \(U_n^{tot}\to 0\) a.s. Lemma 10 then gives \(\norm{\tilde{e}_{n,i}}_2\to 0\) a.s.for every \(i\in\cI^{ob}\), hence \(\hat{\theta}_{n,i}^{ob}\to\theta_i^{*,ob}\) a.s., and continuity of \(\phi^{ob}\) gives \(\tilde{p}_{n,i}\to p_i^{NE}\) a.s. Lemma 6(b) gives \(\hat{\theta}_{n,j}^{in}\to\theta_j^{*,in}\) a.s.for \(j\in\cI^{in}\). Finally, \(U_n^{tot}\to 0\) gives \(m_{n,k}\to p_k^{NE}\) a.s.for every \(k\), so by joint continuity of \(\phi_j^{in}(\cdot;\cdot)\), \[\tilde{p}_{n,j} \;=\; \phi_j^{in}(\hat{\theta}_{n,j}^{in};\, \mathbf{m}_{n,-j}) \;\longrightarrow\; \phi_j^{in}(\theta_j^{*,in};\, \mathbf{p}_{-j}^{NE}) \;=\; p_j^{NE} \quad\text{a.s.}\] for every \(j\in\cI^{in}\).

9.8.0.6 Step 6. Polynomial rate in expectation.

Take unconditional expectation in 79 : \[\EE[\cL_{n+1}] \;\le\; \Bigl(1 - \frac{c^*(\lambda)}{n+1}\Bigr)\,\EE[\cL_n] \;+\; \EE[R_n].\] By Lemma 6(a), \(\sum_j\EE\norm{\tilde{e}_{n,j}^{in}}_2^2 = O(n^{\eta_{\max}-1})\), so \(\EE[R_n] = O(n^{\eta_{\max}-2})\) (the \(O(1/n^2)\) piece is absorbed since \(\eta_{\max}<1/2<1\)). A standard non-homogeneous linear recursion lemma—given \(a_n \le (1 - c/(n+1))\,a_{n-1} + O(n^{-r})\) for constants \(c > 0\) and \(r > 1\), the solution obeys \(a_n = O(n^{-\min\,\left\{c,\,r-1\right\}})\), with an extra \(\log n\) factor at the boundary \(c = r-1\)—applied with \(c = c^*(\lambda)\) and \(r-1 = 1 - \eta_{\max}\) yields \[\EE[\cL_n] \;=\; \begin{cases} O\bigl(n^{\eta_{\max} - 1}\bigr), & c^*(\lambda) > 1 - \eta_{\max},\\[3pt] O\bigl(n^{\eta_{\max} - 1}\,\log n\bigr), & c^*(\lambda) = 1 - \eta_{\max},\\[3pt] O\bigl(n^{-c^*(\lambda)}\bigr), & c^*(\lambda) < 1 - \eta_{\max}. \end{cases}\] Optimizing over \(\lambda\) in the feasible interval 77 replaces \(c^*(\lambda)\) by \(c^* \triangleq \sup_\lambda c^*(\lambda)\), and the closed form for \(c^*\) is given in Lemma 7 (Appendix 9.12). The trichotomy stated in the theorem follows, with the greedy-price MSE rate inheriting both bottlenecks since \[\tilde{p}_{n,j} \;=\; \phi_j^{in}(\hat{\theta}_{n,j}^{in};\, \mathbf{m}_{n,-j})\] inherits the \(\cL_n\) rate through \(\mathbf{m}\) and the informed-OLS rate through \(\hat{\theta}\), with overall \(\min\) being \(\min\,\left\{c^*,\, 1-\eta_{\max}\right\}\).

9.8.0.7 Step 7. Realized-price MSE for informed sellers.

Fix \(j\in\cI^{in}\). Decompose \(p_{n,j} - p_j^{NE} = (\tilde{p}_{n,j} - p_j^{NE}) + z_{n,j}\) with \(\EE[z_{n,j}\mid\cF_{n-1}]=0\), \(\Var(z_{n,j}\mid\cF_{n-1})=\nu_{n,j}^2\), and \(z_{n,j}\) independent of \(\tilde{p}_{n,j}\in\cF_{n-1}\). The cross term vanishes in expectation, giving \[\label{eq:proof-mm-realized-decomp} \EE\!\left[(p_{n,j} - p_j^{NE})^2\right] \;=\; \EE\!\left[(\tilde{p}_{n,j} - p_j^{NE})^2\right] \;+\; \nu_{n,j}^2.\tag{80}\] For the greedy-price piece, the Lipschitz decomposition that produced 65 coordinate-wise gives \[(\tilde{p}_{n,j} - p_j^{NE})^2 \;\le\; 2\,(L_\phi^{in,\theta})^2\,\norm{\tilde{e}_{n,j}^{in}}_2^2 \;+\; 2\,\Bigl(\sum_{k\ne j}\frac{\gamma_{j,k}}{2|\beta_j|}\,|u_{n,k}|\Bigr)^{\!2} \;\le\; C_3\,\bigl(\norm{\tilde{e}_{n,j}^{in}}_2^2 + U_n^{tot}\bigr)\] for a constant \(C_3\) depending only on \(L_\phi^{in,\theta}\) and the projection-box ratios \(\gamma_{j,k}/(2|\beta_j|)\). Taking expectation and applying Lemma 6(a) and Step 6, \[\EE\!\left[(\tilde{p}_{n,j} - p_j^{NE})^2\right] \;=\; O\!\left(n^{\eta_{\max}-1}\right) \;+\; O\!\left(\EE[\cL_n]\right) \;=\; O\!\left(n^{-\min\,\left\{c^*,\, 1-\eta_{\max}\right\}}\right),\] absorbing the \(\log n\) factor at the boundary since the asserted statement is up to absolute constants. The exploration tail is \(\nu_{n,j}^2 = \Theta(n^{-\eta_j})\) by the standing assumption in the opener of Section 6. Combining with 80 , \[\EE\!\left[(p_{n,j} - p_j^{NE})^2\right] \;=\; O\!\left(n^{-\min\,\left\{c^*,\, 1-\eta_{\max}\right\}}\right) \;+\; O\!\left(n^{-\eta_j}\right).\] By condition (i), \(\eta_{\max} < 1/2\), so \(1 - \eta_{\max} > 1/2 > \eta_{\max} \ge \eta_j\), hence \(\eta_j \le 1-\eta_{\max}\) and \(\min\,\left\{1-\eta_{\max},\, \eta_j\right\} = \eta_j\). The two terms therefore combine to \(O(n^{-\min\,\left\{c^*,\, \eta_j\right\}})\), completing the proof of Theorem 5. \(\square\)

9.9 Proof of Proposition [prop:strict-dom]↩︎

By Corollary 2, under each strategy profile \((s_1, s_2)\) the realized surplus-capture at horizon \(T\) converges a.s.: \[\lim_{T\to\infty} S_{T,i} \;\stackrel{\text{a.s.}}{=}\; \begin{cases} -\,|\beta_i|\,\nu_i^2/(\Pi_i^{C}-\Pi_i^{NE}) & \text{if } s_i = \textsf{oblivious},\\[2pt] 0 & \text{if } s_i = \textsf{informed}, \end{cases}\] where the limit depends only on \(s_i\) (not on \(s_{-i}\)). Since \(\lim_T S_{T,i}\) exists a.s., \(S_i = \liminf_T S_{T,i}\) in Equation 5 equals the same limit a.s. Therefore, a.s., \[S_i(\textsf{informed},\, s_{-i}) \;=\; 0 \;>\; -\,|\beta_i|\,\nu_i^2/(\Pi_i^{C}-\Pi_i^{NE}) \;=\; S_i(\textsf{oblivious},\, s_{-i}) \quad \text{for every } s_{-i},\] which is strict dominance of informed over oblivious for each seller. The unique strict pure-strategy Nash equilibrium of the strategy game is therefore \((\textsf{informed}, \textsf{informed})\). \(\square\)

9.10 Lemmas for the proof of Proposition [prop:strict-dom]↩︎

Write \(\epsilon_{n,i}\triangleq\tilde{p}_{n,i} - p_i^{NE}\) for the greedy-price drift error of seller \(i\).

Lemma 4 (Asymptotic per-period revenue). Adopt condition 10 . Suppose \(\tilde{p}_{n,i}\to p_i^{NE}\) a.s.for every \(i\in[N]\). Then, the realized* per-period revenue time-average \[\bar R_{T,i} \;\triangleq\; \frac{1}{T}\sum_{n=1}^T R_{n,i}\] satisfies \[\label{eq:asymp-revenue} \bar R_{T,i} \;=\; \Pi_i^{NE} \;-\; |\beta_i|\,\bar\nu_{T,i}^2 \;+\; o(1) \qquad \text{a.s.\;as T\to\infty,}\tag{81}\] where \(\bar\nu_{T,i}^2 \triangleq T^{-1}\sum_{n=1}^T \nu_{n,i}^2\).*

Proof of Lemma 4. Since \(R_i(\mathbf{p})\) is a quadratic polynomial in \(\mathbf{p}\), Taylor’s expansion around \(\mathbf{p}^{NE}\) is exact: \[\begin{align} R_i(\mathbf{p}) - \Pi_i^{NE} &= \nabla_{\mathbf{p}} R_i(\mathbf{p}^{NE})\cdot(\mathbf{p} - \mathbf{p}^{NE}) + \tfrac12 (\mathbf{p} - \mathbf{p}^{NE})^\top \nabla^2_{\mathbf{p}} R_i\,(\mathbf{p} - \mathbf{p}^{NE}) \\ &= p_i^{NE}\sum_{j\ne i} \gamma_{i,j}(p_j - p_j^{NE}) + \beta_i (p_i - p_i^{NE})^2 + \sum_{j\ne i} \gamma_{i,j} (p_i - p_i^{NE})(p_j - p_j^{NE}), \end{align}\] using \(\partial_{p_i} R_i(\mathbf{p}^{NE}) = \alpha_i + 2\beta_i p_i^{NE} + \sum_{j\ne i}\gamma_{i,j} p_j^{NE} = 0\) (Nash FOC), \(\partial_{p_j} R_i(\mathbf{p}^{NE}) = \gamma_{i,j} p_i^{NE}\) for \(j\ne i\), \(\partial^2_{p_i p_i} R_i = 2\beta_i\), and \(\partial^2_{p_i p_j} R_i = \gamma_{i,j}\) for \(j\ne i\). Under 10 together with \(\tilde{p}_{n,i}\to p_i^{NE}\) a.s., the event \(\{|\tilde{p}_{n,i} - p_i^{NE}| < \min\,\left\{p_i^{NE} - l,\, u - p_i^{NE}\right\} - \delta_i\}\) holds eventually a.s., so clipping is asymptotically inactive and \(p_{n,i} - p_i^{NE} = \epsilon_{n,i} + z_{n,i}\) exactly for all large \(n\), where \(\epsilon_{n,i}\triangleq\tilde{p}_{n,i} - p_i^{NE}\in\cF_{n-1}\). Conditioning on \(\cF_{n-1}\) and using \(\EE[z_{n,j}\mid\cF_{n-1}] = 0\), \(\EE[z_{n,j}^2\mid\cF_{n-1}] = \nu_{n,j}^2\), and \(\EE[z_{n,i} z_{n,j}\mid\cF_{n-1}] = 0\) for \(i\ne j\), \[\label{eq:proof-asymp-revenue-cond} \EE[R_{n,i} - \Pi_i^{NE} \mid \cF_{n-1}] \;=\; p_i^{NE}\sum_{j\ne i}\gamma_{i,j}\,\epsilon_{n,j} \;+\; \beta_i\bigl(\epsilon_{n,i}^2 + \nu_{n,i}^2\bigr) \;+\; \sum_{j\ne i} \gamma_{i,j}\,\epsilon_{n,i}\,\epsilon_{n,j}.\tag{82}\] The martingale residual \(R_{n,i} - \EE[R_{n,i}\mid\cF_{n-1}]\) has uniformly bounded conditional second moments (since prices and demand noise are bounded), so by the strong law of large numbers for bounded martingale differences, \[\label{eq:asymp-rev-mart} \frac{1}{T}\sum_{n=1}^T \bigl(R_{n,i} - \EE[R_{n,i}\mid\cF_{n-1}]\bigr) \;\xrightarrow[T\to\infty]{\text{a.s.}}\; 0.\tag{83}\] Combining 82 and 83 , \[\bar R_{T,i} \;-\; \Pi_i^{NE} \;=\; -\,|\beta_i|\,\bar\nu_{T,i}^2 \;+\; R^{(\epsilon)}_{T,i} \;+\; o(1) \quad\text{a.s.,}\] where \[R^{(\epsilon)}_{T,i} \;\triangleq\; p_i^{NE}\!\sum_{j\ne i}\gamma_{i,j}\,\frac{1}{T}\sum_{n=1}^T\epsilon_{n,j} \;+\; \beta_i\,\frac{1}{T}\sum_{n=1}^T\epsilon_{n,i}^2 \;+\; \sum_{j\ne i} \gamma_{i,j}\,\frac{1}{T}\sum_{n=1}^T\epsilon_{n,i}\,\epsilon_{n,j}.\] By the a.s.greedy-price convergence hypothesis, \(\epsilon_{n,k}\to 0\) a.s.for every \(k\); the Cesàro mean of an a.s.-convergent sequence converges a.s.to the same limit, so each of \(T^{-1}\sum_n \epsilon_{n,j}\), \(T^{-1}\sum_n \epsilon_{n,i}^2\), and \(T^{-1}\sum_n \epsilon_{n,i}\epsilon_{n,j}\) tends to \(0\) a.s. Hence \(R^{(\epsilon)}_{T,i}\to 0\) a.s., and 81 follows. ◻

Lemma 4 sharpens Proposition [prop:costlinearexploration] from a per-period worst-case bound of order \(|\beta_i|\delta^2\) to an exact asymptotic value on any learning protocol whose greedy price converges to Nash a.s.: the asymptotic per-period revenue equals the Nash benchmark minus the Cesàro average of the seller’s own exploration variance, with equality in the leading order. The hypothesis \(\tilde{p}_{n,i}\to p_i^{NE}\) a.s.is exactly what each of our three composition convergence theorems delivers (Theorems 2, 4, and 5), so the lemma applies on the corresponding learning protocols.

Corollary 2 (Asymptotic surplus-capture in the three compositions). Adopt the standing interior assumption \(\mathbf{p}^{NE}, \mathbf{p}^{C}\in(l,u)^N\) and the bounded-exploration condition 10 . Define a finite version of the surplus-capture ratio 5 at horizon \(T\) as \(S_{T,i}\triangleq (\bar R_{T,i} - \Pi_i^{NE})/(\Pi_i^{C} - \Pi_i^{NE})\) with \(\bar R_{T,i}\) from Lemma 4.

  • Under the conditions of Theorem 2 (obob) with constant oblivious exploration \(\nu_{n,i}^2\equiv\nu_i^2\), \[\lim_{T\to\infty} S_{T,i} \;\stackrel{\text{a.s.}}{=}\; -\,\frac{|\beta_i|\,\nu_i^2}{\Pi_i^{C} - \Pi_i^{NE}} \;<\; 0, \qquad i\in[N].\]

  • Under the conditions of Theorem 4 (inin) with decaying informed exploration \(\nu_{n,i}^2\to 0\), \[\lim_{T\to\infty} S_{T,i} \;\stackrel{\text{a.s.}}{=}\; 0, \qquad i\in[N].\]

  • Under the conditions of Theorem 5 (obin) with constant oblivious exploration \(\nu_{n,i}^2\equiv\nu_i^2 > 0\) for \(i\in\cI^{ob}\) and decaying informed exploration \(\nu_{n,j}^2\to 0\) for \(j\in\cI^{in}\), \[\lim_{T\to\infty} S_{T,i} \;\stackrel{\text{a.s.}}{=}\; -\,\frac{|\beta_i|\,\nu_i^2}{\Pi_i^{C} - \Pi_i^{NE}} \;<\; 0 \quad (i\in\cI^{ob}), \qquad \lim_{T\to\infty} S_{T,j} \;\stackrel{\text{a.s.}}{=}\; 0 \quad (j\in\cI^{in}).\]

Proof of Corollary 2. Each of Theorems 2, 4, and 5 delivers \(\tilde{p}_{n,k}\to p_k^{NE}\) a.s.for every \(k\), so the hypothesis of Lemma 4 holds on every cell of the strategy game. Substituting the respective exploration schedule into \(\bar\nu_{T,i}^2 = T^{-1}\sum_n \nu_{n,i}^2\):

(a) Under Theorem 2, \(\nu_{n,i}^2\equiv\nu_i^2\) for every \(i\in[N]\), so \(\bar\nu_{T,i}^2 = \nu_i^2\) exactly. Lemma 4 gives \(\bar R_{T,i}\to \Pi_i^{NE} - |\beta_i|\nu_i^2\) a.s., and dividing by \(\Pi_i^{C} - \Pi_i^{NE} > 0\) yields \(\lim_T S_{T,i} = -|\beta_i|\nu_i^2/(\Pi_i^{C} - \Pi_i^{NE}) < 0\) a.s.

(b) Under Theorem 4, \(\nu_{n,i}^2\to 0\) for every \(i\), so \(\bar\nu_{T,i}^2\to 0\) by the Cesàro-vanishing identity. Lemma 4 gives \(\bar R_{T,i}\to \Pi_i^{NE}\) a.s., hence \(\lim_T S_{T,i} = 0\) a.s.

(c) Under Theorem 5, \(\nu_{n,i}^2\equiv\nu_i^2 > 0\) for \(i\in\cI^{ob}\) and \(\nu_{n,j}^2\to 0\) for \(j\in\cI^{in}\). Applying (a) on \(\cI^{ob}\) and (b) on \(\cI^{in}\) gives the stated formulas. ◻

9.11 Proof of Theorem 6↩︎

We first prove an auxiliary lemma.

Lemma 5 (Growth condition of the design matrix). Fix an integer \(d \ge 1\). Let \((\mathcal{F}_n)_{n \ge 0}\) be a filtration. For each \(n \ge 1\), define the random vector \[x_n = (1, p_{n,1}, ..., p_{n,d})^\top \in \mathbb{R}^{d+1}, \quad \mathcal{J}_n = \sum_{m=1}^n x_m x_m^\top.\] Assume there exist constants \(0 < l < u < \infty\) such that \(p_{n,i} \in [l,u]\) a.s. for all \(n\) and \(i\), and that \[p_{n,i} = \chi_{n,i} + z_{n,i}, \quad i = 1, ..., d,\] where \(\chi_{n,i} \in \mathcal{F}_{n-1}\) are uniformly bounded, and \(z_{n,i}\) are independent across all \(n\) and \(i\), mean-zero, and uniformly bounded. Let \[\nu_{n,i}^2 \triangleq \mathrm{Var}(z_{n,i}) = c_i n^{-\eta_i}, \quad c_i > 0, \; \eta_i \in [0,1).\] Define \[\eta_{\min} = \min_{1 \le i \le d} \eta_i, \quad \eta_{\max} = \max_{1 \le i \le d} \eta_i.\] Then, there exist constants \(C_1, C_2 > 0\) and an \(n_0\) such that for all \(n \ge n_0\), \[\mathbb{P} \left(\lambda_{\min} (\mathcal{J}_n) \le C_1 n^{1 - \eta_{\max}}\right) \le (d+1) \exp \left(-C_2 n^{1 + \eta_{\min} - 2 \eta_{\max}}\right).\] If we assume that \(1 + \eta_{\min} > 2 \eta_{\max}\), then the term on the right-hand side decreases to zero exponentially fast.

Proof of Lemma 5. To set up, let \(\mathbb{E}_{n-1}\) denote the conditional expectation given \(\mathcal{F}_{n-1}\). Denote \[M_n = \sum_{m=1}^n \mathbb{E}_{m-1}[x_m x_m^\top], \quad Y_m = x_m x_m^\top - \mathbb{E}_{m-1}[x_m x_m^\top].\] Then, \((Y_m, \mathcal{F}_n)\) is a self-adjoint matrix martingale difference sequence and \[\mathcal{J}_n = M_n + \sum_{m=1}^n Y_m.\] For any \(t > 0\), \[\label{eq:designmatrixgrowthconditiondecomposition} \left\{\lambda_{\min}(\mathcal{J}_n) \le \lambda_{\min}(M_n) - t \right\} \subseteq \left\{\lambda_{\max}(M_n - \mathcal{J}_n) \ge t\right\} = \left\{\lambda_{\max}\left(-\sum_{m=1}^n Y_m\right) \ge t\right\}.\tag{84}\] So, it suffices to upper bound \(\mathbb{P}\left(\lambda_{\max}\left(-\sum_{m=1}^n Y_m\right) \ge t\right)\). We will apply the Matrix Freedman inequality (Theorem 7) to the matrix martingale \(\left(-\sum_{m=1}^n Y_m, \mathcal{F}_n\right)\). For that, we need:

  • a uniform bound \(R\) on \(\norm{Y}_{op}\);

  • a lower bound on \(\lambda_{\min}(M_n)\); and

  • an upper bound \(\sigma_n^2\) on the predictable quadratic variation process \(\norm{\sum_{m=1}^n \mathbb{E}_{m-1}[Y_m^2]}_{op}\).

The first item is straightforward. Since \(p_{n,i}\) are uniformly bounded, \[\norm{x_m}_2^2 = 1 + \sum_{i=1}^d p_{m,i}^2 \le 1 + d u^2 \eqqcolon C_x^2,\] and so \(\norm{x_m x_m^\top}_{op} = \norm{x_m}_2^2 \le C_x^2\). By Jensen’s inequality, \[\norm{\mathbb{E}_{m-1}[x_m x_m^\top]}_{op} \le \mathbb{E}_{m-1}[\norm{x_m x_m^\top}_{op}] \le C_x^2.\] So, \[\norm{Y_m}_{op} \le \norm{x_m x_m^\top}_{op} + \norm{\mathbb{E}_{m-1}[x_m x_m^\top]}_{op} \le 2 C_x^2\] and we can take \(R = 2 C_x^2\).

Next, we lower bound \(\lambda_{\min}(M_n)\). Write \[\mu_m \triangleq \mathbb{E}_{m-1}[x_m] = (1, \chi_{m,1}, ..., \chi_{m,d})^\top, \quad \xi_m \triangleq x_m - \mu_m = (0, z_{m,1}, ..., z_{m,d})^\top.\] Then, \(\mathbb{E}_{m-1}[\xi_m] = 0\), and by independence across coordinates, \[\mathbb{E}_{m-1}[\xi_m \xi_m^\top] = \mathrm{diag}(0, \nu_{m,1}^2, ..., \nu_{m,d}^2).\] Therefore, \[\mathbb{E}_{m-1}[x_m x_m^\top] = \mu_m \mu_m^\top + \mathrm{diag}(0, \nu_{m,1}^2, ..., \nu_{m,d}^2).\] Summing, we have \[M_n = \sum_{m=1}^n \mu_m \mu_m^\top + \mathrm{diag}(1, S_1(n), ..., S_d(n)), \quad S_i(n) = \sum_{m=1}^n \nu_{m,i}^2.\] Let \(\bar{\mu}_n = \frac{1}{n} \sum_{m=1}^n \mu_m\). Then, \[\sum_{m=1}^n \mu_m \mu_m^\top = \sum_{m=1}^{n} (\mu_m - \bar{\mu}_n)(\mu_m - \bar{\mu}_n)^\top + n \bar{\mu}_n \bar{\mu}_n^\top \succeq n \bar{\mu}_n \bar{\mu}_n^\top.\] Hence, \[M_n \succeq K_n \triangleq n \bar{\mu}_n \bar{\mu}_n^\top + \mathrm{diag}(0, S_1(n), ..., S_d(n))\] and \(\lambda_{\min}(M_n) \ge \lambda_{\min}(K_n)\). Write \(\bar{\mu}_n = (1, b_n)^\top\) with \(b_n \in \mathbb{R}^d\) and \(S(n) = \mathrm{diag}(S_1(n), ..., S_d(n))\). Then, applying Lemma 12 with \(b = b_n\) and \(S = S(n)\), we have \[\lambda_{\min}(M_n) \ge \lambda_{\min}(K_n) \ge \frac{1}{2} \min \left\{\frac{n}{1 + n b_n^\top S(n)^{-1} b_n}, \min_{1 \le i \le d} S_i(n)\right\}.\] Since \(\nu_{m,i}^2 = c_i m^{-\eta_i}\), we have \[\min_{1 \le i \le d} S_i(n) \gtrsim n^{1 - \eta_{\max}}.\] Also, since \(b_n\) is uniformly bounded, we have \[b_n^\top S(n)^{-1} b_n \lesssim \lambda_{\max}(S(n)^{-1}) = \frac{1}{\lambda_{\min}(S(n))} \lesssim n^{-(1 - \eta_{\max})}.\] Thus, \[\frac{n}{1 + n b_n^\top S(n)^{-1} b_n} \gtrsim \frac{n}{1 + n \cdot n^{-(1 - \eta_{\max})}} = \frac{n}{1 + n^{\eta_{\max}}} \gtrsim n^{1 - \eta_{\max}}.\] Together, we have \[\lambda_{\min}(M_n) \gtrsim n^{1 - \eta_{\max}}.\]

Finally, we upper bound the predictable quadratic variation process: \[\sigma_n^2 = \bignorm{\sum_{m=1}^n \mathbb{E}_{m-1}[Y_m^2]}_{op}.\] We claim that there is some constant \(C_V > 0\) such that \[\mathbb{E}_{m-1}\norm{Y_m^2}_{op} \le C_V \sum_{i=1}^{d} \nu_{m,i}^2 \text{ for all } m.\] Given this, we have \[\sigma_n^2 \le \sum_{m=1}^{n} \mathbb{E}_{m-1}\norm{Y_m^2}_{op} \le C_V \sum_{m=1}^{n} \sum_{i=1}^{d} \nu_{m,i}^2 \lesssim \sum_{m=1}^{n} m^{-\eta_{\min}} \lesssim n^{1 - \eta_{\min}}.\] To show the claim, write \[\begin{align} Y_m &= x_m x_m^\top - \mathbb{E}_{m-1}[x_m x_m^\top] \\ &= (\mu_m + \xi_m)(\mu_m + \xi_m)^\top - \left(\mu_m \mu_m^\top + \mathbb{E}_{m-1}[\xi_m \xi_m^\top]\right) \\ &= \mu_m \xi_m^\top + \xi_m \mu_m^\top + \left(\xi_m \xi_m^\top - \mathbb{E}_{m-1}[\xi_m \xi_m^\top]\right). \end{align}\] Since \(\mu_m\) are bounded, \(\mu_m \xi_m^\top + \xi_m \mu_m^\top \lesssim \norm{\xi_m}_2\). The last two terms can be bounded as \[\norm{\xi_m \xi_m^\top - \mathbb{E}_{m-1}[\xi_m \xi_m^\top]}_{op} \le \norm{\xi_m \xi_m^\top}_{op} + \norm{\mathbb{E}_{m-1}[\xi_m \xi_m^\top]}_{op} \le \norm{\xi_m}_2^2 + \mathbb{E}_{m-1} [\norm{\xi_m}_2^2].\] So, \[\norm{Y_m}_{op} \lesssim \norm{\xi_m}_2 + \norm{\xi_m}_2^2 + \mathbb{E}_{m-1}[\norm{\xi_m}_2^2],\] and \[\mathbb{E}_{m-1}\norm{Y_m}_{op}^2 \lesssim \mathbb{E}_{m-1}[\norm{\xi_m}_2^2] + \mathbb{E}_{m-1}[\norm{\xi_m}_2^3] + \mathbb{E}_{m-1}[\norm{\xi_m}_2^4].\] Since \(\xi_m\) are bounded, the higher moment terms can be controlled by the second moment term (since \(\mathrm{Var}(z_{m,i}) = \nu_{m,i}^2 = c_i m^{-\eta_i}\), for large \(m\) the second moment term dominates the higher moment terms anyway). So, \[\mathbb{E}_{m-1}\norm{Y_m}_{op}^2 \lesssim \sum_{i=1}^{d} \nu_{m,i}^2,\] proving the claim.

We are ready to put everything together now. By the Matrix Freedman inequality (Theorem 7), for any \(t \ge 0\), \[\mathbb{P}\left(\lambda_{\max}\left(-\sum_{m=1}^n Y_m\right) \ge t\right) \le (d+1) \exp\left(-\frac{t^2/2}{\sigma_n^2 + R t / 3}\right).\] By 84 , we have \[\mathbb{P}\left(\lambda_{\min}(\mathcal{J}_n) \le \lambda_{\min}(M_n) - t\right) \le (d+1) \exp\left(-\frac{t^2/2}{\sigma_n^2 + R t / 3}\right).\] Since \(\lambda_{\min}(M_n) \gtrsim n^{1 - \eta_{\max}}\) and \(\sigma_n^2 \lesssim n^{1 - \eta_{\min}}\), choose \(t = \frac{1}{2} n^{1 - \eta_{\max}}\) to get \[\mathbb{P}\left(\lambda_{\min}(\mathcal{J}_n) \le \frac{1}{2} n^{1 - \eta_{\max}}\right) \le (d+1) \exp\left(-\frac{n^{2(1 - \eta_{\max})}/8}{C_1 n^{1 - \eta_{\min}} + R n^{1 - \eta_{\max}} / 6}\right)\] for some constant \(C_1 > 0\) for all large \(n\). Since \(\eta_{\min} \le \eta_{\max}\), we have \(n^{1 - \eta_{\min}} \ge n^{1 - \eta_{\max}}\), so \[\frac{n^{2(1 - \eta_{\max})}/8}{C_1 n^{1 - \eta_{\min}} + R n^{1 - \eta_{\max}} / 6} \gtrsim \frac{n^{2(1 - \eta_{\max})}}{n^{1 - \eta_{\min}}} = n^{1 + \eta_{\min} - 2 \eta_{\max}}.\] Redefining the constants, we have shown that there exist constants \(C_1, C_2 > 0\) such that for all large \(n\), \[\mathbb{P}\left(\lambda_{\min}(\mathcal{J}_n) \le C_1 n^{1 - \eta_{\max}}\right) \le (d+1) \exp\left(-C_2 n^{1 + \eta_{\min} - 2 \eta_{\max}}\right).\] ◻

We are now ready to prove Theorem 6. We drop the subscript \(i\) and the superscript \(*\) for notational simplicity. Let \(\mathcal{J}_n = \sum_{m=1}^n x_{m,i}^{in} (x_{m,i}^{in})^\top\) be the empirical Fisher information matrix up to time \(n\), where without loss of generality we can re-order the parameters such that \((x_{n,i}^{in})^\top \theta^*_i\) is the true demand (omitting noise). The compactness of the projection set \(\Theta^{in}_i\) from Section 3.2 provides a finite a.s.bound \(D_\Theta \triangleq \operatorname{diam}(\Theta^{in}_i)^2\) on \(\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2\) that we use below to truncate the integrals at infinity. By Lemma 5, there exist constants \(c_1, c_2 > 0\) and \(N_0 \in \mathbb{N}\) such that for all \(n \ge N_0\), \[\begin{align} \mathbb{E}\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 &= \int_{0}^{D_\Theta} \mathbb{P}\left(\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 > x, \lambda_{\min} (\mathcal{J}_n) > c_1 n^{1 - \eta_{\max}}\right) \, dx \\ & \quad + \int_{0}^{D_\Theta} \mathbb{P}\left(\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 > x, \lambda_{\min} (\mathcal{J}_n) \le c_1 n^{1 - \eta_{\max}}\right) \, dx \\ &\le \int_{0}^{\infty} \mathbb{P}\left(\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 > x, \lambda_{\min} (\mathcal{J}_n) > c_1 n^{1 - \eta_{\max}}\right) \, dx \\ &\quad + D_\Theta\,(N+1) \exp\left(-c_2 n^{1 + \eta_{\min} - 2 \eta_{\max}}\right). \end{align}\] By Lemma 3 of [13], we have \[\mathbb{P}\left(\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 > x, \,\lambda_{\min} (\mathcal{J}_n) > c_1 n^{1 - \eta_{\max}}\right) \le kn \exp \left(- c_1 \rho \min\{x, \sqrt{x}\} n ^{1 - \eta_{\max}}\right)\] for some \(k, \rho > 0\) and all \(x > 0\) and \(n\) large enough. Fix \(\delta > 0\) small. Then, we have \[\begin{align} &\quad\int_{0}^{\infty} \mathbb{P}\left(\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 > x, \lambda_{\min} (\mathcal{J}_n) > c_1 n^{1 - \eta_{\max}}\right) \, dx \\ &\le n^{\eta_{\max} + \delta - 1} + \int_{n^{\eta_{\max} + \delta - 1}}^{1} kn \exp \left(- c_1 \rho \cdot \min\{x, \sqrt{x}\} n ^{1 - \eta_{\max}}\right) \, dx \\ &\quad + \int_{1}^{\infty} kn \exp \left(- c_1 \rho \cdot\min\{x, \sqrt{x}\} n ^{1 - \eta_{\max}}\right) \, dx. \end{align}\] We want to show that the last two integrals are dominated by the first term. For the first integral, we have \[\begin{align} &\quad \int_{n^{\eta_{\max} + \delta - 1}}^{1} kn \exp \left(- c_1 \rho \cdot \min\{x, \sqrt{x}\} n ^{1 - \eta_{\max}}\right) \, dx \\ &= kn \int_{n^{\eta_{\max} + \delta - 1}}^{1} \exp \left(- c_1 \rho \cdot n ^{1 - \eta_{\max}}x \right) \, dx. \end{align}\] For \(a \in (0,1)\) and \(b > 0\), \(\int_a^1 e^{-bx} \, dx = \frac{1}{b}(e^{-ba} - e^{-b})\). So, with \(a = n^{\eta_{\max} + \delta - 1}\) and \(b = c_1 \rho n^{1 - \eta_{\max}}\), the above integral equals \[\begin{align} & \quad kn \cdot \frac{1}{c_1 \rho n^{1 - \eta_{\max}}} \left(\exp\left(- c_1 \rho n ^{1 - \eta_{\max}} \cdot n^{\eta_{\max} + \delta - 1}\right) - \exp\left(- c_1 \rho n ^{1 - \eta_{\max}}\right)\right) \\ &= \frac{k}{c_1 \rho} n^{\eta_{\max}} \left(\exp\left(- c_1 \rho n ^\delta\right) - \exp\left(- c_1 \rho n ^{1 - \eta_{\max}}\right)\right) \end{align}\] For the second integral, we have \[\begin{align} &\quad \int_{1}^{\infty} kn \exp \left(- c_1 \rho \cdot\min\{x, \sqrt{x}\} n ^{1 - \eta_{\max}}\right) \, dx \\ &= kn \int_{1}^{\infty} \exp \left(- c_1 \rho \cdot n ^{1 - \eta_{\max}}\sqrt{x} \right) \, dx. \end{align}\] For \(a > 0\), \(\int_1^{\infty} e^{-a \sqrt{x}} \, dx = 2 \int_1^{\infty} t e^{-a t} \, dt\) with the substitution \(t = \sqrt{x}\). Integration by parts gives \(\frac{2(a+1)}{a^2} e^{-a}\). So, with \(a = c_1 \rho n^{1 - \eta_{\max}}\), the above integral equals \[\begin{align} &\quad kn \cdot \frac{2(c_1 \rho n^{1 - \eta_{\max}} + 1)}{(c_1 \rho n^{1 - \eta_{\max}})^2} \exp\left(- c_1 \rho n ^{1 - \eta_{\max}}\right) \\ &= \frac{2k}{(c_1 \rho)^2} n^{\eta_{\max}} \left(c_1 \rho + n^{\eta_{\max} - 1}\right) \exp\left(- c_1 \rho n ^{1 - \eta_{\max}}\right) \end{align}\] Since \(\delta > 0\) and \(1 - \eta_{\max} > 0\), both integrals, as well as the \(D_\Theta\,(N+1) \exp\left(-c_2 n^{1 + \eta_{\min} - 2 \eta_{\max}}\right)\) term from before, are dominated by the exponential terms, which are in turn dominated by \(n^{\eta_{\max} + \delta - 1}\) for large \(n\). Therefore, \[\mathbb{E}\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 \lesssim n^{\eta_{\max} + \delta - 1}.\] Since \(\delta > 0\) is arbitrary, we have \[\mathbb{E}\norm{\hat{\mathbf{\theta}}_n - \mathbf{\theta}}_2^2 = O\left(n^{\eta_{\max} - 1}\right).\] The proof is complete.

9.12 Auxiliary Lemmas↩︎

This subsection collects two auxiliary lemmas referenced from the proofs of Theorems 4 and 5: a restatement of the informed-OLS rate as both an \(L^2\) and an almost-sure bound, and a closed-form expression for the rate function \(c^*\).

Lemma 6 (Informed-OLS rate, restated). Under \(1 + \eta_{\min} > 2\,\eta_{\max}\) of Theorem 6:

  • \(\sum_{j \in \cI^{in}} \EE\bigl[\norm{\tilde{e}_{n,j}^{in}}_2^2\bigr] = O(n^{\eta_{\max} - 1})\);

  • \(\sum_{j \in \cI^{in}}\norm{\tilde{e}_{n,j}^{in}}_2^2 = O(n^{\eta_{\max} - 1 + \varepsilon})\) a.s.for every \(\varepsilon > 0\), and in particular \(\hat{\theta}_{n,j}^{in} \to \theta_j^{*,in}\) a.s.for every \(j \in \cI^{in}\);

  • \(\sum_{n \ge 1} \tfrac{1}{n+1}\sum_{j \in \cI^{in}}\norm{\tilde{e}_{n,j}^{in}}_2^2 < \infty\) a.s.

Proof. Part (a) is exactly Theorem 6. Part (b) upgrades (a) to an almost-sure rate by Borel–Cantelli applied to the Keskin–Zeevi tail bound of [13]: conditioning on the high-probability design-matrix event \(\{\lambda_{\min}(\cJ_{n,j}^{in}) \ge C_1\, n^{1 - \eta_{\max}}\}\) from Appendix 9.11, the parameter-MSE tail is \(\Pr(\norm{\tilde{e}_{n,j}^{in}}_2^2 > x_n) \le \exp(-c\, n^{1-\eta_{\max}}\, x_n)\). Setting \(x_n = n^{\eta_{\max} - 1 + \varepsilon}\) makes the right-hand side \(\exp(-c\,n^\varepsilon)\), which is summable in \(n\) for every \(\varepsilon > 0\), so Borel–Cantelli gives \(\norm{\tilde{e}_{n,j}^{in}}_2^2 \le n^{\eta_{\max} - 1 + \varepsilon}\) a.s.eventually for every \(\varepsilon > 0\). Taking \(\varepsilon \downarrow 0\) along a countable sequence yields the stated rate, and \(\hat{\theta}_{n,j}^{in} \to \theta_j^{*,in}\) a.s.since the exponent \(\eta_{\max} - 1 + \varepsilon < 0\) for \(\varepsilon < 1 - \eta_{\max}\). Part (c) follows from (a) by Tonelli applied to the non-negative integrand: \[\EE \left[\sum_{n\ge 1}\frac{1}{n+1}\sum_{j\in\cI^{in}}\norm{\tilde{e}_{n,j}^{in}}_2^2 \right] \;=\; \sum_{n\ge 1}\frac{1}{n+1}\,O\!\left(n^{\eta_{\max} - 1}\right) \;=\; \sum_{n\ge 1}O\!\left(n^{\eta_{\max} - 2}\right) \;<\; \infty,\] since \(\eta_{\max} - 2 < -1\). ◻

Lemma 7 (Closed form of \(c^*\)). On the feasible interval \((\Phi^*, \Phi^{**}) \triangleq (C_x\,\bar\Psi/K_2,\, K_1/L_\phi^{ob})\) identified in Step 4 of the proof of Theorem 5, the rate function \(c^*(\lambda)\) of 78 attains its supremum \(c^* \triangleq \sup_\lambda c^*(\lambda)\) at the unique positive root \[\lambda^* \;=\; \frac{[K_1 - C_x^2(K_2 - 1)] + \sqrt{[K_1 - C_x^2(K_2 - 1)]^2 + 4\, L_\phi^{ob}\, C_x^3\,\bar\Psi}}{2\, L_\phi^{ob}},\] and admits the closed form \[\label{eq:cstar-closed} c^* \;=\; c^*(\lambda^*) \;=\; \frac{1 + K_2}{2} \;+\; \frac{K_1 - \sqrt{[K_1 - C_x^2(K_2-1)]^2 + 4\, L_\phi^{ob}\, C_x^3\,\bar\Psi}}{2\, C_x^2}.\qquad{(9)}\]

Proof. The first argument of \(c^*(\lambda) = \min\,\left\{1 + (K_1 - \lambda L_\phi^{ob})/C_x^2,\; K_2 - C_x\bar\Psi/\lambda\right\}\) is a strictly decreasing affine function of \(\lambda\), falling to \(1\) at the upper boundary \(\Phi^{**} = K_1/L_\phi^{ob}\); the second is a strictly increasing function rising from \(0\) at the lower boundary \(\Phi^* = C_x\bar\Psi/K_2\). The supremum of the \(\min\) of an increasing and a decreasing function on a feasible interval is attained at their unique crossing point. Setting the two arguments equal and clearing denominators by multiplying by \(\lambda\, C_x^2\) yields the quadratic \[L_\phi^{ob}\, \lambda^2 \;+\; \bigl[\,C_x^2(K_2 - 1) - K_1\,\bigr]\, \lambda \;-\; C_x^3\,\bar\Psi \;=\; 0.\] Its product of roots is \(-C_x^3\bar\Psi/L_\phi^{ob} < 0\) whenever \(\bar\Psi > 0\), so there is a unique positive root \(\lambda^*\) as displayed. Substituting \(\lambda^*\) into either branch of \(c^*(\lambda)\) and simplifying gives ?? . When \(\bar\Psi = 0\) the quadratic degenerates and \(c^*(\lambda)\) becomes the minimum of a strictly decreasing line and the constant \(K_2\), with optimum \(\min\,\left\{1 + K_1/C_x^2,\, K_2\right\}\), which is exactly ?? evaluated at \(\bar\Psi = 0\). ◻

10 Technical Lemmas↩︎

Lemma 8 (Empirical dispersion lower bounds). Let \((\mathcal{F}_n)_{n\ge 0}\) be a filtration. Suppose \((x_n)_{n\ge 1}\) satisfies \[x_{n} = \mu_n + z_n,\qquad \mu_n \in \mathcal{F}_{n-1},\] where \((z_n)_{n\ge 1}\) are independent across \(n\), satisfy \(z_n \perp \mathcal{F}_{n-1}\), \(\mathbb{E}[z_n]=0\), and have finite second moments \[\sigma_n^2 \triangleq \mathrm{Var}(z_n) = \mathbb{E}[z_n^2]\in (0,\infty).\] Define the empirical mean \(\bar x_n \triangleq \frac{1}{n}\sum_{m=1}^n x_m\) and the empirical dispersion \[J_n \triangleq \sum_{m=1}^n (x_m-\bar x_n)^2.\] Then, for every \(n\ge 1\), \[\label{eq:Jn-exp-lower-general} \mathbb{E}[J_n] \;\ge\; \left(1-\frac{1}{n}\right)\sum_{m=1}^n \sigma_m^2.\qquad{(10)}\] In addition, suppose \(\mu_n\) and \(z_n\) are almost surely uniformly bounded. Then:

(i) If \((z_n)\) are i.i.d.with \(\mathbb{E}[z_n]=0\) and \(\mathrm{Var}(z_n)=\nu^2>0\), then \[\label{eq:Jn-as-liminf-constvar} \liminf_{n\to\infty}\frac{J_n}{n} \;\ge\; \nu^2 \qquad\text{and}\qquad J_n=\Theta(n)\quad\text{a.s.}\qquad{(11)}\]

(ii) If \((z_n)\) are independent with \(\mathbb{E}[z_n]=0\), and there exist constants \(\underline{\nu},\overline{\nu}>0\) such that \[\label{eq:sigma-poly-bounds} \underline{\nu}^2\,n^{-c}\;\le\;\sigma_n^2\;\le\;\overline{\nu}^2\,n^{-c}\quad\forall n \quad \text{for some }c\in(0,1),\qquad{(12)}\] Then, \[\liminf_{n\to\infty}\frac{J_n}{n^{1-c}} \;\ge\; \frac{\underline{\nu}^2}{1-c} \qquad\text{a.s.}.\]

Proof of Lemma 8. We use the identity \[J_n \;=\; \sum_{m=1}^n x_m^2 \;-\; n\bar x_n^2 \;=\; \sum_{m=1}^n x_m^2 \;-\; \frac{1}{n}\left(\sum_{m=1}^n x_m\right)^2.\] Write \(x_m=\mu_m+z_m\). Expanding and taking expectations, \[\mathbb{E}\left[\sum_{m=1}^n x_m^2\right] =\sum_{m=1}^n \mathbb{E}[(\mu_m+z_m)^2] =\sum_{m=1}^n \mathbb{E}[\mu_m^2] + 2\sum_{m=1}^n \mathbb{E}[\mu_m z_m] + \sum_{m=1}^n \mathbb{E}[z_m^2].\] Since \(\mu_m\in\mathcal{F}_{m-1}\) and \(z_m\perp\mathcal{F}_{m-1}\) with \(\mathbb{E}[z_m]=0\), \[\mathbb{E}[\mu_m z_m] =\mathbb{E}\left[\mu_m \mathbb{E}[z_m\mid \mathcal{F}_{m-1}]\right]=0,\] hence \[\label{eq:sumx2-exp} \mathbb{E}\left[\sum_{m=1}^n x_m^2\right] =\sum_{m=1}^n \mathbb{E}[\mu_m^2] + \sum_{m=1}^n \sigma_m^2.\tag{85}\] Next, \[\sum_{m=1}^n x_m = \sum_{m=1}^n \mu_m + \sum_{m=1}^n z_m,\] so \[\left(\sum_{m=1}^n x_m\right)^2 =\left(\sum_{m=1}^n \mu_m\right)^2 +2\left(\sum_{m=1}^n \mu_m\right)\left(\sum_{m=1}^n z_m\right) +\left(\sum_{m=1}^n z_m\right)^2.\] Taking expectations, the cross term vanishes: \[\mathbb{E}\left[\left(\sum_{m=1}^n \mu_m\right)\left(\sum_{m=1}^n z_m\right)\right] =\sum_{m=1}^n \mathbb{E}[\mu_m z_m]=0.\] By independence across \(m\) and \(\mathbb{E}[z_m]=0\), \[\mathbb{E}\left[\left(\sum_{m=1}^n z_m\right)^2\right] =\sum_{m=1}^n \mathbb{E}[z_m^2] + 2\sum_{1\le k<m\le n}\mathbb{E}[z_k z_m] =\sum_{m=1}^n \sigma_m^2.\] Therefore, \[\label{eq:sumx-exp} \mathbb{E}\left[\left(\sum_{m=1}^n x_m\right)^2\right] =\mathbb{E}\left[\left(\sum_{m=1}^n \mu_m\right)^2\right] + \sum_{m=1}^n \sigma_m^2.\tag{86}\] Plugging 85 and 86 into \(\mathbb{E}[J_n]\) gives \[\begin{align} \mathbb{E}[J_n] &= \left(\sum_{m=1}^n \mathbb{E}[\mu_m^2] + \sum_{m=1}^n \sigma_m^2\right) -\frac{1}{n}\left(\mathbb{E}\left[\left(\sum_{m=1}^n \mu_m\right)^2\right] + \sum_{m=1}^n \sigma_m^2\right) \\ &=\left(\sum_{m=1}^n \mathbb{E}[\mu_m^2] - \frac{1}{n} \mathbb{E}\left[\left(\sum_{m=1}^n \mu_m\right)^2\right]\right) +\left(1-\frac{1}{n}\right)\sum_{m=1}^n \sigma_m^2. \end{align}\] By Cauchy–Schwarz, \((\sum_{m=1}^n \mu_m)^2 \le n\sum_{m=1}^n \mu_m^2\) pointwise, hence \[\mathbb{E}\left[\left(\sum_{m=1}^n \mu_m\right)^2\right]\le n\sum_{m=1}^n \mathbb{E}[\mu_m^2],\] so the first bracket is nonnegative. This proves ?? .

Suppose \(|\mu_n|\le B\) a.s.and \(|z_n|\le K\) a.s. We first assume that \((z_n)\) are i.i.d.with \(\mathbb{E}[z_n]=0\) and \(\mathrm{Var}(z_n)=\nu^2>0\). Expand \[\begin{align} J_n &=\sum_{m=1}^n (\mu_m+z_m)^2 - \frac{1}{n}\left(\sum_{m=1}^n (\mu_m+z_m)\right)^2 \notag\\ &=\underbrace{\left(\sum_{m=1}^n \mu_m^2-\frac{1}{n}\left(\sum_{m=1}^n \mu_m\right)^2\right)}_{=:J_n^\mu} +2\underbrace{\left(\sum_{m=1}^n \mu_m z_m-\frac{1}{n}\left(\sum_{m=1}^n \mu_m\right)\left(\sum_{m=1}^n z_m\right)\right)}_{=:C_n} \notag \\ &\quad +\underbrace{\left(\sum_{m=1}^n z_m^2-\frac{1}{n}\left(\sum_{m=1}^n z_m\right)^2\right)}_{=:J_n^z}. \label{eq:Jn-decomp} \end{align}\tag{87}\] Since \(J_n^\mu\) is the empirical dispersion of \(\mu_1,\dots,\mu_n\) and variances are nonnegative, \(J_n^\mu\ge 0\), and we have \(J_n\ge J_n^z + 2C_n\). By the strong law of large numbers, \[\frac{1}{n}\sum_{m=1}^n z_m \to 0 \quad\text{a.s.},\qquad \frac{1}{n}\sum_{m=1}^n z_m^2 \to \mathbb{E}[z_1^2]=\nu^2 \quad\text{a.s.}\] Therefore, \[\frac{J_n^z}{n} =\frac{1}{n}\sum_{m=1}^n z_m^2 - \left(\frac{1}{n}\sum_{m=1}^n z_m\right)^2 \to \nu^2 \quad\text{a.s.}\] Define \[M_n := \sum_{m=1}^n \mu_m z_m.\] Since \(\mu_m\in\mathcal{F}_{m-1}\) and \(\mathbb{E}[z_m\mid \mathcal{F}_{m-1}]=0\), we have \[\mathbb{E}[M_n\mid \mathcal{F}_{n-1}] = M_{n-1},\] so \((M_n,\mathcal{F}_n)\) is a martingale. Moreover, \(|M_n-M_{n-1}|=|\mu_n z_n|\le BK\) a.s. By Azuma–Hoeffding, for any \(\varepsilon>0\), \[\mathbb{P}\left(|M_n|\ge \varepsilon n\right)\le 2\exp\left(-\frac{\varepsilon^2 n}{2B^2K^2}\right),\] and the RHS is summable in \(n\). By Borel–Cantelli, almost surely, \(|M_n|<\varepsilon n\) eventually. Since \(\varepsilon>0\) is arbitrary, this shows that \[\frac{1}{n}\sum_{m=1}^n \mu_m z_m = \frac{M_n}{n} \to 0 \quad\text{a.s.}\] For the second part of \(C_n\), note \(|\sum_{m=1}^n \mu_m|\le nB\), so \[\left|\frac{1}{n}\cdot\frac{1}{n}\left(\sum_{m=1}^n \mu_m\right)\left(\sum_{m=1}^n z_m\right)\right| \le B\left|\frac{1}{n}\sum_{m=1}^n z_m\right| \to 0 \quad\text{a.s.}\] Therefore, \(C_n/n\to 0\) a.s. Together, dividing \(J_n\ge J_n^z + 2C_n\) by \(n\), and taking \(\liminf\) yields \[\liminf_{n\to\infty}\frac{J_n}{n} \ge \lim_{n\to\infty}\frac{J_n^z}{n} + 2\lim_{n\to\infty}\frac{C_n}{n} = \nu^2 \quad\text{a.s.}\] This proves ?? , where the a.s.upper bound \(J_n=O(n)\) follows from the boundedness of \(\mu_n\) and \(z_n\).

Now, suppose instead that \((z_n)\) are independent with \(\mathbb{E}[z_n]=0\) and variances satisfying ?? for some \(c\in(0,1)\). Following the same decomposition 87 , we again have \(J_n\ge J_n^z + 2C_n\). We now analyze the two terms \(J_n^z\) and \(C_n\) separately. Let \[S_n \;\triangleq\; \sum_{m=1}^n z_m, \qquad Q_n \;\triangleq\; \sum_{m=1}^n z_m^2, \qquad V_n \;\triangleq\; \sum_{m=1}^n \sigma_m^2 .\] Then \[J_n^z = Q_n - \frac{1}{n}S_n^2, \qquad \frac{J_n^z}{V_n} = \frac{Q_n}{V_n} - \frac{S_n^2/n}{V_n}.\] We first show that \(Q_n/V_n \to 1\) almost surely. Define the centered variables \[w_m \;\triangleq\; z_m^2 - \sigma_m^2,\] which satisfy \(\mathbb{E}[w_m]=0\) and are independent. Since \(|z_m|\le K\) a.s.and \(\sigma_m^2\le \overline{\nu}^2\) by ?? , we have \[|w_m|\le K^2+\overline{\nu}^2 \quad\text{a.s.}\] Moreover, \[\mathrm{Var}(w_m)\le \mathbb{E}[w_m^2] = \mathbb{E}[z_m^4] - \sigma_m^4 \le \mathbb{E}[z_m^4]\le K^2\mathbb{E}[z_m^2]=K^2\sigma_m^2.\] Because \(c\in(0,1)\), the growth condition ?? implies \(V_n\asymp n^{1-c}\), and hence \[\sum_{m=1}^\infty \frac{\mathrm{Var}(w_m)}{V_m^2} \;\le\; \sum_{m=1}^\infty \frac{K^2\sigma_m^2}{V_m^2} \;\lesssim\; \sum_{m=1}^\infty \frac{m^{-c}}{m^{2(1-c)}} = \sum_{m=1}^\infty m^{-(2-c)} <\infty.\] Therefore, by Kolmogorov two-series theorem, the series \(\sum_{m=1}^\infty w_m/V_m\) converges almost surely. Since \(V_m \rightarrow \infty\), Kronecker’s lemma yields \[\frac{1}{V_n}\sum_{m=1}^n (z_m^2-\sigma_m^2) =\frac{1}{V_n}\sum_{m=1}^n w_m \;\longrightarrow\; 0 \qquad\text{a.s.}\] Equivalently, \[\frac{Q_n}{V_n}\longrightarrow 1 \qquad\text{a.s.}\] Next, we control the term \(S_n^2/n\). Since \[\sum_{m=1}^\infty \frac{\mathrm{Var}(z_m)}{m^{2}} \;\lesssim\; \sum_{m=1}^\infty m^{-(2+c)}<\infty,\] Kolmogorov’s two-series theorem and Kronecker’s lemma again yield \[\frac{1}{n}\sum_{m=1}^n z_m \longrightarrow 0 \qquad\text{a.s.}\] An argument completely analogous to \(\sum_{m=1}^n w_m / V_n \rightarrow 0\) a.s. shows that \[\frac{S_n}{V_n} \;\longrightarrow\; 0 \qquad\text{a.s.}\] Hence, \[\frac{S_n^2/n}{V_n} \longrightarrow 0 \qquad\text{a.s.}\] Combining, we obtain \[\frac{J_n^z}{V_n}\longrightarrow 1 \qquad\text{a.s.}\]

Recall \[C_n = \sum_{m=1}^n \mu_m z_m -\frac{1}{n}\left(\sum_{m=1}^n \mu_m\right)\left(\sum_{m=1}^n z_m\right).\] The second term divided by \(V_n\) goes to zero a.s. because \(\frac{1}{n}\sum_{m=1}^n \mu_m\) is bounded as before and we showed that \(\frac{S_n}{V_n}\to 0\) a.s. The first term divided by \(V_n\) also goes to zero a.s. because \(\mu_n\) are bounded and the same Kolmogorov–Kronecker argument applies. So, \[\frac{C_n}{V_n}\longrightarrow 0 \qquad\text{a.s.}\] Using \(J_n\ge J_n^z + 2C_n\), dividing by \(V_n\), and taking \(\liminf\), \[\liminf_{n\to\infty}\frac{J_n}{V_n} \;\ge\; \lim_{n\to\infty}\frac{J_n^z}{V_n} + 2\lim_{n\to\infty}\frac{C_n}{V_n} = 1 \qquad\text{a.s.}\] Finally, under ?? , \[V_n =\sum_{m=1}^n \sigma_m^2 \ge \underline{\nu}^2\sum_{m=1}^n m^{-c} = \frac{\underline{\nu}^2}{1-c}\,n^{1-c},\] which yields \[\liminf_{n\to\infty}\frac{J_n}{\frac{\underline{\nu}^2}{1-c}\,n^{1-c}} \ge \liminf_{n\to\infty} \frac{J_n}{V_n} \ge 1 \quad \Longrightarrow \quad \liminf_{n\to\infty}\frac{J_n}{n^{1-c}} \ge \frac{\underline{\nu}^2}{1-c} \quad\text{a.s.}\] This completes the proof. ◻

Lemma 9 (Uniform conditional second-moment lower bound). Let \((\mathcal{F}_n)_{n\ge 0}\) be a filtration and define the feature vector \[y_{n, i} \;\triangleq\; \left(1,\,x_{n,i}\right)^\top \in \mathbb{R}^{2}.\] Assume that for each \(i\in\{1,\dots,N\}\), \[x_{n,i} = \mu_{n,i} + z_{n,i},\qquad \mu_{n,i}\in\mathcal{F}_{n-1},\] and conditional on \(\mathcal{F}_{n-1}\) the random variables \(\{z_{n,i}\}_{i=1}^N\) are independent with \[\mathbb{E}[z_{n,i}\mid \mathcal{F}_{n-1}] = 0,\qquad \mathrm{Var}(z_{n,i}\mid \mathcal{F}_{n-1}) = \nu_i^2>0 \quad\text{a.s.}\] Assume further that there exists \(B>0\) such that \(|\mu_{n,i}|\le B\) a.s. for all \(n\) and \(i\). Let \(\nu_* \triangleq \min_{1\le i\le N}\nu_i>0\). Then for all \(n\ge 1\), \[\mathbb{E}[y_{n,i} y_{n,i}^\top\mid \mathcal{F}_{n-1}] \;\succeq\; C_M I_{2} \quad \text{for all } i,\] where \[C_M \;\triangleq\; \frac{1}{2} \min\left\{\frac{\nu_*^2}{\nu_*^2 + B^2}, \nu_*^2\right\} > 0.\]

Proof of Lemma 9. Fix \(n\ge 1\) and \(i\). By zero conditional means of \(\{z_{n,i}\}_{i=1}^N\), \[\mathbb{E}[x_{n,i}\mid\mathcal{F}_{n-1}] = \mu_{n,i}, \qquad \mathbb{E}[x_{n,i}^2\mid\mathcal{F}_{n-1}] = \mu_{n,i}^2 + \nu_i^2.\] Therefore, \[\mathbb{E}[y_{n,i} y_{n,i}^\top\mid\mathcal{F}_{n-1}] = \begin{pmatrix} 1 & \mu_{n,i}\\ \mu_{n,i} & \mu_{n,i}^2 + \nu_i^2 \end{pmatrix} = (1,\mu_{n,i})^\top(1,\mu_{n,i}) + \begin{pmatrix} 0 & 0\\ 0 & \nu_i^2 \end{pmatrix}.\] We now apply Lemma 12 with \(n=1\), \(b=\mu_{n,i}\), and \(S=\nu_i^2\). This yields \[\lambda_{\min}\!\left(\mathbb{E}[y_{n,i} y_{n,i}^\top\mid\mathcal{F}_{n-1}]\right) \;\ge\; \frac{1}{2} \min\left\{ \frac{\nu_i^2}{\nu_i^2 + \mu_{n,i}^2}, \nu_i^2 \right\}.\] By the uniform boundedness assumption \(|\mu_{n,i}|\le B\) a.s.for all \(i\), \[\lambda_{\min}\!\left(\mathbb{E}[y_{n,i} y_{n,i}^\top\mid\mathcal{F}_{n-1}]\right) \ge \frac{1}{2} \min\left\{\frac{\nu_i^2}{\nu_i^2 + B^2}, \nu_i^2\right\} \ge \frac{1}{2} \min\left\{\frac{\nu_*^2}{\nu_*^2 + B^2}, \nu_*^2\right\},\] completing the proof. ◻

Lemma 10 (Spectral lower bound for adaptive covariance matrices). Let \((\mathcal{F}_n)_{n\ge 0}\) be a filtration. Let \((x_n)_{n\ge 1}\) be an \(\mathbb{R}^d\)-valued adapted process such that:

  • There exists \(C_x<\infty\) such that \(\norm{x_n}_2\le C_x\) a.s. for all \(n\ge 1\).

  • There exists \(C_M>0\) such that \(\mathbb{E}[x_n x_n^\top|\mathcal{F}_{n-1}]\succeq C_M I_d\) a.s. for all \(n\ge 1\).

Define \(S_n = \sum_{m=1}^n x_m x_m^\top\), \(M_n = S_n / n\). Then, for any \(\delta>0\), almost surely, \[S_n \succeq (C_M - \delta) n I_d, \quad \text{and} \quad M_n \succeq (C_M - \delta) I_d \quad \text{eventually}.\] As a consequence, for any \(x\) with \(\norm{x}_2 \le C_x\), \[x^\top S_n^{-1} x \le \frac{C_x^2}{(C_M - \delta) n}, \quad \text{and} \quad x^\top M_n^{-1} x \le \frac{C_x^2}{C_M - \delta}\] eventually almost surely.

Proof of Lemma 10. Fix any deterministic unit vector \(v\in\mathbb{R}^d\) with \(\norm{v}_2=1\). Define the scalar process \[Z_m(v) \triangleq (v^\top x_m)^2.\] Then \(Z_m(v)\) is \(\mathcal{F}_m\)-measurable and bounded a.s.: \(0 \le Z_m(v) \le C_x^2\). Define the martingale difference sequence \[\Delta_m(v) \triangleq Z_m(v) - \mathbb{E}[Z_m(v)|\mathcal{F}_{m-1}].\] Since \(\norm{\Delta_m(v)}_2 \le 2C_x^2\) a.s., \(\Delta_m(v)\in L^2\) and \[\sum_{m=1}^\infty \frac{\mathbb{E}[\Delta_m(v)^2]}{m^2} \le \sum_{m=1}^\infty \frac{4C_x^4}{m^2}<\infty.\] By the Martingale Strong Law of Large Numbers, we have \[\frac{1}{n}\sum_{m=1}^n \Delta_m(v) \to 0 \quad\text{a.s.}\] We can compute a conditional mean lower bound: \[\mathbb{E}[Z_m(v)\mid \mathcal{F}_{m-1}] = v^\top \mathbb{E}[x_m x_m^\top|\mathcal{F}_{m-1}] v \ge C_M v^\top v = C_M \quad \text{a.s.}\] Together, we have \[\frac{1}{n}\sum_{m=1}^n Z_m(v) = \frac{1}{n}\sum_{m=1}^n \mathbb{E}[Z_m(v)\mid \mathcal{F}_{m-1}] + \frac{1}{n}\sum_{m=1}^n \Delta_m(v) \ge C_M + o(1) \quad\text{a.s.}\] Hence, for each fixed unit vector \(v\), \[\liminf_{n\to\infty}\frac{1}{n}\sum_{m=1}^n (v^\top x_m)^2 \ge C_M \quad\text{a.s.}\] Since \(v^\top S_n v = \sum_{m=1}^n (v^\top x_m)^2\), we equivalently have \[\liminf_{n\to\infty} \frac{1}{n} v^\top S_n v \ge C_M \quad\text{a.s. for each fixed unit vector } v.\] We next upgrade the bound to hold uniformly over an \(\varepsilon\)-net. Fix \(\varepsilon\in(0,1)\). There exists a finite set \(\mathcal{N}_\varepsilon\subset \mathbb{S}^{d-1}\) (the unit sphere) such that \(|\mathcal{N}_\varepsilon|\le (1+2/\varepsilon)^d\) and there exists \(v\in\mathcal{N}_\varepsilon\) with \(\norm{u-v}_2\le \varepsilon\) for every \(u\in\mathbb{S}^{d-1}\). For each fixed \(v\in\mathcal{N}_\varepsilon\), we have \[\liminf_{n\to\infty}\frac{1}{n}v^\top S_n v \ge C_M \quad\text{a.s.}\] So, for each \(v\in\mathcal{N}_\varepsilon\), there exists an a.s. finite random \(N_v(\omega)\) such that for all \(n\ge N_v(\omega)\), \[\frac{1}{n}v^\top S_n v \ge C_M-\frac{\delta}{2} \quad\text{a.s.}\] Since \(\mathcal{N}_\varepsilon\) is finite, define \(N_0(\omega)\triangleq \max_{v\in\mathcal{N}_\varepsilon} N_v(\omega) < \infty\) a.s., and for all \(n\ge N_0(\omega)\) and all \(v\in\mathcal{N}_\varepsilon\), \[\frac{1}{n}v^\top S_n v \ge C_M-\frac{\delta}{2} \quad\text{a.s.}\] Now, for any unit vector \(u\in\mathbb{S}^{d-1}\), there exists \(v\in\mathcal{N}_\varepsilon\) such that \(\norm{u-v}_2\le \varepsilon\). Then, \[|u^\top S_n u - v^\top S_n v| \le \norm{S_n}_{\text{op}} \norm{uu^\top - vv^\top}_{\text{op}}.\] Since \[uu^\top - vv^\top = (u-v)u^\top + v(u-v)^\top,\] we have \[\norm{uu^\top - vv^\top}_{\text{op}} \le \norm{u-v}_2 \norm{u}_2 + \norm{v}_2 \norm{u-v}_2 \le 2\varepsilon.\] Since each \(x_m x_m^\top\succeq 0\) and \(\norm{x_mx_m^\top}_{\text{op}}\le C_x^2\), we have \[\norm{S_n}_{\text{op}} \le \sum_{m=1}^n \norm{x_m x_m^\top}_{\text{op}} \le n C_x^2.\] Together, we have \[\frac{1}{n}|u^\top S_n u - v^\top S_n v| \le 2 C_x^2 \varepsilon.\] Choose \(\varepsilon\) small such that \(2\varepsilon C_x^2 \le \delta/2\). Then, for all \(n\ge N_0(\omega)\) and all unit vectors \(u\in\mathbb{S}^{d-1}\), \[\frac{1}{n}u^\top S_n u \ge \frac{1}{n}v^\top S_n v - \frac{\delta}{2} \ge \left(C_M-\frac{\delta}{2}\right) - \frac{\delta}{2} = C_M-\delta,\] i.e., \[S_n \succeq (C_M - \delta) n I_d, \quad \text{and} \quad M_n \succeq (C_M - \delta) I_d.\] For sufficiently large \(n\), \(S_n\) is invertible and \(S_n^{-1} \preceq \frac{1}{(C_M - \delta) n} I_d\), so for any \(x\) with \(\norm{x}_2 \le C_x\), \[x^\top S_n^{-1}x \le \frac{C_x^2}{(C_M-\delta)n}, \quad \text{and} \quad x^\top M_n^{-1} x \le \frac{C_x^2}{C_M - \delta}.\] ◻

Lemma 11 (Non-homogeneous linear recursion). For \(a, b > 0\), consider the non-homogeneous linear recursion: \[x_{n+1} \le x_n - \frac{a}{n+1} x_n + \frac{b}{n(n+1)}, \quad x_1 > 0.\] Then, \[x_n = \begin{cases} O(n^{-a}), & 0 < a < 1; \\ O \left(\frac{\log n}{n}\right), & a = 1; \\ O(n^{-1}), & a > 1. \end{cases}\]

Proof of Lemma 11. By standard linear recursion techniques, for \(n \ge 2\) we have \[x_n \le x_1 \prod_{m=1}^{n-1} \left(1 - \frac{a}{m+1}\right) + b \sum_{k=1}^{n-1} \frac{1}{k(k+1)} \prod_{m=k+1}^{n-1} \left(1 - \frac{a}{m+1}\right).\] For the first term, we have \[\prod_{m=1}^{n-1} \left(1 - \frac{a}{m+1}\right) = \prod_{m=2}^{n} \frac{m - a}{m} = \frac{\Gamma(n - a + 1)}{\Gamma(2 - a) \Gamma(n + 1)} \asymp n^{-a}.\] Similarly, \[\prod_{m=k+1}^{n-1} \left(1 - \frac{a}{m+1}\right) = \prod_{m=k+2}^{n} \frac{m - a}{m} = \frac{\Gamma(n - a + 1) \Gamma(k+2)}{\Gamma(k + 2 - a) \Gamma(n + 1)} \asymp \left(\frac{k+1}{n}\right)^{a}.\] Together, we have \[x_n \le C_1 n^{-a} + C_2 n^{-a} \sum_{k=1}^{n-1} \frac{(k+1)^{a}}{k(k+1)} \asymp n^{-a} + n^{-a} \sum_{k=1}^{n-1} (k+1)^{a-2}.\] Here, \[\sum_{k=1}^{n-1} (k+1)^{a-2} = \begin{cases} O(1), & 0 < a < 1; \\ O(\log n), & a = 1; \\ O(n^{a-1}), & a > 1. \end{cases}\] The results follow. ◻

Theorem 7 (Matrix Freedman ([58] Theorem 1.2)). Consider a matrix martingale \((Y_n, \mathcal{F}_n)_{n\ge 0}\) whose values are self-adjoint matrices with dimension \(d\), and let \((X_n)_{n\ge 1}\) be the associated difference sequence. Assume that the difference sequence is uniformly bounded in the sense that there exists \(R > 0\) such that \[\lambda_{\max}(X_n) \le R \quad \text{almost surely for all } n \ge 1.\] Define the predictable quadratic variation process of the martingale: \[W_n = \sum_{k=1}^n \mathbb{E}[X_k^2 \mid \mathcal{F}_{k-1}], \quad n \ge 1.\] Then, for all \(t \ge 0\) and \(\sigma^2 > 0\), \[\mathbb{P}\left(\exists n \ge 0: \lambda_{\max}(Y_n) \ge t, \; \lambda_{\max}(W_n) \le \sigma^2\right) \le d \cdot \exp\left(-\frac{t^2/2}{\sigma^2 + Rt/3}\right).\]

Lemma 12 (Minimum eigenvalue of a diagonal-plus-rank-one matrix). Let \(b \in \mathbb{R}^d\), \(S = \mathrm{diag}(s_1, ..., s_d)\) with \(s_i > 0\) for all \(i\), and define \[K = n(1, b)^\top (1, b) + \begin{pmatrix} 0 & 0 \\ 0 & S \end{pmatrix}.\] Then, \[\lambda_{\min}(K) \ge \frac{1}{2} \min \left\{\frac{n}{1 + n b^\top S^{-1} b}, \min_{1 \le i \le d} s_i\right\}.\]

Proof of Lemma 12. Let \(x = (t, u) \in \mathbb{R}^{1+d}\). Then, \[x^\top K x = n(t + b^\top u)^2 + u^\top S u.\] Also, \(\norm{x}_2^2 = t^2 + \norm{u}_2^2\). So, the Rayleigh quotient is \[\frac{x^\top K x}{\norm{x}_2^2} = \frac{n(t + b^\top u)^2 + u^\top S u}{t^2 + \norm{u}_2^2}\] and \(\lambda_{\min}(K) = \inf_{x \ne 0} \frac{x^\top K x}{\norm{x}_2^2}\). Since \(S \succeq s_{\min} I_d\) where \(s_{\min} = \min_{1 \le i \le d} s_i\), we have \[u^\top S u \ge s_{\min} \norm{u}_2^2.\] The tricky part is to lower bound the term \(n(t + b^\top u)^2\). Note that \[K = \begin{pmatrix} n & n b^\top \\ n b & S + n b b^\top \end{pmatrix}.\] Let \[A = n, \quad B = n b, \quad D = S + n b b^\top.\] Then, we can consider a block form \[K = \begin{pmatrix} A & B^\top \\ B & D \end{pmatrix}.\] We can calculate \[\begin{align} \begin{pmatrix} t \\ u \end{pmatrix}^\top K \begin{pmatrix} t \\ u \end{pmatrix} &= t^2 A + 2 t B^\top u + u^\top D u \\ &= u^\top D u + 2t B^\top u + t^2 B^\top D^{-1} B + t^2 A - t^2 B^\top D^{-1} B \\ &= (u + t D^{-1} B)^\top D (u + t D^{-1} B) + t^2 (A - B^\top D^{-1} B). \end{align}\] Since \(S \succ 0\), \(D \succ 0\), so the first term is non-negative, and \[x^\top K x \ge t^2 (A - B^\top D^{-1} B) = t^2 \left(n - n^2 b^\top (S + n b b^\top)^{-1} b\right).\] We need to upper bound \(b^\top (S + n b b^\top)^{-1} b\). Recall the Woodbury matrix identity: \[(A + U C V)^{-1} = A^{-1} - A^{-1} U (C^{-1} + V A^{-1} U)^{-1} V A^{-1}.\] Take \(A = S\), \(U = b\), \(C = n\), \(V = b^\top\). Then, \[(C^{-1} + V A^{-1} U)^{-1} = \left(\frac{1}{n} + b^\top S^{-1} b\right)^{-1} = \frac{n}{1 + n b^\top S^{-1} b}.\] So, \[(S + n b b^\top)^{-1} = S^{-1} - S^{-1} b \cdot \frac{n}{1 + n b^\top S^{-1} b} \cdot b^\top S^{-1},\] and \[b^\top (S + n b b^\top)^{-1} b = b^\top S^{-1} b - b^\top S^{-1} b \cdot \frac{n}{1 + n b^\top S^{-1} b} \cdot b^\top S^{-1} b.\] Denote \(\alpha = b^\top S^{-1} b\). Then, \[b^\top (S + n b b^\top)^{-1} b = \alpha - \frac{n \alpha^2}{1 + n \alpha} = \frac{\alpha}{1 + n \alpha}.\] So, \[x^\top K x \ge t^2 \left(n - n^2 \cdot \frac{\alpha}{1 + n \alpha}\right) = \frac{n t^2 }{1 + n b^\top S^{-1} b}.\] Combining the two lower bounds, we have \[2 x^\top K x \ge \frac{n}{1 + n b^\top S^{-1} b} \cdot t^2 + s_{\min} \norm{u}_2^2.\] Since \(\frac{n}{1 + n b^\top S^{-1} b} \ge 0\) and \(s_{\min} > 0\), we have \[2 x^\top K x \ge \min \left\{\frac{n}{1 + n b^\top S^{-1} b}, s_{\min}\right\} (t^2 + \norm{u}_2^2) = \min \left\{\frac{n}{1 + n b^\top S^{-1} b}, s_{\min}\right\} \norm{x}_2^2.\] So, for any \(x \ne 0\), \[\frac{x^\top K x}{\norm{x}_2^2} \ge \frac{1}{2} \min \left\{\frac{n}{1 + n b^\top S^{-1} b}, s_{\min}\right\},\] and taking the infimum gives the desired result. ◻

References↩︎

[1]
A. V. den Boer, “Dynamic pricing and learning: Historical origins, current research, and new directions,” Surveys in Operations Research and Management Science, vol. 20, no. 1, pp. 1–18, 2015.
[2]
W. L. Cooper, T. Homem-de-Mello, and A. J. Kleywegt, “Learning and pricing with models that do not explicitly incorporate competition,” Operations Research, vol. 63, no. 1, pp. 86–103, Feb. 2015.
[3]
B. Salcedo, “Pricing algorithms and tacit collusion.” https://brunosalcedo.com/docs/collusion.pdf, 2015.
[4]
E. Calvano, G. Calzolari, V. Denicolo, and S. Pastorello, “Artificial intelligence, algorithmic pricing, and collusion,” American Economic Review, vol. 110, no. 10, pp. 3267–3297, 2020.
[5]
OECD, “Algorithms and collusion: Competition policy in the digital age,” OECD Roundtables on Competition Policy Papers, vol. 206, 2017, doi: 10.1787/258dcb14-en.
[6]
L. Calzolari, “The misleading consequences of comparing algorithmic and tacit collusion: Tackling algorithmic concerted practices under art. 101 TFEU,” European Papers-A Journal on Law and Integration, vol. 2021, no. 2, pp. 1193–1228, 2021.
[7]
A. Klobuchar, “Klobuchar, colleagues introduce antitrust legislation to prevent algorithmic price-fixing.” https://www.klobuchar.senate.gov/public/index.cfm/2024/2/klobuchar-colleagues-introduce-antitrust-legislation-to-prevent-algorithmic-price-fixing, 2024.
[8]
O. of P. A. U.S. Department of Justice, “Justice department sues RealPage for algorithmic pricing scheme that harms millions of american renters.” https://www.justice.gov/archives/opa/pr/justice-department-sues-realpage-algorithmic-pricing-scheme-harms-millions-american-renters, 2024.
[9]
A. V. den Boer and J. M. Meylahn, “A (mathematical) definition of algorithmic collusion,” Available at SSRN 5012923, 2024.
[10]
I. Abada, J. E. Harrington Jr, X. Lambin, and J. M. Meylahn, “Algorithmic collusion: Where are we and where should we be going?” Available at SSRN 4891033, 2025.
[11]
K. T. Hansen, K. Misra, and M. M. Pai, “Frontiers: Algorithmic collusion: Supra-competitive prices via independent algorithms,” Marketing Science, vol. 40, no. 1, pp. 1–12, 2021.
[12]
C. Douglas, F. Provost, and A. Sundararajan, “Naive algorithmic collusion: When do bandit learners cooperate and when do they compete?” in ICIS 2024 proceedings, 2024, p. 15.
[13]
N. B. Keskin and A. Zeevi, “Dynamic pricing with an unknown demand model: Asymptotically optimal semi-myopic policies,” Operations research, vol. 62, no. 5, pp. 1142–1167, 2014.
[14]
S. Fish, Y. A. Gonczarowski, and R. I. Shorrer, “Algorithmic collusion by large language models,” arXiv preprint arXiv:2404.00806, vol. 7, no. 2, p. 5, 2024.
[15]
M. Banchio and G. Mantegazza, “Artificial intelligence and spontaneous collusion,” arXiv preprint arXiv:2202.05946, 2022.
[16]
I. Abada and X. Lambin, “Artificial intelligence: Can seemingly collusive outcomes be avoided?” Management Science, vol. 69, no. 9, pp. 5042–5065, 2023.
[17]
[18]
I. Abada, X. Lambin, and N. Tchakarov, “Collusion by mistake: Does algorithmic sophistication drive supra-competitive profits?” European Journal of Operational Research, vol. 318, no. 3, pp. 927–953, 2024.
[19]
X. Lambin, “Less than meets the eye: Simultaneous experiments as a source of algorithmic seeming collusion,” Available at SSRN 4498926, 2024.
[20]
Q. Bertrand, J. A. Duque, E. Calvano, and G. Gidel, “Self-play \(Q\)-learners can provably collude in the iterated prisoner’s dilemma,” in Proceedings of the 42nd international conference on machine learning, 2025, vol. 267, pp. 3952–3975.
[21]
Á. Cartea, P. Chang, J. Penalva, and H. Waldon, “Algorithmic collusion and a folk theorem from learning with bounded rationality,” Available at SSRN 4293831, 2022.
[22]
Z. Y. Brown and A. MacKay, “Competition in pricing algorithms,” American Economic Journal: Microeconomics, vol. 15, no. 2, pp. 109–156, 2023.
[23]
E. R. Arunachaleswaran, N. Collina, S. Kannan, A. Roth, and J. Ziani, “Algorithmic collusion without threats,” arXiv preprint arXiv:2409.03956, 2024.
[24]
M. Bichler, J. Durmann, and M. Oberlechner, “Algorithmic pricing and algorithmic collusion,” Business & Information Systems Engineering, vol. 67, no. 6, pp. 971–979, 2025.
[25]
A. Aouad and A. V. den Boer, “Algorithmic collusion in assortment games,” Available at SSRN 3930364, 2021.
[26]
J. M. Meylahn and A. V. den Boer, “Learning to collude in a pricing duopoly,” Manufacturing & Service Operations Management, vol. 24, no. 5, pp. 2577–2594, 2022.
[27]
T. Loots and A. V. den Boer, “Data-driven collusion and competition in a pricing duopoly with multinomial logit demand,” Production and Operations Management, vol. 32, no. 4, pp. 1169–1186, 2023.
[28]
L. Musolff, “Algorithmic pricing facilitates tacit collusion: Evidence from e-commerce,” in Proceedings of the 23rd ACM conference on economics and computation, 2022, pp. 32–33.
[29]
S. Assad, R. Clark, D. Ershov, and L. Xu, “Algorithmic pricing and competition: Empirical evidence from the german retail gasoline market,” Journal of Political Economy, vol. 132, no. 3, pp. 723–771, 2024.
[30]
Y. Yang, Y.-C. Lee, and P.-A. Chen, “Competitive demand learning: A noncooperative pricing algorithm with coordinated price experimentation,” Production and Operations Management, vol. 33, no. 1, pp. 48–68, 2024.
[31]
M. Chen and Z.-L. Chen, “Recent developments in dynamic pricing research: Multiple products, competition, and limited demand information,” Production and Operations Management, vol. 24, no. 5, pp. 704–731, May 2015.
[32]
O. Besbes, Y. Gur, and A. Zeevi, “Non-stationary stochastic optimization,” Operations Research, vol. 63, no. 5, pp. 1227–1244, Oct. 2015.
[33]
P. Milgrom and J. Roberts, “Rationalizability, learning, and equilibrium in games with strategic complementarities,” Econometrica: Journal of the Econometric Society, pp. 1255–1277, 1990.
[34]
P. Milgrom and J. Roberts, “Adaptive and sophisticated learning in normal form games,” Games and economic Behavior, vol. 3, no. 1, pp. 82–100, 1991.
[35]
O. Besbes and A. Zeevi, “Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms,” Operations research, vol. 57, no. 6, pp. 1407–1420, 2009.
[36]
V. Goyal, S. Li, and S. Mehrotra, “Learning to price under competition for multinomial logit demand,” Available at SSRN 4572453, 2023.
[37]
S. Li and S. Mehrotra, “Adaptive learning in uncertain and sequential competition,” Operations Research, vol. 74, no. 1, pp. 301–338, 2026.
[38]
R. H. Berk, “Limiting behavior of posterior distributions when the model is incorrect,” The Annals of Mathematical Statistics, vol. 37, no. 1, pp. 51–58, 1966.
[39]
H. White, “Maximum likelihood estimation of misspecified models,” Econometrica, vol. 50, no. 1, pp. 1–25, 1982.
[40]
I. Esponda and D. Pouzo, Berk–Nash equilibrium: A framework for modeling agents with misspecified models,” Econometrica, vol. 84, no. 3, pp. 1093–1130, 2016.
[41]
O. Besbes and A. Zeevi, “On the surprising sufficiency of linear models for dynamic pricing with demand learning,” Management Science, vol. 61, no. 4, pp. 723–739, 2015.
[42]
M. Nambiar, D. Simchi-Levi, and H. Wang, “Dynamic learning and pricing with model misspecification,” Management Science, vol. 65, no. 11, pp. 4980–5000, 2019.
[43]
A. V. den Boer and B. Zwart, “Simultaneously learning and optimizing using controlled variance pricing,” Management Science, vol. 60, no. 3, pp. 770–783, 2014.
[44]
N. Singh and X. Vives, “Price and quantity competition in a differentiated duopoly,” The RAND Journal of Economics, vol. 15, no. 4, pp. 546–554, 1984.
[45]
X. Vives, Oligopoly pricing: Old ideas and new tools. Cambridge, MA: MIT Press, 1999.
[46]
P. Choné and L. Linnemer, “Linear demand systems for differentiated goods: Overview and user’s guide,” International Journal of Industrial Organization, vol. 73, p. 102663, 2020, doi: 10.1016/j.ijindorg.2020.102663.
[47]
M. J. Osborne and C. Pitchik, “Profit-sharing in a collusive industry,” European Economic Review, vol. 22, no. 1, pp. 59–74, 1983, doi: 10.1016/0014-2921(83)90089-2.
[48]
C. Fischer and H.-T. Normann, “Collusion and bargaining in asymmetric Cournot duopoly—an experiment,” European Economic Review, vol. 111, pp. 360–379, 2019, doi: 10.1016/j.euroecorev.2018.10.005.
[49]
M. Stern and J. R. Birge, “Dynamic learning in strategic pricing games,” Available at SSRN 3579123, 2020.
[50]
S. Li, C. Shi, and S. Mehrotra, “Lego: Optimal online learning under sequential price competition,” University of Miami Business School Research Paper, no. 4803002, 2024.
[51]
N. B. Keskin and A. Zeevi, “On incomplete learning and certainty-equivalence control,” Operations Research, vol. 66, no. 4, pp. 1136–1167, 2018.
[52]
A. Goldenshluger and A. Zeevi, “A linear response bandit problem,” Stochastic Systems, vol. 3, no. 1, pp. 230–261, 2013.
[53]
H. Bastani, M. Bayati, and K. Khosravi, “Mostly exploration-free algorithms for contextual bandits,” Management Science, vol. 67, no. 3, pp. 1329–1349, 2021.
[54]
V. S. Borkar, Stochastic approximation: A dynamical systems viewpoint, vol. 48. Springer Nature Singapore, 2023.
[55]
D. Williams, Probability with martingales. Cambridge University Press, 1991.
[56]
D. A. Freedman, On Tail Probabilities for Martingales,” The Annals of Probability, vol. 3, no. 1, pp. 100–118, 1975.
[57]
H. Robbins and D. Siegmund, “A convergence theorem for non negative almost supermartingales and some applications,” in Optimizing methods in statistics, Elsevier, 1971, pp. 233–257.
[58]
J. A. Tropp, “Freedman’s inequality for matrix martingales,” arXiv preprint arXiv:1101.3039, 2011.

  1. The strengthening is not invoked in any of our proofs and all formal results carry over verbatim under the weaker requirement \(H\prec 0\). When cross-price effects are symmetric (\(\gamma_{i,j}=\gamma_{j,i}\) for all \(i\neq j\)), the strengthening reduces to the standing condition \(-\beta_i>\gamma_i\).↩︎

  2. In asymmetric markets, the joint-revenue-maximizing prices may disproportionately benefit some sellers, and richer notions of collusion additionally specify how gains are divided [47], [48]. Since any such outcome is bounded above by \(\mathbf{p}^{C}\), our choice of benchmark keeps the exposition straightforward and leaves the separate question of how collusive revenues are distributed to future work.↩︎

  3. Code for all numerical experiments is available at https://github.com/yw3453/ob-learn.↩︎

  4. The \(\Theta(1)\) spectral lower bound ?? would be satisfied automatically if \(x_{n,i}^{ob}\) are exogenous i.i.d.designs.↩︎

  5. Condition (ii) is a prerequisite for (iv) in the proof, so (iv) is well-posed only when (ii) holds.↩︎

  6. All numerical experiments in this paper use uniform dithering \(z_{n,i}\sim\mathrm{Unif}(-c_n, c_n)\) with \(c_n\) eventually small enough that 10 holds. The convergence theorems of Sections 5 and 6 use only conditional first and second moments of \(z_{n,i}\) and hence carry through verbatim under 10 .↩︎

  7. While the actual estimates \(\hat{\theta}_{n,i}^{ob}\) are projected onto \(\Theta_i^{ob}\), we derive the ODE without projection because it is inactive locally around Nash.↩︎

  8. This is the continuous-time analogue of requiring that the initial design is not degenerate (e.g., the first two prices are not identical) so that the OLS slope is well-defined.↩︎