We show that a deep neural network (DNN) trained to construct a stochastic discount factor (SDF) admits an additive decomposition separating nonlinear characteristic discovery from the pricing rule that aggregates them. This decomposition yields a
linear factor representation governed by the Portfolio Tangent Kernel (PTK), which summarizes the network’s learned features. In population, the implied SDF converges to a ridge-regularized version of the true SDF, with the degree of regularization
determined by spectral complexity. Empirically, using U.S. equity data, the PTK representation delivers economically and statistically significant performance gains, while rising spectral complexity imposes tighter limits on finite-sample pricing.
Modern asset pricing increasingly relies on high-dimensional data and flexible statistical methods. A growing literature shows that deep neural networks (DNNs) can construct stochastic discount factors (SDFs), return forecasts, and portfolio strategies
that outperform traditional linear factor models. These empirical successes raise a fundamental but unresolved question: what, exactly, do neural networks learn when trained on financial data?
From an economic perspective, training a neural network implicitly performs two distinct tasks. First, the network learns what to price: nonlinear features constructed from firm characteristics that summarize economically relevant information.
Second, it learns how to price: a rule that aggregates those features into an SDF. In standard implementations, these two tasks are learned jointly through gradient descent and are therefore tightly entangled. As a result, the network must
simultaneously discover informative features and estimate their optimal combination from finite samples. This joint learning problem obscures economic interpretation and can lead to statistical inefficiencies, reinforcing the perception of neural networks
as black boxes that are difficult to discipline using asset pricing theory.
This paper shows that DNN-based SDFs admit a sharp additive decomposition that separates feature learning from pricing. For networks trained by gradient descent, the estimated SDF can be written as the sum of two components: a feature-based term that
depends only on the nonlinear characteristics learned by the network, and a residual model-dependent term. We show that the feature-based component admits a closed-form representation governed by a new object, the Portfolio Tangent Kernel
(PTK).
Conditional on the learned characteristics, the PTK induces a unique linear pricing rule. The resulting SDF—what we call the PTK-SDF—is the Markowitz portfolio of a very large collection of characteristic-managed factors. This representation
takes the form of a Large Factor Model (LFM). Because this pricing rule is optimal conditional on the learned features, the residual model-dependent component of the original DNN does not improve performance and can be discarded. In this sense, neural
networks are best understood as powerful devices for feature discovery, while the PTK provides a cleaner, transparent, and statistically efficient way to price those features.
Empirically, this distinction is decisive. Using U.S.equity data, we show that the PTK-SDF systematically outperforms the original DNN-SDF, delivering higher Sharpe ratios and economically significant alphas across asset size groups. Conditioning on
learned features while re-optimizing the pricing rule mitigates learning-induced noise that arises when feature discovery and pricing are conducted jointly. This leads to a modular estimation pipeline: a neural network is used solely to learn nonlinear
characteristics; these characteristics are mapped into a large set of characteristic-based factors; and the factors are combined using an explicit, regularized portfolio rule. The PTK-SDF therefore isolates—and improves upon—the economically relevant
content of the original network.
The closed-form PTK representation also allows us to characterize the statistical limits of DNN-based asset pricing models. Because the number of learned factors is extremely large relative to available time-series observations, classical inference
breaks down. Building on recent results in random matrix theory, we show that performance is governed by two forces. The first is alignment: the extent to which risk premia load on statistically strong directions of factor risk. The second is
spectral complexity: the dispersion of factor risk across principal components, which determines the effective degree of regularization faced by the investor. Higher spectral complexity corresponds to a flatter risk spectrum, in which economically
relevant variation is spread across many weak directions. In finite samples, this dispersion amplifies estimation error and limits achievable Sharpe ratios.
We document substantial time variation in the spectral complexity of the PTK. Over the past two decades, it has increased by roughly a factor of six, indicating a pronounced rise in the statistical complexity of the SDF. This increase coincides with the
well-documented decline in factor model performance since the early 2000s. At the same time, alignment between risk premia and factor risk has improved, partially offsetting the adverse effects of rising complexity. These findings highlight a fundamental
trade-off: feature learning improves economic alignment, but rising spectral complexity imposes limits on what can be learned from finite samples.
An important implication of this framework is that feature learning fundamentally reshapes regularization. We find that factors constructed from learned features exhibit substantially higher spectral complexity than factors constructed from random
features, such as those studied by [1]. Quantitatively, the spectral complexity of learned-feature-based factors is roughly an order of
magnitude larger. As a result, these factors are subject to strong implicit regularization even in the absence of explicit ridge penalties. In particular, the PTK-SDF achieves its best performance with little or no additional shrinkage: despite
involving hundreds of thousands of factors, the optimal pricing rule is effectively ridgeless.2
Beyond statistical performance, the PTK representation restores economic discipline. We show that PTK-based SDFs exhibit substantially stronger alignment with long-horizon and future consumption risk—both in the sense of [2] and its micro-founded extension in [3]—than either fully trained DNN-SDFs or random-feature benchmarks.
Taken together, our findings provide a unified framework for understanding deep learning in asset pricing. Neural networks learn economically informative features, but pricing those features efficiently requires separating feature discovery from model
estimation. The Portfolio Tangent Kernel performs this separation by replacing the original DNN with an explicit large-factor pricing representation. This representation allows both the economic content of learned features and the statistical limits
imposed by finite samples to be analyzed using modern asset pricing tools rooted in high-dimensional statistics and random matrix theory.
A central theme in asset pricing is the tension between sparsity and complexity in the construction of stochastic discount factors (SDFs). While classical factor models emphasize parsimony, a large literature documents that richly parameterized models
substantially outperform sparse specifications out of sample [1], [4]–[9]. See [10] for an overview.
These findings motivate Large Factor Models (LFMs), which construct SDFs from extensive collections of characteristic-based factors. With appropriate shrinkage, LFMs perform well even when based on simple linear characteristics [11], [12]. At the same time, approximating the true conditional SDF may require an
extremely large number of factors. [1] show that, in the limit of infinitely many factors, the unconditionally optimal factor portfolio
converges to the true conditional SDF, despite the latter being infeasible to estimate directly.
The use of highly over-parameterized factor models raises natural statistical concerns. Estimation error becomes severe when the number of factors exceeds the sample size, and classical inference breaks down. Nevertheless, [1], building on the virtue-of-complexity framework of [13], show that extreme dimensionality can improve out-of-sample performance. A key insight is that regularization (in the form of effective shrinkage) arises endogenously. [1], [14], and [15] show that effective shrinkage is determined by the spectral properties of covariance matrices, while [15] emphasize the role of alignment between factor risk premia and factor risk in determining achievable Sharpe ratios.
These insights motivate new tools for inference. When the number of factors exceeds the sample size, the classical [16] statistic
diverges. [15] propose a debiased GRS statistic that converges to a ridge-penalized population Sharpe ratio and provides a direct measure of
alignment. We adopt this statistic in our empirical analysis.
Most existing Large Factor Models rely on linear factors or random nonlinear transformations. In particular, the model studied by [1] is
equivalent to a shallow but wide neural network with untrained hidden-layer weights. In contrast, much of the machine-learning-based asset pricing literature employs fully trained, multi-layer networks whose economic content and statistical properties are
difficult to analyze theoretically [17]–[19].
This paper bridges these strands by distinguishing between feature learning and pricing in deep learning–based asset pricing models. We study fully trained neural networks of arbitrary depth and width that remain analytically tractable using Neural
Tangent Kernel methods [20], [21]. We show that, while the DNN learns
economically informative nonlinear features, the associated pricing rule need not be optimal. Instead, the learned features admit a closed-form Large Factor Model representation governed by a novel kernel, the Portfolio Tangent Kernel (PTK), which delivers
an optimal pricing rule conditional on those features. This perspective clarifies how feature learning reshapes spectral complexity, implicit regularization, and alignment, and connects deep learning models to the emerging theory of high-dimensional asset
pricing.
We consider a panel of stocks indexed by \(i = 1, \ldots, N_t\), with excess returns \[R_{t+1} = \left( R_{i,t+1} \right)_{i=1}^{N_t}\] observed at time \(t+1\). Each stock \(i\) is associated with a \(d\)-dimensional vector of characteristics \[X_{i,t} = \left( X_{i,t}(k)
\right)_{k=1}^d,\] and we collect all characteristics in the matrix \[X_t = \bigl[ X_t(1), \ldots, X_t(d) \bigr] \in \mathbb{R}^{N_t \times d}.\] The true tradable stochastic discount factor (SDF) is given by \[M_{t+1}\;=\;1\;-\;\pi_t'R_{t+1}\,,\] where \(\pi_t\) denotes the conditionally efficient portfolio \[\label{mark1}
\pi_t\;=\;E_t[R_{t+1}R_{t+1}']^{-1}\,E_t[R_{t+1}]\,,\tag{1}\] which solves the utility maximization problem \[\label{q-obj1}
\max_{\pi_t}E_t[\pi_t'R_{t+1}\;-\;0.5 (\pi_t'R_{t+1})^2]\,.\tag{2}\] If \(X_t\) contains all relevant conditioning information, then \(\pi_t\) in 1 can be written as \[\label{mark2}
\pi_t\;=\;\pi(X_t)\,,\tag{3}\] for some unknown and potentially highly nonlinear function \(\pi(X)\,.\) Using the law of iterated expectations, the problem of finding the conditionally optimal portfolio policy
can then be rewritten as an unconditional, non-parametric utility maximization problem: \[\label{non-param}
\min_{\pi(\cdot)}E[(1\,-\,\pi(X_t)'R_{t+1})^2]\;=\;1\;-\;2\,\max_{\pi(\cdot)}E\big[U\big(\pi(X_t)'R_{t+1}\big)\big],\tag{4}\] where \(U(x)\;=\;x-0.5 x^2.\) A standard approach to solving 4 for \(\pi(X)\) is to posit a sufficiently rich parametric family of functions \(f(x;\theta)\) and to estimate the parameter vector \(\theta\) using sample analogs of the objective in 4 . Following [1], we consider a
ridge-penalized version of this objective, \[\label{main-1}
\min_{\theta}L(\theta),\;where\;L(\theta)\;=\;\Bigg(\frac{1}{T}\sum_{t=1}^T(1\;-\;f(X_t;\theta)'R_{t+1})^2\;+\;z\,\|\theta\|^2\Bigg)\,,\tag{5}\] where \(z\) denotes the ridge penalty and serves as a
regularization parameter.
Equation 5 highlights the central tension underlying high-dimensional SDF estimation. Richer function families reduce approximation bias by allowing rich functional forms for \(\pi(X)\), but they
simultaneously exacerbate estimation error when the dimensionality of \(\theta\) is large relative to the sample size. As a result, improvements in in-sample fit need not translate into better out-of-sample SDF performance.
Our contribution is to characterize how this bias–variance tradeoff manifests itself in MSRR-based SDF estimation and to show that, in finite samples, limits to learning impose sharp constraints on the achievable Sharpe ratio, even when the true SDF lies
within the span of the chosen function family.
A further complication arises from computation. When \(\theta\) is high-dimensional, solving 5 exactly is generally infeasible. In practice, estimation relies on iterative optimization algorithms
that converge to particular local minima. Recent work in machine learning emphasizes that the choice of optimization algorithm plays a central role in determining out-of-sample performance; see, for example, [22]. Thus, not only the function family \(f(x;\theta)\), but also the manner in which 5 is
solved, matters for the properties of the resulting SDF.
The empirical success of deep learning reflects the fact that certain function families—deep neural networks combined with gradient-based optimization—can achieve good generalization even in regimes where the number of parameters exceeds the available
sample size; see, for example, [23] and [24]. For asset pricing, this raises two distinct questions: what features of the data are learned by these models, and how those features are combined into an SDF. In the sections that follow, we show that addressing these
questions separately is key to understanding both the empirical performance and the statistical limits of neural-network-based SDFs.
Figure 1: The figure above shows \(f(x; \theta; W)\), a mathematical representation of a neural network with a single hidden layer, also known as a shallow network. The weights \(W \in {\mathbb{R}}^{d\times P}\) and \(\theta \in {\mathbb{R}}^{P}\) are randomly initialized. \(\phi: {\mathbb{R}}\rightarrow {\mathbb{R}}\) is an elementwise
non-linear activation function and the vector \(\phi(x'W)\) is referred to as random features.
Before turning to fully trained deep neural networks, it is useful to consider simpler function families that depend linearly on the parameter vector \(\theta\in {\mathbb{R}}^P.\) These models provide a transparent
benchmark in which features are fixed rather than learned, allowing for an explicit characterization of pricing, regularization, and asymptotic behavior. Following [1], we consider the family \[\label{dd1}
f(x;\theta)\;=\;f(x;\theta;W)\;=\;\sum_{k=1}^P \phi(x'W_k)\theta_k \,,\tag{6}\] where \(x \in {\mathbb{R}}^d\) denotes the vector of raw characteristics, \(W_k\in
{\mathbb{R}}^d\) are randomly drawn weights, and \(\phi(\cdot)\) is a nonlinear activation function. The nonlinear functions \(\phi(x'W_k)\) are referred to as random features. As
explained in [1], the specification in 6 is equivalent to a shallow, single-hidden-layer neural network in which the
output-layer weights \(\theta_k\) are estimated, while the hidden-layer weights \(W_k\) are held fixed and not optimized. Figure 1 provides a schematic
illustration.
The linearity of the function family in 6 implies that the optimization problem in 5 admits a unique closed-form solution. Moreover, this solution can be expressed in terms of characteristics-managed
portfolios, or factors. Specifically, defining the vector of factor returns \(F_{t+1}\;=\;(F_{k,t+1})_{k=1}^P\) with \[F_{k,t+1}\;=\;\sum_{i=1}^{N_t}\underbrace{\phi(X_{i,t}',W_k)}_{k\text{-}th\;random\;feature}\,R_{i,t+1}\,,\] we can rewrite the SDF as a factor portfolio, \[\label{sdf61fac}
\sum_{i=1}^{N_t} f(X_{i,t};\theta)R_{i,t+1}\;=\;\theta'F_{t+1}\,.\tag{7}\] Substituting 7 into 5 yields \[\label{msrr-th0}
\theta(z)\;=\;\arg\min_{\theta}L(\theta)\;=\;\left(zI+T^{-1}\sum_{t=1}^T F_tF_t'\right)^{-1}T^{-1}\sum_{t=1}^T F_t\,.\tag{8}\] Thus, within the linear family 6 , the estimated SDF admits a clear and familiar
interpretation as the mean–variance efficient portfolio of a (potentially very large) collection of factors.
Standard intuition based on the Arbitrage Pricing Theory of [25] suggests that the number of factors required to span the SDF should be small. However,
[1] show both theoretically and empirically that this intuition fails in high-dimensional settings. They establish the virtue of
complexity, whereby out-of-sample performance is monotonically increasing in the number of factors \(P\), even in the limit as \(P\to\infty\). This result naturally raises the question
of the limiting behavior of the estimated SDF as the number of random features grows. Addressing this question requires introducing tools from kernel theory.
A positive definite kernel on \({\mathbb{R}}^D\) is a function \(K(x,\tilde{x}):{\mathbb{R}}^D\times {\mathbb{R}}^D\to {\mathbb{R}}\) such that, for any collection of vectors \(x_i\in {\mathbb{R}}^D,i=1,\cdots,n,\) and any coefficient vector \(\alpha\in {\mathbb{R}}^n,\)\[\sum_{i_1=1}^n\sum_{i_2=1}^n\alpha_{i_1}\alpha_{i_2}K(x_{i_1},x_{i_2})\;\ge\;0\,.\] Equivalently, the kernel matrix \(\bar K\;=\;(K(x_{i_1},x_{i_2}))_{i_1,i_2=1}^n\) is positive semidefinite.
Given a collection of nonlinear functions \(\phi_k(x),\;k=1,\cdots,P,\) commonly referred to as features, one can define the associated kernel \[\label{ker61fea}
K(x,\tilde{x})\;=\;\frac{1}{P}\sum_{k=1}^P \phi_k(x)\phi_k(\tilde{x})\,.\tag{9}\] By direct calculation, \(K\) is positive semidefinite.3 Conversely, any positive definite kernel can be approximated by a sufficiently rich collection of nonlinear features: By Mercer’s theorem (see, e.g., [26], [27]), any positive definite kernel \(K(x,\tilde{x})\) admits the representation \[\label{bochner}
K(x,\tilde{x})\;=\;\int \phi(x;\omega)\,\phi(\tilde{x};\omega)\,p(d\omega)\,,\tag{10}\] for some probability distribution \(p(d\omega)\). Sampling \(W_k\) from \(p(d\omega)\) yields the Monte Carlo approximation \[\label{mercer}
K(x,\tilde{x})\;=\;\lim_{P\to\infty}\frac{1}{P} \sum_{k=1}^P \phi(x;W_k)\,\phi(\tilde{x};W_k)\,.\tag{11}\] This representation implies that, as \(P\to\infty\), the SDF based on the optimal \(\theta\) in 8 converges to a well-defined limit determined solely by the kernel \(K\). To characterize this limit in an asset-pricing setting, we introduce the
following definition.
Definition 1 (The Portfolio Kernel). Let \[\label{market-state}
Y_t\;=\;(R_{t+1},X_t)\;\in\;{\mathbb{R}}^{N_t(d+1)}\qquad{(1)}\] denote the market state, encompassing all stock characteristics and all stock returns. Given any two market states \(Y=(R,X),\;\tilde{Y}=(\tilde{R},\tilde{X}),\) with \(R\in {\mathbb{R}}^N,\;\tilde{R}\in {\mathbb{R}}^{\tilde{N}},\) we let \[K(X,\tilde{X})\;=\;(K(X_i,\tilde{X}_j))\;\in\;{\mathbb{R}}^{N\times \tilde{N}}\,,\] and define the Portfolio Kernel* as \[{\mathbb{K}}(Y;\tilde{Y})\;=\;\underbrace{R'}_{1\times
N}K(X,\tilde{X})\underbrace{\tilde{R}}_{\tilde{N}\times 1}\,.\]*
The Portfolio Kernel operates on market states and produces a kernel-based similarity measure \({\mathbb{K}}(Y;\tilde{Y})\). Importantly, this similarity measure is itself a portfolio return, as it aggregates asset
returns using weights determined by the similarity of characteristics across market states.
We denote the in-sample data by \[Y_{IS}\;=\;(R_{t+1}, X_t)_{t=1}^T \in {\mathbb{R}}^{T\times N_t (d+1)},\] and define the corresponding in-sample kernel matrix \[\label{k-is}
{\mathbb{K}}_{IS}\;=\;({\mathbb{K}}(Y_{t_1};Y_{t_2}))_{t_1,t_2=1}^T\;\in\;{\mathbb{R}}^{T\times T}\,,\tag{12}\] which measures similarity across in-sample market states.
Let \(K\) be a positive definite kernel admitting the representation 10 . Let \(\{W_k\}_{k=1}^P\) be independently sampled from \(p(d\omega)\) in 10 , and define the factor vector \(F_{t+1}=(F_{k,t+1})_{k=1}^P\) by \[F_{k,t+1}\;=\;\frac{1}{P^{1/2}}\sum_{i=1}^{N_t}\phi(X_{i,t};W_k)\,R_{i,t+1}\,.\] Equivalently, \[F_{t+1}\;=\;\frac{1}{P^{1/2}}\phi(X_t;W)R_{t+1}\;\in\;{\mathbb{R}}^P\,.\] Let
\[\label{msrr-th}
\theta(z;T)\;=\;\left(zI+T^{-1}\sum_{t=1}^T F_tF_t'\right)^{-1}T^{-1}\sum_{t=1}^T F_t\tag{13}\] denote the ridge-penalized efficient factor portfolio, and let \(\theta'F_{t+1}\) be the corresponding
SDF. By direct calculation, \[\label{lim-1}
\begin{align}
F_{t+1}'F_{\tau+1}
&=\;\frac{1}{P} R_{t+1}'\phi(X_t;W)\phi(X_\tau;W)'R_{\tau+1}\\
&=\;R_{t+1}'\left(\frac{1}{P}\sum_{k=1}^P \phi(X_t;W_k)\phi(X_\tau;W_k)'\right)R_{\tau+1}\\
&\;\underbrace{\to}_{\eqref{mercer}}\;R_{t+1}'K(X_t,X_\tau)R_{\tau+1}\;=\;{\mathbb{K}}(Y_t,Y_\tau)\,.
\end{align}\tag{14}\] Denoting by \(F=(F_t)_{t=1}^T\in {\mathbb{R}}^{P\times T}\) the matrix of in-sample factor returns, it follows that \[\label{lim-2}
F'F\;=\;(F_{t_1}'F_{t_2})_{t_1,t_2=1}^T\;\approx\;({\mathbb{K}}(Y_{t_1},Y_{t_2}))_{t_1,t_2=1}^T\;=\;{\mathbb{K}}_{IS}\,.\tag{15}\] We will require the following technical lemma.
Lemma 1. Let \(F\in {\mathbb{R}}^{P\times T}\) denote the matrix of in-sample factor returns, and let \({\boldsymbol{1}}=(1,\cdots,1)\in {\mathbb{R}}^T\) be the
vector of ones. Then, \[\theta\;=\;T^{-1}(zI+T^{-1}FF')^{-1}F{\boldsymbol{1}}\;=\;T^{-1}F(zI+T^{-1}F'F)^{-1}{\boldsymbol{1}}\,.\]
Lemma 1 follows directly from the regression objective 5 : \(\theta\) is the coefficient vector obtained by regressing
\({\boldsymbol{1}}\) on factor returns. Consequently, \[\begin{align}
F_{T+1}'\theta
&=\;T^{-1}\underbrace{F_{T+1}'F}_{\to\;{\mathbb{K}}(Y_{T+1},Y_{IS})}
\left(zI+T^{-1}\underbrace{F'F}_{\to\;{\mathbb{K}}_{IS}}\right)^{-1}{\boldsymbol{1}}\,,
\end{align}\] and we obtain the following result.
Theorem 1 (LFM-SDF). The Large Factor Model (LFM) SDF is given by \[R^{LFM}(z)_{T+1}\;=\;\lim_{P\to\infty}\theta'F_{T+1}\;=\;{\mathbb{K}}(Y_{T+1},Y_{IS})\xi\;=\;\sum_{t=1}^T
\underbrace{{\mathbb{K}}(Y_{T+1},Y_t)}_{attention\;to\;t}\;\underbrace{\xi_t}_{optimal\;weight}\,,\] where \[\xi\;=\;\frac{1}{T} (zI+T^{-1}{\mathbb{K}}_{IS})^{-1}{\boldsymbol{1}}\,.\]
Theorem 1 characterizes the asymptotic behavior of the random-feature-based Large Factor Model studied in [1]. In the limit of infinitely many random features, the estimated SDF converges to a well-defined object governed entirely by the kernel \(K\).
The LFM-SDF compares the current market state \(Y_{T+1}\) with past realizations \(Y_t,\;t=1,\cdots,T,\) using the Portfolio Kernel \({\mathbb{K}}\).
Periods that are more similar to the present receive higher weight, while less similar periods receive lower weight. These similarity scores are then aggregated in a mean-variance efficient manner through the weight vector \(\xi\in {\mathbb{R}}^T\), which can be interpreted as optimal time fixed effects.
Although the kernel representation emphasizes aggregation over time, the LFM-SDF also admits a conventional portfolio representation in terms of characteristics-managed factors.
Corollary 1. We have \[\lim_{P\to\infty}\theta'F_{T+1}\;=\;R_{T+1}'f^{\mathbb{K}}(X_t)\,,\] where \[f^{\mathbb{K}}(x)\;=\;\sum_{t=1}^T\Big(\sum_{j=1}^{N_t}
R_{j,t+1} K(x, X_{j,t})\Big)\,\xi_t\,.\]
The representation of \(f^{\mathbb{K}}(x)\) provides a transparent interpretation of how the LFM extracts pricing information from the joint distribution of returns and characteristics. For any out-of-sample
characteristic vector \(x\), the model evaluates its similarity to past characteristics through the kernel \(K(x,X_{j,t})\) and forms the similarity-weighted performance measure \[Perf_t(x)\;=\;\sum_{j=1}^{N_t} R_{j,t+1} K(x, X_{j,t})\,.\] The final portfolio weight is obtained by optimally aggregating these performance measures across time, \[f^{\mathbb{K}}(x)\;=\;\sum_{t=1}^T Perf_t(x)\,\xi_t\,.\] In summary, once a kernel \(K\) is specified, the analysis delivers an explicit closed-form SDF that efficiently aggregates an
effectively infinite number of characteristics-based factors. This property motivates the term Large Factor Model. The analysis also highlights a limitation: the kernel \(K\) is exogenously chosen rather than learned. In
the next section, we show how fully trained neural networks endogenously learn such kernels through gradient-based optimization, leading to the Portfolio Tangent Kernel.
We consider a parametric function family \(f(X;\theta):\;{\mathbb{R}}^{N\times d}\to{\mathbb{R}}^N\) mapping characteristics \(X\in {\mathbb{R}}^{N\times d}\) to portfolio weights. For
example, an own-characteristics neural network \(g(X_{i,t};\theta)\) studied in [17], [18] implies \(f(X_t;\theta)\;=\;(g(X_{i,t};\theta))_{i=1}^N\), while a transformer architecture as in [28] yields a mapping that cannot be decoupled across stocks because each coordinate depends on the characteristics of all assets. Throughout, we refer to the function family as a DNN,
although it may originate from other families, such as panel trees [29].
4.1 Optimization, Initialization, and Gradient Flow↩︎
In practice, DNNs are trained by randomly initializing the parameter vector \(\theta=\theta(0)\) and then updating it using gradient-based optimization. This initialization is economically relevant: in over-parameterized
settings, many parameter vectors achieve identical in-sample fit, so the learned SDF depends not only on the objective function but also on the optimization path induced by the algorithm.
Letting \(\eta>0\) denote the learning rate, gradient descent updates the parameters according to \[\label{gr-d}
\theta(s+ds)\;=\;\theta(s) - \eta\, ds\, \nabla_{\theta} L(\theta(s))\,,\tag{16}\] where \(s\ge 0\) indexes (rescaled) training time, or equivalently, the cumulative number of gradient descent steps. In the
limit of an infinitesimal step size, this recursion converges to the gradient flow \[\label{n-ode}
\frac{d\theta(s)}{ds}\;=\;-\,\eta\, \nabla_{\theta} L(\theta(s))\,.\tag{17}\] The path \(\{\theta(s)\}_{s\ge 0}\) therefore defines a sequence of candidate pricing rules generated by the learning
algorithm.
While the dynamics of the high-dimensional parameter vector \(\theta(s)\) are generally intractable, the evolution of the portfolio return generated by the network admits a simpler characterization. Applying the chain
rule yields \[\label{neural-dyn}
\frac{d}{ds}(R'f(x;\theta(s)))\;=\;-\,\eta\, R'\nabla_\theta f(x;\theta(s))
\nabla_{\theta} L(\theta(s))\,.\tag{18}\] Equation 18 shows that gradient descent does not update the SDF directly. Instead, learning operates through the gradient feature portfolio returns \(R'\nabla_\theta f(x;\theta(s))\), which determine how pricing errors translate into changes in portfolio weights. From an asset pricing perspective, these gradient feature portfolios define a large collection of
characteristics-based payoffs that are implicitly traded during training. This observation motivates a kernel representation that measures similarity across market states through the interaction of returns with gradient features.
Definition 2 (The Portfolio Tangent Kernel (PTK)). Let \(R\in {\mathbb{R}}^N,\;\tilde{R}\in {\mathbb{R}}^{\tilde{N}}\) be two return vectors, and let \(X\in
{\mathbb{R}}^{N\times P},\;\tilde{X}\in {\mathbb{R}}^{\tilde{N}\times P}\) denote their characteristics. Define \[{\mathbb{K}}((R,X);(\tilde{R},\tilde{X});\theta)\;=\;
\underbrace{R'}_{1\times N}
\underbrace{\nabla_\theta f(X;\theta)}_{N\times P}
\underbrace{\nabla_\theta f(\tilde{X};\theta)'}_{P\times \tilde{N}}
\underbrace{\tilde{R}}_{\tilde{N}\times 1}\] to be the tangent kernel associated with the portfolio mapping \(f(X;\theta)\). We refer to this object as the Portfolio Tangent Kernel* (PTK).*
The PTK is a pricing-relevant kernel defined on market states. It measures similarity between states by aggregating returns using weights determined by the sensitivity of portfolio payoffs to parameter perturbations. By construction, the PTK is a
Portfolio Kernel in the sense of Definition 1, with the underlying kernel given by the Neural Tangent Kernel (NTK; [20]),4\[\label{ntk-def}
K(X,\tilde{X};\theta)\;=\;\nabla_\theta f(X;\theta)\,\nabla_\theta f(\tilde{X};\theta)'\,.\tag{19}\] Accordingly, \[{\mathbb{K}}((R,X);(\tilde{R},\tilde{X});\theta)\;=\;R'K(X,\tilde{X};\theta)\tilde{R}\,.\] While the NTK captures similarity across characteristics, the PTK lifts this structure to traded returns and therefore governs
pricing dynamics rather than prediction alone.
When \(\theta\) is high-dimensional, the PTK aggregates a large collection of characteristics-managed factors, \[\label{factor-ret}
F_{t+1}\;=\;(F_{k,t+1})_{k=1}^P,\qquad
F_{k,t+1}\;=\;R_{t+1}'\nabla_{\theta_k}f(X_t;\theta)\,.\tag{20}\] For any two market states \(Y_t,Y_\tau\), \[\label{fac-sim-ntk}
{\mathbb{K}}(Y_t;Y_\tau;\theta)\;=\;F_{t+1}'F_{\tau+1}\,,\tag{21}\] so that the PTK measures similarity across time in terms of the realized payoffs of these learned factors.
Equation 18 implies that, under gradient flow, the portfolio return generated by the DNN evolves according to \[\label{neural-dyn-factors}
\frac{d}{ds}\big(R'f(X;\theta(s))\big)
\;=\;-\,\eta\,F_{t+1}'\nabla_{\theta} L(\theta(s))\,.\tag{22}\] Specializing to the MSRR objective yields the following result.
Theorem 2 (MSRR Gradient Flow and the PTK). Suppose that \[\label{main-11}
L(\theta)\;=\;\frac{1}{T}\sum_{t=1}^T(1\;-\;f(X_t;\theta)'R_{t+1})^2\,.\qquad{(2)}\] Under gradient flow, \[\label{ptk-dyn}
\frac{d}{ds}\big(R'f(X;\theta(s))\big)
\;=\;-\,\eta \, T^{-1}\underbrace{{\mathbb{K}}(Y;Y_{IS};\theta(s))}_{PTK}\,
\underbrace{\big({\boldsymbol{1}}_T - R_{IS}'f(X_{IS};\theta(s))\big)}_{in\text{-}sample\;error}\,,\qquad{(3)}\] where \(Y_{IS}=(R_{IS},X_{IS})\) is in-sample (IS) data, and \[{\mathbb{K}}(Y;Y_{IS};\theta)\;=\;\big({\mathbb{K}}(Y;Y_t;\theta)\big)_{t=1}^T\in {\mathbb{R}}^{1\times T}\,.\]
Theorem 2 shows that gradient descent updates portfolio returns by comparing the current market state to past in-sample states through the PTK. Learning therefore operates through
two distinct channels:
Model learning, captured by the evolution of \(f(X;\theta(s))\), which determines the portfolio weight function; and
Feature learning, captured by the evolution of the PTK \({\mathbb{K}}(Y;Y_{IS};\theta(s))\), reflecting learning of the gradient features \(\nabla_\theta f(X;\theta)\) used
to construct the traded factors.
After sufficiently many gradient descent steps, these dynamics stabilize. This occurs, for example, when \(\theta(s)\) approaches a stationary point or when the network is sufficiently wide [20]. Once feature learning is complete, the PTK ceases to evolve, and the dynamics admit a closed-form solution.5
Theorem 3 (Closed-Form Stable-PTK DNN-SDFs). Fix a large \(s_*\) and define \[{\mathbb{K}}^*\;=\;{\mathbb{K}}(\cdot;\cdot;\theta(s_*))\,.\] Suppose that, after \(s_*\) gradient descent steps, the PTK stabilizes in the sense that \[\|{\mathbb{K}}(Y;Y_{IS};\theta(s))-{\mathbb{K}}^*(Y;Y_{IS})\|
+\|{\mathbb{K}}(Y_{IS};Y_{IS};\theta(s))-{\mathbb{K}}^*(Y_{IS};Y_{IS})\|
<\varepsilon\] for all \(s>s_*\). Suppose further that \[{\mathbb{K}}^*_{IS}\;\equiv\;{\mathbb{K}}^*(Y_{IS};Y_{IS})\] is non-degenerate. Then, for all \(s>s_*\), \[\label{decomposition-dnn}
R'f(X;\theta(s))\;=\;
\underbrace{R'f(X;\theta(s_*))}_{learned\;model}
\;+\;
\underbrace{\frac{1}{T}{\mathbb{K}}^*(Y,Y_{IS})\,\xi(s)}_{learned\;features}
\;+\;O(\varepsilon)\,,\qquad{(4)}\] where \[\xi(s)\;=\;
\Big(\tfrac1T{\mathbb{K}}^*_{IS}\Big)^{-1}
\Big(I-e^{-\eta (s-s_*) \tfrac1T{\mathbb{K}}^*_{IS}}\Big)
\big({\boldsymbol{1}}-R_{IS}'f(X_{IS},\theta(s_*))\big)\,.\]
Theorem 3 delivers the central decomposition of the paper. Once the PTK has stabilized, gradient descent no longer alters the feature representation learned by the
network. Instead, subsequent learning acts only through the aggregation of a fixed collection of gradient-based factors. In this regime, the DNN supplies the features, while pricing is governed entirely by the PTK.
This result makes precise the separation between what is learned and how it is priced. The term labeled learned model captures the residual pricing rule associated with the network at the stabilization point \(s_*\). The second term—the PTK-SDF—corresponds to an explicit factor pricing model that optimally aggregates the learned features. Because this aggregation is efficient conditional on the features, the learned-model component
does not improve pricing performance and can be discarded without loss.
The structure of \(\xi(s)\) reveals that the PTK-SDF implements spectral shrinkage closely related to ridge regression. To see this, consider the spectral decomposition of the in-sample kernel matrix, \({\mathbb{K}}^*_{IS}\;=\;UDU',\; D=\operatorname{diag}(\lambda_i)_{i=1}^T\,.\) Then, \[\begin{align}
&\Big(\tfrac1T{\mathbb{K}}^*_{IS}\Big)^{-1}
\big(I-e^{-\tfrac1T\eta {\mathbb{K}}^*_{IS}(s-s_*)}\big)
\;=\;U f_{GD}(D)U'\,,\\
&(zI+\tfrac1T{\mathbb{K}}^*_{IS})^{-1}
\;=\;U f_{ridge}(D)U'\,,
\end{align}\] where \[\begin{align}
&f_{GD}(\lambda)\;=\;\lambda^{-1}\big(1-e^{-\eta (s-s_*) \lambda}\big)\,,\\
&f_{ridge}(\lambda)\;=\;\lambda^{-1}\frac{1}{1+\lambda^{-1}z}\,,
\end{align}\] are the spectral filters induced by gradient descent (GD) and ridge regularization, respectively. They both regularize the factor covariance matrix \({\mathbb{K}}^*_{IS}\) by downweighting directions
associated with small eigenvalues. The key distinction is that, under gradient descent, shrinkage is generated endogenously by the learning dynamics rather than imposed ex ante through an explicit penalty. The effect of this shrinkage is most pronounced
for small eigenvalues \(\lambda\), for which \[\begin{align}
&f_{GD}(\lambda)\;=\;\eta (s-s_*)\;+\;O(\lambda)\,,\\
&f_{ridge}(\lambda)\;=\;z^{-1}\;+\;O(\lambda)\,.
\end{align}\] Thus, gradient descent induces an effective ridge penalty of order \(1/(\eta (s-s_*))\).
If the kernel matrix \({\mathbb{K}}^*_{IS}\) is strictly positive definite, then as \(s\to\infty\), the exponential term vanishes, and the PTK-SDF converges to the ridgeless kernel
regression \[{\mathbb{K}}^*(Y,Y_{IS})\,({\mathbb{K}}^*_{IS})^{-1}{\boldsymbol{1}}\,.\] As we show in the next section, although this estimator interpolates the in-sample data, it remains disciplined out of sample due to
implicit spectral shrinkage induced by the PTK. The strength of this implicit regularization is governed by the spectral complexity of the PTK and plays a central role in shaping out-of-sample SDF performance.
Recent work emphasizes that limits to learning—the gap between the population-optimal SDF and the model learned from finite samples [33]—are
governed by two fundamental properties of high-dimensional pricing models [1], [14]. The first is the spectral distribution of factor risk, which determines the severity of statistical shrinkage. The second is alignment: the extent to which economically relevant risk premia load on
statistically strong directions of the factor space.
Random features are generically poorly aligned with the true SDF. Learned features, by contrast, may become substantially better aligned through training, allowing neural networks to outperform despite operating in extremely high dimensions.
We formalize these ideas using tools from Random Matrix Theory (RMT). Let \(F\in {\mathbb{R}}^{T\times P}\) denote the matrix of factor returns, and define \[\bar F\;=\;\frac{1}{T} \sum_{t=1}^T
F_t,\qquad
\hat{\Sigma}\;=\;\frac{1}{T} \sum_{t=1}^T F_tF_t'-\bar F\bar F'\,.\] Define the effective penalty \[\label{dfn:zstar}
\hat{Z}_*(z)\;=\;\frac{1}{T^{-1}\operatorname{tr}\!\big((zI+FF'/T)^{-1}\big)}\,.\tag{23}\]
Theorem 4. Suppose \(F_{t+1}=\mu_P+{\Sigma}_P^{1/2}X_t\), where \(X_t\) has i.i.d.components with \(E[X_{i,t}^{8+\varepsilon}]<\infty\). Then, \[\label{aym-main}
(zI+\hat{\Sigma})^{-1}\bar F\;-\;\frac{\hat{Z}_*(z)}{z}(\hat{Z}_*(z)I+{\Sigma}_P)^{-1}\mu_P\;\to\;0\qquad{(5)}\] in the weak sense as \(P,T\to\infty\) with \(P/T\to
c>0\).
Theorem 4 implies that all estimation noise collapses asymptotically to the scalar \(\hat{Z}_*(z)\). When \(P>T\), \(\hat{Z}_*(0)>0\), implying nontrivial shrinkage even for ridgeless estimators.
Let \({\Sigma}=UDU'\) be the population eigenvalue decomposition. Then \[(\hat{Z}_*(z)I+{\Sigma})^{-1}\mu
=
\sum_{i=1}^P
\frac{U_i'\mu}{\lambda_i}\,
\frac{1}{1+\hat{Z}_*(z)\lambda_i^{-1}}\,U_i\,,\] and we arrive at the following result.
Theorem 5 (Projection Theorem). If \(\hat{Z}_*(z)\) separates the spectrum so that \(\hat{Z}_*(z)/\lambda_i=O(\varepsilon)\) for \(i\le
I\) and \(\lambda_i/\hat{Z}_*(z)=O(\varepsilon)\) for \(i>I\), then \[(\hat{Z}_*(z)I+{\Sigma})^{-1}\mu
\;=\;P_I({\Sigma}^{-1}\mu)
\;+\;O(\varepsilon)\;,\;P_I\;=\;\sum_{i=1}^I U_iU_i'\,.\]
The economic intuition behind Theorem 5 is as follows. Finite samples force investors to downweight directions of factor risk that are difficult to estimate. This shrinkage
intensifies as dimensionality increases. What matters for performance is therefore not the number of factors, but whether economically relevant risk premia concentrate in directions that survive shrinkage. Random Features spread risk premia across many
weak directions and are poorly aligned. Learned features, by contrast, can rotate the factor space so that economically important variation loads on strong principal components. Alignment determines how much of the population Sharpe ratio survives
finite-sample shrinkage. This naturally raises the question of how to estimate \(\mu'{\Sigma}^{-1}\mu\) when \(P/T\) is non-negligible.
In low dimensions, \(\mu'{\Sigma}^{-1}\mu\) is usually estimated by the [16] (henceforth, GRS)
statistic. In high dimensions, however, this estimator is biased. [15] show that the bias can be corrected using the same effective penalty
\(\hat{Z}_*(z)\) that governs limits to learning: \[\label{debiased}
\hat{W}^{debiased}(z)
=
\frac{\bar F'(zI+\hat{\Sigma})^{-1}\bar F-(z^{-1}\hat{Z}_*(z)-1)}{z^{-1}\hat{Z}_*(z)}
\;\approx\;
\mu'(\hat{Z}_*(z)I+{\Sigma})^{-1}\mu\,.\tag{24}\] The debiased statistic \(\hat{W}^{debiased}(z)\) therefore provides a natural finite-sample measure of economically relevant alignment. In the empirical
analysis below, we apply this statistic to both random features and PTK-based learned features.
This section evaluates the empirical implications of the Portfolio Tangent Kernel (PTK) framework using U.S.equity data. We compare the performance of three classes of stochastic discount factors (SDFs): (i) fully trained deep neural network (DNN) SDFs,
(ii) random-feature SDFs following [1], and (iii) PTK-based SDFs that optimally price features learned by the DNN. Our analysis focuses on
out-of-sample pricing performance, factor spanning, statistical robustness, and economic alignment with long-horizon consumption risk. All results are reported using rolling-window implementations and are evaluated separately across market capitalization
groups.
We use monthly data on 153 stock characteristics for U.S.equities over the period 1963–2024 from [34] hereafter JKP. This open-source
dataset collects a comprehensive set of firm-level characteristics drawn from the asset pricing literature.6 The universe consists of NYSE, AMEX, and NASDAQ stocks with CRSP
share codes 10–12.
To construct a stock–month panel with a stable set of characteristics, we follow the filtering procedure of [1]. First, we restrict
attention to a subset of 132 characteristics with fewer than 30% missing observations over the full sample. Second, we exclude nano stocks whose market capitalization falls below the 1st percentile of NYSE market capitalizations. Third, we drop stock–month
observations with missing values for more than one-third of the characteristics. Finally, each characteristic is rank-standardized cross-sectionally to the interval \([-0.5,0.5]\), and remaining missing values are imputed
using the cross-sectional median, which equals zero by construction. After this procedure, the resulting \(N_t\times132\) matrix of characteristics at each date constitutes the conditioning set \(X_t\). Let \(R_{t+1}=(R_{i,t+1})_{i=1}^{N_t}\) denote the vector of next-month excess returns.
To isolate the role of statistical complexity while abstracting from liquidity concerns, we conduct the analysis separately across market capitalization groups. Following JKP, we study mega (largest 20%), large (80%–50%), small (50%–20%), and micro
(20%–1%) stocks based on NYSE breakpoints each period, as well as the full universe of stocks above the 1% NYSE cutoff, which we refer to as “all.”
As a benchmark, we construct random-feature-based SDFs following [1]. Specifically, we define \[S_t={\rm
ReLU}(X_t\gamma)\;\in\;{\mathbb{R}}^{N_t\times P},\] where \(\gamma=(\gamma_{j,k})\in{\mathbb{R}}^{d\times P}\) with \(P=25{,}000\) and \(\gamma_{j,k}\sim
N(0,1)\) independently. These nonlinear transformations constitute the random features used in the analysis of [1]. We then
form factor returns \[\label{non-linearDKKM}
F^{\mathrm{ReLU}}_{t+1}
\;=\;
N_t^{-1/2}\,R_{t+1}'S_t\,.\tag{25}\]
Normalization. Throughout the paper, all factor returns are normalized by \(N_t^{-1/2}\), where \(N_t\) denotes the cross-sectional sample size at time \(t\). This scaling ensures that factor second moments remain stable across time and are comparable across periods with different cross-sectional dimensions. For the DNN and PTK constructions, this normalization is implemented
directly in the factor definitions, ensuring scale comparability with the random feature benchmark.
We use these normalized factor returns to estimate a candidate stochastic discount factor as the return on a Markowitz portfolio computed over a rolling window of \(T\) months with ridge penalty \(z\): \[\label{R-SDFeLu}
R^{\mathrm{ReLU}}_{t+1}(z,T)\;=\;\theta_t(z;T)'F^{\mathrm{ReLU}}_{t+1},\tag{26}\] where \(\theta_t(z;T)\) is obtained from 13 .
Because the impact of the ridge penalty depends on its magnitude relative to the factor spectrum, we scale \(z\) proportionally to the average eigenvalue of the factor covariance matrix,
\[\label{z61zeff}
z\;=\;z_{eff} P^{-1}\operatorname{tr}(F'F/T)\,.\tag{27}\] We consider rolling estimation windows \(T=12,60,120,240,360\) months and set \(z_{eff}=10^{-5}\). The resulting
ReLU-based SDF returns are combined into an ensemble benchmark using weights inversely proportional to their realized volatility over the preceding 12 months. Importantly, these volatility weights are computed using only information available at time \(t\).
Formally, the ensemble random feature SDF is given by \[\label{the-ReLU-bench}
\bar R^{\mathrm{ReLU}}_{t+1}(z_{eff})
\;=\;
\sum_{T\in\{12,60,120,240,360\}}
\omega_{t,T}\,
R^{\mathrm{ReLU}}_{t+1}(z_{eff} P^{-1}\operatorname{tr}(F'F/T),T),\tag{28}\] where \(\omega_{t,T}\propto \widehat{\sigma}^{-1}_{t,T}\) and \(\widehat{\sigma}_{t,T}\)
denotes the realized volatility of \(R^{\mathrm{ReLU}}_{\tau+1}(z_{eff},T),\;\tau\le t-1,\) computed over the trailing 12 months.
We consider a multilayer perceptron (MLP) architecture with an input layer of dimension 132, corresponding to the number of characteristics, followed by \(D\) hidden layers of width \(w\). We consider depths \(D=1,2,4\) and widths \(w=16,32,64,128,256,512\). The parameter \(w\) controls network width, while
\(D\) controls depth. We denote the resulting function family by \(f(X;\theta;D;w)\).
The dimension of the parameter vector \(\theta\) is given by \[P \;=\;\underbrace{(d_{in}+1)\times w}_{\text{Input layer}}
\;+\;\underbrace{(D-1)\times (w+1)\times w}_{\text{Hidden layers}}
\;+\;\underbrace{(w+1)}_{\text{Output layer}},\] see Definition 3. The deepest and widest network in our analysis contains \(856{,}577\)
parameters. Our DNN–PTK duality implies that \(P\) is the relevant measure of model complexity: it equals both the number of factors in the associated PTK and the interpolation capacity of the network. In particular, by
Theorem 2, the DNN-SDF can interpolate \(T\) observations—achieving in-sample returns equal to one for \(t=1,\ldots,T\)—if and only if \(P\ge T\), provided the PTK matrix is non-degenerate.
We train the DNN using gradient descent as described in Appendix 9, employing a rolling estimation window of 60 months and a learning rate \(\eta=2^{-16}\). At each time \(t\), the network is trained using only data available up to time \(t\). To obtain a common set of learned features, we train the network on the full stock universe (“all”) and obtain a
time-\(t\) parameter vector \(\theta_t^*\). We then evaluate the resulting SDF separately on each size group (mega, large, small, and micro).
The resulting DNN-based SDF is \[\label{def:rdnn}
R^{\mathrm{DNN}}_{t+1}(D;w)
\;=\;
N_t^{-1/2}\,R'_{t+1} f(X_t;\theta^*_t;D;w)\,.\tag{29}\]
We next construct SDFs based on the Portfolio Tangent Kernel. Given a trained network \(f(X;\theta_t^*;D;w)\), we form factor returns using the learned gradient features as in 20 , \[F_{\tau+1}(\theta_t^*;D;w)
\;=\;N_\tau^{-1/2}\,
R_{\tau+1}'\underbrace{\nabla_\theta f(X_\tau;\theta_t^*;D;w)}_{time\text{-}t\;learned\;features}
\;\in\;{\mathbb{R}}^{P},\qquad P=\dim(\theta)\,.\] Thus, \(F_{\tau+1}(\theta_t^*;D;w)\) represents returns on the time-\(t\) vintage of characteristics-managed portfolios constructed
from features learned at time \(t\).
The corresponding time-\(t\) Portfolio Tangent Kernel (PTK) is defined as \[{\mathbb{K}}\big((R_{\tau_1+1},X_{\tau_1});(R_{\tau_2+1},X_{\tau_2});\theta_t^*;D;w\big)
\;=\;F_{\tau_1+1}(\theta_t^*;D;w)' F_{\tau_2+1}(\theta_t^*;D;w)\,.\]
Fixing a rolling window length \(T\), we estimate the ridge-penalized Markowitz portfolio 13 using factor returns \(F_{\tau+1}(\theta_t^*;D;w)\) for \(\tau=t-T,\ldots,t-1\). Denoting the resulting portfolio weights by \(\Theta_t(\theta_t^*;D;w;T;z)\), the associated PTK-based SDF is \[\label{learned-feat-1}
R^{\mathrm{PTK}}_{t+1}(D;w;T;z)
\;=\;
\Theta_t(\theta_t^*;D;w;T;z)'F_{t+1}(\theta_t^*;D;w)\,.\tag{30}\]
As with the random feature benchmark, we construct an ensemble PTK-SDF by averaging across rolling windows, \[\label{the-ptk-bench}
\bar R^{\mathrm{PTK}}_{t+1}(D;w;z_{eff})
\;=\;
\sum_{T\in\{12,60,120,240,360\}}
\omega_{t,T}^{\mathrm{PTK}}\,
R^{\mathrm{PTK}}_{t+1}(D;w;T;z_{eff} P^{-1}\operatorname{tr}(F'F/T)),\tag{31}\] where \(\omega_{t,T}^{\mathrm{PTK}}\propto \widehat{\sigma}^{-1}_{t,T}\) and \(\widehat{\sigma}_{t,T}\) denotes the realized volatility of \(R^{\mathrm{PTK}}_{\tau+1}(D;w;T;z),\;\tau\le t-1,\) computed over the trailing 12 months.
This aggregation exploits potential non-stationarity by allowing the pricing rule to adapt across time scales. Importantly, this flexibility is specific to the PTK construction and entails negligible computational cost, as estimating Markowitz
portfolios for multiple \(T\) is numerically trivial.
6.3 The Performance of \(R^{\rm ReLU}, R^{\rm DNN},\) and \(R^{\mathrm{PTK}}\)↩︎
The random feature portfolios \(R^{\mathrm{ReLU}}_{t+1}\) defined in 26 correspond to a shallow (single hidden layer) neural network with fixed, randomly generated features. Despite the absence
of feature learning, [1] document strong empirical performance of this nonlinear SDF, driven by sheer statistical complexity: the model
effectively prices returns using \(P=25{,}000\) nonlinear factors. This benchmark, therefore, provides a stringent baseline against which to evaluate fully trained neural networks.
Our first objective is to assess whether a trained DNN is able to learn economically relevant features—that is, priced nonlinear characteristics that are not already captured by the random feature SDF.
Figure 2 reports the Sharpe ratios of \(R^{\rm DNN}_{t+1}(D;w)\) from 29 as a function of network depth and width, alongside the ensembled ReLU
benchmark \(\bar R^{\mathrm{ReLU}}\) from 28 . The patterns are similar across size groups; we therefore focus on the full stock universe (“all”).
Several findings stand out. First, depth is essential. Shallow networks (\(D=1\)) perform poorly and substantially underperform the random feature benchmark. Second, deeper networks are capable of extracting more
structure: For example, the depth-2, width-256 network attains a Sharpe ratio of 3.8, slightly exceeding the ReLU benchmark (3.7). This architecture is highly complex, with \(P=100{,}097\) trained parameters. Nevertheless,
even such large networks generally fail to outperform the random feature SDF across all size groups, despite having four times as many parameters.
Taken at face value, these results might suggest that fully trained DNNs add little value for asset pricing. This conclusion, however, is misleading. As emphasized by [14], raw Sharpe ratios alone are not an appropriate metric for evaluating machine learning models. What matters is whether the model uncovers new priced variation—that is, whether
it discovers risk premia that are unspanned by existing models.
To separate what the DNN learns from how it prices, we test the spanning hypothesis by estimating the univariate regressions on the 28 benchmark with \(z_{eff}=10^{-5}:\)\[\label{main-regr0}
R_{t+1}^{\rm DNN}(D;w)
\;=\;
\alpha
\;+\;
\beta\,\bar R^{\mathrm{ReLU}}_{t+1}(10^{-5})
\;+\;
\varepsilon_{t+1}\,.\tag{32}\] Figure 3 reports the \(t\)-statistics of the intercepts.
The results strongly support the feature-learning hypothesis. For moderately complex deep architectures (e.g., \(D=2\) and \(D=4\) with \(w=256\)), the
estimated alpha is positive and highly statistically significant. In contrast, very small networks fail to uncover new priced characteristics, while extremely large networks (depth 4 with a width above 256, corresponding to \(P>200{,}000\)) sometimes fail to do so reliably.7
Economically, these findings imply that moderately complex DNNs (\(P\approx100{,}000\)) are effective at learning what to price: they discover a rich set of priced nonlinear characteristics that are not spanned
by random features. However, they are less effective at learning how to price these characteristics, resulting in suboptimal Sharpe ratios despite strong alpha.
Motivated by these results, we focus on the architecture \(R^{\rm DNN}(D=2,w=256)\) in the remainder of the analysis. This specification exhibits the strongest and most robust feature-learning performance across all size
groups, as illustrated in Figure 3.
We now turn to our central theoretical prediction: the Portfolio Tangent Kernel (PTK) should optimally price the features learned by the DNN. If the learned characteristics subsume those captured by the ReLU benchmark, and if the PTK implements an
efficient pricing rule, then \(\bar R^{\rm PTK}(D=2,w=256)\) should outperform both \(R^{\rm DNN}\) and \(\bar R^{\mathrm{ReLU}}\).
Figure 4 confirms this prediction. The PTK-SDF delivers substantially higher Sharpe ratios than both benchmarks across all size groups. The gains are particularly pronounced for mega and large stocks: PTK
achieves Sharpe ratios of 1.6 and 1.9, compared to 1.3 and 1.6 for the ReLU benchmark.
A striking implication is that the ridgeless PTK attains the highest Sharpe ratios across all size groups. These findings underscore the role of statistical complexity: when low-variance principal components contain an economically meaningful
signal—as suggested by the strong alpha evidence in Figure 3—explicit shrinkage becomes unnecessary.
Figure 5 further shows that the outperformance of PTK is statistically significant. The PTK-SDF delivers large and significant alpha relative to both \(R^{\rm DNN}\) and
\(\bar R^{\mathrm{ReLU}}\), indicating that it fully subsumes their pricing information.
We therefore conclude that feature learning and pricing should be viewed as distinct tasks. DNNs excel at discovering rich nonlinear characteristics, while the PTK provides a statistically efficient mechanism for pricing them. The PTK representation
renders the learned SDF amenable to the Random Matrix Theory analysis developed in Section 5, to which we now turn.
Figure 2: Sharpe Ratio of \(R^{\rm DNN}\) from 29 as a Function of Model Complexity (depth \(D\) and width \(w\)). The lines \(d1,d2,d4\) correspond to \(D=1,2,\) and \(4\), respectively, with widths \(w \in
\{2^4, \dots, 2^9\}\). The dashed line is the Sharpe ratio of the \(\bar R^{\mathrm{ReLU}}\) benchmark from 28 with \(z_{eff}=10^{-5}\).Figure 3: Alpha t-statistics from the regressions 29 as a Function of Model Complexity (depth \(D\) and width \(w\)). The lines
\(d1,d2,d4\) correspond to \(D=1,2,\) and \(4\), respectively, with widths \(w \in \{2^4, \dots, 2^9\}\).Figure 4: Sharpe Ratios of \(\bar R^{\rm PTK}(D=2,w=256;z_{eff})\) from 31 and \(\bar R^{\rm ReLU}(z_{eff})\) from 28 (both are ensembled across \(T\)), as a function of the \(z_{eff}.\) The blue dash-dotted line represents the Sharpe Ratio of \(R^{\rm DNN}(D=2,w=256)\) benchmark from 29 .Figure 5: The t-statistics of alpha from univariate regressions of \(\bar R^{\rm PTK}(D=2,w=256)\) with \(z_{eff}=10^{-5}\) (see 31 ) on \(R^{\rm DNN}(D=2,w=256)\) from 29 (red bars) and \(\bar R^{\textrm{ReLU}}\) from 28 (grey
striped bars). The dashed horizontal line indicates the level of 1% statistical significance.
This out-performance is also statistically significant, as is shown by Figure 5. We therefore conclude that learned features (PTK) entirely subsume the performance of the DNN. In turn, the PTK
representation of the candidate SDF makes it amenable to the RMT analysis from Section 5.
By Theorem 4, a central determinant of LFM performance is the effective shrinkage \(\hat{Z}_*(z)\) defined in 23
. In the PTK setting, we adopt a scaled version of 27 and set \[z\;=\;z_{eff}\,T^{-1}\operatorname{tr}(FF'/T)\,,\] which is proportional to 27 . Since \(\hat{Z}_*(z)\) scales linearly with the kernel matrix \(K=FF'\), we normalize the effective ridge penalty and define \[\label{z431}
\frac{\hat{Z}_*(z)}{T^{-2}\operatorname{tr}(K)}
\;=\;
\frac{1}{T^{-1}\operatorname{tr}\!\big((z_{eff}T^{-2}\operatorname{tr}(K) I+K/T)^{-1}\big)\,\cdot\,T^{-2}\operatorname{tr}(K)}\,.\tag{33}\] By Jensen’s inequality, \(T^{-1}\operatorname{tr}(A^{-1})\;\ge\;(T^{-1}\operatorname{tr}(A))^{-1}\) for any symmetric positive definite \(T\times T\) matrix \(A\), which implies
\[\label{z431ineq}
\frac{\hat{Z}_*(z_{eff}\,T^{-2}\operatorname{tr}(K))}{T^{-2}\operatorname{tr}(K)}
\;\le\;1+z_{eff}\,.\tag{34}\] The normalized quantity in 33 measures the fraction of total factor variance effectively shaded by implicit regularization in an over-parameterized LFM. The upper bound in 34 is attained when \(K\approx I\), corresponding to maximal spectral complexity, where all principal components contribute equally. We therefore refer to 33 as the
spectral complexity of the kernel matrix.
Figure 6 reports the time-series dynamics of spectral complexity (left panel) and the debiased GRS statistic 24 (right panel) for both random-feature-based factors 25 and PTK-based factors. Several striking empirical regularities emerge.
First, PTK factors exhibit spectral complexity roughly an order of magnitude larger than that of random features. Rather than concentrating risk in a small number of dominant directions, the learned representation spreads economically relevant variation
across a large number of principal components. This finding underscores that feature learning does not simplify the factor space; instead, it reorganizes it.
Second, PTK spectral complexity rises monotonically after 2000 and increases by approximately a factor of six over the sample period. In contrast, the spectral complexity of random feature factors remains essentially flat. The increase is therefore
specific to learned, priced features and reflects a gradual rise in the dimensionality of economically relevant risk.
Third, the debiased GRS statistic 24 rises steadily until the early 2000s and declines after 2010. Importantly, this decline is not accompanied by a deterioration in realized SDF performance (Figure 9 in the Appendix). Combined with the sharp increase in \(\hat{Z}_*(z)\), this pattern suggests that the population Sharpe ratio \(\mu'{\Sigma}^{-1}\mu\) has continued to grow, while finite-sample limits to learning have intensified due to rising spectral complexity.
To further investigate alignment in the sense of Theorem 5, we conduct a principal-component truncation exercise. For each \(K\), we construct
truncated versions of \(R^{\mathrm{PTK}}_{t+1}(D_*;w_*;T,z)\) and \(R^{\mathrm{ReLU}}_{t+1}(T,z)\) by retaining only the top \(K\) eigenvectors of the sample
covariance matrix \[\label{eigen-dec}
\hat{\Sigma}_t
\;=\;
T^{-1}\sum_{\tau=t-T}^{t-1}F_{\tau+1}F_{\tau+1}'
\;=\;
\hat{U}_t \hat{D}_t \hat{U}_t'\,,\tag{35}\] and forming portfolios based on \[\label{ftk}
\hat{F}_t(K)\;=\;\hat{U}_t[:,:K]'\hat{F}_t.\tag{36}\] Figure 7 plots Sharpe ratios as a function of \(K\) for learned gradient features and random features.
Consistent with the alignment mechanism of Section 5, learned features concentrate economically relevant signal in far fewer directions. For the full stock universe, the top \(15\) (out
of \(60\)) PTK principal components already achieve a Sharpe ratio of \(3.8\), close to the full-feature maximum of \(3.9\). By contrast, random features
require substantially larger \(K\) to extract the signal and reach only \(2.3\) at the same truncation level.
Taken together, these results highlight a central tension. Feature learning increases spectral complexity and exacerbates limits to learning, yet simultaneously improves alignment by concentrating the economic signal in statistically robust directions.
The empirical success of PTK-based pricing stems precisely from this trade-off.
Figure 6: The dynamics of spectral complexity 33 and the debiased GRS statistic 24 estimated using \(T=360\) for \(R^{\rm PTK}\) and \(R^{\textrm{ReLU}}\); we use \(D=2,\;w=256.\)Figure 7: Factor Alignment. This figure plots the Sharpe Ratio against the number of principal components (\(K\)) for PTK Factors and random feature factors, penalized with \(z_{eff}=10^{-5}\), for different backtest window sizes \(T=12, 60, 120, 240, 360\).
Modern asset pricing theories emphasize long-run and future consumption risk as key drivers of equity premia [36], [37]. In these models, the stochastic discount factor (SDF) loads on future consumption realizations rather than contemporaneous consumption growth.
[2] show that covariances with future realized consumption are priced in the cross-section of stock returns, while [3] provide a micro-founded version based on financial frictions. These frameworks offer economically disciplined benchmarks for
evaluating whether machine-learning-based SDFs capture meaningful macroeconomic risk.
Our alignment framework implies that economically relevant pricing information should concentrate in the leading principal components of the factor space (Theorem 5). We therefore
examine whether the dominant components of PTK-based factors exhibit stronger alignment with future consumption risk than those of random features or fully trained DNN-SDFs.
We project \(\bar R^{\rm PTK}\) and \(\bar R^{\rm ReLu}\) onto their top two principal components 36 and compute correlations and predictive regressions with future
consumption growth, following [2] and [3].8
Figure 8 reports the results. Across all size groups, the PTK-SDF exhibits substantially stronger alignment with both consumption-based benchmarks than either the DNN-SDF or the random-feature benchmark. PTK
delivers larger absolute correlations, higher regression \(t\)-statistics, and materially greater explanatory power. For the aggregate stock universe, the projected PTK-SDF is approximately \(-30\%\) correlated with future consumption growth under [2] and nearly \(-50\%\) under [3], compared to correlations around \(-10\%\) and
\(-20\%\), respectively, for the alternatives.
These findings highlight the economic role of PTK. Gradient-based feature learning rotates the high-dimensional characteristic space so that consumption-related risk loads on statistically strong directions that survive shrinkage. Random features fail
to achieve this alignment, while unconstrained DNN-SDFs do not price it efficiently.
Overall, PTK isolates the economically relevant component of feature learning and translates it into a disciplined pricing rule, yielding tradable SDFs that are tightly linked to long-horizon consumption risk.
Figure 8: (Top 2 PCs for ReLU and PTK based SDFs) Machine Learning SDFs and Consumption Risk. This figure reports the relation between tradeable machine-learning stochastic discount factors (SDFs) and consumption-based SDFs
from [2] (PJ) and [3] (MWZ). The top row shows time-series correlations between portfolio returns from the fully trained DNN SDF \(R^{\mathrm{DNN}}\)29 (red), the random-feature
benchmark \(\bar R^{\mathrm{ReLU}}\)28 (grey, striped), and the PTK-based SDF \(\bar R^{\mathrm{PTK}}\)31 (blue), and the
MWZ future-consumption SDF (left panel) and the PJ long-run consumption SDF (right panel). The middle row reports \(t\)-statistics from univariate time-series regressions of each consumption-based SDF on the corresponding
tradeable SDF return. The bottom row reports the associated adjusted \(R^2\). All statistics are shown separately by size group. Consumption-based SDFs are constructed using future realized consumption growth aggregated
over horizons of 84 months (MWZ) and 36 months (PJ), respectively, which truncates the sample in June 2018. Ridge regularization with \(z_{eff}=10^{-5}\) is applied to the ReLU and PTK portfolios.
This paper studies how modern machine learning models price assets in high-dimensional environments and why their performance is constrained by fundamental limits to learning. We develop a framework that links deep neural networks, kernel methods, and
factor pricing through the Portfolio Tangent Kernel (PTK), which provides an explicit representation of the pricing rule induced by gradient-based training.
Our main conceptual contribution is to separate feature learning from pricing. Fully trained neural networks are effective at discovering rich, nonlinear characteristics that are unspanned by standard factor constructions. However, when used directly as
stochastic discount factors, these models often fail to aggregate the learned features efficiently. The PTK makes this distinction explicit: network architecture and training determine which features are learned, while the PTK governs how these features
are priced through a statistically disciplined factor model.
Using random matrix theory, we show that finite-sample estimation induces endogenous shrinkage even in ridgeless, interpolating regimes. This shrinkage imposes an effective low-dimensional structure on the pricing rule, and its severity depends
critically on the alignment between risk premia and the principal components of the feature space. We formalize this mechanism through a projection theorem that clarifies how limits to learning translate into economic restrictions on achievable Sharpe
ratios.
Empirically, we find that learned gradient features exhibit substantially stronger alignment than random features, allowing the PTK-based SDF to outperform both raw DNN-SDFs and strong random feature benchmarks across asset size groups. The results show
that the gains from deep learning arise not from abandoning structure, but from combining flexible representation learning with explicit and statistically efficient pricing rules.
Overall, our findings suggest that the role of machine learning in asset pricing is best understood as expanding the space of economically relevant characteristics, while leaving pricing to be governed by well-understood principles of risk,
regularization, and finite-sample discipline.
Definition 3 (Multi-Layer Perceptron (MLP)). Fix a neural network architecture given by the widths of layers \((n_1,\cdots,n_L).\) Let \(\theta=(W^{(0)}, b^{(0)}, \cdots,
W^{(L-1)}, b^{(L-1)})\) be a collection of weights and biases. Here, \(W^{(l)}\in {\mathbb{R}}^{n_l\times n_{l+1}},\;b^{(l+1)}\in {\mathbb{R}}^{n_{l}}\), and the total dimension of \(\theta\) is \(P=\sum_{l=0}^{L-1} (n_l+1)n_{l+1}\). Let also \(\phi:{\mathbb{R}}\to{\mathbb{R}}\) be a Lipschitz, twice differentiable nonlinearity with a bounded
second derivative. The MLP neural network \(f(x;\theta)\) is defined as \[\begin{align}
x &= \text{input}\;\in\;{\mathbb{R}}^d \\
y^{(l)}(x) &= \begin{cases}
x & \text{if } l = 0, \\
\phi(z^{(l)}(x)) & \text{if } l > 0.
\end{cases} \\
z^{(l+1)}_i(x) &= \frac{1}{n_l^{1/2}}\sum_j W^{(l)}_{ij} y^{(l)}_j(x) + b^{(l)}_i,\\
\end{align}\] where \(y^{(l)}(x)\), \(z^{(l)}(x) \in \mathbb{R}^{n_l}.\) The output of the network is \[\label{nngp-1}
f(x;\theta)\;=\;z^{(L)}(x)\;=\;\frac{1}{n_{L-1}^{1/2}}\sum_{j=1}^{n_{L-1}}W_j^{(L-1)}y_j^{(L-1)}(x)\,.\qquad{(6)}\] At initialization, all weights \(W^{(\ell)}_{i,j}\) and biases \(b^{(l)}\) in each hidden layer are sampled independently from \(N(0,(\sigma_W^{(l)})^2)\) and \(N(0,(\sigma_b^{(l)})^2)\), respectively, for some layer-specific
parameters \(\sigma_W^{(l)}, \sigma_b^{(l)}\).
As a benchmark example, consider a single-hidden-layer NN \[\label{simple-nn}
\begin{align}
&f(x;\theta)\;=\;z^1(x)\;=\;\frac{1}{n_1^{1/2}}\sum_{k=1}^{n_1}W_k^{1}\phi\left(\frac{1}{n_0^{1/2}}\sum_{i=1}^d W_{k,i}^0 x_i+b_k^0\right)\\
&=\;\frac{1}{n_1^{1/2}}\sum_{k=1}^{n_1}W_k^{1}\phi\left(\frac{1}{n_0^{1/2}}(W_k^0)'x+b_k^0\right)\\
&=\;\frac{1}{n_1^{1/2}} (W^1)'\phi\left(\frac{1}{n_0^{1/2}}(W^0)'x+b^0\right)\,.
\end{align}\tag{37}\] The simple random feature model is therefore equivalent to the shallow (single hidden layer) neural network 37 , with \(\theta=W^1\in {\mathbb{R}}^P,\;W_k=W_k^0\in
{\mathbb{R}}^d,\;P=n_1\). However, there is one important caveat: Neural networks are trained end-to-end; that is, every single weight, including all weights of all hidden layers, is optimized. E.g., for the simple 1-hidden layer net 37 , the whole \(\theta=(W^1,W^0,b^0)\in {\mathbb{R}}^{n_1(d+2)}\) is optimized upon.
Definition 4 (The Infinite Width NTK). Let \[\label{scary-int-1}
\begin{align}
&\widehat\Sigma^{(l)}(x,\tilde{x}) \\
&= (\sigma_w^{(l)})^2\int dz \, d\tilde{z} \frac{d}{dz}\phi(z) \frac{d}{d\tilde{z}}\phi(\tilde{z}) \mathcal{N}\left( \begin{pmatrix} z \\ \tilde{z} \end{pmatrix} ; \mathbf{0}, (\sigma_w^{(l-1)})^2\begin{pmatrix} \Sigma^{(l-1)}(x, x) &
\Sigma^{(l-1)}(x,\tilde{x}) \\ \Sigma^{(l-1)}(\tilde{x},x) & \Sigma^{(l-1)}(\tilde{x},\tilde{x}) \end{pmatrix} +(\sigma^{(l-1)}_b)^2\,{\boldsymbol{1}}_{2\times 2} \right)\,.
\end{align}\qquad{(7)}\] Let also \[\Theta^{(1)}(x,\tilde{x})\;=\;\Sigma^{(1)}(x,\tilde{x})+1,\] and then define recursively for \(l>1:\)\[\label{ntk-closed}
\Theta^{(l)}(x,\tilde{x})\;=\;\Theta^{(l-1)}(x,\tilde{x})\widehat\Sigma^{(l)}(x,\tilde{x})\;+\;\Sigma^{(l)}(x,\tilde{x})+1\qquad{(8)}\]
Theorem 6. Suppose that \(\phi(x)\) is uniformly Lipschitz.9 Then, at initialization, \[\nabla_{\theta}z^{(l)}_i(x)\nabla_{\theta}z^{(l)}_{i_1}(\tilde{x})\;\to\;\delta_{i,i_1}\Theta^{(l+1)}(x,\tilde{x})\] where \(\Theta\) is defined in Definition 4.
The proof is by induction. Recall that \[\begin{align}
x &= \text{input}\;\in\;{\mathbb{R}}^d \\
y^{(l)}(x) &= \begin{cases}
x & \text{if } l = 0, \\
\phi(z^{(l-1)}(x)) & \text{if } l > 0.
\end{cases} \\
z^{(l)}_i(x) &= \frac{1}{n_l^{1/2}}\sum_j W^{(l)}_{ij} y^{(l)}_j(x) + b^{(l)}_i,\\
\end{align}\] where \(y^{(l)}(x)\), \(z^{(l-1)}(x) \in \mathbb{R}^{n^{(l)} \times 1}.\) We have \[z^{(0)}_i(x)\;=\;\frac{1}{n_0^{1/2}}\sum_{j=1}^{n_0}W_{i,j}^{(0)}x_j\;+\;b^0_j\] and, hence, \[\nabla_{\theta}z^{(0)}_i(x)\;=\;(\frac{1}{n_0^{1/2}}x,1),\] so that \[\nabla_{\theta}z^{(0)}_i(x)\nabla_{\theta}z^{(0)}_{i_1}(\tilde{x})'\;=\;\delta_{i,i_1}(\frac{1}{n_0}x'\tilde{x}\;+\;1)\;=\;\delta_{i,i_1}(\Sigma^{(1)}(x,\tilde{x})+1)\;=\;\delta_{i,i_1}\Theta^{(1)}(x,\tilde{x})\,.\] We
now proceed by induction. Suppose we have proved the claim for \(z^{(l-1)}\). We have \[z^{(l)}_i(x)\;=\; \frac{1}{n_l^{1/2}}\sum_{j=1}^{n_l}W_{i,j}^{(l)}\phi(z_j^{(l-1)}(x))+b^{(l)}\] The
vector of coefficients of this neural network, \(\theta,\) can be decomposed into \(\theta=(\theta_{l-1}, W^{(l)},b^{(l)}),\) where \(\theta_{l-1}\) are all
the coefficients except for those of the \(l\)’th layer. Each \(z_j^{(l)}(x),\;j=1,\cdots,n_{l}\). We have \[\begin{align}
&\nabla_{\theta_{l-1}}z^{(l)}_i(x)\;=\;\frac{1}{n_{l}^{1/2}}\sum_{j=1}^{n_l}W_{i,j}^{(l)}\nabla_{\theta_{l-1}}\phi(z_j^{(l-1)}(x))\\
&=\;\frac{1}{n_{l}^{1/2}}\sum_{j=1}^{n_{l}}W_{i,j}^{(l)}\nabla_{\theta_{l-1}}z_j^{(l-1)}(x)\phi'(z_j^{(l-1)}(x))
\end{align}\] and, hence, \[\begin{align}
&\nabla_{\theta_{l-1}}z^{(l)}_{i_1}(x)\nabla_{\theta_{l-1}}z^{(l)}_{i_2}(\tilde{x})'\\
&=\;\frac{1}{n_{l}}\sum_{j_1=1}^{n_{l}}W_{i_1,j_1}^{(L+1)}\nabla_{\theta_{l-1}}z_{j_1}^{(l-1)}(x)\phi'(z_{j_1}^{(l-1)}(x))\sum_{j_2=1}^{n_{l}}W_{i_2,j_2}^{(l)}\nabla_{\theta_{l-1}}z_{j_2}^{(l-1)}(\tilde{x})\phi'(z_{j_2}^{(l-1)}(\tilde{x}))\,.
\end{align}\] By the induction hypothesis, in the limit as \(n_{l-1},\cdots,n_1\to\infty,\) we have \[\begin{align}
&\nabla_{\theta_{l-1}}z^{(l)}_{i_1}(x)\nabla_{\theta_{l-1}}z^{(l)}_{i_2}(\tilde{x})'\\
&\approx\;\frac{1}{n_{l}}\sum_{j_1=1}^{n_{l}}W_{i_1,j_1}^{(L+1)}\phi'(z_{j_1}^{(l-1)}(x))\sum_{j_2=1}^{n_{l}}W_{i_2,j_2}^{(l)}\phi'(z_{j_2}^{(l-1)}(\tilde{x}))\delta_{j_1,j_2}\Theta^{(l)}(x,\tilde{x})\\
&=\;\frac{1}{n_{l}}\sum_{j=1}^{n_{l}}W_{i_1,j}^{(L+1)}\phi'(z_{j}^{(l-1)}(x))W_{i_2,j}^{(l)}\phi'(z_{j}^{(l-1)}(\tilde{x}))\Theta^{(l)}(x,\tilde{x})
\end{align}\] In the limit as \(n_l\to\infty,\) we get \[\begin{align}
&\nabla_{\theta_{l-1}}z^{(l)}_{i_1}(x)\nabla_{\theta_{l-1}}z^{(l)}_{i_2}(\tilde{x})'\\
&\to\delta_{i_1,i_2}(\sigma_W^{(l)})^2E[\phi'(z_{j}^{(l-1)}(x))\phi'(z_{j}^{(l-1)}(\tilde{x}))]\Theta^{(l)}(x,\tilde{x})
\end{align}\] At the same time, \[\begin{align}
&\nabla_{(W^{(l)},b^{(l)})}z^{(l)}_{i_1}(x)\nabla_{(W^{(l)},b^{(l)})}z^{(l)}_{i_2}(\tilde{x})'\\
&=\delta_{i_1,i_2}(\frac{1}{n_l}\sum_j \phi(z_{j}^{(l-1)}(x))\phi(z_{j}^{(l-1)}(\tilde{x}))+1)\\
&\to\;\delta_{i_1,i_2}(E[\phi(z_{j}^{(l-1)}(x))\phi(z_{j}^{(l-1)}(\tilde{x}))]+1)\\
&=\;\delta_{i_1,i_2}(\Sigma^{(l)}(x,\tilde{x})+1)
\end{align}\]\(\square\)
This appendix describes the full empirical pipeline. We first present the MLP training configuration (9.1), then detail the construction of PTK and random feature factors (9.2), followed by performance evaluation metrics and statistical diagnostics (9.3), and finally robustness procedures based on multiple random seeds and ensembling
(9.4).
The MLP models are trained using a rolling-window scheme designed to reflect realistic real-time forecasting and portfolio construction. At each month \(t\), the model is estimated using data from the previous 60 months
and then used to generate portfolio weights for month \(t+1\). The training procedure, architecture, and optimization details are summarized below.
9.1.0.1 Model architecture.
We employ a standard parametrization MLP (SP-MLP) with ReLU activations. Depths: 1, 2, 4 layers; widths: 16–512 neurons (16, 32, 64, 128, 256, 512). Total configurations: 18 (3×6).
9.1.0.2 Initialization.
Network weights are initialized using the default PyTorch Kaiming uniform scheme, where weights are drawn from \(\mathcal{U}(-\sqrt{k}, \sqrt{k})\) with \(k = 1/\text{in\_features}\), and
biases are initialized to zero.
9.1.0.3 Training scheme.
For the initial training window, the model is trained from scratch for 20 epochs. Thereafter, as the training window rolls forward by one month (dropping the oldest observation and adding the newest), the model is warm-started from the previous solution
and fine-tuned for 10 epochs. This rolling fine-tuning scheme substantially reduces computational cost while preserving stability and convergence.
9.1.0.4 Mini-batching and optimization.
Optimization is performed using the Adam optimizer, with learning rate \(\text{lr} = 2^{-\text{lr\_power}}\) and \(\text{lr\_power} = 16\). Training is conducted with a batch size of one,
where each batch corresponds to the entire cross-section of stocks in a given month. Specifically, at each training step, the model simultaneously processes all \(N_t\) stocks observed in month \(t\), producing a single portfolio return forecast. This design aligns the training objective directly with the portfolio construction problem.
9.1.0.5 Loss function and normalization.
The model is trained to minimize the maximal Sharpe ratio regression (MSRR) loss. To ensure scale invariance and stable optimization across months with differing cross-sectional sizes, the loss is normalized by the number of stocks in the cross-section.
Formally, letting \(f_t \in \mathbb{R}^{N_t}\) denote the vector of MLP predictions and \(R_{t+1} \in \mathbb{R}^{N_t}\) the realized excess returns, the loss is defined as \[\mathcal{L}_t
\;=\;
\frac{\bigl(1 - f_t^\top R_{t+1}\bigr)^2}{N_t}.\] This normalization ensures that each monthly training step contributes comparably to the overall optimization objective, regardless of fluctuations in the cross-sectional sample size.
9.1.0.6 Training window.
Throughout, we employ a rolling training window of length 60 months.
The PTK factors are derived from the gradients of the trained MLP with respect to its parameters. For a given month \(t\) with \(N_t\) stocks, let \(f(X_{i,t};
\theta_t^*)\) denote the MLP prediction for stock \(i\).
Gradient Computation: We compute the gradient of the normalized factor return (MSRR loss) with respect to the network parameters:
where \(R_{i,t+1}\) is the realized excess return and \(P\) is the total number of network parameters. Stacking these vectors over time produces the gradient matrix
Kernel Matrix: The kernel matrix \({\mathbb{K}}\) is computed as the uncentered Gram matrix of the gradients: \[{\mathbb{K}}= \frac{1}{T} G G' \in \mathbb{R}^{T \times
T}.\]
Ridge Regression: We solve for the portfolio weights using ridge regression on the kernel matrix with penalty parameter \(z\).
Hyperparameters:
Ridge penalties (\(z_{\mathrm{eff}}\)): 8 values logarithmically spaced from \(10^{-5}\) to \(100\).
Rolling windows: \(\{12, 60, 120, 240, 360\}\) months.
To provide a baseline for the PTK, we generate random-feature factors that approximate infinite-width (dimension \(P=25,000\)).
Weight Initialization: Random weights \(\Omega \in \mathbb{R}^{P \times d}\) are sampled from a scaled Gaussian distribution: \[\Omega_{ij} \sim \mathcal{N}\left(0,
\frac{2\sigma^2}{d}\right)\] where \(d\) is the input dimension and \(\sigma\) is a scaling parameter.
Activations: ReLU: \(\phi(x) = \max(0, x \Omega^T)\).
Factor Construction: The random-feature factors \(F_t\) are constructed as managed portfolios, scaled by \(N_t^{-1/2}\): \[F_{j,t} =
\frac{1}{\sqrt{N_t}} \sum_{i=1}^{N_t} \phi_j(x_{i,t}) R_{i,t+1}\]
To ensure the robustness of our results and account for the variability in initialization:
Random Seeds: All experiments are repeated across multiple random seeds (10 seeds).
For MLP models, the seed controls the initialization of the neural network weights.
For random-feature models, the seed determines the sampling of the random weight matrix \(\Omega\).
Ensembling: In our analysis, we often report the performance of an ensemble of models trained with different seeds. The ensemble return is computed as the volatility-weighted average of the individual seed returns: \[R_{ensemble, t} = \sum_{k=1}^{K} w_{k,t} R_{k,t}\] where \(w_{k,t} \propto 1/\widehat\sigma_{k,t}\) is the inverse rolling volatility of the \(k\)-th seed’s
strategy.
9.4 Performance Evaluation and Statistical Tests↩︎
We evaluate the strategies using the following metrics:
Alpha (\(\alpha\)): The intercept from the OLS regression of the strategy’s excess returns (\(R_{strat}\)) on the benchmark’s excess returns (\(R_{bench}\)): \[R_{strat, t} = \alpha + \beta R_{bench, t} + \epsilon_t\,,\] with Newey–West (1987) standard errors using appropriate lag length.
Sharpe Ratio: The annualized risk-adjusted return, calculated as \(\sqrt{12} \cdot \frac{\mu}{\sigma}\), where \(\mu\) and \(\sigma\)
are the mean and standard deviation of OOS monthly returns.
To characterize model complexity, overfitting, and generalization behavior, we analyze the spectral properties of the kernel matrix using the following statistics:
Implicit Regularization (\(Z_*(z)\)): The effective degrees of freedom of the kernel matrix \(K\) with ridge penalty \(z\): \[Z_*(z) = \frac{T}{\text{tr}((zI + K)^{-1})} = \frac{T}{\sum_{i=1}^T \frac{1}{z + \lambda_i}}\] where \(\lambda_i\) are the eigenvalues of \(K\) and \(T\) is the window size. We often report the normalized metric divided by \(\frac{1}{T}\text{tr}(K)\).
Debiased GRS Statistic (\(\hat{W}_{\text{debiased}}(z)\)): To correct for the bias in the standard GRS statistic when \(P/T\) is large, we compute: \[\hat{W}_{\text{debiased}}(z)\] using formula 24 .
Figure 9: Cumulative Returns of \(\bar R^{\rm PTK}\) from 31 , \(\bar R^{\textrm{ReLU}}\) from 28 , and \(R^{\rm DNN}\) from 29 . We use \(D=2,\;w=256\) and \(z_{eff}=10^{-5}.\) All returns are
normalized to equal full-sample variance.Figure 10: Machine Learning SDFs and Consumption Risk. This figure reports the relation between tradeable machine-learning stochastic discount factors (SDFs) and consumption-based SDFs from [2] (PJ) and [3] (MWZ). The top
row shows time-series correlations between portfolio returns from the fully trained DNN SDF \(R^{\mathrm{DNN}}\)29 (red), the random-feature benchmark \(\bar
R^{\mathrm{ReLU}}\)28 (grey, striped), and the PTK-based SDF \(\bar R^{\mathrm{PTK}}\)31 (blue), and the MWZ future-consumption SDF (left panel) and
the PJ long-run consumption SDF (right panel). The middle row reports \(t\)-statistics from univariate time-series regressions of each consumption-based SDF on the corresponding tradeable SDF return. The bottom row reports
the associated adjusted \(R^2\). All statistics are shown separately by size group. Consumption-based SDFs are constructed using future realized consumption growth aggregated over horizons of 84 months (MWZ) and 36 months
(PJ), respectively, which truncates the sample in June 2018. Ridge regularization with \(z_{eff}=10^{-5}\) is applied to the ReLU and PTK portfolios.
Didisheim, A., S. B. Ke, B. T. Kelly, and S. Malamud(2024): “APT or ‘AIPT’? The Surprising Dominance of Large Factor Models,”
Tech. rep., National Bureau of Economic Research.
[2]
Parker, J. A. and C. Julliard(2005): “Consumption Risk and the Cross Section of Expected Returns,”Journal of Political Economy, 113,
185–222.
[3]
Malamud, S., N. Wang, and Y. Zhang(2025): “Consumer Credit and Asset Prices,”Available at SSRN 4276714.
[4]
Cochrane, J. H.(2011): “Presidential address: Discount rates,”The Journal of finance, 66, 1047–1108.
[5]
Harvey, C. R., Y. Liu, and H. Zhu(2016): “… and the cross-section of expected returns,”The Review of Financial Studies, 29,
5–68.
[6]
McLean, R. D. and J. Pontiff(2016): “Does academic research destroy stock return predictability?”The Journal of Finance, 71,
5–32.
[7]
Kelly, B., S. Pruitt, and Y. Su(2020): “Characteristics are Covariances: A Unified Model of Risk and Return,”Journal of Financial
Economics.
[8]
Gu, S., B. Kelly, and D. Xiu(2020): “Autoencoder Asset Pricing Models,”Journal of Econometrics.
[9]
Hou, K., C. Xue, and L. Zhang(2020): “Replicating anomalies,”The Review of Financial Studies, 33, 2019–2133.
[10]
Kelly, B. and D. Xiu(2023): “Financial Machine Learning,”Working Paper.
[11]
Kozak, S., S. Nagel, and S. Santosh(2020): “Shrinking the cross-section,”Journal of Financial Economics, 135, 271–292.
[12]
Kelly, B. T., S. Malamud, M. Pourmohammadi, and F. Trojani(2023): “Universal Portfolio Shrinkage,”Swiss Finance Institute Research
Paper.
[13]
Kelly, B. T., S. Malamud, and K. Zhou(2022): “The Virtue of Complexity Everywhere,”Available at SSRN.
[14]
Kelly, B. T. and S. Malamud(2025): “Understanding the virtue of complexity,”Available at SSRN 5346842.
[15]
Chernov, M., B. T. Kelly, S. Malamud, and J. Schwab(2025): “A Test of the Efficiency of a Given Portfolio in High Dimensions,” Tech. rep.,
National Bureau of Economic Research.
[16]
Gibbons, M. R., S. A. Ross, and J. Shanken(1989): “A test of the efficiency of a given portfolio,”Econometrica: Journal of the Econometric
Society, 1121–1152.
[17]
Chen, L., M. Pelger, and J. Zhu(2024): “Deep learning in asset pricing,”Management Science, 70, 714–750.
[18]
——— (2020): “Empirical asset pricing via machine learning,”The Review of Financial Studies, 33, 2223–2273.
[19]
Fan, J., Z. T. Ke, Y. Liao, and A. Neuhierl(2022): “Structural Deep Learning in Conditional Asset Pricing,”Available at SSRN
4117882.
[20]
Jacot, A., F. Gabriel, and C. Hongler(2018): “Neural tangent kernel: Convergence and generalization in neural networks,”Advances in neural
information processing systems, 31.
[21]
Yang, G.(2020): “Tensor programs ii: Neural tangent kernel for any architecture,”arXiv preprint arXiv:2006.14548.
[22]
Zhou, Y., T. Pang, K. Liu, C. H. Martin, M. W. Mahoney, and Y. Yang(2023): “Temperature Balancing, Layer-wise Weight Analysis, and Neural Network
Training,”arXiv preprint arXiv:2312.00359.
[23]
Nakkiran, P., G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever(2021): “Deep double descent: Where bigger models and more data
hurt,”Journal of Statistical Mechanics: Theory and Experiment, 2021, 124003.
[24]
Belkin, M.(2021): “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation,”Acta
Numerica, 30, 203–248.
[25]
Ross, S. A.(1976): “The Arbitrage Theory of Capital Asset Pricing,”Journal of Economic Theory, 13, 341–360.
Rahimi, A. and B. Recht(2007): “Random Features for Large-Scale Kernel Machines.” in NIPS, Citeseer, vol. 3, 5.
[28]
Kelly, B. T., B. Kuznetsov, S. Malamud, and T. A. Xu(2025): “Artificial intelligence asset pricing models,” Tech. rep., National Bureau of
Economic Research.
[29]
Cong, L. W., G. Feng, J. He, and X. He(2025): “Growing the efficient frontier on panel trees,”Journal of Financial Economics, 167,
104024.
[30]
Long, P. M.(2021): “Properties of the after kernel,”arXiv preprint arXiv:2105.10585.
[31]
Wei, A., W. Hu, and J. Steinhardt(2022): “More than a toy: Random matrix models predict how real-world neural representations generalize,” in
International conference on machine learning, PMLR, 23549–23588.
[32]
Schwab, J., B. Kelly, S. Malamud, and T. A. Xu(2025): “Training NTK to Generalize with KARE,”arXiv preprint arXiv:2505.11347.
[33]
Chen, Z., B. Kelly, and S. Malamud(2025): “Limits To (Machine) Learning,”arXiv preprint arXiv:2512.12735.
[34]
Jensen, T. I., B. Kelly, and L. H. Pedersen(2023): “Is there a replication crisis in finance?”The Journal of Finance, 78,
2465–2518.
[35]
Yang, G., E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao(2022): “Tensor programs v: Tuning large
neural networks via zero-shot hyperparameter transfer,”arXiv preprint arXiv:2203.03466.
[36]
Bansal, R. and A. Yaron(2004): “Risks for the long run: A potential resolution of asset pricing puzzles,”The journal of Finance, 59,
1481–1509.
[37]
Campbell, J. Y., S. Giglio, C. Polk, and R. Turley(2018): “An intertemporal CAPM with stochastic volatility,”Journal of Financial
Economics, 128, 207–233.
Bryan Kelly is at Yale School of Management, AQR Capital Management, and NBER; www.bryankellyacademic.org. Boris Kuznetsov is at AQR Capital Management. Semyon Malamud is at the Swiss
Finance Institute, EPFL, and CEPR, and is a consultant to AQR. Yuan Zhang is at the Shanghai University of Finance and Economics. Semyon Malamud gratefully acknowledges the financial support of the Swiss Finance Institute and the Swiss National Science
Foundation, Grant 100018-228042. AQR Capital Management is a global investment management firm that may or may not apply similar investment techniques or methods of analysis as described herein. The views expressed here are those of the authors and not
necessarily those of AQR. This work was supported by a grant from the Swiss National Supercomputing Center (CSCS) under project ID lp46. We thank Mohammad Pourmohammadi for his excellent comments and suggestions.↩︎
Here, “ridgeless” does not mean unregularized. The PTK-SDF remains strongly disciplined through the spectral complexity of the learned factors, which endogenously shrinks weak risk directions even in the absence of an explicit penalty.↩︎
Training such large networks may require alternative optimization schemes or architecture-adjusted hyperparameters; see [35].
We leave this issue for future work.↩︎
Results without principal-component projection are weaker, indicating that lower-variance PTK components primarily capture variation orthogonal to consumption risk. See Figure 10 in the Appendix.↩︎
That is, \(|\phi(x)-\phi(y)|\le C |x-y|\) for some \(C>0.\)↩︎
Subjects
Statistical Finance (q-fin.ST)
Computational Engineering, Finance, and Science (cs.CE)