Construction of Nuclear Covariant Energy Density Functional from A Physics-Guaranteed Neural Network Approach


Abstract

Density functional theory is a practical approach for solving quantum many-body problems with available computational resources. The complexity of the nuclear force makes constructing an accurate nuclear energy density functional much more challenging. The feasibility of constructing a nuclear covariant energy density functional with deep neural networks is demonstrated. This physics-guaranteed neural network approach achieves high accuracy in predicting nuclear energy density and exhibits significantly better extrapolation abilities than traditional machine learning methods for binding energies. When combined with the existing covariant density functional, the neural network approach improves the binding energy accuracy from \(644\) keV to \(86\) keV in the known region and also effectively captures the microscopic shell effect. Furthermore, its extrapolation performance is also significantly enhanced, achieving an accuracy of approximately \(5\) MeV even when extrapolating up to \(30\) steps. This work paves the way for the construction of accurate nuclear energy density functionals through machine learning.

GBKsong

1 Introduction↩︎

The exact solution of quantum many-body problems is a significant challenge in various fields of physics [1], including nuclear physics, condensed matter physics, and quantum chemistry. Density functional theory (DFT) provides a practical approach for solving this problem with available computational resources [2], [3], with great success in areas such as materials science [4], [5], quantum chemistry [6], and condensed matter physics [7]. The Hohenberg-Kohn theorem guarantees the existence of such a functional [8]. In practice, the formulation is commonly implemented through the Kohn-Sham equations [9], wherein an interacting many-body system is mapped onto a non-interacting reference system, which provides the same ground-state density with a local Kohn-Sham potential. In nuclear physics, the complexity of the nuclear force makes it much more challenging to construct an accurate nuclear energy density functional. The accuracy of existing energy density functionals is generally greater than \(1\) MeV in the known region, whereas the deviations among them even reach tens of MeV when extrapolated to the region near the neutron-drip line [10]. Consequently, the development of an accurate nuclear energy density functional has become one of the central topics in nuclear theory.

Nuclear covariant density functional theory (CDFT) originates from relativistic quantum field theory, in which a system of Dirac nucleons interacts in a relativistic covariant manner via the meson and the photon fields [11], [12]. Compared with the non-relativistic DFT, the CDFT includes not only inherent Lorentz covariance, but also the automatic incorporation of the nucleon spin degrees of freedom, a natural explanation for the pseudospin symmetry in nucleon spectrum [13][18] and the spin symmetry in antinucleon spectrum [18][20], as well as the inherent inclusion of nuclear magnetic potential [21], among others. The CDFT has attracted much attention and successfully describes many nuclear properties [22], establishing itself as one of the most prominent microscopic frameworks in nuclear theory [23][29].

In recent years, machine learning (ML) methods have provided powerful tools in physics, including particle physics [30][32], condensed matter physics [33], [34], and astrophysics [35], [36]. The ML methods have also been used to predict various nuclear properties, often achieving higher predictive accuracy than those obtained from conventional theoretical models [37][39]. The ML methods have also demonstrated significant potential in enhancing the accuracy of DFT functionals [40]. In electronic structure calculations, the ML has been successfully employed to develop functional forms with high accuracy and transferability, even addressing challenges that are difficult for conventional approaches, such as fitting molecular dissociation curves and describing strongly correlated systems [41], [42]. The DM21 functional, developed by DeepMind, utilizes a neural network architecture trained on large-scale datasets and achieves performance superior to most human-designed functionals in molecular systems, while also exhibiting strong generalization abilities [43]. These advancements demonstrate that ML can not only automatically extract complex features from data but also effectively incorporate physical constraints, thereby driving innovation in functional design [42], [44]. The complexity of nuclear force makes it much more difficult to construct an accurate DFT for self-bound nuclear systems. In this work, we attempt to construct a nuclear energy density functional directly using ML methods. Such ML models guaranteed by physics theory may be promising to achieve higher accuracy and better extrapolation abilities.

The CDFT can be constructed with the relativistic mean-field (RMF) model in the meson-exchange or point-coupling representations. The point-coupling RMF model provides more explicit expressions of energy density in terms of nucleon densities, while the relativistic covariant advantages in CDFT still remain [45][49]. In this letter, a successful attempt for constructing a nuclear covariant energy density functional in the point-coupling representation is made with deep neural networks. This framework establishes a novel paradigm for ML applications in nuclear physics and is promising to achieve better prediction and generalisation abilities. In the following, the basic formulas of CDFT and the architecture of the ML method will be briefly introduced. We will then investigate the feasibility, as well as the prediction and generalization abilities of this physics-guaranteed ML method by comparing it with traditional ML methods.

2 Theoretical framework↩︎

The starting point of the relativistic point-coupling model is the effective Lagrangian density \[\begin{align} \label{Eq:Lagrangian} {\cal L} = {\cal L}_{\rm free}+{\cal L}_{\rm 4f}+{\cal L}_{\rm hot}+{\cal L}_{\rm der}+{\cal L}_{\rm em}, \end{align}\tag{1}\] where \({\cal L}_{\rm free}\), \({\cal L}_{\rm 4f}\), \({\cal L}_{\rm hot}\), \({\cal L}_{\rm der}\), and \({\cal L}_{\rm em}\) represent the Lagrangian densities of the free nucleons, the four-fermion point-coupling terms, the higher-order terms, the derivative terms, and the electromagnetic interaction terms, respectively. Taking the mean-field and “no-sea” approximations, one can get the expectation value \(\langle\Phi|H|\Phi\rangle\) of Hamiltonian \(H\) corresponding to the Lagrangian density [45][47]. Minimizing \(\langle\Phi|H|\Phi\rangle\) with respect to the wave functions, the ground-state binding energy of a nuclear system can be obtained \(B_{\rm CDFT} = -E_{\rm CDFT}\), where \[\begin{align}\label{Eq:er} E_{\rm CDFT} &= \int d^3r {\cal E}(\emph{\boldsymbol{r}}) \\ &= \int d^3r \left\{{\cal E}_{\rm kin}(\emph{\boldsymbol{r}}) + {\cal E}_{\rm int}(\emph{\boldsymbol{r}}) + {\cal E}_{\rm em}(\emph{\boldsymbol{r}})\right\}. \end{align}\tag{2}\] \({\cal E}_{\rm kin}\), \({\cal E}_{\rm int}\), and \({\cal E}_{\rm em}\) represent the kinetic, interaction, and electromagnetic parts of the total energy density \({\cal E}\), respectively. The even-even nuclei with \(8\leqslant Z \leqslant 100\) and \(8\leqslant N \leqslant 156\) are studied in this work, for which the space-like components of the currents vanish due to the time-reversal invariance. The remaining isoscalar-scalar density \(\rho_S\), isoscalar-vector density \(\rho_V\), isovector-vector density \(\rho_{TV}\) are used to describe the energy densities in Eq. (2 ). The CDFT with PC-PK1 parameter set provides an accurate and reliable description of the isospin dependence of nuclear binding energy [47], which has been widely used for studying nuclear properties [50][53]. The deep neural network is employed to construct a nuclear CDFT by taking the PC-PK1 results as the training data, which means that the CDFT with the PC-PK1 is taken as the target DFT in this work, so the deep neural network constructed in this work could serve as a DFT emulator of PC-PK1. The spherical approximation is used in the calculations of the CDFT with PC-PK1, so the densities can then be simplified as functions of the radial coordinate \(r\), which is discretized with \(r_i = 0.1 \times i~{\rm fm}~(i=0,1,...,150)\). The details can be seen in Supplemental Material [54].

Three classes of neural networks with different input layers are constructed to describe nuclear energy \({E}_{\rm CDFT}\) and energy density \({\cal E}_{\rm CDFT}\): The first class (NN1) employs the traditional neural network architecture, i.e., the neural network with the inputs \(I = (N - Z) / (N + Z)\) and \(P = \nu_p\nu_n/(\nu_p+\nu_n)\) besides \(Z\) and \(N\), where \(\nu_p\) and \(\nu_n\) are the differences between the nucleon numbers \(Z\) and \(N\) and the nearest magic numbers (\(8\), \(20\), \(28\), \(50\), \(82\) for protons and \(8\), \(20\), \(28\), \(50\), \(82\), \(126\), \(184\) for neutrons) [55]. The inputs of the neural networks in the second class (NN2) are various densities at all mesh points. The third class (NN3) takes various densities at the coordinate \(r\) as inputs. A composite loss function \(L_{\rm total}\) is designed for the networks with the output of energy density \({\cal E}_{\rm CDFT}\), which is \[\begin{align} L_{\rm total} = a L_E + L_{\cal E}, \end{align}\] with, \[\begin{align} L_E &= \frac{1}{n}\sum_{j=1}^n(E_{{\rm CDFT},j}-E_{{\rm NN},j})^2,\\ L_{\cal E} &= \frac{1}{151n}\sum_{j=1}^n\sum_{i=0}^{150}[ ({\cal E}_{{\rm CDFT},j}(r_i)-{\cal E}_{{\rm NN},j}(r_i))\cdot r_i^2]^2, \end{align}\] where \(a\) is the weight factor, \(L_E\) and \(L_{\cal E}\) are employed to assess the loss of the binding energy and energy density with respect to training data, and \(n\) is the number of nuclei in the learning set. A reasonable parameter \(a\) ensures that the neural network can well describe binding energy and energy density simultaneously, which is automatically adjusted by balancing \(L_E\) and \(L_{\cal E}\). For the neural networks with the output of energy \(E\), the loss function is \(L_E\). The input variables, output variables, and loss functions of various neural networks used in this work are given in Table 1. The neural networks with the output of \({\cal E}\), i.e., NN2-I3\({\cal E}\) , NN2-I6\({\cal E}\), NN3-I3\({\cal E}\), and NN3-I6\({\cal E}\) construct nuclear covariant energy density functionals. Comparing with the NN2 using the global densities as input features, the weight-sharing architecture of NN3 guarantees that the neural network parameters are the same for the densities at different spatial coordinates. Furthermore, the current neural network architecture is implemented using a fully differentiable framework. By utilizing automatic differentiation, the functional derivative of the total energy with respect to the nucleon densities can be obtained efficiently. This differentiability ensures that the proposed framework is theoretically feasible with the variational principle in density functional theory and can be directly integrated into future self-consistent calculations. Therefore, the physics-guaranteed covariant density functional is realized by NN3 through its architectural design and the underlying differentiable programming. The Adaptive Moment Estimation (Adam) [56] optimizer from PyTorch [57] is used for parameter optimization. The network framework and the training process for these neural networks are described in detail in Supplemental Material [54].

Table 1: Input variables, output variables, and loss functions of various neural networks. “size:1*151” or “size:1*1” indicates that the input/output variable is density at all mesh points or density at the coordinate \(r\). All the densities (\(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\), and \({\cal E}\), where \(\Delta\) is the Laplace operator) involved are multiplied by \(r^2\) to improve the performance of neural networks. The results of \(\rho_x\), \(\Delta\rho_x\) (\(x = S, V, TV\)) are calculated by the CDFT with PC-PK1 [47].
 Neural network   Input variables   Output variables   Loss function 
NN1-I4E \(Z\), \(N\), \(I\), \(P\) (size:4*1) \(E\) (size:1*1) \(L_E\)
NN2-I3E \(\rho_S\), \(\rho_V\), \(\rho_{TV}\) (size:3*151) \(E\) (size:1*1) \(L_E\)
NN2-I6E \(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\) (size:6*151) \(E\) (size:1*1) \(L_E\)
NN2-I3\({\cal E}\) \(\rho_S\), \(\rho_V\), \(\rho_{TV}\) (size:3*151) \({\cal E}\) (size:1*151) \(L_{\rm total}\)
NN2-I6\({\cal E}\) \(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\) (size:6*151) \({\cal E}\) (size:1*151) \(L_{\rm total}\)
NN3-I3\({\cal E}\) \(\rho_S\), \(\rho_V\), \(\rho_{TV}\) (size:3*1) \({\cal E}\) (size:1*1) \(L_{\rm total}\)
NN3-I6\({\cal E}\) \(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\) (size:6*1) \({\cal E}\) (size:1*1) \(L_{\rm total}\)
NN1-I4\(\delta\)E \(Z\), \(N\), \(I\), \(P\) (size:4*1) \(E-E_{\rm PC-PK1}\) (size:1*1) \(L_E\)
NN2-I6\(\delta\)E \(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\) (size:6*151) \(E-E_{\rm PC-PK1}\) (size:1*1) \(L_E\)
NN2-I6\(\delta{\cal E}\) \(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\) (size:6*151) \({\cal E}-{\cal E}_{\rm PC-PK1}\) (size:1*151) \(L_{\rm total}\)
NN3-I6\(\delta{\cal E}\) \(\rho_S\), \(\rho_V\), \(\rho_{TV}\), \(\Delta\rho_S\), \(\Delta\rho_V\), \(\Delta\rho_{TV}\) (size:6*1) \({\cal E}-{\cal E}_{\rm PC-PK1}\) (size:1*1) \(L_{\rm total}\)

The nuclei in the known region, i.e., the nuclei with the experimental masses in AME2020 [58], are randomly divided into \(20\) different groups, each containing a learning set (\(80\%\) of the data) and a validation set (\(20\%\) of the data) to reduce potential bias arising from a single randomization. This repeated sampling approach ensures statistical robustness in model development and performance evaluation. Each of these \(20\) groups is trained using five random initial parameters. Neural networks are employed to predict nuclear energy densities and binding energies, and the final predictions are derived from the statistical mean of \(100\) independent realizations.

3 Results and Discussion↩︎

Figure 1: (Color online) Rms deviations \sigma_{\rm rms}(E) of the binding energies predicted by various neural networks with respect to the target data from the calculations with PC-PK1 as a function of the minimum distance d to the isotopes in the known region.

To investigate the extrapolation abilities of different neural networks, the root-mean-square (rms) deviations \(\sigma_{\rm rms}(E)\) of the predicted binding energies with respect to the target data from the PC-PK1 calculations are shown in Fig. 1 as a function of the minimum distance \(d\) to the isotopes in the known region. \(N = 156\) is the maximum neutron number in all learning sets, so only the nuclei with \(N \leqslant 156\) are included in our studies. Clearly, \(\sigma_{\rm rms}(E)\) generally increases as the \(d\) increases for all networks. Compared to NN1, which is widely used in nuclear mass predictions, NN2 and NN3 exhibit much better extrapolation abilities, which indicates that more physics can be extracted by the neural network from various densities. Notably, NN3-I6\({\cal E}\) employing point-to-point mapping achieves the both optimal prediction accuracy with the \(\sigma_{\rm rms}(E)\) of \(644\) keV in the known region and best extrapolation ability by constructing an energy density functional, i.e., the density \(\rho_{S,V,TV}\) and \(\Delta\rho_{S,V,TV}\) at different \(r\) uses the same neural network to get the energy density \({\cal E}(r)\). The best extrapolation ability implies that this neural network architecture can better capture the essential physics correlations in the description of nuclear properties. In principle, this approach could also be used to directly learn the experimental binding energies and then generate a new functional form. However, this requires extending the approach to incorporate deformation degrees of freedom and further introduce various corrections beyond the mean-field approximation, such as vibrational and rotational corrections. Therefore, constructing a new nuclear energy density functional by learning the experimental binding energies remains a significant challenge at present. The densities \(\rho_{S,V,TV}\) are clearly not completely independent of each other, while it is intractable in defining the specific relationship among them. Once the relationship among the densities \(\rho_{S,V,TV}\) is established, the energy density functional constructed in this work can obtain ground-state observables through the variational of density, such as charge radii and density distribution. Due to the complex relationship among the densities \(\rho_{S,V,TV}\), this work calculates the ground-state energies based on the constructed functional and the known ground-state densities instead of solving the ground-state densities via variational methods.

Figure 2: (Color online) Benchmarking of the nuclear energy density predicted by various neural networks against the target PC-PK1 calculations for ^{78,96}Ge. Panels (a,b), (c,d), and (e,f) correspond to the r^2-weighted energy densities r^2{\cal E}, the differences of r^2{\cal E} between PC-PK1 calculations and the neural network predictions, i.e., r^2 \delta{\cal E}, and the cumulative integral of r^2\delta{\cal E} from 0 to r, respectively. The histograms located between coordinates r_1 and r_2 in the panels (g, h) represent the subsection integral of r^2\delta{\cal E} from r_1 to r_2.

Taking \(^{78}\)Ge (a nucleus in the known region) and \(^{96}\)Ge (a nucleus extrapolated by \(10\) steps beyond the known region) as examples, Fig. 2 compares the predictions of energy densities based on four different neural networks. It is clear that NN3-I6\({\cal E}\) best reproduces the target energy densities for both \(^{78}\)Ge and \(^{96}\)Ge, while the other neural networks show remarkable deviations from the PC-PK1 energy densities, particularly near the extrema and inflection points of energy densities. To highlight these behaviors, panels (c) and (d) show the differences \(r^2\delta {\cal E}\) \((\delta{\cal E} = {\cal E}_{\rm PC-PK1} - {\cal E}_{\rm NN})\) between the PC-PK1 calculations and the neural network predictions. There are positive and negative \(r^2\delta {\cal E}\) in different \(r\), so the differences between the binding energies predicted by the PC-PK1 and the neural networks can be reduced due to cancellations. The cumulative integrals of the energy density differences \(E(r)=4\pi\int_0^r r^2\delta {\cal E} dr\) are shown in panels (e) and (f). For \(^{78}\)Ge, \(E(r)\) generally first decreases to a minimum value within \(-6\) MeV, then increases, and gradually stabilizes after \(r \gtrsim 6\) fm. However, the minimum value of \(E(r)\) for NN3-I3\({\cal E}\) is approximately \(-20\) MeV, and it stabilizes only when \(r \gtrsim 8\) fm. For \(^{96}\)Ge, \(E(r)\) fluctuates considerably over the entire \(r\) space. Even when \(r \gtrsim 9\) fm, where the predicted energy density \({\cal E}\) closely matches that of PC-PK1, the contribution of \(r^2\delta {\cal E}\) to the total energy deviation still non-negligible partially due to the amplification of the factor \(r^2\). Therefore, all the densities involved in the neural networks are multiplied by \(r^2\) to improve the performance of the neural networks in this work. This also underscores the importance of accurately describing the tail of the energy density in predicting nuclear binding energies. Panels (g) and (h) show the subsection integral of \(r^2\delta{\cal E}\) from \(r_1\) to \(r_2\), i.e., \(4\pi\int_{r_1}^{r_2} r^2\delta {\cal E} dr\). There are generally larger contribution of \(r^2\delta {\cal E}\) to the total energy deviation at where the \(r^2\)-weighted energy density \(r^2{\cal E}\) is larger. However, these contributions tend to largely cancel each other out, resulting in a relatively small net deviation. The NN3-I6\({\cal E}\) consistently exhibits small deviations across all intervals, both in known and unknown regions. This further supports the feasibility of using neural networks to construct nuclear density functional, demonstrating its significant advantage in describing both nuclear energy densities and binding energies.

Figure 3: (Color online) Same as Fig. 1, but the PC-F1 binding energies are taken as the target data.

Numerous studies have shown that machine learning methods can significantly improve the predictive accuracy of nuclear mass models by learning the differences between experimental masses and theoretical results [55], [59][62]. Following this idea, we take the PC-F1 density functional results as the target data [46] and construct four neural networks to describe the differences of binding energies or the energy densities, whose input and output variables are given in the last four rows of Table 1. Figure 3 shows the rms deviation \(\sigma_{\rm rms}(E)\) of neural network predictions with respect to the PC-F1 binding energies as a function of the minimum distance \(d\) to the isotopes in the known region. Also shown for comparison are the rms deviations \(\sigma_{\rm rms}(E)\) between the PC-F1 and PC-PK1 predictions, which exceed 3 MeV even in the known region. All four neural networks can well describe the differences between PC-F1 and PC-PK1 predictions in the known region. The \(\sigma_{\rm rms}(E)\) are reduced to \(0.198\), \(0.118\), \(0.086\), and \(0.353\) MeV for NN1-I4\(\delta\)E, NN2-I6\(\delta\)E, NN2-I6\(\delta{\cal E}\), and NN3-I6\(\delta{\cal E}\), respectively. In the unknown region, the \(\sigma_{\rm rms}(E)\) of the NN1-I4\(\delta\)E gradually approaches that of the PC-PK1 as \(d\) increases. The neural networks with density information, i.e., NN2-I6\(\delta\)E, NN2-I6\(\delta{\cal E}\), and NN3-I6\(\delta {\cal E}\) achieve better extrapolation abilities, whose rms deviations from the target data are approximately \(5\) MeV even when extrapolating up to \(30\) steps. Therefore, the neural networks can well describe the energy density differences between two theoretical models even extrapolating to the unknown region, which provides a foundation for learning the differences between experimental data and theoretical predictions in the future.

Figure 4: (Color online) Comparison of the target energies E_{\rm CDFT} from the PC-F1 calculations with the PC-PK1 and NN3-I6\delta{\cal E} predictions. (a) Differences between the target energies E_{\rm CDFT} and the PC-PK1 predictions. (b) Same as panel (a) but for the NN3-I6\delta{\cal E} predictions. (c) Rms deviation \sigma_{\rm rms}(E) of target energies E_{\rm CDFT} with respect to the PC-PK1 and NN3-I6\delta{\cal E} predictions as a function of N. (d) Same as panel (c) but as a function of Z. The solid line shows the boundary of nuclei with known masses in AME2020 [58] and the dashed lines denote the traditional magic numbers.

Figure 4 shows the comparison of the target binding energies from the PC-F1 calculations with the PC-PK1 and NN3-I6\(\delta{\cal E}\) predictions. As shown in Fig. 4 (a), PC-F1 predicts larger binding energies than the PC-PK1, with the deviations exceeding \(4\) MeV for the neutron-deficient nuclei with \(Z \gtrsim 50\). On the neutron-rich side, the PC-F1 predicts smaller binding energies than the PC-PK1, with the deviations of larger than \(20\) MeV for the nuclei near the neutron-drip line. From Fig. 4 (b), it is clear that the NN3-I6\(\delta{\cal E}\) significantly improves the description of binding energies, even for the nuclei near the drip lines. It reproduces the target binding energies generally within \(1\) MeV in the known region and around \(6\) MeV when extrapolating to the neutron-drip line. Figures 4 (c) and (d) show the rms deviations of binding energies between NN3-I6\(\delta{\cal E}\) predictions and target data as a function of \(N\) and \(Z\), respectively. Clearly, the NN3-I6\(\delta{\cal E}\) significantly improves the description of binding energies for the medium and heavy mass nuclei in both the known and unknown regions, achieving comparable accuracy across the light, medium, and heavy mass regions. Surprisingly, NN3-I6\(\delta{\cal E}\) shows better agreement with the target data in the unknown regions near the magic numbers. Further study finds that NN3-I6\(\delta{\cal E}\) well describes the abrupt decrease in two-neutron separation energies after the magic numbers, indicating that the microscopic shell effect can be effectively captured by the NN3-I6\(\delta{\cal E}\). Therefore, the neural network approach is a powerful tool for constructing a high-precision nuclear energy density functional if there are enough experimental data to train it.

The uncertainty of the neural network predictions can be assessed using the standard deviation of the \(100\) independent results. It is found that the corresponding uncertainty increases as the distance from the known region increases, which is an inherent property of data-driven models. A quantitative discussion of the uncertainty in the NN3-I6\({\cal E}\) and NN3-I6\(\delta {\cal E}\) predictions is provided in Supplemental Material [54]. Bayesian methods are powerful tools for model selection and uncertainty quantification. The Bayesian analysis of nuclear dynamics framework provides the conceptual scaffolding for model comparison and averaging in nuclear many-body calculations, e.g., nuclear masses and reaction cross sections [63]. Moreover, Bayesian model averaging (BMA) has been successfully used to resolve disagreements between different density functionals when extracting the nuclear symmetry energy [64]. Furthermore, BMA has been widely used to improve nuclear mass predictions and quantify extrapolation uncertainties toward the drip lines [65], [66]. More recently, integrating of Bayesian model mixing with a multireference energy density functional has demonstrated significant improvements in predictive accuracy of two-neutron separation energy compared to individual theoretical models [67]. Therefore, it is promising to quantify uncertainties of our results with the Bayesian method though more computing resources are indispensable.

4 Summary and Perspectives↩︎

The feasibility of constructing a nuclear covariant energy density functional using deep neural networks is demonstrated. By employing point-to-point mapping neural network architecture, this neural network approach guaranteed by the density functional theory achieves high accuracy in predicting both nuclear energy densities and binding energies. It also exhibits significantly superior extrapolation abilities compared to traditional neural network models. When combined with the existing covariant density functional, the binding energy accuracy is improved from \(644\) keV to \(86\) keV in the known regions, and the microscopic shell effect can also be effectively captured. Even when extrapolating up to \(30\) steps beyond the learning region, the deviation remains about \(5\) MeV, highlighting its robust generalization performance. Furthermore, this method significantly improves the description of binding energies for the medium and heavy mass nuclei, achieving comparable accuracy across the light, medium, and heavy mass regions. Although the present theoretical calculations are used as the benchmark, this approach could in principle construct a new functional form by training directly on experimental binding energies, which need the current framework be extended to incorporate deformation degrees of freedom and various beyond-mean-field corrections. Fortunately, the point-to-point mapping nature of our neural network architecture is naturally suited for generalization to deformed nuclei. This can be achieved by replacing the current one-dimensional density with two- or three-dimensional density in deformed nuclei directly, or by utilizing basis expansion coefficients of density to replace the density at coordinate space. This would establish a solid foundation for the data-driven construction of next-generation, high-precision energy density functionals in nuclear theory.

Acknowledgements. This work was partly supported by the National Natural Science Foundation of China under Grants No. 12375109, No. 11875070 and No. 11935001, the Anhui project (Z010118169), and the Key Research Foundation of Education Ministry of Anhui Province (2023AH050095).

References↩︎

[1]
N. H. March, W. H. Young, and S. Sampanthar, The Many-Body Problem in Quantum Mechanics, Cambridge at the University Press (1967).
[2]
R. M. Dreizler and E. K. U. Cross, Density Functional Theory: An Approach to the Quantum Many-Body Problem, Springer-Verlag (1990).
[3]
W. Kohn, Rev. Mod. Phys. 71, 1253 (1999).
[4]
A. G. Souza Filho, V. Meunier, M. Terrones et al., Nano Lett. 7, 2383 (2007).
[5]
P. Wisesa, M. Li, M. T. Curnan, G. H. Gu, J. W. Han, J. C. Yang, and W. A. Saidi, Nano Lett. 25, 1329 (2025).
[6]
N. Mardirossian and M. Head-Gordon, Mol. Phys. 115, 2315 (2017).
[7]
J. Tao and Y. Mo, Phys. Rev. Lett. 117, 073001 (2016).
[8]
P. Hohenberg and W. Kohn, Phys. Rev. 136, B864 (1964).
[9]
W. Kohn and L. J. Sham, Phys. Rev. 140, A1133 (1965).
[10]
M. Bender, P.-H. Heenen, and P.-G. Reinhard, Rev. Mod. Phys. 75, 121 (2003).
[11]
J. Meng, H. Toki, S. G. Zhou, S. Q. Zhang, W. H. Long, and L. S. Geng, Prog. Part. Nucl. Phys. 57, 470 (2006).
[12]
J. Meng, Relativistic Density Functional for Nuclear Structure, World Scientific (2016).
[13]
J. N. Ginocchio, Phys. Rev. Lett. 78, 436 (1997).
[14]
J. Meng, K. Sugawara-Tanabe, S. Yamaji, P. Ring, and A. Arima, Phys. Rev. C 58, R628 (1998).
[15]
J. Meng, K. Sugawara-Tanabe, S. Yamaji, and A. Arima, Phys. Rev. C 59, 154 (1999).
[16]
T. S. Chen, H. F. Lü, J. Meng, S. Q. Zhang, and S. G. Zhou, Chin. Phys. Lett. 20, 358 (2003).
[17]
J. N. Ginocchio, Phys. Rep. 414, 165 (2005).
[18]
H. Z. Liang, J. Meng, and S. G. Zhou, Phys. Rep. 570, 1 (2015).
[19]
S. G. Zhou, J. Meng, and P. Ring, Phys. Rev. Lett. 91, 262501 (2003).
[20]
X. T. He, S. G. Zhou, J. Meng, E. G. Zhao, and W. Scheid, Eur. Phys. J. A 28, 265 (2006).
[21]
W. Koepf and P. Ring, Nucl. Phys. A 493, 61 (1989).
[22]
P. Ring, Phys. Scr. T150, 014035 (2012).
[23]
P. Ring, Prog. Part. Nucl. Phys. 37, 193 (1996).
[24]
D. Vretenar, A. V. Afanasjev, G. A. Lalazissis, and P. Ring, Phys. Rep. 409, 101 (2005).
[25]
T. Nikšić, D. Vretenar, and P. Ring, Prog. Part. Nucl. Phys. 66, 519 (2011).
[26]
J. Meng, J. Peng, S. Q. Zhang, and P. W. Zhao, Front. Phys. 8, 55 (2013).
[27]
J. Meng and S. G. Zhou, J. Phys. G: Nucl. Part. Phys. 42, 093101 (2015).
[28]
S. G. Zhou, Phys. Scr. 91, 063008 (2016).
[29]
S. H. Shen, H. Z. Liang, W. H. Long, J. Meng, and P. Ring, Prog. Part. Nucl. Phys. 109, 103713 (2019).
[30]
P. Baldi, P. Sadowski, and D. Whiteson, Nat. Commun. 5, 4308 (2014).
[31]
L. G. Pang, K. Zhou, N. Su, H. Petersen, H. Stöcker, and X. N. Wang, Nat. Commun. 9, 210 (2018).
[32]
J. Brehmer, K. Cranmer, G. Louppe, and J. Pavez, Phys. Rev. Lett. 121, 111801 (2018).
[33]
J. Carrasquilla and R. G. Melko, Nat. Phys. 13, 431 (2017).
[34]
G. Carleo and M. Troyer, Science 355, 602 (2017).
[35]
F. Villaescusa-Navarro, D. Anglés-Alcázar, S. Genel, D. N. Spergel, R. S. Somerville, R. Dave, A. Pillepich, L. Hernquist, D. Nelson, Paul Torrey et al., Astrophys. J. 915, 71 (2021).
[36]
F. Villaescusa-Navarro, S. Genel, D. Anglés-Alcázar, L. Thiele, R. Dave, D. Narayanan, A. Nicola, Y. Li, P. Villanueva-Domingo, B. Wandelt et al., Astrophys. J. Suppl. Ser. 259, 61 (2022).
[37]
P. Bedaque, A. Boehnlein, M. Cromaz, M. Diefenthaler, L. Elouadrhiri, T. Horn, M. Kuchera, D. Lawrence, D. Lee, S. Lidia et al., Eur. Phys. J. A 57, 100 (2021).
[38]
A. Boehnlein, M. Diefenthaler, N. Sato, M. Schram, V. Ziegler et al., Rev. Mod. Phys. 94, 031003 (2022).
[39]
W. B. He, Q. F. Li, Y. G. Ma, Z. M. Niu, J. C. Pei, and Y. X. Zhang, Sci. China Phys. Mech. Astron. 66, 282001 (2023).
[40]
R. Pederson, B. Kalita, and K. Burke, Nat. Rev. Phys. 4, 357 (2022).
[41]
F. Brockherde, L. Vogt, L. Li, M. E. Tuckerman, K. Burke, and K.-R. Müller, Nat. Commun. 8, 872 (2017).
[42]
L. Li, S. Hoyer, R. Pederson, R. Sun, E. D. Cubuk, P. Riley, and K. Burke, Phys. Rev. Lett. 126, 036401 (2021).
[43]
J. Kirkpatrick, B. McMorrow, D. H. P. Turban, A. L. Gaunt, J. S. Spencer, A. G. D. G. Matthews, A. Obika, L. Thiry, M. Fortunato, D. Pfau et al., Science 374, 1385 (2021).
[44]
R. Nagai, R. Akashi, and O. Sugino, npj Comput. Mater. 6, 43 (2020).
[45]
B. A. Nikolaus, T. Hoch, and D. G. Madland, Phys. Rev. C 46, 1757 (1992).
[46]
T. Bürvenich, D. G. Madland, J. A. Maruhn, and P.-G. Reinhard, Phys. Rev. C 65, 044308 (2002).
[47]
P. W. Zhao, Z. P. Li, J. M. Yao, and J. Meng, Phys. Rev. C 82, 054319 (2010).
[48]
Q. Zhao, Z. X. Ren, P. W. Zhao, and J. Meng, Phys. Rev. C 106, 034315 (2022).
[49]
Z. X. Liu, Y. H. Lam, N. Lu, and P. Ring, Phys. Lett. B 842, 137946 (2023).
[50]
P. W. Zhao, L. S. Song, B. H. Sun, H. Geissel, and J. Meng, Phys. Rev. C 86, 064324 (2012).
[51]
Q. S. Zhang, Z. M. Niu, Z. P. Li, J. M. Yao, and J. Meng, Front. Phys. 9, 529 (2014).
[52]
K. Q. Lu, Z. X. Li, Z. P. Li, J. M. Yao, and J. Meng, Phys. Rev. C 91, 027304 (2015).
[53]
K. Y. Zhang, X. T. He, J. Meng, C. Pan, C. W. Shen, and S. Q. Zhang, Phys. Rev. C 104, L021301 (2021).
[54]
See Supplemental Material at https://xxx for detailed information on the CDFT with a point-coupling interaction and the neural network architecture.
[55]
Z. M. Niu and H. Z. Liang, Phys. Lett. B 778, 48 (2018).
[56]
D. P. Kingma and J. L. Ba, arXiv:1412.6980 [nucl-th].
[57]
PyTorch documentation, https://pytorch.org/docs/stable/index.html.
[58]
M. Wang, W. J. Huang, F. G. Kondev, G. Audi, and S. Naimi, Chin. Phys. C 45, 030003 (2021).
[59]
R. Utama, J. Piekarewicz, and H. B. Prosper, Phys. Rev. C 93, 014311 (2016).
[60]
Z. M. Niu and H. Z. Liang, Phys. Rev. C 106, L021303 (2022).
[61]
X. H. Wu, L. H. Guo, and P. W. Zhao, Phys. Lett. B 819, 136387 (2021).
[62]
Y. Y. Guo, T. Yu, X. H. Wu, C. Pan, and K. Y. Zhang, Phys. Rev. C 110, 064310 (2024).
[63]
D. R. Phillips, R. J. Furnstahl, U. Heinz, T. Maiti, W. Nazarewicz, F. M. Nunes, M. Plumlee, M. T. Pratola, S. Pratt, F. G. Viens, and S. M. Wild, J. Phys. G: Nucl. Part. Phys. 48, 072001 (2021).
[64]
M. Qiu, B.-J. Cai, L.-W. Chen, C.-X. Yuan, and Z. Zhang, Phys. Lett. B 849, 138435 (2024).
[65]
V. Kejzlar, L. Neufcourt, W. Nazarewicz, and P.-G. Reinhard, J. Phys. G: Nucl. Part. Phys. 47, 094001 (2020).
[66]
L. Neufcourt, Y. Cao, S. Giuliani, W. Nazarewicz, E. Olsen, and Oleg B. Tarasov, Phys. Rev. C 101, 014319 (2020).
[67]
A. Sharma, N. Schunck, and K. Wendt, Phys. Rev. Research 7, 023179 (2025).