June 04, 2026
Many areas of astrophysics, including exoplanetary studies, rely on precise and accurate stellar parameters. This demands that uncertainties on these parameters truly reflect all biases and systematics. Within this second work of the
sec:gr8stars collaboration, we take a set of 585 bright FGK dwarfs with high resolution, high signal-to-noise ratio spectra from the SOPHIE spectrograph. We determine stellar effective temperature, surface gravity, and metallicity using five
different spectroscopic methods for each star, with an additional method used for comparisons. We find a typical scatter of 76 K in \(T_{\rm eff}\), 0.14 dex in \(\log\,g\), and 0.07 dex in
\(\rm [Fe/H]\). These deviations are significantly larger than the average precision error on these parameters. We furthermore use isochrone fitting to determine mass, radius, and age for all 585 stars, using input from all
results. We use the radii determined by SED fitting in the first sec:gr8stars paper as a comparison to our isochronal radii from this work, in addition to comparing the isochronal \(\log\,g\) to spectroscopic
\(\log\,g\). The scatter in mass and radius from the use of different spectroscopic methods is investigated and propagated to exoplanetary parameters. The induced fractional uncertainties in planetary radius (\(\lesssim\) 3 %) and mass (\(\lesssim\) 5%) are found to be below those typically found in the literature. We estimate a lower limit on planetary equilibrium temperature fractional uncertainty of
\(\approx\) 4%, a noise floor that is currently not sufficiently represented in the literature.
techniques: spectroscopic – stars: atmospheres – stars: fundamental parameters – stars: solar-type
Solar-like stars, defined in this work as FGK main-sequence stars, play the vital role of the host stars in the search for and study of Earth-like exoplanets. To date, no ‘Earth twin’ exoplanet – being an Earth-like planet orbiting a solar-like star with an orbital period on the order of Earth’s – has been identified. Many studies [1]–[8] demonstrate that the lack of such a discovery is likely not due to an inherent lack of Earth twins, rather instead due to limitations in the disentangling of small planetary signals from the signals of solar-like stars.
Ongoing projects, such as the EXPRES 100 Earths Survey [9] on the Lowell Discovery Telescope [10], [11] and the NEID Earth Twin Survey [12], [13], in addition to upcoming missions such as the Terra Hunting Experiment (THE) [14] on HARPS-3 [15], PLATO [16], [17], and the Second Earth Initiative on the Second Earth Spectrograph on the MPG/ESO 2.2m telescope (2ES;
[18]), aim to overcome the current hurdles in instrumentation to kickstart the discovery and characterisation of Earth twins. In the first
sec:gr8stars paper ([19], henceforth GR8-1), we introduced our sample of 5645 bright, main sequence FGKM stars that form the
sec:gr8stars catalogue, and published homogeneously derived stellar parameters for all Northern-hemisphere targets with available high resolution spectra.
The increasing precision of modern, high-resolution spectrographs has in turn increased the demand for precise and accurate stellar parameters. However, it has been established that quoted precision uncertainties often fail to represent the extent of systematic effects in spectroscopic analyses. There have been numerous studies that show the significant systematics introduced in stellar parameters by the use of different methodologies, model atmospheres, analysis pipelines, and line lists [20]–[23]. Such systematics frequently exceed the quoted uncertainties indicating that they present a fundamental limitation in the derivation of stellar parameters. With this context in mind, we must carefully treat the definition of stellar parameter precision. Instead, agreement between independent spectroscopic methods may provide a more realistic uncertainty estimate than internal precision errors.
The characterisation of spectroscopic systematics has been investigated through previous collaborative endeavours. Through the Gaia-ESO survey, multiple analysis pipelines were combined to derive stellar parameters with realistic uncertainties [24]–[26]. Studies using the iSpec framework investigate controlled comparison between differing techniques [27]–[29]. Additionally, large-scale analyses such as the PASTEL catalogue [30], and investigations of homogeneous parameter determination [31], [32] highlight how diverse literature methodologies are, and the agreement expected between them. Further to this, comparisons of spectroscopic surveys for a common set of stars [33] reinforce the presence of systematic offsets between both observations and methods.
Despite the breadth of studies performed throughout the literature, gaps still remain. Particularly, studies often rely on heterogeneous data sets, differing instrumental qualities, or samples that may only partially overlap. Consequently, there is a necessity for analyses in which multiple state-of-the-art methods are applied to the same set of high-quality spectra. This allows for the isolation of differences arising solely from systematic differences in methodology. Furthermore, although spectroscopic systematics in terms of atmospheric stellar parameters has been widely discussed throughout the literature, their propagation to further derived parameters such as stellar mass, radius, and age, is not fully quantified. This is particularly important in the context of exoplanet studies, as uncertainties in stellar parameters directly propagate to the planetary mass, radius, and density.
In the sec:gr8stars catalogue, we see a uniquely suited framework for such investigations. While in GR8-1 we presented homogenously derived stellar parameters for a large sample of bright FGKM dwarfs, which is essential for a variety of
applications, the data set does not fully capture the systematic uncertainties introduced by modelling choices. In this second paper, the focus is on quantifying these systematics.
A subset of 585 FGK dwarfs observed solely with the SOPHIE spectrograph is employed in this work. We apply multiple independent spectroscopic analysis techniques spanning a range of methodology, model choices, and assumptions, to each of these stars. By comparing the resultant stellar atmospheric parameters, we are able to characterise the agreement between these methods and establish estimates for the scatter introduced by methodological differences. This allows for us to place reliable constraints on the limits of precision in spectroscopic analysis, and to assess the suitability of typically quoted uncertainties to reflect true levels of agreement between independent methods.
The sec:gr8stars catalogue is fully described in Section 2 in terms of its selection, and context in the wider fields of exoplanetary and stellar physics. Section 3 describes the
spectroscopic and photometric data used in this work, along with the instruments employed for the observations. Section 4 details the methods used to determine all sets of stellar parameters presented in this work, with the
results and comparisons explored in Section 5. Section 6 explores the effects of scatter in stellar parameters on the parameters of hosted exoplanets. We conclude our findings in Section 7.
The sec:gr8stars catalogue is a collection of 5645 bright, FGKM stars. As part of GR8-1, we published homogeneously derived atmospheric parameters for 1716 targets in the Northern sample, in addition to photometric stellar parameters for
all targets in the Northern sample. We aim to assemble, and make publicly available, a full complement of uniformly formatted, high resolution spectra, in addition to homogeneously derived stellar atmospheric parameters for all 5645 stars.
As detailed in GR8-1, three requirements were used in the selection of the sec:gr8stars catalogue, with the addition of a declination cut of \(\geq \ang{-15}\) to select the Northern sample that GR8-1
investigates. The first such requirement is that all of the stars are brighter than a G band magnitude of 8, meaning archival and future observations will be to a high SNR (signal-to-noise ratio) standard. The focus of sec:gr8stars on FGKM
dwarfs required both earlier type main-sequence stars and giants/subgiants to be excluded. We achieved these exclusions by requiring the Gaia BP-RP colour to be at least 0.6. We used the Dartmouth isochrones and mass tracks [34] to ensure our catalogue does not include evolved stars, as detailed in GR8-1.
GR8-1 published the first data set as part of sec:gr8stars, the main product of which was the uniformly-formatted spectra and homogeneous stellar atmospheric parameters3. Effective temperature (\(T_{\rm eff}\)), surface gravity (\(\log\,g\)), metallicity (\(\rm [Fe/H]\)), microturbulent
velocity (\(v_{\rm mic}\)), macroturbulent velocity (\(v_{\rm mac}\)), and projected rotational velocity (\(v \sin i_{\star}\)) were provided as part of the
GR8-1 data set for all stars with available spectra. The sec:gr8stars catalogue provides a robust and extensive resource for the stellar and exoplanet communities, with plans to extend to all targets in the catalogue through period
spectroscopic archive searches and observing proposals.
As described in GR8-1, spectra were taken with 6 different high-resolution spectrographs as part of sec:gr8stars. For our method comparison in this work, we use only the spectra from the SOPHIE spectrograph. This was done to ensure no
scatter is introduced by the use of different instruments. The SOPHIE spectrograph covers the most stars in the current version of the catalogue. We additionally removed spectroscopic binaries, variable stars (as per their main SIMBAD object type), and
targets with \(v \sin i_{\star}\), as measured by PAWS [35] in GR8-1, larger than 10 \(\rm km\,s^{-1}\). These cuts were made to ensure we limit the amount of blended spectral lines, which severely impact the equivalent widths (EW) method, and spectral features induced by variability.
The SOPHIE spectrograph (Spectrographe pour l’Observation des Phénoménes des Intérieurs stellaires et des Exoplanètes), on the 1.93 m reflector telescope at the Observatoire Haute Provence, has a wavelength coverage of 387.2 nm to 694.3 nm [36], [37]. As part of sec:gr8stars, we use only observations conducted in
the High Resolution (HR) mode, with a resolving power of \(R \approx\) 75 000. Of the 1286 targets observed with SOPHIE in the sec:gr8stars catalogue, 585 can be classified as single, slow-rotating stars by
using the SIMBAD object classifications and limiting \(v \sin i_{\star}\)to 10 \(\rm km\,s^{-1}\) as measured by [19]. This is the subset analysed in this paper.
Within this work, we employ six different state-of-the-art spectroscopic analysis methods. Reflecting the variety in analysis methods in the literature, these six span not only different methodologies, but also differing atmospheric models, spectral line lists, and assumptions. A summary of the methods and their properties can be found in Table 1.
| Analysis Code | Method Type | Free Parameters | Radiative Transfer Code | Line List | Model Atmosphere |
|---|---|---|---|---|---|
| ARES+MOOG | Equivalent Widths | \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v_{\rm mic}\) | MOOG | [38], [39] | Kurucz [40] |
| FASMA | Spectral Synthesis | \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v \sin i_{\star}\) | MOOG | [41] | MARCS |
| Metal Pipe | Spectral Synthesis | \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v \sin i_{\star}\) | MOOG | Linemake [42] | Phoenix |
| PAWS | Equivalent Width + Spectral Synthesis | \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v_{\rm mic}\), \(\log\,g\) | WIDTH, SPECTRUM | SPECTRUM | ATLAS9 |
| webSME | Spectral Synthesis | \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v_{\rm mic}\), \(v_{\rm mac}\), \(v \sin i_{\star}\) | [43] | Gaia-ESO | MARCS |
| YARARA | Equivalent Widths and Machine Learning | \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\) | N/A | [44] | N/A |
The atmospheric stellar parameters (\(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\) and, \(v_{\rm mic}\)) were derived using the ARES+MOOG methodology described in [45]–[47]. For this we used ARES v24 [48], [49] to consistently measure the EW of selected iron lines for each stellar spectrum, across the entire wavelength range. For this, we used the iron line list presented in [38], however, when we find a temperature that is lower than 5200 K, we switch the list of lines to one that is more appropriate for cooler stars [39]. The best spectroscopic parameters are found by converging into ionisation and excitation equilibrium. In this process, it is used for a grid of Kurucz model atmospheres [40] and the radiative transfer code MOOG [50]. We also derived a more accurate trigonometric surface gravity using the Gaia DR3 data following the same procedure as described in [47].
FASMA5 [41], [51] uses the radiative transfer code MOOG (version 2019, [50]) to compute synthetic spectra on-the-fly and determine stellar parameters through a Levenberg–Marquardt, non-linear, least-squares optimisation. The adopted line list is mostly comprised of iron lines from [41], in the wavelength regions 5399 – 5619 Å and 6470 – 6790 Å. The atomic line data, namely the oscillator strengths, were obtained from the line list of the Gaia-ESO survey [52]. For the rest of the lines, we kept the atomic data obtained either from the VALD database [53] or calibrated empirically from observed spectra (of the Sun and Arcturus, see [41] for details). The damping parameters are based on the ABO theory [54] when available, or in any other case, we used the Unsold approximation [55]. FASMA performs local normalisation for the adopted intervals and a cleaning process for cosmic rays before matching with the synthetic spectrum. We used the standard solar abundances from [56] internally adopted by MOOG, except for iron, for which we adopted A(Fe) = 7.45 dex. The model atmospheres are the grids of the MARCS libraries [57]. The uncertainties on the stellar parameters are derived from the covariance matrix constructed by the non-linear least-squares fit. FASMA delivers \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), and \(v \sin i_{\star}\)while microturbulence and macroturbulence are refined in a second iteration based on empirical calibrations from [39] and [58], respectively.
Metal Pipe [59] is a C++ and Python framework that derives stellar abundances using line-by-line spectral synthesis. To get stellar parameters (\(T_{\rm eff}\) and \(\log\,g\)), it fits a star’s photometry to MIST isochrone models [60] of a given input metallicity (to start, Metal Pipe assumes a metallicity of [M/H] = 0). It then interpolates an appropriate model atmosphere from a grid of Phoenix models [61]. Using this model atmosphere, MOOG [50] is called to generate synthetic spectral lines, which are fit to the observed stellar spectra using a weighted \(\chi^2\) minimisation algorithm. Metal Pipe operates on the entire wavelength range of the spectra.
First, Metal Pipe fits only iron lines. If the output metallicity (derived from synthetic spectral fitting of iron lines) does not match the input metallicity (the metallicity of the model atmosphere), the process iterates. Metal Pipe interpolates a new model atmosphere with a metallicity matching the last output value (i.e. the input metallicity of the next iteration is based on the output metallicity from the previous iteration). Again, Metal Pipe fits the iron lines to get an output metallicity and tests whether it is consistent with the input value. This iteration continues until convergence is reached.
Once [Fe/H] is converged, it fits alpha-enhancement ([\(\alpha\)/Fe]), using a similar method with calcium and titanium lines. Once both [Fe/H] and [\(\alpha\)/Fe] are converged, it fits a number of other elements as specified by the user. At the time of writing, Metal Pipe supports fitting for abundances of C, O, Na, Mg, Al, Si, S, Ca, Ti, and Fe.
Atmospheric parameters were determined as part of GR8-1 using the PAWS pipeline, introduced by [35]. The stellar parameters
presented in this work as derived by PAWS are directly taken from GR8-1. PAWS uses the iSpec framework [27], [62] to employ both the curve-of-growth EW and spectral synthesis methods. For both methods, we used the ATLAS9 model
atmospheres [63]. Initial stellar parameters (\(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v_{\rm mic}\)) were determined using the EW method with the WIDTH radiative transfer code [64]. This set of initial parameters was used as an input to the spectral synthesis method, employing the SPECTRUM radiative transfer code [65] to determine the final stellar atmospheric parameters (\(T_{\rm eff}\), \(\log\,g\), \(\rm
[Fe/H]\), \(v_{\rm mic}\), \(v_{\rm mac}\), \(v \sin i_{\star}\)). Spectral synthesis as part of PAWS covers the wavelength range 480 -
680 nm, using the SPECTRUM line list based on the NIST atomic database [66].
SME (Spectroscopy Made Easy) is a well-established radiative-transfer and fitting code originally coded in IDL and C++ [67], [68]. It runs an adapted and updated version of the line-formation code SYNTH [69] on a pre-computed grid of plane-parallel or spherical MARCS model atmospheres [57]. SME’s frontend and line-formation wrapper was recently recoded in Python [43] and is now referred to as PySME. webSME6 is a server-based interface to PySME that offers several modes. We employ two of webSME’s modes in the current work: the least-squares fitting mode and the forward-modelling mode. These, and the additional modes, are described in Puschnig et al. (A&A, submitted). Output spectra can be downloaded as fits files for further offline comparisons/calculations.
The forward-modelling mode produces a synthetic spectrum for a given set of input parameters: spectral resolution, \(T_{\rm eff}\), \(\log\,g\), metallicity, \(v_{\rm mic}\), \(v_{\rm mac}\), \(v \sin i_{\star}\), using one of several predefined line lists (here the atomic Gaia-ESO line list, [52]) and one of several predefined solar reference compositions (here [70]). As long as the stellar parameters fall within the grid of MARCS model atmospheres, any combination of parameters can be simulated. We note in passing that the grid of MARCS model atmospheres was computed assuming the solar reference abundances of [71]. By assuming [70] for the line-formation abundance zero point, a small inconsistency is introduced in the calculations.
The fitting mode allows the user to upload a normalised spectrum and choose any number of stellar parameters (see above) to be determined using a least-squares metric. For the present work, we focus on the determination of \(\log\,g\)from the Mg I\(b\) lines assuming that the other stellar parameters are known (from runs with the other codes discussed in this section). When strictly fitting only this region, changes in Mg abundance can mimic \(\log\,g\) variations, producing a linear degeneracy – see [72]. To avoid this, we adopt a two step method using webSME. Initially, the full spectrum is fitted with Mg abundance free, constraining it using all spectral information. Secondly, we fix the Mg abundance and perform the synthesis again, limited to the region 5000 - 6000 Å. This approach allows the \(\log\,g\) to be constrained using a region of the spectrum that contains the highly sensitive Mg I\(b\) lines.
Unless a specific Mg abundance is either determined or input, this procedure rests on the scaling of Mg with overall metallicity which works best at solar metallicity and may depart from the scaling relation by as much as 0.2 dex for metal-poor thin-disk stars at \(\rm [Fe/H]\) =\(-0.5\) [73]. For the tests focusing on the six benchmark stars, this potential bias is considered to be of minor importance. The fitting mode needs to make many syntheses and thus takes significantly longer to produce output than the forward-modelling mode.
YARARA is a post-processing methodology initially designed to extract more precise radial velocities by the analysis of spectra time-series [74], [75]. A by-product of the code is the production of a master stellar spectrum free of most known instrumental systematics, including telluric lines. In this context, the pipeline has been recently upgraded [44] in order to derive the stellar atmospheric parameters7 (namely \(T_{\rm eff}\), \(\log\,g\) and \(\rm [Fe/H]\)). For this work, we did not apply the YARARA corrections on SOPHIE spectra, but only used the atmospheric parameter retrieval recipe on them.
The methodology relies on the extraction of 16 EWs from different chemical species (see Appendix E in [44]) after the spectra have been continuum normalised by RASSINE [76]. These spectral lines are all encompassed by the wavelength region 4446 – 6753 Å. The choice of using several species was motivated by the release of some degeneracies between atmospheric parameters [77]. A mapping function \(F\) from \(\mathbb{R}^{16}\) to \(\mathbb{R}^3\): \[F(EW_1,EW_2,...,EW_{16}) = (\text{T_{\rm eff}}, \text{\log\,g},\text{\rm [Fe/H]})\] was then fitted using a non-linear approach such as machine learning regressions. In order to fit the function \(F\), the full HARPS database was reduced to extract the EWs and a collection of several catalogs of stellar atmospheric parameters was obtained from different methods and instruments (see references in [44]). Because the function \(F\) was calibrated using HARPS spectra, we first established the scaling coefficient between the HARPS and SOPHIE EWs that accounts for the different spectral resolution of both instruments. The coefficient was obtained from a star that had been observed with both instruments.
By definition, the YARARA predicted parameters can be understood as the average answer that would have been provided by the cited catalogues/methods. The estimated uncertainties on the different parameters were given by the test sample and are \(\pm\) 70 K, \(\pm\) 0.07 dex, and \(\pm\) 0.07 dex on \(T_{\rm eff}\), \(\log\,g\) and \(\rm [Fe/H]\) respectively. As a drawback of the methodology, the machine learning regressor struggles to predict extreme values for stellar observations that were not in the training sample (such as stars cumulating anomalous values in \(T_{\rm eff}\), \(\log\,g\) or \(\rm [Fe/H]\)).
Using the stellar atmospheric parameters in conjunction with parallax and magnitudes, we used the isochrones [78]
package to determine masses, radii and ages for the sec:gr8stars targets by following the methodology outlined by [79]. As inputs, we
used \(T_{\rm eff}\) and \(\rm [Fe/H]\) from our spectroscopic analyses, in addition to the Gaia DR3 parallax, 2MASS J, H, and K magnitudes [80], and AllWISE W1, W2, and W3 magnitudes [81]. Note that we
did not include spectroscopic \(\log\,g\) in this analysis. Although accurate spectroscopic \(\log\,g\) measurements are achievable, there are potential biases and offsets depending on the
method used at a significance that would be relevant for isochrone fitting [82]–[84]. The use of photometric magnitudes combined with parallax instead provides a more accurate constraint on stellar radius, and hence \(\log\,g\). We employed the MESA Isochrones and
Stellar Tracks (MIST, [60]) stellar evolution model. In this work we present both the individual masses, radii, and ages from running isochrones on the results from each spectroscopic method, and also an inverse-variance-weighted mean of the MAP (Maximum a posteriori) values from all methods. The latter should automatically account for systematic differences
between spectroscopic methods, but still relies on only one set of stellar evolution models which can have systematics of its own. As a result, this work establishes only a lower limit on the estimated systematics.
To understand the systematic differences and accuracies of derived stellar parameters, we compare here the stellar parameters derived from the codes to each other, in addition to comparing to literature results.
From all spectroscopic methods, we obtain values for \(T_{\rm eff}\), \(\log\,g\), and \(\rm [Fe/H]\), in addition to propagating these through isochrones to determine radius, mass, and age. Given the results of all spectroscopic methods, we can estimate the typical scatter on all six of these stellar parameters that arises due to differences in methodology and models. To do so, we calculate the median value of each parameter for each star across all sets of results. Then, for each star, we calculate the RMSD (Root Mean Square Deviation) across the different methods from this median value. The overall scatter is then taken as the median RMSD across all stars. Table 2 displays the results of this, showing the typical scatter that we find in \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), radius, mass, and age.
| Parameter | Typical Scatter |
|---|---|
| \(T_{\rm eff}\)(K) | 76 |
| \(\log\,g\)(dex) | 0.14 |
| \(\rm [Fe/H]\)(dex) | 0.07 |
| Radius (R\(_{\odot}\)) | 0.02 (1.22 %) |
| Mass (M\(_{\odot}\)) | 0.05 (5.50 %) |
| Age (Gyr) | 2.08 (\(>\) 100 %) |
The floors in systematic uncertainties identified in Table 2 correspond to approximately 1.22 % in stellar radius, 5.50 % in stellar mass, and over 100 % in stellar age. These values are, broadly, consistent with the precision requirements for PLATO stellar radius and mass8. This indicates that spectroscopic method differences alone are unlikely to represent the dominant limitation in stellar mass and radius determinations for PLATO targets. In contrast to this, the scatter in age is substantially larger, demonstrating the strong sensitivity of isochrone derived ages to scatter in the employed atmospheric parameters.
The PASTEL catalogue [30] provides a useful resource of stellar atmospheric parameters, collating results from a variety of sources and methods. Within this work we use the 2020-01-30 version of the catalogue for comparison. We note that the \(\rm [Fe/H]\) measurements included in the PASTEL catalogue are all directly from high resolution, high SNR spectroscopy, making them relevant comparison values. The \(T_{\rm eff}\) measurements are predominantly from spectroscopy and photometry, with a small number having \(T_{\rm eff}\) determined from interferometry. In terms of \(\log\,g\), most values result from the ionisation-equilibrium and parallax methods, however the PASTEL catalogue does include \(\log\,g\) values determined through asteroseismology. It is unlikely that any of the stars within this work have a PASTEL \(\log\,g\) measured using asteroseismology, given that we focus on only FGK dwarfs. Given that the median values of the atmospheric parameters, and their associated uncertainties, represent both the consensus and spread in values at the current state of the field, we expect our results in this work to reflect those from PASTEL. Throughout this work, the PASTEL values used for comparison are the median of all PASTEL values for the star. We interpret the inherent scatter in the PASTEL catalogue as the standard deviation present in the results collated for each target. Across all targets in this work with multiple sources of parameters, we see median standard deviations in \(T_{\rm eff}\), \(\log\,g\), and \(\rm [Fe/H]\) of 52 K, 0.09 dex, and 0.04 dex respectively in the PASTEL catalogue.
Figure 1 shows histograms of the differences between atmospheric parameters in this work and those from PASTEL, split into \(T_{\rm eff}\), \(\rm [Fe/H]\), and \(\log\,g\) as the top, middle, and bottom rows, respectively. For each plot, we would expect an unbiased, accurate method to produce a histogram of uncertainties that is Gaussian and centred around 0 difference, with a FWHM (Full Width at Half Maximum) reflective of the typical uncertainties in the methods. This is observed in the majority of cases in Figure 1, with a few notable exceptions. We show the RMSD, median difference, and median uncertainty for all results in Table 3.
Observing the difference between \(T_{\rm eff}\) derived in this work and that from PASTEL in Figure 1, the RMSD is between 89 and 102 K for all targets. This follows investigations performed by [32], in which the PASTEL catalogue was analysed for the Gaia-ESO benchmark stars. From this, it was found that for stars with three or more spectroscopic determinations of \(T_{\rm eff}\) showed a mean standard deviation in \(T_{\rm eff}\) of \(\sim\) 85 K. We repeat such an investigation for the stars in this work with three or more PASTEL entries. This totals 310 stars for \(T_{\rm eff}\), giving a median spread (maximum – minimum) of 131 K. Across the 202 stars with three or more \(\log\,g\) and \(\rm [Fe/H]\) entries, we see a median spread of 0.20 dex in \(\log\,g\), and 0.12 dex in \(\rm [Fe/H]\). Comparing these to the scatter we see in our results through Figure 1 and Table 3, we see that our scatter and that seen in PASTEL is of a similar magnitude.
In Figure 2, we plot the mean results from all analysis codes against the PASTEL values for \(T_{\rm eff}\), \(\log\,g\), and \(\rm [Fe/H]\), respectively. The most striking trend here is the systematically lower \(\log\,g\) values returned in this work compared to the PASTEL values. Given that PASTEL results come from various sources, not limited to spectroscopy, we can assume that this is a spectroscopic effect. The mean difference between our mean \(\log\,g\) and that from PASTEL is -0.09 dex. However, this lies within the mean standard deviation we see in our results of 0.12 dex, suggesting that uncertainties due to method-dependent scatter outweigh this systematic underestimation.
When we consider the median differences shown in Table 3, we see that ARES+MOOG, FASMA, and YARARA results show median differences at absolute values larger than their respective median uncertainties. This could indicate a systematic offset, however as the uncertainties are precision based, and do not account for model inaccuracies, it is likely that uncertainties adjusted for this would encompass the observed median difference. This is further supported by the fact that both PAWS and Metal Pipe have median uncertainties reflective of their scatter (RMSD), whereas scatter significantly outweighs uncertainty for ARES+MOOG and FASMA.
Regarding \(\rm [Fe/H]\), we can see from Table 3 that the RMSD is larger than the median uncertainty for all results, indicating underestimated uncertainties in all cases. We can additionally observe that the median difference between \(\rm [Fe/H]\) from ARES+MOOG and Metal Pipe and that from PASTEL is larger than the median uncertainty.
By far the most obvious issues in Figure 1 lie with the \(\log\,g\). Underestimation of \(\log\,g\) is clear from both PAWS and ARES+MOOG, with median differences of -0.22 and -0.16 dex, respectively, exceeding the median uncertainties in both cases. Both PAWS and ARES+MOOG also show RMSD at scales larger than three times the median uncertainty, strongly suggesting that uncertainties are underestimated for the resulting \(\log\,g\). Both the median differences and RMSD in the results suggest \(\log\,g\) derived from the spectroscopic methods within this work requires larger uncertainties than purely precision based ones, to account for methodological and model differences. In terms of \(T_{\rm eff}\) and \(\rm [Fe/H]\), we see far more significant standard deviations in our work of 64 K and 0.061 dex respectively, than their mean differences from PASTEL of 11 K and 0.001 dex. It is clear that for \(T_{\rm eff}\) and \(\rm [Fe/H]\), method-to-method scatter is a significant source of uncertainty.
| Parameter | RMSD | Median Difference | Median Uncertainty | ||||||||||||
| P | A+M | F | Y | MP | P | A+M | F | Y | MP | P | A+M | F | Y | MP | |
| \(\Delta\) \(T_{\rm eff}\)(K) | 94 | 94 | 90 | 89 | 101 | -8 | 34 | 51 | -14 | -27 | 105 | 30 | 26 | 70 | 93 |
| \(\Delta\) \(\rm [Fe/H]\) | 0.08 | 0.10 | 0.06 | 0.08 | 0.12 | 0.00 | 0.04 | -0.02 | -0.02 | 0.04 | 0.04 | 0.02 | 0.03 | 0.07 | 0.01 |
| \(\Delta\) \(\log\,g\) | 0.29 | 0.19 | 0.17 | 0.13 | 0.12 | -0.22 | -0.16 | -0.05 | 0.01 | -0.02 | 0.12 | 0.05 | 0.09 | 0.07 | 0.04 |
We further investigate differences between \(T_{\rm eff}\) from this work and that in the PASTEL catalogue in Figure 3 by plotting these against the PASTEL \(T_{\rm eff}\), coloured by the \(\rm [Fe/H]\) from each method. Each plot is repeated below coloured by the PASTEL \(\rm [Fe/H]\), to allow for an alternative comparison independent of method. For each comparison of \(T_{\rm eff}\) difference to PASTEL \(T_{\rm eff}\), we calculated the Pearson correlation coefficient (PCC) to identify any trends. Out of all methods, only FASMA shows little-to-no correlation, with a PCC of 0.02. Additionally, ARES+MOOG is the only method showing a positive correlation, with a value of 0.39. whereas PAWS, YARARA, amd Metal Pipe have negative correlations of -0.37, -0.35, and -0.40, respectively. This evaluation suggests that FASMA determines \(T_{\rm eff}\) without bias relative to the PASTEL values across the FGK regime. Limiting the sample to solar and cooler stars, based on the PASTEL \(T_{\rm eff}\), the correlations for the latter three results change significantly. For PAWS and Metal Pipe, the PCCs become -0.05 and 0.01, respectively, in the solar and cooler regime, suggesting an improved suitability here as opposed to with hotter stars. Interestingly, the YARARA \(T_{\rm eff}\) difference becomes positively correlated with PASTEL \(T_{\rm eff}\) at PCC of 0.22.
Figure 3 additionally shows trends with \(\rm [Fe/H]\). Particularly prevalent here are the results from Metal Pipe. Colouring the points in this plot by both the Metal Pipe and PASTEL \(\rm [Fe/H]\) values reveals an influence of \(\rm [Fe/H]\) on \(T_{\rm eff}\), or vice versa. This is not unexpected, given that changes in one stellar parameter can be at least partially compensated by apparent changes in another, particularly when using small numbers of lines [77]. High metallicity targets tend to have an overestimated \(T_{\rm eff}\) compared to PASTEL when determined by Metal Pipe, whereas lower metallicity stars see an underestimated \(T_{\rm eff}\) compared to PASTEL. Figure 3 is replicated for \(T_{\rm eff}\) and \(\log\,g\) in Appendix 8.
As detailed in Section 4.2, we used the spectroscopic \(T_{\rm eff}\) and \(\rm [Fe/H]\), together with photometric magnitudes, to determine mass
and radius for each star, amongst other parameters. Figure 4 shows histograms of the resulting stellar mass distributions from PAWS, ARES+MOOG, FASMA, YARARA, and Metal Pipe, in addition to the mass distribution from the combined isochronal posterior. This shows that our sample is weighted towards stars heavier
than solar, with all distributions peaking higher than solar mass. This is to be expected, given that the mass tracks used to select the whole sec:gr8stars sample (see GR8-1) retain metal rich F dwarfs, however not metal poor ones. This biases
the F type population, and therefore in part the entire sec:gr8stars population, to heavier stars.
In terms of stellar ages, typical uncertainties on the isochronal findings are on the order of \(\approx\) 30% [85], owing to the slow timescales of change of observables. Isochrones for different ages are very close for main sequence stars. We see these uncertainties reflected in our own results from isochrones in this work. The distribution of stellar ages determined through isochrones and spectroscopic inputs is shown in Figure 5. The individual distributions from the use of different spectroscopic methods follow a broad shape that peaks towards a region between 1 and 5 Gyr, however the value of this peak varies significantly between the codes used. For any given star, we see a typical scatter in the determined age of 2.08 Gyr solely from the use of \(T_{\rm eff}\) and \(\rm [Fe/H]\) inputs from different spectroscopic methods.
For mass, radius, and age, all difference in results comes from the scatter in \(T_{\rm eff}\) and \(\rm [Fe/H]\) from the spectroscopic methods. As shown in Table 2, we see scatter in mass of 0.05 M\(_{\odot}\), radius of 0.02 R\(_{\odot}\), and age of 2.08 Gyr. It is important to contextualise these in terms of the typical precision uncertainties we see on these parameters using isochrones. Across all of our results, we see a median uncertainty in stellar mass of 0.03 M\(_{\odot}\), stellar radius of 0.01 R\(_{\odot}\), and stellar age of 1.66 Gyr. Given that these uncertainties are lower than the scatter we observe when using spectroscopic inputs from different methods, we suggest that these do not fully reflect the true uncertainties in these parameters.
A useful tool for our method comparison is the availability of stellar parameters for benchmark stars. [86] provide interferometric fundamental \(T_{\rm eff}\) and radii measurements for 192 Gaia FGK benchmark stars. While these measurements have minimal reliance on spectroscopy or evolutionary models, they do not represent complete model independence. The bolometric fluxes used are determined through SED fitting, and limb-darkening models are applied to the stellar angular diameters. Cross-matching with [86], we find a small overlap of 6 targets that are also in our sample of 585 stars in this paper.
Figure 6 displays the difference in \(T_{\rm eff}\) between the spectroscopic methods in this work, and the interferometry from [86] for these six stars. We see overall agreement between all methods and the interferometric effective temperatures. Across all methods, the mean difference from the interferometric effective temperature is 78 K, sitting comfortably within the interferometric mean uncertainty in \(T_{\rm eff}\) of 135 K. Considering the differences in \(T_{\rm eff}\) between the methods used in this paper, we see a mean dispersion across the six stars of 211 K. [30] compare the fundamental \(T_{\rm eff}\) values for 151 stars to those from photometry and spectroscopy, finding a median absolute deviation of 96 K. Our results here are therefore within the broad expectations of comparing spectroscopic to interferometric \(T_{\rm eff}\).
An additional observation from Figure 6 is that for \(\beta\) CVn we see strong agreement between \(T_{\rm eff}\) results from our spectroscopic methods, however at an offset when compared to the fundamental \(T_{\rm eff}\) from [86]. All spectroscopic results are within 60 K of each other for \(\beta\) CVn, hence indicating that the spectroscopic methods agree for this star, however differ somewhat significantly compared to the interferometric \(T_{\rm eff}\). In terms of other spectroscopic parameters, we see a standard deviation from the six codes of 0.17 dex in \(\log\,g\), and 0.18 dex in \(\rm [Fe/H]\).
In performing a literature search for \(\beta\) CVn, we found two other interferometric \(T_{\rm eff}\) measurements. [87] gives a \(T_{\rm eff}\) of 5653 \(\pm\) 72 K, while [88] measures 5966 \(\pm\) 117 K, both values cooler than the 6013 \(\pm\) 91 K obtained from [86]. Of particular note is the \(T_{\rm eff}\) from [87] being cooler than both other interferometric values by over 300 K. Differences in interferometric values could result from uncertainty in the derivation of angular diameter, or differing treatments of limb darkening and bolometric flux determination. Given the small sample size, it is not possible to truly disentangle systemic effects in spectroscopic methods from those that are inherent in the interferometric analysis.
| Source | \(\beta\) CVn \(T_{\rm eff}\)(K) |
|---|---|
| Spectroscopy (This Work) | |
| webSME | 5872 \(\pm\) 279 |
| ARES+MOOG | 5863 \(\pm\) 20 |
| FASMA | 5866 \(\pm\) 19 |
| PAWS | 5841 \(\pm\) 101 |
| YARARA | 5901 \(\pm\) 70 |
| Metal Pipe | – |
| Literature Interferometry | |
| [86] | 6013 \(\pm\) 91 |
| [87] | 5653 \(\pm\) 72 |
| [88] | 5966 \(\pm\) 117 |
| Literature Spectroscopy | |
| [89] | 5897 |
| [90] | 5806 \(\pm\) 24 |
| [91] | 5875 \(\pm\) 30 |
| Literature Photometry | |
| [92] | 5865 |
| [93] | 5779 |
| [94] | 5879 |
Given the agreement between spectroscopic methods we observe for \(\beta\) CVn, it is of interest to see if this propagates to including literature values aside from the interferometry from [86]. Table 4 displays the stellar effective temperature determined as part of this work, in addition to literature values from interferometry, spectroscopy, and photometry. We see a general agreement between our results and those of literature spectroscopy and photometry that suggest the star is cooler than the interferometric methods of [86] and [88], however hotter than that from [87].
As part of GR8-1, photometric \(T_{\rm eff}\) and radii were obtained following the spectral energy distribution (SED) fitting methodology first applied by [95]–[97] – see GR8-1 for a detailed description of how this was used. Comparing the results from this work to these photometrically derived values allows for a comparison with \(T_{\rm eff}\) that is independent of our own in terms of both data and models.
A comparison of the \(T_{\rm eff}\) from each method in this work to the SED results is shown in Figure 7. The top row of Figure 7 shows the difference between the \(T_{\rm eff}\) from each spectroscopic method and that from SED fitting, plotted against the SED \(T_{\rm eff}\), and the bottom row shows the histograms of these differences. Observing the top row, we see that for PAWS, YARARA, and Metal Pipe there is a negative trend between the \(T_{\rm eff}\) difference and the SED \(T_{\rm eff}\), with PCCs of -0.49, -0.48, and -0.52, respectively. Restricting to the cooler-than-solar regime, as in Section 5.1.1, transforms these correlation coefficients to -0.03, 0.08, and -0.06, respectively. This replicates the observation that the negative correlation drastically reduces, or disappears entirely, when comparing to the PASTEL catalogue in the solar-and-cooler regime in Section 5.1.1, suggesting that further calibration for these methods is required for hotter stars. We again see a positive correlation in difference between \(T_{\rm eff}\) results when comparing ARES+MOOG in Figure 7, as was observed when comparing to PASTEL, however to a less significant effect with a PCC of 0.26. When we compare the difference between FASMA and SED \(T_{\rm eff}\) to the SED \(T_{\rm eff}\) we see a negative correlation, with a PCC of -0.18. Given the strong agreement between the FASMA \(T_{\rm eff}\), \(\rm [Fe/H]\), and \(\log\,g\) and those from PASTEL shown in Section 5.1.1, this may reflect an inherent systematic in the comparison of spectroscopic to SED fitting results. We additionally see more negative PCCs when comparing PAWS, YARARA, and Metal Pipe \(T_{\rm eff}\) to SED \(T_{\rm eff}\) than when comparing to PASTEL. Furthermore, although we again see a positive correlation when comparing the ARES+MOOG results in Figure 7, this is to a smaller magnitude at 0.26, rather than 0.39 when comparing to the PASTEL catalogue.
As expected from observing the top row of Figure 7, the histograms in the bottom row reveal the difference between spectroscopic and SED \(T_{\rm eff}\) have a negative peak for PAWS, YARARA, and Metal Pipe, indicating overall underestimation in \(T_{\rm eff}\) compared to SED results from these methods. Despite the trend observed in the top row, the histogram for ARES+MOOG shows a peak close to 0 K, indicating the majority of targets have an ARES+MOOG \(T_{\rm eff}\) that agrees well with the SED fitting result. The same is observed for the FASMA histogram, indicating that ARES+MOOG and FASMA are the methods that produce a \(T_{\rm eff}\) in the best agreement to SED fitting.
It is stressed in the literature that spectroscopic \(\log\,g\) often shows biases and offsets [82], [83], [98]. This is, however, dependent on the spectroscopic method, not a broad limitation of spectroscopy itself. With specialised analysis using extensive wavelength coverage and spectral line constraints, spectroscopic methods can recover \(\log\,g\) to within \(\approx\) 0.05 dex of asteroseismic values [72]. The disagreements in \(\log\,g\) determined by the methods in this work reflect the different choices in line lists, model atmospheres, and spectral content employed. Due to this, we selected the 30 targets for which spectroscopic \(\log\,g\) shows the largest disagreement between spectroscopic method results to perform comparisons within this section.
In addition to PAWS, ARES+MOOG, FASMA, YARARA, and Metal Pipe, we applied webSME to the 30 \(\log\,g\) discrepant stars. To save computation time we ran only the region 5000 - 6000 Å within webSME, with all targets beginning with solar input parameters. This spectral region includes the Mg I b triplet, known to have high sensitivity in the wings to \(\log\,g\) via pressure broadening [82], [99]. Given that webSME directly computes \(\log\,g\) from this region, we ran it on all 30 stars in this section to provide an additional, homogeneously derived spectroscopic comparison.
In Figure 8, for each method we have plotted the difference in \(\log\,g\) when compared to the webSME value against the webSME \(\log\,g\) for each of the 30 stars. It is clear from this that PAWS underestimates \(\log\,g\) significantly for these targets when compared to the webSME results, with the median difference from webSME \(\log\,g\) being -0.57 dex, far exceeding the median uncertainty of 0.13 dex. To a lesser degree, however still with significance, ARES+MOOG, FASMA, and Metal Pipe also underestimate \(\log\,g\) compared to webSME, with median differences of -0.34, -0.28, and -0.24 dex, respectively. The \(\log\,g\) values from YARARA differ from webSME results by a median value of -0.05 dex, indicating better agreement on average, however we observe a systematic trend in the relationship between the difference in \(\log\,g\) and the webSME value itself. Splitting this sample into the lower half of webSME \(\log\,g\)(\(\leq\) 4.32), we see a median difference between YARARA and webSME \(\log\,g\) of -0.02 dex. For the higher \(\log\,g\) half, this median difference increases to -0.18 dex.
To further assess the accuracy of the spectroscopic \(\log\,g\) values returned by each method, we used webSME’s forward modelling capabilities to compare synthesised spectra based on the results to the observed spectra for all 30 of the discrepant \(\log\,g\) stars. The forward modelling takes \(T_{\rm eff}\), \(\log\,g\), \(\rm [Fe/H]\), \(v_{\rm mic}\), \(v_{\rm mac}\), and \(v \sin i_{\star}\) as inputs for synthesising the spectra. Multiple options are available for the line list and solar composition employed in the synthesis – we opted for the Gaia-ESO (Y,Y|U) atomic linelist [52] and solar reference composition from [70]. We note that this line list was optimised using the solar composition from [71]. The small scale differences introduced by using a different solar composition here, however, do not impact the relative comparisons between methods that we perform. All forward modelling was performed within the region 5000 - 6000 Å, containing the Mg I b triplet, rather than the whole spectrum, to save computation time. Due to the sensitivity of these lines to \(\log\,g\), inaccurate \(\log\,g\) values would be strongly visible in the synthesised lines not representing the observed ones.
For results from methods in which not all broadening parameters (\(v_{\rm mic}\), \(v_{\rm mac}\), and \(v \sin i_{\star}\)) are determined, we used the inverse variance weighted mean from PAWS, FASMA, and webSME results. ARES+MOOG provides \(v_{\rm mic}\) for all targets, hence the mean was only used for \(v_{\rm mac}\) and \(v \sin i_{\star}\), whereas all three mean broadening parameters were required for both Metal Pipe and YARARA.
We used a reduced chisquared metric to directly compare how well the synthesised spectra match the observed one. The result of this is shown in Figure 10, with the median reduced chisquared from each code shown in Table 5. The nature of the reduced chisquared is that a value of 1.0 fits the data in accordance with its uncertainties, a value less than 1.0 is overfitting, and a value larger than one indicates a poorer fit. As seen in Table 5, webSME produces synthetic spectra fitting best to the observed ones around the Mg I b triplet, whereas PAWS performs poorly in this region, confirming the likely largest inaccuracy in the \(\log\,g\) determined by PAWS.
Given that we see median reduced chisquared values close to 1.0 in Table 5 for the EW based methods (ARES+MOOG and YARARA) we would also expect the EW section of PAWS to produce a well-fitting \(\log\,g\) value. Figure 9 displays a box-and-whisker plot of the \(\log\,g\) values from each code in this sample of 30 stars, along with the \(\log\,g\) produced by the EW section of PAWS. Although displaying a much larger range of \(\log\,g\) values, the results from PAWS EW align far more with those from webSME than the full PAWS results, indicating these may be a more reliable measurement.
| Code | Median reduced chisquared |
|---|---|
| PAWS | 1.805 |
| ARES+MOOG | 1.057 |
| FASMA | 1.408 |
| YARARA | 1.076 |
| Metal Pipe | 1.255 |
| webSME | 0.995 |
As isochrones determines mass and radius, we can also obtain \(\log\,g\) values from this method. As with Figure 1, we can again compare to the PASTEL catalogue, allowing systematic trends to be made apparent. This comparison is shown in Figure 11, with the results of inputting the results from each code to isochrones and the combined isochrone posteriors shown through histograms of their difference from the PASTEL \(\log\,g\). Using these isochrone \(\log\,g\) values, the trends of Figure 1 have been largely mitigated. Within the context of this sample and the spectroscopic methods tested, this indicates that the combined isochrone posteriors provide the most consistent set of \(\log\,g\) values. We do, however, note that radii determined through isochrones, and hence \(\log\,g\), are significantly model dependent and may not agree with empirical measurements [100].
We additionally investigate whether the use of the equivalent width or synthesis methods influences the observed underestimation of \(\log\,g\). We show results from this in Figure 12. The left hand panel shows the difference between the full PAWS spectroscopic \(\log\,g\) and the combined isochronal \(\log\,g\), plotted against the PAWS \(T_{\rm eff}\). The right hand panel repeats this, albeit comparing the \(T_{\rm eff}\) from only the EW of PAWS, rather than the full pipeline. Although large scatter is seen in both plots, there is a striking negative correlation between the difference in \(\log\,g\) from EW and the stellar effective temperature, replicating the results shown by [83]. Given this correlation, and the large scatter we see in spectroscopic \(\log\,g\) values throughout Section 5.3, we conclude that the combined isochrone \(\log\,g\) values provide the most self-consistent surface gravities from this sample. It is important, however, to remain aware that \(\log\,g\) values derived from isochrones are model dependent, and retain the influence of other stellar parameters [100], [101].
An important consideration in our analysis is the influence of spectroscopically determined \(v_{\rm mic}\) on other derived atmospheric parameters. For this investigation, comparison is limited to the codes that provide \(v_{\rm mic}\), being ARES+MOOG, FASMA, and PAWS.
Figure 13 displays the histograms of the resulting \(v_{\rm mic}\) determined by each respective code, in addition to the median and standard deviation of each sample. From this, we see all three methods clustering in a region of \(v_{\rm mic}\) \(\sim\) 1.0 - 1.2 km s\(^{-1}\). This is consistent with expectations for a sample of FGK dwarfs [39], [83], [84], indicating a general physical agreement between the methods.
There are, however, evident systematic differences between the methods. The median \(v_{\rm mic}\) values across this common stellar sample vary by up to \(\sim\) 0.17 km s\(^{-1}\), with FASMA producing the highest median \(v_{\rm mic}\) of 1.23 km s\(^{-1}\) and ARES+MOOG producing the lowest of 1.06 km s\(^{-1}\). These offsets are on a comparable scale to the internal scatter in the individual methods, highlighting that \(v_{\rm mic}\) is indeed sensitive to the choice of methods and models.
A difference in the dispersions of the resulting \(v_{\rm mic}\) distributions is also apparent, with PAWS showing the lowest standard deviation (\(\sigma\) = 0.20 km s\(^{-1}\)), and ARES+MOOG showing the largest (\(\sigma\) = 0.31 km s\(^{-1}\)). These differences in how tightly \(v_{\rm mic}\) is constrained could be attributed to a variety of reasons, potentially reflecting choices in line lists, models, methodology, or parameter degeneracies.
Overall, the differences in \(v_{\rm mic}\) distributions from each method reflect method-dependent systematics of \(\sim\) 0.1 - 0.3 km s\(^{-1}\), exceeding the median uncertainties of 0.06 km s\(^{-1}\), 0.03 km s\(^{-1}\), and 0.06 km s\(^{-1}\) for PAWS, FASMA, and ARES+MOOG respectively. This is further evidenced by the median RMSD in \(v_{\rm mic}\) across the three methods being 0.10 km s\(^{-1}\). Given the degeneracies between \(v_{\rm mic}\) and other spectroscopic parameters, these systematics may contribute to the differences observed between \(T_{\rm eff}\), \(\log\,g\), and \(\rm [Fe/H]\) in the methods, however are unlikely to be the sole cause.
In addition to the overall \(v_{\rm mic}\) distributions, we also observe a clear positive trend between \(v_{\rm mic}\) and \(T_{\rm eff}\) across all three methods, shown in Figure 14. This is consistent with expectations for FGK dwarfs, for which higher \(T_{\rm eff}\) is associated with higher velocity convective motions.
Although this positive trend is broadly followed for all three methods, Figure 14 shows that the exact relationship differs. This indicates that while the understood physical relationship between \(v_{\rm mic}\) and \(T_{\rm eff}\) is concretely recovered, its absolute calibration remains sensitive to the choice of method.
As previously described, we obtained stellar radii from SED fitting as part of GR8-1. In Figure 15, we show the difference between the radii from using each spectroscopic method plus isochrones, as described in Section 4.2, and the radii from SED fitting, plotted against the SED fitting radii. In all subplots, we see a slight upwards trend, with the difference between spectroscopic+isochrones and SED radii increasing with the SED radius.
An excellent agreement between the values is seen, with the RMSD between spectroscopy+isochrones and SED fitting being under 0.10 R\(_{\odot}\) for all sets of results. Figure 15 also shows the results from the combined posterior, shown as the ‘Isochronal’ results in the bottom right hand panel. From combining all spectroscopic+isochrones results, we observe an RMSD of 0.08 R\(_{\odot}\) from the SED radii.
In the context of exoplanet characterisation or demographic studies, it is important to understand the effect that the scatter in stellar parameters has on exoplanetary parameters. Here, we consider the planetary radius, mass, and equilibrium temperature.
In terms of the effect on exoplanetary radii, R\(_{p}\), we can use the transit depth \[\mathrm{Depth} = \left(\frac{R_{p}}{R_{\star}}\right)^{2}\] to investigate the effect of scatter in stellar radius. Considering only the uncertainty induced by the scatter in stellar radius, we see an uncertainty in planetary radius of \[\label{radius95unc} \sigma_{R_{p}} = \sigma_{R_{\star}} \sqrt{D}\tag{1}\] where \(\sigma_{R_{\star}}\) is uncertainty in stellar radius, and \(\sigma_{R_{p}}\) is the resultant uncertainty in planetary radius. We calculate the scatter in radius as the RMSD of the individual results from the median value. Across all targets, the median scatter in radius is 0.02 R\(_{\odot}\). Investigating the Earth-Sun system, we see an uncertainty in planetary radius of 0.02 R\(_{\oplus}\) (2.00 %) induced purely by the median scatter in stellar radius. As we extend to the cooler regime of K dwarfs, we use the case study of HD219134 [102] – a K dwarf with 4 orbiting planets. Looking at HD219134 b, we see an induced uncertainty in planetary radius of R\(_{\oplus}\) (2.57 %). We additionally test on CoRoT-11b [103], a planet orbiting at 0.0436 AU around an F dwarf with \(R\) = 1.37 \(R_{\odot}\) and \(T_{\rm eff}\)= 6440 K. In this system, we see an uncertainty in planetary radius purely from scatter in stellar radius of 0.23 R\(_{\oplus}\) (1.46 %). For HD219134 b and CoRoT-11 b, the quoted uncertainties in planetary radii are 0.09 R\(_{\oplus}\) and 0.34 R\(_{\oplus}\), respectively. From the NASA Exoplanet Archive [104], the median fractional uncertainty on exoplanetary radius in the literature is 7.36%, higher than uncertainty predicted to arise from scatter in stellar radius alone. It is important to note, however, that this prediction comes from scatter due to differing spectroscopic models only – the effects of varying evolutionary models using isochrones is not explored here.
To determine the significance of the effect of stellar mass scatter on the exoplanetary mass, we use the equation \[M_{p} \sin i \approx K\left(\frac{P}{2\pi G}\right)^{\frac{1}{3}}M_{\star}^{\frac{2}{3}}\] where M\(_{p}\) is planetary mass, P is planetary orbital period, G is the gravitational constant, and M\(_{\star}\) is the stellar mass. We assumed a circular orbit and a planetary mass that is negligible compared to the stellar mass. Isolating the contribution to uncertainty of scatter in stellar mass, we obtain the following equation to determine uncertainty in planetary mass \[\sigma_{Mp} \approx \frac{2}{3} \frac{\sigma_{M\star}}{M_{\star}}M_{p}\] where \(\sigma_{Mp}\) and \(\sigma_{M\star}\) are the uncertainties in planetary and stellar mass, respectively. Again taking the median scatter as RMSD from the median of the combined isochronal posteriors, we find this to be 0.05 M\(_{\odot}\) in stellar mass. For the Earth-Sun system, this scatter induces an uncertainty of 0.03 M\(_{\oplus}\) (3.00 %) in planetary mass. For HD219134 b, with a mass of 4.36 \(\pm\) 0.44 M\(_{\oplus}\), a scatter of 0.05 M\(_{\odot}\) would contribute an uncertainty of 0.19 M\(_{\oplus}\) (4.36 %). CoRoT-11 b, with a mass of 740.47 \(\pm\) 108.05 M\(_{\oplus}\), would have an uncertainty contributed from stellar mass scatter of 18.01 M\(_{\oplus}\) (2.43 %). Typical uncertainties in planetary mass appear to be significantly higher than these fractional uncertainties due to stellar mass scatter – see [105] (20.95 %), [106] (7.25 %), [107] (9.71 % & 40.68 %), [108] (8.79 %).
Taking all results from this work for the targets analysed, we calculate the scatter on each target as the RMSD from the median value. In \(T_{\rm eff}\), the median scatter across all targets is 76 K. Taking the situation of an Earth-Sun twin, with all orbital and physical parameters equal to that of the Earth and Sun, we used the following equation to calculate planetary equilibrium temperature \[T_{\textrm{eq}} = T_{\textrm{eff}} \sqrt{\frac{R}{2a}}(1-A_{B})^{\frac{1}{4}} \label{eqt}\tag{2}\] where \(R\) is stellar radius, \(a\) is orbital semi-major axis, and \(A_{B}\) is the Bond albedo of the planet. When propagating the uncertainty due to scatter in stellar \(T_{\rm eff}\) from using different methods, we consider that stellar radius is often not directly inputted into equation 2 , rather \(\frac{R_{\star}}{a}\) is generally found from transit information, and assume this to be the case for this purpose. Using the NASA Exoplanet Archive [104], we find the median fraction uncertainty in \(\frac{R_{\star}}{a}\) from transit measurements to be 0.16. We can use the following equation to translate the scatter in stellar \(T_{\rm eff}\) from the use of different methods to a fractional uncertainty in the planetary \(T_{\textrm{eq}}\); \[\frac{\sigma_{T_{\textrm{eq}}}}{T_{\textrm{eq}}} = \sqrt{(\frac{\sigma_{T_{\rm eff}}}{T_{\rm eff}}) + \frac{1}{4}R_{\rm frac}^2 } \label{teq95uncert}\tag{3}\] where \(R_{\rm frac}\) is the fractional uncertainty in \(\frac{R_{\star}}{a}\), which we set to 0.16 as per the median value in the literature. From equation 3 , all three case studies as previously investigated return a fractional uncertainty of 0.04 in the planetary equilibrium temperature. This suggests a noise floor of at least 4% fractional uncertainty should be present in all T\(_{\textrm{eq}}\) measurements. In recent literature, planetary equilibrium temperatures appear to often have uncertainties in the range 2.0% - 3.5% (e.g. [105] – 2.13%, [106] – 3.10%, [107] – 2.96% & 2.98%, [109] – 2.40%, [108] – 3.27%), suggesting that the noise floor introduced by scatter in \(T_{\rm eff}\) is not fully represented in the current state of the field.
From these investigations, it is strongly suggested that the scatter in spectroscopic stellar parameters from differing model choices is not a dominating factor in exoplanetary mass and radius uncertainties. However, as previously stated, this comparison does not reflect the scatter that can also be introduced by the use of differing evolutionary models in the determination of mass and radius. For planetary equilibrium temperature, the noise floor from the combination of scatter in \(T_{\rm eff}\) and median uncertainty in \(\frac{R_{\star}}{a}\) is higher than typical uncertainties in the literature, suggesting uncertainties in T\(_{\textrm{eq}}\) are generally underestimated. Additionally, as the precision and accuracy of exoplanet instrumentation and methodology increases, the noise floors investigated in this work will become increasingly prevalent and require careful attention.
By providing uniformly formatted, high resolution, high SNR spectra, the sec:gr8stars catalogue forms a uniquely useful aid in studying the systematic effects of different spectroscopic methods. This work shows the results of applying five
spectroscopic methods, differing in methodology, models, and assumptions, to the same data for a subset of 585 sec:gr8stars targets.
For the stellar effective temperature, surface gravity, and metallicity, we show that the precision errors often quoted for spectroscopic parameters are far outweighed by the scatter that comes from the use of different methods. Given that we consistently see disagreements between results at a significance larger than the precision uncertainties, we can see that spectroscopic atmospheric parameters are still influenced heavily by the disagreement between methodologies and models, thus require uncertainties reflective of this. From this work, suitable uncertainties for spectroscopically derived \(T_{\rm eff}\), \(\log\,g\), and \(\rm [Fe/H]\) require inflations of at least 76 K, 0.14 dex, and 0.07 dex respectively.
As part of this work, we also study the spectroscopic \(\log\,g\) using webSME for a direct comparison of 30 stars. As expected, we see large scatter in the spectroscopic \(\log\,g\) returned from each method. We demonstrate that the isochronal \(\log\,g\), which incorporates information from photometry and parallax, has improved agreement with literature values from the PASTEL catalogue. Due to this, we conclude that the isochronal \(\log\,g\) should be relied upon, rather than the spectroscopic one.
We also investigate the masses, radii, and ages of the stars, using isochronal fitting with the spectroscopic effective temperature and metallicity as part of the input. We propagate the scatter induced in these parameters from using \(T_{\rm eff}\) and \(\rm [Fe/H]\) from different spectroscopic methods to exoplanetary equilibrium temperature, radius, and mass. We find that the induced scatter does not outweigh the typical fractional uncertainties on exoplanetary mass and radius in the literature, whereas fractional uncertainties in planetary equilibrium temperature present a noise floor above those typically found in the literature.
While we found that the scatter in planetary radii and masses induced by different spectroscopic inputs is typically lower than uncertainties reported in the literature, it is important to recognise that this represents only part of the total uncertainty sources. In this work, stellar mass, radius, and age were determined using a single set of stellar evolution models. As a result, the noise floors that we estimate due to scatter do not account for systematic uncertainties resulting from model choice.
In addition, the demands of next-generation exoplanet surveys and studies require increasingly precise stellar properties. When we look particularly at the detection and characterisation of Earth-sized planets, stellar properties require uncertainties at a level for which the scatter explored in this work becomes significant.
While we focus on FGK dwarfs within this work, the expansion of this to cooler dwarfs would be particularly valuable. Spectroscopic modelling becomes increasingly difficult for stars with lower effective temperatures, and these are important targets for upcoming exoplanet surveys, including PLATO.
Future work should aim to extend this analysis with the incorporation of multiple stellar evolution models and grids. This will allow the full quantification of systematic uncertainties in derived stellar and planetary properties. As observational precision continues to improve, a thorough understanding of these systematics will be essential in reaching the full potential of next-generation exoplanet science.
We are grateful to the anonymous referee for their constructive feedback that improved the quality of this work.
AVF acknowledges the support of the IOP through the Bell Burnell Graduate Scholarship Fund.
A.M. acknowledges funding from a UKRI Future Leader Fellowship, grant number MR/X033244/1 and a UK Science and Technology Facilities Council (STFC) small grant ST/Y002334/1.
The Flatiron Institute is a division of the Simons Foundation.
This work benefited from support from the European Research Council (ERC) under the European Union’s Horizon 2020 Framework Programme (grant agreement no. 865624).
This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 101008324 (ChETEC-INFRA).
This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (CartographY GA. 804752)
JIGH and ASM acknowledge financial support from the Spanish Ministry of Science, Innovation and Universities (MICIU) project PID2023-149982NB-I00.
S.G.S. acknowledge support from FCT through FCT contract nr. CEECIND/00826/2018 and POPH/FSE (EC)
NCS is co-funded by the European Union (ERC, FIERCE, 101052347). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council. Neither the European Union nor the granting authority can be held responsible for them. This work was supported by FCT - Fundação para a Ciência e a Tecnologia through national funds by these grants: UIDB/04434/2020 DOI: 10.54499/UIDB/04434/2020, UIDP/04434/2020 DOI: 10.54499/UIDP/04434/2020.
This research has made use of the NASA Exoplanet Archive, which is operated by the California Institute of Technology, under contract with the National Aeronautics and Space Administration under the Exoplanet Exploration Program.
C.A.W. would like to acknowledge support from the UK Science and Technology Facilities Council (STFC, grant number ST/X00094X/1)
All stellar parameters derived through this work will be available through Vizier CDS. Spectra are available upon request to the authors.
Figures 16 and 17 display the differences between values determined spectroscopically in this work, and those from the PASTEL catalogue. Figure 16 compares \(\log\,g\) from this work to PASTEL, coloured by the \(\rm [Fe/H]\), whereas Figure 17 compares \(\rm [Fe/H]\) to PASTEL, coloured by \(T_{\rm eff}\).
E-mail: axf859@student.bham.ac.uk↩︎
NASA Sagan Fellow↩︎
ARES v2 can be downloaded at https://github.com/sousasag/ARES↩︎