August 03, 2025
While color harmony has long been studied in art and design, a clear consensus remains elusive, as most models are grounded in qualitative insights or limited datasets. In this work, we present a quantitative, data-driven study of color pairing preferences using controlled hue-based palettes in the HSL color space. Participants evaluated combinations of thirteen distinct hues, enabling us to construct a preference matrix and define a combinability index for each color. Our results reveal that preferences are highly hue dependent, thereby challenging classical color harmony theories. Yet, when averaged over hues, statistically meaningful patterns of aesthetic preference emerge, with certain hue separations perceived as more harmonious. Strikingly, these patterns align with hue distributions found in natural landscapes, pointing to a statistical correspondence between human color preferences and the structure of color in nature. Finally, we analyze our color-pairing score matrix through principal component analysis, which uncovers two complementary hue groups whose interplay underlies the global structure of color-pairing preferences. Together, these findings offer a quantitative framework for studying color harmony and its potential perceptual and ecological underpinnings.
Early efforts to understand the nature of color can be traced back to Antiquity with Aristotle’s chromatic theories [1]. This said, throughout much of history, color was primarily associated with symbolic and religious meanings. In Western civilizations, for example, saturated colors and polychromy were often associated with superficiality, vulgarity, or foreignness, while whiteness was imbued with connotations of purity and reassurance. A scientific turn occurred in the seventeenth century with the emergence of a more intellectual and systematic approach to color, exemplified by the creation of various color charts and tables, such as Richard Waller’s Catalogue of Simple and Mixt Colours [2]. A major scientific milestone was reached in 1704 with Newton’s famous treatise, Opticks [3].
Since then, the ambition to build consistent and rigorous color theories has continued to grow, drawing interest from a wide range of disciplines including mathematics, psychology, neuroscience, philosophy, and design, to name a few. A considerable body of research has focused on the construction of well-defined color spaces, whether through strictly mathematical frameworks [4]–[8] or more artistically driven approaches [9]. One particularly prominent area of color theory that has attracted sustained interest over the centuries is the study of color harmony [10]–[13]. Goethe [14], Chevreul [15], Ostwald [16], among others, first sought to capture the essence of harmonious color pairings by linking aesthetic appreciation to the perceptual distance between colors. More recently, data-driven approaches have gained traction. Some studies have focused on color emotion [17]–[21], aiming to identify correlations between specific hues and the emotions they trigger. Other works have investigated color preferences in terms of biological adaptations, highlighting marked gender differences [22], [23], or through the Ecological Valence theory [24], which links color preferences to associations with objects perceived positively or negatively. Additional studies have examined how preferences vary with age [25], [26], art education [27], or seasonality [28]. Building on these perspectives, researchers have also sought to develop generalizable models of color harmony using machine learning techniques [29], [30] and to build computational toolkits for extracting image features relevant to aesthetic judgments [31]. Complementing these scientific developments, a distinctly historical-cultural line of work has emphasized the enduring importance of colours across history and the ways in which the social, moral, and symbolic values attributed to them evolve over time, shaped by historical contexts and collective practices [32], [33].
In spite of these developments, most approaches to the color pairing problem still lack a truly quantitative foundation. Landmark perception-based studies such as those by Moon and Spencer [34] or Granger [35] have produced widely cited harmony principles, but were based on a very limited number of survey participants. A more recent and ambitious study was conducted by Nemcsics [36] in the Coloroid space among Technology and Economics students at Budapest University. However, the high level of complexity in the study—both in the geometrical patterns presented to participants and in the selection of colors combining hue, saturation, and lightness—may limit the interpretability of the results.
In summary, to this day, there remains no clear consensus on the principles underlying color harmony [37], thus leaving the topic open to debate and inquiry. Our approach here is inspired by the work of Lakhal et al. [38], on the structural complexity of black-and-white images, in which some of us provided evidence for a certain level of universal quantitative criteria for aesthetic judgment, see also [39]. Interestingly, the preferred level of complexity correlated strongly with that found in natural images, suggesting that what we find aesthetically pleasing may be shaped by what we are most frequently exposed to. Related work in the color domain has similarly connected aesthetic preferences to statistical regularities in natural images and paintings, pointing to a degree of universality in preferred chromatic compositions [40]–[43]. In line with this approach, we conduct a large-scale survey targeting a sufficiently diverse participant panel and grounded in simple evaluation procedures, aiming both to gain insights into preferences for color pairings and to compare these preferences with color combinations found in natural images.
The remainder of this paper is structured as follows. We first introduce our survey methodology, which is based on direct comparisons of simple color pairs. Next, we analyze the resulting data and propose a metric that quantifies how harmoniously each
color pairs with others. We then explore how this metric aligns both with absolute color preferences reported in independent studies and with hue distributions observed in natural images. By averaging across hues, we identify angular distances between
colors that are consistently preferred or rejected, and compare these perceptual trends to signals found in natural landscapes. Finally, we examine how different color groups contribute to the emergence of configurations of appreciation and
rejection.
We conduct a large-scale survey in which participants are asked to select their three most and three least preferred color combinations from sets of predefined color palettes, described in details in the following sections. Our study operates within the
HSL (Hue, Saturation, Lightness) color space [44], focusing exclusively on the hue component while keeping saturation and lightness fixed
at \(\text{L} = 0.5\) and \(\text{S} = 0.8\) 1. This choice also supports our large-scale, quantitative
objective: by focusing on hue—which is less affected by between-device variability than saturation and lightness—we were able to conduct the survey online and recruit a large, heterogeneous participant pool. Hues are represented as angular values on the
HSL color wheel [45], with red as reference color positioned at 0°, following the usual convention. We sample 18 equally spaced
hues at \(20^{\circ}\) intervals, starting from 0°, providing a reasonably fine coverage of the full hue wheel while keeping the set manageable for a preference task; then we remove five of them (corresponding to 80°, 100°,
140°, 260°, and 340°) to keep only clearly distinguishable colors (verified across different display devices). Indeed, the HSL color space lacks perceptual uniformity [46], as some colors comprised within large angular ranges (such as greens) cannot be unmistakably separated by the eye. In the following we refer to the 13 retained colors with indices \(i \in \{1,
\dots, 13\}\), with 0° red corresponding to \(i=1\). We then generate 13 sets of 12 color pairs by taking each color \(i\) in turn as the reference color, and pairing it with each of
the remaining 12 colors \(j\), hereafter referred to as the counterpart colors. The color pairs are presented as checkerboards consisting of 8\(\times\)8 square grids, with the reference and
counterpart colors distributed in equal amounts, though randomly in space to avoid recognizable patterns that might bias participants’ responses. An example of set with \(\text{H}=200\)° blue as reference color is shown in
the Supplemental information (Fig. 6). The survey is structured as follows: for each of the 13 sets, a question is created in which participants are shown, in a random order, one of the sets. Participants are then asked
to select the three most and three least harmonious color combinations. To limit survey duration and maintain engagement, each participant was asked to complete 6 randomly assigned questions (out of 13), rather than the full set. The subset of questions
varied across participants, ensuring a comparable number of responses for each of the 13 reference hues. The survey was conducted via the Qualtrics platform [47]. We collected 346 responses from colleagues at École Polytechnique and CFM (France), OIST (Japan), as well as volunteers across Europe who participated without financial incentives. The sample included both male and
female participants aged between 20 and 65, mostly with a high level of education but from diverse academic backgrounds.
For each reference color \(i\) and counterpart color \(j\), we define \(f^\text{B}_{ij}\) as the frequency with which the color combination \((i, j)\) was selected among participants as one of the three most harmonious combinations, normalized to the number of survey responses per question. Analogously, \(f^\text{W}_{ij}\) is the normalized frequency with which \((i, j)\) was selected as one of the three least harmonious combinations. A high value of \(f^\text{B}_{ij}\) suggests that color \(i\) pairs well with color \(j\), though a low value does not necessarily imply a negative judgment; it may simply reflect infrequent selection among the top three. Conversely, a high \(f^\text{W}_{ij}\) suggests unpleasantness while a low value does not imply positive judgment. We thus define the overall score of each color pair as: \[\label{eq:score} \text{S}_{ij} = f^\text{B}_{ij} - f^\text{W}_{ij} .\tag{1}\] The resulting score matrix \(\text{S}\) is shown in Fig. 1, where the \(i^{\text{th}}\) row corresponds to responses to the question in which color \(i\) was used as the reference. Note that the rows of the matrix naturally sum to zero (participant answers for each question were only retained if they provided all three best and three worst choices). Remarkably, \(\text{S}\) reveals an underlying structure, with discernible clusters of color combinations that are perceived as either harmonious or inharmonious. The nature of these clusters is examined in greater depth later in the paper.
Reassuringly, the score matrix appears to be quasi-symmetric (see Fig. 4a), indicating that the appreciation or dislike of a color combination does not strongly depend on the specific set in which it is presented to
survey participants. In other words, interchanging the reference and counterpart colors leaves the results largely unchanged to first order. That said, the matrix is not perfectly symmetric, and the residual asymmetry carries meaningful information. While
the row-wise normalization only allows to capture relative preferences for specific color pairs, the columns reflect judgments of absolute color combinability—that is, the ability of a given color \(j\) to harmonize with
all other colors across the set. For example columns 8, 9, or 10–corresponding to 200°, 220°, and 240° (shades of blue)–are composed of a majority of positive scores. This indicates that these hues tend to be selected more frequently as part of the best
than worst combinations. On the contrary columns of red and purple (columns 1, 12 or 13) carry mostly negative values, suggesting that these colors tend to produce less harmonious combinations when paired with others.
To quantify how harmoniously a color tends to combine with others, we define, for each color \(j\), the Combinability index as the sum of its pairing frequencies with all the other colors, across all questions: \[\label{eq:combin} \text{C}(j) = \sum_{ i\neq j} \text{S}_{ij} .\tag{2}\] The combinability indices for the 13 colors of our study are plotted in Fig. 2. Following the shades of blue (200°, 220°, and 240°) discussed above, yellow (60°), closely followed by orange (40°), exhibit the highest Combinability indices. Conversely, red (0°), green (120°) and purple (300°) tend to produce combinations that are more frequently perceived as inharmonious. The top and bottom gray lines delimiting the shaded area represent, respectively, the sum \(\text{C}^\text{B}(j)\) and \(\text{C}^\text{W}(j)\) of the best-only frequencies \(f^\text{B}_{ij}\) and the worst-only frequencies \(f^\text{W}_{ij}\) for each color \(j\). Interestingly, although certain colors like 180° cyan (\(j=7\)) have a Combinability index close to zero, they show substantial values of \(\text{C}^\text{B}\) and \(\text{C}^\text{W}\), indicating that such hues are frequently selected by participants—appearing as often in the most appreciated combinations as in the least.
Building on the intuition that combinability might be linked to absolute color preference, we compare our results with findings from the Berkeley Color Project [48], which investigated human preferences for eight highly saturated individual colors (triangular markers and dashed line in Fig. 2). Remarkably, individual hue preferences closely mirror the Combinability index, suggesting that a color’s ability to form harmonious combinations aligns with its overall aesthetic appeal.
Finally, leveraging previous studies that have highlighted a strong correlation between human aesthetic preferences and structures found in nature [38], [49], [50], we analyzed a dataset of 12,000 landscape images [51] spanning various natural biomes, including coasts, deserts, forests, glaciers, and mountains. The count of hue occurrences in the entire database is
shown as a background histogram in Fig. 2. Strikingly, the peaks and valleys of the distribution match remarkably well that of the Combinability index and absolute preferences (maximum for blue and orange-yellow tints,
minimum for greens and purples), suggesting that our aesthetic preferences may be influenced by the colors we have most frequently been exposed to. To test the robustness of our findings and avoid potential biases in image selection, we also measured the
hue distribution using two additional non-overlapping datasets of 4,319 and 15,501 natural images respectively [52], [53]2; the results were found to be virtually identical.
Classical color harmony theories, such as those of Moon and Spencer [34] or Itten [11], tend to rely on a strong assumption of hue independence, meaning that preferences for a given hue pair depend only on the angular distance separating the two colors on the hue wheel, regardless of their absolute positions. To examine the relative positions of the combined colors \((i,j)\), we rotate the hue wheel in each set such that the reference color is set at \(0^\circ\) (see Fig. 8), and assign to each comparison color a value function of its angular distance from the reference. Our findings challenge this hue-independence assumption to a large extent, as discussed further below. Nevertheless, in order to confront our findings with existing theories, we compute, for each angular distance, the average selection frequency across all reference colors. The results are shown in Fig. 3, where green bars indicate preference and red bars signify dislike, with \(180^\circ\) corresponding to the maximum angular distance between two hues. Higher preferences are observed in the contrast region (between \(160^\circ\) and \(220^\circ\)), flanked by regions of lower preference at smaller angular distances. Such global preference for contrast aligns rather well with the universal harmony model proposed by Moon and Spencer[34], which remains influential today 3. However, the other preference regions they proposed—coined similarity and located near the reference color—are not consistently supported by our hue-averaged data. The universality of such theories is further challenged when examining the standard deviation at each angular distance, shown as error bars in Fig. 7 (see Supplemental information). The substantial variability observed reflects pronounced differences across individual hue preference wheels (see Fig. 8). Consequently, the assumption of hue independence cannot be reasonably upheld. For instance, the preference wheels for green (e), cyan (g), and light purple (l) in Fig. 8 display markedly distinct patterns 4.
To compare our results with hue combination occurrences in natural images, we identify the dominant hue of each landscape in our dataset using a Gaussian kernel density estimation [54]. For each image, we then count the number of pixels falling at each angular distance from the dominant hue. Summing these counts over the entire dataset yields the gray histogram shown in Fig. 3. Note that angular regions between 0°–20° and 340°–360° were excluded from the analysis, as they are artificially overrepresented: since the dominant hue is always aligned to 0°, nearby shades of the same tint are, by construction, more likely to appear. At sufficiently large angular distances from the reference color, dominant hues in natural scenes are most often separated by approximately 180°, suggesting that natural environments are largely characterized by strong color contrasts. The overlay of the empirical preferences and the natural color distance histogram reveals a compelling alignment: the most appreciated color pairs in our survey tend to coincide with those that occur more frequently in nature, while the least appreciated combinations correspond to rarer configurations in natural landscapes. This suggests that our aesthetic preferences for color pairings may be shaped—at least in part—by repeated exposure to common visual patterns in our environment. This finding strongly echoes earlier results by some of us [38] on black-and-white image complexity, where aesthetic appeal was likewise found to correlate with structural patterns prevalent in nature. Note that the alignment between preferences and natural color pair occurrences can also be analysed on a hue-by-hue basis (see Fig. 8).
To better understand the structure of color pairing preferences—and to assess whether certain colors or groups of colors contribute more significantly to the overall signal—we analyze the symmetric part of the score matrix, defined as \(\text{S}^s = \frac{1}{2}(\text{S} + \text{S}^\top)\), and study its principal components. As previously noted, the original score matrix \(\text{S}\) exhibits a near-symmetric structure, which
is quantitatively supported by the proximity of its eigenvalues to the real axis, see Fig. 4a. This observation justifies our focus on the symmetric component \(S^s\), the eigenvalues of
which are plotted in Fig. 4b. In Fig. 4c, we display the outer product matrices \(\lambda_i^{\text{s}} \boldsymbol{\text{v}_i}
\boldsymbol{\text{v}_i}^\top\) corresponding to the six eigenvalues \(\lambda_i^{\text{s}}\) with the highest absolute values (see Fig. 4b), where \(\boldsymbol{\text{v}_i}\) denotes the eigenvector associated with the \(i^{\text{th}}\) eigenvalue. The first matrix, associated with the largest eigenvalue, reveals a clear emergence of
positive and negative clusters, indicating a division of the hue wheel into two main groups: one spanning from orange to cyan (group 1), and the other from blue to dark orange (group 2). Colors from group 1 tend to combine harmoniously with those from
group 2, but not with others within their own group, and vice versa. This structured clustering highlights which segments of the hue spectrum predominantly drive the universal appreciation seen in the contrast region of Fig. 3, as well as the adjacent zones of disfavor.
Interestingly enough, when our color groups are represented in the standard CIE 1931 xy chromaticity diagram [55] (see Fig. 5), they are found to be linearly separable. One can also clearly see that while group 1 includes colors from the central region of the visible spectrum, group 2 is composed of colors from two extremes of such spectrum, together
with the purples (which correspond to a mix of blue and red). Finally note that the decision boundaries shown in Fig. 5 pass close to white, at the center of the gamut.
Further structure is visible in the subsequent components. In particular, the second matrix exhibits a more localized negative signal, concentrated among hues from light purple to dark orange (bottom right red cluster), indicating internal
incompatibilities within this range. Additionally, the fifth matrix reveals another pronounced negative contribution, pointing to poor combinability between green and cyan.
Let us summarize the main contributions of this work. We designed and conducted a large-scale survey in which participants were asked to select the most and least appealing color combinations from carefully curated sets. Our first finding was that
certain hues—such as blue and yellow—consistently form more harmonious pairings with other colors. We then introduced a combinability index to quantify this behavior, and showed that it correlates well with absolute color preferences reported in an
independent study [48]. More notably, this index also aligns closely with the distribution of hue occurrences in natural landscapes. Shifting
focus from individual hues to angular distances between paired colors, we uncovered a robust preference for combinations in the contrast region, centered around complementary hues. This pattern mirrors color distributions commonly found in nature,
suggesting that aesthetic appeal may be shaped, at least in part, by frequent exposure to naturally occurring visual stimuli (see also [38]). That
said, our results do not fully align with the universal view of classical color harmony theories [11]–[16], which assumes that harmony depends mainly on fixed color distances (e.g. simple angular relations) rather than on the absolute colors involved. While contrast is
favored on average, preferences vary strongly with the specific hues paired, so fixed-distance rules do not consistently predict human preferences. Finally, a principal component analysis of the score matrix provided further insight, revealing clusters of
hues that tend to combine either particularly well or poorly with others, thus offering a more nuanced understanding of the structure underlying human color pairing preferences.
To simplify the interpretation of color combinability and offer an accessible entry point into the question of color harmony, we deliberately limited our analysis to variations in hue only. This choice is also well suited to a large-scale online survey,
as hue is comparatively more robust than saturation and lightness to differences across display devices. While this approach offers clarity and tractability, it overlooks the potentially rich interactions with saturation and lightness. A natural direction
for future work would be to extend the study to full-color combinations across all HSL dimensions, and to examine how such variations relate to patterns found in natural imagery. Another important avenue lies in broadening the participant panel—not only in
size but also in cultural and demographic diversity—to test the generalizability of our findings across different populations. Finally, our use of the HSL color space, though convenient for isolating hue, deserves further discussion. HSL suffers from
perceptual non-uniformity and device dependence, limiting its fidelity in capturing human color perception. Future research could explore alternative, perceptually uniform color spaces such as CIELAB (CIE \(\mathrm{L}^{*}\mathrm{a}^{*}\mathrm{b}^{*}\) defined by the Commission Internationale de l’Éclairage) and OKLAB [59]. However, this comes with its own challenges: in such models, hue is no longer a single scalar dimension but arises from interactions between multiple components, complicating direct comparisons.
We are deeply grateful to all the participants of our survey, whose contribution was foundational to this work. We also thank Jean-Philippe Bouchaud, Pierre Bousseyroux, Samy Lakhal, Elia Moretti and Mirko Polato for fruitful discussions. This research was conducted within the Econophysics & Complex Systems Research Chair, under the aegis of the Fondation du Risque, the Fondation de l’Ecole polytechnique, the Ecole polytechnique and Capital Fund Management.
Note that the HSL color space was chosen because it allows hues to be easily isolated, which would not be possible in other commonly used spaces such as RGB, which requires the use of all three channels. Interestingly enough, the choice of color space does not appear to affect the results on color preferences, see e.g. [22]. Most importantly, hue preferences have been found to remain largely unaffected by variations in lightness and saturation [35].↩︎
These latter correspond to the validation split of the Places database [53], restricted to 31 categories associated with natural scenes.↩︎
Although Moon and Spencer used the Munsell color system [9], while we rely on HSL, the hue wheels are broadly comparable, and our conclusions do hold to first order.↩︎
Also note that, contrary to Moon and Spencer’s theory [34], the preference wheels are not symmetric with respect to angular distance–preferences at \(\Delta\theta\) are different from those at \(-\Delta\theta\). See, for instance, cyan (g) in Fig. 8 for a striking illustration.↩︎