The Noisy Work of Uncertainty Visualisation

Harriet Mason
Department of Econometrics and Business Statistics, Monash University
and
Dianne Cook
Department of Econometrics and Business Statistics, Monash University
and
Sarah Goodwin
Department of Human Centred Computing, Monash University
and
Emi Tanaka
Biological Data Science Institute, The Australian National University
and
Susan VanderPlas


Abstract

Better representation of the uncertainty in a data visualisation is a focus of recent research activity. A problem with the current literature is that there is a lack of clarity about the definition of uncertainty and what it means to represent it in a plot. This confusion results in a significant amount of conflicting results in the literature, especially in experiments that assess the effectiveness of different uncertainty representations. In this review, we summarise the current literature, provide workable definitions, and illustrate these definitions with examples. In doing so, we ask what it really takes to achieve transparency in statistical graphics. It is hoped that it will be useful for guiding new graphics methodology and experimental research.

data visualisation, statistical graphics, data science, information visualisation

1 Introduction↩︎

What do we mean when we talk about “uncertainty visualisation”? The phrase can feel contradictory to anyone familiar with the term. Among statisticians, “uncertainty” is often discussed as an omnipresent spectre touching every stage of our analysis without ever being fully seen. Authors will often mention that the phrase is vague [1], [2], or avoid defining it by describing a list of things uncertainty could be [3], [4], but rarely do authors attempt to discuss what uncertainty actually is. By contrast, visual statistics (information visualisations, data plots) are one of the most powerful tools in the statistician’s toolbox, allowing for quick and memorable communication that identifies quirks in our data that we didn’t even know to look for. We see this in datasets such as Anscombe’s quartet [5] or the Datasaurus Dozen [6], [7], where visual statistics are able to highlight elements of the data that are invisible to the typical summary statistics. We also see this in recall experiments, where simply sketching a distribution before recalling statistics or making predictions can greatly increase the accuracy of those measures [8], [9]. Taken together, uncertainty visualisation implies a need to pull back the curtain and explore the unknowns of our analysis.

As nice as this sentiment is, it turns out to be easier said than done. Reviews on uncertainty visualisation rarely offer tried and tested rules for effective uncertainty visualisation, instead commenting on the difficulties faced when trying to summarise the field. [3] found most experimental methods to be ad hoc, with no commonly agreed upon methodology, formalisations, or a greater goal of describing general principles. [4] noticed there is a serious noise issue in the field, with noise from participants misunderstanding visualisations, misinterpreting questions, and incorrectly applying heuristics, overwhelming any information we can glean from studies. [10] identified so much contradicting evidence that they spent an entire page discussing the conflicting evidence for the question “Should I map uncertainty to colour hue?” [1] concluded that different plots are good for different things, arguing against a universal best plot for all people and circumstances. [11] summarised several cognitive effects that repeatedly arise in uncertainty visualisation experiments; these effects were each discussed in isolation as a list of considerations rather than an overarching theory for effective uncertainty visualisation.

“Science is built up of facts, as a house is built of stones; but an accumulation of facts is no more a science than a heap of stones is a house.” - Henri Poincaré (1905)

While these reviews are thorough in scope, none discuss how the existing literature contributes to the broader goal of uncertainty visualisation – that is, despite the wealth of reviews, the field of uncertainty visualisation remains a heap of stones. There is a mountain of work that identifies common heuristics found in uncertainty visualisations, evaluates competing plot designs, or starts a theoretical discussion on a niche aspect of the field. While important, each of these papers offers up its own bespoke motivation and methodology, with little reference to the uncertainty visualisation papers outside their fiefdom. The field is in desperate need of a unifying theory that can tie the conflicting and siloed research together. This review attempts to address this issue by offering a novel perspective on the uncertainty visualisation problem. That is, we will use the wealth of established stone to construct a foundation to build a house.

2 The purpose of uncertainty visualisation↩︎

Mentions of “uncertainty visualisation” start springing up around 1990, across several different fields [12], [13], each with its own motivation for the work. In computer science, the area appears to be motivated by issues in the public’s perception of random variables, with the hope that visualisations would give laypeople the ability to extract important information from graphical representations [12]. With similar concerns about the public’s understanding of randomness, the fields of psychology, statistics, and economics used “uncertainty visualisations” as a communication tool to mitigate the psychological bias associated with the communication of risk, a topic of concern since the early 1980s [1]. In cartography, it was motivated by the inherent uncertainty of geoscience data, the practical use of visualisation as an exploratory tool, and the constrained visual channels from map representations [13]. These disparate motivations have blended together, and today, uncertainty visualisation is usually motivated by the vague goal of “decision-making”. This term has been used to mean the mitigation of psychological bias to ensure economically rational decisions [14][16], the facilitation of trust or confidence [17], [18], the ability to extract values related to a distribution [19], the prevention of false discovery in plots [20], [21], or the extraction of some other metric that is vaguely related to “uncertainty” [22], [23]. This gradual scope creep of the field, motivated by the hazy definition of “decision-making”, is the most likely culprit for the jumbled literature that makes up the field today.

Given that there is so much subliminal disagreement in uncertainty visualisation, how did these disparate motivations come to be seen as interchangeable? All discussions on uncertainty visualisation seem to have a common thread that connects them: the belief that the ultimate goal of uncertainty visualisation is not trust, rationality, or value extraction, but transparency. We see it said directly in reviews of the field [11], or when authors claim that failing to include uncertainty is akin to fraud or lying [24], [25]. We see it when authors assert that uncertainty communicates the legitimacy (or illegitimacy) of the conclusion drawn from visual inference [2], [26]. We see it when authors say uncertainty visualisations should communicate a degree of confidence [27], [28] or validity [2], [24] in our conclusions. We see it when authors suggest uncertainty visualisation should “guide, qualify, or soften our judgements of uncertain data” (e.g. [29] in his seminal work on the grammar of graphics). These authors are not wrong about the need for transparency in science communication: a six-month survey of anti-mask groups on Facebook during the COVID-19 pandemic showed that anti-maskers made persuasive arguments by exploiting inherent uncertainty ignored by pro-maskers [30].

Uncertainty visualisation is motivated by the need for a sort of “visual hypothesis test”, a sentiment expressed by some authors directly [13], [31]. A successful uncertainty visualisation would act as a “statistical hedge” for any inference we make using the graphic. Since the purpose of a visualisation is to give a quick gist of the information [1], this hedging should be communicated visually without the need for complicated mental calculations. Therefore, an effective uncertainty visualisation should not just “show” uncertainty; untrustworthy conclusions should not be visible. If we refer to the conclusion we draw from a graphic as its signal, where uncertainty should make this signal harder to read as the “noise” increases, we can summarise the above information into three key requirements. A good uncertainty visualisation should:

  1. Reinforce justified signals to encourage confidence in results.

  2. Hide spurious signals that are overwhelmed by noise.

  3. Perform tasks 1) and 2) in a way that is proportional to the level of confidence in those conclusions.

Usually, visualisations that are unconcerned with uncertainty have no issue showing justified signals, but struggle with the display of unjustified signals. Therefore, we suggest calling this approach to uncertainty visualisation “signal-suppression” since it primarily differentiates itself from the normal “noiseless” visualisation approach through criterion (2). This is the main criterion we will use to assess the current literature on uncertainty visualisation.

3 Current Approaches↩︎

3.1 Ignoring uncertainty↩︎

The most common way to visualise uncertainty is to simply not. A study conducted by [24] found that only a quarter of authors surveyed included uncertainty in 50% or more of their visualisations, in part because authors are not sure how to calculate uncertainty. This is not entirely unreasonable, given that even visualisation authors themselves seem to be in conflict about what exactly uncertainty is. We will start with visualisations that ignore uncertainty with the hope that by looking at where uncertainty isn’t, we can better understand where it is.

3.1.1 What is uncertainty?↩︎

It is surprisingly hard to describe what uncertainty is. Most authors avoid the problem and describe the many characteristics of uncertainty. Often, uncertainty is split by factors such as whether it is due to true randomness or a lack of knowledge [1], [4], [11], [32][34]; quantifiable or unquantifiable [1], [11], [34]; scientific or human [33], [35]; systematic or random [36]; statistical or bounded [37], [38]; accuracy or precision [2], [4], [35]; etc. There are enough qualitative descriptors of uncertainty to fill a paper, but none of this is particularly helpful in understanding how to integrate it into a visualisation.

Rather than trying to define uncertainty by looking at the myriad ways in which it does appear in an analysis, we may find it easier to look at where it does not. Descriptive statistics describe our sample as it is and summarise large data into a usable format, but they are not seen as the primary goal of modern statistics. In 19th-century England, positivism was the popular philosophical approach to science (positivists included famous statisticians such as Francis Galton and Karl Pearson). Practitioners of the approach believed statistics ended with descriptive statistics, as science must be based on actual experience and observations [39]. In order to make statements about population statistics, future values, or new observations, we need to perform inference, which requires the assumption of the “uniformity of nature”, that is, we need to assume that unobserved phenomena should be similar to observed phenomena [39]. Positivists believed referencing the unobservable was bad science, embracing descriptive statistics due to the inherent certainty associated with them. Since uncertainty is nonexistent in descriptive statistics, it is clear that uncertainty is a by-product of inference: uncertainty is the noise that is both inseparable from our inference and meaningless without it.

Rather than extracting just one element of the distribution, if you can retain the whole distribution, that not only allows the uncertainty calculation to be reproduced, but also makes it possible to derive other estimates as well. I think it’s also important to acknowledge that if you use this approach, you need to know the “chicken” (i.e. distribution) that gave birth to the “egg” (uncertainty estimate). In practice, uncertainties are sometimes calculated by other people or organisations, and the process used to derive them may not be known.

If we consider uncertainty to be a by-product of statistical inference, then uncertainty visualisations are the plots that depict an estimate, and therefore have an associated uncertainty. The most complete description of these estimates is their distributions. Rather than extracting just one element of the distribution, if you can retain the whole distribution, that not only allows the uncertainty calculation to be reproduced, but also makes it possible to derive other estimates as well. Suggesting distributions as a representation of uncertainty is not new. [40] originally suggested thinking about uncertainty visualisations as visualisations with distribution inputs, to replace the commonly used mean and standard deviation, removing the assumption of a Gaussian distribution. In practice, uncertainties are sometimes calculated by other people or organisations, and the process used to derive them may not be known. While some researchers believe these abstract notions of uncertainty, such as credibility [41], forecaster confidence [15], or uncertainty about uncertainty [42], are too complex to be quantified, this is not necessarily true. Abstract concepts such as human belief or credibility are regularly quantified by Bayesians, and hierarchical approaches are often used to model uncertainty about uncertainty.

3.1.2 Example: ignoring uncertainty↩︎

If visualising uncertainty is fundamentally visualising a set of random variables, what does “ignoring” uncertainty look like? Figure 1 shows plots of data from three different scenarios, each with either high or low uncertainty. The plots show the expected value of the input distribution. In plots R1 and R2, the expected value is the line from a simple linear regression on a car’s miles per gallon (mpg) and weight (wt) using the mtcars data, or a subset of it (available in ggplot2 [43], and originally from [44]), with differing sample sizes creating differing levels of uncertainty. Plots S1 and S2 show a choropleth map of Iowa, where counties are coloured according to a simulated temperature measurement, with measurement error being the uncertainty associated with the measuring instrument. Plots G1 and G2 show samples simulated from five populations (A-E). In G1, distributions have the same low variance, and in G2, they have the same high variance. In the linear regression, we see a downward trend for both levels of uncertainty. Here, we elected to plot the data under the fit, to provide context for this example, and help follow the thinking through to the later illustrations. Often, only the fitted line is shown. In the map, you should see a sine wave spatial trend, and in the univariate distributions, an incremental increase in treatment. If we were to ask a reader, “Can you see a difference between the plots in the top row versus the bottom row? Is the strength of the trend communicated through the visualisation?” The answer to both of these questions will be no for S1, S2, G1, and G2, as the high and low variance cases are identical. For R1 and R2, the answer would likely only relate to the points in the plots, not the regression line.

a

Figure 1: Three example types of data and associatedplots that will be used to illustrate various choices of uncertaintyrepresentation throughout the paper: scatterplot and regression line(R1, R2), spatial choropleth map (S1, S2), grouped dotplot (G1, G2). Twolevels of uncertainty (low, high) are used with each example. Ignoringthe points in R1 and R2, there are no differences between the high andlow uncertainty versions. Ignoring uncertainty can lead tomisrepresentation of data..

These examples will serve as touchstones for a discussion of uncertainty visualisation, focusing on the approaches suggested in the literature.

3.2 Uncertainty as a statistic↩︎

Uncertainty is also often treated as another statistic, for example, the exceedance probability map of spatial data (see [45] with software available in [46]), where the probability of exceeding a specified value is displayed on the map, placing the focus on extreme events. For risk communication, [1] recommends showing probability with a set of coloured icons to communicate the relative number of affected individuals. The summary plot [47] was also developed to address concerns about uncertainty in data displays. This dense display combines the plot of a single set of numeric values with a density estimate, boxplot, and statistical moments.

Some authors take this approach because they explicitly believe uncertainty is a variable of importance [48], while others straddle the line, asserting uncertainty is acting as signal and noise, and should fulfil both roles [49]. Simply put, these approaches might be described as swapping out the statistic for an “uncertainty statistic” to get an “uncertainty visualisation”. Is it really that easy?

3.2.1 Example: visualising variance↩︎

Figure 2 depicts the six plots showing this approach for the data introduced in Figure 1. The original central value estimate has been replaced with an uncertainty statistic. Plots R1 and R2 show the residual plot of our linear regression instead of the scatterplot and the regression line. The visual patterns we are looking for in this plot are distinct from the linear regression, so it is hard evaluate it relative to the trend. Sometimes the variance approach is still related to our original display, as we can see in the exceedance probability maps depicted in plots S1 and S2. These map \(P(temperature>27)\) to colour for each county. The sine wave trend is clearly visible when the error is low, but barely visible with high error. Our visualisation of the univariate groups does reveal something interesting: the variance is constant in G1 but different for each group in G2. It is a nonsensical display, though, because the trend has completely disappeared. In practice, two plots are typically presented: one showing the main estimate and the other showing the uncertainty.

a

Figure 2: Treating the uncertainty as a statistic,with the same six examples. The regression (R1 and R2) is a scatterplotof residuals vs explanatory variable, separated from the regressionmaking comparison of uncertainty related to the trend more difficult.For the choropleth (S1, S2): instead of temperature, the probability ofexceeding 27\(^o\)C is shown. This has the effect of highlighting (sinewave) trend in the low error data, and de-emphasising it in the higherror data. The plots G1 and G2 have replaced the treatment effect withthe variance on our treatment effect. It is a bit nonsensical, but welearn something interesting that was not seen earlier: the variance inG2 is not uniform like that in G1..

3.2.2 What is an uncertainty visualisation, then?↩︎

What an uncertainty visualisation is or is not is one of the most pervasive divides in the literature. For example, [29] mentions that popular graphics, such as pie charts and bar charts, omit uncertainty. [50] suggests their product plot framework, for area plots like bar charts and also histograms, needs to be extended to include uncertainty representation. However, pie charts, bar charts and histograms have all been used in a significant number of experiments as examples of an “uncertainty visualisation” [12], [17], [38], [51]. What is going on here?

This conflict stems from a subconscious disagreement about the purpose of uncertainty visualisation. If you believe uncertainty visualisation is about communicating risks or random variables, uncertainty visualisations are just visualisations of “uncertainty statistics” or distributions. On the other hand, if you believe uncertainty visualisation is about suppressing false signals visually, then you see an uncertainty visualisation as a transformation of an existing graphic that adds the uncertainty in. The former has no limitation on the visual appearance of an “uncertainty visualisation”, allowing pie charts, bar charts or histograms, so long as the graphic is visualising “uncertainty”, while the latter believes uncertainty visualisations only exist in relation to some “normal” visualisation. When we refer to the graphics depicted in Figure 2 as “uncertainty visualisations”, we are classifying visualisations by the data they display, not their visual features. This is not the standard approach in statistical graphics. A scatter plot that compares means and a scatter plot that compares variances are both scatter plots.

Unlike plots, which are not defined by a statistic, uncertainty is only defined in relation to our visual statistic, a topic that frequently appears in the literature to be dependent on the “goals” of our analysis. [52] commented that what is kept as data and what is tossed away is determined by the motivation of an analysis - what was previously noise can become signal depending on the question. [39] suggested that the process of observing data to calculate statistics is largely dependent on our goals, because the process of boiling real-world entities down into probabilities depends on the relationships we seek to identify within our data. [53] argue that the best method for evaluating or combining subjective probabilities depends on the uncertainty the decision-maker wants to represent, and why it matters. [54] suggested we should have methods for communicating uncertainty depending on what the user is supposed to do with it. [1] says we “cannot assess the quality of risk communication unless the objectives are clear”. [49] asserted that whether or not uncertainty is a source of doubt depends on the context. The sentiment behind this repeated point is clear: the role of uncertainty or signal is not dependent on the “type” of statistic, on the source of the information, or the methods we use; it is determined by the statistic we wish to draw inference on. Therefore, the fundamental problem with the “uncertainty statistic” approach is that the uncertainty in the plot isn’t acting as noise; it is acting as signal.

If the uncertainty in a graphic is acting as a signal, there isn’t an interesting perceptual challenge associated with the visualisation: the uncertainty can be displayed using standard principles of graphic design. In changing the inferential statistic, we also haven’t dealt with the original problem of integrating noise, as these “uncertainty statistics” also have associated uncertainty in the estimates (e.g. variance of standard deviation estimate) that is being ignored. There is nothing wrong with explicitly visualising variance, error, bias, or any other statistic. These metrics provide important and useful information for analysis and decisions. The problem with this approach is that it means everything is an uncertainty visualisation, and if everything is an uncertainty visualisation, nothing is.

3.3 Uncertainty as a variable↩︎

Another common characterisation of uncertainty is as just another variable to be integrated into the visualisation, which means uncertainty visualisation is, at its core, a high-dimensional visualisation problem [2], [49]. This occurs within computer science [3], cartography [10], and statistical graphics [29].

Discussion of the visualisations focuses on how “integrated” the uncertainty is with the estimate. [3] identified a split between intrinsic plots, where we map uncertainty to the colour or size of the geometric object of our estimate, and extrinsic plots, where uncertainty is mapped to a separate geometric object, such as glyphs or error bars. Similarly, [11] classified uncertainty visualisations as graphical annotations (extrinsic), and probability mapped to a visual encoding channel (intrinsic), or a hybrid of the two. It is unclear if these levels of “integration” in a plot design affect its ability to suppress signals.

3.3.1 Example: mapping two independent variables↩︎

Figure 3 shows the six examples introduced in Figure 1. Here, uncertainty is mapped to a spare aesthetic in the plot. In the grammar of graphics, variables are mapped to aesthetics, like position, colour, and size, within a plot. In plots R1 and R2, the standard error of the slope is represented by the width of the line. This is common, but it fails to represent the standard error of the intercept alongside the slope. We might think we can examine the width of the line at wt=0, but this would be incorrect. Here, it has been included, less obviously, by mapping the standard error of the intercept to transparency. While these help to give a gestalt of the uncertainty, they are not exacting representations. One would expect that the width above and below the line is one standard error, but why not map the width to two standard deviations, or even three? Interpreting transparency into a numerical quantity is also virtually impossible, so using this aesthetic mapping is effectively useless.

A bivariate colour palette map is shown in plots S1 and S2. With two dimensions of the plot reserved for spatial position, it is difficult to incorporate the error. The bivariate colour palette maps the estimate to hue and the error to saturation, keeping the signal and noise contained to one visual aesthetic. This has the unfortunate effect of making the signal appear stronger when the error is higher (S2). Colour perception is a wild beast that is hard to tame.

The extrinsic approach, shown in plots G1 and G2, has the estimate represented by a point, and uncertainty computed as a 95% confidence interval mapped to line length. The trend is still visible in both displays. It could be argued that where the variance is high, the trend remains a main focus; that is, the display fails to sufficiently suppress the signal.

Because all of these graphics visualise the distribution’s estimate and uncertainty as two separate pieces of information, the message of the plot is “here is the trend and here is the uncertainty”. It is worthwhile to examine why this occurs, to see if we can move towards a version of this plot where we are able to communicate signal and noise simultaneously.

a

Figure 3: Treating the uncertainty as a variable,with the same six examples. In plots R1 and R2, the standard error ofthe slope is mapped to the line width. The standard error of theintercept is mapped to the transparency, which is less conventional. Inplots S1 and S2, a bivariate colour palette is used with mean mapped tothe hue, and error mapped to saturation. Plots G1 and G2 represent themean of each group as a point, and the variance as an interval. Theuncertainty is integrated well with the signal, but for the choropleth,the result is undesirable: the signal is easier to see when the error ishigher..

3.3.2 Can we visualise a “single integrated uncertain value”?↩︎

The reality is, based on our discussion on inferential statistics, uncertainty isn’t a separate variable: it is a component of the random variable that is indistinguishable from the random variable itself. Similarities between the “as a variable” approach and the “as a statistic” approach are apparent when we read motivations for visualising an estimate (i.e. Figure 1) and variance (i.e. Figure 2) side-by-side using two separate graphics. This organisation is vulnerable to change blindness [55] as one needs to switch focus between two displays.) The preference for the methods utilised in Figure 3 is usually motivated by the difficulties in combining information on two separate graphics [27], [56], rather than an understanding that the “uncertainty statistic” approach is not philosophically sound. To achieve signal suppression, we need to visualise noise and signal together as a “single integrated uncertain value” [3] rather than as two separate statistics.

Just because we can still see the signal in Figure 3, that does not mean the reading of the estimate is completely independent of the uncertainty. When making any visualisations, we usually want the visual channels to be separable, that is, we don’t want the data represented through one visual channel to interfere with the others [57]. It is also interesting to know whether readers can see all structures when there are multiple structures present in a data plot, or whether they systematically fixate on one [58]. Separability may be desirable in standard data visualisation, but in uncertainty visualisation, it allows the estimate and its variance to be read independently, potentially leading to the uncertainty being ignored [11]. Therefore, rather than trying to maintain visual separability, the goals of uncertainty visualisation align far better with the pursuit of visual integration. In an ideal system, our estimate and uncertainty would be manipulated separately, but would be so well-integrated that they are read as a single channel by the human brain. The problem is that even if we can implement the most extreme versions of integrable, our methods fall short, as illustrated by the bivariate colour palette map in Figure 3. Colour hue and brightness are one of the classic examples of integrable variables [59], and decreasing saturation should make the colours harder to distinguish, but the signal is still clearly visible with the high variance. This is to say nothing of the fact that multi-dimensional colour palettes can make the graphics harder to read and less accessible [60].

3.3.3 Another example: mapping combined variables↩︎

The Value Suppressing Uncertainty Palette (VSUP) [27] was designed with the intention of preventing high uncertainty values from being extracted from a map by blending colours together as they become less certain. Figure 4 shows the spatial example (S1, S2) using the VSUP approach. Since the palette was designed with the extraction of individual values in mind and it has only been tested on simple value extraction tasks [27] or search tasks [23], we can see that, at least for our example, when the uncertainty is high, the spatial trend has functionally disappeared.

a

Figure 4: The spatial examples displayed with achoropleth map using a VSUP colour palette, where hue is blended whenincreased uncertainty. Plot B has successfully produced signalsuppression. Look closely at the scales, though: it may have anotherexplanation..

Finally, we have signal suppression! Well, not really, sorry, we tricked you. The two plots depicted in Figure 4 actually show the exact same data; they are both the low variance case. If you look closely, you can see that the two plots have a different scale, where plot A has been scaled according to our existing knowledge about this data, while plot B has been scaled using the range of the data passed to the plot. Since the variance and estimate are scaled independently, arbitrary differences in the range of our variance, unrelated to the estimate itself, will have significant impacts on the visual appearance of our plot. The scale issue in VSUP maps was also recognised by [61], who noted that the suppression of any one hypothesis largely depends on the methods we use to combine the palette, and the variance levels at which the blending occurs. This means that, for us to know that our plot will successfully perform signal suppression, we need to already know what signal we are trying to suppress and set up the VSUP palette accordingly. This means that VSUP maps are not suitable for exploratory data analysis.

3.3.4 Uncertainty and exploratory data analysis↩︎

The lack of uncertainty in descriptive statistics is due to the lack of inference. Descriptive statistics are actually a small piece of a much larger field, exploratory data analysis (EDA), that tends not to perform statistical inference. [62] described EDA as the process of searching for interesting hypotheses (“the greatest value of a picture is when it forces us to notice what we never expected to see”), and defined it in relation to confirmatory data analysis (CDA), the process of verifying a hypothesis. There are more subfields of EDA: initial data analysis [63], [64], which involves checking assumptions and data quality prior to CDA, and model diagnostics (e.g. [65]), including posterior checks of model fit. What binds these pursuits together is their reliance on visual summaries for making assessments and an absence of formal inference.

[66] argued that the EDA and CDA are not entirely distinct, as it is often difficult to draw a hard line. Our belief, as with many concepts, is that these approaches exist on a continuum, where we have an inherent trade-off between the number of hypotheses we can look for and the certainty of any conclusions reached. It can help to think of the knowledge-generating process of EDA and CDA as the nozzle on a hose with multiple spray options, where EDA is a fine misting spray that touches everything in the room, and CDA is a high-pressure jet capable of obliterating any and all debris from any single spot.

Viewing EDA and CDA as a dichotomy can create some confusion when it comes to understanding the source of uncertainty in our analysis. This is why we have avoided the topic until now, despite the fact that an uncertainty visualisation system for EDA is one of the most discussed topics in the field [2], [10], [20], [42], [49]. The EDA versus CDA dichotomy can be compared to the dichotomy between induction, for building theories, and deduction, for testing theories, from formal logic. One of the trade-offs in the two methods is that deductive conclusions provide certainty, while conclusions from induction are inherently uncertain. This means that EDA, the inductive counterpart, has uncertainty in any conclusions reached, with the requirement to follow up our newfound hypothesis with CDA if we want true certainty. This adds another layer of confusion to the study of uncertainty visualisation: conflating the uncertainty in our data with the uncertainty that is inherent to an exploratory process (EDA).

Authors often over-compensate for the inherent EDA uncertainty, pre-emptively hedging against every false inference that could possibly be drawn from a graphic. [66] argues there is no such thing as a “model-free” visualisation. [67] provides a grammar for visualising statistical model checks. The lineup protocol [68] provides the viewer with plots of the data in a field of plots of null data where any patterns seen are due to sampling variability. The Rorschach protocol, from the same paper, shows only null plots to give the reader some intuition for what spurious sampling patterns exist. [69] provides a statistical super-test against multiple comparisons driven by probabilistic arguments. There is a CDA quality to these approaches. The sum total of a lot of CDA is not EDA, just as swinging a high-pressure jet around a room is not equivalent to using a misting spray. While EDA and CDA may be along a continuum, we cannot simultaneously perform EDA and CDA as the approaches are, philosophically speaking, perpendicular to one another.

To truly create an uncertainty visualisation approach that is capable of EDA, we need to accept that uncertainty is inherent to the method, and it cannot be pre-emptively removed from the visualisations. If we accept this fact, then the only possible source of uncertainty in an uncertainty visualisation system that performs EDA is from the data itself. That is to say, the only possible way for there to be uncertainty in a visualisation designed for EDA is that the data itself represents inference that has been done earlier in our analysis. We can see this in the original description of Figure 1, where the uncertainty in all cases, even measurement error, represents inference that was performed earlier in our analysis.

In the VSUP approach, our ability to arbitrarily decide which values to blend at, or which suppression approach to use, means that the “uncertainty” we are visualising will be informed by the conclusions we are drawing, and not a product of the data itself – an antithetical approach to EDA. For a visualisation to be suitable for EDA, it should always look the same regardless of what hypothesis we plan to draw from it. As much of the transparency in data visualisation comes from this feature in EDA, it is reasonable to set it as a requirement of our visualisations. Ensuring this property means we cannot treat signal and noise as separate variables, but rather as a single integrated unit.

3.4 Uncertainty as a distribution↩︎

Rather than trying to reduce random variables to a single value or pair of values, why not visualise the whole distribution? This approach is found in computer science’s hypothetical outcome plots (HOPs), which animate a sequence of potential outcomes of a distribution [70], in geoscience’s pixel-maps [46], [48], and in statistics multiple forecasts [71].

3.4.1 Example: visualising samples↩︎

Figure 3 shows the six examples introduced in Figure 1, shown as distributions. The distribution is represented by either a sample of outcomes or a quantile dot plot. Plots R1 and R2 show a linear regression as a sample of possible outcomes from the distribution. The distribution used for both the slope and intercept is normal with the conventional mean and standard error. Plots S1 and S2 show a pixel map, which is a choropleth map where each county is coloured by a sample of outcomes from the distribution. Plots G1 and G2 show each univariate distribution as a quantile dot plot. We can see that the strong downward trend in the linear regression, the sine wave in the choropleth map, and the incremental increase in the univariate distributions are all clearly visible in the low variance case, but disappear in the high variance cases. The graphics have achieved signal suppression. Visualising the random variable as a distribution gives additional information, such as the previously hidden bimodality of the univariate distributions in G1 and G2.

a

Figure 5: Treating the uncertainty as adistribution, with the same six examples. Plots R1 and R2 show a linearregression as a sample of possible outcomes from the distribution. PlotsS1 and S2 show a pixel map. Plots G1 and G2 show the groups as quantiledot plots. In each case, the signal (regression line, sine wave,increasing trend) has disappeared with high uncertainty..

3.4.2 Quantified versus unquantified uncertainty↩︎

By showing our distribution as “data”, we are able to read the “uncertainty” plots using the same perceptual mechanisms we use to read the “non-uncertainty” plot. This should lead to more effective communication, as people tend to read more complicated visualisations like the bivariate and VSUP plots the same way they read the simple choropleth counterpart [23]. This approach also does not significantly hinder our ability to extract the individual values mapped by the previous plots, as extracting global statistics from a sample can be done with relative ease [72]. We are able to include more information by offloading more computation to visual processing, but what is the limitation on this approach? Which aspects of uncertainty should be processed by our visual system, and which should be processed by our statistical computation? It is not obvious from the question, but this is actually a question about how much of our uncertainty should be quantified.

Quantified uncertainty usually focuses narrowly on concepts such as probability, confidence intervals, variance, error, or precision [8], [41], [73], while unquantified uncertainty often includes a broader range of concepts like missing values, reliability, model validity, or source integrity [2], [29], [74][76]. When discussing uncertainty, we typically include these unquantified uncertainties, not because these things are uncertainty, but because they can create uncertainty when we perform inference. This is often because these unquantified uncertainties violate our assumptions of the uniformity of nature [39].

Sometimes we are able to visualise these assumption violations directly. For example, we can check for structure in our missing data using the naniar package [77] that allows us to include missing values as a “shadow” alongside our usual visualisations. This approach amounts to just “showing the data”, which is a simple but effective option for uncertainty visualisation that is largely overlooked. While this approach is useful for better understanding data, it will not eliminate trends that have become invalid due to structure in missing data or an invalid model. We can only integrate uncertainty as noise when that uncertainty has been quantified as an effect on the estimates we are visualising. This is not to say one method is preferable; visualising both quantified and unquantified uncertainty is necessary for a healthy analysis. Data analysis often works in cycles, where we find assumption violations using EDA, quantify the effect of these violations on inference, and then visualise the output of that inference using uncertainty visualisation.

4 Evaluating uncertainty visualisations↩︎

Unfortunately, the conflicting results in the field are not limited to plot design and extend to the experimental findings as well [3], [4], [10]. There are as many explanations for the noisy evaluation studies as there are contradictions in the research itself. [78] believes there is some interference in results from participants’ prior beliefs; [4] believes the noise in the literature could come from visual heuristics, subjective probabilities, unknown participant utility functions, or a misunderstanding of statistical concepts (such as confidence intervals). [3] suggest the perception of visualisation changes by audience, so we cannot expect the same results between different subpopulations. [79] attributes evaluation difficulties to cognitive load from complicated uncertainty visualisations, as well as the participants’ prior experience in the topic. While these issues will certainly have some impact on our ability to synthesise, none of them is unique to uncertainty visualisation. Rather, the issue is likely due to a disconnect between the evaluation methods used and the stated goals of each experiment, a common issue in visualisation evaluation experiments [59].

4.1 Current evaluation methods↩︎

4.1.1 Value extraction↩︎

Uncertainty visualisations are most commonly evaluated based on how accurately viewers can extract an estimate and its variance [3], [80]. This is not unusual, as direct observation is the simplest way to verify that information can be accurately read from a graph [59]. Unfortunately, this approach doesn’t work for uncertainty visualisation. The second we ask a specific question about a statistic, that statistic becomes inferential, even if the plot was not the intent behind the question. By shifting the focus from \(\hat{X}\) to \(Var(\hat{X})\) or \(P(\hat{X}<x)\), we end up evaluating visualisations on their ability to convey uncertainty statistics, rather than their ability to perform signal suppression. Even if the authors do not realise it themselves, there is nothing unique to uncertainty in these studies, so when we boil the findings down to generalised results, they simply restate existing principles within information visualisation. Some of the findings are obvious: participants were more accurate when reading a probability expressed as text than when they had to extract it from a graphic [81], [82]. Other studies replicate existing research, such as the finding that a probability mapped to a position is more accurate than one mapped to an area [12], [37], established as part of the hierarchy of perceptual tasks more than 40 years ago [83], replicated by [84]. This extends beyond simple accuracy evaluations: [36] found that colour was more effective than size when searching for extrema in variances; we have known that pre-attentive aesthetics, such as colour, are more efficient for search tasks since the 1980s [59]. By classifying these studies as evaluations of “uncertainty visualisation” while evaluating uncertainty as a signal, we are encouraged to see successful examples of signal suppression as failure. This approach leads authors to advise against particular aesthetic mappings for uncertainty, because they cause participants to have more difficulty extracting values [48]. This conclusion is antithetical to the goals of signal suppression and occurs because these methods evaluate uncertainty as a signal, not as noise.

4.1.2 Trust, confidence, and risk aversion↩︎

If we cannot directly measure uncertainty for fear that it turns into a signal, we might then assume we can measure the secondary benefits of increased transparency. This seems to be the approach of many visualisation authors, as secondary benefits such as trust, confidence, and risk aversion are all frequently used in uncertainty evaluation studies [80]. Unfortunately, measuring these secondary effects often leads to confusing conclusions that simultaneously argue for and against the inclusion of uncertainty.

This is most commonly noticed in the use of trust as a measure, as several authors have commented that measuring trust, and not transparency, can lead to a questionable subtext that argues against transparency [1], [85]. We see this directly play out in the visualisation literature, where surveyed visualisation authors explicitly said they didn’t include uncertainty due to the fact that they might decrease trust in their conclusions [24]. This sentiment is also true for confidence, as [48] commented that visually integrable depictions of uncertainty should be avoided, as they decrease the viewer’s confidence in their extracted data values.

Another secondary effect that is similar to trust and confidence is risk aversion. Risk aversion is an economic term used to describe an agent who would choose a random variable with a lower expected payout because it also has a lower variance. Risk aversion’s role in the uncertainty visualisation is unclear, as authors will argue uncertainty should elicit more risk aversion in one paper [80], and argue for less risk aversion (by proxy of suggesting rational agents as a benchmark) in the next [86]. Ultimately, these approaches have similar issues to value extraction studies, except they are slightly more confusing in their goals, leading them to simultaneously argue for and against the inclusion of uncertainty in a visualisation.

4.1.3 Alternative approaches↩︎

Often, authors understand that the effects of uncertainty are more complicated than simple value extraction. These studies indicate that accurately capturing uncertainty will be more complicated than simply avoiding value extraction or trust as a measure.

One approach is to ask indeterminate questions, such as asking participants for the “best estimate” [12], or to select which distribution is the “furthest to the right” in a lineup [51]. In both cases, the ground truth is based on the mean of the distribution, which is not as indeterminate as the question. This approach can lead to inconclusive results, as we are left unclear whether it was the phrasing of the question or the plot design that caused the participants to answer incorrectly.

On the other hand, questions that are incredibly specific about the distribution information can confuse the participants and induce noisy results. For example, [70] asked participants to compare two normally distributed groups, A and B, and had many participants say that group A was more likely to be bigger, despite group B having a higher mean. [37] asked participants the “probability that the interval has already ended at the marked point in time?” and participants replied with the probability that the interval had already started.

The confusion around trying to capture the effects of uncertainty can also (understandably) extend to the authors of the study itself. We can see an example of this in [87]. In order to answer the question correctly, the first experiment required participants to assume that an oil rig being “more likely to be hit” by a hurricane would not translate to the rig sustaining “more damage”. The second experiment required participants to assume the opposite.

4.1.4 Effective methods↩︎

This is not to say all evaluation studies fail to properly evaluate uncertainty as noise. There are several studies that ask participants to identify a particular signal that the noise is trying to obfuscate [26], [31], which seems to be an effective method. The only problem with these studies is that most uncertainty visualisation methods are not integrated into the grammar of graphics [29], [43], so we regularly see comparisons between disparate plots that would never be considered substitutes for one another outside the artificial “uncertainty visualisation” framework they are placed within. For example, [26] compared static bar charts with error bars to a bar chart with animated samples, meaning that any difference in participants’ ability to read the plot could be due to the statistic (confidence interval versus sample), the geometry (bars versus intervals), or the use of animation (static versus animated plots). This issue was rectified in their second experiment, where they compared overlayed and animated samples, giving us an insight into the types of visualisations that are appropriate to compare in uncertainty visualisation experiments. This means that even when evaluations are done correctly, there is no generalisable theory we can take from the results.

The point here is not to accuse the authors of poor academic rigour. The papers are (usually) logically consistent and well-formulated pieces of work. Rather, the point is to illustrate that evaluating uncertainty as noise is surprisingly difficult. Designing tests for signal suppression will require a formalisation of uncertainty within the grammar of graphics, as well as improved evaluation methods.

4.2 Implicit Hypothesis Testing↩︎

The main problem with current uncertainty visualisation evaluations is that they often require explicit (or convoluted) questions about the variance. Asking direct questions about the statistics or outcomes is not an explicit requirement of visualisation evaluations. In their review of testing statistical graphics, [59] drew a distinction between explicit tests, where participants are asked direct questions about specific features of a plot, and implicit testing, where users identify both the purpose and function of the plot. The lineup protocol is the most salient example of the implicit approach. Lineups are a confirmatory visualisation tool where participants are shown a set of \(M\) plots, and asked to identify the plot that is the “most different”, leaving participants to decide what “most different” means to them, even if it is not what the authors intended [58]. The implicit test does not limit the versatility of the approach, with the lineup being used to evaluate the effectiveness of different types of plots [51], colour palettes [88], and design decisions [58].

Lineup protocols are not only useful for implicit testing: they also have parallels to hypothesis testing that can be leveraged in uncertainty visualisation. The concept of signal suppression is, at its core, an assertion of statistical validity: the visibility of signals should be directly proportional to \(p\)-values or some equivalent measure. This comparison is not new in uncertainty visualisation, where parallels have been drawn to frequentist statistics by [31], who compared results to Cohen’s D, and to Bayesian statistics by [78], who evaluated plots based on their impact on the users’ prior beliefs. The comparison to hypothesis testing is far more natural for the lineup protocol, which has a visual test statistic [89] and can be compared to standard statistical tests using power curves [90]. The connections between lineups and uncertainty visualisation are numerous and have been previously identified in the development of HOPs [70].

The lineup protocol and uncertainty visualisations are similar: lineups were designed for checking if perceived patterns are real or merely the result of chance [68], [91]. As both approaches are attempting to do the same thing, it is likely that we are unable to leverage the lineup protocol directly to evaluate uncertainty visualisation, but a new evaluation methodology should try to learn from the success of the lineup approach. Designing an implicit testing method for uncertainty visualisation that allows us to draw parallels to standard notions of statistical significance would solve many of the issues with the current evaluation approaches.

5 Conclusions and Future Work↩︎

This paper examines the literature and provides suggestions for a structural framework to support uncertainty visualisation. Particularly, we propose that uncertainty visualisation should accomplish signal suppression, dampening weak signals and amplifying strong signals. We have also highlighted several gaps in the existing literature.

Experimental practices on uncertainty visualisation need standards. Some existing evaluation experiments treat uncertainty as a signal, while others treat uncertainty as noise. As a result, it is difficult to combine results from papers to get a meaningful sense of how uncertainty information is understood by a viewer. Researchers need to ensure that when they identify the motivation behind their visualisation technique, their evaluation methods align with the stated goals of the paper.

Experimental methods that evaluate uncertainty as noise need to be developed. Research into separability and integrability of signal and noise is of particular interest to uncertainty visualisation, as it allows assessment of the interference between the two. When designing experiments, authors often choose aesthetics that are visually distinguishable; uncertainty visualisation authors should consider doing the opposite.

Uncertainty needs to be formalised within the grammar of graphics. Some of this formalisation was done by [40], but it focuses only on the visualisation of univariate distributions. Giving authors the ability to describe uncertainty visualisations in terms of statistics, geometries and aesthetics will support evaluation experiments that can build towards a cohesive theory of visualising uncertainty.

Software that allows users to easily perform signal suppression is needed. Existing uncertainty visualisation methods view a distribution as its own object, and there are no software options treating “an uncertainty visualisation as a function of an existing visualisation” philosophy.

Signal suppression is an undeveloped area of visualisation research, and developing methods for the practice may require us to challenge our entire notion of what makes a good visualisation.

Reproducibility↩︎

The R packages were used for this work were: tidyverse [92], RColorBrewer [93], scales [94], sf [95], urbnmapr [96], flextable [97], colorspace [98], ggdist [40], ggdibbler [99], patchwork [100], distributional [101], ggthemes [102], broom [103], and rgeos [104]. The GitHub repository for this paper can be found at https://github.com/harriet-mason/ARSA-UncertaintyLitReview, which contains the files required to reproduce this article in full.

References↩︎

[1]
D. Spiegelhalter, “Risk and uncertainty communication,” Annual Review of Statistics and Its Application, vol. 4, pp. 31–60, 2017, doi: 10.1146/annurev-statistics-010814-020148.
[2]
H. Griethe and H. Schumann, “The visualization of uncertain data: Methods and problems,” in SimVis, 2006, vol. 6, pp. 143–156.
[3]
C. Kinkeldey, A. M. MacEachren, and J. Schiewe, “How to assess visual communication of uncertainty? A systematic review of geospatial uncertainty visualisation user studies,” Cartographic Journal, vol. 51, no. 4, pp. 372–386, 2014, doi: 10.1179/1743277414Y.0000000099.
[4]
J. Hullman, “Why evaluating uncertainty visualization is error prone,” ACM International Conference Proceeding Series, vol. 24–October, pp. 143–151, 2016, doi: 10.1145/2993901.2993919.
[5]
F. J. Anscombe, “Graphs in statistical analysis,” The American Statistician, vol. 27, no. 1, pp. 17–21, 1973, [Online]. Available: https://www.tandfonline.com/doi/abs/10.1080/00031305.1973.10478966.
[6]
J. Matejka and G. Fitzmaurice, “Same stats, different graphs: Generating datasets with varied appearance and identical statistics through simulated annealing,” in Proceedings of the 2017 CHI conference on human factors in computing systems, 2017, pp. 1290–1294, doi: 10.1145/3025453.3025912.
[7]
S. Locke and L. D’Agostino McGowan, R package version 0.1.4datasauRus: Datasets from the datasaurus dozen. 2018.
[8]
J. Hullman, M. Kay, Y. S. Kim, and S. Shrestha, “Imagining replications: Graphical prediction discrete visualizations improve recall estimation of effect uncertainty,” IEEE Transactions on Visualization and Computer Graphics, vol. 24, no. 1, pp. 446–456, 2018, doi: 10.1109/TVCG.2017.2743898.
[9]
D. G. Goldstein and D. Rothschild, “Lay understanding of probability distributions,” Judgment and Decision Making, vol. 9, no. 1, pp. 1–14, 2014.
[10]
A. M. MacEachren et al., ISBN: 1523040054738“Visualizing geospatial information uncertainty: What we know and what we need to know,” Cartography and Geographic Information Science, vol. 32, no. 3, pp. 139–160, 2005, doi: 10.1559/1523040054738936.
[11]
L. Padilla, M. Kay, and J. Hullman, “Uncertainty visualization,” in Computational statistics in data science, W. W. Piegorsch, R. A. Levine, H. H. Zhang, and T. C. M. Lee, Eds. Hoboken, NJ: John Wiley & Sons, 2022, pp. 405–426.
[12]
H. Ibrekk and M. G. Morgan, “Graphical communication of uncertain quantities to nontechnical people,” Risk Analysis, vol. 7, no. 4, pp. 519–529, 1987, doi: 10.1111/j.1539-6924.1987.tb00488.x.
[13]
A. M. MacEachren, “Visualizing uncertain information,” Cartographic Perspectives, no. 13, pp. 10–19, Jun. 1992, doi: 10.14714/CP13.1000.
[14]
L. Padilla, H. Hosseinpour, R. Fygenson, J. Howell, R. Chunara, and E. Bertini, “Impact of COVID-19 forecast visualizations on pandemic risk perceptions,” Scientific Reports 2022 12:1, vol. 12, no. 1, pp. 1–14, Feb. 2022, doi: 10.1038/s41598-022-05353-1.
[15]
L. Padilla, M. Powell, M. Kay, and J. Hullman, “Uncertain about uncertainty: How qualitative expressions of forecaster confidence impact decision-making with uncertainty visualizations,” Frontiers in Psychology, vol. 11, Jan. 2021, doi: 10.3389/fpsyg.2020.579267.
[16]
A. Kale, M. Kay, and J. Hullman, “Visual reasoning strategies for effect size judgments and decisions,” IEEE Transactions on Visualization and Computer Graphics, vol. 27, no. 2, pp. 272–282, 2021, doi: 10.1109/TVCG.2020.3030335.
[17]
J. Zhao, Y. Wang, M. V. Mancenido, E. K. Chiou, and R. Maciejewski, “Evaluating the impact of uncertainty visualization on model reliance,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 7, pp. 4093–4107, 2023, doi: 10.1109/TVCG.2023.3251950.
[18]
F. Yang et al., “Swaying the public? Impacts of election forecast visualizations on emotion, trust, and intention in the 2022 us midterms,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 1, pp. 23–33, 2023.
[19]
A. Sarma et al., “Evaluating the use of uncertainty visualisations for imputations of data missing at random in scatterplots,” IEEE Transactions on Visualization and Computer Graphics, vol. 29, no. 1, pp. 602–612, 2023, doi: 10.1109/TVCG.2022.3209348.
[20]
A. Sarma, X. Pu, Y. Cui, M. Correll, E. T. Brown, and M. Kay, “Odds and insights: Decision quality in exploratory data analysis under uncertainty,” in Proceedings of the CHI conference on human factors in computing systems, 2024, doi: 10.1145/3613904.3641995.
[21]
R. Koonchanok, G. Y. Tawde, G. R. Narayanasamy, S. Walimbe, and K. Reda, “Visual belief elicitation reduces the incidence of false discovery,” Conference on Human Factors in Computing Systems - Proceedings, 2023, doi: 10.1145/3544548.3580808.
[22]
S. Chakraborty, P. Kiefer, and M. Raubal, “The influence of uncertainty visualization on cognitive load in a safety- and time-critical decision-making task,” International Journal of Geographical Information Science, vol. 38, no. 8, pp. 1583–1610, Aug. 2024, doi: 10.1080/13658816.2024.2348747.
[23]
A. Ndlovu, H. Shrestha, and L. T. Harrison, “Taken by surprise? Evaluating how bayesian surprise & suppression influences peoples’ takeaways in map visualizations,” in 2023 IEEE visualization and visual analytics (VIS), 2023, pp. 136–140.
[24]
J. Hullman, “Why authors don’t visualize uncertainty,” IEEE Transactions on Visualization and Computer Graphics, vol. 26, no. 1, pp. 130–139, Jan. 2020, doi: 10.1109/TVCG.2019.2934287.
[25]
C. F. Manski, “The lure of incredible certitude,” Economics and Philosophy, vol. 36, no. 2, pp. 216–245, 2020, doi: 10.1017/S0266267119000105.
[26]
A. Kale, F. Nguyen, M. Kay, and J. Hullman, “Hypothetical outcome plots help untrained observers judge trends in ambiguous data,” IEEE transactions on visualization and computer graphics, vol. 25, no. 1, pp. 892–902, 2018.
[27]
M. Correll, D. Moritz, and J. Heer, “Value-suppressing uncertainty palettes,” Conference on Human Factors in Computing Systems - Proceedings, vol. 2018–April, pp. 1–11, 2018, doi: 10.1145/3173574.3174216.
[28]
N. Boukhelifa, A. Bezerianos, T. Isenberg, and J. D. Fekete, “Evaluating sketchiness as a visual variable for the depiction of qualitative uncertainty,” IEEE Transactions on Visualization and Computer Graphics, vol. 18, no. 12, pp. 2769–2778, 2012, doi: 10.1109/TVCG.2012.220.
[29]
L. Wilkinson, The grammar of graphics. Berlin, Heidelberg: Springer-Verlag, 2005.
[30]
C. Lee, T. Yang, G. D. Inchoco, G. M. Jones, and A. Satyanarayan, “Viral visualizations: How coronavirus skeptics use orthodox data practices to promote unorthodox science online,” in Proceedings of the 2021 CHI conference on human factors in computing systems, 2021, pp. 1–18.
[31]
M. Correll and M. Gleicher, “Error bars considered harmful: Exploring alternate encodings for mean and error,” IEEE Transactions on Visualization and Computer Graphics, vol. 20, no. 12, pp. 2142–2151, 2014, doi: 10.1109/TVCG.2014.2346298.
[32]
S. H. Begg, M. B. Welsh, and R. B. Bratvold, “Uncertainty vs. Variability: What’s the difference and why is it important?” in SPE hydrocarbon economics and evaluation symposium, May 2014, vol. SPE Hydrocarbon Economics and Evaluation Symposium, doi: 10.2118/169850-MS.
[33]
A. Gustafson and R. E. Rice, “The effects of uncertainty frames in three science communication topics,” Science Communication, vol. 41, no. 6, pp. 679–706, 2019, doi: 10.1177/1075547019870811.
[34]
W. E. Walker et al., “Defining uncertainty,” Integrated Assessment, vol. 4, no. 1, pp. 5–17, 2003, [Online]. Available: https://www.narcis.nl/publication/RecordID/oai:tudelft.nl:uuid:fdc0105c-e601-402a-8f16-ca97e9963592.
[35]
D. M. Benjamin and D. V. Budescu, “The role of type and source of uncertainty on the processing of climate models projections,” Frontiers in Psychology, vol. 9, no. MAR, pp. 1–17, 2018, doi: 10.3389/fpsyg.2018.00403.
[36]
J. Sanyal, S. Zhang, G. Bhattacharya, P. Amburn, and R. J. Moorhead, “A user study to compare four uncertainty visualization methods for 1D and 2D datasets,” IEEE Transactions on Visualization and Computer Graphics, vol. 15, no. 6, pp. 1209–1218, 2009, doi: 10.1109/TVCG.2009.114.
[37]
T. Gschwandtner, M. Bögl, P. Federico, and S. Miksch, “Visual encodings of temporal uncertainty: A comparative user study,” IEEE Transactions on Visualization and Computer Graphics, vol. 22, no. 1, pp. 539–548, Jan. 2016, doi: 10.1109/TVCG.2015.2467752.
[38]
C. Olston and J. D. Mackinlay, “Visualizing data with bounded uncertainty,” Proceedings - IEEE Symposium on Information Visualization, INFO VIS, vol. 2002–Janua, pp. 37–40, 2002, doi: 10.1109/INFVIS.2002.1173145.
[39]
J. Otsuka, Thinking about statistics: The philosophical foundations, 1st ed. New York: Routledge, 2023, p. 204.
[40]
M. Kay, ggdist: Visualizations of distributions and uncertainty in the grammar of graphics,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 1, pp. 414–424, 2023.
[41]
J. Thomson, E. Hetzler, A. MacEachren, M. Gahegan, and M. Pavel, “A typology for visualizing uncertainty,” Visualization and Data Analysis 2005, vol. 5669, no. March 2005, p. 146, 2005, doi: 10.1117/12.587254.
[42]
A. Hadjimichael, J. Schlumberger, and M. Haasnoot, “Data visualisation for decision making under deep uncertainty: Current challenges and opportunities,” Environmental Research Letters, vol. 19, no. 11, p. 111011, Nov. 2024, doi: 10.1088/1748-9326/ad858b.
[43]
H. Wickham, “A layered grammar of graphics,” Journal of Computational and Graphical Statistics, vol. 19, no. 1, pp. 3–28, 2010, doi: 10.1198/jcgs.2009.07098.
[44]
Henderson and Velleman, “Building multiple regression models interactively,” Biometrics, vol. 37, pp. 391–411, 1981.
[45]
P. M. Kuhnert, D. E. Pagendam, R. Bartley, D. W. Gladish, S. E. Lewis, and Z. T. Bainbridge, “Making management decisions in the face of uncertainty: A case study using the Burdekin catchment in the Great Barrier Reef,” Marine and Freshwater Research, vol. 69, no. 8, pp. 1187–1200, 2018, doi: 10.1071/MF17237.
[46]
L. Lucchesi, P. Kuhnert, and C. Wikle, Vizumap: An R package for visualising uncertainty in spatial data,” Journal of Open Source Software, vol. 6, no. 59, p. 2409, 2021, doi: 10.21105/joss.02409.
[47]
K. Potter, J. Kniss, R. Riesenfeld, and C. R. Johnson, “Visualizing summary statistics and uncertainty,” Computer Graphics Forum, vol. 29, no. 3, pp. 823–832, 2010, doi: 10.1111/j.1467-8659.2009.01677.x.
[48]
S. Blenkinsop, P. Fisher, L. Bastin, and J. Wood, “Evaluating the perception of uncertainty in alternative visualization strategies,” Cartographica, vol. 37, no. 1, pp. 1–13, 2000, doi: 10.3138/3645-4v22-0m23-3t52.
[49]
V. Peña-Araya, C. M. Fontaine, X. Wei, G. Delpech, and A. Bezerianos, “Uncertainty in science is malleable. Advocating for user-agency in defining uncertainty in visualizations: A case study in geology,” in Proceedings of the 2025 CHI conference on human factors in computing systems, 2025, pp. 1–18.
[50]
H. Wickham and H. Hofmann, “Product plots,” IEEE Transactions on Visualization and Computer Graphics, vol. 17, no. 12, pp. 2223–2230, 2011, doi: 10.1109/TVCG.2011.227.
[51]
H. Hofmann, L. Follett, M. Majumder, and D. Cook, “Graphical tests for power comparison of competing designs,” IEEE Transactions on Visualization and Computer Graphics, vol. 18, no. 12, pp. 2441–2448, Dec. 2012, doi: 10.1109/TVCG.2012.230.
[52]
X. L. Meng, “A trio of inference problems that could win you a nobel prize in statistics (if you help fund it),” Past, Present, and Future of Statistical Science, pp. 537–562, 2014, doi: 10.1201/b16720-52.
[53]
T. S. Wallsten, D. V. Budescu, I. Erev, and A. Diederich, “Evaluating and combining subjective probability estimates,” Journal of Behavioral Decision Making, vol. 10, no. 3, pp. 243–268, 1997, doi: 10.1002/(sici)1099-0771(199709)10:3<243::aid-bdm268>3.0.co;2-m.
[54]
B. Fischhoff and A. L. Davis, “Communicating scientific uncertainty,” Proceedings of the National Academy of Sciences of the United States of America, vol. 111, pp. 13664–13671, 2014, doi: 10.1073/pnas.1317504111.
[55]
D. J. Simons and D. T. Levin, “Change blindness,” Trends in Cognitive Sciences, vol. 1, pp. 261–267, 1997, doi: 10.1016/S1364-6613(97)01080-2.
[56]
D. Moritz, D. Fisher, B. Ding, and C. Wang, “Trust, but verify: Optimistic visualizations of approximate queries for exploring big data,” in Proceedings of the 2017 CHI conference on human factors in computing systems, 2017, pp. 2904–2915.
[57]
S. Smart and D. A. Szafir, “Measuring the separability of shape, size, and color in scatterplots,” Conference on Human Factors in Computing Systems - Proceedings, pp. 1–14, 2019, doi: 10.1145/3290605.3300899.
[58]
S. VanderPlas and H. Hofmann, “Clusters beat trend!? Testing feature hierarchy in statistical graphics,” Journal of Computational and Graphical Statistics, vol. 26, no. 2, pp. 231–242, 2017.
[59]
S. Vanderplas, D. Cook, and H. Hofmann, “Testing statistical charts: What makes a good graph?” Annual Review of Statistics and Its Application, vol. 7, no. 1, pp. 61–88, 2020.
[60]
S. VanderPlas and H. Hofmann, “Signs of the sine illusion—why we need to care,” Journal of Computational and Graphical Statistics, vol. 24, no. 4, pp. 1170–1190, 2015.
[61]
M. Kay, “How much value should an uncertainty palette suppress if an uncertainty palette should suppress value? Statistical and perceptual perspectives.” OSF Preprints, Oct. 2019, doi: 10.31219/osf.io/6xcnw.
[62]
J. W. Tukey et al., Exploratory data analysis, vol. 2. Springer, 1977.
[63]
M. Huebner, W. Vach, and S. le Cessie, “A systematic approach to initial data analysis is good research practice,” The Journal of Thoracic and Cardiovascular Surgery, vol. 151, no. 1, pp. 25–27, 2016, doi: 10.1016/j.jtcvs.2015.09.085.
[64]
C. Chatfield, “The initial examination of data,” Journal of the Royal Statistical Society. Series A (General), vol. 148, no. 3, pp. 214–253, 1985, doi: 10.2307/2981969.
[65]
D. A. Belsley, E. Kuh, and R. E. Welsch, Regression diagnostics: Identifying influential data and sources of collinearity. New York: Wiley, 1980.
[66]
J. Hullman and A. Gelman, “Designing for interactive exploratory data analysis requires theories of graphical inference,” Harvard Data Science Review, pp. 1–70, 2021, doi: 10.1162/99608f92.3ab8a587.
[67]
Z. Guo, A. Kale, M. Kay, and J. Hullman, VMC: A grammar for visualizing statistical model checks,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 1, pp. 798–808, 2025, doi: 10.1109/TVCG.2024.3456402.
[68]
A. Buja et al., “Statistical inference for exploratory data analysis and model diagnostics,” Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, vol. 367, no. 1906, pp. 4361–4383, Nov. 2009, doi: 10.1098/rsta.2009.0120.
[69]
R. Savvides, A. Henelius, E. Oikarinen, and K. Puolamäki, “Significance of patterns in data visualisations,” in Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, Jul. 2019, pp. 1509–1517, doi: 10.1145/3292500.3330994.
[70]
J. Hullman, P. Resnick, and E. Adar, “Hypothetical outcome plots outperform error bars and violin plots for inferences about reliability of variable ordering,” PLoS ONE, vol. 10, no. 11, Nov. 2015, doi: 10.1371/journal.pone.0142444.
[71]
R. J. Hyndman and G. Athanasopoulos, Forecasting: Principles and practice, 3rd edition. Melbourne, Australia: OTexts, 2021.
[72]
S. L. Franconeri, “Three perceptual tools for seeing and understanding visualized data,” Current Directions in Psychological Science, vol. 30, no. 5, pp. 367–375, 2021, doi: 10.1177/09637214211009512.
[73]
A. M. Maceachren, R. E. Roth, J. O’Brien, B. Li, D. Swingley, and M. Gahegan, “Visual semiotics & uncertainty visualization: An empirical study,” IEEE Transactions on Visualization and Computer Graphics, vol. 18, no. 12, pp. 2496–2505, 2012, doi: 10.1109/TVCG.2012.279.
[74]
A. T. Pang, C. M. Wittenbrink, and S. K. Lodha, “Approaches to uncertainty visualization,” Visual Computer, vol. 13, no. 8, pp. 370–390, 1997, doi: 10.1007/s003710050111.
[75]
B. Pham, A. Streit, and R. Brown, Visualization of information uncertainty: Progress and challenges,” in Advanced information and knowledge processing, vol. 36, Springer-Verlag London Ltd, 2009, pp. 19–48.
[76]
N. Boukhelifa, M.-E. Perrin, S. Huron, and J. Eagan, “How data workers cope with uncertainty: A task characterisation study,” in Proceedings of the 2017 CHI conference on human factors in computing systems, 2017, pp. 3645–3656, doi: 10.1145/3025453.3025738.
[77]
N. Tierney and D. Cook, “Expanding tidy data principles to facilitate missing data exploration, visualization and assessment of imputations,” Journal of Statistical Software, vol. 105, no. 7, pp. 1–31, 2023, doi: 10.18637/jss.v105.i07.
[78]
Y. S. Kim, L. A. Walls, P. Krafft, and J. Hullman, “A bayesian cognition approach to improve data visualization,” Conference on Human Factors in Computing Systems - Proceedings, pp. 1–14, 2019, doi: 10.1145/3290605.3300912.
[79]
A. Brennen and S. Tuerk, ISBN: 9781450356206“An instrument for evaluating uncertainty visualization techniques,” Conference on Human Factors in Computing Systems - Proceedings, vol. 2018–April, pp. 1–6, 2018, doi: 10.1145/3170427.3188649.
[80]
J. Hullman, X. Qiao, M. Correll, A. Kale, and M. Kay, “In pursuit of error: A survey of uncertainty visualization evaluation,” IEEE Transactions on Visualization and Computer Graphics, vol. 25, no. 1, pp. 903–913, Jan. 2019, doi: 10.1109/TVCG.2018.2864889.
[81]
L. Cheong, S. Bleisch, A. Kealy, K. Tolhurst, T. Wilkening, and M. Duckham, “Evaluating the impact of visualization of wildfire hazard upon decision-making under uncertainty,” International Journal of Geographical Information Science, vol. 30, no. 7, pp. 1377–1404, 2016, doi: 10.1080/13658816.2015.1131829.
[82]
S. Savelli and S. Joslyn, “The advantages of predictive interval forecasts for non-expert users and the impact of visualizations,” Applied Cognitive Psychology, vol. 27, no. 4, pp. 527–541, 2013.
[83]
W. S. Cleveland and R. McGill, “Graphical perception: Theory, experimentation, and application to the development of graphical methods,” Journal of the American Statistical Association, vol. 79, no. 387, pp. 531–554, 1984, doi: 10.1080/01621459.1984.10478080.
[84]
J. Heer and M. Bostock, “Crowdsourcing graphical perception: Using mechanical turk to assess visualization design,” in Proceedings of the SIGCHI conference on human factors in computing systems, 2010, pp. 203–212, doi: 10.1145/1753326.1753357.
[85]
O. O’Neill, “Linking trust to trustworthiness,” International Journal of Philosophical Studies, vol. 26, no. 2, pp. 293–300, 2018, doi: 10.1080/09672559.2018.1454637.
[86]
Y. Wu, Z. Guo, M. Mamakos, J. Hartline, and J. Hullman, “The rational agent benchmark for data visualization.” 2023, [Online]. Available: https://arxiv.org/abs/2304.03432.
[87]
L. Padilla, I. Ruginski, and S. Creem-Regehr, “Effects of ensemble and summary displays on interpretations of geospatial uncertainty data,” Cognitive Research: Principles and Implications, vol. 2, no. 1, Dec. 2017, doi: 10.1186/s41235-017-0076-1.
[88]
K. Reda and D. A. Szafir, “Rainbows revisited: Modeling effective colormap design for graphical inference,” IEEE Transactions on Visualization and Computer Graphics, vol. 27, no. 2, pp. 1032–1042, Feb. 2021, doi: 10.1109/TVCG.2020.3030439.
[89]
M. Majumder, H. Hofmann, and D. Cook, “Validation of visual statistical inference, applied to linear models,” Journal of the American Statistical Association, vol. 108, no. 503, pp. 942–956, Sep. 2013, doi: 10.1080/01621459.2013.808157.
[90]
W. Li, D. Cook, E. Tanaka, and S. VanderPlas, “A plot is worth a thousand tests: Assessing residual diagnostics with the lineup protocol,” Journal of Computational and Graphical Statistics, pp. 1–19, May 2024, doi: 10.1080/10618600.2024.2344612.
[91]
H. Wickham, D. Cook, H. Hofmann, and A. Buja, “Graphical inference for infovis,” IEEE Transactions on Visualization and Computer Graphics, vol. 16, pp. 973–979, 2010, doi: 10.1109/TVCG.2010.161.
[92]
H. Wickham et al., “Welcome to the tidyverse,” Journal of Open Source Software, vol. 4, no. 43, p. 1686, 2019, doi: 10.21105/joss.01686.
[93]
E. Neuwirth, R package version 1.1-3RColorBrewer: ColorBrewer palettes. 2022.
[94]
H. Wickham, T. L. Pedersen, and D. Seidel, R package version 1.3.0scales: Scale functions for visualization. 2023.
[95]
E. Pebesma and R. Bivand, Spatial data science: With applications in R. Chapman; Hall/CRC, 2023.
[96]
S. Strochak, K. Ueyama, and A. Williams, R package version 0.0.0.9002Urbnmapr: State and county shapefiles in sf and tibble format. 2024.
[97]
D. Gohel and P. Skintzos, R package version 0.9.6Flextable: Functions for tabular reporting. 2024.
[98]
R. Stauffer, G. J. Mayr, M. Dabernig, and A. Zeileis, “Somewhere over the rainbow: How to make effective use of colors in meteorological visualizations,” Bulletin of the American Meteorological Society, vol. 96, no. 2, pp. 203–216, 2009, doi: 10.1175/BAMS-D-13-00155.1.
[99]
H. Mason, D. Cook, S. Goodwin, and S. VanderPlas, Ggdibbler: Add uncertainty to data visualisations. 2026.
[100]
T. L. Pedersen, R package version 1.3.2patchwork: The composer of plots. 2025.
[101]
M. O’Hara-Wild, M. Kay, A. Hayes, and R. Hyndman, R package version 0.5.0distributional: Vectorised probability distributions. 2024.
[102]
J. B. Arnold, R package version 5.1.0Ggthemes: Extra themes, scales and geoms for ’ggplot2’. 2024.
[103]
D. Robinson, A. Hayes, S. Couch, and E. Hvitfeldt, R package version 1.0.12broom: Convert statistical objects into tidy tibbles. 2026.
[104]
R. Bivand and C. Rundel, R package version 0.6-3Rgeos: Interface to geometry engine - open source (’GEOS’). 2023.