July 07, 2026
The uncanny valley is a long-standing empirical rule in humanoid robot design: making robots more human-like can reduce, rather than increase, affinity. Yet existing guidelines, such as adopting robot-like appearances, avoiding excessive realism, and reducing cross-modal mismatches, remain difficult to use for algorithmic design because they are not expressed as manipulable variables. Here, we propose a hierarchical Bayesian generative model that operationalizes these guidelines as mathematical design variables. The model represents affinity toward humanoid robots as posterior-weighted negative category-conditional surprise and explains category ambiguity and perceptual mismatch as increases in surprise. It maps uncanny-valley mechanisms onto four variables: deviation from the predicted robot-category mean, inconsistency in human likeness across modalities, prediction uncertainty, and observational uncertainty. Simulations showed that category ambiguity and appearance–motion mismatch can produce affinity reductions, and that uncertainty reshapes the valley. In a human-subject experiment with robot–human morphing images, we manipulated prediction uncertainty using blurred prior robot stimuli and observational uncertainty using blurred evaluation stimuli. Increased observational uncertainty attenuated the decrease in familiarity ratings at intermediate human likeness, whereas low prediction uncertainty increased ratings for robot-like appearances. This framework turns empirical uncanny-valley heuristics into a computational basis for algorithmically evaluating and optimizing humanoid robot appearance and behavior.
For humanoid robots that interact closely with people, designers must make appearance and behavior feel familiar and acceptable. Perceived human likeness, social cues, appearance, and behavior shape trust in robots, intention to use them, emotional responses, and interaction quality[1]–[4]. However, increasing human likeness does not always increase affinity. The uncanny valley phenomenon shows that entities that appear human-like but not fully human can sharply reduce affinity or positive impressions[5]. Figure1 illustrates Mori’s uncanny valley curve. Thus, humanoid robot design must consider how to avoid the uncanny valley[6].
Many studies have proposed design guidelines for avoiding the uncanny valley. For example, rather than aiming for a fully human appearance, researchers recommend adopting a stylized, robot-like appearance corresponding to the “first peak” before the uncanny valley[5], [8]–[10]. Findings that cartoon-like or non-realistic faces can elicit more trust than realistic human faces[11], and that excessive human likeness can reduce liking for consumer robots[12], support this guideline. Researchers also emphasize aligning human likeness across appearance, motion, voice, touch, and other modalities to avoid cross-modal mismatches[13], [14]. In addition, previous work highlights the need to align appearance-based expectations about competence and warmth with actual robot behavior[15].
Although these guidelines have empirical support, they do not yet specify design variables mathematically. For example, the guideline of aiming for the “first peak” does not specify where this peak lies on the human-likeness axis. Similarly, the guideline of avoiding cross-modal mismatches does not specify how strictly designers should align modalities. As a result, designers still rely on experience and intuition when searching for appearances and behaviors that avoid uncanniness.
This problem poses an engineering challenge: infer the perceptual mechanisms behind empirical design guidelines and translate them into variables that designers can manipulate. Human-centered design emphasizes systems based on user requirements, ergonomics, and usability[16]. Kansei engineering offers methods for translating users’ emotions and impressions into product design elements[17], [18]. More recently, HRI has increasingly used human-in-the-loop optimization to optimize robots and devices based on human responses and performance measures[19].
Robot design widely uses mathematical models for kinematics, dynamics, control, and optimization[20], [21]. Yet few models can guide the design of appearance and behavior to avoid the uncanny valley. Researchers have also noted the lack of models that predict the uncanny valley curve in advance[22]. Moore proposed a model of the uncanny valley curve based on Bayesian category perception[23], and Ueyama applied this model to robot therapy[24]. However, existing models do not primarily map uncanny-valley design guidelines onto concrete variables that designers can manipulate.
Here, we construct a hierarchical Bayesian generative model that represents affinity toward humanoid robots as posterior-weighted negative category-conditional surprise[25]. The model draws on findings that appearance–motion mismatches and deviations from human norms relate to prediction-error-like responses and negative evaluations[26]–[28]. The model also builds on the Bayesian brain hypothesis, which views perception as probabilistic inference about the latent causes of sensory inputs[29], and on the free-energy principle[30].
This formulation explains category ambiguity and perceptual mismatch, two major hypotheses for the uncanny valley, within a common framework as increases in surprise. Category ambiguity refers to a state in which an observation fits neither the human category nor the robot category sufficiently[31]–[33]. Perceptual mismatch refers to a state in which cues such as appearance, motion, and voice indicate inconsistent levels of human likeness[34]–[38].
This formulation organizes guidelines for reducing uncanniness around four design variables. The first is the distance between observed human likeness \(y\) and the predicted mean \(\mu_R\) for a typical robot, \(|y-\mu_R|\). This variable operationalizes the guideline of aiming for the “first peak” as bringing appearance closer to the predicted mean of the robot category. The second is the difference in human likeness indicated by appearance and motion, \((y_a-y_m)^2\), which represents perceptual mismatch across modalities. The third is prediction uncertainty for the robot category, \(\sigma_R^2\), which represents the breadth of the observer’s belief about robot appearance. The fourth is observational uncertainty, \(\sigma_l^2\), which represents how precisely the observer processes appearance as sensory evidence. This mapping embeds existing design guidelines into model components and provides a basis for deriving underexplored design strategies.
Among these variables, we focus on \(\sigma_R^2\) and \(\sigma_l^2\), which concern observer-side uncertainty, and test their effects on familiarity in a human-subject experiment. Few uncanny valley studies have separately manipulated prediction uncertainty and observational uncertainty within the same experiment. We therefore used morphing stimuli between robot and human images, manipulated prediction uncertainty by blurring the robot image presented as a prior stimulus, and manipulated observational uncertainty by blurring the evaluation image. This design allowed us to test whether the model-predicted uncertainty effects appear in familiarity with humanoid robot appearances and in the shape of the uncanny valley.
This study aims to formulate affinity perception toward humanoid robots as a Bayesian generative model composed of designer-manipulable variables and to systematize design guidelines for reducing uncanniness.
This study makes three contributions. First, we construct a Bayesian generative model that explains category ambiguity and perceptual mismatch as increases in Shannon surprise within a unified framework. This model interprets existing design guidelines, such as approaching the first peak and aligning modalities, as computational components such as \(|y-\mu_R|\) and \((y_a-y_m)^2\). Second, model simulations and a human-subject experiment show that prediction uncertainty \(\sigma_R^2\) and observational uncertainty \(\sigma_l^2\) have distinct effects on the shape of the uncanny valley. In the experiment, increased observational uncertainty attenuated the decrease in familiarity at intermediate human likeness, and low prediction uncertainty increased familiarity for robot-like appearances. However, high prediction uncertainty did not increase familiarity at intermediate human likeness. Third, we systematize design guidelines for reducing uncanniness based on the model components. This framework organizes existing empirical guidelines and provides a theoretical basis for underexplored design strategies, such as adjusting observational precision and guiding attention. The model may also provide a mathematical basis for future algorithms that evaluate and optimize robot appearance and behavior.
We model affinity as posterior-weighted negative category-conditional surprise. We first formulate a unimodal model for appearance alone and then extend it to a multimodal model that incorporates appearance and motion.
We first constructed a unimodal hierarchical Bayesian generative model for appearance. The model assumes that an observer infers the latent human likeness \(x\) of an entity from an appearance observation \(y\) and infers whether the entity belongs to the robot category \(R\) or the human category \(H\) (Fig.2A). The generative model is
\[p(c,x,y)=p(c)p(x|c)p(y|x). \label{eq:unimodal95genmodel}\tag{1}\]
We define the observer’s affinity toward an entity by how well the generative model explains the observation \(y\) after category recognition. Specifically, we define the Shannon surprise for observation \(y\) under category \(c\) as
\[S_c(y) = -\ln p(y|c) \label{eq:category95conditional95surprise}\tag{2}\]
and define its negative as the affinity measure under category \(c\):
\[A_c(y) = -S_c(y) = \ln p(y|c). \label{eq:category95conditional95affinity}\tag{3}\]
Thus, as category \(c\) explains observation \(y\) less well, conditional surprise \(S_c(y)\) increases and the category-specific affinity measure \(A_c(y)\) decreases. We define the final affinity measure by integrating these category-specific affinity measures according to posterior category probabilities based on \(y\).
We assume the category-specific predictive distribution \(p(x|c)=\mathcal{N}(x;\mu_c,\sigma_c^2)\) and the observation process \(p(y|x)=\mathcal{N}(y;x,\sigma_l^2)\). Here, \(\mu_c\) denotes the predicted mean of category \(c\), \(\sigma_c^2\) denotes prediction uncertainty for that category, and \(\sigma_l^2\) denotes observational uncertainty. Marginalizing over the latent variable \(x\) gives the predictive distribution of \(y\) under category \(c\):
\[p(y|c)=\mathcal{N}(y;\mu_c,\sigma_c^2+\sigma_l^2).\]
The affinity measure under category \(c\) is therefore
\[A_c(y) = \log p(y|c) = - \frac{1}{2}\log\{2\pi(\sigma_c^2+\sigma_l^2)\} - \frac{1}{2(\sigma_c^2+\sigma_l^2)} \underbrace{(y-\mu_c)^2}_{\text{distance from the category mean}}. \label{eq:category95affinity}\tag{4}\]
When \(c=R\), the term \((y-\mu_R)^2\) in Eq.@eq:eq:category95affinity represents how far the observed human likeness \(y\) deviates from the observer’s predicted mean \(\mu_R\) for a typical robot category. This term therefore provides a design variable for interpreting the “first peak” as a region close to the predicted mean of the robot category.
We assume that the observer does not assign \(y\) completely to either the robot category \(R\) or the human category \(H\). Instead, the observer integrates category-specific affinity measures according to posterior category probabilities after observation. Letting the prior category probability be \(\pi_c=p(c)\), Bayes’ rule gives the posterior category probability for observation \(y\) as
\[p(c|y) = \frac{ p(c)p(y|c) }{ \sum_{c' \in \{R,H\}} p(c')p(y|c') } = \frac{ \pi_c \exp\{A_c(y)\} }{ \sum_{c' \in \{R,H\}} \pi_{c'} \exp\{A_{c'}(y)\} }. \label{eq:category95posterior}\tag{5}\]
Here, \(A_c(y)=\log p(y|c)\). Equation@eq:eq:category95posterior represents how well the robot and human categories explain observation \(y\).
We define the affinity measure as the category-specific affinity measures weighted by posterior category probabilities:
\[A(y) = \mathbb{E}_{p(c|y)}[A_c(y)] = \sum_{c \in \{R,H\}} p(c|y)A_c(y). \label{eq:affinity95posterior95weighted}\tag{6}\]
In this formulation, when the robot category explains observation \(y\) well, \(p(R|y)\) becomes large and \(A(y)\) mainly reflects \(A_R(y)\). When the human category explains \(y\) well, \(p(H|y)\) becomes large and \(A(y)\) mainly reflects \(A_H(y)\). Near the category boundary, both posterior probabilities take intermediate values, and the model smoothly integrates the category-specific affinity measures.
Figure2B shows the simulated affinity measure \(A(y)\) under representative parameters. The model generated a curve in which affinity first increased with human likeness, then decreased at intermediate human likeness, and increased again near the fully human region. This result supports the category ambiguity hypothesis: affinity decreases when neither the robot nor human category sufficiently explains an intermediate appearance.
The baseline model formulates the observation process as a Gaussian likelihood. Under this assumption, surprise increases quadratically with prediction error. Thus, especially under high prediction precision, the model may overestimate surprise for observations that deviate substantially from category predictions[39]. Following our previous study[39], we therefore also considered an \(\epsilon\)-floor likelihood that adds a small floor to the observation likelihood.
This model replaces only the observation process:
\[p_\epsilon(y|x) = \frac{ \mathcal{N}(y;x,\sigma_l^2)+\epsilon }{ 1+\epsilon }. \label{eq:epsilon95observation95likelihood}\tag{7}\]
After marginalizing over \(x\), the predictive distribution under category \(c\) takes the form of the baseline distribution \(\mathcal{N}(y;\mu_c,\sigma_c^2+\sigma_l^2)\) with the same \(\epsilon\) floor. Thus, when observation \(y\) deviates substantially from the category prediction and the Gaussian component approaches zero, the likelihood approaches \(\epsilon/(1+\epsilon)\), which bounds surprise.
Let \(A_c^\epsilon(y)\) denote the category-specific affinity measure obtained from the \(\epsilon\)-floor likelihood, and let \(p_\epsilon(c|y)\) denote the corresponding posterior category probability. The final affinity measure becomes
\[A^\epsilon(y) = \sum_{c \in \{R,H\}} p_\epsilon(c|y)A_c^\epsilon(y). \label{eq:epsilon95affinity95posterior95weighted}\tag{8}\]
This formulation reduces to the baseline model when \(\epsilon=0\).
We next extended the unimodal model to a multimodal model that incorporates appearance and motion. In the multimodal model, the observer obtains two types of sensory evidence: an appearance observation \(y_a\) and a motion observation \(y_m\). We assume that a common latent variable, the entity’s human likeness \(x\), generates both observations (Fig.3A). The generative model is
\[p(c,x,y_a,y_m) = p(c)p(x|c)p(y_a|x)p(y_m|x). \label{eq:multimodal95genmodel}\tag{9}\]
We assume the category-specific predictive distribution \(p(x|c)=\mathcal{N}(x;\mu_c,\sigma_c^2)\) and the observation processes \(p(y_a|x)=\mathcal{N}(y_a;x,\sigma_a^2)\) and \(p(y_m|x)=\mathcal{N}(y_m;x,\sigma_m^2)\). Here, \(\mu_c\) denotes the predicted mean of category \(c\), \(\sigma_c^2\) denotes prediction uncertainty for that category, \(\sigma_a^2\) denotes observational uncertainty for appearance, and \(\sigma_m^2\) denotes observational uncertainty for motion. Under these assumptions, the affinity measure under category \(c\) is
\[\begin{align} A_c(y_a,y_m) = &-\log(2\pi) - \frac{1}{2}\log D_c \\ &- \frac{1}{2D_c} \left\{ \underbrace{(y_a-y_m)^2}_{\substack{\text{mismatch}\\\text{across modalities}}} \sigma_c^2 + \underbrace{(y_a-\mu_c)^2}_{\substack{\text{distance from}\\\text{the category mean (appearance)}}} \sigma_m^2 + \underbrace{(y_m-\mu_c)^2}_{\substack{\text{distance from}\\\text{the category mean (motion)}}} \sigma_a^2 \right\}. \end{align} \label{eq:multimodal95affinity}\tag{10}\]
where
\[D_c = \sigma_a^2\sigma_m^2 + \sigma_a^2\sigma_c^2 + \sigma_m^2\sigma_c^2.\]
The final affinity measure weights category-specific affinity measures by posterior category probabilities after observation:
\[A(y_a,y_m) = \sum_{c \in \{R,H\}} p(c|y_a,y_m)A_c(y_a,y_m). \label{eq:multimodal95affinity95posterior95weighted}\tag{11}\]
Here, \(p(c|y_a,y_m)\) extends \(p(c|y)\) in the unimodal model to the multivariate case using the category-specific affinity measure \(A_c(y_a,y_m)\).
The term \((y_a-y_m)^2\) in Eq.@eq:eq:multimodal95affinity represents inconsistency in human likeness between appearance and motion. Thus, when the appearance is human-like but the motion is robot-like, or vice versa, this term increases and the category-specific affinity measure \(A_c(y_a,y_m)\) decreases.
Figure3B shows a heatmap of the affinity measure \(A(y_a,y_m)\) in the multimodal model. Affinity was high when appearance human likeness \(y_a\) and motion human likeness \(y_m\) were both robot-like or both human-like. In contrast, affinity decreased when appearance and motion differed substantially in human likeness.
Figure3C shows a cross-section along the diagonal line \(y_a=y_m\). Because appearance and motion match in human likeness along this section, the affinity reduction arises mainly from category ambiguity. Figure3D shows a cross-section along the line \(y_a+y_m=10\). This condition holds the average human likeness of appearance and motion constant while varying only their difference. Affinity decreased as the inconsistency between appearance and motion increased.
Together, these results show that the multimodal model predicts a valley from category ambiguity when appearance and motion are consistent, and an affinity reduction from perceptual mismatch when they are inconsistent.
Among the four design variables, we next examined how prediction uncertainty for the robot category, \(\sigma_R^2\), and observational uncertainty, \(\sigma_l^2\), alter the affinity curve. Figure4 shows affinity curves obtained by varying \(\sigma_R^2\) and \(\sigma_l^2\) in the unimodal baseline model. These simulations yielded four theoretical predictions for the human-subject experiment.
TH1-1: For robot-like appearances, lower prediction uncertainty increases affinity. Varying prediction uncertainty for the robot category, \(\sigma_R^2\), mainly changed the robot-side curve (Fig.4A). When \(\sigma_R^2\) was small, appearances close to \(\mu_R\) produced high affinity. Thus, for robot-like appearances, lower prediction uncertainty for the robot category should increase affinity.
TH1-2: For intermediate human likeness, higher prediction uncertainty increases affinity. In contrast, when \(\sigma_R^2\) was small, affinity sharply decreased as appearance moved away from \(\mu_R\). When \(\sigma_R^2\) was large, the robot-side curve became flatter, and affinity decreased more gradually for appearances slightly distant from the robot category (Fig.4A). Thus, for appearances with intermediate human likeness, higher prediction uncertainty for the robot category should increase affinity.
TH2: Higher observational uncertainty increases affinity. Varying observational uncertainty, \(\sigma_l^2\), changed the sharpness of the overall affinity curve (Fig.4B). When \(\sigma_l^2\) was small, affinity substantially decreased near the category boundary, producing a deep valley. When \(\sigma_l^2\) was large, this decrease weakened and the valley became shallower. Thus, higher observational uncertainty should increase affinity.
TH3: Higher observational uncertainty attenuates the decrease in affinity with respect to human likeness. Under large \(\sigma_l^2\), the affinity curve became smoother, and the decrease in affinity at intermediate human likeness weakened (Fig.4B). Thus, higher observational uncertainty should attenuate the decrease in affinity with respect to human likeness.
These simulation results yielded theoretical hypotheses on prediction uncertainty (TH1-1 and TH1-2) and observational uncertainty (TH2 and TH3). These THs concern the model’s affinity measure \(A(y)\). In the next section, we map these model-based predictions onto experimentally manipulable stimulus conditions and test them as experimental hypotheses.
We next tested the theoretical hypotheses from the unimodal baseline model in an experiment that manipulated appearance human likeness. We mapped prediction uncertainty \(\sigma_R^2\) and observational uncertainty \(\sigma_l^2\) onto image-blur conditions. Specifically, we used the blur level of the prior robot stimulus to manipulate prediction uncertainty \(\sigma_R^2\) for the robot category, and the blur level of the evaluation stimulus to manipulate observational uncertainty \(\sigma_l^2\) (Fig.5). Because the prior robot image provided category-level information about the robot endpoint, blurring this image reduced the precision of the robot-category prior rather than the sensory precision of the subsequent evaluation stimulus. Participants rated the familiarity of each evaluation stimulus. Thirty-three adults participated. We determined the sample size based on feasibility and on the within-participant design used in previous uncanny valley studies. We excluded no participants from the analysis.
Table1 summarizes the theoretical hypotheses (THs), experimental manipulations, and experimental hypotheses (EHs). The THs concern the model’s affinity measure \(A(y)\), whereas the EHs test these predictions as familiarity ratings.
| Symbol | Experimental hypothesis | Experimental manipulation | Corresponding theoretical hypothesis |
|---|---|---|---|
| EH1-1 | For robot-like appearances, weak blur of the prior stimulus yields higher familiarity ratings than strong blur. | Blur applied to the prior robot image. The low-blur condition corresponds to low prediction uncertainty, and the high-blur condition corresponds to high prediction uncertainty. | For robot-like appearances, lower prediction uncertainty for the robot category increases affinity (TH1-1). |
| EH1-2 | For appearances with intermediate human likeness, strong blur of the prior stimulus yields higher familiarity ratings than weak blur. | Blur applied to the prior robot image. The high-blur condition corresponds to high prediction uncertainty. | For appearances with intermediate human likeness, higher prediction uncertainty for the robot category increases affinity (TH1-2). |
| EH2 | Strong blur of the evaluation stimulus yields higher familiarity ratings than weak blur. | Blur applied to the evaluation stimulus. The low-blur condition corresponds to low observational uncertainty, and the high-blur condition corresponds to high observational uncertainty. | Higher observational uncertainty increases affinity (TH2). |
| EH3 | Strong blur of the evaluation stimulus attenuates the decrease in familiarity ratings at intermediate human likeness compared with weak blur. | Interaction between blur applied to the evaluation stimulus and appearance human likeness. | Higher observational uncertainty attenuates the decrease in affinity with respect to human likeness (TH3). |
2pt
Figure6 shows the familiarity ratings. Figure6A compares the low- and high-blur conditions for the prior robot stimulus, corresponding to the manipulation of prediction uncertainty. Figure6B compares the low- and high-blur conditions for the evaluation stimulus, corresponding to the manipulation of observational uncertainty.
Appearance human likeness significantly affected familiarity ratings (\(F=43.700\), \(P < 0.001\)). Familiarity varied nonlinearly with appearance human likeness and decreased for stimuli with intermediate human likeness. This result supports the basic premise of the uncanny valley: affinity decreases near the category boundary between robots and humans.
For the prediction uncertainty manipulation, participants gave higher familiarity ratings when the prior stimulus had weak blur than when it had strong blur (Fig.6A; \(F=4.486\), \(P = 0.034\)). However, blur applied to the prior stimulus did not significantly interact with appearance human likeness (\(F=1.036\), \(P = 0.404\)). Thus, the results supported EH1-1, which predicted higher familiarity for robot-like appearances in the low-blur prior-stimulus condition, but did not support EH1-2, which predicted higher familiarity for appearances with intermediate human likeness in the high-blur prior-stimulus condition.
For the observational uncertainty manipulation, participants gave higher familiarity ratings when the evaluation stimulus had strong blur than when it had weak blur (Fig.6B; \(F=53.187\), \(P < 0.001\)). This result supports EH2, which predicted that higher observational uncertainty would increase familiarity ratings. Blur applied to the evaluation stimulus also significantly interacted with appearance human likeness (\(F=2.108\), \(P = 0.040\)). Simple main-effect tests showed that the high-blur evaluation-stimulus condition significantly increased familiarity ratings for stimuli with human likeness levels from \(h=2\) to \(h=6\) (\(h=2\): \(F=6.031\), \(P = 0.015\); \(h=3\): \(F=20.394\), \(P < 0.001\); \(h=4\): \(F=20.933\), \(P < 0.001\); \(h=5\): \(F=9.493\), \(P = 0.003\); \(h=6\): \(F=8.989\), \(P = 0.003\)). We found no significant differences for \(h=1\), \(h=7\), or \(h=8\).
Taken together, the observational uncertainty manipulation shown in Fig.6B supported EH2 and EH3. The prediction uncertainty manipulation shown in Fig.6A supported EH1-1 but not EH1-2. A supplementary \(\epsilon\)-floor model analysis showed that this model attenuated, but did not eliminate, the predicted reversal of prediction-uncertainty effects (fig. 7). We return to this discrepancy in the Discussion.
Our experiment supported three of the four hypotheses, except EH1-2. Increased observational uncertainty attenuated the decrease in familiarity at intermediate human likeness, and low prediction uncertainty increased familiarity for robot-like appearances.
Below, we organize design guidelines for reducing uncanniness according to the four design variables introduced in the Introduction. In our model, guidelines discussed in previous studies, such as approaching the first peak, aligning modalities, and avoiding expectation mismatch, correspond to operations on distinct computational components. Table2 summarizes the design variables, design guidelines, and model-based interpretations for reducing uncanniness.
| Design variable (model term) | Design guideline | Examples | Supporting studies | ||||
|---|---|---|---|---|---|---|---|
| Distance from robot-category mean (\(|y-\mu_R|\)) | Design appearances close to what observers expect as “robot-like.” | Pepper, NAO[40], [41] | First-peak guidelines[5], [8]–[10]; non-realistic and stylized appearances[11], [12] | ||||
| Context- or user-specific robot-category mean (\(|y-\mu_R|\)) | Make faces, expressions, and voices adjustable according to context and user. | Furhat[42] | Adjustable robot appearance and expressions; context-dependent appearance design[14] | ||||
| Mismatch across modalities (\(|y_a-y_m|\)) | Align the level of human likeness across appearance, motion, voice, and touch. | Geminoid HI-2[43] | Cross-modal consistency[13], [14]; negative responses from expectation mismatch[15] | ||||
| Prediction uncertainty (\(\mu_R\), \(\sigma_R^2\)) | Promote repeated contact through routines, care, learning tasks, and personalized responses. | PARO, Moxie[44], [45] | Reduced uncanniness through repeated exposure[43], [46]; sustained interaction in care contexts[44] | ||||
| Observational uncertainty (\(\sigma_l^2\)) | Use color, surface texture, and simplified shapes to guide attention away from fine details. | LOVOT, PARO[44], [47] | Color use and liking[48]; texture and shape perception[49]; baby-schema features, trust, and cuteness[50], [51] | ||||
| Local-feature weighting (attention) | Distribute attention toward decorations and bodily features. | Hats, sunglasses, face paint, etc.[52] | Effects of shape and decoration on gaze allocation and user experience[52] |
3pt
The first design variable is the distance between observed human likeness \(y\) and the human likeness \(\mu_R\) that an observer predicts for a typical robot. The guideline of aiming for the first peak of the uncanny valley curve[5], [8]–[10] helps designers avoid full human likeness, but it does not specify which level of human likeness corresponds to that first peak. Our model interprets this guideline as bringing the appearance closer to the predicted mean \(\mu_R\) of the robot category. In other words, it reframes the target “first peak” as the observer’s representation of a typical robot appearance. Designing an appearance close to this representation should reduce uncanniness.
Observers’ predictions about the robot category may also vary across users and contexts. Therefore, making faces, expressions, and voices adjustable to context and user can bring the appearance \(y\) closer to that observer’s prior belief. Social robots such as Furhat, which can change their face, expressions, gaze, and voice, provide an example of this design strategy[42].
The second design variable is inconsistency in human likeness across appearance, motion, voice, touch, and other modalities. In our multimodal model, the term \((y_a-y_m)^2\) represents mismatch between appearance and motion, and the affinity measure decreases as this difference increases. Thus, even if designers make the appearance human-like, mechanical motion or voice can create perceptual mismatch and reduce affinity. Studies using Geminoid HI-2 have increased uncanniness by combining a human-like appearance with incongruent voice or jerky motion[43]. Findings that uncanniness increases when visual and vocal naturalness mismatch[37], as well as design suggestions emphasizing consistency across appearance, motion, sound, and touch[13], [14], align with our model. In addition, negative responses caused by mismatch between appearance-based expectations and actual competence or warmth[15] can be interpreted as inconsistency between predicted human likeness and observed behavior.
The third design variable is prediction uncertainty for the robot category. Prediction uncertainty represents how narrow or broad an observer’s prior belief about robot appearance is. Our results suggest that, for robot-like appearances, clearer predictions about the robot category can increase familiarity.
To strengthen beliefs about a robot, designers can create mechanisms that encourage repeated contact. Previous studies show that repeated exposure reduces uncanniness toward robots[43], [46]. For example, robots such as PARO and Moxie encourage sustained contact through care, learning tasks, and everyday conversation. Such robots can draw users’ typical robot representations toward the specific robot and reduce initial prediction errors[44], [45].
The fourth design variable is observational uncertainty. In our experiment, increasing observational uncertainty by blurring the evaluation stimulus attenuated the decrease in familiarity at intermediate human likeness. In actual robot design, designers need not visually blur the robot itself. Instead, they can use surface texture, facial shape, color, and decoration to prevent excessive attention to fine details and thereby suppress affinity reductions.
For example, vivid colors can attract visual attention and may reduce liking for robots[48]. Thus, white or pale colors may reduce excessive attention to fine details and mitigate uncanniness. Surface texture and reflectance also affect the precision of shape perception[49], so matte textures may prevent fine geometric details from standing out. In addition, baby-schema features increase cuteness and trust[50], [51]. Rounded and soft appearances, such as those of LOVOT and PARO, can direct attention away from realistic facial details and toward tactile or animal-like familiarity[44], [47]. Moreover, because shape and decoration influence gaze allocation and user experience[52], hats, sunglasses, face paint, and similar features may divert attention from local features, such as the eyes and mouth, that often trigger uncanniness.
EH1-1 was supported: low prediction uncertainty increased familiarity for robot-like appearances. In contrast, EH1-2 was not supported: high prediction uncertainty did not increase familiarity for appearances with intermediate human likeness. This result suggests that the baseline model may have overestimated the affinity reduction for stimuli that deviate substantially from category predictions.
In the baseline model, low prediction uncertainty for the robot category makes the likelihood of appearances far from the robot-category mean drop sharply. However, actual observers may not treat stimuli that deviate from category predictions as completely implausible. We therefore examined an \(\epsilon\)-floor model, which prevents likelihoods for highly deviant observations from becoming extremely small (fig. 7). This model still predicted lower affinity under low prediction uncertainty in the intermediate human-likeness region, but the difference between low and high prediction uncertainty became smaller than in the baseline model. Thus, the \(\epsilon\)-floor model does not fully explain the lack of support for EH1-2, but it suggests that the baseline model may have overestimated the effect of prediction uncertainty and that the reversal effect may be difficult to detect experimentally.
Applying our model to real robot design requires three extensions. First, future work should test the model in interactions with real robots and in real-world use contexts. This study used image stimuli and focused on appearance-based affinity evaluations. In real robot interactions, however, appearance, motion, voice, embodiment, and the surrounding environment act together[53], [54]. Future studies should therefore test whether the model-derived design guidelines improve affinity and reduce uncanniness not only with image stimuli but also in real interactions and use contexts.
Second, the model should extend to personalized model fitting. Uncanniness and affinity may vary with observers’ attitudes toward robots[55], [56], interest in technology[55], and cultural background[57]. The model may express these differences as parameters. For example, observers from cultures familiar with robots or observers with strong interest in technology may have lower prediction uncertainty \(\sigma_R^2\) for the robot category. By estimating parameters such as \(\mu_R\) and \(\sigma_R^2\) for each observer, future work could connect the model to personalized robot design[58], [59] and human-in-the-loop design optimization[19].
Third, the model should incorporate context dependence. Expected appearance and behavior for humanoid robots vary with use context[60]. In contexts where people already have clear images of humanoid robot use, such as reception, guidance, or care, observers may hold relatively clear robot-category expectations, resulting in low \(\sigma_R^2\). In contrast, in unfamiliar applications, such as a humanoid robot guiding mourners at a funeral ceremony, observers may not know what appearance or behavior to expect, resulting in high \(\sigma_R^2\). The predicted mean of the robot category, \(\mu_R\), may also shift by context. For humanoid robots working in factories or logistics warehouses, people may expect mechanical and efficient motion as typical robot-like behavior, placing \(\mu_R\) at a lower value. For reception or customer-service robots, people may expect more human-like behaviors, such as gaze shifts and nodding, placing \(\mu_R\) closer to the human side. Future work should examine how designers should adjust model parameters according to use context.
This study constructed a hierarchical Bayesian generative model that represents affinity toward humanoid robots as posterior-weighted negative category-conditional surprise and explains the uncanny valley as increased surprise arising from category ambiguity and perceptual mismatch. This formulation organizes design issues such as approaching the first peak, aligning modalities, prediction uncertainty, and observational uncertainty as model variables: \(|y-\mu_R|\), \((y_a-y_m)^2\), \(\sigma_R^2\), and \(\sigma_l^2\).
The significance of this study lies in reframing empirical design guidelines for avoiding the uncanny valley as operations on model variables that increase the affinity measure, namely negative category-conditional surprise. The model also provides a theoretical framework for deriving underexplored design strategies, such as adjusting observational precision and guiding attention. In the future, extensions to real robot interaction, personalized parameter estimation, and context-dependent evaluation functions may connect this model to design algorithms that automatically optimize humanoid robot appearance and behavior to avoid uncanniness.
This experiment tested whether prediction uncertainty and observational uncertainty, as defined in the model, influence familiarity ratings for humanoid robot appearances. We used a within-participant design with three factors: appearance human likeness, prediction uncertainty, and observational uncertainty. The primary outcome was the z-scored familiarity rating for each evaluation stimulus.
Thirty-three adults aged 18–39 years participated in the experiment (16 women and 17 men). We recruited participants who had no visual diseases or disabilities and no strong psychological resistance to humanoid robots. We recruited participants through an online participant recruitment platform[61].
The Research Ethics Committee of the Graduate School of Engineering, The University of Tokyo, approved this study (approval number: KE24-62). We conducted the experiment from November 18 to December 5, 2024. Before participation, we explained the study content, data to be collected, handling of personal information, voluntary participation, and the right to withdraw during the experiment without disadvantage. All participants provided written informed consent.
In the unimodal experiment, we used static morphing stimuli between robot and human face images to examine affinity evaluations for humanoid robot appearance. We selected four robot images from the Anthropomorphic roBOT Database (ABOT Database)[7]: Roboy, Lgus, JD Humanoid, and Robina. We created four human face images using Face Generator, a face-image generation service provided by Generated Photos[62]: Japanese male, Japanese female, U.S. male, and U.S. female.
We morphed each robot image and human face image using Abrosoft FantaMorph 5[63]. We defined the appearance human-likeness level as \(h\), set the robot image to \(h=0\) and the human image to \(h=9\), and created a 10-level morphing sequence. In the experiment, we used the eight intermediate levels, \(h=1\) to \(h=8\), as evaluation stimuli, excluding the robot and human endpoints.
We manipulated prediction uncertainty by changing the blur level of the prior robot image presented before the evaluation stimulus. We manipulated observational uncertainty by changing the blur level of the morphing image used as the evaluation stimulus. For both manipulations, we used low- and high-blur conditions.
We applied blur using the cv2.GaussianBlur function in OpenCV. In the low-blur condition, we set the kernel size to \(3 \times 3\) and the standard deviation to \(-1\). In
the high-blur condition, we set the kernel size to \(81 \times 81\) and the standard deviation to \(-1\). We used the same blur settings for the prior robot image and the evaluation
stimulus.
We used Face Generator by Generated Photos to create four synthetic human face images used as source images for the morphing stimuli. The images corresponded to the following stimulus categories: Japanese male, Japanese female, U.S. male, and U.S. female. The exact prompts or interface settings used to generate these images were not retained at the time of stimulus creation. To document the stimuli used in the experiment, the generated source images, derived morphing stimuli are provided in the OSF repository. These images were used only as experimental stimuli and examples of stimuli, not as evidence or data generated by the proposed model.
AI-assisted tools were also used for limited support in manuscript editing, presentation, and code generation. Specifically, ChatGPT 5.5 was used for language refinement, wording suggestions, and formatting support during manuscript preparation, and Codex 5.5 was used to assist with generating and revising simulation, analysis, and figure-generation code. The authors reviewed and edited all AI-assisted outputs and verified the accuracy of the manuscript, citations, code, data analysis, and conclusions.
We conducted the experiment in a laboratory at the Hongo Campus of The University of Tokyo. We presented stimuli and collected responses using a web application implemented in React.js and displayed in full-screen mode on a Microsoft Surface Laptop 4. Participants sat approximately 60 cm from the screen and wore earmuffs during the experiment. We used automatic voice instructions created with VOICEVOX.
The experiment used a within-participant design that manipulated appearance human likeness, prediction uncertainty, and observational uncertainty. Appearance human likeness had eight levels (\(h=1,\ldots,8\)), prediction uncertainty had two levels (low/high blur of the prior robot image), and observational uncertainty had two levels (low/high blur of the evaluation stimulus). Thus, the main session comprised \(8 \times 2 \times 2 = 32\) conditions.
Each participant completed three practice trials followed by 32 main trials. During the main trials, participants took a 1-min break after every 11 trials. We randomized the stimulus order for each participant. We counterbalanced the left–right positions of the human and robot images presented as prior stimuli. We also randomized the order and left–right arrangement of the subjective rating items on each trial.
In each trial, participants first viewed the human and robot images as prior stimuli for 7 s. Then, they viewed the morphing image used as the evaluation stimulus. Participants could not respond for 7 s after the evaluation image appeared, during which they observed the image. They then rated the familiarity of the presented image using a 7-point semantic differential scale. We treated the “familiar/unfamiliar” rating as the behavioral measure corresponding to the model’s affinity measure.
We performed statistical analyses using R version 4.4.2. To adjust for individual differences in response tendencies, we z-scored familiarity ratings within each participant. We used the Shapiro–Wilk test to assess normality in each condition. Because some conditions violated normality, we applied an aligned rank transform (ART) to the z-scored ratings and then conducted a three-way analysis of variance.
The fixed factors were appearance human likeness (eight levels), prediction uncertainty (low/high blur of the prior stimulus), and observational uncertainty (low/high blur of the evaluation stimulus). The main tests targeted the main effects of appearance human likeness, prediction uncertainty, and observational uncertainty, as well as their interactions. When the interaction between observational uncertainty and appearance human likeness reached significance, we tested the simple main effect of observational uncertainty at each level of appearance human likeness. Error bars in the figures represent 95% confidence intervals, and we set the significance level at \(P < 0.05\).
S.H. was supported by The University of Tokyo project “Advanced AI Talent Development to Lead the Next-Generation Intelligent Society (BOOST NAIS),” funded by the Japan Science and Technology Agency (JST) through the Broadening Opportunities for Outstanding Young Researchers and Doctoral Students in Strategic Areas (BOOST) program.
S.H. conceived the study. S.H., R.S., and H.Y. developed the mathematical model and interpreted the model. S.H. introduced and analyzed the \(\epsilon\)-floor model. S.H. and R.S. wrote the simulation code and analyzed the simulation results. S.H. prepared the figures and graphical model. S.H. and R.S. designed the human-subject experiment. R.S. created the morphing and blurred stimuli, recruited participants, conducted the experiment, and collected the data. S.H. implemented the experimental application and stimulus presentation system. S.H. and R.S. designed the statistical analysis. S.H. performed the statistical analysis. S.H. and R.S. interpreted the experimental results. S.H. wrote the original draft. S.H., R.S., and H.Y. revised the manuscript, figures, and structure. H.Y. supervised the project and provided critical intellectual input.
The authors declare that they have no competing interests.
The authors used AI-assisted technologies to support manuscript preparation, code generation, and creation of synthetic human face stimuli. Details are provided in the Methods section. The authors reviewed all AI-assisted outputs and take full responsibility for the work.
Anonymized data, analysis and simulation code, figure-generation code, and stimulus images are available at OSF: https://osf.io/qc8za/?view_only=db83bf89ad534ae485f02b03b791ba80. Stimulus images are shared in accordance with the original providers’ usage conditions. No physical materials were generated in this work.
| Source | Sum of Squares | df | Mean Square | F value | P value |
|---|---|---|---|---|---|
| Blur (pre) | 426972.307 | 1 | 426972.307 | 4.486 | 0.034 |
| Blur (eval) | 4836813.470 | 1 | 4836813.470 | 53.187 | \(<0.001\) |
| Human likeness | 22561394.636 | 7 | 3223056.377 | 43.700 | \(<0.001\) |
| Blur (pre) x Blur (eval) | 117012.742 | 1 | 117012.742 | 1.225 | 0.269 |
| Blur (pre) x Human likeness | 688568.030 | 7 | 98366.861 | 1.036 | 0.404 |
| Blur (eval) x Human likeness | 1390881.303 | 7 | 198697.329 | 2.108 | 0.040 |
| Blur (pre) x Blur (eval) x Human likeness | 219810.303 | 7 | 31401.472 | 0.329 | 0.941 |
3pt
| Human likeness | F value | P value |
|---|---|---|
| 1 | 2.538 | 0.114 |
| 2 | 6.031 | 0.015 |
| 3 | 20.394 | \(<0.001\) |
| 4 | 20.933 | \(<0.001\) |
| 5 | 9.493 | 0.003 |
| 6 | 8.989 | 0.003 |
| 7 | 3.196 | 0.076 |
| 8 | 0.225 | 0.636 |
8pt