September 14, 2025
While Virtual Reality (VR) systems have become increasingly immersive, they still rely predominantly on visual input, which can constrain perceptual performance when visual information is limited. Incorporating additional sensory modalities, such as sound and scent, offers a promising strategy to enhance user experience and overcome these limitations. This paper investigates the contribution of auditory and olfactory cues in supporting perception within the portal metaphor, a VR technique that reveals remote environments through narrow, visually constrained transitions. We conducted a user study in which participants identified target scenes by selecting the correct portal among alternatives under varying sensory conditions. The results demonstrate that integrating visual, auditory, and olfactory cues significantly improved both recognition accuracy and response time. These findings highlight the potential of multisensory integration to compensate for visual constraints in VR and emphasize the value of incorporating sound and scent to enhance perception, immersion, and interaction within future VR system designs.
<ccs2012> <concept> <concept_id>10003120.10003121.10003125</concept_id> <concept_desc>Human-centered computing Interaction devices</concept_desc> <concept_significance>500</concept_significance> </concept> <concept> <concept_id>10003120.10003121.10011748</concept_id> <concept_desc>Human-centered computing Empirical studies in HCI</concept_desc> <concept_significance>500</concept_significance> </concept> </ccs2012>
The portal metaphor in Virtual Reality (VR) presents a powerful technique for interaction and navigation by visually connecting spatially separated virtual regions. Through portals, users can observe and even traverse distant locations, enabling efficient access to remote scenes without the need for continuous physical or virtual travel [1], [2]. This approach is particularly valuable in confined or complex VR environments where physical space is limited. However, portals are often constrained in size, offering only a narrow field of view (FoV) as shown in Figure [fig:teaser]. Such spatial restrictions limit users’ ability to perceive and interpret the context of remote scenes before engaging with them directly, thereby reducing spatial awareness and increasing cognitive effort [3]–[6].
To address these challenges and examine how cross-modal priming can support scene recognition when visual input is restricted, we explore multisensory portals, portals augmented with auditory and olfactory cues. Although most VR systems primarily rely on visual and auditory modalities, real-world perception is inherently multisensory. Human cognition is shaped by the integration of sensory inputs, including vision, hearing, and smell, which operate in conjunction to build a coherent understanding of the environment [7].
While many VR systems support visual and auditory channels, prior research has emphasized the need for additional sensory modalities to better simulate real-world experiences [8], [9]. In particular, olfactory cues have demonstrated unique benefits: they capture user attention [10], evoke affective and contextual memory [11], and enhance autobiographical memory retrieval when combined with visual stimuli [12]. Similarly, auditory cues provide temporal and spatial signals that help direct user attention and support environmental understanding [13]–[15].
This work investigates whether augmenting portals with olfactory and auditory cues can enhance users’ perception of scenes beyond the portal in VR through cross-modal priming, particularly under conditions of constrained visual information. While prior research has largely employed multisensory integration to improve realism and immersion [13]–[15], we reframe the portal as a perceptual bottleneck and examine whether sensory augmentation can expand its informational utility. We conducted a formal user study to evaluate how the integration of sensory cues influences spatial recognition and object identification through portals. In our study, participants selected the correct portal from four alternatives based on a brief textual description of a remote scene. The results show that participants more accurately and efficiently identified target scenes when portals were enhanced with multisensory cues, demonstrating the potential of sensory augmentation to mitigate visual limitations and improve perceptual performance in constrained VR interfaces.
Although many VR systems still focus on providing users with an immersive experience through visual and auditory stimuli, growing studies highlight the importance of incorporating additional senses like olfaction and haptics in VR experience as multisensory cues can enhance user performance in VR tasks [16]–[18]. Moreover, multisensory cues enhance users’ experiences by fostering a greater sense of naturalness, immersion, and realism [13]–[15]. Previous research has shown that increasing the variety of sensory stimuli contributes to a more realistic virtual experience, which in turn leads to significantly higher user satisfaction [19], [20].
While many earlier works explore tactile cues, olfactory cues have also been explored to enhance the sense of immersion and the perceived presence of virtual objects in VR [21], [22]. Olfactory cues also play a crucial role in enhancing autobiographical memory recall and emotional responses [12], and have also been shown to improve cognitive retention during language learning [23]. Given these benefits, previous studies have proposed VR applications that incorporate olfactory cues for the treatment of post-traumatic stress disorder (PTSD) [24] and cognitive therapy for individuals with autism spectrum disorder (ASD) [25]. Additionally, the combined use of olfactory and tactile cues has been found to effectively elicit specific emotional responses, such as trust or disgust [26].
Previous studies have shown that multisensory cues help improve users’ task performance, and that smell in particular aids memory recall and emotional responses [27]–[30]. There have also been studies showing that smell influences user behavior [31]. Our study uses olfactory cues as one of the multisensory cues and focuses on whether participants can recognize a space or object and select the corresponding button. This is similar to the study by Persky and Dolwick [22], which found that olfactory cues influence participants’ food choices, but our study differs in that we provide olfactory cues to help participants select the corresponding object.
Researchers have developed various olfactory devices that release scents through methods such as vapor diffusion [13], [32], [33], as well as by heating solid [34] or liquid materials [35]. The olfactory design space includes four key aspects: chemical, emotional, spatial, and temporal [36], which correspond to scent types, users’ emotional reactions, the spatial origin of the scent, and how long the scent is perceived. In everyday environments, these aspects help individuals navigate spaces and interpret scent-related information such as type, strength, and blends [29], [34]. Recreating such olfactory dynamics in VR, however, poses significant challenges due to technological limitations, particularly in managing scent intensity and offering a broad scent range. Nonetheless, recognizing the value of olfactory input is essential for improving the quality of VR experiences [37], [38]. In this study, we designed an olfactory device that is attachable to the bottom of the head-mounted display (HMD) based on a previous approach [33]. This allows scents to be delivered instantly and ventilated within the VR environment.
In VR, users typically navigate virtual environments through physical movements, but these are limited by the real-world physical space. To address these constraints, a range of virtual locomotion techniques have been developed, including steering, teleportation, and portal [39]–[41]. Among them, portal facilitates navigation of a broader virtual space by connecting disparate virtual environments, even within confined spaces, thereby enabling users to experience transitioning between virtual environments while traversing a confined space [41].
Research on portal movement is being conducted with a focus on operability, motion sickness prevention, and maintaining immersion. However, in complex environments or those with multiple users, portal-based movement can cause disorientation or confusion regarding spatial awareness [3], [42]. In addressing this issue, various methods have been proposed and studied, including minimizing physical movement, limiting the FoV, or providing visual effects [41]–[44].
The impact of portals on users is also subject to variation depending on the portal’s dimensions. However, smaller portals may impose limitations on user immersion due to their constrained FoV, necessitating physical movement to verify the space beyond the portal [3], [43], [45]. Despite their comparatively restricted field of view, small-sized portals demonstrated a higher degree of efficacy in mitigating directional confusion when compared to their larger counterparts. Conversely, the use of larger-sized portals has been demonstrated to enhance immersion and improve spatial awareness [41], [46], [47]. However, extant studies have indicated that users experience an increase in cognitive load due to the increased amount of information they must process concurrently [2]. This paper addresses the challenge of the limited FoV inherent in small-sized portals by integrating auditory and olfactory stimuli into the navigation process to support object and location identification.
Our user study investigates whether augmenting a VR portal with auditory and olfactory cues enhances users’ understanding of a scene viewed through the portal. We compare a visual-only condition against multisensory conditions to explore the potential effectiveness of multimodal cues in improving the portal experience under limited FoV constraints. This study addresses the following research questions (RQs):
How do multisensory cues help users grasp the context of a scene that is only partially visible through the portal?
What are the specific contributions of auditory and olfactory cues to portal-based interaction?
Our study employs a within-subjects design based on four sensory cue conditions augmented in the portal. When a participant approaches within 1 meter of the portal, they can perceive additional auditory and/or olfactory cues in addition to the visual cue, depending on the assigned condition. The conditions are as follows:
Visual only (V): This is the baseline condition where participants only see the visual scene inside each portal. No additional auditory or olfactory cues are provided.
Visual plus olfactory (VO): In this condition, participants can see the scene inside the portal and perceive a scent corresponding to the scene when they are within 1 meter of the portal.
Visual plus auditory (VA): In this condition, participants are presented with both visual and auditory cues. As they approach the portal, they hear sounds associated with the scene inside.
Visual plus auditory plus olfactory (VAO): This is the fully multisensory condition, where participants receive visual, auditory, and olfactory cues simultaneously.
In this study, we assumed that the portal presents a scene with a limited FoV, following the approach adopted in prior work [1], [2]. The portal was configured as a square measuring 1 meter by 1 meter—twice the size used in Ablett et al.’s design [2] and comparable in scale to the circular portal used by Han et al. [1]. This dimension was selected to provide a constrained yet sufficiently detailed view, balancing visual limitation, essential for evaluating multisensory cue effects, while ensuring that the content within the portal remains both perceptible and semantically interpretable to participants.
We designed four visually distinct virtual environments, Desert, North Pole, City, and Cafe, as illustrated in Figure 2. They were selected for their easily recognizable visual characteristics. Each environment could include various virtual objects. Among them, some objects presented their corresponding auditory or olfactory cues, called olfactory and auditory elements in this work. The auditory elements consisted of four types: a dog vocalization, a bird vocalization, a person coughing, and music, represented in the scene by a virtual dog, bird, human, and speaker, respectively. The olfactory elements consisted of virtual objects such as coffee, orange, flowers, and pizza, with each contributing a distinct scent1. Participants were able to see virtual environments and objects through portals. Participants could move around the VR scene to view the portals from different angles and clearly identify these objects within the portals. This could be important in the sole V condition, where no auditory or olfactory cues were available.
In each task, four portals were situated within a virtual room, as shown in Figure 1A. Participants began exploring the portals from the center of the room (i.e., the green circle in Figure 1A). Two randomly selected environments out of the four were displayed through two portals each. Each auditory and olfactory element was placed in a separate portal along with various virtual objects.
Depending on the sensory condition, participants could perceive a scent and/or hear a sound when they were within 1 meter of a portal, or they might not receive any additional cues other than a visual cue. For each trial, participants received the following prompt: “Identify the [virtual environment] where the [object 1] is heard and the [object 2] is present.” While both objects could be verified visually, object 1 was additionally represented through auditory cues and object 2 through olfactory cues, depending on the multisensory condition. They were instructed to select the portal that matched the description in the prompt, then either proceed to the next task or complete the study.
The experimental design aimed for simplicity and intuitiveness, with users selecting one of four portals. The decision was predicated on two considerations. First, the provision of the same sensory information to all four portals was implemented to prevent the mixing of types and to facilitate a clear understanding of the effects of sensory information. Secondly, the provision of information regarding various objects for each portal was intended to facilitate user comprehension of the inquiries and enable intuitive selections. The objective of this study was to ensure that the research findings were both clear and intuitive in terms of user perception, with the emphasis on sensory information.
Quantitative measures are the participants’ response accuracy for portal selection, task completion time, and confidence rating. Two subjective indicators are the NASA Task Load Index (NASA-TLX) [48] and the Simulator Sickness Questionnaire (SSQ) [49].
Accuracy Rate: The proportion of correct responses out of 10 trials per sensory condition, reported as a percentage (0–100%). Each trial required participants to select the correct portal based on a given cue, and accuracy reflects the effectiveness of cue interpretation.
Task Completion Time: The time (in seconds) taken by participants to complete each trial, measured from the onset of the task to the submission of a response. Each participant completed 10 trials per sensory condition, and the average completion time was computed per condition.
Confidence Rating: After each trial, participants rated their confidence in their choice using a 7-point Likert scale (1 = "Not confident at all", 7 = "Extremely confident"). These ratings were collected across 10 trials per sensory condition and averaged to evaluate perceived certainty.
NASA-TLX: A standardized subjective workload assessment tool comprising six subscales: mental demand, physical demand, temporal demand, performance, effort, and frustration. Participants completed the NASA-TLX after finishing all trials within each sensory condition to evaluate perceived workload.
SSQ: A 16-item validated instrument used to assess symptoms of simulator or motion sickness (i.e., nausea, oculomotor discomfort, disorientation) experienced during VR tasks. The SSQ was administered after completing all trials in each condition to monitor adverse effects related to VR exposure.
After completing the task under each sensory condition, participants were asked to rate how helpful the sensory cues were using a 7-point Likert scale (1 = Not at all helpful, 7 = Extremely helpful). In addition, participants were also prompted to provide open-ended feedback describing how or why the cues were (or were not) helpful during the task. This experiment was designed to prove the following hypothesis.
Participants are expected to have higher accuracy rates and faster completion times in multisensory conditions (VO, VA, and VAO) compared to the V condition. The presence of congruent auditory and/or olfactory cues is anticipated to provide additional contextual information, thereby facilitating quicker and more accurate scene recognition.
Participants will report higher confidence ratings in their responses under multisensory conditions compared to the visual-only condition. This is because supplementary sensory inputs are likely to reduce uncertainty and reinforce the correctness of participants’ interpretations, resulting in greater perceived confidence during decision-making.
Participants are expected to show lower mental and physical load in multisensory conditions (VO, VA, and VAO) than in the V condition. This is because they will recognize the location more quickly and feel less burden when other senses are provided than when only vision is provided to recognize the location.
The VAO condition was expected to yield the highest SSQ scores, with VA and VO producing similar but slightly higher scores than the V condition. This expectation is based on the assumption that the addition of auditory and olfactory stimuli may introduce mild sensory conflicts or increase perceptual load, potentially leading to elevated simulator sickness symptoms compared to the baseline condition.
This study used a Meta Quest 3, a desktop to run the VR scene, and an olfactory device. Meta Quest 3 has a 110\(^\circ\) of horizontal FoV and a 96\(^\circ\) of vertical FoV, with a resolution of 2064 × 2208 pixels per eye. The virtual scene for the study was developed in Unity version 2021.3.19f1 using a high-performance PC equipped with an AMD Ryzen 5 4600H CPU, 16GB RAM, and a GeForce RTX2060, and Windows 11 was installed as the operating system.
The olfactory device used in this study was adapted for integration with the Meta Quest 3, based on the design proposed by Myung et al. [33] (Figure 3). It comprises two main components: the Scent Control Unit and the Scent Delivery Unit, which are attached to the side and lower sections of the Quest 3, respectively. It is capable of delivering up to four distinct scents and includes a built-in ventilation system to manage scent dispersion and clearance. The total weight of the device is approximately 102 grams.
The Scent Control Unit includes an Arduino Nano [50], a Bluetooth module (HC-06) [51], five 5V fans, and four scent modules. The Arduino Nano features a compact size, lightweight design, and low power consumption. It is powered directly via a USB connection to Meta Quest 3, requiring no external power supply. Communication between the device and the VR application—for both scent emission and ventilation—is handled via Bluetooth serial communication through the Arduino Nano.
The Scent Delivery Unit consists of five fans, four scent modules, and supporting frames as shown in Figure 3A. Each fan measures 30mm 30mm (width height) and operates at 5V. Among the five fans, four are used for dispersing scents, while the central fan is dedicated to ventilation. Unlike the scent fans, the ventilation fan is installed in reverse to help clear residual odors and maintain air circulation. Each scent module contains a cotton pad soaked in scented oil (Figure 3-a). When the system is activated, the fans generate airflow that disperses the scent into the environment, providing users with olfactory feedback.
Initially, 28 participants were recruited through university-affiliated social media channels. One participant was excluded from the analysis due to dizziness experienced during the task, resulting in withdrawal from the study. Consequently, data from 27 participants were included in the final analysis (21 males, 6 females; M = 23.04, SD = 2.89, age range = 18–27). This number of participants satisfied the requirements of our power analysis conducted with G*Power 3.1 for a within-subject ANOVA study. The parameters were: effect size f = 0.40, α error probability = .05, power = .80, number of groups = 4, number of measurements = 2, correlation among repeated measures = .50, and nonsphericity correction ε = 1. This analysis yielded a minimum sample size requirement of N = 24. All participants reported normal vision, hearing, and sense of smell. Nineteen participants had prior experience with VR. Each participant received approximately $7.27 for their participation, which reflected the local hourly minimum wage in the country where the study was conducted.
The study lasted approximately 60 minutes (IRB: HIRB-2024-038). Upon arrival, participants sign an informed consent form and complete a demographic questionnaire. An instructor then provides an overview of the study, including its objectives, procedures, scent types, and the devices used. To ensure participants were familiar with the scents, they were asked to smell the four scents that would be used in the task. Following that, participants engaged in a training session to practice viewing and selecting the portal and learning how to use the VR controllers to complete the task. The training session took about five minutes. Before proceeding to the main session, participants completed the SSQ questionnaire to assess their baseline state.
The main session consisted of four sub-sessions. It was conducted with participants standing. In each sub-session, participants performed ten tasks under one of the four sensory conditions. After completing each sub-session, they responded to the NASA-TLX and SSQ questionnaires. Then, they took a 3-minute break before proceeding to the next sub-session. During these breaks, the windows in the room where the experiment was conducted were opened for ventilation. And, the participants took off their HMD and took a break. After completing all tasks, participants fill out a post-questionnaire. At the end of the study, the instructor ventilates the room and sanitizes the olfactory device with alcohol to eliminate any residual scents.
We report the results analyzed with a one-way repeated measures Analysis of Variance (ANOVA) test at the 5% significance level. The degrees of freedom are corrected using the Greenhouse-Geisser correction to protect against violations of the sphericity assumption. A post-hoc test using the Bonferroni correction for multiple comparison is employed when a significant effect was observed at the 5% significance level. We report the mean, minimum, and maximum values for each condition.
There was a main effect on the accuracy rate (\(F(2.239, 58.213)\)= 5.467, \(\eta_p^2\)=.174, \(p\)=.005). The accuracy rate was lowest in V (M = 90.741 [86.502, 94.980]), followed by VA (M = 95.926 [92.970, 98.882]), VO(M = 96.667 [93.197, 100.136]), and highest in VAO(M = 98.519 [97.086, 99.951]). Pairwise comparisons revealed that V had a lower accuracy rate than VAO (\(p=.01\)) by 7.778 [1.432, 14.123]. There were no significant differences between the other conditions. The results are shown in Figure 4A. This result supports H1. This is because the accuracy rate was lowest in V and highest in VAO.
The results are shown in Figure 4B. There was a main effect on the task completion time (\(F(2.362, 61.4)\) =3.084, \(\eta_p^2\)=.106, \(p\)=.045). Task completion time was slowest for V (M = 17.333 s [14.275, 20.392]), followed by VO (M = 16.296 s [13.415, 19.178]), VA (M = 13.889 s [11.061, 16.717]), and VAO (M = 12.926 s [10.612, 15.240]) being the fastest. Pairwise comparisons revealed that V was slower than VAO (\(p=.027\)) by 4.41 [1.498, 7.317]. These results support H1, showing that task completion time was slowest under the VAO condition.
The results are shown in Figure 4C. There was a main effect on the Confidence Rating (\(F(1.842, 47.889)\)= 14.965, \(\eta_p^2\)=.365, \(p\)<.001). Confidence Rating was lowest for V (M = 5.641 [5.110, 6.172]), followed by VO (M = 5.919 [5.472, 6.365]), VA (M = 6.615 [6.448, 6.782]), and VAO (M = 6.844 [6.733, 6.955]). Pairwise comparisons showed that V had a lower confidence rate than VA (\(p=.002\)) and VAO (\(p<.001\)). In addition, VO (M = 5.92 [5.47, 6.37]) had a lower confidence rate than VA (\(p=.028\)) and VAO (\(p=.002\)).
These results support H2 in the Confidence Rating, indicating that the multisensory conditions yield higher confidence ratings for responses than in the visual-only condition.
There was a main effect of Performance (F(3, 78) = 3.013, \(\eta_{p}^2\) = 0.104, p = 0.035). Pairwise comparisons revealed that V had higher performance rate (M = 3.222 [2.452, 3.993]) than VAO (M = 2.296 [1.448, 3.145], \(p=.041\)).
A main effect of Effort (F(3, 78) = 4.548, \(\eta_{p}^2\) = 0.149, \(p = .005\)) was also disclosed. Pairwise comparisons showed that V (M = 3.519 [2.813, 4.224]) had a higher effort rate than VAO (M = 2.556 [1.762, 3.349], \(p=.040\)). In addition, VO (M = 3.370 [2.693, 4.048]) had a higher effort rate than VAO (\(p=.031\)). These findings partially support H3.
There were no significant main effects across SSQ subscales (Nausea (\(p\)=0.542), Oculomotor (\(p\)=0.668), and Disorientation (\(p\)=0.488) or in participants’ perceived helpfulness ratings (\(p=.059\)) among conditions. Therefore, H4 is rejected.
Our results showed that providing multisensory cues (i.e., the VAO condition) significantly improved accuracy rates, reduced task completion time, and increased participants’ confidence levels during scene recognition tasks. Moreover, all VO, VA, and VAO conditions yielded higher accuracy rates than the visual-only condition. Interestingly, although VO and VA did not significantly enhance these behavioral metrics beyond accuracy, they were associated with reductions in perceived effort, as shown in the NASA-TLX results. This suggests that even a single additional sensory cue can reduce cognitive load, potentially by helping participants filter and interpret ambiguous visual information. However, it is the combination of auditory and olfactory cues in VAO that likely provided a richer, more reliable context, reinforcing visual input and reducing perceptual uncertainty.
Notable differences emerged in subjective ratings and participant feedback regarding the use of auditory and olfactory cues in portals. Although there was no significant difference in recognition accuracy between VA and VO, participants consistently rated audio cues as more effective and easier to distinguish than olfactory cues. Participants reported that olfactory cues caused increased fatigue and were difficult to differentiate, whereas auditory cues were perceived as more intuitive and less cognitively demanding. These findings suggest that while both modalities can aid performance, auditory cues may offer a more practical and user-friendly enhancement for portal interactions in VR.
These findings are consistent with multisensory integration theory, which suggests that the brain combines information across modalities to enhance perception and reduce uncertainty [16]. The VAO condition provided congruent auditory and olfactory signals that reinforced visual input, allowing participants to bind cues into a coherent representation and thus respond more quickly and accurately. This pattern is also compatible with cross-modal priming, where exposure to one modality (e.g., sound or scent) may activate expectations in another, supporting faster recognition and lower cognitive load [52].
Participants reported higher confidence ratings in conditions that included additional sensory cues compared to the V condition. Notably, the VAO condition elicited the highest confidence ratings, suggesting that the combined use of visual, auditory, and olfactory cues helped participants feel more certain about their decisions. While both VO and VA conditions also led to increased confidence compared to V, their effects were less pronounced than the full multisensory VAO condition. These findings suggest that multisensory integration, particularly when multiple senses are simultaneously engaged, can enhance users’ perceived certainty during recognition tasks in VR environments, even when visual information is limited.
NASA-TLX results revealed partial support for H3. Specifically, participants in the VAO condition reported lower effort than in V and VO, and lower performance workload ratings than in V. These results indicate that multisensory cues can help offload some of the cognitive burden associated with interpreting limited visual information. Interestingly, participants in the V condition reported the highest self-rated performance, despite performing worst in terms of accuracy, suggesting a potential overestimation of ability when relying solely on visual information. This discrepancy highlights the value of objective metrics in evaluating performance and workload.
Qualitative feedback and post-task interviews revealed a persistent reliance on visual cues. In the post questionnaire, 16 participants commented they preferred VAO the most, followed by V (8), VA (3), and VO (1). Participants often used the visually displayed elements inside the Portal to cross-reference with audio or scent cues. While multisensory input aided recognition, the visual modality remained the anchor for decision-making. This suggests that while multisensory cues enhance performance, their role may be more supportive than primary in visually constrained tasks. Their feedback was consistent with earlier research findings on the effects of multisensory cues on human perception. Previous research has shown that when visual and auditory stimuli are presented simultaneously, people tend to prioritize the visual component [53], known as the Colavita Visual Dominance Effect. Similarly, earlier research examining vision and olfaction stimuli together suggested that vision exerts a stronger influence on human perception than olfaction [52], [54]. Future work should consider adding visual-free conditions to better isolate and understand the true effects of sensory augmentation.
Contrary to concerns that adding olfactory and auditory stimuli might increase discomfort, our results showed no significant differences in SSQ scores across the four conditions. This rejects H4, indicating that the introduction of multisensory stimuli did not lead to increased symptoms such as nausea, oculomotor strain, or disorientation. This finding is consistent to earlier findings demonstrated the minimal effects of audio and olfaction on simulator sickness [55], [56]. In contrast, Keshavarz et al. reported that pleasant odors may even help alleviate simulator sickness. This evidence collectively supports the viability of incorporating scent and sound in VR environments without compromising user comfort or safety.
A primary limitation of this work lies in the constrained design of the experimental tasks. The scenarios used in our study may not fully capture the complexity of real-world VR portal interactions, particularly features such as dynamically resizing portals, interacting with objects within portals, or navigating directly through them. These simplified tasks, while useful for controlled measurement, may introduce learning effects in completing our tasks in repeated trials and may not reflect how users engage with multisensory cues during more complex and interactive VR experiences. To address this, future work should explore the use of auditory and olfactory cues in more ecologically valid and dynamic contexts, such as collaborative environments, narrative-driven scenarios, or large-scale virtual spaces that demand ongoing interaction, spatial reasoning, and information retrieval. Moreover, while this work primarily focused on quantitative metrics, asking participants to verbally describe scenes visible through portals and their associated multisensory cues could provide deeper insight into how these cues shape comprehension of what lies beyond the portal. Such approaches would help assess the scalability and practical effectiveness of multisensory feedback in supporting immersive portal-based navigation.
A second limitation involves the olfactory feedback system. Several participants reported that lingering scents from previous trials occasionally interfered with their ability to detect new scent cues. This issue may have arisen from limitations in scent management, despite providing rest periods between trials. Because olfactory cues are more difficult to clear than visual or auditory stimuli, future work should consider more robust dispersion and removal strategies, such as active ventilation, the use of scent-neutralizing agents, or longer inter-trial intervals, when designing VR studies with olfactory stimuli.
Lastly, this study’s participant pool was constrained in diversity. In particular, the number of male participants was substantially higher than the number of female participants (21 vs. 6). Prior research has shown that females generally exhibit greater olfactory sensitivity and discrimination ability than males [57], which could have influenced how multisensory cues were perceived in our study. Moreover, the participants were all young adults, which limits the ability to generalize our findings across other age groups. Olfactory sensitivity and multisensory integration are known to change with age, with declines reported in older adults and developmental differences observed in younger populations. These demographic constraints represent potential sources of bias. Future research should therefore recruit larger and more heterogeneous participant groups, spanning gender, age, and cultural backgrounds, to enhance the robustness and generalizability of the results.
This work investigated the effectiveness of incorporating multisensory cues—specifically auditory and olfactory stimuli—into Portal, which is a visually constrained VR interaction method. We conducted a formal user study to evaluate how these cues influence users’ spatial understanding within VR portals. The findings demonstrated that the addition of multisensory cues improved participants’ ability to quickly and accurately grasp the context of virtual scenes, compared to visual-only conditions. Importantly, while both auditory and olfactory cues contributed to enhanced accuracy, subjective responses revealed a clear preference for auditory cues. Participants reported that auditory stimuli were easier to perceive, less fatiguing, and more intuitively linked to scene content. In contrast, some found olfactory cues harder to distinguish and more mentally taxing, suggesting that the effectiveness of scent-based interaction may be constrained by current scent delivery technologies and scent recognizability. Finally, we concluded by outlining our study limitations and suggesting future research directions to explore portal-based interactions in more diverse and practical VR scenarios.
This research was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. RS-2023-00254695).
All scent products were procured from https://www.esfood.kr.↩︎