Last updated: 2026-08-28 7:14PM EDT
Wednesday, 7 October
Thursday, 8 October
Friday, 9 October
Wednesday – 7 October 2026
PS1 - PS3 11:15–12:15
PS4 - PS6 11:45–12:45
PS7 - PS9 13:45–14:45
PS10 - PS12 14:15–15:15
PS13 - PS15 17:15–18:15
PS16 - PS18 17:45–18:45
Paper Sessions 1-311:15–12:15
PS1: XR Learning Orione + Perseo
-
1107 Mobility Learning in Microgravity: AR-guided Spacewalk Exploration for Education and Spatial Cognition
Spacewalk during extravehicular activities (EVAs) is uniquely challenging due to microgravity and the complex spatial structure of space stations. These characteristics also make EVA an inherently engaging yet difficult to grasp learning scenario for spacewalk education, where novices often struggle to balance high-fidelity experiential immersion with effective spatial cognition and learning. Therefore, we introduce augmented reality (AR) as a form of guidance for educational exploration, supporting learners in understanding station structure, orientation, and mobility principles. We begin by conducting a formative study with aerospace experts to identify domain knowledge and key learning barriers for AR-guided spacewalk education. Guided by these insights, we develop a high-fidelity AR simulation prototype for exploratory spacewalk touring that incorporates EVA constraints and multiple movement modes, providing visual waypoint guidance, posture-aware feedback, and real-time route replanning. We evaluate the prototype through a controlled within-subject user study with 16 participants. Results show that AR-guided exploration improves route comprehension and embodied orientation awareness, enhances performance during exploration, and enhances cognitive map construction during tasks. Finally, we discuss design implications for optimizing the prototype across contexts ranging from in-depth education to operational astronaut support systems. -
1129 Olfactory Context Cues for Procedural Skill Transfer from Digital Twin-Based Mixed Reality to Real-World Execution
Mixed Reality (MR) learning systems employ Digital Twins (DTs) to replicate real-world tasks, yet sensory discontinuity between DT-based MR environments and physical task contexts can hinder skill transfer. We investigate whether olfactory contextual cues can bridge this gap by serving as cross-environment anchors for memory encoding and retrieval. In a two-day study, 48 participants learned Rubik's Cube rotation notations in a DT-based MR environment and were tested on a physical cube (Day 1), then completed a transfer test approximately 24 hours later (Day 2). A 2 × 2 between-subjects design manipulated rosemary odor presence during learning and testing. Results showed that rosemary during DT-based MR learning increased spatial presence and motivation while producing equivalent short-term task completion time compared to the no-odor condition. For overnight transfer, when rosemary was present during learning, reinstating the same odor at testing yielded higher transfer performance and lower cognitive load, and reduced recall difficulty compared to the no-odor condition at retrieval. More broadly, matched olfactory conditions between learning and testing produced higher accuracy and reduced cognitive load relative to mismatched conditions. These findings provide the first evidence that olfactory contextual cues can support procedural skill transfer from DT-based MR environments to real-world execution, and suggest that when real-world tasks involve characteristic ambient odors, replicating them in virtual training may help preserve encoding–retrieval congruence. -
1328 The Importance of Pedagogical Agent Design in Gamified Immersive Virtual Reality for Vocational Education and Training
This study investigates how different pedagogical agent designs can affect learning performance and user experience in gamified Immersive Virtual Reality (IVR) for Vocational Education and Training (VET) instructional activities. Although theory predicts benefits from combining gamification and pedagogical agents, it remains unclear whether such integration enhances learning and user experience or introduces extraneous stimulation. To fill this gap, we conducted a user study with 84 VET students comparing four pedagogical IVR conditions within an engineering-oriented VET learning task: non-gamified, gamified without a companion, gamified with a side-by-side companion, and gamified with an in-situ companion as part of the learning activity. We assessed the educational impact of these conditions, measuring learning outcomes, alongside psychological and behavioral metrics related to learners’ well-being in line with Education 5.0 principles, including cognitive load, attention, engagement, and motivation. Findings indicate that learning performance was preserved across conditions, while the inclusion of an active in-situ pedagogical agent enhanced engagement, optimized cognitive load distribution, and improved user experience. -
2051 Exploring Effects of Physical and Cognitive Demands: Simulating AR-based Real-world Situations for Procedural Tasking in UH-60 Helicopters
Prior research on augmented reality (AR) and virtual reality (VR) task guidance has largely focused on environments and experiences that are typically free of external distractions or conditions that impact user performance and cognitive load. However, this presents a gap the literature with regard to understanding how simulating adverse conditions (e.g., poor lighting, rain, loud noises) from the real world affect user perception and performance. Therefore, in our work, we present one of the few studies to investigate the usage of interaction cues in task guidance contexts that involve these adverse, distracting conditions. We conducted a 3 × 3 within-subjects experiment in which 22 participants performed procedural button-pressing tasks in a simulated UH-60 Blackhawk helicopter while also completing a 2-back cognitive task. Our study systematically varied interaction cue jitter and environmental conditions as to simulate potential situations co-pilots may find themselves in while piloting a helicopter. The results of our work indicate the jitter has a significant, adverse effect on task completion time and performance. Furthermore, the storm conditions led to significant increases in mental effort, physical effort, and user discomfort despite not being directly related to the primary task. In our paper, we further detail our findings and discuss how our results highlight the importance of accounting for potential adverse conditions that occur in real-world experiences. We also provide avenues of future work to further explore this nuanced area of research.
PS2: Crowd Interaction Sezione 4+5
-
1902 Who Feels the Crowd? Predicting Individual Cognitive Load Sensitivity to Virtual Crowds
Virtual crowd simulations underpin applications from social VR to exposure therapy, yet their design rarely accounts for wide individual variation in crowd-induced cognitive demands. We investigated which virtual crowd characteristics impose the greatest cognitive load and whether distinct individual profiles predict differential vulnerability. Fifty-four participants completed a dual-task paradigm (spatial navigation + auditory 2-back) across 13 conditions in a 2×2×3 within-subjects factorial design manipulating crowd density (sparse vs. dense), gaze behavior (neutral vs. evaluative staring), and movement pattern (static, unidirectional, bidirectional), plus a no-crowd baseline. Crowd presence impaired 2-back accuracy by 11.8% and elevated physiological arousal relative to baseline. Among crowd factors, evaluative gaze emerged as the dominant cognitive load driver, with a significant three-way interaction revealing synergistic impairment when density and socialevaluative scrutiny converged in dynamic configurations. Latent profile analysis of 12 indicators across trait psychology, baseline physiology, and crowd reactivity identified four distinct sensitivity phenotypes: Highly Sensitive (22%), Moderately Sensitive (28%), Stable/Adaptive (28%), and Resilient (22%), each showing qualitatively different vulnerability patterns. A nested cross-validation pipeline using pre-exposure markers classified profile membership with 91.6% accuracy, demonstrating that individual crowd sensitivity is predictable before any VR exposure, enabling adaptive systems to personalize crowd parameters to individual tolerance thresholds. -
1909 Asymmetric Interaction in VR Game Live Streaming: Effects of Identity and Device on Audience Experience and Collaborative Efficiency
Virtual Reality (VR) game live streaming has emerged as a fast-growing interactive entertainment format, yet the asymmetric interaction between streamers and audiences remains an under-explored core challenge. Grounded in asymmetric interaction theory, this study systematically examines how two core dimensions of asymmetric interaction—interaction capability asymmetry and device asymmetry —shape audience experience and collaborative efficiency in VR game live streaming. We conducted a 2×2 within-subjects experiment. Results show that: (1) Participant identity (higher interaction capability) significantly improves all dimensions of audience experience and collaborative efficiency; (2) VR device significantly enhances player experience, user behavior intention, and especially social presence; (3) Significant interaction effects exist between identity and device on player experience and social presence, with the Participant-VR condition yielding the optimal experience and highest efficiency. This work extends asymmetric interaction theory to the emerging VR live streaming context, provides empirical evidence for the mechanism of multi-dimensional asymmetric interaction, and offers actionable design guidelines for VR live streaming platforms, game developers, and content creators. -
1928 NPCRadar: NPC-Centered VR Multi-User Crowd Forecasting
Multi-user virtual reality (VR) enables immersive entertainment experiences in which multiple users share a virtual space and continuously influence one another through their movements and interactions. These crowd-state changes directly affect how non-player characters (NPCs) should respond, yet existing VR behavior understanding research mainly focuses on individual-level modeling. To address this gap, we introduce the task of NPC-centered crowd forecasting in multi-user VR, which aims to predict future crowd states relative to the NPC over a short time horizon. We represent crowd states with three complementary dimensions: Flow, Density and Intent. Based on this formulation, we propose CrowdRadar, establishing a complete pipeline from historical multi-user behavior observation to group-level future-state prediction. By extending VR behavior understanding from individual behavior modeling to crowd-level forecasting, our work provides a structured basis for proactive NPC control and supports more intelligent, natural, and socially aware interaction in multi-user VR systems. -
1969 Psychophysiological Effects of Violent and Non-Violent Crowd Behavior Simulation in Immersive and Non-Immersive VR
Understanding human reactions to violent incidents involving crowds is crucial for crowd management, evacuation, and emergency response. Unfortunately, studies involving real-world simulation of such scenarios are often infeasible due to ethical, practical, and safety constraints. Virtual Reality (VR) is increasingly used to provide controlled simulation of such scenarios while preserving a degree of ecological validity. This work investigates how violent crowd behavior is experienced in immersive VR (iVR) and non-immersive VR (niVR). We conducted a user study with 74 participants in a virtual train station, manipulating crowd behavior (violent vs. non-violent) and VR technology (iVR vs. niVR). We collected self-reported measures about affect, sense of presence, social presence, and perceived realism of the crowd. Moreover, we recorded physiological indicators of stress and arousal (heart rate variability and electrodermal activity). Results showed that violent crowd behavior induced more negative emotional valence and lower perceived realism and sense of presence compared to non-violent crowd behavior, regardless of VR technology. The effect of VR technology was significant for arousal, user involvement, and stress, as reflected in heart rate variability metrics (SDNN, HF, RMSSD). Overall, this work shows that despite niVR being sufficient to elicit self-reported emotional responses to violent crowd behavior, only iVR can also trigger physiological emotional responses, which could be relevant in many VR applications such as crowd management training systems.
PS3: Finger Touch Input Sezione 1
-
1370 One Finger Warrior: Thumb-to-Index Finger Interaction Technique for Mobile Use
As Head-Mounted Displays (HMDs) become increasingly prevalent, designing User Interfaces (UIs) for quick, on-the-go access is essential. We introduce One Finger Warrior, a thumb-to-index finger interaction technique that enables mobile input on HMDs. The technique segments the index finger into six zones: three on each phalanx of the volar (front) and radial (top) sides. Taps and swipes between these zones allow for a high-density input vocabulary. We examine how different mobility conditions affect interaction speed, accuracy, and comfort, and find that they do not impact user performance in One Finger Warrior. Based on this, we design a text-entry application and evaluate its performance envelope by comparing two layout configurations: Optimized and Sub-Optimal. We use a text-entry task as it provides a standardized and cognitively demanding rapid target-selection task that helps characterize the performance envelope of the technique. The performance of both layouts was evaluated under varying mobility conditions. Results show that mobility has no significant impact on performance, with users achieving text entry speeds of 8.31 WPM on the Sub-Optimal layout and 12.05 WPM on the Optimized layout, demonstrating One Finger Warrior’s effectiveness for on-the-go input. We present a set of applications that demonstrate One Finger Warrior’s ability to enable mobile input on HMDs. -
1680 GeoCtrl: Finger-based Geometric Mapping and Speed-Adaptive Gain for In-Hand Object Rotation in AR/VR
We propose a geometry-based virtual object rotation technique that leverages triangular configuration of fingertips and a speed-responsive Control-Display (CD) gain. Rather than replicating real-world physics which often lead to unnatural mid-air behavior in the absence of physical feedback, we leverage finger-level geometry to enable more intuitive object rotation. We map the orientation and morphology of a triangle formed around fingertips directly to the object's rotation, assuming that users tend to trace the object's intended motion when physical constraints are missing. To enhance the proposed interaction and accommodate natural motor limits, we implement an adaptive CD gain that scales geometric input according to the fingertip movement speed. A user evaluation revealed that our method significantly reduces task completion time compared to physics-based rotation and physical demand compared to both reference conditions, without increasing overall workload. We demonstrate that geometry-driven mapping and dynamic gain adaptation enable more efficient dexterous manipulation of virtual objects in immersive environments. -
1722 Effects of Spatial Constraint on Bare-Hand Logographic Character Handwriting Performance in Virtual Reality
Spatial constraints significantly influence motor performance in virtual reality (VR), as interactions often involve hands hovering in mid-air. Yet their specific impact on logographic handwriting, such as Chinese and similar characters, remains underexplored. While VR text entry research is predominantly alphabetic-centric, especially English, handwriting support for complex scripts is fragmented across systems that treat physical or virtual guidance as implementation details rather than analytical or design variables. To address this, we compare three spatial constraint form as a distinct interaction design variable under a unified input pipeline: Unconstrained Mid-Air, Virtual Collision Surfaces, and Anchored Physical Surfaces. In a 3 × 3 within-subjects study (N = 24) with users writing Chinese characters of various complexity, we identified a critical stability-efficiency trade-off. While constrained conditions (virtual and physical) reduced trajectory length and rewrite counts, they increased normalized writing and total input time compared to mid-air input. Subjectively, users rated Unconstrained Mid-Air as more usable, whereas Anchored Physical Surfaces uniquely mitigated posture-maintenance workload without sacrificing stabilization benefits. We also contribute design implications on when freer mid-air input may be preferable and when planar or physically supported writing may be more beneficial in VR handwriting. -
1818 Tactile Search: Enhancing Targeting in 3D Space
Visual search is crucial in daily life, from scanning for relevant information to spotting signs of danger. When sensory channels are overloaded or degraded, cognitive tasks can be supported by crossmodal information representations through vibrotactile cues. We introduce Tactile Search, an approach that uses modulation of frequency and amplitude of vibrations to the hands, for guiding attention to the location of objects in 3D space. We evaluated this approach in a competitive VR game where participants searched for targets using both vision and touch. Across two studies -- an in-the-wild demonstration (n=55) and a controlled laboratory experiment (n=28) -- we found that vibrotactile feedback significantly improved performance and increased user confidence. Frequency modulation supported vertical targeting by leveling performance across target heights. We further analyzed participants' subjective experiences and search strategies highlighting the benefits of the tactile cues. Our findings establish Tactile Search as an effective means of enhancing interaction and provide design considerations for integrating haptic search into interactive systems.
Paper Sessions 4-611:45–12:45
PS4: Scene Reconstruction Cigno + Auriga
-
1008 Mix3R: Zero-shot Interactive 3D Scene Reconstruction from a Single Image in Mixed Reality
We present Mix3R, a training-free 3D scene reconstruction framework that enables immersive Mixed Reality (MR) interaction from a single real-world image. Despite recent progress in single-image 3D scene generation, existing methods often struggle with background recovery under occlusion and lack evaluations regarding natural interaction as well as immersion in MR environments. We address these challenges through a structured integration of large-scale model priors, enabling geometrically consistent 3D scene reconstruction without additional training. The framework combines vision–language-guided foreground–background separation with depth alignment from back-projected point clouds to achieve reliable background completion while preserving detailed object geometry. Experiments on the 3D-FRONT dataset show that Mix3R improves reconstruction quality, reducing Chamfer Distance by over 23% and increasing object-level F-score by more than 6%, with additional gains in scene-level perceptual metrics. User studies on in-the-wild images further demonstrate higher user preference for the proposed method across diverse real-world scenes. These results indicate that our approach enables practical reconstruction of interactive MR spaces from an arbitrary single image. Our project page is https://anonymous-17482.github.io/Mix3R/ -
1744 MR-Compare: A Mixed-Reality Framework for Spatially Grounded Visual Comparison of Heterogeneous 3D Reconstructions with Reality
We introduce MR-Compare, a spatially grounded mixed reality framework that registers heterogeneous 3D reconstructions to the live video see-through (VST) view for visual comparison. Implemented on a Meta Quest 3, it combines a two-stage registration pipeline with a 3D Slider to enable seamless cross-media comparison between the reconstructed past and the live present. Through a real-world benchmarking and exploratory user study involving 30 participants, we evaluated five representative workflows spanning both mesh-based and 3D Gaussian Splatting (3DGS)-based reconstructions across desktop pipelines (Nerfstudio 3DGS, 3DGS-MCMC, and RealityScan mesh) and mobile capture pipelines (Polycam mesh and Scaniverse 3DGS). The results showed that MR-Compare achieved centimetre-level registration in all cases. Desktop 3DGS workflows outperformed the other methods overall, with 3DGS-MCMC showing the lowest registration error (0.91 cm and 0.86 degrees) and the highest visual consistency with the VST. These subjective findings were strongly corroborated by our objective evaluations. MR-Compare also demonstrated high overall usability and low cognitive workload across all tested workflows. To further optimise registration precision for 3DGS workflows and complement our two-room user study, we propose the anisotropy filter, a zero-shot module that leverages Gaussian anisotropies to extract geometric surface proxies. Evaluations on the Replica dataset show that moderate pruning improves robustness and reduces alignment errors. For standard 3DGS, the filter reduced translation and rotation errors by 0.35 cm (37.6%) and 0.08 degrees (38.7%), respectively. For 3DGS-MCMC, it reduced translation and rotation errors by 0.12 cm (16.9%) and 0.04 degrees (29.1%), respectively. -
1878 SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes
Estimating the physical properties of objects is as important as reconstructing their geometry for achieving physically coherent interactions in mixed reality. Prior work typically analyzes object dynamics observed in video or object semantics inferred from a single image, relying on vision-language models (VLM) for contextual reasoning. These methods are either computationally costly or lack geometric grounding, and consequently fail to capture spatial and contact relationships across objects. To address these challenges, we present SnapPhysics, a framework that reconstructs 3D objects with their physical properties, such as mass, friction or center of gravity, from a single image through instance-level 3D object reconstruction, spatial alignment, and reasoning based on a scene graph. The physics-aware scene graph is central to our approach. It is provided to a VLM along with per-object geometry to infer physical properties. Our experiments on the 3D-FRONT dataset demonstrate that SnapPhysics improves the F-Score on a scene level by 18.6% over previous approaches and achieves competitive performance in physical property estimation, reducing mean absolute log difference error (mALDE) by up to 20.5% and improving linear correlation (r2 ls) by up to 19.6% compared to VLM-only estimation. SnapPhysics enables physically interactive MR experiences from an image without any manual parameter tuning. -
10.1109/TVCG.2026.3698654 Semantic Scene Graphs for Creating a Localization-Ready Internet of Things
Controlling devices connected to the Internet of Things often requires juggling multiple smartphone apps or physical remote controls, creating a fragmented user experience. Augmented Reality (AR) can afford superior control by automatically presenting virtual user interfaces that are spatially aligned with networked devices. However, before such user interfaces can be delivered, physical devices must be localized in the environment. This paper introduces LORIOT (LOcalization-Ready Internet Of Things), a novel end-to-end system that uses a semantic scene graph and a large language model to map the identities of the networked devices to physical objects, given a pre-filtered set of IoT-device candidate nodes. A declarative UI specification enables automatic generation of device control panels for AR and non-AR clients. We evaluate the mapping component on a controlled synthetic-room benchmark of 100 randomly generated rooms. Using device network metadata alone, we achieve a baseline macro-averaged F1 score of 0.80 for digital → physical associations. When device metadata is enriched with physical attributes (mounting location, materials, color, and size), performance improves to 0.88. Moreover, we evaluate the benefit of spatially registered AR control in a within-subject user study (N=20), comparing in-situ AR panels against conventional non-AR control with smartphone apps or physical remote controls. AR yields significantly faster task completion, lower mental demand, and higher usability.
PS5: Haptic Interaction Glasshaus
-
1322 Direction-Aware 3D Haptic Retargeting in VR
Passive haptic retargeting effectively addresses real-world limitations in virtual reality by repurposing physical proxies. However, existing methods typically rely on single proxies, 2D planar arrangements, or fixed 3D proxies, failing to provide authentic tactile feedback for complex 3D interaction tasks. This limitation stems from the challenges of managing mobile proxies in volumetric space while ensuring directional consistency between virtual and physical contacts. To address these problems, this paper proposes a direction-aware 3D haptic retargeting method. We first design a magnetic physical proxy system that enables stable 3D stacking and robust omnidirectional pose tracking. Central to our approach is the introduction of a Haptic-Direction Score, which quantitatively evaluates non-occupied proxies and available placement positions with Local Exposure, Global Haptic Direction Quality, and Non-occupied Proxy Availability. We also propose a haptic retargeting method based on Haptic-Direction Score, which dynamically maps virtual targets to optimal physical targets during the retargeting process. A user study comparing our approach with state-of-the-art methods demonstrates that our method significantly reduces the system reset rate, while improving task completion efficiency, the consistency of haptic perception, and overall user experience in complex 3D environments. -
1525 Shaping Redirected Walking with Ankle-Based Vibrotactile Feedback
Redirected Walking (RDW) is a locomotion technique where users can explore large virtual environments while physically moving within a limited real-world space. Existing haptic redirection techniques often rely on infrastructure-heavy setups that limit deployability in real-world XR systems. In this work we investigate whether lightweight, ankle-based vibrotactile feedback can modulate perceptual sensitivity to RDW curvature manipulations. We present a wearable, wireless, gait-synchronized ankle-tendon vibrotactile system that delivers unilateral or bilateral cues during the swing phase of walking. In a within-subject study (N=28) spanning 14 experimental conditions, we systematically evaluate how stimulation side, temporal profile (constant vs. gait-velocity modulated), and curvature direction affect curvature detection thresholds. Results show that RDW effectiveness is strongly state-dependent: slow walking consistently permits larger curvature gains, and spatially asymmetric (unilateral) ankle stimulation produces greater perceptual bias than bilateral cues. Movement-contingent stimulation further improves perceived naturalness, presence, and user acceptance. Together, these findings provide the first systematic investigation of ankle-based vibrotactile augmentation for RDW and demonstrate that portable lower-limb haptics can meaningfully extend redirection capabilities without environmental instrumentation. We derive design guidelines highlighting the importance of locomotor-state adaptation, spatial asymmetry, and gait synchronized feedback for scalable redirected walking in XR. -
1939 Look Forward, Walk Naturally: Effects of an Extended Downward Field of View on Posture and Immersion in Virtual Stair Walking
Avatar motion remapping enables users to experience walking on virtual stairs while physically walking on a flat surface. In stair walking, the downward vertical field of view (FOV) plays an important role in ground recognition and foot placement; however, its perceptual and behavioral effects remain poorly understood, largely due to the FOV constraints of conventional head-mounted displays (HMDs). An extended downward FOV may enhance postural naturalness by supplying peripheral visual cues, yet it could also heighten awareness of the visuo-proprioceptive mismatch inherent in avatar remapping. To address this gap, we investigated the effects of downward FOV on posture control, height perception, and applicability of virtual stair walking using an HMD with an extended downward viewing range. Our results reveal that, without an extended downward FOV, participants actively compensated by tilting their heads downward to acquire necessary visual information, yet this behavior yielded no significant differences in subjective height perception or awareness of the spatial manipulation. In contrast, providing an extended downward FOV significantly suppressed this compensatory head tilt, promoting a natural, forward-facing posture more consistent with real-world stair walking. Subjective evaluations further indicated enhanced experienced realism during ascent and greater general presence during descent. Together, these findings demonstrate that an extended downward FOV effectively fosters natural locomotion behavior and deepens immersion in virtual stair walking, highlighting downward FOV as a key design parameter for locomotion interfaces in VR. -
10.1109/TVCG.2026.3664327 Discrete Virtual Rotation in Pointing vs. Leaning-Directed Steering Interfaces: A Uni vs. Bimanual Perspective
In this work, we explore the integration of Orientation Selection into Steering interfaces, aiming to preserve the seamless sensation of real-world movement while mitigating the risk of inducing cybersickness. Our implementation encounters conflicts in standard input mappings, prompting us to adopt bimanual interaction as a solution. Recognizing the complexity of interaction that may arise from this step, we also develop unimanual alternatives, e.g., utilizing a Human-Joystick, commonly referred to as a Leaning interface. The outcomes of an empirical study centered around a primed search task yield unexpected findings. We observed a sample of users spanning multiple levels of gaming experience and a balanced gender distribution exhibit no significant difficulties with the bimanual, asymmetric interfaces. Remarkably, the performance of Orientation Selection is, as in prior work, at least on par with Snap Rotation. Moreover, through a subsequent exploratory analysis, we uncover indications that Pointing-Directed Steering outperforms embodied interfaces in usability and task load in the given setting.
PS6: Avatar Identity Sezione 2
-
1011 The Impact of Point-Cloud Rendering on User Perceptions of Human- and AI-Controlled Avatar Embodiment
Avatars for social VR are increasingly driven by imperfect tracking or AI-generated motion, where mismatches between appearance fidelity and behavioral realism can harm user perception. We investigate whether a simulated point-cloud rendering (PCR) can improve social perception and increase tolerance to motion artifacts compared to standard shaded mesh rendering. We present a real-time VR avatar system that supports (i) a streaming conversational loop, (ii) human-controlled embodiment via hybrid tracking using a head-mounted display for upper body and an external camera for lower body and root motion, and (iii) AI-controlled embodiment via context-aware selection from an offline-generated animation library. We evaluate a 2x2x2 within-subjects study crossing rendering style (PCR vs. mesh) and control source (human vs. AI) across two sequential scenarios: a conversation-focused desk interaction and a movement-focused guided exercise. We measure social presence, human-likeness and affinity, and appearance–behavior plausibility, and collect post-study rendering-style preference rankings. Our results show that PCR significantly improves perceived human-likeness and plausibility, with particularly strong benefits for AI-controlled avatars. Crucially, PCR successfully closed the plausibility gap between human and AI avatars during high-movement interactions, acting as a perceptual buffer against motion artifacts. While social presence was driven by the interaction scenario rather than rendering style, post-study rankings revealed a trade-off between perceptual coherence and visual comfort. -
1078 When Avatars Get Under the Skin: Exploring Relationships between Neuroticism, Avatar Perceptual Features and Physiological Responses across Emotional Virtual Environments
Embodied photorealistic avatars are increasingly used in emotionally rich and diverse virtual reality (VR) applications due to their perceptual benefits. However, it is unclear how specific avatar perceptual features (APFs) and individual differences may relate to user physiology across different emotional virtual environments (VEs). In a within-participants study, we investigated the relationship between user neuroticism, APFs and physiological metrics, measured across four emotion VEs systematically manipulating emotional valence and arousal (N=51). Users embodied photorealistic personalised avatars (i.e., mirroring their appearance) and generic gender-matched avatars. We demonstrate that individual APFs have distinct relationships with user physiology, which are generally multiplied and reversed by higher user neuroticism. Moreover, APFs such as familiarity, appeal, appearance similarity and perceived realism can support emotional regulation by increasing heart rate variability (HRV) for higher neuroticism individuals. Overall, our findings highlight how avatar design can be strategically leveraged to optimise affective outcomes across diverse VR contexts. -
1090 Doppelgänger with Character: Effects of Self-similarity and Persona Valence on VR Conversation
Personalized, doppelgänger-like virtual agents may increase social realism, yet it remains unclear how users calibrate their perceptions and responses when an agent both resembles them and conveys distinct persona traits. In a 2 x 2 within-group study (N = 24), we examined how multimodal self-similarity and persona valence jointly affect study participants during open-ended conversations with an embodied virtual agent. After each condition, participants reported measures of perceived identity and persona, relational and social experiences, and affective responses; we also extracted conversational behavior from logged data, including word and turn counts and awareness-probing questions. We found that self-similarity increased perceived self-similarity, rapport, and emotional reactivity, while trust showed a valence x similarity interaction. Persona valence produced stronger overall shifts. Positive valence increased perceived agent persona, trust, co-presence, rapport, and willingness for future interaction, whereas negative valence increased emotional reactivity and elicited more words, turns, and awareness-probing questions. These findings clarify how resemblance and persona valence jointly shape user perceptions and interactions with embodied conversational agents and inform the design of self-similar virtual agents. -
2031 Shape-shifter avatar: How to embody a constantly-changing virtual body?
While virtual reality allows users to embody diverse non-human avatars, continuous body transformation such as growing new limbs or reshaping the body remains largely unexplored. We present a shape-shifter avatar based on metaball based implicit surfaces and examine how control mapping design and prior character knowledge influence embodiment, usability, cybersickness, preference, and plausibility. In a study with 24 participants, we compared three control mappings, Button, Task-Oriented, and Full-Control, across two tasks: extending a tentacle to reach objects and changing body height to search for keys. Prior knowledge was manipulated between participants using animations in which the character either transformed its body or remained in its default form. The Full-Control condition achieved the highest embodiment, plausibility, and preference. The Task-Oriented condition produced higher embodiment than button input by using task relevant body movements as triggers. Prior knowledge increased plausibility related to appearance and world fit, but did not affect evaluations of control correspondence or embodiment. Based on these findings, we discuss interaction design for shape-shifter avatars.
Paper Sessions 7-913:45–14:45
PS7: 3D Asset Modeling Orione + Perseo
-
1366 Learning Visual and Motion Aware 3D Model Representations for Semantic Retrieval
Virtual Reality (VR) content creators today face lengthy trial-and-error cycles when searching for 3D models, which relies heavily on manually annotated metadata. However, this approach struggles to capture the rich visual and dynamic details inherent to 3D models. Prior work has attempted to address this by rendering models as sets of images, but has largely remained restricted to category-level search and static geometry. In VR environments where asset behavior is often as important as appearance, we present a visual and motion aware 3D retrieval method that represents each model through multiview renders and short animations, extending search beyond static geometry to include dynamic behaviors. Pretrained image–text and video–text encoders extract features that are combined into a unified embedding by training a lightweight multimodal fusion model. This enables 3D model retrieval from text, images, and videos without requiring abundant and detailed ground truth data, supporting more nuanced and semantically meaningful search. On text queries referencing both appearance and motion, our method achieves a 23-percentage-point improvement in top-5 retrieval accuracy on the Objaverse dataset, while also attaining competitive accuracy on the ModelNet40 dataset and maintaining robustness as the model library size scales. A user study further shows that participants strongly preferred our retrievals compared to the one used by Sketchfab, one of today's most widely used 3D repositories. To support future work, we release 26k multi-view renders, animation clips, and 500 manually annotated visual and motion aware captions for the Objaverse dataset. -
1681 AtlasLC: Local-Competition Pruning and Deterministic Atlas Packing for Deployment-Aware Object-Centric 3D Gaussian Splatting
3D Gaussian Splatting (3DGS) enables photorealistic novel-view synthesis with real-time rendering, but deploying compressed object-centric 3DGS in XR requires more than image-space rate-distortion. In practical XR asset pipelines, reusable objects are repeatedly packaged, transmitted, decoded, and instantiated, making asset-preparation cost, codec compatibility, decoding latency, and preservation of depth and silhouette cues first-class concerns. Existing 3DGS compression methods are largely developed for scene-scale captures and often rely on heavy layout generation or aggressive global pruning, assumptions that transfer poorly to semantically concentrated foreground objects. We present AtlasLC, a source-free, training-free compression pipeline for object-centric 3DGS that operates directly on released Gaussian assets, without original images, camera poses, or per-asset optimization. AtlasLC couples local-competition pruning with deterministic atlas packing to remove the mapping/remapping bottleneck while preserving object-wide foreground support; a lightweight single-pass sort based conditional transport is used as a shared coordinate backbone for these stages. Across the evaluated assets, AtlasLC reduces atlas-preparation time by up to 25x and end-to-end compression time by up to 5x, while achieving the smallest average payload, the fastest decoding time, the highest FPS, and the best 3D F1 among compressed baselines. Relative to similarly compact structured baselines, it uses about 6-8% fewer bits while maintaining comparable or better perceptual and geometric quality. These results show that object-centric 3DGS compression should be optimized for a deployment-aware operating point, not bitrate alone, enabling scalable XR asset libraries. -
1732 OcclusionGS: Occlusion-Guided Gaussian Splatting for Interactive Object Segmentation in Virtual Reality
High-fidelity 3D scene reconstruction and object-level scene understanding are crucial to virtual reality and augmented reality applications, enabling tasks such as immersive scene interaction and object-level manipulation. However, the joint achievement of both tasks in open-world settings remains a significant challenge. To address this, we propose OcclusionGS, a framework that leverages a occlusion-aware scene structure to guide the 3D rendering process, enabling both high-fidelity reconstruction and precise object-level segmentation. The occlusion graph among adjacent objects in the scene is constructed by projecting a semantically-labeled 3D point cloud into all camera views and subsequently analyzing the depth relationships at each pixel. By aggregating all such pairwise occlusion relationships, we construct a view-dependent object topology graph for the scene. We translate the inferred occlusion graph into a powerful supervisory signal for the Gaussian optimization. This is achieved via a dual loss function, comprising an alpha loss on the full object silhouette for visibility learning and an RGB loss on the visible pixels for appearance reconstruction. Extensive experimen- tal validation on two challenging benchmarks demonstrates that our method achieves state-of-the-art performance in both object-level reconstruction and open-vocabulary semantic segmentation, outperforming existing NeRF-based and 3DGS-based approaches. Furthermore, by importing the reconstructed Gaussians into a VR headset, our method enables intuitive object selection in virtual space and cross-scene editing, allowing objects from one scene to be seamlessly transferred and integrated into another. -
10.1109/TVCG.2026.3651640 ESGaussianFace: Emotional and Stylized Audio-Driven Facial Animation via 3D Gaussian Splatting
Most current audio-driven facial animation research primarily focuses on generating videos with neutral emotions. While some studies have addressed the generation of facial videos driven by emotional audio, efficiently generating high-quality talking head videos that integrate both emotional expressions and style features remains a significant challenge. In this paper, we propose ESGaussianFace, an innovative framework for emotional and stylized audio-driven facial animation. Our approach leverages 3D Gaussian Splatting to reconstruct 3D scenes and render videos, ensuring efficient generation of 3D consistent results. We propose an emotion-audio-guided spatial attention method that effectively integrates emotion features with audio content features. Through emotion-guided attention, the model is able to reconstruct facial details across different emotional states more accurately. To achieve emotional and stylized deformations of the 3D Gaussian points through emotion and style features, we introduce two 3D Gaussian deformation predictors. Futhermore, we propose a multi-stage training strategy, enabling the step-by-step learning of the character's lip movements, emotional variations, and style features. Our generated results exhibit high efficiency, high quality, and 3D consistency. Extensive experimental results demonstrate that our method outperforms existing state-of-the-art techniques in terms of lip movement accuracy, expression variation, and style feature expressiveness.
PS8: Remote Collaboration Sezione 4+5
-
1228 Eye Contact Detectability Thresholds for Mixed Reality Telepresence
As remote and hybrid telepresence become more pervasive, preserving eye contact between interlocutors becomes paramount for regulating social interactions. Although prior work has explored eye contact dynamics in videoconferencing systems, further investigation is needed to determine whether these findings apply to mixed reality (MR) telepresence systems. This work contributes to this research by analyzing whether gaze direction misalignments disrupt eye contact perception in MR hybrid telepresence, which can occur when systems require users to focus on where interlocutors are displayed rather than on the acquisition camera. We calculate eye contact detectability thresholds through a two-alternative forced choice (2AFC) user study (N = 30 per set), analyzing how physical camera offsets along the X and Y axes influence eye contact perception. Participants in an MR environment viewed pairs of videos of interlocutors and indicated which showed better eye contact, one with no camera offset, or another with a camera offset in the [-10 cm, 10 cm] interval. We measured detectability thresholds for interlocutors with and without glasses, with dark and light irises. Findings suggest that iris color has a greater impact on the thresholds than glasses, with effects differing between the horizontal and vertical axes. These insights provide guidelines for understanding eye contact dynamics in MR telepresence. We also report a preliminary exploration on geometric transformations---rotating (yaw) the virtual panel displaying the interlocutor---can compensate for gaze misalignments in MR. Preliminary results suggest that panel rotation may partially restore perceived eye contact for moderate camera-display offsets. -
1358 MURMR: A Multimodal Sensing Pipeline for Automated Group Behavior Analysis in Mixed Reality
When teams coordinate in immersive environments, collaboration breakdowns can go undetected without automated analysis, directly affecting task performance. Yet existing methods rely on external observation and manual annotation, offering no annotation-free method for analyzing temporal collaboration dynamics from headset-native data. We introduce MURMR, a passive sensing pipeline that captures and analyzes multimodal interaction data from commodity MR headsets without external instrumentation. Two complementary modules address different levels of analysis: a structural module that generates automated multimodal sociograms and network metrics at both session and intra-session granularities, and a temporal module that applies unsupervised deep clustering to identify moment-to-moment dyadic behavioral phases without predefined taxonomies. An exploratory deployment with 48 participants in a co-located object-sorting task reveals that intra-session structural analysis captures significant within-session variability lost in session-level aggregation, with gaze, audio, and position contributing non-redundantly. The temporal module identifies five behavioral phases with 83% correspondence to video observations. Cross-tabulation shows that behavioral transitions consistently occur within structurally stable states, demonstrating that the two modules capture complementary dynamics. These results establish that passive headset sensing provides meaningful signal for automated, multi-level collaboration analysis in immersive environments. -
1438 CrossAtlas: Evaluating Projection Techniques for Spatial Referencing in Cross-Reality Collaboration
Cross-reality collaboration increasingly connects immersive and desktop users within synchronized workspaces, yet little is known about how bidirectional projection techniques between immersive 3D layouts and desktop 2D views influence communication. Spatial referencing depends on shared spatial understanding, but different mappings preserve and distort geometric relationships in different ways, altering perceived adjacency, orientation, and coverage across collaborators' views. We present CrossAtlas, a synchronized VR–PC collaboration platform that integrates multiple bidirectional projection techniques, including three planar projection variants and equirectangular, a spherical projection variant, across layouts of varying curvature. In a controlled study with 24 dyads, collaborators completed spatial referencing tasks under different layout–projection conditions while we collected performance and subjective measures. Our results show that projection choice strongly shaped collaboration, with the spherical variant consistently outperforming planar projections and remaining robust across layouts. -
1959 Investigating Points of View and Their User Control in Collaborative Heterogeneous Mixed Reality
Advances and increasing diversity in interactive Mixed Reality (MR) technologies foster heterogeneous collaborative environments, giving rise to multiple points of view (PoVs) on shared physical and virtual information. To address this diversity, we introduce a design space that enables precise descriptions of PoVs, their coupling and control by users during synchronous collaboration in MR. We provide an analytical exploration of this design space by classifying 28 existing remote-assistance MR systems in which a distant expert supports one or more on-site operators. Our analysis reveals that coupled PoVs are predominantly controlled by the operator. Guided by the design space, we introduce two new forms of control—expert control and dual-user mixed control—and present a first experimental evaluation of these three strategies: operator, expert, and mixed control of coupled PoVs. Results from a user study with 18 participant pairs indicate that expert control yields lower performance, usability, and preference as compared to mixed control, which offers greater flexibility for collaborative users. Moreover, mixed control achieves performance comparable to the commonly used operator control while being preferred by participants and perceived as less demanding by experts. This work advances the design of collaborative MR by introducing a PoV-centered perspective, PoV being an increasingly central feature in heterogeneous MR design. To do so, we characterize a PoV and we experimentally investigate design options for controlling coupled PoVs.
PS9: Gaze Gesture Input Sezione 1
-
1034 UniScale: Exploring Unimanual Gesture Mapping Strategies for Gaze+Pinch-based Scaling Interaction
Object scaling serves as a fundamental spatial manipulation that enables complex and productive tasks in XR environments. This paper investigates unimanual scaling techniques for XR using gaze and hand interactions. We propose UniScale, a set of unimanual alternatives to the standard bimanual pinch, allowing users to scale objects while preserving hand availability for concurrent spatial manipulations. We design five distinct mapping strategies based on physical metaphors, exploring unimanual control that varies depth, angle, micro-gestures, and finger-distance input. We then compare these techniques against a standard bimanual baseline, in which users adjust the inter-hand distance via a bimanual pinch gesture. In a user study, we evaluate their effectiveness in a 3D object scaling task under both clutching and clutching-free conditions. The results indicate that while bimanual scaling relies on clutching for stable control, unimanual techniques excel in clutching-free conditions, significantly reducing physical hand movement. From the results, we derive valuable design implications for developing efficient 3D multimodal interactions in XR. -
1265 MagicPitch: 3D Object Manipulation with Gaze, Pinch, and Head Pitch
Gaze has been integrated to assist 3D object manipulation with the hand for fast and effortless selection and radial movement, yet depth translation still largely depends on hand input. In this work, we investigate using gaze for radial redirection and head pitch for depth control together to complement 6DoF hand manipulation by introducing MagicPitch - a trimodal technique that integrates gaze, pinch, and head pitch. We conducted a 3D docking task (N=24) to evaluate \MagicPitch alongside its variant, which scales the pitch-depth effect via hand movement to improve stability. Our results show that redistributing depth control to head pitch significantly reduces hand effort with a trade-off to slightly increased neck motion, highlighting the potential of head–hand cooperation for fatigue-aware multimodal XR interaction. Overall, the findings demonstrate the effectiveness of trimodal integration of gaze, hand, and head pitch input for 3D object manipulation and suggest a general principle of complementing 6DoF hand input with gaze and head input to offload manual effort. -
1447 Point&Spawn: Mid-Air Reference-Free Object Instantiation Using Gaze and Hand Gestures in Extended Reality
Mid-air object instantiation in XR is challenging without physical references, often forcing users into weary ''instantiate-then-reposition" cycles. We present Point&Spawn, a technique that enables reference-free instantiation via a single continuous gesture flow. Point&Spawn utilizes a three-stage pipeline: (1) macro-level direction setting via Gaze or Non-Dominant Hand (NDH); (2) depth setting via a Dominant Hand (DH) semi-pinch, evaluating ray-casting against relative methods (Relative Gain, Drag&Hold); and (3) a full pinch for continuous micro-refinement before finalizing the spawn. A user study (N=24) revealed that relative depth settings significantly outperform absolute ray-casting in efficiency, accuracy, and workload, especially for distant targets. Furthermore, while NDH offers faster and more stable initial direction setting, Gaze effectively reduces hand fatigue. We conclude with actionable design guidelines for optimizing multimodal spatial placement in XR. -
10.1109/TVCG.2025.3615198 Manual-Free Gaze Interaction via Bayesian-Based Implicit Intention Prediction
Eye gaze is regarded as a promising interaction modality in extended reality (XR) environments. However, to address the challenges posed by the Midas touch problem, the determination of selection intention frequently relies on the implementation of additional manual selection techniques, such as explicit gestures (e.g., controller/hand inputs or dwell), which are inherently limited in their functionality. We hereby present a machine learning (ML) model based on the Bayesian framework, which is employed to predict user selection intention in real-time, with the unique distinction that all data used for training and prediction are obtained from gaze data alone. The model utilizes a Bayesian approach to transform gaze data into selection probabilities, which are subsequently fed into an ML model to discern selection intentions. In Study 1, a high-performance model was constructed, enabling real-time inference using solely gaze data. This approach was found to enhance performance, thereby validating the efficacy of the proposed methodology. In Study 2, a user study was conducted to validate a manual-free technique based on the prediction model. The advantages of eliminating explicit gestures and potential applications were also discussed.
Paper Sessions 10-1214:15–15:15
PS10: Assembly Workflows Glasshaus
-
1071 Delegate the Details: Flexible Assembly Constraint Authoring for Collaborative XR
Extended reality (XR) interfaces for assembly specification typically require users to specify the 6DoF pose of every part. This works well when exact poses are necessary, but can impose unneeded constraints when not. We present a collaborative XR constraint-authoring interface that supports a spectrum of specification authority. In our interface, an author can specify groups of parts and transfer them to a human or robot executor that carries out the author's instructions in the physical world. In prescribed mode, the author specifies the 6DoF pose of a part or group. In delegated mode, the author entrusts the executor to determine the 6DoF pose of a part or group. Both modes can be mixed within a single task. We conducted a within-subjects study comparing prescribed-only authoring against dual-mode authoring. Results show that dual-mode authoring significantly reduced authoring task completion time and perceived workload. Most participants significantly preferred dual-mode, with delegation rate scaling with the complexity of authored structures. We also evaluated both conditions offline with a robot executor, feeding authored constraints through a workspace optimizer and motion planner. Delegated constraints achieved significantly higher inter-group clearance than prescribed-only specifications. -
1449 Object-Driven Event Retrieval with Assembled Object Interface for Mixed-Reality Assembly Playback
We present object-driven event retrieval and an assembled object interface for Mixed Reality (MR) assembly playback, enabling users to select parts on the completed assembly and automatically retrieve and play back the corresponding joining process. When users want to review how a specific part was assembled, time-based navigation requires repeatedly scrubbing through the timeline and existing object-based approaches are limited to single-object queries that cannot address assembly scenarios where inter-part joining relationships are central. The proposed approach extracts inter-part joining events from a scene-graph-based recording and returns the relevant event segment by traversing the shortest path on the assembly graph, while presenting the completed assembly as a scaled-down 3D model so that part selection directly serves as a retrieval query. In a within-subjects experiment with 24 participants across drone, chair, and bicycle assembly scenarios, the proposed Object Interface(OB) significantly reduced navigation time compared to time-based comparison conditions, with the advantage most pronounced in scenarios with many parts or wide spatial extent, and the majority of participants selected OB as the most suitable interface for part-specific lookup. Based on these results, we expect object-based retrieval to support part-specific lookup that is difficult to achieve with time-based navigation alone, particularly in large-scale industrial assembly or maintenance environments with expansive workspaces. -
1900 ZCAR: Zero-Annotation CAD-Driven Assembly Recognition for Mixed Reality Guidance
We present ZCAR, a zero-annotation CAD-driven assembly recognition framework for Mixed Reality (MR) guidance that eliminates the need for manual training labels. Unlike generic object recognition, assembly-state tracking must align idealized 3D CAD assets with complex real-world observations, resulting in a substantial sim-to-real domain gap. ZCAR addresses this problem by using the LDraw CAD standard as a supervision source and rendering automatically labeled synthetic training images directly from 3D models, thereby bypassing the conventional data collection and annotation pipeline. The deployed system extracts visual embeddings with a frozen MobileCLIP2-S0 encoder and learns a lightweight projection head (0.33M parameters in the MobileCLIP2-S0 setting) using Supervised Contrastive Learning (SupCon). This design targets fine-grained state discrimination, enabling the model to capture subtle geometric differences between consecutive assembly steps. For resource-constrained spatial computing hardware, we further introduce a K-means-based embedding compression scheme that reduces memory usage while suppressing transient viewpoint noise. We also formulate a human-in-the-loop active perception strategy to resolve viewpoint-dependent ambiguities in which distinct assembly states appear similar from particular angles. In a controlled stress-test replay, the multi-frame Active Vision procedure improves average accuracy from 62.7% to 75.0% under ambiguous viewpoints. By favoring low-latency inference over heavyweight foundation models, ZCAR demonstrates the feasibility of real-time deployment on passthrough MR platforms such as the Meta Quest3. -
1461 Debugging Cyber-Physical Systems: The Impact of Spatial Registration and Degree of Virtuality
Debugging Cyber-Physical Systems like Digital Twins requires developers to mentally bridge the gap between two-dimensional source code and three-dimensional physical behavior. While Augmented Reality (AR) offers potential for immersive debugging by spatially registering information, Virtual Reality (VR) might serve as a suitable substitute when physical entities are not (yet) available or are inaccessible. In this paper, we investigate the impact of additional situated visualizations, placed next to the physical entity, versus additional stationary visualizations, placed next to the IDE. Furthermore, we evaluate the degree of virtuality by comparing these situated visualizations across AR and VR environments to assess VR's suitability as a substitute. We conducted a within-subject laboratory study (N=36) using a distributed robotic scenario to evaluate debugging performance and experience across three conditions: Stationary AR, Situated AR, and Situated VR. Results indicate that situated visualizations significantly enhance usability and user experience compared to stationary visualizations, even when objective performance metrics were comparable. Qualitative analysis reveals a trade-off between the spatial awareness of AR and the abstracted focus of VR, as well as a preference for the clear spatial assignment provided by situated panels. Our findings suggest that while situated visualizations do not necessarily speed up the resolution of bugs, they provide a more intuitive and satisfying workflow for identifying and fixing them.
PS11: Vibrotactile Haptics Cigno + Auriga
-
10.1109/TVCG.2026.3683951 HaptiCraft: A Modular Multimodal Haptic Controller for Immersive Virtual Reality Interactions
This paper presents HaptiCraft, a modular handheld haptic controller designed to replicate the appearance, mass distribution, and multimodal haptic properties of virtual objects. HaptiCraft consists of many modules that provide distinct functions for assembly and multimodal haptic feedback, supporting five types: vibration, impact, thermal, variable inertia, and variable stiffness. Users can readily assemble these modules to create a controller tailored to their needs. To simplify the design process, especially for novice users, we also provide a graphical authoring tool and assess its usability. Finally, we evaluate the system’s effectiveness through various virtual reality (VR) scenarios, demonstrating that HaptiCraft significantly enhances user experience. HaptiCraft makes a substantial contribution to the field of VR handheld controllers by offering comprehensive support for shape-changing, variable inertia, and rich multimodal haptic feedback, coupled with an effective authoring method. -
1068 Posture-Adaptive Azimuthal Guidance via a Forearm Vibrotactile Interface for VR Navigation
While forearm-worn interfaces are effective for tactile navigation in VR, the high degrees of freedom in arm posture-exacerbated by handheld controller use-introduce distortions in directional perception. This paper presents a posture-adaptive compensation model designed to deliver consistent 2D azimuthal cues regardless of arm posture, elbow flexion, and wrist rotation during handheld interaction. In our first experiment, we quantified how different arm postures shift perceived azimuth, establishing a foundation for our computational model. Notably, the perceptual shift remained substantially smaller than the physical change in elbow angle, indicating partial compensation between body-centered and device-centered coordinates. Based on these findings, we developed a posture-adaptive compensation algorithm combining representative-posture modeling with lightweight user adaptation. Subsequent VR navigation experiments showed that our model significantly outperforms methods that do not consider posture. This work contributes a methodology for eyes-free navigation, providing reliable spatial guidance in virtual environments by decoupling directional communication from user posture. -
1203 GroundedReach: Enabling Body-grounded Haptic Experience in Virtual Reality with an Elbow Wearable Haptic Device
Grounded haptic experience is essential for realistic virtual reality (VR) experiences, but existing solutions face a fundamental trade-off. Complex grounded devices and exoskeletons provide convincing force feedback, but limit user mobility and can be challenging to wear. In contrast, handheld and hand-wearable devices can maintain user mobility but do not provide sufficient grounded resistance. We identify a subset of scenarios common to many VR applications in which the user stands or sits in a fixed location (e.g., in a cockpit). In such situations, useful interactions with grounded objects can be realistically rendered by controlling their arm extension or flexion. We introduce GroundedReach, an elbow-wearable device that controls arm extension and flexion with a single-joint architecture. The system combines brake-based Dynamic Passive Haptic Feedback (DPHF) with a servo-controlled ratchet to render both continuous resistance and rigid stopping forces for interactions such as opening heavy doors and contacting hard surfaces. Our evaluation shows that GroundedReach can effectively simulate grounded feedback without applying force to the palm or shoulder. This design enables a compact, lightweight, and easily donnable peripheral that can be readily shared across users. -
1549 HaptoFlow: High-Fidelity Real-Time Vibrotactile Generation via Flow Matching for Virtual Reality
Haptic feedback is widely employed to enhance immersion in Virtual Reality (VR) environments. However, designing haptic stimuli that cover diverse interaction conditions remains a significant scalability challenge. Data-driven haptic generation has emerged as a promising approach, yet existing models face an inherent trade-off between waveform expressiveness and inference responsiveness, which becomes increasingly critical as training data grow in scale and diversity. To address this challenge, we propose HaptoFlow, a vibrotactile generative model based on Flow Matching, designed for interactive real-time haptic rendering in VR. Flow Matching learns a continuous vector field that transforms a base distribution into the target data distribution, enabling efficient representation of complex haptic data distributions and thereby facilitating both high-quality generation and computational efficiency. We train HaptoFlow conditioned on material labels and interaction parameters (stroking velocity and applied force), and integrate it into a VR system. Technical evaluation demonstrates that HaptoFlow outperforms all baseline methods in both waveform reproduction accuracy and inference latency. Furthermore, user studies confirm that the system latency falls well within the perceptual threshold of visual-haptic delay, and statistically significant improvements in perceived haptic quality are observed for a subset of materials. These findings establish a practical foundation for scalable, data-driven haptic content creation in VR, and provide latency benchmarks that inform the design of future real-time haptic rendering systems.
PS12: Body Ownership Sezione 2
-
1124 Having Dog Ears "for Real": Effects of Active and Passive Haptics on Embodying Non-Human Body Parts in VR
Embodying non-human body parts in VR is a prevalent practice among certain subcultures, and is a personally important creative outlet to many individuals. However, the discrepant morphology between real and virtual bodies can decrease Sense of Embodiment (SoE). Haptic feedback can compensate by increasing SoE felt towards non-human body parts, but there is a literature gap in comparing the effects of different haptic modalities, and their combinations, on SoE. Through an online survey sent out to social VR communities (n=63), we determined that animal ears are a commonly embodied and ecologically valid non-human body part to study. We then ran a 2x2 within-subjects user study (n=28) with two independent variables: active haptics, delivered through vibrotactile gloves, and passive haptics, delivered through a physical headband, for when participants reach up to touch virtual dog ears appended to their avatar in VR. Our findings show that (1) passive haptics outperform active haptics, (2) combining the two modalities produces an antagonistic effect on SoE, and (3) SoE towards non-human body parts positively correlates with SoE towards the entire avatar. We discuss implications of our findings in various domains, and on embodiment literature. -
1145 How Avatar Appearance Reshapes Internal Representation of Hand Geometry
Avatars whose appearance deviates from that of the physical body are widely used in virtual reality to enable experiences that go beyond biological limitations. To extend prior avatar research beyond subjective and performance-based measures, we systematically examined how variations in hand avatar appearance influence the internal representation of hand geometry. Using a structured set of 26 hand avatars, participants adapted to each avatar and subsequently reported their hand geometry. Results show that internal hand geometry does not change in direct correspondence with avatar geometry and exhibits a higher-dimensional structure than the avatar designs themselves. Furthermore, changes in hand geometry were organized along multiple interpretable dimensions rather than reflecting unsystematic variability. These findings suggest that avatar appearance is reinterpreted through internal bodily constraints. The proposed framework provides a basis for understanding avatar appearance in relation to body representation, supporting more informed avatar design beyond simple geometric mapping. -
1779 Multimodal Guidance under IK-Mediated Embodiment: Topological Limits and Gender-Moderated Effects
Combining, otherwise known as `stacking', auditory and haptic cues with visual guidance is a common strategy for improving online error correction in VR training. However, under sparse 6-DoF tracking with inverse kinematics (IK), multimodal benefits may be limited by solver-induced kinematic ambiguity and may not generalize across users. We investigated this question in a VR motor-imitation study ($N=104$) with four feedback conditions: visual only (V), visual+audio (VA), visual+haptic (VH), and visual+audio+haptic (VAH). Results show that multimodal efficacy is constrained by IK topology and learning-stage trade-offs. Along the kinematic chain, feedback effects were absent at solver-driven joints (elbows), exhibited trend-level divergence at sensor-driven joints (wrists), and were strongest at hierarchy-driven distal segments (fingers). Visuo-haptic guidance (VH) accelerated online adaptation relative to the visual baseline, but this benefit did not carry over to the full stack (VAH). Instead, VAH was associated with lower offline cued recall, indicating a trade-off between execution support and short-term retention. We also observed a distal Condition$\times$Gender interaction --- haptic inclusion improved performance in the male group but degraded performance in the female group under the same mapping, consistent with a confirmation-versus-interference pattern. Based on these findings, we derive three design guidelines for IK-mediated VR training --- Topology-Aware Masking, Interference Detection, and Scaffold-and-Wean Scheduling. -
1793 A Hand Age-Swap Proteus Effect on Memory and Movement
Running virtual reality (VR) user studies with older participants is challenging due to mobility and cognitive limitations prevalent in this age group. Therefore, most VR studies opt to bypass this participant demographic, with the consequence that little is known in terms of whether, which, and under what conditions VR interventions can benefit the older members of our society. This paper presents a first step towards improving age representation in VR, not by actually enrolling older participants in the study, but rather by attempting to modify the cognitive and motor behavior of younger participants for it to resemble that of older participants. The behavior modification relies on the Proteus effect, a well-documented change in the behavior of a user of a VR application when provided with an incongruent self-avatar. A user study asked N = 30 participants in the 18 to 35 age group to play a memory game in VR, twice: once with virtual hands with a young appearance, in line with the actual participant age, and once with virtual hands with an old appearance, in contradiction with the actual participant age. The game board is a torus divided into four buttons of different colors. The buttons light up one at the time in random order and the participant has to recall the sequence and play it back by pressing the buttons in the same order. The sequence grows longer each time it is played back correctly. The results show that when presented with older looking hands participants recalled shorter sequences and their hands did not reach as far as when presented with age-congruent hands. The study brings evidence of simultaneous cognitive and motor manipulation of VR user behavior. Specifically, the study shows that both the cognitive and motor behavior of young VR users can be shifted towards behavior commonly associated with older participants. This opens the door to collecting initial data on possible older participant responses to a VR application using only younger participants, by manipulating their avatar.
Paper Sessions 13-1517:15–18:15
PS13: Applied XR Evaluation Glasshaus
-
1428 User-Centered Evaluation of an AR Virtual Triage Tag: Responder Perceptions and Performance with Virtual and ‘Real’ Patient Actors
Prior work proposed AR tools to facilitate novice emergency responders in MCI triage, concluding with an introductory lab study of the AR tool’s usability with actual emergency responders. Given the high risk and consequential nature of MCI triage, 1) sterile, controlled lab settings are not sufficient in isolation to evaluate the actual scalability of these AR tools, and simultaneously, 2) it is impossible to jump straight to actual field evaluation without further proof of concept. Thus, this work offers additional AR tool evaluation with emergency responders through two contextually embedded simulations—one with virtual patients and one with ‘real’ patient actors situated in a physical space. This paper describes the two simulations created via a partnership with a national training facility for MCI triage (Disaster City), discusses the AR tool’s usability in progressively embedded real context via subjective semi-structured interviews and objective performance measures, and finally compares the effectiveness of the simulations themselves for evaluative purposes. While responders were excited about the AR and found it generally usable, they had broader scalability concerns about things like network connection and security. -
1957 GenAI for Usability Testing AR Apps
Augmented reality (AR) testing can be challenging, especially since the lack of regular user feedback is an issue in many domains. In this paper, we explore Generative AI (GenAI) as a method for testing AR usability without human users. We answer the question of whether GenAI can support usability testing of AR apps and, if so, how to utilize it effectively. We provided the GenAI (here: GPT-5.1) with app key frames under two interaction protocols that differ in the timing of information delivery, illustrating how interaction affects the model's comprehension of the app. We evaluated the workflow by comparing the usability issues generated by the GenAI with those reported by human users. For this, we developed a representative task and an AR app, and conducted 48 evaluations (24 human users and 24 independent GenAI trials). We analysed issue content in terms of emerging themes, three AR experts rated the severity of the issues, and outputs were compared across sources. The results indicate that the GenAI produces more unique issues, but also issues with a significant lower mean severity score compared to issues produced by humans. The primary benefit of GenAI lies in its ability to generate well-articulated usability issues quickly. Beyond the studied task and application, the proposed workflow and interaction protocol offer a basis for transfer to other AR development contexts, but requires validation across additional applications and settings. GenAI is best positioned as a complementary tool for rapid, early feedback, rather than a replacement for human usability testing. -
10.1109/TVCG.2026.3703957 Elucidating the Effects of Augmented Reality Head-Worn Display Factors on Performance, Trust, and Behavior
Augmented reality (AR) display characteristics have the potential to either enhance or impair users' spatial abilities and performance. While previous work included studies of spatial performance with various display factors, evidence for objective performance differences is limited due to compensatory behaviors employed by users and overall behavioral differences. In general, it is challenging to document the effects of display factors on task performance as they depend on users' task behaviors, which in turn depend on users' reliance and trust in the technology, which are also affected by the display factors. In this paper, we present two within-subjects experiments (each N=20) in which we aim to elucidate some of the interrelations between two AR display factors (field of view and visual contrast) with objective task performance and subjective assessments of reliance and trust, while controlling for different behaviors. Participants performed a 360° search-and-selection task in a unique hybrid setup, in which we simulated a controlled task environment by having participants stand inside an immersive CAVE-like space while at the same time wearing a head-worn display that overlaid AR tags over the simulated environment. Specifically, we evaluated three fields of view (43°, 93°, and 143°) and three visual contrasts (0.05, 0.25, and 0.5). We controlled for four different behaviors: AR-Only (only relying on AR), AR-First (prioritizing AR over real world), Real-First (prioritizing real world over AR), and Real-Only (only relying on real world). By controlling for these behaviors, we were able to show objective and subjective benefits of larger fields of view and visual contrast. We illustrate how the controlled behaviors relate to users' subjective reliance and trust in an AR system, and why it is important for researchers and practitioners to understand these subjective and behavioral aspects. -
1019 [EXTENDED] Effects of False Positive and False Negative Early Warnings for Hand Tracking Failures on User Trust and Confidence in Virtual Reality
Early-warning systems (EWS) mitigate hand-tracking failures in Virtual Reality (VR); yet, their effectiveness depends on reliability. Predictive warning mechanisms are prone to errors as they produce false positives (FP) and false negatives (FN). In VR, these errors directly interfere with continuous sensorimotor control and user agency when the hand is the main input channel, potentially affecting user trust and confidence in the system while changing user behaviour. Thus, we conducted a within-subject study (N = 24) comparing a reliable EWS baseline with 50% FPs and 50% FNs. Results showed that both error types reduced system trust; however, FP caused a larger and immediate decline (48.7%) compared to FN (25.7%), which emerged progressively. Self-confidence remained stable across FP and FN conditions but decreased in the absence of warnings. FP led to rapid habituation and reduced compliance, whereas FN increased mental demand. These findings indicate that reliability trade-offs in VR EWS design should be carefully managed, as FPs and FNs affect user experience in different ways.
PS14: XR Authoring Cigno + Auriga
-
1429 Drawing From Life: Interaction Techniques for 3D Drawing From Observation in Mixed Reality
We present our vision for Drawing From Life in mixed reality (MR) along with interaction techniques developed to support observational 3D drawing. While 3D drawing in (semi-)immersive environments has been studied for decades, the work has not emphasized learning the skill of observation, which, in contrast, is foundational to traditional drawing instruction. We demonstrate novel interaction techniques to address this new emphasis, including a multi-viewpoint 3D ``sighting'' technique and proportional ruler inspired by the traditional sighting artists do with a physical pencil or paintbrush. The approach demonstrates how real-world observations of a subject can be mirrored in a virtual Artwork Space that provides a scaffold for 3D drawing with correct proportions. The new tools were deployed in two semesters of a university course, 3D Drawing in eXtended Reality. Results include a characterization of expected sighting accuracy, example observational 3D drawings, analysis of proportional deviation in drawings and methods to compare to ground truth, and rich feedback from student 3D drawers. -
2157 Comparing Pinch and Point Poses for Single Stroke Drawing in Virtual Reality
This paper presents an investigation of bare hand input for drawing single strokes in extended reality. Through two user studies, we compared two hand poses: pinch and point, across three surface types: a physical surface, a virtual surface, and no surface (free space). We evaluated the participant-produced strokes based on geometric and algebraic properties, performance, and subjective experience. Despite the common use of point poses for direct interaction with 2D content, our results showed that the pinch pose outperformed the point pose across most metrics. We discuss the strengths and limitations of each hand pose for bare hand stroke input and provide insights into the role of surfaces in drawing tasks within extended reality. -
1728 SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user’s final edited scenes. A mixed-design user study (N=42) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield highly personalized layouts. Ultimately, SPHERE centers humans to shape machine intelligence, establishing a reliable, governed collaboration framework for immersive spatial design. -
10.1109/TVCG.2025.3640423 DynAvatar: Dynamic 3D Head Avatar Deformation With Expression Guided Gaussian Splatting
Generating high-fidelity, expressive, and realistic 3D head avatars remains a fundamental challenge for immersive applications such as virtual reality, gaming, and telepresence. This task requires not only precise modeling of non-rigid facial deformations but also semantically controllable expression synthesis under diverse viewpoints and motion contexts. We present DynAvatar, a novel framework that integrates expression-guided deformation into the 3D Gaussian splatting pipeline to produce photorealistic and emotionally resonant head avatars. Our method introduces two key innovations: (1) an expression-guided Gaussian deformation module that tightly couples geometric displacement with high-level semantic cues, enabling fine-grained and anatomically meaningful facial animation; and (2) a spatial context embedding mechanism that encodes the canonical position of each Gaussian to preserve semantic coherence and spatial consistency during expression generation. Extensive experiments on both controlled and in-the-wild datasets demonstrate that DynAvatar significantly outperforms state-of-the-art methods in terms of visual realism, expression fidelity, and rendering quality.
PS15: Shared Spatial Cues Sezione 4+5
-
1126 Gaze-Contingent Facial Windows to Support Task-based VR Interaction
Facial expressions facilitate human communication by conveying intent, emotion, and social connection. However, maintaining facial contact is difficult when tasks require sustained visual attention. Virtual Reality (VR) allows facial cues to appear anywhere in the environment, such as near a user’s gaze, which may increase attention to facial cues and strengthen co-presence. Yet, its influence on social experience remains unclear. Self-view is also widely adopted in modern videoconferencing systems, though prolonged exposure can negatively affect users’ psychological states. This research investigates gaze-contingent facial windows that allow VR users to maintain facial contact while performing a collaborative decision task. Three conditions were compared: Baseline (BL) with no facial window, Shared Single (SS) with a partner’s facial window that follows the user’s gaze, and Shared Dual (SD), which adds a self-view next to the partner’s window during mutual attention. Results show that SD increases arousal and reduces perceived psychological distance compared to BL. Gaze results further show that facial windows increase attention to facial cues, with attention to the partner’s face highest in SS, followed by SD and BL, while attention to one’s own face is greatest in SD. These findings confirm that gaze-contingent facial windows influence social experience and attentional allocation in VR collaboration. -
1450 Sharing Roughness with Hand-Outline Visualization to Reduce Sensory Asymmetry in VR Collaboration
In collaborative VR, asymmetric access to haptic hardware creates a critical information gap: tactile evidence remains private to the haptic user, hindering the shared understanding needed for joint decision-making. While prior work has explored crossmodal sensory cues in virtual environments, it remains unclear how such cues should be designed for asymmetric collaboration, where collaborators receive information through different modalities. In our setting, the haptic user feels roughness through fingertip vibration, whereas the non-haptic user relies on vision alone. To reduce this asymmetry, we propose externalizing an object’s tactile state through a glanceable hand-outline visual proxy. Specifically, we examine whether abstract visual roughness cues based on line shape and motion can encode three discrete roughness levels for both haptic and non-haptic users. Two preliminary studies establish a shared visual semantics by identifying visually distinguishable cues for non-haptic users and validating their visuo-haptic correspondence for haptic users. In a main study of a collaborative sorting task, showing this visualization on both users’ hands significantly reduced completion time relative to a no-visualization baseline. Moreover, showing the cue on the non-haptic user’s own hand, rather than only on the partner’s hand, significantly increased perceived contribution and confidence. Together, these findings show that hand-anchored abstract visual cues provide a lightweight means of externalizing object tactile state, reducing information asymmetry without compromising social presence. -
1472 From Asymmetric Guidance to Shared Evidence: Governing the Visibility of Spatial Cues in Multi-User Virtual Reality
In asymmetric multi-user VR, collaborators often rely on private, uncertain guidance that must become shared spatial evidence. Persistent visibility keeps user-placed anchors visible once created, but collapses proposal and commitment, accumulates outdated cues, and can promote uncritical acceptance of uncertain output. We propose a governed visibility policy that structures the transition from proposal to persistence through transient previews, propagated uncertainty, and bounded persistence. Evaluated in a dyadic navigation testbed across three studies, the policy outperforms a persistent baseline: it accelerates alignment, reduces false-positive commitment to unreliable cues, and keeps shared evidence closer to operational sufficiency baselines. These findings show the value of treating visibility not merely as a display setting, but as a coordination policy for managing how uncertain private guidance becomes persistent shared evidence for joint action. -
1524 The Three-Body Problem of Collaboration: Effects of Asymmetric Physical Co-Presence in Augmented Reality
Augmented Reality (AR) enables hybrid collaboration, where some participants share a physical space while others join remotely. However, it remains unclear how the presence of even a single remote participant reshapes the collaborative dynamics of hybrid AR groups. We present a controlled user study examining how spatial configuration influences collaboration in shared AR environments. Thirty-six participants worked in triads to solve a cooperative puzzle task across five configurations, ranging from fully co-located to fully remote, using Meta Quest head-mounted displays. We measured task performance, user experience, social presence, and communication behavior. Task performance remained stable across conditions, indicating that AR can support collaborative task execution even when participants are physically distributed. In contrast, spatial configuration significantly affected the collaborative experience: fully remote collaboration increased workload and reduced pragmatic usability, while perceived co-presence was highest when all collaborators were co-located. Notably, the presence of a single remote participant significantly reduced co-presence to levels indistinguishable from fully distributed collaboration. This reveals a structural asymmetry in hybrid AR collaboration: even a single remote participant substantially lowers perceived co-presence across the group. We interpret this group-level effect as the three-body problem of collaboration.
Paper Sessions 16-1817:45–18:45
PS16: Spatial Audio Orione + Perseo
-
1150 DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics (FOA) from video therefore remains an important but challenging problem. In complex scenes, sound perception depends not only on sound source locations but also on scene geometry, materials, and dynamic interactions with the environment. However, existing approaches only rely on visual cues and fail to model dynamic sources and acoustic effects such as occlusion, reflections, and reverberation. To address these challenges, we propose DynFOA, a generative framework that synthesizes FOA from 360-degree videos by integrating dynamic scene reconstruction with conditional diffusion modeling. DynFOA analyzes the input video to detect and localize dynamic sound sources, estimate depth and semantics, and reconstruct scene geometry and materials using 3D Gaussian Splatting (3DGS). The reconstructed scene representation provides physically grounded features that capture acoustic interactions between sources, environment, and listener viewpoint. Conditioned on these features, a diffusion model generates spatial audio consistent with the scene dynamics and acoustic context. We introduce M2G-360, a dataset of 600 real-world clips divided into MoveSources, Multi-Source, and Geometry subsets for evaluating robustness under diverse conditions. Experiments show that DynFOA consistently outperforms existing methods in spatial accuracy, acoustic fidelity, distribution matching, and perceived immersive experience. -
1355 The influence of room-adapted spatial audio rendering on a holographic calling experience
Augmented reality (AR) telemeetings, also referred to as holographic calls, promise a stronger sense of presence and a more natural experience than regular video calls. Using point-cloud transmission and see-through AR displays, visual rendering aims to create the impression that a remote interlocutor is present in the user's room. Room-adapted, dynamic spatial audio rendering should provide the corresponding auditory impression. However, it is still unclear whether room-adapted spatial audio enhances perceived quality and social presence in AR telemeetings, and how its effectiveness is influenced by practical constraints, such as suboptimal room-acoustic match, increased audio capture distance, and lossy audio compression. In our experiment, pairs of participants engaged in conversations on pre-defined topics using an AR holographic calling system under eight spatial audio rendering conditions, which varied in room-acoustic rendering (no room, matched room, too dry, too wet), audio capture distance (close vs. far), and compression (lossless vs. low bitrate). After each condition, participants completed an eight-item questionnaire, rating audiovisual coherence, audio quality, overall experience, and social presence. The results show that room-adapted spatial audio can significantly enhance overall quality and social presence in AR telemeetings compared to no room acoustic rendering. Surprisingly, we also found improved experiences with relatively large mismatches between rendered and physical room acoustics. Furthermore, increasing capture distance or applying lossy audio compression degraded the experience and eliminated the benefits of room acoustic rendering. These findings provide design guidelines for spatial audio rendering in future AR telemeeting systems: to maximize presence, room-acoustic rendering should be added to high-quality audio from close-microphone capture, even when moderate room acoustic mismatches are expected. -
1547 Switched Reading: Toward Seamless Visual-Auditory Switching When Reading Text in Augmented/Mixed Reality
Augmented/mixed reality (AR/MR) wearable glasses now permit information interaction anywhere, but visual displays can be inappropriate when real-world awareness is essential. We propose \textit{Switched Reading}, a novel interaction framework for reading text in AR/MR that supports switching between visual and auditory modalities as needed. Specifically, we explore two key interaction techniques within this framework: (1) gaze-based voice playback and (2) a correspondence-aware transition effect. We implemented them on an MR headset through a parameter-tuning user test. Next, we conducted a user study (N=16) to investigate the impact of the two techniques on reading performance and overall user experience with simulated modality switching in virtual reality. The results show that the condition combining both techniques was the most preferred among four conditions. Moreover, we found that the gaze-based voice playback reduced gaze offsets when switching modalities and improved reading speed over the baseline condition using scroll position. Finally, we implemented a Switched Reading application for reading while walking and collected user feedback, yielding further design implications for practical use. -
1743 Audiovisual Discrepancies in Large-Scale VR: Impact of Physical Sound Source Distance
Large-scale immersive virtual reality (VR) systems often rely on physical loudspeakers for spatial audio. However, practical constraints such as venue geometry, hardware placement, and multi-user safety requirements frequently hinder the precise co-location of physical sound sources and virtual visual targets. This leads to inherent audiovisual (AV) directional discrepancies. Although AV spatial alignment is extensively studied, the perceptual limits of such discrepancies in real acoustic environments remain poorly understood. In this work, we investigate how physical sound source distance affects the detection threshold (DT) for AV directional offsets. We conducted a within-subject psychophysical experiment using the method of constant stimuli. Horizontal azimuth deviations were systematically manipulated across multiple source distances in a physical acoustic environment. Participants heard the audio directly from physical loudspeakers while simultaneously visually viewing a virtual avatar. Our results indicate that DTs increase significantly with source distance. This trend is accompanied by broader perceptual tolerance and flatter psychometric functions, reflecting a distance-dependent decline in directional localization sensitivity. These findings offer empirical insights into the perceptual limits of spatial audio. Furthermore, they provide actionable guidelines for loudspeaker deployment and AV calibration in large-scale VR systems.
PS17: Social Avatars Sezione 2
-
1321 Beyond Single-Perspective Ground Truth: Systematic Biases and Inter-Perspective Misalignment in Avatar-Mediated Emotion Perception
Affective computing research in social Extended Reality (XR) has largely relied on passive emotion elicitation paradigms and single-perspective annotations, leaving open how emotional states are transmitted and perceived across different observational perspectives in avatar-mediated interaction. We introduce Remote Embodied Improvisation (REI), a dataset collected through active dyadic emotion elicitation in a full-body avatar XR environment, annotated simultaneously from three perspectives (self, interaction partner, and third-party observer) along the Valence, Arousal, and Dominance (VAD) dimensions of emotion. Analysis of inter-perspective agreement reveals a systematic dimension-dependent pattern: Valence and Arousal are moderately to well preserved across perspectives, while Dominance exhibits markedly poor agreement and a consistent observer overestimation bias. Mixed-effects modeling of behavioral cues further demonstrates that the association between each behavioral signal and inter-perspective misalignment varies across affective dimensions: mouth dynamics reduced disagreement in Arousal but increased disagreement in Dominance and Valence, whereas upper-body motion was uniquely associated with disagreement in Dominance. These findings challenge the assumption that affective annotations from any single perspective constitute reliable ground truth, and suggest that Dominance is difficult to transmit through avatar mediation due to its dependence on relational context unavailable to external observers. -
1958 When the Group Plays Along: How Avatar Appearance Commonality and Group Size Influence Drumming Behavior and Group Perception in VR
User behavior in virtual reality (VR) is influenced by the self-avatar's appearance through the Proteus effect, which can be modulated by social comparisons between one's own and others' avatars. However, it remains unclear how these effects are moderated in large-scale groups. Grounded in Social Identity Theory, this study investigated how appearance commonality between the self-avatar and surrounding avatars, combined with group size, influences user behavior. 32 participants performed a VR Taiko drumming task in a mixed design of the following factors: embodying either a festival attire (Happi) or a business suit avatar, surrounding avatars' appearances either matching the participant's attire or wearing diverse casual clothing, and group sizes of 3, 25, and 100 avatars. Results revealed that when surrounded by matching avatars, participants embodying Happi avatars exhibited significantly greater strike intensity, and participants embodying Suit avatars exhibited marginally significantly greater movement synchrony. Furthermore, larger group sizes decreased movement synchrony but increased subjective unity, while matching avatar appearance facilitated perceived group membership over diverse clothing. These findings offer design insights for social VR experiences: matching avatar appearance can enhance group identity and amplify task-relevant behavior when the appearance is contextually congruent with the activity, while increasing group size facilitates felt togetherness and decreasing it boosts behavioral coordination. -
10.1109/TVCG.2026.3668294 Effects of Social Contextual Variation Using Partner Avatars on Memory Acquisition and Retention
This study investigates how partner avatar design affects learning and memory when an avatar serves as a lecturer. Based on earlier research on the environmental context dependency of memory, we hypothesize that the use of diverse partner avatars results in a slower learning rate but better memory retention than that of a constant partner avatar. Accordingly, participants were tasked with memorizing Tagalog-Japanese word pairs. On the first day of the experiment, they repeatedly learned the pairs over six sessions from a partner avatar in an immersive virtual environment. One week later, on the second day of the experiment, they underwent a recall test in a real environment. We employed a between-participants design to compare the following conditions: the varied avatar condition, in which each repetition used a different avatar, and the constant avatar condition, in which the same avatar was used throughout the experiment. Results showed that participants in the varied avatar condition recalled significantly worse during the learning trials on the first day. However, we found no significant difference between conditions in the delayed recall test on the second day. We discuss these effects in relation to the social presence of the partner avatar. This study opens up a novel approach to optimizing the effectiveness of instructor avatars in immersive virtual environments. -
1459 Too Close for Comfort: Continuous Anxiety Decoding from Physiology and Embodied Signals in Social VR
Immersive social VR enables fine-grained manipulation of anxiety-inducing encounters---crowds, proxemic intrusions, and confrontational agents---yet how physiological, behavioral, and environmental signals co-vary with and could decode moment-to-moment anxiety in these settings remains largely unexplored. To address this, here we present a study of N=108 participants navigating five self-paced, multi-agent social VR scenes, under synchronized physiological (EDA, PPG, RSP, eye-tracking), locomotion (speed, head rotation), and proxemics (agent distance, count) recording, with continuous anxiety rated retrospectively. Four converging validity analyses validate this continuous retrospective self-report as a sensitive measure of state social anxiety, with consistent physiological co-variation confirming genuine affective responses. Importantly, we systematically evaluate three settings for anxiety prediction: within-person decoding, cross-participant generalization, and causal prospective prediction. Within individuals, deep sequence modeling (BiLSTM) achieves excellent prediction performance. In the cross-participant setting, however, performance drops substantially. In both settings, combining all modalities yields the best overall performance. In the most challenging prospective prediction setting, we chart how predictability varies across time provided for training and different classifiers. Our results establish a quantitative baseline for continuous anxiety decoding in ecological social VR and highlight cross-participant transfer as a key open challenge.
PS18: Gaze Selection Sezione 1
-
1017 Evaluating Vergence–Accommodation Conflict in Gaze-Based 3D Target Selection
State-of-the-art head-mounted displays (HMDs) enable gaze-based selection in virtual environments. Yet, these HMDs suffer from the vergence–accommodation conflict (VAC), which is known to affect interaction performance. The VAC might influence gaze-based selection performance as it directly affects eye movement behavior. Thus, in this paper, we investigate how the VAC influences \gaze-based 3D target selection across varying depth conditions. Our results show that as the (visual) depth increases, user performance significantly decreases with gaze-based selection. Moreover, a previously-suggested ``\textit{Variation in Diopter}'' Fitts’ law model captured this performance change better relative to a linear model. These findings provide evidence that gaze-based pointing is negatively affected by the VAC and highlight the importance of accounting for depth-dependent factors when designing gaze-based interaction in 3D environments. -
1032 Investigating Depth-Based Target Expansion for Gaze Selection with Dwell Activation in Virtual Environments
While dwell-based selection enables interaction with little physical effort, it is prone to (eye) gaze jitter. Target expansion methods mitigate this problem by increasing the effective selectable target area in 2D environments. Yet depth-related factors may limit their effectiveness in 3D virtual environments due to stereo-deficiencies. In this paper, we investigate how depth affects target expansion methods for dwell-based gaze selection, and propose Depth-based expansion, a method that adapts target expansion based on depth. We conducted two user studies to evaluate our proposed method with an ISO 9241:411 multidirectional selection study and a 3D selection study in a complex environment. Our results show that while the effectiveness of 2D target expansion methods decreases significantly with increasing depth, our proposed method consistently improved user performance and experience. Our findings provide design insights for effective dwell-based selection of expanding targets in 3D virtual environments. -
1111 Dynamic Vergence–Accommodation Conflict: Effects on 3D Selection Performance When Target Depth Changes
The Vergence–Accommodation Conflict (VAC) is well-studied for 3D interaction performance in state-of-the art head-mounted displays. These interaction studies examined VAC in scenarios where the selected target remains at a fixed depth during acquisition. Yet, many VR applications involve targets moving continuously toward or away from the user, causing vergence demand to change over time while accommodation remains fixed at the display focal plane. We study this scenario using an ISO 9241-411 multidirectional target selection task with: noVac (targets at the display focal plane), Constant VAC (targets further from the focal plane), Varying VAC (targets appear at both NoVAC and ConstantVAC depths across selections), and Dynamic VAC (target depth changes continuously during acquisition) conditions. In a within-subject study with 24 participants, Dynamic VAC produced slower selections, higher error rates, lower accuracy, and reduced throughput compared to the noVAC condition. These findings suggest that continuously changing depth during target acquisition introduces additional performance costs beyond static VAC scenarios. -
2059 Eye-vergence Depth Estimation for Hands-free 3D Selection in VR
Eye tracking is a key enabler of natural and intuitive interaction. In virtual reality (VR), gaze depth estimation through vergence offers promise for precise interaction in real 3D space, not limited to the image plane. We propose a set of gaze-based techniques that utilize vergence in various ways for interaction tasks within visually cluttered environments. We conducted a series of experiments to assess accuracy, robustness, and user comfort. We report results from the evaluation of vergence reliability, selection pointing, and selection confirmation methods. Our findings show that, while gaze direction estimation is highly stable, depth precision decreases with distance, requiring adaptive error tolerance. These findings advance the design of gaze interaction systems toward seamless, controller-free experiences in immersive environments.
Thursday – 8 October 2026
PS19 - PS21 08:30–09:30
PS22 - PS24 09:00–10:00
PS25 - PS27 11:45–12:45
PS28 - PS30 12:15–13:15
PS31 - PS33 14:15–15:15
PS34 - PS36 14:45–15:45
Paper Sessions 19-2108:30–09:30
PS19: Embodied Agents Sezione 2
-
1035 Signals of AI Hallucination: Designing Hallucination-Aware Cues for Embodied Conversational Agents in VR
LLM-powered conversational agents (CAs) often present uncertainty and provenance cues alongside their responses to help users assess reliability and detect potential hallucinations. In immersive environments such as Virtual Reality (VR), CAs usually take embodied form as speech-led embodied conversational agents (ECAs), in which hallucination-awareness cues no longer benefit from persistent inline text. Delivered only through speech, these cues are easy to miss and may disrupt comprehension. We conducted a within-subjects study (N = 24) to compare three designs for presenting the hallucination-awareness information (uncertainty and provenance) in ECAs in VR against a no-cue baseline: embodied cues using gestures and posture, icon cues using visual indicators, and text cues using color-coded text with inline citations. We evaluate how these designs affect users' ability to identify hallucination-related information, trust in ECA, and interaction experience (immersion and task load). Our results show that all three designs support users in identifying hallucinations. Embodied cues were associated with higher trust and immersion, text cues offered clearer interpretability, and icon cues provided a middle ground across these qualities. This work contributes to the ISMAR community by comparing different designs of hallucination cues in immersive ECA settings and examining how they affect users' ability and experiences to recognize hallucinations. It also offers practical insights and design implications for developing future hallucination-awareness interfaces for ECA. -
1663 The Effects of an Agent's Clarifying Questioning Behavior During Task-guidance in Virtual Reality
With the rapid growth of artificial intelligence (AI), intelligent agents are becoming increasingly prevalent, guiding people through various tasks. When coupled with virtual and augmented reality devices, these task-guiding agents offer significant potential to train users and assist with performing real-world procedures. Inspired by their promise and imminent ubiquity, we conducted a study focusing on scenarios where an agent may have to ask clarifying questions to its human partner to resolve its uncertainty and proffer correct guidance accordingly. In a mixed factorial experiment, we manipulated the difficulty of the agent's questions and how frequently it asked questions, assessing how task performance levels, subjective perceptions of the agent and workload, and user behaviors were affected. Results suggested that the impact of the agent’s question-difficulty was influenced by how often the questions were posed, with counterproductive effects on task performance levels and perceptions emerging, most notably when difficult questions were asked frequently. While frequent questioning can degrade task-performance, it appears to be somewhat tolerated when the demands involved in providing clarifications are low. We discuss how task-guiding virtual agents can question users when uncertain and need clarifying information. -
2201 Self-Resemblance and Activity-Aware Responses Shape Relational Openness and Reflection in XR Companion Interaction
As virtual companions increasingly populate extended reality (XR) environments, they are evolving from simple interactive interfaces into socially co-present entities that accompany users during everyday activities. This shift raises broader questions about how meaningful relationships between users and digital agents can emerge in physically situated XR interaction. Prior research in psychology suggests that people tend to disclose more personal thoughts and emotions when relational closeness is established and when conversational partners demonstrate awareness of their behaviors or experiences. These findings imply that both relational identification and behavioral responsiveness can encourage users to move from simple expression toward deeper self-disclosure. In XR environments, such relational cues may be conveyed through how companions embody identity-related signals and how they perceive and respond to users’ ongoing actions. To address this gap, we introduce a modular XR companion framework that independently manipulates two design dimensions: self-resemblance and activity-aware responsiveness. We evaluate these dimensions within a physical–virtual expressive interaction context using LEGO flower arrangement, where users’ affective states are naturally externalized through open-ended physical construction. Our findings reveal that self-resemblance forms the relational grounding necessary for self-disclosure, while activity-aware responsiveness primarily supports reflective engagement once relational grounding is established. This work reframes XR companion design from enhancing presence to cultivating relational openness and provides design principles for supporting meaningful user–agent relationships in physical–virtual environments. -
2222 VAGENT: A Vision-Language Embodied Agent for Attention Awareness and Guidance in Virtual Reality
Effective and engaging communication and collaboration in virtual reality (VR) require not only understanding others’ attentional cues, but also guiding attention through one’s own behavior. Although prior work has explored LLM- or VLM-driven agents for VR interaction, it remains unclear how embodied agents can both interpret users’ nonverbal attentional cues and actively guide user attention in spatially situated interaction. To address this gap, we present VAGENT (Vision-Language Attention-aware and Guidance Embodied ageNT), a vision-language embodied agent framework that integrates attention awareness from multimodal user cues and spatial understanding with attention guidance through coordinated verbal and nonverbal behaviors grounded in VR spatial context, helping users align their attention more effectively during conversation and collaboration. In a within-subject controlled study with 54 participants, we compared VAGENT with a baseline embodied agent that had access to scene spatial information but no attention awareness or guidance. Across two representative VR communication and collaboration tasks, VAGENT led to more target-oriented gaze behavior, significantly improved social presence, communication satisfaction, collaborative experience, and trust. Beyond these gains, our behavioral analyses revealed that when multiple non-verbal attention cues coexist, users tend to prioritize pointing gestures over head movements and eye gaze for attentional guidance.
PS20: Teleoperation Orione + Perseo
-
1221 On-Demand Gaze-Triggered SSVEP: Toward Comfortable Brain–Computer Interaction in AR/MR Environments
Brain–computer interfaces (BCIs) offer a promising pathway for hands-free input in augmented and mixed reality (AR/MR), yet the adoption of steady-state visual evoked potential (SSVEP) systems in head-mounted displays remains limited by a fundamental usability barrier: the persistent flickering of multiple targets causes visual fatigue that is unacceptable for prolonged use. We propose an on-demand gaze-triggered paradigm in which flickering stimuli activate only when the user gazes at a designated menu region, replacing the conventional always-on stimulus presentation. To isolate the effect of stimulus timing from display-specific confounds, we conducted a controlled study with 14 participants (324 trials per participant) on a 27-inch LCD, comparing three onset strategies: immediate onset, fixed delay (500 ms), and adaptive delay calibrated to individual gaze behavior. Classification accuracy (TRCA: 82–84%) was comparable across all three conditions, confirming that on-demand timing does not compromise decoding performance. The 8 Hz stimulus consistently outperformed 9 Hz and 11 Hz regardless of spatial position. A preliminary layout preference study identified the triangular target arrangement (M=4.29/5) as significantly preferred over horizontal (M=3.29, p < 0.05) and vertical layouts (M=2.71, p < 0.01, Cohen's d=1.89); this preferred layout then served as the spatial configuration for the paradigm validation. To assess the feasibility of the proposed paradigm in a wearable setting, we present the adaptation to and preliminary observations from an optical see-through AR prototype (Jorjin J7EF Gaze). Our results demonstrate that on-demand activation can replace always-on flickering without accuracy loss, offering a more comfortable interaction model for SSVEP-based BCIs in AR/MR environments. -
1369 Understanding User Adaptation to Network Impairments and Training in Immersive VR Teleoperation
Network impairments such as delay, jitter, and packet loss remain a major challenge for streaming and remote-rendered Virtual Reality (VR) systems. While prior work primarily focuses on system-level mitigation, less is known about how users behaviorally adapt their control strategies under degraded network conditions. We investigate user adaptation in a high-precision VR teleoperation task under controlled network perturbations through a pick-trace-place task. Thirteen participants performed the same manipulation task across three training days under three impairment conditions—delay, jitter, and packet loss—each with six severity levels. From trial-level trajectory features, we infer latent control strategies using Gaussian mixture models and categorize them post hoc as Continuous, Move-wait, and Predictive. Our results reveal that impairment severity shifts strategy composition in condition-specific ways, while repeated practice consistently drives users toward Continuous control. Within packet-loss conditions, we further identify interpretable sub-strategies within the Predictive strategy and observe systematic reweighting across training days. These findings provide a behavior-centered account of how users adapt to network-impaired interaction in VR and highlight the importance of strategy-aware system design for teleoperation and remote-rendered VR applications. -
1385 Look to Move and Grasp: A Gaze-Only VR Framework for Robot Teleoperation
While robots hold immense potential, conventional teleoperation interfaces predominantly rely on manual input, creating substantial barriers for users with severe motor impairments or in hands-busy scenarios. Gaze interaction in virtual reality (VR) offers a natural, hands-free alternative that eliminates the need for physical movements. In this work, we propose a gaze-only VR teleoperation framework designed to provide an accessible, hands-free interaction paradigm. The framework incorporates three gaze-based methods for robot locomotion and two methods for robot manipulation. To systematically evaluate the usability, interaction efficiency, and user experience of the proposed approach, we conducted a user study. The results demonstrate that the gaze-driven methods enable effective control of both robot locomotion and manipulation, exhibiting favorable ease of use and feasibility in teleoperation tasks. This study validates the feasibility of gaze-driven robotic teleoperation within virtual reality environments, providing a promising hands-free solution that lowers physical barriers and contributes toward the development of more natural and inclusive forms of human–robot interaction. -
1781 Exploring Cross-Reality Transitions between Projections and Head-Mounted Displays for Immersive Digital Art
Immersive exhibitions increasingly combine large-scale projection displays and mixed reality (MR) head-mounted displays (HMDs), yet their integration into a coherent experience remains underexplored, particularly in terms of how users perceive transitions across heterogeneous visualization environments. This paper investigates cross-reality (CR) object- and scene-level transitions between projection and an MR HMD through a hybrid immersive art installation spanning projection, augmented reality (AR), and virtual reality (VR). In a within-subjects study (N=24), we compare a calibrated condition with a deliberately degraded condition that introduces bundled cross-display inconsistencies in spatial alignment, color, and latency. Rather than using this contrast only to measure a decline in experience quality, we use it as a probing strategy to make transition-relevant cues more perceptible and discussable. Quantitative results show that the degraded condition lowers presence and increases workload, while post-session interviews further reveal how disruptions in spatial alignment, color consistency, and system latency affected participants' experience of continuity across transitions, with these effects varying across asset types. These findings provide empirical insight into how users perceive projection--MR transitions and inform the design of coherent hybrid immersive art experiences.
PS21: Hand Gestures Sezione 1
-
1100 Switching Between Worlds: A Comparative Study of Hand-Tracking Gestures and System UIs for XR Application Transition
As virtual reality (VR) systems mature, users increasingly interact with multiple immersive applications across virtual environments (VEs), making program switching a crucial requirement in XR systems. In this work, we investigate how hand-tracking-based XR systems can support minimally disruptive application switching. We model program transition via hand tracking as a two-stage process: an activation gesture followed by a lightweight graphical interface for selecting running applications. We thus explore two design factors: Interface (Grid vs. Scroll vs. Dial) and Gesture (Static vs. Dynamic), using a within-group user study (N = 26) to evaluate switching efficiency, workload, accuracy, and subjective user experience. Results show that interfaces that enable direct selection outperformed those that support sequential browsing. Dynamic gestures increased efficiency and perceived engagement but also required greater physical effort and a higher workload. Our findings highlight trade-offs between gesture activation and interface structure, providing design guidance for OS-level application-switching mechanisms in XR systems. We open-source our implementation at: [link is omitted due to the double-blind review process]. -
1521 Did I Do It Differently? Comparing Gesture Elicitation in Virtual and Augmented Reality
Gesture elicitation is a common way to map commands, e.g., actions or referents, to gestures by reaching an agreement between participants. However, it is uncertain how gestures elicited in Virtual Reality (VR) transfer to Augmented Reality (AR) and vice versa. Using a wind-simulation virtual environment that supports gesture interaction with concepts, we aim to explore and compare gesture-elicitation results from participants in AR and VR for the 33 referents presented. By analyzing gesture elicitation for different types of referents across the two technologies within five referent categories, this study shows that, except for the translation category, participants' gestures in AR and VR were highly similar. The specific result for the translation category is likely due to its reliance on the physical surface. This also provides a foundational context for deeper understanding in future research. -
1223 Tackling Open Set Gesture and Voice Interaction in 3D with Runtime Code Generation
Creative problems like 3D modeling and scene-editing are naturally open set problems since a deployed system will likely encounter novel situations due to user creativity that designers failed to anticipate beforehand. However, advancements with large language models (LLMs) mean that systems can infer user intent and enable runtime code generation to define novel interactions. We propose a method called GestCoGen that provides on-demand generation and execution of application-agnostic 3D scene actions based on user speech commands and hand gestures. GestCoGen enables flexible scene interactions and promotes fluid sequencing of actions using real-time speech and gesture recognition. In a user study, we compare GestCoGen to a traditional menu-based system in a 3D application. We find that GestCoGen promotes novel scene interactions that concatenate multiple discrete actions with a single articulation of gesture-speech command. To help analyze and describe this observed phenomenon, we introduce the concept of the equivalent command chain. The equivalent command chain enables a direct comparison between low-level menu interactions and the concatenation approach promoted by GestCoGen. -
10.1109/TVCG.2026.3690947 SelfBlending: Artificial Intelligence-driven Augmentation with Hand Interactions for Seamless Reality Blending in Virtual Environments
Accessing real-world objects during immersive virtual reality (VR) experiences remains challenging, as current cross-reality systems often rely on predefined interaction steps, tracking devices/markers, or fixed object setups. They also lack support for personalized object recall, where users can add, remove, or modify real-world items blended into the virtual environment (VE). Many head-mounted displays (HMDs) include passthrough technology to switch between virtual and real worlds, but it often disrupts immersion by requiring a full shift from virtual to real. Thus, maintaining an optimal balance between virtuality and reality is difficult. To address these challenges, we developed SelfBlending, a framework that uses AI-based hand tracking to let users label physical objects through freehand gestures, then blends the selected item into the VE using object recognition, enabling interaction with the relevant real-world object. SelfBlending was evaluated against two common interaction conditions: the default passthrough feature in VR HMDs and the conventional approach of physically removing the HMD to access real-world objects. Results from seated, single-object interactions with tabletop-placed items showed that SelfBlending enhanced user experience by boosting presence, supporting efficient physical interaction, and improving cross-reality continuity. It also enabled selective interaction with real objects while minimizing the disruption of VR experience.
Paper Sessions 22-2409:00–10:00
PS22: XR Infrastructure Cigno + Auriga
-
1307 SemanticXR: Low Power and Real-time Queryable Semantic Mapping with an Object-Level Distributed Architecture
Semantic mapping is a core service that enables grounded interactions in emerging Extended Reality (XR) applications such as AI assistants and spatial object search. Deploying this capability on mobile XR devices requires a system that is open-vocabulary, real-time, and low-power. Existing approaches are compute-intensive and assume server-class resources. Cloud offloading offers a practical path, but no existing system implements distributed semantic mapping, and current approaches do not address how to manage communication, execution, and memory footprint across a device-cloud boundary. We present SemanticXR, the first device-cloud system for real-time, open-vocabulary semantic mapping and querying under XR power, bandwidth, and memory constraints. Our key insight is to elevate semantically identifiable objects to first-class units of distributed system design, governing how the system communicates, executes, and manages memory across the device and the server. On the server, object-level parallelism and geometry downsampling improve mapping latency, while object-level depth-mapping co-design reduces upstream bandwidth. On the device, an object-level sparse local map with incremental updates and update prioritization enables network-robust querying with bounded memory and downstream bandwidth. Object-level configurable resource usage vs. quality trade-offs allow applications to adapt the system to their needs. Evaluation against a distributed baseline using the same perception models shows that object-level system organization improves server-side mapping latency by 2.3×at equivalent semantic quality. Object-level depth-mapping co-design maintains upstream bandwidth under 2.5 Mbps. On the device, SemanticXR enables sub-100 ms query latency even under network drop while supporting thousands of objects within 500 MB, and scales downstream bandwidth with map changes rather than total scene size. The system adds only 2% device power during normal operation. -
1960 DELUGE: Decomposed Entropy-coded Live Unstructured Geometry Exchange for Real-time Particle Streaming
Particle-based physics simulations, including fluids, smoke, and granular media, are fundamental to visual realism in immersive VR and AR. With the growing adoption of social VR and digital twins, demand is increasing for shared experiences in which multiple users interact with the same simulation in real time. Realizing such experiences requires low-latency streaming of large-scale particle data from a server to each client, yet existing point cloud compression methods such as G-PCC (TMC13) and Draco assume static geometric structures; when applied to dynamic particle streaming, their encoding latency exceeds the frame period, failing to meet real-time delivery requirements. We propose DELUGE, a streaming compression architecture that exploits the temporal coherence and velocity predictability inherent in physics simulation particles, achieving sub-frame-latency encoding and decoding through three complementary techniques. Evaluation on dynamic point cloud datasets demonstrates that DELUGE achieves 19 times faster decoding than G-PCC (TMC13) and 6 times faster encoding than Draco, when all codecs are measured in-process via FFI. We further build end-to-end client implementations for both web browsers and Apple Vision Pro, and confirm through a within-participants perceptual quality evaluation and a two-person collaborative task study on Vision Pro that the proposed method supports real-time collaborative experiences with hand-tracked fluid interaction. -
2130 IMRSIVE: Inverse Mixed Reality for Autonomous Systems Interacting in Virtual Environments
Testing autonomous robots in hazardous scenarios presents a fundamental dilemma: pure simulation cannot capture real-world physical dynamics such as surface friction and center-of-gravity effects, while physical test environments are dangerous, expensive, and slow to reconfigure. We present IMRSIVE, an inverse mixed reality framework that enables physical robots to navigate through virtual hazards while preserving real-world physics. The system provides dual synchronized perspectives: a first-person augmented reality view rendered on an AR headset for the autonomous agent, and an extended third-person spectator view composited from a live camera feed for external observers. To achieve correct visual compositing in the spectator view, we introduce a stencil buffer inverse occlusion pipeline that renders motion capture-tracked physical objects behind virtual geometry in real time. To our knowledge, this is the first framework to unify first-person AR and third-person composited perspectives for autonomous system evaluation. We evaluate the system through sustained rendering performance analysis, end-to-end latency measurement, occlusion quality assessment, and a parametric physical fidelity experiment in which an 11.43 cm center-of-gravity shift under identical control inputs produces trajectory deviations up to 79.6 cm, amplifying the input perturbation by a factor of seven. These results demonstrate that the framework captures physical phenomena that conventional simulation cannot predict, while maintaining real-time performance suitable for interactive mixed reality. -
10.1109/TVCG.2026.3695303 Grand Challenges in Cross Reality
Cross Reality (CR) is a new emerging field based on the current developments in Mixed Reality hardware, especially supported by the broad market penetration of video-based see-through Head-Mounted Displays. It refers to applications that span across different stages (real, Augmented Reality, Augmented Virtuality, Virtual Reality) of the reality-virtuality continuum, where users are interconnected between different stages and/or are able to transition between these stages. This publication follows the concept of other grand challenges publications and reflects the discussion of various researchers invested in CR. After an initial discussion at the 1st Joint Workshop on Cross Reality at IEEE ISMAR 2023, six topic groups have been identified, leading to 22 challenges, which were discussed in groups over the period of multiple months. The discussion of these challenges should act as a road map for future research in the area of CR.
PS23: Near-Eye Optics Glasshaus
-
10.1109/TVCG.2025.3617940 Comparison of User Performance and Experience between Light Field and Conventional AR Glasses
Light field AR glasses can provide better visual comfort than conventional AR glasses; however, studies on user performance comparison between them are notably scarce. In this article, we present a systematic method employing a serial visual search task without confounding factors to quantify and compare the user performance and experience between these two types of AR glasses at two different viewing distances, 30 cm and 60 cm, and in two modes, purely virtual VR mode and virtual-real integration AR mode. The results show that the light field AR glasses led to a significantly faster reaction speed and higher accuracy than the conventional AR glasses at 30 cm in the AR mode. The participant feedback also shows that the former led to better virtual-real integration. User performance and experience of the light field AR glasses remained consistent across different viewing distances. Although the conventional AR glasses had a better search efficiency than the light field AR glasses at 60 cm in both AR and VR modes, it had more negative feedback from the participants. Overall, the design of this experiment successfully allows us to quantify the effect of VAC and underscores the strength of the evaluation method. -
1140 Effects of Vergence–Accommodation Conflict Magnitude and Direction on Near-field Augmented Reality Depth Perception
Several augmented reality (AR) applications (e.g., AR-assisted surgery) require users to judge the depth of AR objects relative to real objects at near-field distances. Although near-field AR depth perception tasks demand high accuracy and precision, the perceived depth of AR objects is not veridical when compared to real objects. In addition, the vergence–accommodation conflict (VAC) can potentially lead to distorted visual perception, increased visual fatigue, eye strain, and other adverse effects when using AR devices. In this paper, we present the first systematic study examining how varying magnitudes and directions of VAC affect the accuracy and precision of near-field AR depth perception, using a custom-built AR Haploscope. Two focal planes were examined: 2D and 2.25D. Five levels of VAC magnitude and direction were tested: -1.0D, - 0.5D, 0D, +0.5D, and +1.0D. A perceptual depth-matching task was carried out with 24 participants. Participants subjectively rated various visual discomfort symptoms related to perceived depth. Our results showed that as the value of VAC magnitude increased, the accuracy and precision of near-field AR depth perception declined. In line with the theory of human perception under VAC, we observed that the perceived depth of the AR object was biased toward the focal plane, was underestimated when positioned behind the focal plane, and was overestimated when positioned in front of the focal plane. Our findings indicated that both the accuracy and precision of depth perception, as well as the degree of visual discomfort with near-field AR objects, were influenced by the focal plane position. Overall, our research will advance our understanding of near-field AR depth perception and expand the currently limited empirical evidence on how VAC influences near-field AR depth perception. -
1798 Impact of Inter-Pupillary Distance, its Discrepancy with Inter-Camera Distance, and Viewing Methods on Near-Field Size Perception in eXtended Reality
In Video-See-Through (VST) Head-Mounted Displays (HMDs) the cameras are fixed at a static distance which causes discrepancy with user's inter-pupillary distance (IPD). This discrepancy could differentially affect users with varying IPDs. To examine how this discrepancy affects size perception in near-field XR, a mixed-factorial study was designed, with Virtual Reality (VR), VST - Augment Reality (AR) and VST-Real viewing conditions manipulated between-subjects. Participants, with IPDs ranging from 58 mm to 70 mm, visually observed and provided size estimates of 9 different target discs of graspable sizes. Data from a large study involving 150 participants was gathered, providing us with 6750 data points for analysis of perceptual accuracy. The results revealed that compared to VR, in VST-AR, larger discrepancy involving wider IPDs resulted in more pronounced underestimation, especially for larger target objects. Also, when the interplay between IPD and the IPD-ICD discrepancy was considered, augmented targets were underestimated more than the physical targets. Over the course of trials, participants' size estimations tended to improve despite the variation in IPDs, suggesting perceptual calibration over time. These shed novel insights on the impact of IPD and IPD-ICD difference, and have the potential to guide the future research and design of hardware and software to facilitate improved size perception. -
1482 Egomotion Amplifies Vertical Disparity Discomfort in See-Through Augmented Reality
The rapid expansion of Augmented Reality (AR) technologies has increased the need to ensure visual comfort during extended, real-world use. In optical see-through AR systems (OST-AR), users perceive the real environment without mediation, whereas inaccuracies in the rendered virtual content can lead to perceptual conflicts with the real-world view. One such issue is Binocular Vertical Misalignment (BVM), which adds vertical disparity to augmented content while leaving the real-world view unaffected. Existing research on BVM relies on restricted stereoscopic displays and static stimuli, limiting ecological validity for real-world AR applications. In this study we investigate the effects of BVM directly on a see-through AR device, during prolonged (30-minute), interactive usage sessions. Our methodology incorporates individually calibrated BVM thresholds and evaluates discomfort across two distinct scenarios: a search and navigation task involving ego-motion and a stationary manipulation task. Visual discomfort was assessed using questionnaires and in-task Likert-scale ratings collected throughout the session. Experimental results indicate that perceived discomfort scales with a user’s individual tolerance threshold rather than the absolute magnitude of misalignment. Notably, the navigation task significantly exacerbated discomfort compared to the manipulation task, highlighting the critical role of ego-motion in BVM perception. These findings suggest that visual comfort in AR is a multi-faceted interaction between hardware-induced disparity, user tolerance, and specific task demands. This work provides foundational insights for designing more comfortable, long-term wearable AR experiences and informs calibration requirements for OST-AR displays.
PS24: Accessible Navigation Sezione 4+5
-
1122 Poros: Perception-Enabled Interactive Assistance for Appliance Use by Blind and Low Vision Users
Modern home appliances pose significant accessibility barriers for individuals with visual impairments, driven by increasing functional complexity and minimalist design trends. Drawing on formative interviews with blind and low vision (BLV) participants, we propose Poros, a novel system providing perception-enabled interactive assistance to facilitate seamless appliance interaction. Our approach features an automated instruction generator that models appliance operation logic directly from user manuals to synthesize executable guidance for arbitrary tasks. Integrated into an Augmented Reality (AR) wearable, Poros provides real-time auditory feedback synchronized with the user's hand gestures and the appliance's display state. A user study with quantitative analysis and qualitative insights shows that our approach enhances user independence, enabling the successful execution of complex, multi-step operations that were previously considered inaccessible. -
2110 How VR Systems Impact Virtual Navigation for Blind People
Virtual Reality (VR) offers immersive experiences through embodied interactions like head tracking, going beyond traditional controls like joysticks or keyboards. Prior research has studied the impact of VR affordances on people’s experience and performance, but exclusively through sighted people's perspectives. While blind people may also leverage the physical nature of VR interactions, there is no knowledge on how VR compares with traditional controls in nonvisual tasks. We conducted a study where 26 blind participants performed navigation tasks with two conditions: VR and Game Controller. Additionally, as virtual environments can offer various degrees of complexity, we observed how participants performed under different cognitive demands. Our findings showed that VR improved navigation performance, whereas adding a simultaneous cognitive task decreased performance in both conditions. In VR, participants benefited from direct rotation, but with game controllers, the added complexity of rotating often led to reliance on translational 2D movements (e.g., moving sideways). -
2177 Navigating the Last Mile: Evaluating Head- and Cane-Mounted Cameras for Egocentric Spatial Awareness
Robust navigational guidance is an important XR application for both sighted and non-sighted populations. In this paper, we mainly focus on blind pedestrians, who continue to face “last-mile” challenges such as locating entrances and navigating cluttered spaces. While smartglasses and wearables are maturing, a foundational design question remains underexplored: where on the body should cameras be placed to best support navigation? We present a mixed-methods investigation that focuses on the question of optimal sensor placement for generating spatial data supporting ego-centric navigation. A survey of 10 blind cane users surfaced practices for last-mile navigation and perceptions of body-mounted XR devices. A controlled case study with a blind co-author compared head- and cane-mounted cameras using synchronized Project Aria glasses while traversing five real-world environments. Using Simultaneous Localization and Mapping (SLAM) and Neural Radiance Fields (NeRFs) as diagnostic probes, we find that head-mounted cameras offer stability, cane-mounted cameras capture complementary ground-level detail, and fusing both yields increased robustness in scene reconstruction. We synthesize these findings into design guidance for hybrid XR systems that extend the cane without interfering with tactile and auditory cues. -
2190 Identifying VR Accessibility Barriers: A Mixed-Methods Study of Controllers, Hand Tracking, and 360-Video
Despite the rapid growth of Virtual Reality (VR), many applications remain inaccessible to individuals with disabilities. This study addresses this gap by identifying accessibility barriers and proposing inclusive design practices through the evaluation of three interaction modalities: controllers, hand tracking, and passive 360-degree video. Using a Meta Quest 3 headset, 37 participants, including 19 individuals with motor, visual, auditory, or neurological impairments, tested the First Steps with Hand Tracking application. A mixed-methods approach was employed, combining usability evaluation with the SUS, NASA-TLX, a custom accessibility questionnaire, and qualitative feedback. The results revealed significant accessibility barriers across the tested modalities. Findings show that hand tracking presents more barriers than controllers. Furthermore, qualitative data highlighted critical issues such as sensory overload, rigid interface heights, insufficient haptics for deaf users, and camera instability affecting wheelchair users' balance. The research demonstrates that even introductory VR tutorials contain accessibility limitations, concluding that comprehensive personalization of interfaces and input methods is essential for true inclusivity.
Paper Sessions 25-2711:45–12:45
PS25: Task Reliability Glasshaus
-
1523 AR on Ladders: Augmented Reality Tasks on Instability While Working on a Ladder
The use of augmented reality (AR) is growing in industrial applications such as within the construction industry [1, 2]. While AR provides convenience such as providing hands-free displays and presenting time-sensitive information, it may also introduce unintended adverse consequences. For example, a potential application of AR within the construction industry is when working on a ladder conducting installation, inspection, maintenance, and repairs. Falls off ladders are already a major source of workplace injuries [3], and any adverse effects of AR on balance while working on a ladder could exacerbate the risk of falling. Therefore, the purpose of this study was to investigate the effects of AR on balance during simulated construction tasks while on a ladder. The four tasks investigated involved: (1) a reading task with AR input, (2) an assembly task with AR input, (3), an assembly task with verbal input, and (4) a button-poking task with AR input. We also investigated screen-relative and world-relative AR display styles. Balance was assessed by measuring mean speed of postural sway of the center of pressure using a force plate under the ladder, with increases in sway speed from a baseline condition involving standing as still as possible on the ladder being compared between tasks and display styles. Regarding the four tasks, results indicated postural sway during the assembly task with verbal input (6.39 cm/s; p<0.05), assembly task with AR input (6.33 cm/s; p<0.05), and button-poking task with AR input (7.35 cm/s; p<0.05) all increased more than during the reading task with AR input (2.90 cm/s). Additionally, postural sway increased more during the button-poking task with AR input than the assembly task with AR input. Regarding the two display styles, postural sway with a world-relative AR display (9.27 cm/s) increased more than with a screen-relative AR display (4.95 cm/s) during the poking task with AR input (p < 0.05), but did not differ between display styles during the assembly task with AR input (p > 0.05) or the reading task with AR input (p > 0.05). Our results imply that care must be taken when designing AR applications for construction applications, as the type of task to be completed and the display style can affect a worker’s balance. Further research into input types (gestures vs. voice commands) and tasks involving AR menu navigation is proposed. -
1594 ReAlign: Closed-Loop Support for Continuous Virtual Reality Tasks
Continuous virtual reality (VR) tasks require sustained interaction, but human task stability naturally fluctuates. We frame support in these environments as a closed-loop policy problem, using passive head-mounted display (HMD) signals to estimate task-directed stability and trigger brief, response-free cues. We instantiate this in the ReAlign pipeline, an engine-aware system driven by gaze and head kinematics. In a within-subjects study (N=48), we compared an Adaptive policy against a Periodic policy across three evaluation axes: event-locked recovery, efficacy maintenance across repeated exposures, and bounded false-alarm cost. The tested Adaptive policy was associated with more favorable event-locked recovery, flatter efficacy decay, and lower perceived disruption. Furthermore, catch trials delivered during high-stability periods incurred a reaction-time delay that remained within a preregistered equivalence bound. These findings suggest that the tested closed-loop policy can serve as lightweight task scaffolding, highlighting the value of evaluating adaptive systems through integrated policy-level effects rather than average performance alone. -
10.1109/TVCG.2026.3715380 Signal or Noise? How User Movement and Sensor Placement Impact Physiological Data Collection in 6DoF Virtual Reality
To evaluate user experience in Virtual Reality (VR), researchers are increasingly using physiological data, such as electrocardiograms (ECG) and electrodermal activity (EDA), in conjunction with questionnaires. However, in 6-Degrees-of-Freedom VR, user movements and interactions can introduce noise on captured signals, making them potentially unusable. This noise depends on the movements performed and on sensor positions. Unfortunately, no study has addressed this issue, significantly hindering the effective use of physiological data in VR. To address this gap, we conducted a user study to evaluate the effects of sensor positions and user movements on signal quality, as well as sensor comfort and intrusiveness. We compared three positions for EDA and two for ECG on 36 users, who performed a series of five gestures. Results indicate that positioning EDA sensors on the palm offers the best balance between signal quality and user comfort. For ECG, a position closer to the heart is better. Most tested movements affect quality, with the worst being when users close their hand (for ECG and EDA) and move their arm (for ECG), where sensors are located. Based on these findings, we provide recommendations for the effective use of EDA and ECG sensors in 6DoF VR. -
10.1109/TVCG.2026.3704898 Revenge of the Sick: A Meta-Analysis of Washout Periods in Cybersickness Research
Cybersickness remains a major barrier to the widespread adoption of virtual reality (VR) technologies, motivating researchers to investigate its causes and mitigation strategies through comparative human-subjects studies. These experiments may employ either within- or between-subjects designs. Although within-subjects designs require fewer participants and potentially offer higher statistical power, researchers need to wash out carryover effects of cybersickness symptoms before each session to avoid confounding the results. Prior studies have employed washout periods of varying lengths; some required testing conditions on different days, while others allowed only short breaks of less than 15 minutes. Although shorter washout periods are more convenient for experimenters, their impact on study outcomes has not been systematically investigated. In this work, we conducted a meta-analysis to evaluate the effects of study design and washout period length on cybersickness self-reports measured with the Simulator Sickness Questionnaire (SSQ). We found that short washout periods reduce statistical power compared to long washout periods or between-subjects designs. Based on these findings, we provide guidelines to help researchers design more reliable cybersickness studies and improve the generalizability of their results.
PS26: Cybersickness Mitigation Cigno + Auriga
-
1246 Gaze-Contingent Counter-Vection Noise for Cybersickness Reduction
Cybersickness remains a common side effect of many immersive experiences, occurring particularly when fast, virtual-only movements induce high levels of optical flow. These visual movements can induce the perception of self-motion, denoted vection, which leads to a neurological sensory mismatch with the balance system, causing the body to react with sickness. Many techniques have been proposed to reduce cybersickness, aiming either to align the signals of both sensory systems or to suppress the motion perception of one system. One of the simplest, unobtrusive, and popular techniques is foveated blurring, which applies a gaze-contingent low-pass filter to the visual field. However, while the filtering removes high-frequency motion information responsible for high vection, the preserved low-frequency information of the content still carries significant motion cues. Therefore, the effectiveness of the approach varies depending on the visual content. In this paper, we introduce a novel method to enhance foveated blurring by integrating Gabor noise to actively counteract optical flow. Specifically, by leveraging the wave-like characteristics of Gabor functions, we dynamically shift their phase to generate localized motion signals. When these functions are oriented against the scene's optical flow, they induce a ``counter-vection'' effect. Through spatial pooling, this conflicting motion information cancels out primary sickness-inducing stimuli without degrading the user's overall motion perception. In our real-time implementation, the noise is fine-tuned to induce most counter-vection without being actively perceptible to the average user. In a naturalistic VR experiment, we validate the effectiveness of the approach and show that it can significantly improve the performance of regular foveated blurring. -
1700 Towards General Motion Sickness Reduction via Semi-random Galvanic Vestibular Stimulation
Cybersickness remains a major barrier to the widespread adoption of virtual reality (VR) across professional domains and public acceptance. This work explores semi-random Galvanic Vestibular Stimulation (GVS) as a scene-independent method to mitigate cybersickness by reducing vestibular sensitivity and, consequently, visual-vestibular conflicts. In contrast to prior approaches that align vestibular input with scene motion (e.g., omnidirectional GVS), our method introduces stochastic vestibular disruption without relying on any motion data. We developed a custom VR application with active user control to evaluate semi-random GVS in comparison with a state-of-the-art technique. Results show that semi-random GVS not only significantly reduces cybersickness but also matches the effectiveness of scene-dependent GVS in reducing negative symptoms. Thereby, its independence from the visual content makes semi-random GVS a robust and broadly applicable solution for both VR and real-world motion sickness scenarios. -
1961 Feeling Visual Motion: Mitigating Cybersickness in VR Using Tactile Motion Cues
Despite rapid advancements in VR, cybersickness remains a critical barrier to comfortable user experiences. Cybersickness is often attributed to sensory conflict between perceived visual self-motion and the lack of corresponding vestibular feedback. We hypothesized that providing tactile motion cues as an alternative for vestibular stimuli could mitigate this conflict. To test this, we conducted an experiment using a haptic vest to deliver torso-based tactile motion cues synchronized with visual rotations in VR, focusing on yaw rotation as a first step. We evaluated sickness levels and UX across different directions (congruent and incongruent) and velocities. Results indicated that tactile motion opposing the visual direction--mimicking natural vestibular feedback--significantly reduced cybersickness and improved UX. Interestingly, tactile cues congruent with the visual direction also yielded positive outcomes when delivered at twice the visual velocity. These findings suggest that illusory tactile motion can suppress cybersickness without direct vestibular stimulation, highlighting the potential of haptic-based motion compensation in immersive environments. -
2252 Cybersickness Reduction Effects of Visual and Vestibular Player-Fixed Rest Frames
Player-fixed rest frames (RFs) such as virtual noses have been proposed as cybersickness mitigation cues, yet their underlying mechanisms remain unclear. This study investigates how a player-fixed RF influences cybersickness, gaze behavior, and perceived orientation during passive visually induced motion. In two experiments using a virtual roller coaster with eye tracking, we found that a player-fixed RF reduced overall cybersickness without compromising presence or vection, and promoted a more centralized gaze distribution that progressively diverged from the control condition over repeated trials. However, because the RF's presence and its alignment with upright were confounded in this setting, a second experiment introduced a visual--vestibular conflict through scene tilt. Under this conflict, only an RF aligned with gravitational (vestibular) upright suppressed the cumulative buildup of sickness and preserved veridical orientation perception, whereas an RF aligned with the tilted visual scene did not. These results suggest that a player-fixed RF operates through at least two channels: redistributing gaze away from destabilizing peripheral motion, and serving as an egocentric orientation reference whose effectiveness may be further enhanced by its congruence with vestibular upright. The findings inform the design of rest frame cues by highlighting that both visual placement and spatial alignment contribute to their effectiveness.
PS27: Immersive Experiences and Culture Sezione 4+5
-
1240 ImmPres: Exploring the Potential of Mixed Reality Presentations in Real-world Office Settings
Mixed Reality (MR) can potentially transform workplace presentations and enhance knowledge sharing in office environments. Traditional slide-based presentation tools often constrain presentations to linear, screen-bound structures, which limit spatial exploration, audience engagement, and personalization. In contrast, MR supports three-dimensional content, augmenting real environments, and personalized views. We envision a future where MR presentations become a central medium for our workplaces. To systematically explore this potential, we conducted an expert workshop with MR and presentation specialists, including lecturers. Through thematic analysis, we identified key design aspects for creating MR presentations, including spatial content distribution, audience interaction, and maintaining narrative cohesion. Additionally, we demonstrate their use through an exemplified scenario of designing an MR presentation, illustrated by an implemented prototype. With the discussion of challenges and opportunities, we hope to lay the foundations for integrating MR into office presentations, reimagining how stories and ideas can be shared in professional settings. -
1533 LLM-Think-Alouds: Multimodal Player Experience Analysis for VR Playtesting
Virtual reality (VR) games can evoke strong and dynamic player responses due to their immersive nature. To design and refine these experiences, developers often rely on playtesting to understand how players’ emotions and perceived difficulty evolve throughout gameplay. However, manually analyzing gameplay footage and think aloud data is time-consuming and difficult to scale. We propose an LLM-based approach for automated player experience analysis in VR games using synchronized gameplay video and player audio. Based on estimated rating trajectories, the approach supports gameplay analysis, comparison of player experiences, and deeper inspection of how experience unfolds over time. We evaluated the approach through a user study with three VR platformer games. Our findings suggest that LLMs can recover meaningful experiential signals, particularly perceived difficulty, while performance across emotional dimensions varies depending on game context and data expressiveness. Overall, the results demonstrate the potential of LLM-assisted methods as scalable tools for VR playtesting and game design analysis. -
1534 More Than Immersion: How Virtual Cultural Experiences Influence Identity, Reflection and Presence Among Migrants
Loneliness is a subjective emotional state that arises from unmet expectations of social relationships. For migrants, loneliness is often intensified by disconnection from cultural identity and familiar lifestyles. This can hinder formation of meaningful social connections and trap individuals in a recurring loneliness cycle. Empirical studies reveal that missing traditional festivals can be a major contribution to emotional disconnection among migrants in their host country. Despite the ethnicity, missing own New Year festival emerges as a significant factor. In this research, insights from empirical work, together with findings from existing literature and expert consultations, informed the design of a culturally sensitive Virtual Reality (VR) experience. Given the subjective nature of both loneliness and culture, at this phase the prototype was designed targeting one specific ethnic community. Qualitative and quantitative evaluation conducted with first generation migrant adults (N=23) of this community demonstrated that the experience supported identity expression, self-reflection, and a strong sense of presence. This paper presents findings from a mixed method descriptive analyses, and discusses design implications for culturally sensitive VR experiences. In particular, the paper highlights the importance of everyday cultural artefacts, ritualized activities, multisensory cues, and opportunities for personal and social meaning-making supported through immersive experiences. The work contributes to immersive technologies research by showing that cultural presence in VR extends beyond visual realism to encompass identity, memory, and emotional resonance. -
1641 A Decade of Extended Reality in Cultural Heritage (2015-2025): A Systematic Literature Review
Cultural Heritage plays a crucial role in shaping human identity and transmitting knowledge across generations. Over the past decade, immersive technologies such as Augmented Reality, Virtual Reality, Mixed Reality and Extended Reality (XR) have increasingly been adopted to enhance cultural experiences and support the preservation and communication of heritage. However, while many XR applications have been proposed in this field, the literature often emphasizes the technological novelty and the “wow experience” of immersive experiences rather than systematically evaluating their effectiveness from a user-centered perspective. It remains unclear to what extent these systems are designed and assessed in terms of knowledge transfer and learning outcomes. To address this gap, this paper presents a systematic literature review of XR applications in cultural heritage published between 2015 and 2025. A total of 84 peer-reviewed studies were selected from five major academic databases and analyzed to investigate technological trends, application domains, and user evaluation practices across four areas: Museums and Exhibitions, Conservation and Reconstruction, Tourism and Exploration, and Intangible Cultural Heritage. The results show that Virtual Reality is the most widely adopted technology, while many applications are evaluated in controlled laboratory settings rather than real heritage contexts. Furthermore, although usability and user experience are frequently assessed, only a minority of studies explicitly evaluate knowledge transfer, and very few investigate knowledge retention. These findings highlight important methodological gaps and suggest directions for future research on user-centered XR design for Cultural Heritage.
Paper Sessions 28-3012:15–13:15
PS28: Projection and Display Techniques Orione + Perseo
-
1108 VIPA: View-Invariant Projector-Based Adversarial Attack
Projector-based adversarial attack aims to physically manipulate real-world scenes by projecting adversarial patterns, thereby causing deep image classifiers to produce incorrect predictions. However, existing stealthy projector-based adversarial attack methods model the project-and-capture process in a single and static view, limiting their view-invariant capability and practical applicability in dynamic environments. In this paper, we introduce View-Invariant Projector-Based Adversarial Attack (VIPA), a novel method designed to overcome these limitations by achieving both true view-invariant and classifier-agnostic adversarial attacks in the physical world. We first leverage a view-invariant projector-camera system simulation method to model the physical interactions between projected patterns and real-world surfaces. To ensure robustness and stealthiness of the attack across different views, VIPA optimizes adversarial projections by aggregating simulated attack losses from multiple views. This joint optimization across diverse views helps maintain robustness regardless of the viewpoints. Finally, the optimized patterns are physically projected into real-world scenes, where they successfully fool various classifiers from different viewpoints, thereby enabling robust and practical view-invariant adversarial attacks. Our experiments in both targeted and untargeted attacks demonstrate that VIPA consistently achieves higher attack success rates than existing methods, while also enhancing stealthiness and ensuring minimal perceptual degradation across all tested views. -
1176 Text2Makeup-DPM: A Text-Driven Dynamic Projection Mapping Makeup System for Interactive Facial Augmentation
Dynamic Projection Mapping (DPM) has revolutionized cosmetic consulting by transcending 2D screen overlays to be realized in the real world. However, existing DPM systems often rely on rigid interfaces and predefined textures, limiting creative expression. Therefore, we explore natural language as an intuitive interaction modality to maintain user focus on the projected results. Nevertheless, applying free language input remains challenging, particularly in achieving precise localized makeup generation and appropriately rendering the generated makeup on the face through projection. To address these limitations, we present Text2Makeup-DPM, the first fully text-driven DPM makeup system transforming free-form descriptions into facial makeup projections. Our system features a two-scale makeup generative pipeline and efficient radiometric compensation. First, we conduct text parsing to align free-format text input into the global style and local demand of makeup. The global style module captures the overall makeup style and atmosphere through contrastive learning. A local component module optimizes latent direction compositions to enable precise, disentangled editing of specific regions (e.g., eyeshadow, lipstick). Subsequently, an asynchronous Hadamard-division compensation ensures high-fidelity projection without compromising temporal interactivity. Quantitative evaluations demonstrate that our two-scale pipeline aligns more accurately with natural language descriptions than baseline models. Furthermore, a user study (N=12) confirms the system's effectiveness in enhancing the immersion of interaction flow, proving its contribution to intuitive DPM authoring. -
1079 Fast 3D Gaussian Splatting with Blue-Noise Stochastic Sampling for Dynamic Projection Mapping
Dynamic projection mapping (DPM) enables spatially aligned projection onto moving objects, but requires low-latency rendering and accurate appearance reproduction. Meanwhile, 3D Gaussian Splatting (3DGS) has recently attracted considerable attention as an image-based approach for reconstructing 3D scenes solely from images and rendering them from arbitrary viewpoints. This paper proposes a framework for realizing DPM using 3DGS. Compared with conventional DPM approaches based on physically based rendering (PBR), the proposed 3DGS-based framework enables simpler, high-fidelity reproduction of real object appearance by eliminating the need for complex prior measurement of material parameters. However, the low-latency rendering required for DPM necessitates the use of fast 3DGS variants, which introduce rendering noise that leads to visible projection artifacts. To address this issue, we introduce blue-noise sampling into stochastic transparency for 3DGS to reduce perceptually noticeable noise. We further reduce noise by using high-frame-rate projection that leverages persistence of vision and adapting perceptual-quality-preserving efficient rendering techniques from the PBR literature to 3DGS. Furthermore, assuming a rigid target object, we exploit the equivalence between viewpoint motion in 3DGS and object motion in DPM to obtain viewpoints consistent with the pose of the projection target, enabling the direct use of existing 3DGS renderers. Experiments on synthetic scenes and real-world projections demonstrate that the proposed method maintains rendering speeds comparable to conventional fast variants while achieving LPIPS reductions of 57.7 % and 45.2 %, respectively, effectively suppressing visual artifacts. -
1088 [EXTENDED] LowPowAR: Power-Constrained Tone Mapping for Augmented Reality
Everyday-wearable Augmented Reality (AR) glasses must meet strict power limits, making displays a key target for optimization. We cast display power optimization as a power-constrained tone-mapping problem and propose a human-vision--grounded, learning-based framework that maximizes perceptual quality under a given power budget. We introduce an optimization-friendly tone-mapping operator (TMO) parameterization along with a progressive optimization strategy to effectively navigate the quality-vs-power landscape. We distill the iterative optimization into a lightweight feed-forward neural network for real-time deployment. Subjective experiments show that our method yields better perceptual quality than prior work at the same power budget.
PS29: Navigation Cues Sezione 2
-
1248 Effects of Field-of-View Restriction and Peripheral Teleportation on Path Integration during Virtual Locomotion
Most software-based cybersickness mitigation techniques, such as field-of-view (FOV) restrictors and rest frames, alter visual content and peripheral optical flow, often negatively affecting various locomotion quality factors. Path integration, a critical component of spatial updating, relies on optical flow during virtual locomotion and may therefore be particularly affected by these techniques. Despite a substantial body of work evaluating cybersickness mitigation methods, relatively few studies have examined their effects on spatial updating or path integration. Moreover, prior studies on spatial updating have not yielded conclusive results, despite theoretical predictions from the perception literature. In this work, we conducted a preregistered experiment with 48 participants to compare the effects of two mitigation techniques, the dynamic FOV restrictor and peripheral teleportation, on path integration, relative to an unrestricted control condition. By minimizing potential confounding factors identified in previous work, we found that participants made significantly larger homing errors with the FOV restrictor than in the control condition, whereas performance with peripheral teleportation did not differ significantly from either of the other conditions. These findings provide direct evidence that FOV restriction impairs path integration, while the impact of peripheral teleportation remains inconclusive. Our results further highlight the trade-off between mitigation effectiveness and spatial updating, providing VR designers with deeper context for selecting mitigation techniques. -
1546 Effects of Spatial Perspective and Frame of Reference Integration on Gaze Behavior and Spatial Learning
As spatial computing offers expanded design spaces and brings digital information seamlessly into a physical environment, it redefines how we perceive space and navigate our surroundings. This study investigates enhancing human navigation and spatial knowledge acquisition in large space environments, focusing on the integration of spatial visual cue display in 3D design space. Three cues were designed (world-fixed cue, body-fixed world-in-miniature map, and a combination of both) and their impact on spatial learning, eye movements, cognitive load, and user experience was assessed. The findings reveal that displaying cues across frames of reference (combining the world- and body-fixed cues) significantly enhances spatial information processing and learning without the cost of cognitive load. The study also sheds light on how individuals' eye movement patterns differ with cue design and while learning, which advances our understanding of the relationship between gaze behavior and spatial learning outcomes and informs cognition-aware adaptive 3D interface design. -
2046 Central or Peripheral? Investigating AR Navigation Cue Placement in Multitasking Environments
Augmented reality (AR) often presents navigational cues directly within the user’s field of view. While centrally placed cues provide explicit direction, they may compete with task-relevant perception and increase visual clutter during multitasking. Peripheral cues are often considered a less visually intrusive alternative by leveraging the human visual system’s sensitivity to motion and orientation, but their effectiveness for guidance remains insufficiently understood. This paper investigates how cue placement within the visual field (foveal 3D Arrow vs. peripheral segmented halo) influences navigation performance, attention management, and user experience in AR environments involving divided attention and visually complex surroundings. In a controlled study (N=22), participants navigated between targets while performing concurrent reading tasks. Results show that foveal cues enable significantly faster navigation and spatial localization across the majority of target destinations. In contrast, peripheral cues led to a significant 2.12-second delay in secondary task response times, indicating increased cognitive processing demands during multitasking. Interestingly, despite this performance cost, peripheral cues were perceived by some participants as visually comfortable, and user preference was divided (55% foveal cues vs. 45% peripheral cues). These findings highlight a critical trade-off between visual presentation and cognitive efficiency: while peripheral cues were at times experienced as less visually intrusive, their abstract representation requires additional interpretation effort. This challenges the assumption that relocating guidance cues to the periphery inherently improves attentional resource allocation in AR systems. -
2264 Actionable Guidance Outperforms Map and Compass Cues in Demanding Immersive VR Wayfinding
Navigation aids are central to immersive virtual reality (VR) experiences that involve physical locomotion. Their effectiveness depends not only on how much spatial information they provide, but also on how directly that information supports movement decisions. We compared three common guidance techniques for immersive VR wayfinding: a directional arrow, a minimap, and a compass. In a controlled room-scale VR study with 42 participants completing 1008 trials, participants navigated to target landmarks in a time-pressured maze with reduced visibility and forced route replanning. Across behavioral and eye-tracking measures, arrow guidance produced the strongest navigation performance, minimap guidance yielded intermediate performance, and compass cues performed worst, suggesting that during immersive locomotion users benefit from guidance that can be interpreted rapidly while moving. These results suggest that in demanding immersive locomotion tasks, interfaces that translate spatial information directly into actionable movement cues can outperform richer but more interpretive spatial representations. Our findings highlight the importance of designing XR navigation interfaces that minimize the cognitive translation between spatial information and movement decisions.
PS30: 3D Manipulation Sezione 1
-
1049 Modeling Throwing Selection in Virtural Reality
Throwing is a natural and embodied interaction for selecting distant targets in virtual reality (VR). Unlike conventional pointing, throwing is ballistic: once released, the trajectory can no longer be corrected. Despite its practical importance, throwing-based target selection has received relatively little attention in human behavior and performance modeling, making systematic design and evaluation still difficult. In this paper, we investigate throwing-based target selection in VR and propose predictive models for its behavior and performance. Through a controlled user study manipulating target width ($W$) and movement amplitude ($A$), we measured planning time ($PT$) and endpoint distribution. Our results showed that $PT$ increased systematically with task difficulty, following a pattern similar to movement time in Fitts' law. Throwing endpoints were also generally well approximated by a bivariate Gaussian distribution, and target parameters significantly affected both endpoint bias and variability. Based on these findings, we derived a probabilistic accuracy model that predicts target acquisition performance. The proposed models showed high goodness-of-fit, and the accuracy model was robust under leave-one-condition-out cross-validation. Our findings provide a quantitative understanding of throwing-based target selection and support the design of throwing interactions in immersive environments. -
1677 Investigating Scale-Control Strategies for Portal-Based Remote Object Manipulation in Virtual Reality
Portal-based interaction enables users to access and manipulate distant content without physical navigation. However, effective remote manipulation remains challenging due to mismatches between object and interaction scale, reducing control precision and increasing coordination effort. In this study, we investigate how different scale-management strategies in portal-based interaction affect task performance and user experience in virtual reality (VR). We designed four portal techniques that vary in how interaction scale is controlled during manipulation: fixed-scale, manually scalable, overview+detail, and transfer-based configurations. We conducted a within-subject VR experiment (N = 31) using a remote 6DoF docking task with objects of varying sizes. Results show that scale-management strategies significantly influence manipulation performance, particularly for smaller objects. Multi-portal configurations improved efficiency by providing task-appropriate interaction scales, whereas single-portal approaches offered simpler workflows but required additional user effort. These findings highlight the importance of structuring interaction scale across task phases and provide design implications for portal-based manipulation in VR. -
10.1109/TVCG.2026.3703476 Rubbing Interaction with Two-Handed Virtual Reality Controllers
As a common action in our daily lives, rubbing has been a viable technique for various devices, including touchscreens, digital pens, and skin surfaces. However, there has been little research on the design of rubbing in virtual reality (VR), particularly with VR controllers. This study explores the design space of rubbing with two-handed VR controllers. We started with a one-on-one interview with 10 participants to understand four components of rubbing with VR controllers: hand use, rubbing direction, device type and rubbing contact. In Experiment 2, we proposed a method to classify rubbing with three hand uses and six rubbing directions. Results showed that the overall classification accuracy reached 96%. Experiment 3 revealed that our method was able to successfully reject all seven tested non-rubbing actions. Results from Experiment 4 reflected user preference for using our technique to perform five typical VR tasks. Our study contributes to VR interaction design with controllers. -
1237 RadBoxing: A Virtual Reality Environment for Detecting and Annotating Abnormalities in Volumetric Medical Imaging
A major challenge in radiological diagnoses is the ability to correctly and quickly identify abnormalities from Computed Tomography (CT) images. Although some research exists on the volumetric rendering and visualization of CT scans for training, interaction with and testing of 3D visualizations for diagnosis and educational training in radiological imaging is still limited. To address this gap, we present a 2D + 3D interactive system that allows users to locate potential fractures simultaneously on 2D and 3D CT visualizations and to annotate them using a controller-based encapsulation technique. Using algorithms we built to automatically filter 2D CT images and display a bounding box on the 2D images from the 3D model's equivalent location, participants can simultaneously view both 2D and 3D anomalies to learn to recognize fractures. In addition to this technical framework, we conducted an experiment (N=55) to compare 2D, 2D3D, and 3D viewing conditions to evaluate the effects of visualization methods on abnormality detection using VR-based annotation. Our findings show that interactive methods comprised of 2D images in a 3D space lead to a higher detection sensitivity and reaction time when compared to 3D only. These results suggest that embedding 2D diagnostic imagery within an interactive 3D spatial context may improve perceptual training in novices, offering a promising direction for using virtual reality to improve radiological training in fracture identification.
Paper Sessions 31-3314:15–15:15
PS31: Affective XR Sezione 1
-
1141 CyberSelf: Embodied Self-Distancing for Emotional Support in Virtual Reality
Self-distancing is an effective emotion regulation strategy; however, it may fail during personal crises due to its cognitive demands. Virtual Reality (VR) provides a novel approach to externalizing psychological distance by enabling embodied self-representation. In this paper, we present CyberSelf, a VR system for emotional support that integrates a visually self-resembling avatar, a cloned self-voice, and Large Language Model (LLM)-driven real-time dialogue. The system enables users to engage in multi-turn conversations with their self-representations in immersive VR, enabling embodied self-distancing while maintaining a strong sense of self-relevance. We evaluated CyberSelf in a short-term study that compares three levels of self-representation richness (Text, Text+Voice, and Text+Voice+Appearance). The results demonstrated robust pre-post improvements across affective and coping measures, specifically increased valence, arousal, hope, and resilience, as well as reduced anxiety and simulator sickness. Richer representations increased conversational engagement, and full embodiment produced the strongest physiological indicators of emotional regulation. A subsequent four-week long-term study demonstrated that these benefits are both sustainable and cumulative. Additionally, users rated the reconstructed avatar and the cloned voice as highly recognizable and acceptable. Collectively, these findings suggest that embodied, self-resembling conversational agents provide a viable mechanism for externalizing self-distancing and supporting emotional regulation in VR. -
1660 To Understand Changes of Emotion and Cognitive Efforts in VR
Understanding how cognitive and emotional states evolve during everyday activities is important for designing effective Virtual Reality (VR) systems for training and rehabilitation, such as for recovering from traumatic brain injury (TBI). However, less is known about how these internal states fluctuate during naturalistic VR tasks. In this work, we investigate cognitive and emotional state dynamics during a VR shopping activity designed for TBI treatment. Participants performed budgeting, item selection, and decision-making tasks while interacting with a virtual coach in either a low-stimulus (quiet) or high-stimulus (busy) environment. During the task, participants reported on momentary mental workload while physiological signals were continuously recorded. After the session, participants retrospectively annotated perceived cognitive and emotional state changes over time. Our results showed higher perceived workload and physiological arousal in the busy environment, with GSR activity positively correlated with reported cognitive workload. We describe implications of the research, limitations, and directions for future work. -
2064 Towards Calibration-Free Affective Modeling from Physiological and Behavioral Signals in Immersive VR Interaction
Understanding users’ affective states is important for enabling adaptive Virtual/Extended Reality (VR/XR) interaction systems. However, reliable cross-subject affect prediction remains challenging due to substantial individual variability in physiological responses and behavioral patterns during interaction. We propose AMDA, an Asymmetric Modality Disentanglement Architecture for continuous Valence–Arousal–Dominance prediction from multimodal physiological and behavioral signals in VR teleoperation. AMDA models the two modalities asymmetrically: physiological signals are processed via personality-aware residual disentanglement to separate stable traits from affective dynamics, while behavioral signals are modeled through interaction-driven personality conditioning to capture task-dependent emotional expressions. Experiments on a multimodal VR teleoperation dataset collected under diverse network degradation conditions show that AMDA improves cross-subject continuous affect prediction, particularly for Valence and Dominance. The results further suggest a calibration-light direction for affect-aware VR/XR systems through structured multimodal representation learning. -
2320 Asymmetric Threat Manipulation and Counterbalanced Leadership Cues: Effects on Decision Making, Evacuation, and Visual Attention in Virtual Emergency Scenarios
This study employed an immersive virtual reality paradigm to investigate how smoke exposure, leadership signals, and the congruence or incongruence of hazard–leadership cues shape evacuation behavior and patterns of visual attention. Participants interacted with two virtual humans: a familiar guide and an unfamiliar newcomer. At a critical intersection, they decided whether to follow one of them or avoid smoke at a key intersection. Smoke presence was manipulated to align with the guide’s path, the newcomer’s path, or be absent, creating congruent and incongruent threat–leadership conditions. Evacuation outcomes were analyzed through following behavior and smoke avoidance across the simulation, while visual attention (fixation time) and decision area duration were measured at the key intersection. Additionally, participants were grouped by their overall tendency to follow others, enabling subgroup analyses of evacuation behavior. Results revealed that reliance on the guide varied systematically by participant gender and following tendency. Smoke presence redirected visual attention toward social cues, reducing attention to environmental hazards. Smoke avoidance was jointly shaped by leadership cues and prior smoke exposure within the simulation, highlighting the dynamic interplay between social influence and threat perception. These findings suggest that emergency training should integrate both social and environmental cues, account for individual differences in following behavior, and use VR simulations to expose participants to conflicting leadership and hazard signals, offering insights for the design of effective emergency training and simulation systems.
PS32: Redirected Walking Sezione 4+5
-
1395 Predictor-Dependent Reference-Frame Selection for Trajectory Forecasting under Redirected Walking
Modern pedestrian trajectory predictors achieve strong performance on standard benchmarks and are increasingly considered for Virtual Reality (VR) locomotion. However, Redirected Walking (RDW) introduces a domain gap: the same user motion can be represented either as a perceived trajectory in the virtual environment (VE)or as an executed trajectory in the physical environment (PE), with potentially different statistics under RDW gains. This raises a fundamental question: which representation is more favorable for trajectory prediction under RDW? To study this question, we collected an RDW trajectory dataset with paired VE/PE trajectories from 25 participants under four RDW conditions, and constructed a controlled evaluation framework tailored to reference-frame comparison under RDW. Within this framework, we evaluated four predictors spanning diffusion-based, goal-conditioned, Transformer-based, and physics-based methods. Using a subject-wise five-fold protocol with hierarchical aggregation, we compared zero-shot transfer with fine-tuning on VE, PE, and mixed data. Using length-normalized errors, we find that the three learning-based predictors evaluated in this study tend to favor PE, whereas the physics-based baseline EKF favors VE. Adaptation strategy affects the magnitude of the VE--PE gap, but does not overturn this overall pattern. These findings provide bounded but practical guidance for prediction-aware RDW: for the learning-based predictors evaluated in this study, PE is a strong default representation in the evaluated setting, while mixed fine-tuning offers a practical adaptation strategy when RDW-specific data are limited. -
2040 Global Mapping-Guided Redirected Walking for Continuous Locomotion in Irregular Spaces
Redirected walking (RDW) has primarily focused on accommodating virtual environments that are much larger than the tracked physical space. In many location-based entertainment and room-scale VR deployments, however, the physical and virtual environments are comparable in extent but differ in boundary geometry, making frequent turn-in-place resets undesirable rather than inevitable. We introduce a global mapping-guided (GMG) controller for this comparable-scale setting. The method uses a dual-layer design. First, a hard mapping layer establishes a bijective correspondence between an irregular physical domain and a near-rectangular virtual reference domain using a discrete Laplacian formulation with force-directed refinement. This layer encodes global reachability without prescribing a fixed route. Second, a soft control layer regulates translation and strafing gains within perceptual thresholds so that runtime motion tracks the hard mapping while remaining perceptually natural. The resulting controller supports continuous locomotion without predefined trajectories and avoids explicit turn-in-place reset operations during normal operation. Simulations and a user study show that GMG substantially reduces boundary contacts relative to representative gain-based baselines in comparable-scale layouts, while improving comfort and subjective continuity of walking. -
2041 Coupled Redirected Walking for Socially Aligned Co-Located VR
Location-Based Entertainment (LBE) VR has rapidly gained global momentum, exemplified by large-scale free-roam experiences such as The Disappearing Pharaoh. This global boom has motivated many companies to explore Redirected Walking (RDW) as a cost-saving technique, enabling expansive virtual worlds within limited physical venues. However, while RDW saves space, the mismatch it creates between physical and virtual motion leads to a key challenge: when partners meet in the virtual world, their physical bodies may fail to co-locate, breaking immersion and producing a frustratingly unrealistic experience. This issue cannot be solved by simply assigning identical gain values to all users. We propose a control framework that ensures synchronization among participants during RDW. We formalize two complementary concepts: distance-coupling—the bounded ratio of inter-user distances across physical and virtual spaces—and angle-coupling—the mutual reachability of participants after heading adjustments. Based on these definitions, we derive differential constraints on translation and curvature gains within perceptual thresholds. An arctangent-based controller regulates distance ratios, while an angular controller minimizes orientation divergence. For multi-user conflicts, a priority mechanism classifies temporarily decoupled users and selects gain adjustments that reduce group-level disruption. Our method transforms multi-user RDW from obstacle avoidance into socially coherent redirection, allowing participants to meet and interact naturally in both spaces while maintaining the scalability and cost efficiency of LBE VR. -
10.1109/TVCG.2025.3595181 Towards Walkable and Safe Areas: DRL-Based Redirected Walking Leveraging Spatial Walkability Entropy
Redirected walking (RDW) expands the virtually reachable areas within confined physical spaces by real-walking locomotion. However, existing RDW controllers struggle with extracting spatial features, hindering the improvement for physical obstacle avoidance. To overcome this, we propose a novel spatial walkability-aware redirection controller utilizing deep reinforcement learning (DRL), which learns to enhance obstacle avoidance capability by leveraging comprehensive spatial features. Based on information entropy, we innovatively introduce the spatial walkability entropy (SWE) metric to characterize the walkability and safety of each physical position by assessing the difficulty of reaching its surroundings. Guided by this, we design a novel joint reward that considers both the SWE distribution and the user's virtual-physical alignment, providing ample guidance for learning. Moreover, unlike existing controllers employing traditional reset strategies, we propose a novel reset method that maximizes regional entropy to guide users towards more open areas, reducing the re-collision risk. Extensive simulation experiments compare our controller with state-of-the-art (SOTA) redirection controllers. The results demonstrate that our controller significantly reduces physical collisions across various virtual-physical scenarios. Moreover, live user experiments confirm that our controller offers a superior roaming experience in practical settings.
PS33: Social Impact Orione + Perseo
-
1273 Do We Need to “Be There” to Say Goodbye? Disentangling Participatory Agency and Immersive Presence in Digital Farewells
Digitally mediated attachment is increasingly common, yet the mechanisms of symbolic closure in hybrid physical--digital relationships remain underexplored. We investigate how graded levels of participatory agency and immersive presence shape symbolic closure and emotional processing in an experimental farewell context. Using ``cloud petting'' as a controlled proxy of digitally mediated attachment, we compare three farewell modalities: (1) an observational messaging interaction as a low-agency, low-presence baseline (``Messaging''), (2) a 2D memorial website enabling symbolic participation with moderate presence (``Website''), and (3) a 6-DoF Virtual Reality ritual environment combining active participation with embodied spatial presence (``VR''). In a mixed-methods study ($N = 45$), both the Website and VR produced significantly higher perceived symbolic closure than Messaging, with no significant difference between Website and VR. While the VR condition was associated with significantly stronger grief-related responses on the Pet Bereavement Questionnaire, generalized negative affect measured by the Positive and Negative Affect Schedule did not differ significantly across modalities. This divergence suggests an Agency-Amplification account: symbolic closure is supported by a combination of active participatory affordances and moderate experiential presence, whereas maximal immersive embodiment functions as a condition associated with stronger loss-oriented emotional processing. By comparing participatory agency and immersive presence as distinct design factors, this study advances XR research by clarifying their distinct contributions to symbolic closure and emotional intensity in digital farewell experiences and by providing preliminary design guidance. -
1661 The Missing Link of XR: Empathy-Driven Reality for XR and Beyond
Extended reality (XR) for socialising is becoming increasingly popular. However, unlike conventional social platforms, XR prioritises embodiment and immersion, factors that strongly impact one's physical and mental states. We envision a future for XR where all users, regardless of abilities and backgrounds, can understand one another, participate, and find safe socialisation spaces. An Empathy-Driven Reality (EDR) is a space where understanding each other's emotional, physical, and cognitive states takes centre stage. It has the potential to enhance empathy beyond how we normally perceive it. To explore this concept, we conducted a hybrid-style workshop over two months with 27 industry and academic researchers in XR, emotion, physiology, assistive technology, and social science. This paper reports on the findings and aims to establish a structure and reference for the 1) design guidelines, 2) research challenges, and 3) potential applications for the future of XR as an EDR. -
2161 Harassment Isn’t Virtual When It Feels Real: Understanding Emotional Impact of Gendered Embodiment in VR
Harassment in Social Virtual Reality (SVR) can feel more intense than in traditional online platforms due to embodiment and real-time interaction. However, how gender cues shape users’ perception of harassment in immersive environments remains unclear. We investigate how avatar gender, harasser gender, and participant gender influence perceived discomfort during social VR interactions. Using a customized Unity-based VR environment, we present participants with harassment scenarios enacted by non-player characters representing female, male, and neutral genders. The behaviors are derived from a prior online survey with 61 respondents. In a mixed-methods user study (N=46), participants experienced seven predefined harassment behaviors and reported their emotional discomfort on a 5-point Likert scale after each interaction, followed by post-study interviews. Results show that gender cues influence how harassment is perceived in immersive environments, highlighting the role of embodiment in shaping users’ emotional responses. We contribute empirical evidence on gendered harassment in SVR and discuss implications for mitigation strategies, avatar design, and ethical safeguards in immersive social systems. -
10.1109/TVCG.2026.3699425 Sharp Body, Sour Taste: Changing Taste Experiences through Cross-Modal Correspondence with the Bouba/Kiki Self-Avatar in VR
Cross-modal correspondence, exemplified by the Bouba/Kiki effect, which refers to the association between linguistic sounds and visual shape impressions, is known to influence taste perception. Previous studies have shown that taste perception varies with the visual features of foods, beverages, and tableware. However, it remains unclear how embodying a self-avatar in virtual reality (VR) whose visual features are associated with Bouba/Kiki influences taste perception. We conducted two studies to investigate the effects of cross-modal correspondence between VR self-avatar characteristics and taste perception using lemon-flavored carbonated beverages that elicit both sweetness and sourness. Study 1 directly compared Bouba avatars (round shapes, red colors, slow animations) and Kiki avatars (angular shapes, yellow colors, fast animations) in a within-participants design (N=36). Study 2 examined the effects of each avatar type separately by comparing them with neutral control conditions (N = 72). In Study 1, the Bouba avatar led participants to perceive the beverage as mellower compared to the Kiki avatar. However, this effect was not significant when compared with a neutral control avatar in Study 2. In contrast, the Kiki avatar produced robust effects, increasing perceived sourness and sharpness in both studies. These findings suggest that the observed changes in taste are primarily driven by the congruence between avatar characteristics and the intrinsic sensory properties of the beverage. This study demonstrates the potential for avatar-mediated congruence effects to alter perceived taste in VR, highlighting the importance of the alignment between the user's virtual body and the consumed stimulus.
Paper Sessions 34-3614:45–15:45
PS34: Panoramic Media Cigno + Auriga
-
1267 SGMRS: Spherical Gaussian Guided Multi-resolution Sampling for Fast Foveated 360-degree Video Streaming
With the rapid development of virtual reality technology, 360◦ panoramic video has become increasingly widespread. To ensure high visual quality across the wide field of view, the resolution of such videos continues to increase, which brings significant challenges for video storage, processing, and especially trans- mission. Previous research has proposed foveated image transformation techniques to reduce the resolution in non-foveal regions, thereby lowering transmission bandwidth. However, these approaches operate on ERP-projected panoramic video frames without accounting for the severe geometric distortions introduced by ERP projection, particularly in high-latitude regions. This paper proposes a spherical Gaussian-guided multi-resolution sampling (SGMRS) method. Specifically, we use a normalized spherical Gaussian distribution to generate a Fibonacci-based sampling point set that conforms to the spherical visual acuity model, and perform foveated multi-resolution sampling directly in the spherical coordinate domain. We then precompute the spherical sampling map for a canonical gaze direction and exploit the rotational symmetry of the spherical Gaussian distribution to achieve fast recovery of 360◦ foveated video frames. This enables real-time foveated 360◦ video streaming. We conduct quantitative evaluations on existing 360◦ video datasets and compare our method against state-of-the- art techniques. The results demonstrate that our method substantially improves visual consistency across different gaze directions. Moreover, under comparable transmission bandwidth, our method significantly improves both the perceptual quality of the foveated video and its temporal stability. -
1516 Saliency-aware Foveated Sparse Volume Rendering
Foveated sparse volume rendering is a key direction for enabling real-time immersive volume visualization on resource-constrained standalone headsets. However, existing methods typically ignore the human visual system’s sensitivity to salient content in peripheral vision and often suffer from severe temporal flickering artifacts under sparse sampling. In this paper, we present Saliency-aware Foveated Sparse Volume Rendering, a lightweight pipeline that enables real-time immersive raycasting on standalone headsets. The key idea is to jointly model gaze-dependent acuity falloff and content-driven visual saliency, and to use the resulting screen-space saliency map to guide sparse raycasting. We further introduce an artifact-aware reconstruction strategy that leverages two inexpensive, shading-simplified auxiliary renderings as guidance, effectively mitigating flickering artifacts induced by sparse sampling. We evaluate our method on multiple biomedical volumes; both objective metrics and a user study demonstrate improved perceptual quality over representative foveated sparse volume rendering methods under real-time constraints. -
1538 A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence
360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardware. Many telepresence installations feature fixed viewpoints and largely static scenes, yet optimization strategies tailored to such setups have received limited attention. We present a multi-layer, ultra-high-resolution system for static 360-degree telepresence that combines an 8K panoramic camera with a rotatable 4K pan-tilt-zoom (PTZ) camera. Our approach builds a three-layer representation: (1) a tile-based ultra-high-resolution panoramic background, generated by offline stitching high-detail 4K PTZ scans onto the base 8K panorama to achieve effective resolution beyond native capture, and represented as a set of spatial tiles; (2) a dynamic update layer that composites foreground motions from the 8K stream via real-time high-resolution background matting; and (3) a region-of-interest 4K layer that streams a real-time PTZ view of the selected region and additionally updates the corresponding background tiles over time. We evaluate the proposed system through comparisons with representative video super-resolution approaches and a user study assessing perceived detail and immersive experience. Our results indicate that tile-based background refinement, together with user-guided updates, provides a practical way to balance panoramic fidelity and interactivity in static 360-degree telepresence. -
2296 Hear What You See: Depth-Aware Spatial Audio Rendering for Immersive Panoramic Video
Panoramic video has emerged as a core medium for virtual reality (VR) and immersive media, yet its audio typically remains stereo, lacking spatially coherent soundfield reproduction consistent with the visual content and severely compromising immersion. Although multi-loudspeaker spatial audio systems can reconstruct realistic soundfields in physical spaces, existing solutions remain heavily reliant on manual production, and mainstream algorithms such as Vector Base Amplitude Panning (VBAP) handle only directional information while inadequately modeling distance perception. This paper proposes the Vision–Acoustic Coupled Spatialization (VACS) framework, which automatically extracts three-dimensional (3D) motion trajectories of sounding objects from panoramic videos containing only stereo audio tracks and generates spatially consistent multichannel audio end-to-end. The core innovation is the Depth-Aware Spatial Audio Rendering (DASAR) method, which leverages monocular depth estimation to integrate sound pressure level (SPL) attenuation, spectral adjustment, and reverberation control within a unified framework, overcoming the angle-only limitation of conventional VBAP. Experiments show 62.6% localization accuracy within 10° tolerance, with Pearson correlations of 0.877 and 0.843 (p < 0.001) for azimuth and elevation. Subjectively, compared with VBAP, spatial depth and immersion improve by 13.2% (p < 0.05) and 19.5% (p < 0.01), respectively. Comparison with professional manual production shows comparable quality across all dimensions (p > 0.05) while substantially lowering the production barrier.
PS35: Cybersickness Prediction Glasshaus
-
1180 One Size Does Not Fit All: Personalized Cybersickness Forecasting with Asymmetric Few-Shot Meta-Learning
Cybersickness remains a critical barrier to the widespread adoption of virtual reality (VR), making accurate prediction of its onset an important research priority for enabling comfortable VR use. However, cybersickness varies across individuals, and models trained on population-level data often fail to generalize to new users. Most existing approaches also focus on detecting or classifying cybersickness after symptoms appear rather than forecasting how discomfort will evolve, which limits timely mitigation. To address these challenges, we propose an asymmetric few-shot meta-learning approach that uses asymmetric temporal encoding to forecast cybersickness severity 30 and 60 seconds ahead using multimodal data, including heart rate, eye tracking, and head tracking, with only 30 seconds of user calibration. We validate the approach on two public datasets: a real walking VR task (Maze, 36 participants) and seated VR tasks (Simulation21, 30 participants). The model is evaluated for both in-domain performance and cross-domain generalization. Compared to state-of-the-art methods, it achieves a 53.7% reduction in RMSE for in-domain prediction on the Maze dataset and a 49.1% reduction under cross-domain transfer at the 60-second forecast horizon. Memory and computational analysis show an inference latency of 2.5 milliseconds and a memory footprint below 1.0 MB, indicating suitability for standalone VR deployment. Overall, the results provide a practical foundation for personalized cybersickness forecasting to support proactive mitigation and improve user comfort in VR environments. -
1263 Not Everyone Gets Sick the Same Way: Personalized Cybersickness Prediction via Bayesian Networks with Individual Susceptibility Modeling
Predicting the onset of cybersickness enables proactive mitigation strategies that minimize user discomfort, which is essential for the widespread adoption of virtual reality (VR) in education, healthcare, training, and entertainment. Current prediction models leverage multimodal data such as visual scene parameters and physiological signals, but often overlook individual differences that strongly influence cybersickness susceptibility, limiting generalization across diverse users. To address this limitation, we propose a personalized cybersickness prediction framework based on Bayesian Networks (BN) that incorporates individual factors, including age, gender, and prior VR experience. We introduce three personalization strategies: direct integration of user attributes, a two-stage pipeline that learns a latent susceptibility score, and stratified models trained on homogeneous user groups. We validate the approach on two publicly available datasets, Maze and Simulation21. On the Maze dataset, the method achieves a 14.27% relative improvement over the non-personalized Bayesian baseline and outperforms DeepTCN and LSTM by 44.59% and 48.07%, respectively. It also improves accuracy by 37.83% over the transfer learning strategy. Similar trends are observed on Simulation21, where the framework achieves a 13.40% improvement over the non-personalized baseline, outperforms DeepTCN and LSTM by 9.16% and 10.82%, and improves accuracy by 3.89% over transfer learning. Consistent gains are also observed in F1 score, Cohen’s kappa, and Spearman correlation. The method maintains low inference latency, supporting interpretable and efficient probabilistic predictions in VR settings. -
1352 AdaptVR: Leveraging Reinforcement Learning and LLMs for Personalized Cybersickness Mitigation and Reasoning in VR
Cybersickness remains a persistent barrier to the widespread adoption of virtual reality (VR), often undermining user comfort and immersion. Existing cybersickness mitigation techniques are mostly static and may fail to adapt to diverse user needs and dynamic VR contexts. For instance, aperture size in the dynamic field of view (DFOV) and blur intensity in the dynamic Gaussian blur (DGB) may need to vary across users and applications based on real-time cybersickness severity. To address this, we introduce AdaptVR, a reinforcement learning (RL)–based adaptive cybersickness mitigation framework that predicts, explains, and mitigates cybersickness. Unlike static methods, our RL agent optimizes long-term user comfort by adaptively selecting and adjusting mitigation strategies through continuous interaction with VR environments. We train the RL agent using Proximal Policy Optimization (PPO) in domain-randomized simulations and deploy the learned policy on a consumer-grade VR headset (HTC Vive Pro). At runtime, AdaptVR continuously processes real-time user signals (e.g., eye-tracking, head motion, scene context) to detect cybersickness severity and appropriate mitigation (e.g., DFOV and DGB) with exact intensity. We also design a large language model (LLM)–powered interactive dialogue engine that enables users to engage in natural, voice-based conversations to understand detection outcomes, explore mitigation choices, and provide feedback. This feedback is incorporated into a human-in-the-loop training process, allowing the RL agent to fine-tune its policy for personalized cybersickness detection and mitigation. We evaluate AdaptVR through a user study in a VR maze simulation. Results show that our RL-driven adaptive mitigation significantly reduces cybersickness with minimal impact on immersion. Moreover, 90% of participants reported that the framework is intuitive, effective, and improved their understanding of the mitigation process through the dialogue engine. The open-source implementation of the AdaptVR framework can be found at https://anonymous.4open.science/r/AdaptVR-D676 -
2237 Graph Anomaly Detection on Eye-Tracking for Cybersickness Assessment in Virtual Reality
Cybersickness remains a major barrier to the widespread adoption of virtual reality (VR) systems. Although many prior studies have attempted to predict cybersickness using objective sensor data, most approaches rely on subjective self-reported measures such as the Fast Motion Scale (FMS) or the Simulator Sickness Questionnaire (SSQ) as ground truth labels. The highly subjective nature of these measures introduces variability across participants, making it difficult to reliably relate objective signals to cybersickness levels and limiting the generalizability of prediction models. In this work, we propose an alternative approach that extracts cybersickness-relevant information directly from objective eye-tracking data by detecting anomalies in gaze behavior. Raw eye-tracking signals are first converted into scanpath graphs that represent fixation and saccade patterns, enabling the capture of both structural and temporal characteristics of visual exploration. These graphs are then used to train several anomaly detection models, including Isolation Forest, One-Class SVM, Dense Autoencoder, LSTM Autoencoder, and Spatio-Temporal Graph Autoencoder. The models learn patterns of normal gaze behavior and identify deviations that may indicate progression of cybersickness. We evaluate the proposed framework on three VR datasets: VRWalking, Terrain, and VRNet, comprising 5,357 scanpath graphs derived from $\sim$443 sessions. Using segments with minimal sickness (FMS ≤ 1) as the baseline behavior, anomaly scores are generated to simulate the temporal progression of cybersickness within each session. Across all datasets and participants, our approach achieves an RMSE of 1.188 and an MAE of 0.927 when approximating FMS trajectories. These results suggest that gaze-based anomaly detection can provide an objective signal for cybersickness assessment and enable the development of alternative scoring frameworks that reduce reliance on subjective self-reported measures.
PS36: Spatial Interfaces Sezione 2
-
1178 Designing Spatial User Interfaces in Mixed Reality: A Systematic Literature Review
We present a systematic review of spatial user interfaces (SUIs) in mixed reality (MR) environments. SUIs organize interface content and interaction through 3D spatial relationships. However, despite the rapid proliferation of prototypes and system implementations, the existing contributions are largely confined to narrowly scoped application contexts, isolated interface components, or discrete interaction techniques, with limited cross-dimensional synthesis. As a result, we still lack a systematic understanding of how core design dimensions interact and co-evolve within MR environments. This absence of an integrative perspective impedes the identification of recurring design patterns and constrains the field’s capacity to surface emerging innovation opportunities. To address this gap, we conducted a systematic review of 108 publications across four key dimensions: Display Representation, Spatial Organization, Generation and Update Strategy, and Interaction Method. We further analyzed intra- and inter-dimensional relationships, synthesizing recurring cross-dimensional patterns and distilling their broader design implications. Building on our findings, we propose a four-part research agenda to advance the development of next-generation SUIs for real-world contexts and complex task environments. Our review and supplementary materials are available through an interactive online literature platform (https://spatialuserinterface.github.io/) to support continuous updates and community engagement. -
1342 An Evaluation of Spatial Window Switching Techniques for Mixed Reality Environments
Mixed Reality interfaces let users anchor virtual windows to specific, meaningful locations in the physical space. When the association between a window and a location is strong, accessing the window often entails users physically going to that location rather than bringing the virtual window to them. This task, which we call Spatial Window Switching (SWS), consists of two steps: 1) identifying the window of interest from an overview; and 2) reaching that window in physical space. We compare four design variations of SWS techniques, based on two overviews: a 3D world-in-miniature that preserves spatial relationships between windows but is prone to occlusion; and a 2D tiled overview inspired by desktop interfaces that is occlusion-free but breaks spatial relationships between windows. Both overviews are evaluated with and without animated visual guidance. We report on a comparative study across low and high window-density conditions, quantifying the performance improvements of SWS techniques compared to a baseline condition without assistance. Results indicate that 2D overviews perform best in high-density scenarios while 3D overviews better support spatial awareness and one-handed interaction. -
1465 User Preferences for UI Anchoring in XR: Effects of Task Mobility and Interface Properties
Anchoring - the choice of frame of reference for mixed reality (MR) interface elements - is a critical design decision involving trade-offs between accessibility, interaction comfort, and visual interference. Despite its importance, user preferences for anchoring across different mobility contexts and interface properties remain poorly understood, as prior work has largely focused on specific tasks or fixed interface configurations. We address this through a mixed-methods user study in which participants configure anchoring strategies across different mobility conditions and interface types. Combining behavioral analysis with structured qualitative inquiry, we analyze how participants select and reason about anchoring modes. Our results show a clear transition from world-anchored interfaces in stationary contexts to body-anchored interfaces during locomotion. However, no single body anchor consistently dominates, highlighting the personal nature of anchoring strategies. Our qualitative analysis reveals the factors users consider in their anchoring decision, including interface accessibility, stability during interaction, visual clutter, and individual mental models. These findings inform the design of adaptive and controllable MR interfaces and highlight the importance of supporting user customization. -
10.1109/TVCG.2026.3683941 Hybrid User Interfaces: Past, Present, and Future of Complementary Cross-Device Interaction in Mixed Reality
We investigate hybrid user interfaces (HUIs), aiming to establish a cohesive understanding and to adopt consistent terminology for this nascent research area. HUIs combine heterogeneous devices in complementary roles, leveraging the distinct benefits of each. Our work focuses on cross-device interaction between 2D devices and mixed reality environments, which are particularly compelling, leveraging the familiarity of traditional 2D platforms while providing spatial awareness and immersion. Although prior work has prominently explored such HUIs in the context of mixed reality, we still lack a cohesive understanding of the unique design possibilities and challenges of such combinations, resulting in a fragmented research landscape. We conducted a systematic survey and present a taxonomy of HUIs that combine conventional display technology and mixed reality environments. Based on this, we discuss past and current challenges, the evolution of definitions, and prospective opportunities to tie together the past 30 years of research with our vision of future HUIs.
Friday – 9 October 2026
PS37 - PS39 08:30–09:30
PS40 - PS42 09:00–10:00
PS43 - PS45 11:45–12:45
PS46 - PS48 12:15–13:15
PS49 - PS51 14:15–15:15
PS52 - PS54 14:45–15:45
Paper Sessions 37-3908:30–09:30
PS37: Multisensory Perception Sezione 1
-
1478 Perceiving Compliance in Virtual Reality: Effects of Interaction Direction and Sensory Discrepancy on Visuo-Haptic Weighting
In Virtual Reality (VR), users judge object compliance by combining visually displayed deformation with physically delivered force, yet it remains unclear whether both modalities require equally high rendering fidelity across interaction conditions, particularly when rendering resources are limited. We investigated visuo-haptic weighting in compliance perception across interaction directions and visuo-haptic discrepancy magnitudes. Two discrepancy magnitudes were tested across three representative interaction directions: a smaller magnitude comparable to values commonly used in sensory-weighting studies and a larger magnitude that remained perceptually acceptable in VR. Under the larger discrepancy, visual influence was greater in vertical top-down interaction, whereas haptic influence was greater in front-back interaction, while side-to-side interaction showed more balanced weighting. Directional effects were minimal under the smaller discrepancy. Together, these findings suggest that compliance rendering in VR may benefit from accounting for interaction direction, rather than assuming a single uniform visuo-haptic balance across interaction conditions. -
1709 Exploring the Effects of Olfactory Cues and Ventilation on Teleportation-based Navigation in VR
Humans interact with the world using various sensory cues, yet Virtual Reality (VR) experiences remain primarily limited to visual and auditory stimuli. The potential benefits of incorporating olfactory cues in 3D user interactions within VR environments have not been thoroughly explored. In this work, we investigate the usability of integrating olfactory cues into teleportation-based navigation in VR. We first introduce a wearable olfactory device prototype designed for head-mounted display (HMD) VR systems, capable of delivering olfactory cues and ventilating lingering scents to prevent scent overlap. Next, we conduct a user study to evaluate its effectiveness in teleportation techniques, specifically examining how olfactory cues support users’ spatial awareness, with a focus on object recognition during navigation. Our findings suggest that incorporating olfactory cues with ventilation in Teleportation and Dash enhances object recognition in VR environments. Finally, we discuss study limitations and propose future research directions for the effective integration of olfactory cues in VR. -
1883 Smelling the Way: Olfactory Modulation of Spatial Estimation and Path Integration in Virtual Reality
Spatial cognition enables individuals to perceive and interpret spatial relationships, estimate locations, and navigate within their surroundings. While vision plays a dominant role, other sensory modalities, particularly olfaction, can support spatial processing when visual cues are limited. This paper explores the contribution of olfactory cues to spatial cognition within virtual reality (VR) environments. We conducted two formal studies to evaluate the effectiveness of scent-based spatial information. Study 1 (N = 24) examined the participants' ability to localize a scent source while walking a linear path using two different olfactory devices. Study 2 (N=19) employed a triangle completion task, in which participants attempted to return to their starting point under varying rotation angles and olfactory conditions. Results from both studies suggest that olfactory cues can support localization accuracy, demonstrating the potential of integrating olfactory feedback into VR systems to enhance user localization in immersive environments. -
1898 Quantitative Analysis of Force Feedback-Based Stiffness Perception and Sensitivity in Video See-Through and Optical See-Through AR Displays
Human perception is influenced by visual dominance, making display type and occlusion important considerations in AR environments. This paper investigates stiffness perception in AR environment that provides the spatially aligned visual and haptic feedback of deformable objects to analyze how force feedback and corresponding visual cues influence the stiffness perception of a deformable object. This study employs 2 × 2 factorial design with display type (optical see-through (OST) and video see-through (VST) HMD) and the visual rendering of the stylus (haptic stylus). Participants interacted with the virtual deformable object with force feedback from a haptic device. A visuo-haptic dual-channel model was developed to provide realistic visual cue of deformation while maintaining the accurate control of force feedback. A psychophysical experiment evaluated perceptual thresholds, response behavior, and subjective ratings. The results revealed significant effect of stylus rendering, with higher perceptual thresholds indicating lower sensitivity, particularly in the highly immersive VST HMD. However, this condition also resulted in lower inter-participant variability due to the presence of a visual collision cue between the stylus and the virtual object. These results suggest a design trade-off between excluding stylus rendering for precise stiffness and including it for consistent interaction.
PS38: Pedagogical Agents Orione + Perseo
-
1097 Enhancing Electromagnetism Learning in Virtual Reality with an LLM-Prompted Pedagogical Agent Explaining Interactive Simulations
Electromagnetism concepts are difficult for students to learn because many underlying phenomena are not directly observable. While interactive simulations can support learning, learners often struggle to interpret observed behaviors and connect them to underlying physical principles. This paper introduces a pedagogical agent that delivers a lecture in immersive virtual reality (IVR) and explains the behavior of an interactive charged particle simulation during the lesson. Explanations are generated using a prompt-engineered large language model (LLM) that interprets the learner's recent interactions with charged particles and the current simulation state. The system uses a static system-level prompt combined with a dynamic runtime prompt constructed from lecture context, learner actions, and simulation state. We implemented this approach in a VR classroom where students attend a lecture on charged particles and Coulomb's law while interacting with a charged particle simulation that visualizes electric field interactions. As an evaluation, we conducted a between-groups study (N=60) compared three conditions: VR lecture-only (Lec), VR lecture with simulation (LecSim), and VR lecture with simulation plus LLM-prompted explanations from the pedagogical agent (LecSimExpl). Results show that LecSimExpl produced significantly higher conceptual learning gains than Lec and higher knowledge transfer gains than both Lec and LecSim. LecSimExpl also increased germane cognitive load, intrinsic motivation, co-presence, perceived knowledge, and perceived anthropomorphism. Participants in the simulation conditions also spent more time interacting with the lesson and exhibited higher head-oriented attention toward the agent. These findings suggest that pedagogical agents that dynamically explain simulation behavior using LLM-prompted explanations grounded in simulation state and user interaction can improve learning performance, user experience, and perceptions of the pedagogical agent in IVR learning environments. -
1098 Transforming Lecture Videos into LLM-Enhanced Immersive Learning Experiences
Video lectures have become a dominant medium for distance learning due to their accessibility and flexibility. However, their largely passive format limits opportunities for interaction, making it difficult to sustain learner engagement and support active learning. To address this limitation, we describe a workflow for transforming video lectures into interactive and immersive learning experiences. Following this workflow, a lecture setting is replicated from an input video and augmented with a pedagogical agent driven by a prompt-engineered large language model (LLM). The lecture experience includes a Q\&A mechanism that enables students to ask questions and receive responses from the agent. We evaluated the immersive learning experience developed with our workflow in a study against a video lecture (Video) and a desktop-based interactive version (Desktop). Our results show that the interactive conditions (i.e., Desktop and Immersive) significantly increased engagement compared with the passive condition (i.e., Video), while only the Immersive condition increased intrinsic interest. However, both interactive conditions also increased reported mental effort and perceived pressure relative to the passive condition. Importantly, participants in the Immersive condition achieved significantly higher learning gains than those in the Video condition, whereas no significant differences emerged between the Desktop and the other conditions. These findings highlight both the potential benefits and the practical trade-offs of transforming passive instructional media into interactive and immersive learning experiences. Based on these insights, we derive design guidelines for converting conventional video lectures into interactive and immersive learning experiences. -
1425 The Use of Pedagogical Agents in Virtual and Augmented Reality: A Scoping Review of Motivation through Self-Determination Theory
Pedagogical agents (PAs) are increasingly embedded in augmented/virtual reality (AR/VR) learning and training applications, yet their motivational impact is unevenly characterized. We present a scoping review that organizes how embodied PA designs support Self-Determination Theory (SDT) needs-autonomy, competence, and relatedness-and identifies methodological gaps. Following PRISMA guidelines, we searched IEEE Xplore, ACM DL, and SpringerLink for 2020–2025 publications combining agent, immersive medium, educational/training domain, and SDT terms. Of 2,701 records, 37 met inclusion after screening and full-text review. We synthesized results with a three-layer lens: Design, Theory, and Evaluation. The review highlights that current PA work is often positioned as motivational supporting; however, researchers frequently treat motivation as an assumed benefit rather than explicitly targeting it as a design objective grounded in motivation theory. Furthermore, current PA systems are designed mostly to support competence, revealing an opportunity to augment their support for autonomy and relatedness. Opportunities include embedding SDT explicitly in design (e.g., choice structures, non-controlling feedback, socially contingent cues) and broadening modalities (e.g., nonverbal, haptic, adaptive mechanisms) to strengthen competence and relatedness. -
1570 Student and Educator Perspectives on Embodied AI Agents in Education
As AI is beginning to take embodied form in classrooms (e.g., through avatars or immersive companions), it may fundamentally reshape how students learn. However, little is known about how such AI embodied agents might be used, what roles they could take, and what risks they carry. We addressed this gap through a three-phase study: \emph{(a)} scoping review of Extended Reality educational applications with embodied AI agents; \emph{(b)} participatory workshops with 23 students and educators; and \emph{(c)} a workshop with 9 HCI and education experts. Participants envisioned agents as tutors, accessibility companions, and teacher-training partners, with educators being more enthusiastic about adoption while students stressing the irreplaceable value of human teachers; yet both groups agreed that embodied AI should augment rather than replace teaching. These findings suggest that embodiment should be understood not as a technical property solely but as a relational quality that shapes trust, authority, and inclusion in education.
PS39: Hybrid Mobile Interfaces Sezione 2
-
1056 Challenges of Interactive Content Offloading for AR-enhanced Mobile Web Browsing
Lightweight Augmented Reality eyewear will soon offer an opportunity to extend mobile Web browsing beyond the physical phone, enabling users to interactively offload parts of a page into the space around them. But developing a system that effectively supports such user-driven content offloading poses both design and technical challenges. Informed by a set of mobile Web browsing scenarios, we identify three main challenges: Expressive Selection -- let users easily offload any Web page part they want; Stable Anchoring -- let users seamlessly place those elements in their surroundings, not only in their field of view or in the world, but around the phone as well; and Adaptable Rendering -- enable changes to the layout, styling and interactive behavior of offloaded elements to optimize them for display in mid-air. We rely on these challenges to structure the space of existing design options. Using two complementary probes, we explore design options addressing these challenges and assess their feasibility and integration within the constraints of current AR eyewear and Web technologies. We discuss the strengths and weaknesses of the different options to individual challenges, as well as the trade-offs involved in addressing all challenges jointly. -
1313 Modeling Across Realities (MAR): Extending Desktop 3D Modeling with Augmented Reality
Three-dimensional modeling (3D Modeling) is an essential part of creating digital content across various industries. However, most industry-standard tools rely on conventional desktop setups with 2D screens. One possible solution is to use Augmented Reality (AR), which allows users to perceive 3D shapes in a body-centric space better, but this increased spatial understanding comes at the cost of precision. To solve this problem, we present the MAR (Modeling Across Realities) system, a cross-reality 3D modeling system that combines the precision of conventional desktop-based 3D modeling with immersive AR visualization and interaction. We systematically evaluated MAR by asking 20 participants to model an abstract shape and an everyday object. Our results show that MAR outperformed desktop-based 3D modeling with lower workload, higher usability, and user preference without compromising the model quality in the abstract shape task, and improving model quality in the everyday-object task. These outcomes indicate that MAR is a promising cross-reality approach with the potential to inform the design of future 3D modeling systems. -
1557 Phonaze: Gaze-Hover with Torso-Supported Phone Confirmation for Ergonomic Interaction in Supine Mixed Reality
Mid-air gestures during supine extended reality (XR) use can cause rapid arm fatigue, limiting comfort and sustained engagement. We present Phonaze, a cross-device interaction technique for ergonomic supine XR that decouples pointing from confirmation: users acquire targets via the headset’s native gaze hover and confirm selections by tapping a smartphone resting on their torso. This configuration enables fully supported arms, reduces occlusion-related hand-tracking issues, and requires only minimal thumb motion. Phonaze is implemented as two native applications connected over local Wi-Fi, achieving 52 ms mean end-to-end tap-to-feedback latency. In a controlled within-subjects study on Apple Vision Pro (N = 18, Latin-square counterbalanced), participants performed discrete selection, list scrolling, and naturalistic video-platform browsing under Direct Touch, Gaze+Pinch, and Phonaze. Compared to Direct Touch, Phonaze reduced perceived arm fatigue by 75% (Borg CR-10: 1.9 vs. 7.6), improved selection time by 31%, and achieved high usability (SUS: 85.2). Behavioral analysis further showed 2.6× longer uninterrupted interaction flows. These findings demonstrate that posture-aware cross-device task allocation, using gaze for pointing and a torso-supported phone for confirmation, can substantially improve comfort in supine XR while maintaining competitive task efficiency. -
1751 DuoZone: A User-Centric, LLM-Guided Mixed-Initiative XR Window Management System
Extended Reality (XR) offers portable, private, and extensive workspaces by extending virtual displays beyond desktops. Current XR systems offer users limited manual control over placing, arranging, and resizing virtual windows for productivity work, leading to increased interaction costs and fatigue over time. We propose a mixed-initiative, human–AI window management system, DuoZone, that preserves users' agency while automating efficient window management. Through DuoZone's spatial zones, users can design the layout of a workspace where a cost-model-driven large language model (LLM) further places, orders, and adjusts relevant applications for multi-scenario efficiency. A user study (N=16) with Apple's Vision Pro found that DuoZone leads to significantly faster adjustments and lower effort and cognitive load than the existing baseline, while its method supports better scalability and may be beneficial for future automatic virtual window management.
Paper Sessions 40-4209:00–10:00
PS40: XR Privacy Glasshaus
-
1102 Who Are You in Virtual Reality? Suppressing Demographic Leakage via Progressive Adversarial Privacy
Virtual reality (VR) systems continuously capture multimodal sensor streams, including head motion, gaze direction, and physiological signals, to enable interaction and immersive experiences. However, this multimodal telemetry carries persistent biometric signatures that adversaries can exploit to re-identify users and infer sensitive demographic attributes such as age and gender, even from anonymized datasets. Existing privacy-preserving approaches, including noise injection and anonymization, degrade the temporal structure of the data, which is essential for downstream utility in VR. To address this gap, we present a Progressive Privacy Preserving method designed specifically for multimodal VR telemetry while maintaining downstream utility. Rather than optimizing against all demographic attributes jointly, the method separates physiological and kinematic feature domains and applies adversarial optimization sequentially. This progressive strategy prevents gradient interference between demographic traits that originate from distinct physiological signal sources, enabling targeted suppression of sensitive attributes while preserving behavioral patterns required for application utility. We evaluate the method on three publicly available VR datasets: Mazed and Confused, Simulation21, and WhoIsAlyx. Results show that demographic inference accuracy drops to near random chance, approximately 50% for binary attributes, while downstream utility task performance, including cybersickness severity classification, cognitive load classification, physical load classification, and action intent classification, remains strong, with reductions of at most 14.5% compared to unprotected data. On average, the method modifies approximately 11% of the original signal magnitude, indicating that the temporal structure remains largely intact. These findings demonstrate that sensitive demographic attributes in VR data can be effectively suppressed through progressive adversarial optimization while preserving signal fidelity for downstream utility. -
1250 Privacy on Autopilot: Examining XR Assistant Autonomy and Its Influence on Privacy Awareness
As Augmented Reality (AR) devices become embedded in everyday environments, their outward-facing cameras introduce continuous privacy risks by capturing sensitive information. While intelligent privacy assistants can mitigate these risks, little is known about how varying levels of autonomy reshape user responsibility and engagement in embodied AR contexts. We developed an AR privacy assistant with three autonomy roles and evaluated them in a within-subjects study (N=30) where participants performed a timed search task in a virtual home while managing privacy concerns. We collected objective performance metrics, subjective workload and trust ratings, and qualitative interview data. The Autonomous assistant significantly improved privacy concern detection accuracy without increasing subjective workload or degrading trust. Rather than eliminating human effort, autonomy redistributed cognitive responsibility, from environmental scanning to supervisory verification, prompting participants to adopt an “auditor” role when system judgments conflicted with their privacy expectations. While autonomy heightened situated perceptual awareness, it did not produce measurable gains in global privacy literacy, suggesting AR privacy assistants must balance transparency and alignment to support human agency. -
1475 Challenge-Response Virtual Reality User Authentication Using a Consumer-Grade EEG Headband
Virtual reality (VR) is increasingly used in security-sensitive domains, but user authentication in immersive environments remains underexplored. Current methods (e.g., passwords) are impractical and more susceptible to observation-based attacks. We present a challenge-response authentication pipeline for VR based on consumer-grade electroencephalography (EEG). Our pipeline combines a behavioral liveness module with an EEG authentication module built on an interactive VR task. Using a four-channel EEG headband, our dataset achieved an Equal Error Rate (EER) of 3.37% in single-session evaluation and 11.48% in cross-run evaluation, and remained competitive with several public datasets under a shallow machine learning benchmark. It also delivered the strongest channel-normalized cost-performance among all datasets considered, offering roughly 8× greater efficiency than 30–32-channel benchmarks. Usability evaluation results further indicate that the combined VR-EEG setup was generally well tolerated in short sessions. These findings support the feasibility of lightweight, deployable EEG-based authentication for VR applications. -
2067 Right-to-Be-Forgotten: Machine Unlearning–Based Privacy Protection in Artificial Intelligence-Driven XR Systems
The convergence of artificial intelligence (AI) and extended reality (XR) technologies (AI-XR) holds great promise for innovative applications in many domains. However, XR data (e.g., eye-tracking) is highly sensitive, and trained AI models can retain user-specific information that attackers can exploit. Hence, privacy regulations such as the European Union’s General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) grant individuals the \emph{right to be forgotten}, requiring the removal of personal data not only from databases but also from AI models trained on them when requested by a user. This makes post-hoc deletion of XR user data a critical requirement for deployed AI-XR systems. Existing state-of-the-art privacy mechanisms in AI-XR applications (e.g., differential privacy (DP)) primarily provide pre-training or sample-level protection and do not explicitly remove the influence of revoked XR users from already-trained AI models; satisfying deletion by retraining from scratch is often computationally prohibitive in real-world, large-scale AI-XR pipelines. Motivated by this, we propose a novel machine-unlearning-based framework for AI-XR that enables post-training revocation compliance by efficiently removing the influence of revoked XR users from trained AI models. Specifically, we design and evaluate two practical unlearning strategies: SISA (sharded, isolated, sliced, and aggregated training) and knowledge distillation (KD) to update AI-XR classifiers by removing revoked users’ learned contribution from the trained model parameters, producing an updated model that is equivalent to one trained model without those users, while preserving high utility on retained users and minimizing update cost compared to retraining the model from scratch. We evaluate the effectiveness of our proposed method on three SOTA AI-XR models and three different XR datasets (i.e., cybersickness, emotion, and activity classification). Our experimental results show that the proposed method preserves both privacy and utility (i.e., low accuracy on the forgotten-user set and high accuracy on the retained-user set) while substantially reducing model update cost compared to retraining from scratch. For instance, using a Transformer model, SISA reduces forgotten-set accuracy to 18\%, 15\%, and 19\% while maintaining retained-set accuracy of 95\%, 97\%, and 93\% for cybersickness, activity, and emotion classification, respectively, and achieves up to an $\approx 12\times$ speedup in update time over retraining from scratch. Privacy auditing via membership inference attack further shows that our proposed method reduces the attack success rate by up to $38\%$, $36\%$, and $40\%$ for the same classification tasks compared to the no-unlearning baseline. Finally, we validate these findings by constructing a new XR cybersickness dataset collected from 40 participants in a user study, which shows similar trends in privacy and utility, demonstrating that the proposed approach generalizes beyond benchmark datasets to a realistic AI-XR setting.
PS41: Cybersickness and Locomotion Sezione 4+5
-
1417 Attention Drives Sickness in AR
Cybersickness, a phenomenon of negative symptoms arising from visual motion, remains an obstacle to prolonged use of immersive displays. While cybersickness in virtual reality (VR) has been extensively studied, its mechanisms and mitigation remain underexplored in optical see-through augmented reality (AR). Unlike VR, AR preserves a direct view of the physical world while overlaying virtual content. This stable visual context fundamentally alters the visual–vestibular conflicts that contribute to cybersickness. In this paper, we investigate the causes of cybersickness in optical see-through AR and examine whether mitigation techniques developed for VR transfer to this setting. Specifically, we analyze how attention to virtual and physical content influences cybersickness under visual–vestibular conflict and whether reducing the virtual field of view (FOV) lowers sickness when users do attend to the virtual layer. Our results show that optical see-through AR can induce severe cybersickness when users attend to dynamic virtual content despite the presence of a grounding physical environment. In contrast, sickness levels are attenuated when attention is allocated towards the physical environment. We further find that reducing the global motion cues by a reduced FOV does not reliably mitigate AR cybersickness, suggesting that established cybersickness mitigation strategies from VR may not directly transfer to AR. Together, these findings demonstrate that AR cybersickness is an attention-dependent phenomenon with significant implications for AR interface design. -
1809 NPR Locomotion: Improving Spatial Awareness and Cybersickness with Sketch-Like Continuous Transitions
Architectural design review in Virtual Reality (VR) requires clients to move in the environment to understand building layouts. However, locomotion methods present a trade-off: continuous movement supports spatial awareness but causes cybersickness, while teleportation prevents cybersickness but does not support spatial orientation as much. We investigated whether applying sketch-style non-photorealistic rendering (NPR) during continuous locomotion in guided architectural tours reduces cybersickness while maintaining the information necessary for spatial comprehension of the design. In a within-subjects study (N = 36), participants completed guided tours in three virtual museums on three separate days using full color continuous locomotion (FCL), NPR continuous locomotion, or jump teleportation (JMP). Spatial awareness was measured via different tasks and tests including: point-to-landmark, egocentric distance estimation, map drawing, and object recognition tasks. Pointing accuracy in NPR was comparable to FCL (11.6◦ vs 10.6◦, p = 0.143, and Bayesian equivalence testing BF01 > 1000) confirming statistical equivalence. NPR significantly outperformed JMP (15.1◦, p < 0.001) and reduced post- exposure cybersickness compared to FCL(∆SSQ: FCL M = 28.6 vs NPR M = 21.6, p = 0.018). Real-time Fast Motion Sickness (FMS) scores showed that JMP produced the lowest discomfort (M = 0.88), NPR had an intermediate level (M = 1.19), and FCL showed the highest discomfort (M = 1.48). These findings demonstrate that deploying sketch-style rendering during guided transitions enables comfortable architectural explorations without losing the spatial comprehension for design evaluation. -
1896 To Blur or Not to Blur? A Cybersickness Dataset with Real-Time Reduction Techniques
Cybersickness, a form of motion sickness experienced in virtual reality, remains a significant challenge that limits prolonged user engagement. Researchers have explored various mitigation techniques to reduce user discomfort in VR; however, although several public datasets on cybersickness exist, datasets that include data collected under mitigation techniques are rarely publicly available. To address this gap, we present a user study with 38 participants in a VR roller coaster simulation where they experienced three visual conditions: no blur (NB), static maximum blur (MB), and dynamic blur (DB), where the blur intensity adapts in real time based on users’ self-reported sickness levels. We present a novel multimodal dataset that includes eye tracking, head tracking, and physiological signals such as electrocardiogram and galvanic skin response collected under conditions with and without cybersickness reduction techniques. In addition, the dataset contains subjective measures including fast motion sickness ratings reported every 30 seconds, simulator sickness questionnaire scores, and igroup presence questionnaire responses for evaluating users’ sense of presence in the virtual environment in different visual conditions. In addition to data recorded during the experience, the dataset contains baseline physiological recordings collected before the VR simulation. As an initial use case, we conducted a case study of training several deep regression models to predict cybersickness in which we achieved 1.58 RMSE. This open dataset allows future researchers to explore the relationship between physiological responses, user behavior, and cybersickness in VR. It can also support the development of predictive models and adaptive systems that select the most appropriate cybersickness reduction technique for each user, leading to more comfortable and personalized VR experiences. -
1280 Less is More: Step-Synchronized Mastoid Vibration for Comfort-Preserving Redirected Walking
Redirected Walking (RDW) enables users to explore large virtual environments within confined physical spaces by leveraging perceptual tolerances in self-motion perception. Previous research has shown that adding vestibular noise through electrical or continuous vibratory stimulation can increase redirection detection thresholds (DTs). However, electrical stimulation may compromise gait stability, while continuous vibration often induces discomfort. Consequently, alternative approaches that balance effectiveness and user comfort are needed. To address this issue, this study investigates whether temporally synchronizing vestibular vibration with gait events can enhance redirection tolerance while preserving comfort and locomotor stability. Thirty-six participants performed a curvature-based RDW task under three conditions: no vibration, continuous vibration, and step-synchronized vibration. Detection thresholds were estimated using a two-alternative forced-choice paradigm combined with psychometric modeling. In addition, subjective discomfort, simulator sickness, and presence were assessed alongside objective gait stability measures obtained through markerless motion analysis. Overall, both vibration conditions significantly increased curvature DTs compared to the no-vibration baseline. Notably, step-synchronized vibration produced curvature DTs comparable to continuous vibration while significantly reducing discomfort. Furthermore, markerless gait analysis revealed no significant differences in trunk sway or vertical body displacement across conditions. Taken together, these findings indicate that gait-synchronized vestibular vibration can effectively enhance redirection tolerance while maintaining user comfort and without introducing observable gait instability during walking-based RDW.
PS42: Medical Procedures Cigno + Auriga
-
1016 Dynamic, Static, or None? Virtual Magnification Loupe Designs for High Precision Tool Manipulation: A Dental Case Study
Precise tool manipulation with magnification loupes is fundamental in dentistry, yet the design of virtual magnification in immersive environments remains underexplored. This paper presents two Virtual Magnification Loupes (VMLs) co-designed with two senior dentists over six months: a Dynamic Virtual Loupe (DVL) providing head-coupled stereoscopic magnification and a Static Virtual Loupe (SVL) providing world-fixed magnification anchored near the intra-oral cavity. Both were integrated in a realistic immersive dental training environment featuring a physical dental treatment room with a phantom-based patient replica, a drill tool, and foot pedal-task confirmation to preserve the clinical workflow. We conducted a within-subjects study (N = 25) in which dental students performed high-precision drill positioning tasks under three conditions: DVL, SVL, and a No Loupe (NL) baseline. We measured positional error, rotational error, task completion time, cognitive load (NASA-TLX), and user preference. Results showed no significant differences in completion time or spatial error across conditions. However both VMLs significantly reduced perceived cognitive workload compared to NL. DVL yielded the lowest temporal demand and highest perceived performance, while SVL produced the lowest physical demand. In preference rankings; DVL was preferred by 15 dental students, SVL by nine, and NL by only one. These results indicate that VMLs can achieve equivalent task precision while significantly reducing cognitive load. The complementary strengths of both VML designs suggest that future immersive systems should support switchable loupe modes to accommodate varying task demands. -
1520 In Situ vs. Monitor-Based: Evaluating Augmented Reality for Ultrasound-Guided Peripheral Venous Catheter Insertion
Ultrasound-guided placement of peripheral venous catheters (PVCs) is a routine medical procedure that involves inserting a needle in a patient’s vein with the visual guidance of an ultrasound scanner to monitor the progression of the needle towards the vein. This procedure is cognitively demanding, as it requires operators to constantly divide their attention between the ultrasound display and the spatial relationship among the probe, the needle, and the vein. As a result of this difficulty, the procedure is associated with a high complication rate for patients. The goal of this paper is to evaluate the potential of augmented reality (AR) to facilitate the procedure for clinicians by displaying ultrasound (US) images in situ. To do so, we conducted a user study that compares two visualization modes: a conventional monitor-based procedure replicating the standard ultrasound display, and an AR-assisted procedure where the ultrasound image is spatially anchored directly under the probe and augmented with real-time distance cues (needle-to-skin and needle-to-ultrasound plane). The study was conducted within a virtual reality (VR) environment to provide a standardized and controlled setting. Our user study involved 30 participants from the healthcare domain, including 9 experts and 21 novices. Each participant performed ten insertions in both conditions. We analyzed procedure time, successful insertion attempts, and needle-beam alignment accuracy. We also assessed perceived workload (NASA-TLX) and system usability (SUS) using standardized questionnaires. Experts completed the procedure in 524.2 ± 154 s with AR, compared to 1125.7±472 s using the traditional monitor-based layout, whereas novices completed the procedure in 746 ± 410 s, compared to 1165 ± 597 s in the traditional setup. Spatial accuracy in novices increased from 50.6 ± 8.2% to 65.9 ± 13.6%, corresponding to a reduction in mean targeting error from 1.26 mm to 0.79 mm. The AR condition also greatly reduced tissue perforations in novices, decreasing from 29.4 ± 45.6 to 3.0 ± 2.7. Participants reported a lower perceived workload with AR, which may indicate that placing the ultrasound image directly under the probe helps reduce attention shifts and supports spatial understanding of the gesture. These findings indicate that AR-based in situ visualization can meaningfully enhance the efficiency, accuracy, and safety of ultrasound-guided procedures. -
2145 Visual Cue Interactions in AR-Guided Needle Insertion: A Prostate Biopsy-Inspired Phantom Study
Percutaneous needle procedures are minimally invasive interventions that involve inserting a needle through the skin to collect tissue samples, deliver medication, or treat specific conditions. Despite the apparent simplicity of the motor action, these tasks are challenging because clinicians cannot directly visualize anatomical targets and must instead rely on ultrasound images, which increase cognitive demand. Augmented reality (AR) has emerged as a valuable resource to provide assistance during the performance of these tasks, allowing for visualization of pertinent information in the clinician’s field of view. However, the use of naïve representations of augmented content frequently ignores meaningful visual cues, complicating depth perception and spatial understanding. In this work, we use and combine a set of visualization techniques designed to provide relevant information during the performance of percutaneous procedures: a localized focus-and-context window that exposes internal structures, a color-based proximity encoding that provides information regarding the distance between a tool and the organs, and an explicit trajectory overlay that provides directional information. These techniques, developed during design sessions with medical experts, were evaluated in a user study involving 26 participants, 7 of them clinical experts, using a prostate-biopsy-inspired phantom. Across objective and subjective measures, cue effects depended on both the surrounding cue configuration and user expertise. For novices, trajectory guidance improved targeting accuracy and reduced retreat behavior, but its effect on completion time depended on cue context and introduced a speed–precision trade-off when presented alone. For experts, the focus-and-context window reduced completion time and retreat events, while color feedback also improved completion time. Subjectively, proximity color cues were often perceived as helpful even when their effects on accuracy were not consistently positive. Overall, these results show that AR guidance cues are not uniformly additive and should be designed with attention to both user expertise and visual context in needle-based interventions. -
2163 The Role of Mixed and Augmented Reality in Medical Visualization: Literature Review and A Context-Aware Taxonomy
The discovery and evolution of medical imaging technologies have enabled non-invasive visualization of internal anatomy that has become essential for supporting diagnosis, monitoring, and treatment. However, because medical imaging relies on complex physical processes and contrast mechanisms for image formation, imaging alone is not sufficient to enable humans to fully leverage the resulting information. In addition, traditional methods to visualize the resulting information use two-dimensional displays to present three-dimensional anatomical structures. The introduction of Augmented and Mixed Reality (AR/MR) technologies offers an opportunity to provide valuable paradigms for medical imaging visualization, allowing users to observe, explore, and interact with anatomical information in more spatially intuitive ways. However, naive implementation without careful design considerations can lead to perceptual inconsistencies, potentially compromising utility and effectiveness. In this paper, we present a structured taxonomy of medical AR/MR visualization strategies aimed at providing clearer insight into how visualization design varies across clinical use cases. The taxonomy organizes techniques based on four core design components: image modality, data dimensionality, display technology, and clinical application. In addition, we introduce two critical dimensions that are often overlooked in the literature: visualization anchoring --the spatial relationship between virtual content and the physical world -- and perceptual awareness -- the use of visual cues to support spatial interpretation. Together, these components form a comprehensive taxonomy, offering a detailed framework for selecting appropriate visualization techniques in medical applications.
Paper Sessions 43-4511:45–12:45
PS43: XR Trust Glasshaus
-
1359 Believe It. Engage It.: A Mixed Reality Dataset of Trust and Engagement Under Attack
Mixed Reality (MR) systems are increasingly used in safety-critical domains such as medical training and emergency response, where users must interpret automated guidance while performing complex tasks. In these settings, misleading system feedback can disrupt engagement and degrade performance. Such disruptions can arise from incorrect guidance delivered through textual instructions, visual overlays, or spoken prompts, which may increase cognitive load and interfere with correct actions. Existing research lacks resources that systematically vary deceptive guidance modalities and their intensity to study their behavioral impact. To address this gap, we introduce a mixed reality dataset and machine learning benchmark that capture user behavior under deceptive system feedback during a Tactical Combat Casualty Care training scenario. We collected behavioral telemetry from 29 participants exposed to misleading textual, visual, and auditory cues, including synchronized eye gaze, field of view traces, and interaction logs. The dataset also includes psychometric profiles from 17 validated instruments measuring traits such as self-efficacy and locus of control, along with continuous workload and presence scores. Statistical analysis reveals a significant behavioral adaptation effect (p < .001), showing that users progressively learn to filter misleading guidance and recover task efficiency. As a predictive baseline, an Explainable Boosting Machine classifier achieves over 80% accuracy and a micro AUC of 0.92 in detecting engagement degradation using gaze and interaction signals. By releasing this dataset with strong baselines, this work supports research on engagement dynamics and anomaly-aware modeling in immersive systems. -
1528 CAESAR: Mixed Reality Search-and-Rescue Platform for Cognitive Attack Evaluation
Mixed Reality (MR) enables immersive training by tightly coupling virtual content with the physical environment and supporting embodied interaction. In safety-critical collaborative tasks such as search-and-rescue, however, users rely heavily on virtual elements for perception, coordination, and decision-making, making them vulnerable to cognitive attacks that subtly distort visual feedback, virtual content, or system synchronization. While prior MR security research has examined individual attack mechanisms, evaluations are typically conducted in isolated or single-user settings, and there is a lack of standardized collaborative testbeds for systematic adversarial experimentation. In this paper, we develop the CAESAR testbed, an end-to-end multi-user MR platform designed for systematic evaluation of cognitive attacks in search-and-rescue training. CAESAR integrates four classes of cognitive attacks within a synchronized collaborative scenario and incorporates a multimodal data collection pipeline that captures behavioral, performance, and subjective measures under adversarial conditions. We conduct a controlled user study with in-depth analysis of the latency-based attack, demonstrating its impact on task performance, user behavior, and subjective experience. -
1588 Hidden Ethics in XR Research: A Five-Year Scoping Review of ISMAR
Extended Reality (XR) technologies are increasingly deployed across domains. However, despite growing calls for Responsible XR, it remains unclear whether ethical considerations are explicitly addressed in XR research or appear implicitly within technical problem formulations. We conducted a scoping review of 567 papers from the IEEE International Symposium on Mixed and Augmented Reality (ISMAR) proceedings (2021-2025) and identified 45 papers containing ethically relevant themes. We analyzed how ``ethically charged'' concepts appear in XR research and whether they correspond to explicit ethical reflection or are framed as technical challenges. We found that explicit ethical engagement is rare. Ethical concerns are typically embedded within technical discourse, particularly in work on user experience optimization, safety and risk mitigation, behavioral influence, and data practices. We synthesized these patterns into a framework that surfaces the ``hideouts of ethics'' in XR research, and outlined recommendations for integrating ethical reflection into XR design and evaluation. -
2093 Robust Engagement Estimation in Mixed Reality Under Cognitive Attack
Engagement is a critical operational state in immersive systems, where users must sustain attention to virtual content despite distraction, workload, and degraded system responsiveness. In MR, one important source of disruption is a cognitive attack: an adversarial or system-level perturbation to the interaction loop that interferes with attention, situational awareness, or stable task execution. In this paper, we study induced latency as a concrete MR-relevant cognitive attack and present a two-stage model for estimating engagement from multimodal time-series data, including eye tracking, head motion, scene dynamics, and system latency. The first stage uses contrastive learning to encode variable-length behavioral trajectories into compact embeddings that suppress nuisance variation while preserving engagement-relevant structure. A lightweight regression head then maps these embeddings to a continuous engagement score. We evaluate the framework in a controlled hostage-rescue MR study with step-wise latency perturbations and minute-level engagement self-reports. Across cross-validation and held-out-user evaluation, the proposed model performs favorably relative to traditional regressors on pooled features, indicating that temporally structured multimodal representations provide a measurable benefit for engagement estimation under attack. We further show how formal verification can be used to derive conservative safe operating intervals for controllable MR parameters, connecting engagement estimation to runtime reasoning in safety-aware MR systems.
PS44: Agents for Guidance and Persuasion Sezione 4+5
-
1356 Can Virtual Agents Sell? Effects of Advertisement Delivery Method and Message Type in VR Shopping
As virtual reality (VR) shopping environments become more widespread, understanding how advertising messages are most effectively delivered in immersive contexts has become increasingly important. Prior research has examined communication modalities such as text, audio, and video, but largely in non-immersive media settings. In VR shopping environments, however, advertisements (ads) can also be delivered through embodied virtual agents, introducing spatial co-presence and richer social cues. Despite these benefits, little is known about the comparative effects of ads delivered via text, audio, single agent, and multiple agents within the same advertising context. To address this question, this study compares these four delivery methods and two message types (informational vs. emotional) in a VR shopping environment. We evaluate their effects through a within-subject experiment (N = 72) in a virtual shopping mall using a 4 × 2 factorial design. Results show that agent-based delivery increases purchase intention, trust, and user engagement compared with text or audio delivery. Informational messages increase trust, whereas emotional messages increase engagement. We also observed an interaction effect in which the advantage of agent-based delivery becomes more pronounced for emotional messages. Ranking comparisons further support this pattern, with the emotional multiple agents condition most frequently ranked highest. These findings highlight the importance of socially rich delivery methods and provide guidance for aligning delivery methods with message strategies in VR commerce. -
1437 Not All Pedestrians Are the Same: Effects of Real-World Pedestrian Visualization on Intent Recognition and Task Distraction
Users wearing head-mounted displays in immersive virtual reality (VR) often lose awareness of people in the surrounding physical environment. This reduced awareness can lead to safety issues, such as collisions with pedestrians, as well as missed opportunities for real-world interaction. Prior work has attempted to notify VR users of nearby people through auditory alarms or by switching to passthrough views, but these approaches can disrupt immersion and interfere with ongoing tasks. Moreover, existing work rarely considers pedestrians’ interaction intent, and visualizing all pedestrians indiscriminately may introduce unnecessary distraction during VR tasks. We investigate how pedestrians with different intents should be visualized within the virtual environment. We conducted a user study manipulating pedestrian intent, visualization method (shape and transparency), and task-space size. Our results show that humanoid representations improve users’ ability to identify pedestrian intent but also increase task distraction compared to capsule representations. Additionally, we propose intent-dependent visualization strategies to preserve users’ task engagement: humanoid representations for pedestrians likely to interact with the user or stand close (Interactors and Bystanders), and capsule representations for pedestrians who are merely passing by (Passersby). These findings highlight the importance of designing intent-aware visualizations of real-world pedestrians in immersive VR environments to balance situational awareness and task distraction. -
1593 How Self Embodiment and Agent Position Affect Safety-Conscious Behaviors and Performance in Guiding Peer Agents Across Traffic in VR
This study investigates how embodiment and safety-conscious behaviors shape participants' strategies and performance when guiding virtual peer agents across traffic in immersive road-crossing scenarios. Participants were embodied in the virtual environment, physically walking in tracked space while motion sensors and the HTC VIVE system mapped their movements into the virtual environment. They were tasked with safely leading same-gender virtual peers across a one-way road with continuous traffic. Participants needed to remain vigilant and carefully guide their peers to ensure safety for both. Two between-subject factors were also included: Participant Gender (Male vs. Female) and Self-Avatar Presence (Full-Body Avatar vs. Iconic Avatar). Two within-subject factors were manipulated: Agent Position (Trailing vs. Beside) and Vehicle Proximity Relationship (participant positioned closer to approaching vehicles vs. peer positioned closer). Results revealed that crossing strategies and safety outcomes were significantly shaped by agent position, vehicle proximity, gender, and avatar presence. Walking side‑by‑side with a virtual peer supported safer, more conservative behaviors with larger time margins and fewer collisions, whereas trailing disrupted coordination and increased risk. Vehicle positioning relative to traffic also mattered: peers opposite the vehicle (AUV) led to hesitation and smaller safety buffers, while same‑side positioning (UAV) promoted earlier entry with larger margins and safer crossings. Gender and avatar presence further influenced strategies, with females generally maintaining larger safety margins and avatars amplifying cautious behavior, while males showed more collisions in trailing conditions and distinct timing adjustments depending on peer positioning. These findings have important implications for VR applications in road‑crossing training and education, underscoring the need to account for peer positioning, traffic context, and user characteristics when designing effective and safe learning environments. -
1683 Inside the VRdrobe: How AI Agent Embodiment and Plurality Shape Persuasion in Virtual Outfit Selection
Recent advances in generative Artificial Intelligence (AI), particularly Large Language Models (LLMs) and Vision–Language Models (VLMs), enable intelligent assistants in interactive environments. In Virtual Reality (VR), such agents communicate through text, speech, or embodied interaction and may influence user decisions. However, the effects of agent form (chat vs. embodied avatar) and plurality (single vs. multiple agents) on persuasion in immersive settings remain underexplored. We study how embodiment and plurality affect responses to persuasive AI in VR using VRDrobe, a virtual try-on system where users select an outfit and receive persuasive recommendations under four conditions: Single Chat, Multi Chat, Single Avatar, and Multi Avatar. We measure behavioral revision, perceived persuasion, social presence, decision time, and workload. Results show that embodiment increases perceived social presence, while plurality increases perceived persuasion and decision time. Social presence is strongly associated with persuasion, suggesting embodiment contributes to persuasive experience even when plurality has the stronger direct effect. Behavioral revision shows a weaker pattern, leaning more toward plurality. These findings highlight that social presence and persuasive impact are related but distinct design outcomes, providing guidance for XR systems in which AI agents act as socially active recommendation interfaces.
PS45: Immersive Analytics Cigno + Auriga
-
1010 Understanding Organizational Strategies Across Multimodal Artifacts in Immersive Computational Notebook
Immersive Computational Notebooks (ICoN) extend traditional notebook environments into immersive spaces, enabling analysts to interact with multimodal artifacts, including code, narratives, data tables, and visualizations. In ICoN, analysts can seamlessly transition between analytical tasks by integrating multimodal artifacts within a single immersive workspace. Meanwhile, understanding organizational strategies is critical for designing effective interactions to further support analysts. However, prior research on immersive computational notebooks has primarily examined organizational strategies centered on single-modality artifacts. Systematic investigations of how analysts spatially arrange and coordinate multimodal artifacts in immersive environments remain underexplored. To address this gap, we conducted a user study to examine organizational strategies for multimodal artifacts in immersive computational notebooks. Our findings show that participants predominantly adopted depth-based layouts, and their spatial organization was largely structured around cell-based artifacts. -
1574 DP-LENS: A Hybrid AI-Assisted Density-Aware Lens for Occlusion Management in 3D Immersive Analytics
Immersive environments, spanning Virtual and Augmented Reality (VR/AR), offer a unique approach to exploring complex 3D datasets, where data is often heavily occluded and exploration incurs a high cognitive load. We proposed a Hybrid AI-Assisted DP-LENS framework accordingly. While preserving peripheral context through geometric deformation and 3D perspective techniques, it enables users—with the assistance of AI and algorithms—to explore 3D data with a lower cognitive load. We conducted two user studies with 34 participants to investigate the potential benefits of combining a density-aware polyfocal fisheye lens with intent-driven AI navigation. Our first study (N=18) compared the DP-LENS against two industry-standard baselines (i.e., World-in-Miniature coupled with manual zooming, and volumetric slicing) in heavily occluded 3D datasets. The results show that the DP-LENS significantly reduced cognitive load, decreased completion time, and improved user preference. Following these results, the second study (N=16) compared a hybrid-controlled system featuring AI-assisted navigation with a fully manual DP-LENS. The results show that the AI-assisted hybrid system improved task efficiency, further reduced cognitive load, and garnered higher user preference. Furthermore, the hybrid system partially decoupled exploration efficiency from the physical dimensions of the data and mitigated physical fatigue to some extent. Based on the findings, we proposed design implications to inform the development of more spatially scalable and low-fatigue interactions for future 3D visual analytics systems. -
10.1109/TVCG.2026.3672314 I Feel Like Iron Man”: Authoring, Exploring, and Presenting Data Visualizations in Immersive AR
A truly anytime, anywhere approach to immersive analytics is predicated on the ability for users to author and explore visualizations in an immersive fashion and without a desktop computer. In this paper, we present a generalizable visual programming technique for immersive authoring, exploration, and presentation of data visualizations through direct manipulation, and implement it in a web-based augmented reality (AR) system. In a preregistered user study, we use this implementation to explore how analysts author and use data visualizations in immersive AR accessed using a head-mounted display (HMD). The study focuses on the authoring technique, spatial organization strategies, and sensemaking and presentation practices in such immersive analytics environments for HMD-based AR. Our findings suggest that participants particularly liked the direct manipulation of our approach, and that participants’ spatial organization is not based on physical landmarks in the room. -
1421 A Design Study on Voice-based Interaction for Immersive Network Visualization and Analysis
Visual network analysis leverages network visualization authoring techniques to facilitate sensemaking, serendipitous discovery, and hypothesis verification on network data. However, transferring the same paradigm to immersive environment is non-trivial due to insufficient UI affordance for authoring operations. Researchers have studied combining multiple modalities for interactions, but the high learning curve of such input systems limits their adoption by typical data analysts, let alone for network analytics. In this work, we investigate the advantages and limitations of voice as the primary input modality with a research-by-design (RtD) study, in which we design a system that supports voice-based interactions for immersive network visualization facilitated by Large Language Models (LLMs). Through a user study on social network data analysis with participants from social science and computer science backgrounds, we find that voice interactions provide improved usability over controllers and reduce cognitive efforts in formulating intended verbal instructions as concisely as possible. We discuss design implications for immersive visualizations, highlighting how usability limits adoption while simplified interactions and voice-based controls enhance fluidity and support complex operations.
Paper Sessions 46-4812:15–13:15
PS46: Passthrough Perception Sezione 1
-
1456 The Perceptual Cost of Passthrough: How Video See-Through HMDs Degrade Human Visual Perception of Acuity, Contrast, and Color
Video see-through (VST) technology aims to seamlessly blend the virtual and physical worlds by reconstructing reality through cameras. However, while manufacturers promise high perceptual fidelity, it remains unclear how accurately these systems replicate natural human vision, particularly across varying environmental conditions. In this work, we introduce a low-cost testing method to quantify the perceptual gap between the naked eye and three popular VST headsets: Apple Vision Pro, Meta Quest 3, and Meta Quest Pro. Using adapted psychophysical measures, we evaluated participants' visual acuity, contrast sensitivity, and color vision under both normal and low-light conditions. Our results demonstrate that despite recent hardware advancements, all tested VST systems fail to match the dynamic range and adaptability of natural human vision. While high-end devices can approach human performance in ideal lighting, they exhibit significant visual degradation in low-light environments, particularly concerning contrast sensitivity and visual acuity. By mapping these specific perceptual limitations, this work establishes a definitive benchmark for current VST capabilities and highlights the need for experience design compensations to achieve truly indistinguishable mixed reality experiences. -
1507 Passthrough Rigidity: The Behavioral and Visuomotor Costs of Mediated Perception
Broad public adoption of video passthrough head-mounted displays remains elusive despite significant market investment, and as a community we lack a precise understanding of why users experience persistent discomfort even as hardware factors such as resolution and latency have dramatically improved. This paper establishes the first mechanistic link between passthrough usage and discomfort, specifically regarding the impact on visual behavior and physiology. We developed a novel protocol to capture synchronized oculomotor, kinematic, and physiological data during a block assembly task requiring complex hand-eye coordination. Using a within-subject design (N=110), we evaluated both natural and passthrough viewing conditions. Our results reveal a four-fold suppression of rotational head velocity (33.4°/s vs. 133.3°/s, p < 0.05) and a fundamental decoupling of head-gaze coordination – which we coin “Passthrough Rigidity”. This phenomenon appears to shift the information-gathering burden to the oculomotor system, resulting in significantly longer fixation durations (304 ms vs. 279 ms, p < 0.05) and restricted visual search patterns. These kinematic shifts directly correlate with poorer task performance (natural accuracy 85.5%, passthrough 82.1%) and measurable physiological cost, evidenced by a significant reduction in blink duration (229ms vs. 261ms, p < 0.05) and increased reports of ocular strain and cognitive load. We conclude that current passthrough implementations force a transition from flexible exploration to motor caution, where task performance is preserved at the cost of user comfort and biomechanical efficiency. These findings provide a critical quantitative framework for evaluating and improving future XR devices, establishing that resolving "comfort" for passthrough requires addressing the deep-seated biomechanical compensations caused by mediated perception. All data is made publicly available at https://removed.for.double.blindness. -
1806 How Opacity and Background Affect Surface Contact Perception in Optical See-Through AR
For virtual objects to appear spatially anchored in Augmented Reality (AR), users must be able to tell whether they rest on a real surface. Humans judge this primarily through cast shadows and contact edges (the visible boundaries where a virtual object meets the physical surface). However, additive blending in Optical See-Through (OST) displays weakens cast shadows, leaving contact edges as the main available cue. How rendering parameters and environmental factors jointly shape this edge-based judgment remains poorly understood. We investigated this on an OST display with segmented dimming (Magic Leap 2) by manipulating virtual object opacity (20\%--100\%), background spatial frequency (0.5, 1.5, 4.5 cpd), and ambient lighting (500 lx, 1500 lx). A within-subjects study (N = 39) assessed contact judgment accuracy, signal detection sensitivity (d′), confidence ratings, and eye-tracking behaviour. Results show that 60\%–80\% opacity yields the highest accuracy and efficiency, as this range keeps the virtual object visible without occluding the surface texture needed to judge contact. This advantage was most pronounced over mid-frequency backgrounds (1.5 cpd). Notably, users feel most confident at 100\% opacity despite lower objective performance, revealing a disconnect between perceived and actual judgment quality. Eye tracking reveals two distinct failure modes: at 100\% opacity, relatively minimal gaze effort at the contact edge suggests overconfidence-driven underinspection; at 20\% opacity, prolonged fixations indicate that increased effort cannot compensate for insufficient surface texture. Together, these findings show that the most visually solid rendering is not the most perceptually effective, pointing toward contact-state-driven opacity adaptation as a design strategy for OST-AR. -
1817 When Registration Error Becomes Ambiguous: Behavioural Effects of Static Misalignment in Augmented Reality
Static registration error can impair user performance in augmented reality, but its behavioural and subjective effects remain insufficiently characterised. While existing works often evaluate acceptable registration quality as a purely quantitative geometric problem, this perspective overlooks how spatial misalignment creates perceptual ambiguity for the user. In this paper, we demonstrate that AR registration requirements might not be defined by absolute offset magnitude alone, but also by the physical layout and the boundaries of cue-target ambiguity. We report a controlled video see-through augmented reality study with 24 participants using a reproducible 3 × 3 button-acquisition task, in which a visual cue was displaced relative to physical targets across six offset magnitudes. User performance remained close to the zero-offset reference at the smallest offsets, but degraded significantly under larger misalignments. Participants also showed short-term adaptation after the switch to misalignment, although larger offsets retained substantial residual costs. Subjective ratings showed parallel increases in rating with performance. These findings provide a basis for setting task-specific registration requirements by identifying offset ranges that remain behaviourally tolerable, rather than those that produce ambiguity-driven breakdown under explicit performance criteria.
PS47: VR Locomotion Sezione 2
-
1077 Teleporting to Any Height: Comparative Evaluation of Steering-Free 3D Locomotion Techniques for Immersive VR
Teleportation is widely used in immersive virtual reality, as it enables efficient travel and reduces cybersickness compared to continuous steering. However, most teleportation techniques are designed for horizontal ground-based movement and offer limited support for navigation to arbitrary 3D positions. In this paper, we comparatively evaluate three teleportation techniques for mid-air destinations: a previously proposed Two-Step technique in which horizontal position and elevation are specified sequentially, a ray-based approach that extends conventional pointing teleportation with controllable placement of safe mid-air destinations, and a novel technique in which users place teleport targets on a rotating plane dynamically coupled to head pitch. A within-subject study with 25 participants compared the three techniques in a target-based flying task and a large-scale wayfinding task. The results show that the ray-based technique achieved the best overall performance across most measured dimensions, leading to lower workload, higher usability, faster task completion, and fewer teleport actions than the other techniques. At the same time, the head-coupled plane technique was more frequently associated with precision and spatial orientation, revealing a trade-off between efficiency and spatial control. These findings demonstrate the potential of 3D teleportation for immersive VR and contribute to ongoing research on alternative approaches for arbitrary-height locomotion. -
1096 EMSWalker: Integrating Gait-Synchronized Electrical Muscle Stimulation for Seated Walking-in-Place in Virtual Reality
We present EMSWalker, a system that provides proprioceptive feedback during seated walking-in-place by delivering gait-synchronized electrical muscle stimulation (EMS) to the leg muscles involved in walking (quadriceps, tibialis anterior, and gastrocnemius). By mapping foot tapping to the human gait cycle, EMSWalker renders virtual walking sensations for level, uphill, and downhill locomotion while users remain seated. Through three user studies, we evaluate (1) cognitive remapping from seated walking-in-place to walking, (2) discrimination of terrain-specific EMS patterns, and (3) EMSWalker’s effect on virtual walking experience. Results show that EMSWalker significantly increased proprioceptive drift and embodiment relative to other conditions, enabled above-chance discrimination of slope-specific stimulation patterns, and improved presence in a VR walking experience with flat and sloped terrain. These findings demonstrate that gait-synchronized, terrain-adaptive EMS can enrich seated VR locomotion by providing more coherent proprioceptive feedback and a more realistic, spatially immersive walking experience. -
1879 A Low-Cost Prism-Based Module for Walking-in-Place Locomotion in Mobile VR with a Trajectory-Based Metric for Continuous Locomotion
To enrich interaction in constrained mobile VR systems, this paper introduces a low-cost hardware module that uses a simple prism to enable vision-based Walking-in-Place (WIP) on standard smartphones. We evaluated this controller-free system in a user study against Look-Down-To-Move (LDTM). Methodologically, we propose and apply Fréchet distance as a more robust metric for analyzing continuous and trajectory-based locomotion, moving beyond traditional accumulative metrics like total distance traveled. The empirical findings both align with and extend prior research. We replicate the key result that WIP is significantly more immersive than LDTM without a cost to objective performance. To further investigate subjective experience, we utilized a 7-point Likert scale questionnaire. The results were consistent with findings from categorical preference questionnaires but captured more nuance, revealing a critical trade-off: while confirming WIP's superior immersion, participants found both techniques reliable, although LDTM was rated as more reliable, highlighting a tension between embodiment and system efficiency. -
2335 Looking Around by Looking Around: Omnidirectional Gaze-based VR Viewport Control
Traditional VR viewport control relies on head and torso movement, which can be effortful and limiting in both constrained and extended-use settings. We introduce Looking Around by Looking Around (LALA), a gaze-based VR pitch-and-yaw viewport control technique designed for natural and effortless omnidirectional exploration via eye movements, without requiring or obstructing movement of the head, hand, or body, offering a low-effort and highly accessible interaction method. Because gaze is primarily used for perception and exhibits oculomotor and perceptual asymmetries, using it directly for control is difficult. To address this, we designed an asymmetric omnidirectional control profile for the eye, then built on it to exploit tendencies for eyes to stay within comfortable regions for viewport control. We evaluated LALA in a user study (N=18) featuring two contrasting tasks: alignment towards known directions and open-ended visual search towards unknown directions. LALA was strongly preferred over the traditional baseline, achieving competitive performance while enabling fully hands-free interaction with minimal physical movement
PS48: Human Performance and Assessment Orione + Perseo
-
1934 EgoBalanceFormer: Predicting Human Balance from Egocentric Pose and VR-Native Motion
Maintaining balance is important for safe interaction in immersive virtual reality (VR), where visual–vestibular conflicts can affect postural stability. Most prior works on VR-based balance analysis rely primarily on motion signals from the head-mounted display (HMD), which capture head movement but provide only limited information about full-body biomechanics. As a result, directly predicting biomechanical balance indicators, such as the Center of Pressure (CoP), from VR sensing data remains challenging. In this paper, we present EgoBalanceFormer, a multimodal framework for predicting human balance during VR interaction by leveraging egocentric pose reconstruction as a body-aware sensing modality together with VR-native HMD motion signals. Unlike prior VR balance approaches that rely primarily on sparse device motion such as HMD trajectories, our approach uses egocentric vision to reconstruct full-body pose and derive biomechanical cues—including lower-body kinematics and center-of-mass (CoM) dynamics that are critical for modeling balance control. These pose-derived biomechanical cues are fused with VR-native HMD motion signals and processed through temporal encoders to capture motion dynamics over time. We further apply bidirectional cross-attention between lower-body joint kinematics, CoM dynamics, and rotational features to capture complementary biomechanical relationships. The refined multimodal representations are fused and decoded using a transformer with a learnable CoP query to regress CoPX and CoPY. We evaluate EgoBalanceFormer against several machine learning and deep learning baselines for CoP prediction. Experimental results show improved accuracy and stronger robustness across participants, demonstrating that egocentric pose-derived biomechanical cues significantly enhance balance prediction compared with using VR-device motion signals alone. -
2024 BoredomVRoom: A TOVA-Inspired VR Task for Boredom Induction
Boredom is a recognized but empirically underexplored factor affecting performance and safety in professional environments. In industrial workplaces, its consequences pose significant risks to health and well-being. Despite growing interest in using Virtual Reality (VR) for studying occupational phenomena, the field lacks standardized, ecologically valid procedures for reliably inducing and analyzing boredom in controlled experimental settings. This paper presents the exploratory evaluation of BoredomVRoom, a VR-based boredom-induction task inspired by the Test of Variables of Attention (TOVA), adapted as a repetitive quality control operation within a neutral virtual environment. In a user study (N=32), where the task was administered as a baseline condition within a broader, multi-condition VR protocol, a 15-minute task led to a significant increase in self-reported boredom, as measured by the Multidimensional State Boredom Scale (MSBS). This induction was accompanied by a significant decrease in affective valence and arousal, consistent with the theoretical affective signature of boredom. The raw NASA-TLX revealed low cognitive demand, high self-rated performance, and moderate frustration. This pattern aligns with the characterization of boredom as a state arising from cognitive underload. Furthermore, participants significantly overestimated task duration, a known temporal marker of the construct. These converging results suggest the proposed task's potential as a boredom-induction procedure for VR-based human-computer interaction research in industrial contexts, while acknowledging that its embedded administration necessitates future isolated normative validation. The application and Unity source code will be released as an open-source package to serve as a foundational tool for boredom induction in XR research. -
2251 Toward Postural State Classification in Immersive VR with Multimodal data and Explainability Analysis
Ensuring a safe virtual reality (VR) experience requires effective systems that can predict and respond when users lose their balance. While prior research has focused on predicting falls and motion sickness in VR, most approaches have been regression-based and have not fully explored postural state classification. This study presents an exploratory comparative analysis of machine learning (ML) and deep learning (DL) algorithms for classifying postural states in VR environments affected by visual perturbations. We used a publicly available multimodal dataset that contains kinematic, electromyographic (EMG), and electrodermal activity (EDA) signals. The data were prepared for a binary classification task that distinguishes balanced and imbalanced postural states. Participant-wise downsampling was applied to address class imbalance in the dataset. The ML and DL models that we compared, were then evaluated under a Leave-One-Participant-Out (LOPO) cross-validation protocol. Among the models, the Mamba-inspired CNN (MI-CNN) achieved the highest accuracy of 96.76%. The most important factors influencing the postural state classification were identified using SHapley Additive exPlanations (SHAP) analysis, which enhanced the interpretability of the model. The results from SHAP analysis indicate that kinematic data features are the most dominant in classifying postural states. We also evaluated the MI-CNN model using only the top two-third of features ranked by SHAP importance. Despite a 66% reduction in input dimensionality, the model retained strong classification performance (0.957 ± 0.022 accuracy and 0.957 ± 0.0218 F1-score), with only a marginal decrease of approximately 1% compared to the full-feature model. This study is an initial step toward using multimodal sensing and advanced learning models to better understand balance-related instability in immersive VR environments. The ability to accurately classify an imbalanced postural state may support early awareness of potential fall risks. Also, the findings contribute to the design of safer and more adaptive VR systems. Such systems may respond to early signs of instability and help improve safety and user experience in immersive applications. -
2254 Towards the Threshold: A Fall-Risk Anchored Pareto Framework for Virtual Reality Gait Feedback Selection for Individuals with Multiple Sclerosis
Reduced walking speed in people with multiple sclerosis (MS) is closely associated with a higher risk of falls. But VR rehabilitation studies often overlook whether improvements in performance come with acceptable cognitive and physical effort. This study introduces a threshold-anchored efficiency framework that evaluates both performance and overall effort simultaneously. Here, a normative velocity target (T=1.29 m/s) was established from the mean walking speed of non-fallers in an independent MS gait dataset, representing the speed at which MS patients walk when falls are not occurring. 34 adults with MS were then tested across eight VR feedback conditions (unimodal, bimodal and multimodal). Each condition was evaluated on how much it closed the velocity gap to T and at what cognitive and physical burden cost. Pareto efficiency analysis identified five non-dominated conditions : Static Visual, Spatial Auditory, Auditory+Visual, Auditory+Vibrotactile, and Multimodal -- while Spatial Vibrotactile and Vibrotactile+Visual were dominated. Spatial Auditory offered the best efficiency, closing 72.7% of the gap at below-average burden. Multimodal closed the most gap (95.6%) but imposed the highest burden. These findings give clinicians a principled basis for selecting VR feedback strategies matched to each patient's capacity.
Paper Sessions 49-5114:15–15:15
PS49: Object Physicality Sezione 1
-
1448 How Physics Fidelity Shapes Perceived Plausibility in Mixed Reality: A Case Study of Bounce Behavior
Mixed Reality (MR) systems increasingly place virtual objects on visible real-world surfaces, yet it remains unclear how users judge the physical plausibility of these interactions when rebound behavior departs from reality. We study this question through ball bouncing, a simple but expressive MR interaction in which plausibility depends on both object properties and surface-dependent response. In a 4 × 2 within-subjects study, 20 participants interacted with virtual tennis and pingpong balls on three real surfaces under four rebound models, and evaluated each model before and after direct familiarization with the corresponding real balls and surfaces. Bounce realism significantly affected perceived plausibility (p $<$ .001), and this effect depended on the familiarization phase (p $<$ .001), whereas phase alone showed no overall effect (p = .399). Before familiarization, participants showed only weak differences across modes, but after familiarization they clearly distinguished the calibrated model from the alternatives. These findings suggest that, in familiar MR interactions, preserving the expected surface-dependent structure of motion may matter more than strict physical accuracy, offering practical guidance for designing future MR experiences that balance realism and system complexity. -
1625 Modulating Perceived Affordances in Augmented Reality Through Object Physicality and Thermal Appearance
In Augmented Reality (AR), users interact with objects whose visual appearance can be decoupled from their physical properties. This can create a discrepancy between the actions enabled by the AR system and the actions users perceive as possible, commonly referred to as perceived affordances. While AR overlays virtual content onto real environments or objects, it remains unclear how such augmented visual information shapes perceived affordances during interaction with virtual objects or real objects augmented with virtual effects. To address this question, a 2 × 4 within-subjects study was conducted with 32 participants, who were asked to grasp and move a cup under different object modalities (real, virtual) and visual content conditions (empty, neutral, hot, cold). Kinematic measures of reaching and grasping behavior were analyzed to assess how visual information related to object physicality and content representation influences movement. Results show that augmented visual cues, particularly liquid and thermal appearances, produced more pronounced effects in the real cup condition during the pre-grasp phase of movement. Although the task remained identical across conditions, systematic variations were observed in both reaching and grasping phases depending on the visual characteristics of the object. These findings demonstrate that visual signifiers in AR can subtly but consistently modulate interaction kinematics, providing insight into how perceived affordances are shaped in augmented environments and informing the design of AR systems that leverage visual cues to guide physical interaction. -
1632 Grasping Real and Virtual Objects in AR: Study of Carry-Over Effects on Grasping Anticipatory Behavior
When alternatively interacting with virtual and real objects in an Augmented Reality (AR), users experience sensory discrepancies, notably differences in visual and tactile feedback. While prior work has demonstrated that interacting with both virtual and real objects leads to a decrease in grasping performance for real objects, the underlying impact on the grasping process itself remains unclear. To explore how virtual object interactions can influence subsequent interactions with real objects, we conducted a pick-and-place experiment (N = 28), where participants alternated interactions between real and virtual cylinders using two common virtual grasping methods, a pinch gesture or a physics-based grasping method. We measured the grasp anticipatory process, i.e. pre-shaping, on subsequent interactions with real objects (maximum grip aperture and time to peak aperture) and transition time. Results show that virtual interactions, no matter the virtual grasping method used, influenced pre-shaping towards subsequent real-object reach motion. Moreover, the physics-based grasping method significantly showed the largest grip aperture and the earliest preparation for subsequent interactions with real objects, suggesting greater grasp uncertainty. This study contributes to the understanding of interactions with both virtual and real objects, highlighting the importance of analyzing, not only traditional performance metrics, such as task completion time, grasping dynamics, such as pre-shaping. -
2258 Effects of Virtual Hand Anthropomorphic Fidelity and Controller End-Effectors on a Near-Field Complex Dynamic Interception Task in VR
As immersive virtual reality (VR) rises in popularity and is applied in various fields, it will be used to simulate increasingly complex and dynamical interactions. In order to facilitate interaction with these environments, users are often provided with representations of their end-effectors in the virtual world. However, the design of users' self-representations can affect how they perceive and interact with virtual objects. The interplay between the fidelity of users' virtual hand representation and the standard controller based interaction may differentially affect user performance. Inspired by prior work on simple and dynamic affordances in VR, we examined a unique phenomena of how the anthropomorphic fidelity of the virtual hand representation, as compared to controller based end-effector representation, affects near-field or personal space complex dynamical interception task (CDIT) performance in VR. In a novel experiment scenario, participants had to successfully intercept and retrieve a moving target object, while avoiding moving non-target obstacles in the near-field or personal space. The speed and size of the target object randomly varied between trials. Furthermore, the frequency of the opening and closing of an obstacle doorway that participants had to avoid also varied between trials. Thus, the task involved a CDIT affordance for participants to interact with in the near-field, or close-quarters, in VR. Additionally, when the users' end-effector was represented as a virtual hand, we examined the effects of anthropomorphic fidelity of the virtual hand as either high-fidelity realistic or low-fidelity iconic representation on users' perception-action coordination in CDIT performance as a between-subjects factor. Results indicate that the anthropomorphic fidelity of the virtual hand representation and the interaction method of controller versus the virtual hand affects the near-field dynamical target interception and retrieval performance in VR.
PS50: Adaptive and Personalized XR Sezione 4+5
-
1168 TeAPOT: Smartphone-Assisted Terrain-Aligned Geospatial Registration for Head-Mounted Mixed Reality in Outdoor Environments
Many outdoor Mixed Reality (MR) applications require placing virtual content at geographic locations. However, most standalone Extended Reality (XR) headsets lack built-in Global Navigation Satellite System (GNSS) and magnetometer sensors, forcing developers to rely on specialized external hardware. Moreover, for terrain-aligned situated visualization, correct geospatial placement alone is insufficient. The system must also reconstruct the local ground surface so that virtual content conforms to uneven terrain. We present TeAPOT (Terrain-Aligned Positioning and Orientation Toolkit), which uses a commodity smartphone as the sole external sensor to provide XR headsets with GNSS position and compass heading, while performing real-time, on-headset reconstruction of a 2.5D terrain patch for terrain-aligned placement of virtual content. In outdoor experiments with a Meta Quest 3 and iPhone 16 Pro Max, TeAPOT achieves a median end-to-end position error of 2.04 m (P95 = 5.52 m), with up to 10 m standing position from the target reference, across 150 trials at three locations. For local terrain alignment over a 1 × 1 m patch, evaluated across 30 trials across three terrain types, TeAPOT achieves a median absolute height error of 0.6 cm (P95 = 1.1 cm), produces a stable terrain estimate after a median of 1.1 s, and incurs a median GPU time of 5.6 ms on the headset. These results demonstrate that smartphone-assisted geospatial registration, combined with on-headset terrain reconstruction, enables practical outdoor geospatial MR with meter-scale global placement and centimeter-level local terrain alignment, without specialized hardware. -
1344 FlowAvatar: Real-Time Full-Body Avatars from Sparse Egocentric Inputs on Consumer XR Devices
Prior avatar embodiment research has studied tracking completeness and personalized appearance in isolation, each requiring specialized lab equipment. How these factors interact on consumer hardware remains an open question, because no existing system unifies body pose, face, hands, and personalized appearance on a standalone headset. We present FlowAvatar, a system enabling real-time personalized look-alike avatars with full-body pose, hand articulation, facial expression, and eye gaze tracking from sparse egocentric inputs on consumer XR devices (Meta Quest Pro). Our pose model achieves state-of-the-art accuracy on the AMASS benchmark against prior works with comprehensive spatial and temporal losses validated through architecture and loss ablations, while running in real-time. Two user studies (N=12, N=10) provide converging evidence that tracking completeness and personalized appearance serve complementary roles - tracking completeness broadly improves perceived motion quality, while only their combination achieves positive total embodiment. -
2218 User Profiling and Modeling in Extended Reality: A Scoping Review
Extended reality (XR) systems increasingly seek to adapt to individual users by leveraging behavioral and contextual data captured during immersive interaction. Central to this capability is the construction of user profiles or user models, which represent user characteristics such as identity, states, or preferences to guide system behavior. Despite growing interest in XR personalization, research on how user profiles and models are constructed and applied remains fragmented across disciplines and application domains. This paper presents a scoping review of 161 studies on user profiling and modeling in XR. Our analysis shows that most XR user profiles are derived from behavioral interaction traces, particularly head, hand, and eye movements, and are primarily used for user identification, state inference, and adaptive system responses. We further observe that modeling methods are closely tied to the computational tasks they support, with learning-based models dominating behavioral inference and rule-based approaches often used for descriptive attributes. The findings suggest that XR user profiling and modeling is evolving from isolated recognition tasks toward behavior-aware infrastructures that support adaptive immersive systems. -
2200 Exploring Object Recall in an Object-Rich AR Office Scene
Augmented reality changes how users perceive and attend to their surroundings. Prior work suggests AR can reduce attention to the physical world, but it remains unclear whether this holds in environments dense with both virtual and physical objects, and which factors most affect object recall. This study examines recall in a mixed-reality office scene after an object-selection task, testing object-level factors (virtuality, size), context-level factors (semantic congruency, saliency), and user-level factors (cueing, dual tasking). Cueing informed participants of a later recall test, while dual tasking added cognitive load through a secondary audio task. Virtual objects were recalled more accurately than physical ones. Cueing produced a modest improvement in recall of present objects, especially under added audio-task demands. Using playback software to reconstruct each participant’s view throughout the trials, we computed object-level saliency and found a modest positive correlation with recall, slightly stronger than the effects of cueing or virtuality. These findings have implications for AR interface design and for building mixed-reality environments that better support cognition during real-world tasks.
PS51: Rehabilitation Screening Orione + Perseo
-
1558 Beyond Score-Based Gamification: Designing Spatiotemporal and Musical Experiences for VR Neck Rehabilitation
Pain-related anxiety and fear of movement are major barriers to adherence and therapeutic outcomes in rehabilitation exercises for chronic neck pain. Virtual reality (VR) enables the design of immersive experiences that can transform repetitive therapeutic movements into engaging and emotionally supportive interactions. In this work, we investigate how experience-oriented gamification can reduce anxiety and improve user experience during VR-based neck range-of-motion (ROM) exercises. We introduce two novel interaction paradigms that embed therapeutic neck movements within multisensory VR experiences. The first paradigm, Spatiotemporal Progression, couples head-tracked trajectories with environmental progression in a tropical island setting, where movement segments dynamically transform time of day, weather, and spatial location as experiential rewards. The second paradigm, Musical Interaction, maps movement segments to meditative musical notes layered with relaxing ambient soundscapes. We evaluate these designs against a conventional score-based gamification baseline in a controlled user study with 20 participants. We assess usability and user experience through subjective measures, exercise performance with motion tracking, and anxiety modulation using the Subjective Units of Distress Scale (SUDS), heart rate, and skin conductance. Our findings suggest that, in comparison with traditional score-based gamification design, immersive environmental and musical feedback have better potential to reduce anxiety and improve user experience, with little to no impact on successful performance of the exercise. Our findings highlight the value of experience-based interaction design for VR rehabilitation, suggesting an alternative to performance-centric gamification that prioritizes emotional engagement without compromising therapeutic efficacy. -
1965 VRScout for AD: An Immersive Virtual Reality System Towards Early Screening of Alzheimer’s Disease
Alzheimer's Disease is a growing global health challenge, with early detection critical to effective intervention. Current screening methods such as biomarker detection and paper-based cognitive scales face limitations including cost, invasiveness, and cultural or age-related biases. In this work, we introduce VR Scout for AD, an immersive virtual reality system specifically designed for the early screening of AD. The system comprises two scenario modules: a test-oriented module incorporating tasks adapted from clinically validated tools like the Montreal Cognitive Assessment and Color Trails Test, and a life-like module simulating daily activities to evaluate memory, attention, and executive function in a naturalistic environment.We conducted five iterative prototyping sessions with neurologists and older adults with and without cognitive decline, collecting qualitative feedback and System Usability Scale scores. Through repeated revisions, we improved system accessibility, usability, and comprehensiveness. The final version received a score of 91.9, indicating excellent usability and strong user acceptance. Our findings suggest that VR Scout is a promising, user-friendly tool for scalable, non-invasive early AD screening. -
2058 StrokeLens XR: Clinician Perspectives on an AI-Enabled Immersive System for Early Stroke Assessment
Context: Early stroke assessment is highly time-critical, yet Neurologist involvement in the acute phase is often constrained by fragmented communication and delayed patient visibility. Objective: To investigate clinicians’ perceptions of StrokeLens XR, an immersive AI-supported system for early remote stroke assessment, focusing on perceived clinical usefulness, usability, workflow integration, adoption readiness, and trust. Method: We conducted a mixed-method evaluation comprising focus groups with five neurology residents who interacted with a functional prototype in simulated scenarios, followed by a TAM-based questionnaire completed by 13 clinicians. Qualitative and quantitative results were analyzed and integrated through triangulation. Results: Clinicians reported strong perceived usefulness, particularly in enabling earlier Neurologist visibility, improving communication with first responders, and utilizing transport time for meaningful assessment. The system was generally viewed as intuitive and compatible with existing workflows. Adoption readiness was primarily influenced by reliability, rapid activation, training requirements, and protocol integration. Participants consistently emphasized that the system should function as a clinician-support tool that augments judgment rather than replaces it. Conclusion: StrokeLens XR shows promise as an immersive decision-support system for time-critical stroke care when it enhances shared situational awareness, preserves clinician oversight, and integrates smoothly into existing workflows. These findings provide early empirical insights into clinician acceptance and implementation considerations for immersive AI-supported assessment systems. -
1881 Evaluation of 3D Locomotion and Cybersickness Mitigation in Cycling-Based VR Exergames
Cycling-based VR exergames are a promising approach to motivate sedentary users to exercise. However, cybersickness (CS) concerns often limit the design of such games, particularly the incorporation of rotational movements. Current evidence on the relationship between rotational motion and CS during exergaming is primarily limited to 2D locomotion in biking contexts. Additionally, previously evaluated CS countermeasures for cycling-based exergames have shown limited effectiveness. This research evaluated three 3D locomotion methods with varying degrees of rotational freedom and a dynamic rest frame CS countermeasure for cycling-based VR through a controlled user study with 35 participants. Results indicate that while introducing rotational control may increase CS, it does not necessarily affect motivation, exercise intensity, or preference for the locomotion method. The dynamic rest frame effectively reduced CS overall and was most effective in conditions with full rotational freedom. Despite this benefit, many users disliked the rest frame because it obstructed their view during gameplay. These findings suggest that incorporating rotational control into cycling-based VR exergames can expand design possibilities without substantially hindering player motivation or exercise performance. The results may inform exergame designers on how to implement 3D locomotion and CS mitigation techniques to create more comfortable and engaging VR exergames.
Paper Sessions 52-5414:45–15:45
PS52: Visual Cueing Glasshaus
-
1030 Magnification Lenses for Small-Scale Augmented Reality Instructions
Small-scale objects that are difficult to see or distinguish from their surroundings are common in many assembly and maintenance procedures. Augmented reality (AR) instructions involving such small parts may be difficult to perceive, potentially causing misinterpretations. This paper explores the use of AR visualizations for small-scale instructions. In particular, we present the design and evaluation of AR magnification visualizations. Through iterative design, we explore how magnified AR views impact user behavior, visual attention, and subjective perception in the context of AR instructions. Our results demonstrate the effectiveness of live AR magnification in improving task guidance, providing a strong foundation for future work on AR-assisted instruction and precision interaction design. -
1736 Seeing Through the Surface: Evaluating Visualization Techniques for Depth Perception in Augmented Microscopy
The ability to see through the surface of the human body is central to medical augmented reality (AR) applications ranging from surgical planning to intraoperative guidance. However, accurately perceiving the depth of virtual structures relative to real anatomy remains a persistent challenge. This is particularly pronounced under the operating microscope, where high magnification, limited viewpoints, and tight error tolerances make depth and occlusion both essential and difficult to infer. Prior work has proposed methods for visualizing subsurface content using Video See-Through (VST) and Optical See-Through (OST) head-mounted displays (HMDs). However, the implementation of these approaches in surgical microscopes remains an open question due to the additive characteristics of microscope optics and display systems. To the best of our knowledge, this work is the first to investigate the perceptual accuracy of AR microscope visualization techniques for depicting virtual content inside the human body. We introduce a custom depth perception measurement apparatus based on a half-silvered mirror that enables alignment between a real measurement probe and subsurface virtual targets while viewing through a stereoscopic surgical microscope. We then propose a custom alpha-blending method and a virtual aperture that conforms to the three-dimensional surface geometry. Using our apparatus, we conducted a user study evaluating depth estimation accuracy and confidence across visualization techniques. Results showed that both the virtual aperture and alpha-blending significantly reduced absolute depth error compared to standard opaque rendering. Alpha-blending without an aperture achieved localization accuracy comparable to opaque rendering with a virtual hole while eliminating the systematic depth underestimation bias observed in traditional renderings. Furthermore, the presence of a virtual aperture increased users' confidence and was preferred over other methods. These findings provide insights for the design of mixed-reality visualization systems for microsurgical applications. -
1737 GAAC: Gaze-Aligned AR Cueing for Mitigating Inattentional Blindness
Inattentional blindness (IB) is a perceptual-attention phenomenon in which visual information enters the eye but fails to reach conscious awareness, thereby impairing situational awareness and leading to missed detection of low-probability yet high-risk events. In this paper, we present GAAC, a gaze-aligned augmented reality (AR) cueing framework that integrates scene perception and eye tracking to estimate IB occurrence and provide warning cues with AR display. GAAC combines an eyewear-based eye tracker with a forward-facing camera to localize anomalies, align gaze evidence with target ROIs, and enable attention-contingent AR cueing with latency suitable for interactive on-device use. To evaluate the framework, we design a controllable IB-induction paradigm in an assisted-driving task that enables repeatable measurement across groups. The user study shows that both non-contingent and attention-contingent cueing reduced IB relative to a no-cue baseline, while the attention-contingent condition achieved a lower IB rate. Subjective workload increased under non-contingent cueing but remained at baseline levels under attention-contingent cueing. These findings suggest that attention-contingent AR cueing is a promising approach for mitigating IB while limiting unnecessary user burden in attention-demanding settings. -
2066 Beyond the Black Patch: A Composable, Gaze-Contingent Framework for Simulating Central Vision Loss in XR
Central vision loss (CVL) from macular diseases such as age-related macular degeneration and Stargardt disease is commonly simulated in extended reality (XR) using a single mechanism, typically a dark patch or blur overlay. Yet clinical reports suggest that people with CVL reject this depiction and instead report a broad range of perceptual effects, including spatial distortion, perceptual invisibility, dynamic noise, dimming, filling-in, and temporal disturbances. We present XR-VIEU, a modular, real-time, gaze-contingent framework for simulating CVL in XR through seven clinically grounded, independent, composable shader modules implemented in Unity HDRP, with formal mathematical specifications and over 120 configurable parameters per eye. To probe whether our simulation can capture lived CVL experiences, we conducted a co-participatory design study with two individuals with late-stage Stargardt disease. Despite sharing the same diagnosis, the two co-designers required significantly different module configurations: one profile was dominated by content displacement and dimming, while the other relied primarily on dense fractal noise. These cases provide initial evidence that modular, composable simulation is needed to represent the heterogeneity of CVL experience. The full seven-module pipeline adds only 0.20 ms of rendering overhead (2.9%) while maintaining approximately 90 Hz on the Varjo XR-4 at native resolution. XR-VIEU provides a foundation for more faithful and personalized CVL simulation in research, clinical education, and accessibility-oriented design.
PS53: Viewpoint Transitions Sezione 2
-
1243 Cognitive Performance During Display Switching in Optical See-Through Augmented Reality
Optical see-through (OST) augmented reality (AR) displays can render virtual content directly in the user's view of the physical world. In certain use cases, the physical world can include fixed emissive displays. For example, pilots, ship captains, and power plant operators equipped with AR displays might switch their visual attention between AR elements, physical interface components, and imagery on emissive flat panel displays fixed in the environment. Switching attention between these targets involves aspects of dynamic perception and cognition. We conducted a controlled study using XREAL One Pro AR glasses to examine the cognitive effects of switching attention between content displayed on a nearby flat panel display (FPD)/ screen and content displayed on a head-worn AR display (HWD), under two ambient brightness levels (10 lx and 150 lx). Participants completed an adapted Paced Visual Serial Addition Test (PVSAT), a continuous cognitive task that measures information processing speed and working memory under sustained demand, while switching between displays on every trial. Contrary to expectation, we found no general cost of switching over repeating a display. Instead, performance costs were specific to the direction of the transition: transitioning to the screen was significantly slower than transitioning to the HWD, no matter where the transition began. Higher ambient brightness improved response speed, reduced missed trials, and did not increase subjective visual fatigue. These findings suggest that in unpredictable multi-display environments, what display you transition to matters as much as the act of switching itself, with implications for the design of OST AR interfaces where transitions between head-worn and external displays are frequent. -
1615 Between Places with Transitional Spaces: A Walkable WIM in Transitional Spaces for VR Navigation
Teleportation techniques for Virtual Reality (VR) locomotion that use a World-in-Miniature (WIM) for target selection typically rely on instant relocation. Although efficient, this abrupt transition can impair spatial orientation and reduce users’ sense of presence. To address this issue, this paper introduces transitional spaces as VR navigation concept. Users enter a transitional space containing a scaled-down walkable WIM, move (walking/steering) to the in- tended destination, and confirm their target position. They then exit the transitional space through a second portal, creating a percepti- ble and continuous transition rather than a direct relocation. In an exploratory user study, we compared classic WIM-based telepor- tation, portal-Augmented WIM, and the proposed walkable WIM. The results show a trade-off between efficiency and orientation: the walkable WIM achieved the best spatial orientation performance, but also has the longest completion times and a higher cognitive load due to additional navigation overhead. At the same time, both portal-based approaches provided a more positive user experience than the classic WIM, including a stronger sense of presence and higher hedonic quality. In addition, the Walkable WIM showed the significantly shortest reorientation time. -
1730 Contextual Blocking Entities: Transformation of Guardian Boundaries in Virtual Reality
Conventional VR guardian systems communicate physical boundaries through grid-based overlays that effectively signal spatial limits, but disrupt narrative coherence and presence. This tension between safety and immersion reflects a trade-off that existing approaches have not fully resolved. To address this, we propose Contextual Blocking Entities (CBEs), a boundary design approach that transforms guardian boundaries into narrative-integrated virtual objects by embedding spatial constraints within the internal logic of the virtual world. We define CBE design through three dimensions: Function Match, which signals impassability through visual form, Context Consistency, which determines how naturally the object belongs to the world, and Constraint Type, which determines how the boundary geometry engages spatial processing. To operationalize these dimensions, we first conducted a formative study classifying 80 candidate virtual objects, then carried out a controlled within-subject user study using a 2x2x2 factorial design to compare CBE conditions against a grid baseline across presence, spatial awareness, and boundary compliance. All CBE conditions produced significantly higher presence than the grid baseline, while four conditions showed comparable safety to the baseline. The three dimensions operated asymmetrically, each governing a distinct stage of boundary interaction. Function Match governed initial approach deterrence, Context Consistency supported corrective behavior following a violation, and Constraint Type shaped movement range independently of conscious spatial judgment. These findings show that the safety-presence trade-off is not inherent to boundary systems, but depends on how spatial constraints are expressed within the virtual world, and provide guidelines for virtual object selection and placement in room-scale VR. -
2142 Transition Techniques for Externally-Guided Multi-Scale Viewpoint Changes
Extended reality (XR) is increasingly used to help users understand complex virtual environments through multiple viewpoints across different levels of immersion, positions, and scales. While numerous techniques address viewpoint transitions for self-guided exploration, many scenarios require externally-guided transitions where a system or presenter controls the user's viewpoint, leaving the user with limited spatial knowledge and control over the transition process which can increase susceptibility to disorientation and discomfort. We present three transition techniques for externally-guided multi-scale XR viewpoint changes and evaluate them against a fade-to-black baseline in a within-subjects study (N=20). Participants transitioned between world-in-miniature, street-level, and indoor destination views. We combined spatial recall measures, standardized questionnaires, and semi-structured interviews to assess orientation, workload, comfort, and continuity. Results showed that techniques externalizing reference frames and the user's pose improve multi-scale spatial recall relative to a fade, while same-scale recall was insensitive to transition type. We conclude with design implications for multi-scale XR viewpoint transitions.
PS54: XR Analytics and Measurement Cigno + Auriga
-
1031 In Sight of Insights: Virtual Representation of Boxing for Interactive Viewpoint in Tactical Analysis through 3D Human Motion Reconstruction
Combat sports such as boxing require rapid tactical decision-making and precise spatio-temporal awareness, yet current analysis primarily relies on video replays and statistical charts. Fixed viewpoints and occlusion constrain video replays, while statistical data are displayed separately, requiring users to correlate information across sources mentally. This separation constrains effective reflection on athlete movements and tactical interactions. Although prior sports analytics research has explored event detection and action recognition, few systems support motion reconstruction and integrated spatial analysis using real match data in combat sports. To address this gap, we conducted a formative study with professional boxing coaches and developed In Sight of Insights, a 3D interactive tactical analysis system that reconstructs athlete motion and integrates match events within a unified virtual environment. By embedding contextual information directly within reconstructed scenes, the system enables spatially coherent observation and interaction. We evaluated our system through two user studies with professional athletes and experienced spectators. Results show that Free-viewpoint Virtual Representation improves observation accuracy and analytical efficiency compared with conventional viewing modes. Participants also reported that free-viewpoint perspectives and integrated contextual visualization enhanced tactical understanding. These findings demonstrate the potential of reconstructed 3D environments with embedded analytics to support more effective combat sports analysis. -
1529 Zero-Install WebXR for City-Scale Environmental AR: Deployment and Empirical Evaluation of Situated Comprehension
Urban environmental sensing networks generate dense spatial data describing atmospheric conditions across city neighborhoods, yet this information remains largely inaccessible to residents within the environments it describes. We present a WebXR-based handheld augmented reality system that enables walk-up public access to city-scale environmental visualization through the mobile browser without application installation. The system integrates minute-level data from approximately 280 active weather stations across a 138-square-mile metropolitan area and visualizes atmospheric conditions as geospatially anchored particle fields within a browser-delivered urban model. Users access the experience by scanning QR codes, launching the AR interface in under ten seconds. A 48-hour deployment audit characterizes end-to-end pipeline latency, spatial binding accuracy, and rendering performance across heterogeneous mobile devices. A controlled in-situ within-subjects user study ($N = 16$) comparing the AR interface against a browser-based map presenting identical sensor data demonstrates that situated AR visualization significantly improves self-reported spatial comprehension of neighborhood-scale environmental conditions ($p = 0.005$, $d = 0.81$). The comprehension advantage is explained by reduced spatial translation effort: map users reported significantly greater cognitive work to relate the display to their physical surroundings ($p = 0.031$, $d = 0.89$). Preliminary behavioral comprehension probes with an independent sample provide initial convergent support for the self-report findings. These results position standards-based WebXR as a viable civic interface layer for distributed urban sensing infrastructures. -
1611 Collaborative Immersive Analytics: A Systematic Review
Collaborative Immersive Analytics (CIA) has emerged as a promising approach for supporting collective sense-making of data using immersive technologies. Over the past decade, a growing number of systems have been proposed, yet the design space and collaborative practices in CIA remain fragmented across the literature. This paper presents a systematic review of CIA research published between 2015 and 2025. Following PRISMA guidelines, we analysed 19 empirical studies across six dimensions: time--space configuration, analytical tasks, system-centred symmetry, user-centred symmetry, collaboration dynamics, and data characteristics and domain applications. Our results show that CIA systems predominantly supported synchronous collaboration and shared exploratory analysis coordinated through natural communication mechanisms such as verbal communication and deictic gestures. However, many systems were evaluated using datasets selected primarily to demonstrate system capabilities rather than real-world analytical applications. We conclude by outlining future directions for CIA research. -
2289 From Motion to Meaning: A Systematic Review of Motion-Tracking Measures in XR-Based Experimental Psychology
Immersive technologies such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) are being increasingly used in experimental psychology because they enable controlled environments while simultaneously producing high-frequency motion-tracking data from headsets, controllers, and body trackers. These motion traces provide continuous behavioral signals that may reflect affective, behavioral, and cognitive processes during human–human, human–computer, and human–virtual human interactions. Despite the growing use of motion measures in XR (Extended Reality) experiments, there is limited synthesis of how such data are collected, transformed, analyzed, and interpreted. This paper presents a systematic literature review of studies collected from the ACM Digital Library, IEEE Xplore, PsycINFO, PubMed, and Web of Science that use motion-tracking data from XR devices in experimental psychology, focusing on affect, behavior, cognition, or their intersections. The review examines the devices used, the body parts tracked, and the data formats employed in these studies. More importantly, it analyzes how raw motion-tracking data streams are transformed into kinematic or trajectory-level features and the analytical methods used to link them to psychological constructs. Our findings highlight substantial variability in device configurations and motion data extraction methods, but we observe a common pattern of using descriptive statistical measures to summarize motion patterns or applying machine learning models to learn from processed motion features. Based on these observations, we outline practical recommendations for selecting tracking setups, collecting and validating motion data, and designing analysis pipelines in experimental psychological studies. This review aims to support researchers seeking to incorporate motion-tracking data into XR-based psychological experiments and to promote more consistent and interpretable use of motion measures in immersive research.
For questions, contact: papers2026@ieeeismar.net