Hero Image

Posters


Last updated: 2026-08-28 7:14PM EDT


All accepted posters, including those invited from the Doctoral Consortium and TVCG Journal papers, are assigned to a specific day, which has multiple presentation timeslots throughout that day. We highly recommend being at your posters during the allocated timeslots! This is a great opportunity to network and talk to other conference attendees.

This year, there are no traditional poster fast-forward sessions. Instead, fast-forward poster videos that you submitted will be showcased at the exhibition hall.

** Please put your poster on the board in the exhibition hall before the first presentation timeslot and clean up your poster after the last presentation timeslot.

WednesdayWed Oct 7 ThursdayThur Oct 8 FridayFri Oct 9

Wednesday – October 7, 2026

Posters 1ACassiopeaVisualization, Rendering & Immersive Experience Wednesday10:30–12:45

  • P1342 Texturized Surfel-Based Dynamic Reconstruction: Comparing Visual Perception with Meshed-TSDF Bruno Caby, Guillaume Bataille, Florence Danglade, Jean-Rémy Chardonnet
    Real-time dynamic reconstruction is a cornerstone technology to maintain spatial awareness over time in collaborative scenarios. Most of existing methods do not handle dynamic changes in real-time, while neural based approaches require extensive preparation time. We present a reconstruction method that addresses both limitations. We performed a user study (N=23) comparing our method against a meshed baseline using signed distance fields in a practical VR scenario, aiming to show comparable user experience. We found no significant differences for perceived quality, task completion time, and task load, with most variations remaining within a 15\% reference margin, indicating a comparable user perception.
  • P1544 What is Lost in Immersion? How Virtual Reality Reshapes Spatial Understanding of 1920s German Expressionist Cinema? Rafaella Siagkri, Alexandra Covaci, Gordana Korolija Fontana-Giusti, Ambrose Gillick, Marios Constantinides, Fotis Liarokapis
    Virtual Reality (VR) is increasingly used for cultural heritage preservation and reconstruction of lost architectural spaces. However, existing work has focused primarily on physically built heritage and has largely treated spatial accuracy as the primary measure of reconstruction fidelity. This assumption remains untested for artistic and cinematic heritage, where distortion, perspective, and controlled viewing conditions are often central to the original experience. To test this, we conducted a within-subjects study (N=30) to examine how VR reshapes the experience of German Expressionist film architecture. Participants experienced Jane's Bedroom from The Cabinet of Dr Caligari across three conditions: (a) film, (b) VR-static, and (c) VR-moving. Our findings show that VR significantly improves spatial perception and visualisation, but simultaneously reduces key Expressionist qualities, including claustrophobia and cinematic disorientation. We argue that spatial accuracy alone is an insufficient measure of fidelity for artistic heritage and discuss implications for future preservation workflows.
  • P1795 XR4DAvatar: Photorealistic 4D Avatar for Extended Reality Yanling Hua, Stefanie Zollmann, Tobias Langlotz
    We present a novel framework for reconstructing photorealistic 4D avatars from monocular videos for XR applications. Unlike existing 3D Gaussian Splatting (3DGS)-based avatar reconstruction methods that rely on constrained human poses and camera setups and often lack animation rigs, our approach supports unconstrained real-world monocular videos as input and produces animatable, photorealistic avatars. We integrate a diffusion-based image generation model to generate multi-view images and a textured SMPL-X-guided 3D Gaussian Splatting model to construct 3D avatars. Our results show superior performance over existing approaches in standard image quality metrics. We also show practical application of our results by demonstrating 4D human performance replay rendered as 4DGS from monocular videos in XR based on our Unity-based replay pipeline.
  • P2383 Beyond Visual Fixes: How Users Experience Anticipatory Audio Cues for VR Cybersickness Yuxue Bao, Xin Wang, Dominik Lange-Nawka, Elliott Wen, Burkhard C. Wünsche
    Visual VR cybersickness mitigation often reduces immersion, motivating non-visual alternatives. Although prior work has quantified that anticipatory audio cues lower cybersickness, how users experience them in immersive VR remains little understood. We addressed this through semi-structured interviews and reflexive thematic analysis with 24 participants in a roller-coaster simulator, comparing directional (turn direction), intensity (turn magnitude), and combined audio cues. Participants read directional cues effortlessly, but found intensity rhythms harder to interpret. Combining both felt the most complete and reassuring. Users experienced them as easing cybersickness by making motion predictable and restoring control, with intensity encoding an open design question.
  • P2607 Forgetting the Binders: A Pervasive Augmented Reality Platform for Supporting Maintenance on the Shop Floor Pedro Reisinho, Bernardo Marques, Fábio Barros, João Alves, Tiago Araújo, Duarte Almeida, Paulo Dias, Beatriz Sousa Santos
    In modern industrial environments, maintenance is hindered by the growing complexity of machinery and the dispersed nature of operational information, forcing technicians to consult data away from the equipment and sacrifice spatial context. The present study proposes a pervasive AR platform for supporting SV of maintenance information, anchoring relevant data to the components of a hydraulic press for consultation in situ. The results of a 20-participant user study demonstrated that the proposed platform achieved strong task completion effectiveness, alongside high levels of attentional allocation, information understanding, and satisfaction, with low levels of visual complexity, mental and physical effort, and frustration.
  • P2885 Exploring Dynamic-Eye-Dominance-Guided Foveated Rendering Michal Kubirita, Hans Gellersen, Stefanie Zollmann
    Foveated rendering is a well-established performance optimisation technique in virtual reality, concentrating rendering quality in the foveal region while reducing it in the periphery. Recent eye-dominance-guided methods apply stronger foveation to the non-dominant eye but incorrectly assume dominance is static. To better imitate the human visual system, we propose a dynamic gaze-angle-based approach. Our method switches the assigned dominant eye at extreme horizontal eccentricities, keeping higher rendering quality aligned with the eye that becomes dominant as gaze shifts peripherally. We evaluated this novel approach through a user study comparing dynamic dominance switching against a static dominance and a symmetric baseline.
  • P3001 Step Inside the Match: Embodied Spatial FIFA World Cup Analytics in Mixed Reality Jai Sachdeva, Prashant Rawat
    Televised football flattens the game's three-dimensional structure (space, pressure, passing networks) into fixed 2D views. We present Step Inside the Match, a mixed-reality design that anchors FIFA World Cup match analytics in the viewer's room as walkable spatial representations: passing networks as 3D node-link structures, pressure as occupiable fields, and momentum as a walkable timeline on a room-scale or tabletop pitch, built from open StatsBomb event and 360 data. We describe the interaction design, a proposed Quest 3 implementation, and a planned evaluation with fans and analysts, exploring how immersive analytics can turn passive viewing into embodied tactical inspection.
  • P3264 A Body-Proportional Constraint Architecture for Kinematically-Verified Full-Body VR Exergaming Pheelip Raipure, suryansh Bakshi, Abhijit Patil, Christian Eichhorn, David A. Plecher, Annushree Bablani, Anish Chand Turlapaty, Johanna Pirker, Himangshu Sarma
    Current full-body VR exergames lack the ability to determine whether players are actually exercising during gameplay. Players can satisfy game objectives through minimal head bobs and wrist flicks without ever engaging in genuine physical movement. We built TiltFit, a VR exergame where every in-game action is tied to real, measurable body movement through a novel Body-Lock constraint architecture. The system calibrates personalised squat and lunge thresholds from each player’s own body proportions, integrates a Dynamic Difficulty Adjustment algorithm, and includes an anti-cheat system that catches players who try to fake squats by simply nodding their heads. We evaluated TiltFit with 15 participants using independent body-worn sensors to verify that what the game detects as a squat is genuinely a squat. The answer is yes: HMD vertical displacement and knee flexion angles were strongly correlated (mean |r| = 0.79 across 23 sessions). We also discovered that when players were dropped into the game with no instructions, they could figure out the body-mapped controls in about 7 minutes with the help of a progressive hint system, and their perception of the system’s usability remained positive throughout. Participants’ heart rates climbed steadily into the moderate-intensity exercise zone during gameplay, and they reported feeling physically challenged but not frustrated. TiltFit provides the first published kinematic proof that a VR exergame produces genuine exercise.
  • P3394 Mirror-Based Avatar Transition Visualization During VR Scene Changes Dong-Yun Kim, Ho Jung Lee, Taewoo Jo, In-Kwon Lee
    Scene transition techniques for moving users across virtual environments are widely studied. Such transitions often require changing the avatar to match the destination's context. However, how to present these avatar changes remains underexplored. We introduced a waiting room between transitions to make avatar changes perceivable. Within this room, we compared how the changes were perceived with and without a mirror under cut or dissolve transition techniques. Mirror-based visualization supported aspects of user experience and embodiment, whereas transition type showed more limited effects. The results show the necessity of explicitly presenting avatar transitions in scene transition design.
  • P3566 Multimodal Augmented Reality Guidance for Object Localization and Wayfinding in Visually Challenging Environments Rafael Damouni, Muhammad Haj Ali, Julian Kreimeier, Hannah Schieber, Daniel Roth, Sarit Szpiro, Ilan Shimshoni
    Object localization is difficult in cluttered environments, particularly for people with visual impairments like tunnel vision. Augmented reality (AR) and spatial audio offer promising avenues for assistive cueing, yet little is known about how auditory, visual, and combined cues compare in guiding object search in naturalistic conditions. Using a video-see-through head-mounted display and camera-based object pose estimation, we compared these conditions in a user study where participants searched objects in an everyday environment under simulated tunnel vision. Our system provides distance-adaptive, spatially localized guidance toward the objects. Our study revealed the potential of dual-modality spatialized cues in supporting daily tasks.
  • P3628 Immersive Point Cloud Inspection: A Hierarchical Octree Streaming Add-on for the Godot Game Engine Florian Richter, Béatrice Gagnon, Bernhard Jung
    Game engines enable the fusion of scientific 3D data with interactive applications to create highly immersive and engaging user experiences. However, large-scale scenes generate a large amount of data and real-time capabilities can be limited. Sophisticated point cloud renderers therefore use optimizations, such as hierarchical octree data structures to stream data on demand and ensure high rendering performance while preserving important details. We propose a minimalist add-on for the free and open-source (FOSS) Godot game engine as an alternative to commercial engines. Our open-source add-on seamlessly integrates point cloud data while being compatible with various desktop and VR platforms.
  • P3651 Comparing 360-Degree Panoramic Videos and Peripheral Projection for Immersive VR Streaming Experiences Yuma Kokubu, Tomokazu Ishikawa
    Converting existing flat video content into immersive VR experiences offers a scalable path to leveraging vast archives of media for XR entertainment. However, whether AI-generated 360-degree conversion meaningfully enhances presence without increasing VR sickness remains unclear. We evaluate an approach using Argus, a deep learning-based panoramic generation method, to convert flat videos into 360-degree panoramas presented in VR space via Apple Vision Pro. A within-subjects experiment with 20 participants across six video stimuli and three conditions compared this approach against standard flat viewing and peripheral projection using ExtVision. IPQ results showed significantly higher presence for all six videos ($p < 0.05$ or better), while VRSQ results indicated no significant increase in VR sickness except for third-person game content. These findings provide content-type-specific design guidelines, offering practical contributions toward converting legacy video content for XR consumption.
  • P5344 VR-CineCam: Intent-Aware Automatic Spectator Camera for VR Gameplay Sungnam Kim, Mincheng Piao, Jong-Il Park
    When external spectators view VR gameplay on a 2D monitor, first-person and basic third-person follow views may not clearly convey what the player is doing and why. We present VR-CineCam, an intent-aware automatic spectator camera system for VR gameplay. It infers player intent from gameplay context and the object or character relevant to the player's current action, then selects a third-person shot that balances visibility and stable camera transitions. In a user study, VR-CineCam improved spectators’ understanding of scene events, player intention, and target relationships. These results suggest the value of intent-aware 2D spectator views for VR gameplay.
  • P5535 Finetuning Prompt-Based Image Editor for Geometry-Consistent Room Defurnishing Prakhar Kulshreshtha, Brian Pugh, Zachary Nussdorfer, Salma Jiddi
    Erasing all furniture from a room to generate its empty version is valuable for various AR and virtual staging applications. The challenge in achieving this is twofold: inpainting the furniture regions with realistic content, and doing so without distorting the room's 3D geometry. Prompt-driven image editors, such as FLUX.2, handle realism impressively, but their structural faithfulness is inferior to explicit layout conditioned inpainting. Hence, to improve geometric consistency, we propose a training recipe for finetuning a LoRA on 2,500 furnished-empty room pairs. Flux.2 with our LoRA adaptor significantly improves geometric consistency while retaining the high texture fidelity.
  • P5814 AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov, Jeffrin Sam, Dmitrii Iarchuk, Yara Mahmoud, Daniia Zinniatullina, Mikhail Konenkov, Dzmitry Tsetserukou
    Human egocentric video is a scalable supervision source for humanoid policy learning, but current pipelines struggle with hand-object occlusion, oversimplified motion, or specialized capture hardware. We introduce \textit{AgenticFocus}, a Mixed Reality synthesis pipeline that converts ordinary first-person-view human videos into robot-trainable demonstrations by restoring occluded object geometry, reconstructing full-hand motion, and retargeting it to a humanoid embodiment through camera-relative alignment and layered compositing. The resulting dataset pairs focused visual observations with synchronized robot actions and states. AgenticFocus achieves lower trajectory error and smoother wrist motion than cross-embodiment baselines, with SPARC scores of −5.18 versus −5.56 and −6.05.
  • P6213 StereoBrake: Towards Mitigating VR Sickness by Suppressing Stereopsis with a 2D Projection Yuto Ogata, Kodai Hatakeyama, Eunhee Chang, Guanghan Zhao, Yuhui Wang, Kazuyuki Fujita, Yoshifumi Kitamura
    Virtual reality (VR) sickness remains a critical barrier to comfortable user experiences. To address this, we propose StereoBrake, a dynamic stereopsis control system that mitigates sickness during virtual locomotion without obstructing the field of view. By modulating binocular disparity based on user movement it restricts stereopsis during locomotion and restores full stereoscopic vision when stationary. A within-subjects user study demonstrated that StereoBrake successfully maintained presence without significant degradation compared to control and continuous restriction conditions. StereoBrake demonstrates strong potential for future investigation and methodological evolution despite not significantly mitigating sickness on its own.
  • P6743 OrthoPaint: Photorealistic 3D Mesh Defurnishing via Orthographic Generative Inpainting Sunil Kumar Narayanan, Prakhar Kulshreshtha, Brian Pugh, Salma Jiddi
    Photorealistic defurnished meshes are vital for mixed-reality virtual staging. However, consumer-grade room-scans retain furniture remnants, projection seams, and texture holes. Naive per-frame inpainting restores each captured view independently but often produces 3D-inconsistent textures. We propose OrthoPaint, a generative mesh-texturing pipeline that inpaints on surfaces rather than frames. We render each surface with an orthographic camera. A Wall-Graph then groups narrow walls with nearby broader walls into Wall Island composites. An MLLM captions each composite, and an inpainter restores its constituent surfaces with photorealistic textures. OrthoPaint outperforms naive per-frame and video-diffusion baselines, yielding fully textured meshes with higher perceptual quality scores.
  • P6888 Feeling Together Without Talking: Behavioral Signatures of Co-presence in Silent Cooperative VR Alex Fuentes-Raventos, Ana Zappa, Xim Cerda-Company, Michael Wiesing, David Cucurella, Shunkai Wang, Marc Ballestero-Arnau, Mel Slater, Antoni Rodrigues-Fornells
    How do people come to feel together in a shared virtual space when speech and gesture are unavailable? We examine this question in a silent three-player VR card-ordering task where coordination must emerge through action and timing. Across 12 teams, stronger team co-presence ratings were linked to higher team task effectiveness: accurate card play and repeated error-free completion. This pattern remained visible in difficulty-validation checks, suggesting it was not simply due to easier card states. Higher co-presence teams also showed more stable early timing, pointing to shared task fluency as a behavioural signature of feeling together in silent VR.
  • P6893 ROGS-VR: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration Timofei Kozlov, Dmitrii Maliukov, Andrey Marchenko, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
    We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods assuming clean depth or external poses, our system combines ORB-SLAM3-based pose estimation with online Gaussian reconstruction for noisy real-world data. A VR pipeline enables immersive exploration of incremental reconstructions; a semantic module transcribes voice commands, generates scene descriptions, and records points of interest. Against state-of-the-art online Gaussian splatting methods, we improve image quality on our and TUM-RGBD datasets, with comparable or superior frame rates via quality–speed configurations. We achieve 88% VLM object-recognition rate.
  • P6909 Large-Aperture Flat Diffractive Optics for Ghost-Suppressed Full-Color Holographic AR Displays HYUN EUI KIM
    Large-area spatial holographic AR requires optical apertures whose space-bandwidth product scales with image size and field of view, yet centimeter-scale refractive optics are bulky. We present a closed-form carrier-allocation rule for a 100 mm × 100 mm full-color multilevel diffractive optic that co-locates the three RGB signal orders at a common focal-plane coordinate while displacing cross-channel parasitic orders outside the signal aperture, without iterative optimization. For f = 450 mm and a 10° carrier, the RGB orders co-locate at u = 79.3 mm and the nearest parasitic order is displaced by 10.8 mm. The fabricated optic forms ghost-suppressed, depth-resolved real images over a 0–600 mm range in a spatial holographic AR testbed.
  • P7126 Enhanced Legibility of Out-of-Focus Virtual Latin and East Asian Characters with a Commercial See-through Augmented Reality Display Junhwan Kim, Carlos Montalto, J. Edward Swan, Mohammed Safayet Arefin
    Out-of-focus effects degrade text legibility in optical see-through (OST) augmented reality (AR). Previous research using SharpView fonts has shown that precorrection techniques can improve out-of-focus text legibility. However, the text was limited to Latin characters, and the evaluation environment was limited to a laboratory OST AR device. This research extends the SharpView approach in two ways: (1) it extends the evaluation to East Asian characters (Korean, Japanese, and Chinese) and (2) it conducts synthetic simulation and camera-based optical evaluations using a Microsoft HoloLens 2. With both evaluations, the SharpView precorrected font shows quantitatively improved sharpness and text legibility. These results suggest that the SharpView font method can generalize to multilingual environments and commercial AR devices.
  • P7189 Geospatial Augmented Reality as a New Medium for Outdoor Spatial Storytelling Si Jung Kim, Byeong-hee Roh, Adrienne Landerville
    This paper presents a geospatial augmented reality system for large-scale outdoor theatrical performance, instantiated as (anonymized) The system superimposes GPS-anchored 3D virtual content over an (anonymized), accessible via mobile devices. We address three core geospatial computing challenges: (1) mitigating Global Navigation Satellite System multipath error in an urban canyon environment through a hybrid localization pipeline combining raw GPS, and a GeoJSON point-of-interest map; (2) establishing a viewer-relative local coordinate frame that maintains sub-3-degree angular registration across 10 distributed viewpoints spanning the full promenade perimeter; and (3) synchronizing virtual spatial events to physical (anonymized) choreography. Findings establish empirically grounded design constraints for GPS-anchored spatial interaction in dense urban outdoor venues.
  • P7640 The Whispering Wood: Gamified Age Gate for immersive experiences with Spatialized-Audio and Anti-Spoofing game mechanics Olufunke Faderera Ajenifuja, Riccardo Bovo, George Loukas
    Mixed reality platforms require a method to determine whether a user is a minor or an adult without disrupting the experience. However, existing age checks are based on face images, which require the user to remove the headset and break immersion while also creating privacy concerns. We explore audiometry as an alternative approach that leverages headset audio output, allowing users to remain immersed. We use spatialized audio and XR holograms to turn the hearing test into a head-tracking game, where the same head movements that report whether a tone was heard also act as a spoofing check. We test this method on a simulated population and demonstrate a working prototype.
  • P7815 Lighting-Aware Presentation of Real Objects on a 3D Display Using Relightable 3D Gaussian Splatting Aki Yato, Taishi Iriyama, Takashi Komuro
    We propose a method for reproducing a real object's material appearance by relighting its reconstructed 3D Gaussian representation according to the observer's surrounding illumination and presenting the relit result on a 3D display. Using relightable 3D Gaussian Splatting, which estimates per-Gaussian normals and material parameters to separate object appearance from lighting, we obtain an illumination-independent representation and relight it with an HDR environment map captured by an omnidirectional camera. We qualitatively demonstrate that, under bright and dim lighting conditions, the diffuse and specular reflections of the rendered object change consistently with the color and intensity of the surrounding light source.
  • P7832 Inner Mirror: Exploring Projection-Based Biofeedback for Multisensory Breathing Meditation in Home Environments Lisa Timm, Michael A. Herzog, Danny Schott
    Stress regulation is increasingly important in everyday life. We present a projection-based multisensory meditation experience for home environments that combines breathing and pulse biofeedback, real-time visualization, sound, and mudra-inspired microgestures. A qualitative user study with five participants explored user experience, body awareness, and biofeedback perception following controlled stress induction using questionnaires and semi-structured interviews. Participants perceived the dynamic projection either as a breathing pacing cue or as a representation of their own breathing when synchronization appeared credible. These findings inform the design of multisensory biofeedback systems that support relaxation, self-awareness, and well-being in everyday domestic contexts through credible embodied interaction.
  • P8439 Human Factors in Implementing Virtual Showrooms: Insights from an Industry Focus Group and a Survey Melisa Tasliarmut, Katharina Precht, Ronja Schneider, Chris Calum Pankerl, Philipp Klimant, Angelika C. Bullinger
    Virtual showrooms integrating generative artificial intelligence and spatial computing offer new opportunities to enhance user experiences in industrial contexts. Despite growing interest, adoption remains limited due to insufficient understanding of industry requirements and human factors. This study explores requirements and barriers using a mixed-methods design within the user-centered design framework. A focus group with industry representatives collected requirements and barriers before and after experiencing a VR showroom, while survey participants without prior showroom experience ranked literature-based factors. Triangulated results provide an overview of critical requirements and barriers for practical implementation, particularly regarding maintainability, usability, cost-benefit analyses, and simulator sickness.
  • P8762 FocusStack2XR: Provenance-Preserving Depth-Aware Visualization of Focal-Stacked Specimen Images Bingxi Li
    We present FocusStack2XR, a prototype pipeline and early browser-based viewer for depth-aware inspection of focal-stack specimen images. Rather than claiming metric 3D reconstruction, the system preserves focal-plane provenance by linking displayed regions to source image slices. From a captured stack, it estimates sharpness maps, focus-index maps, confidence cues, foreground masks, and relative focus-depth proxies. A 24-slice insect-specimen prototype demonstrates how these representations can support all-in-focus viewing, source-slice inspection, and uncertainty-aware scientific browsing of extended-depth-of-field image data without hallucinating occluded or uncertain structure.
  • P8961 InnerFlow: A VR Tai Chi Guidance System with Movement-Responsive Self-Avatar Visualization Sirui Xiao, Zekun Qin, Zhaohan Wang, Yezi He, Zihan Gao, Xin Lyu
    Tai Chi is a mind–body practice cultivating posture, rhythm, and internal flow, yet existing AR/VR Tai Chi training systems mainly emphasize posture accuracy. We present InnerFlow, a VR guidance system that visualizes movement-responsive inner flows within a semi-transparent self-avatar. Using motion capture and real-time particle visualization, InnerFlow transforms users’ movement quality into coherent or disturbed visual flows. A pilot study with 12 participants showed good usability, strong embodiment, and moderate-to-strong presence. InnerFlow contributes culturally grounded self-avatar feedback that shifts VR Tai Chi learning from posture correction toward embodied perception.

Posters 1BCassiopeaXR Training, Education & Responsible Use Wednesday13:45–16:00

  • P1059 Semantic Dissonance Detection: A Contrastive Kinematic Transformer for Predictive Support for Errorless Learning in VR Vedanta Hazra, Kaushal Kumar Bhagat
    High consequence procedural training in Virtual Reality currently relies on reactive Intelligent Tutoring Systems that correct novices only after a physical error is committed. Allowing erroneous physical execution fosters negative training transfer and incorrect motor schema consolidation. We introduce Semantic Dissonance Detection, a cross-modal machine learning framework designed to predict and intercept cognitive action slips before physical contact occurs. By mapping a user's verbalized intent into a semantic latent space using a frozen Vision-Language Model, and simultaneously projecting their real-time 13-element OpenXR kinematic state vector through a custom Contrastive Transformer, the system continuously estimates intent-action divergence. To solve spatial overfitting, we apply Relative Position Normalization across 45-frame rolling temporal windows. The kinematic encoder is trained using Supervised Contrastive Learning to mitigate false negative penalties among kinematically similar linear motions. In a controlled study (N=50, 1500 trajectories) evaluating a four-action spatial protocol, our framework achieved a 93.7% Leave-One-Subject-Out (LOSO) cross-validation accuracy. During applied pedagogical testing, the framework isolated a stable congruence baseline that successfully identified action slips in our experimental setting within the first 0.25 seconds of movement. This 750ms predictive safety buffer allowed users to correct their trajectories mid-reach, associated with a 95.3% relative reduction in completed physical errors compared to reactive baselines. This work offers a focused yet rigorous demonstration of supervised cross-modal kinematic-semantic alignment within a constrained spatial training environment for early predictive modeling in immersive environments.
  • P1860 TurniBlocks: A Head-Mounted Mixed Reality System for Turn-taking Training in Children with ASD Zi di Yu
    TurniBlocks is a head-mounted mixed reality system designed to support turn-taking training for children with autism spectrum disorder (ASD). Inspired by the role collaboration and turn-taking mechanisms of LEGO®-Based Therapy (LBT), it transforms block-based cooperative roles into virtual building tasks embedded in a familiar real therapeutic space, forming a multi-party collaboration model. To evaluate its usability, we conducted a user study. Results suggest that TurniBlocks can partially enhance children’s engagement, collaborative turn-taking behaviors, and social participation, highlighting the potential of head-mounted MR for structured autism social training.
  • P2438 Adaptive Coaching with High-Fidelity Physics for Virtual Reality Badminton Training Suryansh Bakshi, Chaitanya Anand, Pheelip Raipure, Vaishnav Kameswaran, Himangshu Sarma, Annushree Bablani, Mrinmoy Ghorai
    Traditional sports training often struggles to deliver real-time, personalized, and biomechanically accurate feedback during rapid physical execution. This paper presents XR Badminton, a system architecture engineered for adaptive coaching with high-fidelity physics for virtual reality badminton training running on standalone head-mounted displays. The framework introduces three core contributions: (1) a real-time tracking pipeline that utilizes a Holt-Linear filtered state machine to segment complex overhead strokes into a seven-phase kinematic sequence, scoring performance across eleven expert-derived biomechanical rules via distinct mathematical mapping profiles; (2) a high-fidelity training environment featuring a procedural, aerodynamic shuttlecock simulation with nose-heavy physical balance and quadratic drag trajectory prediction; and (3) an adaptive coaching layer that isolates a player's most critical mechanical flaw via priority-weighted deficit selection and dynamically routes tailored corrective feedback through seamless modality substitution across visual HUDs, spatial audio, and haptic cues. We outline the core system components, demonstrate a non-blocking telemetry logging infrastructure used for high-frequency player tracking, and present a verification protocol validating the framework's tracking calculations and trajectory physics
  • P2798 AR Bouldering Replay with Contact-Aware Human-Scene Alignment Yanling Hua, Stefanie Zollmann, Tobias Langlotz
    We present a 4D human-scene reconstruction method for precise human-scene alignment in bouldering scenarios. 4D bouldering replay can convey complex climbing movements from different perspectives and support asynchronous collaborative training. However, it is nontrivial to accurately reconstruct and align human motion with climbing-wall geometry from monocular video. To address this problem, we introduce a contact-aware fine-tuning method that builds upon state-of-the-art human-scene reconstruction methods. By incorporating human-scene contact and penetration constraints, our method improves alignment between reconstructed climbers and climbing walls. We further develop an AR bouldering system for optimal climbing-wall anchoring and motion replay.
  • P3108 Expertise-based Adaptive AR Instructions with Motion Data Tania Kaimel, Denis Kalkofen, Lucchas Ribeiro Skreinig, Ana Stanescu
    Augmented Reality has been used to support users in a range of scenarios, often showing advantages over traditional approaches such as paper-based instructions. However, existing systems rarely consider the user's expertise when presenting instructional content. We investigate how expertise can be predicted from data available in a consumer-grade head-mounted display and use these predictions to implement an adaptive augmented reality assembly tutorial that tailors instructional support to the user. Expertise classification is performed on a public dataset using different machine learning algorithms. Results show motion-based expertise classification is feasible, demonstrating the potential for adaptive instructional systems.
  • P3493 Evaluating LLM-Driven Virtual Parent Dialogues in VR for Negative Emotion Elicitation in Pre-Service Early Childhood Teachers Chunyue Yan, Pengxiang Wang, Fangbing Qu, Xiaohui Tan
    Communication with emotionally expressive parents may induce negative emotions in early childhood teachers, challenging emotion-management training. Although VR can simulate such interactions, prior work relies on passive observation or scripted dialogues. We evaluated an LLM-driven virtual parent dialogue approach in VR for eliciting negative emotions in pre-service early childhood teachers. Twenty-eight participants completed dialogues with anxious and angry virtual parents while subjective and physiological data were collected. Both dialogues elicited negative subjective changes and physiological responses; participants recognized affective profiles, and angry parents produced stronger negative responses. Findings provide an experimental basis for teacher training and VR emotion-elicitation research.
  • P3630 Spatializing Algorithmic Bias in XR: A Bio-Inspired Mechanism for User Awareness and Intervention Yiding Wang
    As social media has become a central part of daily life, algorithmically driven recommender systems can now shape many human decisions, including what they view, what they focus on, and what they do. These recommendations introduce algorithmic bias based on the assumptions and choices made during the system's development and data use. The benefits of using recommender systems to enhance the social media experience are substantial, but so too is the potential to affect users' cognitive processes and decision-making through embedded algorithmic biases. Most current studies addressing algorithmic bias focus on model design, auditing, and regulation, with little to no research on how users interact with and dispute these systems. To address this gap, this study proposes to develop an interdisciplinary Human-Computer Interaction (HCI) method using Extended Reality (XR) to create a fully immersive virtual environment that renders abstract algorithmic biases as tangible, manipulable spatial elements. Drawing on self-organization in ant colonies as my model, I am developing the A-N-T-S (Apprehension–Negation–Transition–specialization) model of user behavior to understand what causes users to detect and oppose algorithmic bias. The ANTS model is guiding the development of an XR prototype that brings together geographic information systems and multimodal interaction to represent weighted algorithmic decisions as dynamic "crack zones" and bias maps. In the XR environment, users can navigate a virtually distorted city, experience the spatial manifestation of recommender logic, and test different paths to escape or intervene in the recommendation process (see Fig. 1). My planned user study will investigate whether using this XR environment based on the ANTS model can improve users' ability to perceive, resist, and interpret algorithmic bias, providing a design framework for more transparent, participatory, and user-focused algorithms.
  • P4166 A Virtual Reality Cognitive Training Platform for Brain Cancer Patients with Mild Cognitive Impairment: Clinical Study-in-Progress Anthony Faiola, Lalanthica Yogendran, Rhonna Shatz
    Brain cancer affects 76,000 Americans annually. Approximately 80% of brain cancer survivors (BCSs) experience mild cognitive impairment (MCI), affecting memory and executive function (MEF). Although studies show that paper/computerized cognitive training mitigates MCI, meta-analyses confirm that virtual reality can improve MEF in BCSs. We developed a Virtual Reality Cognitive Rehabilitation Training (VR-CRT) platform with immersive cognitive exercises. As a follow-up to our preliminary study, our current four-week intervention/control clinical study (N = 36) will evaluate the feasibility and effect of VR-CRT. Cognitive performance is measured at baseline/pre-post using HVLT/COWA/TMT tests. Anticipated outcomes are modest-to-significant MEF improvements in BCSs with MCI.
  • P4315 A Research Agenda for Responsible Neuro-XR Marios Constantinides, Kyriaki Kostoglou, Horia A. Maior, Fotis Liarokapis, Max L Wilson, Christos Fidas
    As neurotechnology is converging with Extended Reality (XR), advances in neurophysiological monitoring are expanding the ability of XR systems to gain insight into users' attentional, cognitive, and affective states. Existing Responsible XR frameworks address risks associated with behavioural and physiological data, yet Neuro-XR (i.e., XR systems augmented with neural sensing) introduces a different challenge; that is, access to information about users' internal mental processes, raising concerns around mental privacy, cognitive liberty, neuro-surveillance, and algorithmic influence. We argue that this convergence necessitates a new research agenda centred on Responsible Neuro-XR. Drawing on literature in XR ethics, neuroethics, and neurotechnology governance, we identified four core ethical challenges and propose four design considerations for the ISMAR community. We also highlighted research opportunities in neuroadaptive learning, collaboration, wellbeing, and inclusive design where the community can (re)shape responsible development of future brain-aware immersive systems.
  • P4416 Towards Conversational Industrial Mobile Augmented Reality Andreas Andreou, Marios Constantinides, Fotis Liarokapis
    Industrial assembly and maintenance tasks require access to complex technical knowledge, yet consulting manuals or experts is often impractical during hands-busy and safety-critical work. These challenges particularly affect non-knowledge workers (e.g., technicians and operators) who must resolve issues in situ under time pressure. While Augmented Reality (AR) has been used to support industrial training and guidance, many existing systems rely on predefined instructions that cannot adapt to workers’ ad hoc questions. We present a prototype AR system that integrates an LLM-powered conversational assistant to provide hands-free assistance, and illustrate it through a pipe assembly and maintenance use case.
  • P5139 Striving for Stability: Evaluating the Feasibility of Dynamic Balance Skills Across the Reality–Virtuality Continuum Alexander Schwadtke, Stefan Pastel, Matthias Kunz, Patrick Müller, Rüdiger Braun-Dullaeus, Christian Hansen, Kerstin Witte, Danny Schott
    Dynamic balance assessment is important in sports and rehabilitation. Mixed Reality (MR) offers new opportunities for portable diagnostics. This study investigated whether the standardized Y-Balance Test (YBT) can be reliably and validly administered in MR environments. Twenty-four participants performed the YBT in a real environment (RE), Augmented Reality (AR), Augmented Virtuality (AV), and Virtual Environment (VE) on two test days. Results confirmed validity across all MR conditions (p > 0.05), with good test-retest reliability in AR and VE and moderate-to-good reliability in AV. Despite higher perceived demands, MR-based YBT assessment showed outcomes comparable to RE, supporting its feasibility for diagnostics.
  • P5228 Co-Designing Conversational AR Avatars for Low-Pressure Classroom Support Manjeet Singh, Sajjanhar Atul, Yi Wang, Shaun Bangay, Rasnaam Kaur, Dr. Harjit Kaur, Harjinder Kaur
    Conversational and augmented reality technologies offer new opportunities for classroom learning support, yet their design often overlooks the social and emotional conditions that shape students’ willingness to ask for help. This paper presents a co-design study of a conversational AR avatar intended to support low-pressure help-seeking and emotional regulation in classrooms. Through workshops and surveys with 47 students and 13 teachers, we found that the main barriers to help-seeking were social rather than purely cognitive, especially fear of judgement, embarrassment, and unwanted visibility. Participants preferred the avatar to act as a low-authority peer or guide, provide scaffolded support rather than direct answers, and offer multimodal interaction options that preserve privacy. Based on these findings, we derive design requirements for conversational AR avatars as socially aware, classroom-integrated support tools.
  • P5880 Designing Sustainable XR Content Pipelines for Industrial Training Services: A Case Application in Airport Vehicle Checkpoint Security Screening Jeonghoon LEE
    eXtended Reality (XR) is a key technology for industrial digital twins, yet adoption is constrained by inefficient content authoring and management. We propose an XR content authoring and management strategy based on an XR Contents Fabric architecture, which extends the data fabric concept to XR by separating functional and semantic content elements to enable reusable, maintainable pipelines. We demonstrate it through an XR-based airport vehicle checkpoint security screening training service. A pilot with airport security instructors confirms that the framework supports instructor-centered authoring, continuous content operation, and service sustainability, providing a foundation for future AI-driven intelligent XR content platforms.
  • P5907 Mixed Reality Endovascular Thrombectomy Training via Holographic Patient CT-A Visualization Ehtiram Shukurov, Alex Berg
    Endovascular procedures such as thrombectomy require hands-on training constrained by limited access to expensive simulation hardware. We present a mixed reality training simulation for endovascular thrombectomy on Meta Quest 3. Real patient CT-A scan data is rendered holographically using a custom HLSL raymarching shader, enabling trainees to visualize patient-specific vascular anatomy in their physical environment. The system supports interactive guidewire and catheter navigation through reconstructed vascular geometry, with haptic feedback during procedural steps. A fluoroscopy simulation mode and radiation exposure monitoring replicate clinical workflow conditions, demonstrating feasibility of patient-specific holographic anatomy visualization on consumer MR hardware for accessible endovascular training.
  • P6406 Claws of Change: An Embodied Educational VR Experience for Visualizing the Effects of Ocean Climate Change Through Time Travel Danielle Cole, Aude Nguyen, David Bruce, Anthony Scavarelli
    Claws of Change is an immersive virtual reality educational experience that places learners in the body of an American lobster exploring the Atlantic Ocean. Participants travel through time to witness the effects of climate change using a novel claw-based swimming locomotion system that strengthens embodiment. The experience aims to foster empathy for marine life through non-human perspective-taking and environmental storytelling. A pilot study in a university sustainability classroom found the experience engaging, easy to learn, moderately immersive, and effective for climate education. Future work will improve accessibility, locomotion, environmental interactions, and educational content.
  • P6452 A Retention-Centered Review of XR Training Systems in Agriculture Haris Psallidopoulos, Fotis Liarokapis
    Extended Reality (XR) is an umbrella term that includes various immersive technologies, including Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) \cite{whatisxr,mrdef}. Recently, XR has gained tremendous traction and application in the educational sector, particularly for training professionals aiming to refine their skills and deepen their expertise in specific disciplines. This transformative tool has proven to be exceptionally beneficial across numerous industries, with agriculture standing out as a crucial area of application \cite{agrivr}. By offering immersive and interactive simulations that replicate diverse farming environments and the operational workflows inherent to them, XR significantly enhances the educational experience.
  • P6981 Instruct360: A 360° Video Framework for Egocentric Object Capture and Tracking for AR Instructions Max Rütz, Tobias Langlotz, Stefanie Zollmann
    We present an approach for extracting 3D object representations and their spatial positions from egocentrically captured panoramic video recordings. The system processes panoramic videos captured by a head-worn camera, with a focus on instructional scenarios in which object configurations and interactions are central to learning. The extracted 3D objects are tracked over time and replayed in augmented reality (AR). By decoupling instructional information from the original camera viewpoint, learners can observe and follow procedures from their own perspectives. The primary motivation is to reduce constraints during capture, permitting natural head movement while maintaining spatial fidelity required for AR-based instructions.
  • P7693 Cooperative Feedforward Avatars in a VR Rowing Exergame: Effects on Exercise Motivation and Performance Lujia Yang, Tony Bosch, Junxiao Liao, Lijie Wang, Yaomin Zhu, Zixuan Wang, Burkhard C. Wünsche
    We present a VR rowing exergame that uses static cooperative feedforward avatars to support pacing, motivation, and performance. After data cleaning, 26 participants were included in a repeated-measures study across baseline, +3\%, and +5\% conditions. Both avatar conditions significantly increased stroke rate and power output compared with baseline. Interest/Enjoyment and Effort/Importance differed significantly across conditions, while Pressure/Tension did not. Qualitative feedback suggested that +3\% felt smoother and more motivating, whereas +5\% was sometimes perceived as more effortful or tiring. These findings suggest that cooperative avatar pacing can support exercise without relying on competitive pressure.
  • P7713 A Predict–Observe–Explain Real-Time Parametric Virtual Reality Environment for Teaching Column Buckling in Civil Engineering Education Alyssa Lubrano, Ali Darejeh, Sara Mashayekh, Samad M.E. Sepasgozar
    This paper investigates the effectiveness of a real-time parametric virtual reality (VR) environment for teaching the buckling behaviour of concrete columns in civil engineering education. The system enables users to interactively modify structural parameters, including column height, cross-section, applied load, and boundary conditions, while observing real-time deformation and failure behaviour. A between-subject experiment compared VR-based learning with traditional video instruction, evaluating conceptual understanding, memory retention, cognitive load, and learner confidence. Results showed that the VR group experienced lower perceived cognitive load and greater increases in confidence and engagement. Although no statistically significant differences were found in learning outcomes between the two methods, the confidence measures and qualitative findings indicated that the interactive features of the VR environment provided meaningful pedagogical benefits. The work contributes a Predict–Observe–Explain-based VR approach that integrates symbolic and visual representations to support the teaching of structural behaviour that is difficult to demonstrate through physical experimentation.
  • P8007 Resilience Protocol: A VR Stress Inoculation Training System with MetaHuman-Driven Adaptive AI Ehtiram Shukurov
    Stress inoculation training (SIT) builds psychological resilience by progressively exposing individuals to controlled stressors. We present Resilience Protocol, a first-person virtual reality application designed as an immersive SIT experience built in Unreal Engine 5.6. A MetaHuman-driven AI antagonist pursues the participant through three escalating stages of increasing challenge, incorporating physics-based interactive puzzles and an adaptive audio system that responds dynamically to in-game events. The system translates established cognitive-behavioral SIT methodology into an interactive VR format, offering an engaging and scalable alternative to traditional stress inoculation approaches. We describe the system architecture, design rationale, and potential applications in resilience training contexts.
  • P8896 Spatially Anchored Knowledge Graphs: Connecting Everyday Contexts to Study Materials in Mixed Reality Qingzhu Zhang
    Professional study often requires learners to connect abstract concepts with concrete experience, but conventional review rarely preserves the everyday cues that could support those connections. We present spatially anchored knowledge graphs, a mixed-reality authoring workflow that turns captured objects into source-grounded memory anchors. The prototype captures object snapshots in passthrough MR, generates role-aware association paths through exploratory AI or textbook-grounded retrieval, and transfers selected paths into a VR knowledge graph for spatial organization and source inspection. It positions MR as situated capture and VR as embodied graph authoring, providing a concrete basis for studying associative recall and review workload.
  • P9448 Personalised 3D Mnemonics for Vocabulary Learning in Context Karolina Trajkovska, Matjaž Kljun, Maheshya Weerasinghe, Klen Čopič Pucihar
    Augmented reality (AR) can support vocabulary learning by placing mnemonic cues in the physical contexts where new words are encountered. Prior work has augmented real objects with bilingual labels, audio, keywords, and 3D visualisations of those keywords interacting with the objects, but these cues are typically predefined by system designers. This limits scalability and reduces learners' ability to construct personal mnemonic links. By grounding mnemonics in objects that learners already recognise, use, and associate with prior experiences, the system can connect new vocabulary to existing mental models rather than requiring learners to adopt designer-defined associations. We propose a generative AR approach in which learners create personalised 3D mnemonics by transforming familiar objects in their surroundings into situated mnemonic anchors. We position this concept within a timeline of related work and argue that recent advances in generative 3D asset creation enable a new design point: personalised 3D mnemonics that do not merely place content in the physical learning environment, but modify familiar real-world objects so that they become part of the mnemonic itself.

Thursday – October 8, 2026

Posters 2ACassiopeaHuman Perception, Cognition & XR Methods Thursday11:15–13:45

  • P1076 Toward Interoperable and Reproducible XR: Key Insights from the IEEE VR 2026 Standardization Panel Takeshi Kurata, Jen-Shuo Liu
    The IEEE VR 2026 Standardization Panel, "Standardization in XR and VR: Challenges and Priorities Beyond Terminology,'' brought together representatives from academia, industry, and standards organizations to discuss the role of standardization in the future of immersive technologies. Rather than focusing solely on terminology, the panel addressed how standardization can support interoperability, reproducibility, research practices, and emerging technologies while maintaining close communication between the research community and standards organizations.
  • P2466 Exploring a Virtual Pet as an Adaptive Agent in Personalized VR Rehabilitation: A Country Fair Serious Game for Stroke Survivors David Palricas, Bernardo Marques, Inês Figueiredo, Sérgio Oliveira, Bianca Guerreiro, Beatriz Sousa Santos
    Traditional stroke rehabilitation methods fail to sustain long-term motivation due to their repetitive and non-adaptive nature. This work presents a VR-based country fair serious game designed to support stroke survivors during upper-limb rehabilitation through personalized and engaging interactions. The game integrates a virtual pet as the central engagement mechanism. Survivors complete frisbee-throwing tasks toward predefined targets in the virtual environment, after which the pet retrieves the object and returns it, creating a continuous interaction loop. Visual and auditory feedback, including celebratory cues such as barking, reinforces progress and maintains attention. A personalization module continuously collects performance data and dynamically adjusts target placement according to the survivor’s ability, ensuring an appropriate challenge level while reducing monotony. The game was evaluated with 32 participants, who reported positive feedback, describing it as fun, easy to use, and appropriately challenging. Results suggest that adaptive personalization enhances motivation and provides useful behavioral data for new features.
  • P2789 Beyond Non-Human Embodiment: An Affordance Analysis of Non-Anthropocentric Experience in VR Games Yinshuai Zeng, Chenyu Zhang, Ana-Despina Tudor
    Virtual Reality is increasingly used to support environmental storytelling and non-human perspective-taking (Zong, 2026). However, embodying non-human characters does not necessarily result in a non-anthropocentric experience.This qualitative study aims to uncover how affordances in two VR games support non-anthropocentrism through a comparative analysis of Moss and Stray VR. Gameplay recordings and post-experience questionnaires were inductively coded in NVivo. Four affordance patterns emerged: Species-Scaled Movement, Environmental Navigation, Species-Consistent Interaction and Relational Agency, revealing that non-human embodiment is often supported, while ecological relations remain structured through anthropocentric control logic. Future work will include developing design guidelines for non-anthropocentric VR experiences.
  • P2811 Human-centric Evaluation of Enhanced Legibility of Out-of-Focus Virtual Fonts During the Incomplete Eye Accommodation Junhwan Kim, Carlos Montalto, J. Edward Swan, Mohammed Safayet Arefin
    During focal distance switching in Optical see-through (OST) augmented reality (AR) systems, users need to continuously change the shape of their eyes (accommodation) to see information sharply at different focal distances. In addition, when the eye adapts to a new focal distance (which takes around 360 to 480 milliseconds), the view of the information becomes blurred, resulting in transient out-of-focus blur. In this research, we investigate a perception-driven approach based on image pre-correction, in which displayed text is optimized to appear sharper during out-of-focus periods (SharpView Font). We conducted a controlled user study with a commercial OST AR headset, explicitly accounting for accommodation latency and a range of perceptual design parameters. Consistent with expectations, our results revealed that the performance of SharpView text under defocused viewing conditions is determined by the specific combination of perceptual design parameters and the content’s geometric characteristics. Together, these insights yield user-focused guidelines for optimizing pre-correction strategies for out-of-focus AR text.
  • P3068 Same Environment, Opposite Errors: Asymmetric Effects of Rendering Type on Spatial Internalization and Externalization in VR Jungyoon Kim, inhi kim
    VR spatial perception research has traditionally concentrated on individual users' perception. However, real-world VR applications increasingly demand interactions involving multiple users, requiring bidirectional spatial information processing. Nevertheless, studies examining both internalization and externalization simultaneously remain limited. To address this, we conducted an experiment to investigate how these two directions of spatial information processing differentiate, while also examining whether rendering type (Point Cloud vs. Mesh) differentially affects each direction. Results revealed a clear asymmetry between conditions suggesting that rendering type affects spatial encoding and decoding. This highlights the need to consider directionality of spatial processing in VR system design.
  • P3145 Does Full-Body Motion Reveal More? Comparing Human Inference of Personal Attributes from Upper-Body and Full-Body Motion Jiayi Liu, Arran Zeyu Wang, Mark Roman Miller, Danielle Albers Szafir, William Christopher Payne
    Motion tracking in extended reality (XR) enables immersive experiences but also raises privacy concerns. Prior work has shown that motion data enables model-based personal attribute inference, with richer body tracking improving accuracy. However, less is known about whether tracking more body regions can also improve human perceptual inference. This study compares upper-body and full-body motion tracking in a between-subjects experiment (N = 120), where participants inferred gender, age, height, weight, and ethnicity from skeletal motion clips. While participants reliably inferred certain attributes in both conditions, there was no statistically significant difference in accuracy between conditions.
  • P4773 Or Is It Just Coincidence: Towards Multimodal Interpersonal Synchrony in Real-World and Virtual Interaction Tasks Zane Zarathushtra Karai, Frederik Kalle, Nastaran Saffaryazdi, Alireza Farrokhi Nia, Kunal Gupta, Benjamin Tag, Mark Billinghurst
    When people interact with each other, they unconsciously produce movements, emotional reactions, and physiological signatures that synchronize with their partner. In this poster, we present a methodology for collecting eye-gaze and brain activity to examine synchronization between pairs of people during competitive gameplay in real-world and Virtual Reality (VR) environments. We investigate how measures differ when players correctly or incorrectly predict their opponent's move, testing if competitive predictions emerge from interpersonal synchrony. By comparing environmental dynamics, we explore whether VR preserves ecologically-valid social functions. This work gives insight into how VR can be used to explore interpersonal synchronization.
  • P5343 The Extended Six-Dimensional Framework for Intelligent XR Learning Environments (6D-XR Framework) Fotis Liarokapis, Ourania Miliou, Yiorgos Chrysanthou
    The integration of Artificial Intelligence (AI) and Extended Reality (XR) is transforming immersive learning environments into adaptive, intelligent educational ecosystems. Existing XR learning frameworks provide valuable foundations but do not fully address recent advances in AI-driven personalization, intelligent tutoring, and responsible innovation. This paper proposes an Extended Six-Dimensional Framework that expands the four-dimensional framework by introducing two additional dimensions: Intelligence and Adaptation, and Ethics, Trust, and Sustainability. The framework integrates learner characteristics, pedagogy, immersive representation, context, intelligent adaptation, and ethical considerations into a unified structure for designing, evaluating, and guiding next-generation XR learning environments.
  • P5374 Do Elliptical Avatar Distortions Shape 3D Reaching Movements? Iris Willaert, Valentin Vallageas, David R Labbe
    Embodied avatars can influence movement when their movements are distorted. We tested whether an ovalized self-avatar hand trajectory alters unconstrained 3D reaching during straight back-and-forth movements, and whether this effect depends on whether moving toward the avatar can reduce the spatial discrepancy between the participant’s real hand and the displayed avatar hand. Participants performed reaches under No-Distortion, Adaptive Elliptical Distortion, and Fixed Elliptical Distortion conditions. The distortion produced a small but significant increase in trajectory ovalization and reduced embodiment, but ovalization did not differ between distortion methods. These results suggest modest visuomotor interference rather than strong avatar following.
  • P5640 Extending Games for Evoking Emotions in Research Database for Virtual Reality and Positive Emotions Paweł Jemioło, Bartosz Chochół, Konrad Reczko, Eliza Petrycka, Dominika Bujnarowska, Antoni Ligęza
    Emotion elicitation is a central challenge in affective computing, as reliable measurement requires stimuli capable of inducing genuine emotional responses. While prior research has largely relied on passive stimuli, interactive experiences such as video games may offer greater ecological validity. The aim of this paper is to extend the Games for Evoking Emotions in Research (GEER) database by introducing games designed to elicit Ekman's positive emotions - enjoyment and surprise - and by adapting the existing negative emotion scenarios for virtual reality (VR). A quantitative study (N = 30) demonstrated that the new games evoked their target emotions, with VR amplifying emotional intensity, immersion, and positive affect compared to desktop conditions. A complementary preliminary qualitative study (N = 6) was conducted with negative stimuli on PC and VR. It highlighted both the strengths and limitations of VR-based emotion elicitation. Together, these results position GEER as a convenient, open-access repository for interactive emotion and affective computing research.
  • P5841 From Intent to Action: A Multimodal VR Task Dataset João Alves, Teresa Matos, Daniel Mendes
    This paper presents a multimodal dataset for analyzing user behavior during task performance in Virtual Reality. We collected data during a user study with 36 participants that performed free exploration, object finding, and navigation tasks across 4 virtual scenes with guided and free task structures. The dataset includes synchronized gaze, head, hand/controller, eye geometry, pupil, interaction, and task progression data. Subjective data about cybersickness, presence, and workload was also collected before, during, and after the VR experience. The dataset was analyzed to explore how behavior changes across tasks, guidance levels, and virtual environment contexts. Results showed that object finding produced stronger visual search and interaction behavior, while navigation showed a planning pattern where gaze tended to appear before movement. Overall, the dataset supports the study of intention-related behavior in VR and shows that multimodal signals should be interpreted together with the active task and scene context.
  • P6699 Multi-modal Augmented Spatial Environmental Data Asset Valuation Zhenggui Xiang
    Spatial environmental data assets have evolved into a core strategic asset for ecological governance, carbon neutrality, urban planning, and climate risk management. However, previous Data Asset Valuation (DAV) methods suffer from single-modal data dependence, insufficient spatial-temporal dynamic characterization, and lack of augmented feature fusion. In this paper, we propose a Multi-modal Augmented Spatial Environmental Data Asset Valuation (MASE-DAV) approach, which integrates multi-source heterogeneous data fusion (remote sensing images, ground monitoring, IoT sensing, text-based environmental policies, and social media perception), spatial-temporal knowledge graph augmentation, and a hybrid DAV model combining machine learning and income-cost method. Experiments verifies that our proposed MASE-DAV outperforms traditional DAV methods in accuracy, robustness, and adaptability.
  • P6822 A Look At Our Bodies: Automated Eye-Tracking Data Acquisition on Virtual Avatars Lisa Rieder, Marie Luisa Fiedler, Carolin Wienrich, Marc Erich Latoschik, Mario Botsch, Nina Döllinger
    Harnessing gaze patterns as an objective diagnostic tool for eating disorders and for guiding and assessing mirror exposure therapy (MET) is currently hindered by resource-intensive manual coding and intrusive physical markers. This paper presents a Virtual Reality (VR) application for automated, markerless gaze acquisition using photorealistic, personalized embodied avatars. Our framework dynamically generates skeletal-anchored Areas of Interest (AoIs) that are scaled based on the participant's Body Mass Index (BMI). It supports gaze mapping in both direct observation and virtual mirror-based self-observation through reflection-aware raycasting. In a technical validation study (N=5), the framework achieves high dwell-time registration for large anatomical regions, with values approaching the expected performance range of consumer-grade VR eye tracking (80-90% dwell time). However, results also identify challenges regarding torso segmentation, granularity, and smaller extremities. This work provides a proof-of-concept scalable gaze annotation in VR-based body-image research and identifies key requirements for future clinical applications, including improved full-body tracking and sex- and body-shape-sensitive AoI models.
  • P7099 Voice-to-3D: Generative AI for real time generation of 3D objects for Augmented Reality Rafael Conceição, Daniel Mendes, Daniel Sá Pina
    While augmented reality (AR) increases worker productivity, ap- plication development is heavily bottlenecked by the need to pre- create 3D assets. This paper presents an end-to-end pipeline gener- ating 3D assets in real time on an AR application: the user verbally describes an asset, and within four seconds the corresponding GLB mesh appears in their field of view. The backend components were selected via an objective benchmark of four image-to-3D models and a text-to-image preference study (n = 35). A System Usability Scale (SUS) evaluation was conducted ($n=18$) on the final AR application, yielding a score of 82.64.
  • P7318 Understanding Task Resumption Strategies after Physical Interruptions in 360 VR Videos Junho Jang, Jung Who Nam
    When real-world demands force users to remove their headsets, they can lose their spatial and narrative context upon returning to a suspended virtual task. To understand this disconnect, we conducted a formative study (N=8) exploring how users recover their place in 360-degree VR videos. By evaluating different video characteristics and interruption durations during unstructured viewing, we establish design criteria to inform future adaptive context-recovery systems.
  • P7455 A Dual-System Decision-Making Framework for Virtual Pet Behavior Generation Yidan Pan, Dongdong Weng, Haiyan Jiang, Nan Gao
    This poster introduces a dual-system decision-making framework for generating virtual pet behaviors in virtual environments (VEs). Within this framework, the virtual pet functions like an intelligent agent, capable of moving within the virtual environment, perceiving objects, and performing various behaviors, thus providing users with a more immersive interactive experience. The framework generates executable behaviors based on the pet's internal states (including hunger, thirst, energy, toilet, tension, comfort, and curiosity) and the VE, by combining a rule-based fast-thinking system with a large language model (LLM)-based slow-thinking system. It improves the rationality, naturalness, and diversity of virtual pet behaviors.
  • P7600 ViT-Split+: In-Place Adaptation of Vision Foundation Models for Efficient Semantic Segmentation Andrei Chubarau, Prakhar Kulshreshtha, Brian Pugh, Salma Jiddi
    Semantic segmentation of indoor scenes is key for augmented reality (AR) in interior design. Among recent methods, ViT-Split achieves competitive accuracy with fewer trainable parameters by directly adapting a DINOv2 backbone for dense prediction, fine-tuning a task head initialized from its final transformer layers. However, ViT-Split duplicates layers for the task head, inflating compute, while its feature fusion draws from uniformly spaced, potentially redundant frozen layers, diluting informative features. We further find that upgrading ViT-Split to DINOv3 degrades accuracy, despite the stronger backbone. We address these limitations with ViT-Split+, which (i) splits the transformer in place at a single pivot index, (ii) selects the pivot and fused layers via a Centered Kernel Alignment (CKA) analysis, and (iii) realizes significant accuracy gains with DINOv3. On ADE20k at 512×512, ViT-Split+ improves over ViT-Split by +1.4 mIoU (58.2 → 59.6) while reducing inference FLOPs by ~35%. For AR deployment, ViT-Split+ shows comparable gains on our in-house high-resolution indoor dataset, improving downstream defurnishing and scene editing pipelines.
  • P7883 Peripheral Motion for Attention Guidance in Virtual Reality Carolina Gonçalves, Teresa Matos, Rui Rodrigues
    Virtual Reality applications require attention guidance mechanisms that balance effectiveness and immersion. This work investigates the design of a covert attentional cue through the use of peripheral motion to direct users toward targets outside their field of view, while an animated outline helps locate the target once it becomes visible. A user study with 55 participants compared the proposed approach with an overt arrow cue and Subtle Gaze Direction in static and dynamic virtual environments. Results showed that the motion cue effectively guided attention while being perceived as less distracting, more subtle, and comfortable.
  • P8082 Investigating Human Perception and Spatial Learning During Active Locomotion: A Modular XR Framework Lior Maman, Tsvi Kuflik, Sarit Szpiro
    Traditional screen-based cognitive paradigms offer high experimental control, but require participants to remain seated. Yet in everyday life, perception and cognition occur during active physical movement, limiting ecological validity and leaving open fundamental questions about how these processes differ during locomotion. We present a framework for designing cognitive experiments adapted to 3D space and active locomotion, implemented in XR headsets, illustrated through three use cases: spatiotemporal statistical learning, dual-task cognitive load during walking, and visual prior violation through a shape-from-shading paradigm. Together, these cases demonstrate XR's potential as a powerful platform for cognitive neuroscience, enabling investigation of perception and cognition.
  • P8727 Visual Fatigue During Context Switching in Naturalistic Tasks for AR Interaction Nora Jane Castner, Yannick Sauer, Björn Severitt, Rajat Agarwala, Siegfried Wahl
    This study examines how augmented reality (AR) focal distance affects visual comfort and performance during naturalistic tasks. Using a commercial binocular headset (RayNeo X3 Pro), participants reproduced 3D toy-brick models on a desk within 45 seconds, repeated shifting attention between virtual instructions and physical bricks. Vergence distance was matched to the desk while four focal distances (0D,0.5D,1D,1.5D) were presented counterbalanced. There were no significant effects of eye tiredness, concentration, focus switching, double vision, or performance, but significantly increased reported blur at 0D versus 1.5D, indicating focal distance systematically modulates blur-related discomfort in AR-guided near-work tasks.
  • P9220 Visibility and Perceptual Impact of Distortions in Augmented Reality (AR) Image Composition Mahzar Eisapour, Zhou Wang
    Mobile and wearable augmented reality (AR) systems operate under rendering, computation, and transmission constraints that can introduce visible composition artifacts. This paper analyzes distortion visibility using the ARC-IQA dataset, which contains 193 AR compositions with subjective mean opinion scores (MOS). We study controlled appearance and structural degradations, including luminance mismatch, chromaticity variation, contrast modification, texture compression, and geometric perturbation, and relate them to quantitative content descriptors. Results show that luminance and contrast inconsistencies cause the strongest perceptual degradation, whereas textural complexity can partially mask compression artifacts. The analysis further shows that geometric complexity, texture entropy, and colorfulness influence how strongly the same distortion level is perceived across different virtual objects. These findings provide practical guidance for perceptual optimization, AR quality assessment, and resource-aware multimedia rendering.
  • P9697 The Ventriloquist Effect as a Design Mechanism for VR Masking Sound Therapy Tristan Gabriel Mona, Luca Eastwood, Meng Du, Philip J Sanders, Burkhard C. Wünsche
    Tinnitus is the perception of sound without an external source and can severely reduce quality of life. Masking sound therapy aims to reduce attention to tinnitus by blending it with external sounds. In VR, masking sounds can be paired with visual objects, but it remains unclear how these objects influence perceived sound location. This paper investigates how visual masking-sound objects affect sound localisation in VR. A user study with 20 participants shows that visual objects can bias perceived sound location, demonstrating a ventriloquist effect. These findings suggest that visual objects can help shape masking-sound perception in therapeutic VR environments.

Posters 2BCassiopeaCollaborative XR, Avatars & Navigation Thursday14:15–16:30

  • P1178 Sea-Scale Situation Awareness for Subsea Infrastructure Security in XR: Design Insights for Immersive and Non-Immersive Tools Andrea Nordwall, Giulia Wally Scurati
    Subsea infrastructure faces an increased risk of damage as geopolitical tensions rise in the Baltic Sea region, necessitating cooperation between public and private actors. Current monitoring systems used by national coast guards and cable solutions providers are based on 2D sea chart data visualization, with limited capability for displaying contextual information, and do not effectively support collaboration and knowledge sharing. Based on an iterative user-centered design process, we introduce a concept design for sea-scale situation awareness, integrating 3D bathymetry, satellite, ship transponder, geodata, and cable system sensor data into an integrated view, evaluating immersive vs non-immersive XR modalities.
  • P1710 Adaptive Mixed-Reality Spatial Prompts for Navigation in Cultural Environments Meng Xia
    Wayfinding support in mixed reality (MR) is usually delivered as fixed, turn-by-turn overlays that substitute a visitor’s spatial reasoning. This is efficient for routing but poorly suited to cultural environments, where orientation and exploration matter alongside reaching a target. We present an early-stage framework that treats embodied navigation behaviour (dwell, hesitation, backtracking, heading change from head-mounted display tracking) as signals for lightweight inference of a visitor’s navigation state, then adapts MR spatial prompts accordingly, from in-situ cues to withheld guidance. We describe the pipeline, a museum walkthrough, and a planned within-subjects evaluation. No user study has been conducted.
  • P2010 HARI-AR: A Heritage Agent Reasoning Interface for Autonomous Outdoor Navigation and Semantic Spatial Reasoning Prathik M Hadagali, Himangshu Sarma, Mrinmoy Ghorai
    Outdoor heritage and pilgrimage sites present navigation challenges that existing AR systems have not addressed: GPS-grade localization without structured geometric models, culturally specific destinations expressed in informal or vernacular language, and landmark-rich but semantically unannotated spaces. We present HARI-AR, an AR navigation system deployed at the Tirumala Venkateshwara temple complex, integrating a five-agent LLM orchestration pipeline over a custom 3,650-node OSM navigation graph enriched with 155 hand-curated semantic POIs to deliver landmark-anchored AR guidance via natural language voice queries. Rather than proposing a new AR rendering paradigm or novel agent architecture, the primary contribution is the deployment, empirical evaluation, and validation of semantically grounded landmark-anchored guidance in a real outdoor heritage environment at scale. A within-subjects study (N=24) showed landmark-referenced instructions significantly reduced directional errors compared to GPS-arrow-only guidance (21/24 vs. 14/24 correct turns; Fisher's exact p<0.05), with a SUS score of 74.2, establishing that natural language landmark navigation is both feasible and measurably effective for first-time users in complex cultural environments.
  • P2489 Asymmetric AR Access for Co-Located Collaborative Work Michael Anderson, Rajas Kejriwal, Nelusa Pathmanathan, Kuno Kurzhals, Alexander Plopski
    Collaboration is a core component of augmented reality (AR) usage, especially in co-located settings where shared visualizations can improve performance and communication. While symmetric AR collaboration has been widely studied, asymmetric settings where only one or a subset of users have access to AR content remain underexplored. To investigate this phenomenon, we conducted a user study with 24 participants working in pairs on a tangram puzzle task. Participants either used an AR headset (Hololens 2) with AR guidance in the form of puzzle outlines in the workspace or, when having no access to the AR instructions, utilized Aria glasses for non-intrusive eye tracking. Results demonstrate that AR improves objective performance; however, the asymmetric constellations significantly change team communication, gaze behavior, and subjective feelings about the collaborative process. This situation reveals challenges for supporting effective collaboration when AR access is unequal.
  • P2802 From the Lab to the Shop Floor: Investigating eXtended Reality Use in Industrial Scenarios through a Human-Centered Lens Bernardo Marques, Fábio Barros, Pedro Reisinho, João Alves, Mariana Leite, Carlos Ferreira, Tiago Coelho, Jorge Neves, Jorge Figueiredo, Duarte Almeida, Paulo Dias, Beatriz Sousa Santos
    Most eXtended Reality (XR) applications for industrial scenarios are developed and validated under controlled laboratory conditions, where environmental variability, operational constraints, and real- life workflows are partially represented. Although such evaluations provide important evidence of feasibility and usability, they often limit understanding of how XR performs when integrated into industrial routines, where safety requirements, task interruptions, and contextual factors can influence adoption. This work reports an XR framework designed for industrial application domains, exploring individual and collaborative contexts (co-located/remote), through virtual and augmented features interconnected with a factory digital twin to support human technicians. The framework was developed through a human-centered process conducted in cooperation with industry partners across successive evaluation stages, from laboratory studies to shop floor deployment. Early findings indicate promising potential for XR adoption while revealing key design considerations for transferring XR solutions beyond controlled settings.
  • P2823 VRSlate: In-Headset Recording Management for Social VR Filmmaking Wang Haoyu, Xin Lyu, Zhaohan Wang, Yurun Chen, Zihan Gao
    Social VR creators film inside VRChat while recording controls remain on the desktop. VRSlate brings recording control, tally feedback, scene/shot/take labels, and post-take handoff into the camera operator's in-headset avatar-menu workflow. It bridges VRChat OSC parameters with OBS WebSocket, records crash-resilient MKV takes, remuxes H.264/AAC media to MP4, and writes a CSV shot-list index. In one author-led project timing, an off-HMD manual control step took 20--34 s, versus about 3 s with VRSlate.
  • P3590 First Comes Trust, Then Comes the Slip-of-Tongue: How Social Cues and Reciprocity Drive User Self-Disclosure with Virtual Agents Natalia Lipp, Fabian Schunk, Cornelia Wrzus
    This experimental study investigates how characteristics of virtual avatars build trust and encourage user (excessive) self-disclosure within a VR-based digital interview context. Thus, we manipulate agents’ social cues and conversational reciprocity in an adult lifespan sample. We specifically focus on older adults, whose age-related changes in motivational priorities and cognitive control may increase their susceptibility to oversharing. Self-disclosure is operationalized via response rates and elaboration during conversations about personal information. By evaluating these dynamics, this research clarifies privacy boundaries and rapport-building in human-agent interactions across different age groups. Data collection is ongoing, with final results expected in August 2026.
  • P3853 Preliminary Evaluation of Transition-Aware Locomotion Adaptations for Supporting Postural Stability in VR Ramisa Fariha Joyee, M. Rasel Mahmud
    We present three locomotion adaptation approaches: Motion Acceleration, Turn Acceleration, and Motion Deceleration to improve postural stability during body-state transitions in virtual reality (VR). The system detects standing-to-walking, turning, and walking-to-stopping transitions and applies adaptative locomotion smoothing. Motion Acceleration gradually increases locomotion speed when users begin walking, Turn Acceleration smooths rotation while turning, and Motion Deceleration gradually reduces movement speed before stopping. We evaluated these techniques in a virtual navigation task using objective balance measures. Preliminary results show reduced center of pressure (COP) velocity and improved balance confidence. These findings suggest that locomotion adaptations can improve balance and navigation experience.
  • P4196 “If Today Were a Place”: Designing Agent-Guided Journaling for AI-Generated Reflective Virtual Spaces Ruoyu Wen, Daksitha Senel Withanage Don, Xiaohan Wei, Binyang Han, Niklas Heimerl, Simon Hoermann, Elisabeth André, Alaeddin Nassani, Mark Billinghurst, Thammathip Piumsomboon
    Journaling has been widely used in mental well-being contexts to support self-awareness, emotional reflection, and emotion regulation. Digital journaling extends this practice by reducing writing effort and enabling interactive, guided, and more engaging reflection. In this paper, we present DayScape, a digital journaling system that combines an LLM-driven virtual agent with generative VR. Instead of relying only on text-based entries, DayScape transforms users’ daily reflections into immersive, supportive virtual spaces with generated environments, background soundscapes, and music. We conducted a pilot study with eight participants to evaluate system usability, the perceived quality of the generated VR experiences, and users’ views of VR-based journaling. Our findings provide early insights into how generative VR could support daily reflection and inform the design of a future full-scale study and the development of generative VR systems for mental well-being contexts.
  • P5348 Physically Interactable Gaussian Splatting for Shared MR–VR Remote Inspection Semin Bae, Sungwook Choi, Chul Min Yeum, Jongseong Brad Choi
    Inspecting large or hazardous facilities requires scarce on-site experts. Video inspection locks the expert's viewpoint; point clouds lack the fidelity to judge fine surfaces. We present a shared MR–VR teleinspection system where an on-site MR user and a remote VR expert co-inhabit one photorealistic 3D Gaussian Splatting replica—shared, metric, and touchable at once, unlike single-user or unregistered 3DGS. A lightweight offline pipeline (single-marker scaling, 2DGS collider extraction, image-based localization) registers users and assets into one common metric frame, enabling on-surface pointing, metric measurement, and cross-reality annotation. We validate each component in a proof-of-concept room.
  • P6196 From Smart Devices to Home Companions: Designing Anthropomorphic Avatars in Mixed Reality Chenyue Zheng, Hu Yong
    In MR environments, virtual avatars play an important role in supporting users’ understanding of smart home device states, communicating behavioral intentions, and enhancing emotional experience. However, existing research has offered relatively limited discussion on the visual expression of virtual avatars for smart home products.This study first conducted semi-structured interviews to identify the suitable application scope of virtual avatars for smart home devices. It then employed participatory design to explore a design framework for anthropomorphic avatars in MR environments from the perspectives of product attributes and user emotions. Based on these findings, an initial MLLM-based prototype was developed to explore the workflow from physical device recognition to virtual avatar generation. This study aims to move beyond the design of individual visual forms in smart home anthropomorphism and develop a strategy-oriented framework grounded in design metaphor. By providing structured rule constraints for MLLMs, the proposed approach supports more systematic and context-aware avatar generation, offering a new perspective for the design and application of smart home avatars in MR environments.
  • P8041 Effects of Avatar Appearance on Users' Perception in Social VR: A User Study with Seated Acquainted Dyads Selina Palige, Katharina Precht, Angelika C. Bullinger
    This study investigates the effect of avatar design on users’ perception in social virtual reality (VR). While it is a common understanding that avatar usage enhances user perception in VR applications, little is known about the effect of different avatar designs on the user perception within social scenarios with seated positions. We conducted a between-subjects user study with 70 participants (35 acquainted dyads) attending a virtual opera event. Bayesian statistical analysis revealed equivalence for hand-and-hands vs. full-body avatars for all measures of user perception applied. Systematic content analysis of post-experiment interviews show 90% of participants preferred full-body representations, nonetheless.
  • P8235 Remembering Together: How Community-Based Design Turns Memory into a Social VR Experience Karolina Wylężek, Patricia de Torres Coll, Irene Calvis, François Matarasso, Ibrahim El Shemy, Katrien De Moor, Sergio Cabrero Barros, Pablo Cesar
    As history shapes current perspectives, the way we present past events should be carefully curated, enabling responsible societies. In this work, together with the relevant communities (ex-inmates, contemporaries, experts), we co-design and evaluate a social VR experience showcasing the history of the Franco regime’s political prisoners. The results, analysed alongside current literature, show that depending on the type of memory the groups hold, they focus on different aspects of the experience during the co-design and evaluation. This finding suggests that to create meaningful memory-based experiences, all groups shaped by that memory should be engaged in the design process and decisions.
  • P8236 Hearing Through Portals: Preliminary Results on Auditory Cues for Navigation in VR Impossible Spaces Paulo Benety Mendes de Lima, Elizabeth Rose Urquhart, Jack Freeth, Oliver Mitchell, Samuel Ou, Dominik Lange-Nawka, Samuel E R Thompson, Burkhard C. Wünsche, Steffan Hooper, Elliott Wen
    Impossible Spaces in virtual reality enable navigation of large virtual environments within limited physical areas by connecting overlapping spaces through portals. We investigated whether spatial audio propagating through such portals helps users to navigate efficiently. We conducted a user study with 31 participants comparing navigation in impossible spaces with and without propagated audio. No significant differences were observed, however, qualitative feedback varied: some participants found audio helpful, and others perceived it as overwhelming. This suggests that spatial audio through portals may not consistently improve navigation in impossible spaces, and benefits may depend on individuals' cognitive and auditory processing skills.
  • P8867 CrisisX: Emotionally Adaptive VR Training for Healthcare Crisis Communication Viet-Tham Huynh, Ngoc-Thu Thi Tran, Anh Cong Phan, Nu Tra-Mi Le, VI KY LE, Tam Nguyen, Minh-Triet Tran
    Existing clinical communication training is limited by cost, scalability, and static interactions. We present CrisisX, an AI-driven VR system featuring an emotionally adaptive virtual patient. Driven by a novel Emotional State Engine and large language models, the agent dynamically escalates or de-escalates in response to the trainee’s voice interactions and provides explainable, rubric-based feedback. A two-phase evaluation, culminating in a clinical pilot with 14 health-profession students, demonstrated high immersion, perceived usefulness, and good usability (UMUX-LITE 76.8/100). Crucially, workload analysis revealed a high-effort, low-frustration cognitive "sweet spot," proving particularly effective at bridging the practical experience gap for novices. These results suggest CrisisX offers a scalable, emotionally realistic approach to healthcare communication training.
  • P8884 The Role of Human-Like Embodiment and Conversational Cues of Intelligent Virtual Agents in VR for Mental Health Support Lukas Oberfrank, Lucie Wittig, Frank Steinicke
    The increasing demand for mental healthcare, and recent advances in large language models (LLMs), have led to the widespread adoption of intelligent virtual agents (IVAs). IVAs can be designed with a diverse range of visual features, and LLMs offer customization options for their communication. In our 2x2 within-subject study, we examined how visual human-likeness and conversational style of IVAs in virtual reality (VR) influence perceived trust and rapport. 35 participants engaged in psychotherapeutical self-help exercises with four different IVAs that varied in embodiment (human-like vs. abstract) and conversational style (human-like vs. mechanical). Human-like conversational style had the strongest positive effect.
  • P8915 Empathic Concern and Task Performance in VR Tangram Collaboration Heesook Shin, Yongho Lee, Gun A. Lee, Seungwon Kim, Youn-Hee Gil
    Even with the same VR collaboration system and task, collaboration outcomes can vary by participant characteristics. This study examined how empathic concern relates to performance and dialogue structure in a VR director–matcher Tangram task. Thirty-two dyads completed the task, and completion time and matcher turn count were measured per round. Matcher's empathic concern was found to be positively associated with longer completion time, and this relationship was statistically explained by increased matcher turn count. These findings suggest that empathic concern may be associated with greater conversational involvement, which can appear as a communication cost in time-sensitive VR collaboration.
  • P9065 Translating Text to Italian Sign Language Using 3D Avatars Alessandro Clocchiatti, Flavio Dumitru Roman, Agata Marta Soccini
    Communication barriers for Deaf individuals persist in everyday scenarios where interaction relies primarily on speech or text, limiting accessibility and increasing the risk of social exclusion. While translation frameworks exist for widely documented sign languages, real-time pipelines for Italian Sign Language (LIS) remain scarce. We present a real-time text-to-LIS framework combining a Large Language Model-driven multi-agent architecture with a motion-capture dataset for 3D avatar animation. Our approach handles out-of-vocabulary signs, supports multiple avatars, and is highly extensible to other sign languages and vocabularies. A user evaluation with Deaf users provided overall promising feedback, suggesting potential extension to Extended Reality applications.
  • P9844 A Pilot Study to explore Gaze and Speech in Avatar and Robot Mediated Communication with Older Adults Stephanie Arevalo Arboleda, Felix Immohr, Jakob Hartbrich, Florian Weidner, Melisa Conde, Veronika Mikhailova, Söhnke Benedikt Fischedick, Bea Vorhof, Christoph Gerhardt, Kay Richter, Christian Kunert, Nicola Döring, Horst-Michael Gross, Wolfgang Broll, Alexander Raake
    Older adults are a growing demographic group who could bene fit from communication technologies such as telepresence robots and augmented reality avatars. However, they constitute a distinct user group due to age-related changes and life-course experiences. We present a pilot study exploring how older adults perceive (self-reported measures) and behave (gaze and speech behavior) in avatar- and robot-mediated communication, and contrast these results with face-to-face (F2F) interaction. Our initial findings suggest that representation-specific characteristics shape gaze and speech behavior: in robot-mediated communication, multiple robot components may compete for gaze, whereas in avatar-mediated communication, behavioral realism and novelty may influence gaze. We further observe a divergence between self-reported and observed gaze behavior, alongside subtle intra- and inter-individual differences. Together, these initial insights point to representation- and user-dependent effects relevant to the design of future mediated communication systems for older adults and that possibly extend to other demographics.

Friday – October 9, 2026

Posters 3ACassiopeaEmbodied Interaction, Multimodal Control & Assistive Systems Friday11:15–13:45

  • P1387 Effects of Body-Movement Remapping and Attachment Position of a Virtual Supernumerary Arm on Embodiment, Workload, and Performance Harin Hapuarachchi, Michiteru Kitazaki, Hiroaki Shigemasu
    Virtual supernumerary limbs require effective mappings between the user’s body movements and the additional limb. We investigated how body-movement remapping and attachment position influence embodiment, workload, and reaching performance of a virtual supernumerary arm. Twenty-four participants controlled the arm using shoulder or foot movements while the arm was attached to the shoulder or hip. Ownership and agency were not significantly modulated by either factor. However, shoulder attachment reduced subjective workload, whereas foot control enabled faster reaching. These results reveal a design trade-off: anatomically proximal attachment may reduce workload, while foot-based remapping can improve task performance without reducing embodiment.
  • P1714 An Augmented Reality Prototype for Spatial Attention Assessment and Rehabilitation in Visuospatial Neglect Rim Yu, Eui Jung, Hyo Suk Nam, JoonNyung Heo, Eunjeong Park
    Visuospatial neglect following stroke significantly affects patients’ functional recovery and their ability to perform daily activities. However, traditional assessment methods often require considerable administration time and have limitations in quantitatively capturing behavioral metrics of neglect. This study aims to investigate the technical feasibility of an augmented reality (AR)–based system for assessing and rehabilitating spatial attention deficits. We developed an AR-based system using lightweight AR glasses that present visual stimuli and enable the measurement of user responses through interactive tasks. The proposed system includes multiple spatial attention assessment and rehabilitation tasks, including a Ball Tracking task, Line Bisection task, Target Search task, and Line Following task. Behavioral data collected during task performance enable quantitative assessment of spatial attention functions. The system utilizes lightweight AR glasses with 6 Degrees of Freedom (6-DoF) head tracking to capture user interactions through head movements alone. This approach enables the extraction of behavioral metrics, such as reaction time, spatial bias, and visual search patterns, within an interactive virtual environment. The results suggest that diverse behavioral metrics related to spatial attention can be quantitatively derived through 6-DoF head-movement interaction. The proposed AR-based system demonstrates the potential of a digital rehabilitation platform for the quantitative assessment and rehabilitation of spatial attention functions.
  • P2486 Stepping Beyond the Stones: A Mixed Reality Stonehenge Experience with Interactive Augmentations and Virtual Portals Kennard Nicolai, Christian Geiger
    We present a mixed reality (MR) Stonehenge experience based on the stones in England that combines a spatially registered digital twin reconstructed from LiDAR and photogrammetry with a physical replica. The system enables embodied hand tracked interaction, allowing users to apply and trigger virtual augmentations directly grounded on the stone structures, including painting, animation, and video playback visible through a Head Mounted Display (HMD). Because the physical setup can only represent a subset of the monument, we introduce a virtual portal mechanism that seamlessly connects localized MR with a full-scale virtual reconstruction of Stonehenge and its surrounding landscape. A preliminary user study with 14 participants reports positive usability and user experience, indicating that virtual portals can effectively extend small-scale embodied interaction into large-scale cultural heritage experiences.
  • P3002 Scene-Graph-Driven Olfactory Rendering: A Metadata Rule Engine for Large VR Scenes Rodrigo M. de Carvalho, Pedro Koziel Diniz, Gustavo Peres de Faria, Thalles H. B. da Silva, Cleiver Batista da Silva, Marcelo Ramos Jordão, Rui Gonçalves de Oliveira Júnior, Melina Mottin, Arlindo Rodrigues Galvão Filho, Carolina Horta Andrade
    Deploying olfactory systems in large authored virtual reality (VR) scenes is limited by the manual cost of assigning odor sources. Perobject placement tools such as the Smell Engine and bubble-trigger frameworks require each odor source to be configured individually. Automated computer-vision (CV) pipelines avoid manual placement, but add per-frame inference latency. We present a metadata-driven rule engine that queries object names, classes, and materials already exposed by the Unreal Engine 5 scene graph at runtime. The engine replaces per-object placement with a compact set of reusable semantic rules. Across three 500×500 m scenes with 200+ heterogeneous objects, the system sustains 58.7±2.1 Hz updates at 52.3±8.7 ms control-loop latency, a 2.8× reduction over a reimplemented visionbased baseline, while requiring a one-time ruleset per odor taxonomy instead of per-scene manual configuration.
  • P3026 Augmented Reality for In-Situ Soil Moisture Visualization to Support Irrigation of Sports Fields Lara Flaig, Julia Hertel, Frank Steinicke
    Large grass sports fields are often watered manually, requiring greenkeepers to know the moisture content of the soil in order to identify which areas need watering. This requires experienced greenkeepers to visually inspect the grass, but can also be supported by the use of sensors to measure moisture. However, such moisture data is typically displayed on 2D interfaces, requiring greenkeepers to mentally transfer this data to the actual area. This work investigates whether greenkeepers see potential in Augmented Reality (AR) to visually integrate such data directly into the real world. A human-centered design is presented that includes a requirement analysis, the development of a prototype, and an evaluation of the prototype in a preliminary field study involving greenkeepers. The results show that greenkeepers see potential and expect their work to be facilitated by seeing moisture data directly on the grass. Further, additional feature wishes and current challenges are presented.
  • P3665 Spatially Grounding VLA Robot Manipulation with Immersive Grasp and Place Targets Daniia Zinniatullina, Iaroslav Kolomiets, Mikhail Konenkov, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
    Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets users specify grasp and placement regions and renders them as image-space overlays for VLA control. We collect 200 Unity pick-and-place demonstrations and fine-tune OpenVLA-7B with LoRA on temporally subsampled annotated observations. The policy predicts tokenized 7-DoF incremental actions from marked RGB observations and language. We evaluate the policy in closed-loop Unity trials, achieving a grasp success rate of 91.25% and mean grasp and placement errors of 0.5 cm and 0.7 cm, respectively.
  • P4560 An AI-Assisted XR Framework for Continuous Design Validation in Engineering Prototyping Ali Darejeh, Eliza Lee
    Engineering prototyping workflows frequently separate design manipulation, compliance checking, and design rationale across CAD, simulation, and documentation tools. This separation delays feedback and can increase the cost of early decisions. We present an AI-assisted XR framework for continuous design validation in engineering prototyping. The framework integrates spatial inspection, in-situ design manipulation, spatial constraint visualization, and AI-supported design reasoning in a unified immersive workflow. We demonstrate the framework through a VR-based automotive engineering application that integrates AI-supported design reasoning with real-time constraint validation. The system supports full-scale inspection, exploded assemblies, semantic labels, interactive component modification, real-time constraint feedback, and AI-generated explanations of engineering trade-offs. The work reframes XR prototyping as an active validation workspace rather than a passive visualization environment.
  • P5056 VR-GPT: VLM-based Task Assistant for Natural Language User-Computer Interaction in Virtual Environments Mikhail Konenkov, Daria Trinitatova, Artem Lykov, Dzmitry Tsetserukou
    Industrial VR is increasingly used for operator training, work-cell simulation, and procedural guidance, yet VLM integration in interactive VR remains underexplored. We present VR-GPT, a Unity-based system that couples a domain-fine-tuned Qwen2-VL-7B model with voice interaction for context-aware task assistance without visual text overlays. The HMD view is streamed to a VLM server via WebSocket, while Whisper STT and Jets/ESPnet TTS support conversation. Fine-tuning on 400 VR-scene images using QLoRA improved task-guidance success by 1.97× over the baseline. In a pilot study (n=6), VR-GPT reduced execution time by 13.7% and wrong actions from 2.6 to 1.7, indicating improved practical usability.
  • P5498 Vid2Spatial: Controllable Spatial Audio Authoring by Parameter Extraction from Video and Sketch Seungryeol Paik, Kyungsu Kim, Kyogu Lee
    Spatial audio is central to presence in AR/MR, yet authoring it means entering long sequences of 3D positional parameters by hand. Vid2Spatial reframes authoring as spatial parameter extraction: a per-frame stream of azimuth, elevation, and relative distance is extracted from monocular video or a freehand 3D sketch, stored as an editable artifact, and routed to binaural, Ambisonics, and OSC consumers without re-extraction. A controlled study (N=12) shows 41-49% faster authoring and 87-90% fewer edits; tracking reaches 1.36-degree and generalizes across 24 non-COCO categories, while an end-to-end audio-visual baseline recovers no usable azimuth; a perceptual study (N=20) adds support.
  • P5609 WavePlay: Static and dynamic hand gestures for playback controls in VR 360° films Tanushree Pillai, Jayesh S. Pillai
    360° VR video playback controls rely on spatial UIs, adapted from desktop interfaces, which creates occlusion and disrupts the sense of presence. WavePlay is a novel media control interface using mid-air hand gestures, eliminating UI overlay entirely. A pilot usability study (N=24) compared WavePlay against the spatial UI of SkyboxVR across task completion speed, presence, and perceived agency. WavePlay achieved comparable performance on all measures, however, hand tracking limitations introduced minor fatigue. Findings reveal core trade-offs between spatial UI familiarity and direct gesture interaction in controller-free contexts. These insights inform design guidelines for hand gesture-based VR media playback control.
  • P5779 Developing a Voice-Driven Mixed Reality Visual Analytics System for Immersive Data Exploration Joel Saji Varghese, Tarang Rana, Muhammad Minhajuddin, Niya Jose, Somang Nam
    In this paper, a mixed-reality visual analytics tool using a conversational user interface for queries and code generation with a Large Language Model. The query text is combined with data metadata to generate results to display as 2D data visualizations with a digestible summary. Users can convert the visualizations in a 2D window into 3D spatial elements, affording immersive interactions to show how the data is structured. The system allows users to ask questions to learn further in layman's terms. This work aims to contribute to building an accessible data interpretation tool for non-experts who have no data analytic knowledge.
  • P6164 OmniAI: Spatially Adaptive Aerial Agent with Surface-Aware Projection, Gesture Control, and Web-Augmented Interaction Nikita Kuzmin, Yuhua Jin, Georgii Demianchuk, Mariya Lezina, Fawad Mehboob, Ivan Valuev, Nikolai Lutsenko, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
    Drones in human environments often lack spatially grounded interfaces for situated communication. We present OmniAI, an embodied aerial agent that supports surface-adaptive interaction by switching projection between an onboard screen and nearby environmental surfaces. A servo-actuated MEMS laser projector renders text-and-image responses from a web-augmented LLM pipeline. Projection surfaces are detected online using RGB-D sensing and RANSAC plane fitting, without pre-mapped geometry. OmniAI provides functionally equivalent voice and gesture control for both drone motion and projected content. By combining speech, mid-air gestures, adaptive projection, and aerial mobility, OmniAI demonstrates a mobile spatial AR interface for context-aware human-drone interaction.
  • P6950 Manipulating Objects in DeskVR using a Continuous Interaction Space Hugo Almeida, Daniel Mendes
    Virtual Reality (VR) can enhance visualization and enable natural interactions, but it can cause physical fatigue during extended use. DeskVR mitigates this by enabling seated interaction. This work introduces MOCIS (Manipulating Objects in DeskVR using a Continuous Interaction Space), a hybrid technique that combines touch-based input and mid-air gestures to reduce fatigue during object selection and manipulation, featuring input scaling and adaptive filtering. A user study compared MOCIS to a mid-air baseline (Scaled HOMER). Results showed successful task completion, but MOCIS was slower, more complex, and harder to learn. The findings reveal design limitations and highlight areas for future improvement.
  • P8988 Comparison of 3D Sketching Modes in Mobile Augmented Reality Lennard Grimm, Christoph Holtmann, Johann Habakuk Israel
    With the increasing ubiquity of mobile devices and increasingly accurate location services, mobile augmented reality (MAR) is gaining in importance. Some MAR systems use drawing and sketching techniques to create 3D objects on site. While prior work has explored 3D sketching in virtual reality (VR) with head-mounted displays and dedicated controllers, MAR sketching on smartphones remains underexamined. In this paper, we present the design and implementation of three modular sketching tools within a smartphone AR prototype, enabling six distinct sketching modes, and an empirical evaluation comparing these configurations. Our study investigates user experience, efficiency, and usability, revealing that canvas-based sketching (2D surfaces placed in 3D space) is perceived as fastest and most natural, while a combination of freehand “precision” drawing and “connect” snapping aids produces the best results for fully 3D sketches. We discuss challenges in MAR evaluation - such as depth perception and input-display coupling - and provide guidelines for future mobile 3D sketching applications.
  • P9886 Tactonality: Expressing Character Personalities through Wrist-Worn Vibrotactile Feedback Lim Junbeom, YuJune Kang, Gyeore Yun
    Haptic feedback in XR has been mapped to perceptual attributes and affective states, but whether it can convey a character’s personality remains unclear. We examine this question using a wrist-worn vibrotactile device and 108 parametrically designed stimuli. Twenty participants rated each vibration along the OCEAN dimensions. All five vibration parameters significantly affected perceived personality, and a three-component PCA explained 96.6% of the variance. Relating the ratings to the PCA space, we interpret the components as personality axes and show how vibration parameters shift stimuli along them. This mapping supports personality-oriented vibrotactile design for character-rich XR applications.

Posters 3BCassiopeaInteraction & XR Interfaces Friday14:15–16:30

  • P1854 Effects of Contextually Competitive Virtual Environments on Selection Studies in Virtual Reality Eric DeMarbre, Hamza Bashandy, Robert J. Teather
    We present a preliminary study on how the visual presentation of a virtual environment influences performance on a Fitts’ law selection task, which is known to be repetitive and tedious. We compare a "neutral" or sterile environment with a detailed arcade environment suggestive of a competitive atmosphere. Results suggest that the environment had a limited impact on performance, but participants had a more positive impression of the more detailed environment.
  • P2380 Cross-reality XR for Situation Awareness in Drone Operations Luljeta Sinani, Florian Baugé, Frank Steinicke, Andreas Gerndt, Georgia Albuquerque
    This work-in-progress presents an eXtended Reality (XR) interface for Unmanned Traffic Management (UTM). As new airspace regulations emerge to accommodate automated drone flights, representing dense traffic on 2D displays limits spatial understanding of complex urban environments. Designed based on stakeholder input, we developed a cross-reality XR interface that combines Mixed Reality (MR) tabletop overviews with immersive Virtual Reality (VR) to support visualization, monitoring, and planning of drone operations. We integrate our prototype with a real UTM system to show technical feasibility. Preliminary expert feedback indicates promising potential for XR in unmanned airspace management and reveals important considerations for future work.
  • P2707 Couch-AR: Soft Furniture as Interactive Augmented Reality Surfaces Alexandru-Tudor Andrei, Radu-Daniel Vatavu
    We present Couch-AR, a spatial augmented reality system that transforms soft furniture into interactive surfaces for lean-back seated interaction. While prior spatial AR systems have focused on rigid surfaces such as tables, floors, and walls, soft furniture remains underexplored. In this work, we showcase seated interaction via touch, swipe, and mid-air gestures for couch-based content navigation and manipulation. A preliminary evaluation revealed a positive user experience across UMUX, UEQ+, and NASA-TLX, which creates the premises for future investigations on seated AR experiences.
  • P2783 EmpathicMirror: Adaptive Mirror for Sharing Emotions through Reflections in Virtual Reality Theophilus Teo, Allison Jing, Heesook Shin, Yongho Lee, Youn-Hee Gil, Mark Billinghurst, Gun A. Lee
    VR collaboration allows users to share non-verbal expressions through visual cues. Existing techniques often project these expressions to various locations within the VR space, which can cause confusion as this is not common in real-world interactions. To address this, we introduce EmpathicMirror, a mirror that orients itself near the user's perspective to enable sharing expressions. It uses reflections to reproject VR users’ gestures and emotions to support a VR collaboration. It operates in three modes: Player Dominant, Fixed, and Partner Dominant. These modes manipulate the mirror's position based on the users' view directions or a fixed location in the VR space. Additionally, it rotates itself to allow users to see each other through reflections. Finally, it can facilitate sharing emotions through facial expressions by rendering visual cues and effects on or around the mirror.
  • P2860 A Design Framework for VR-Native Game Development with In-Headset Level Editing and Visual Scripting Ali Darejeh, Dylan Beretov
    VR game development remains largely desktop-centric, requiring developers to repeatedly switch between conventional authoring tools and head-mounted displays for testing. This paper proposes the VR-Native Game Development Framework, comprising four design principles: In-situ Spatial Editing, World-Embedded Visual Logic, Continuous In-Headset Iteration, and Scaffolded Complexity. We instantiated the framework through a functional Unreal Engine 5 environment supporting immersive level editing, visual scripting, integrated Build–Link–Play workflows, and persistent world management, and evaluated the resulting environment through an exploratory user study with 20 participants. Results showed promising usability (SUS = 73.88), low workload (NASA-TLX = 5.03/20), and high ratings for immersion, responsiveness, and satisfaction. The findings suggest that the proposed framework provides a complementary approach to immersive authoring and early-stage gameplay prototyping.
  • P3655 On Designing Task-Assisting HUDs for EVAs on the Moon Aurelia Föllner, Marius Grießhammer, Maic Masuch, Aidan Cowley
    This work explores the early design of Augmented Reality Head-Up Displays (HUDs) to support astronauts during lunar surface missions, using a simulated Extravehicular Activity task in Virtual Reality for user-centered development. Based on preliminary interviews, two HUD concepts were designed, implemented, and tested in a simulated deployment task on the lunar surface. Both HUDs were well received and considered propitious for providing assistance, particularly through a task procedure step list. Early stage feedback provides a foundation for future system iterations and indicates the target group's desire for user customization.
  • P3767 ORCESTRA: Immersive No-Code and Language-Guided Robot Programming on Mixed-Reality Digital Twins Ivan Snegirev, Elizaveta Semenyakina, Mikhail Konenkov, Artem Lykov, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
    ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and language-guided control. In a passthrough mixed-reality workspace, users place robot twins on real surfaces, teach trajectories, save robot-relative episodes, or issue spoken/typed commands that a vision-language model converts into structured digital-twin plans. Both interaction modes share a backend for metric grounding, embodiment-aware validation, preview, confirmation, and digital-twin execution. The system supports heterogeneous robot embodiments, including fixed-base manipulators, a mobile base, and a humanoid robot, demonstrating MR validation as a safety layer for language-guided robot programming before physical deployment.
  • P4119 Hands-Busy Queries: Comparing Voice and Virtual Keyboard Input for an LLM-Based AI Agent in AR Industrial Quality Inspection Martim Latas, Gabriel Marques, Patrícia Macedo, Pedro Albuquerque Santos, Rui Neves Madeira
    This paper reports a controlled laboratory user study of input modalities for a functional LLM-based AI agent integrated into an Augmented Reality (AR) quality-inspection prototype. The system combines spatial AR inspection guidance with a conversational assistant grounded through Retrieval-Augmented Generation. A within-subject study with 32 participants compared virtual keyboard and speech-to-text input across three inspection-related query tasks. Voice input was significantly faster, achieved higher usability scores, and was preferred for hands-busy scenarios. Results suggest that speech-to-text should serve as the primary modality for rapid operator queries, while virtual keyboard input remains useful as a precision fallback.
  • P4397 RoomAnchor: Object-Level Room Memory for Embodied XR Assistants Edward Scott Carvalho Johnson, Arthur Galdino Dangoni, Bianca Mirtes Araújo Miranda, Wilson Maranhão Ramos Filho, Pedro Engelberg Silva Borges, Diogo Fernandes, ARLINDO GALVÃO
    Embodied XR assistants need room knowledge that can be reused during dialogue and connected to spatial behavior. A room image can describe visible objects, but the assistant also needs a representation that links object descriptions to approximate locations so it can retrieve, answer, and point or move toward relevant objects later. We propose RoomAnchor, an object-level room memory pipeline for embodied XR assistants. RoomAnchor takes a single egocentric RGB room capture, produces object-level 3D hypotheses, labels detected crops, stores the resulting object records in a retrieval index, and uses an agent to answer room-related questions and select reference anchors for embodied behavior. We evaluated RoomAnchor on 24 room scenes and 1440 manually checked room-related question-answer trials. Across this evaluation, RoomAnchor reliably answered questions about mapped room content while usually retrieving the supporting mapped objects and selecting appropriate embodied reference anchors. Memory construction took 2.381 s on average. The results indicate that compact object-level room memory can support reliable room-question answering without continuous reconstruction or a device-specific depth pipeline.
  • P4516 GeminiPainter: Real-Time AI-Guided Robotic Sketch Generation for Live Portraiture Miguel Altamirano Cabrera, Aleksey Fedoseev, Iana Zhura, Dzmitry Tsetserukou
    We present an autonomous robotic portrait-generation system combining real-time face detection, AI-based sketch generation, and robotic drawing. The system captures video frames, extracts facial regions, converts them into minimalist single-line sketches using the Gemini Vision API, optimizes stroke order through graph-based path planning, and executes smooth trajectories on a 6-DoF collaborative manipulator. This perception-cognition-action pipeline integrates computer vision, neural artistic abstraction, motion optimization, and robot control. User ratings on a 5-point scale were high for sketch quality (4.33), perceived execution (4.53), and user experience (4.65), indicating recognizable, appealing, and engaging robotic portraits.
  • P4539 Augmented Reality Action Summaries for Situated Instructions of Repetitive Tasks Lucchas Ribeiro Skreinig, Ana Stanescu, Peter Mohr, Christoph Ebner, Markus Tatzgern, Dieter Schmalstieg, Denis Kalkofen
    Augmented reality (AR) instructions help people understand how to operate or assemble objects by overlaying visual guidance directly onto the physical environment. However, repeatedly presenting detailed instructions can negatively impact performance. We introduce an interactive AR instruction system that groups instructions based on recurring action patterns, providing detailed guidance on the first occurrence of actions and abstracting repeated sequences. We introduce a novel workflow that automatically organizes narrated expert demonstrations into instruction structures using a multimodal large language model, improving the scalability of AR guidance systems. Our user study shows that grouped instruction visualization is clearly preferred.
  • P4658 Layered Head-Intention Interaction for Low-Cost Mobile VR Museum Tours Lesi Hu, Hu Yong, Xukun Shen
    Low-cost mobile virtual reality (MVR) devices, such as Google Cardboard, support accessible virtual museum tours, but their limited input channels make interaction challenging. In such systems, head movement serves as both natural viewing and control input, which may cause intention ambiguity and false triggers. This paper proposes a Layered Head-Intention Interaction method for low-cost MVR museum tours. The method identifies five recurring interaction types and translates them into five head-intention layers: natural viewing, spatial movement, viewpoint adjustment, system operation, and exhibit information access. These layers are separated through spatial, angular, temporal, and object cues as design rules. We implemented the method in a Cardboard-based virtual museum prototype and conducted a preliminary user study, showing acceptable usability and feasibility.
  • P4737 A Speech-Driven AR Concept Board for CMF Exploration in Early Product Design Qing Ao, Suihuai Yu
    Early CMF exploration in product design often begins with informal design intent but becomes fragmented across generated images, comparison boards, and downstream 3D previews. This fragmentation makes it difficult to trace how color, material, finish, and perceived product qualities are interpreted, compared, selected, and transformed across media. This work presents a speech-driven AR concept-board workflow that structures speech-like CMF intent into a design brief, supports LLM-assisted shortlisting and designer-led board comparison, and links a selected candidate to 3D preview inspection. A portable-speaker walkthrough and preliminary designer feedback illustrate its potential for traceable CMF reasoning in early design.
  • P5020 VR-Based Sketching of Crochet-Inspired Openwork Structures for 3D Fabrication Bixun Chen, Rami Ghannam
    We present an early-stage VR-based system for embodied crochet-style designs through freeform line sketching in 3D space. Unlike existing crochet modeling approaches with surface tiling or pattern-based workflows, our method focuses on sparse, lace-like openwork structures characterized by branching lines and spatial voids. The system translates user-drawn strokes into procedural crochet arrangements and generates models suitable for FDM and SLA 3D printing. Our initial exploration suggests that VR enables intuitive spatial creation of decorative crochet forms, supporting rapid prototyping of open structures. We discuss fabrication limitations and highlight opportunities for embodied digital craft design in VR environments.
  • P5213 Toward WorldMCP: An Agent-First Approach for Spatial-Grounding of AR VLMs Emeric Marguet, Sebastian Hubenschmid, Niklas Elmqvist
    Current LLM-powered augmented reality systems rely on dedicated perception pipelines to bridge semantic understanding and real-world coordinates. However, recent advances in AI increasingly render this division obsolete, enabling scene analysis to be reasoned about entirely within the model itself --- given the right tools. We present WorldMCP, an approach that lets VLMs view a raw scene, spatially ground references within it, and then register and anchor virtual objects directly in an AR environment. A preliminary evaluation using frontier models confirms the feasibility of our approach, paving the way for general-purpose assistants that leverage unrestricted scene understanding through the use of tools.
  • P7030 Computationally Efficient Geometry-Based 3D Scene Understanding for Mobile Mixed Reality Andrzej Czajkowski, Marek Kowal, Krzysztof Patan, Rafael Greszczynski
    We present a lightweight geometry-based pipeline for segmentation of depth-sensor meshes on standalone mixed reality headsets. The method combines mesh simplification, normal-based plane extraction, concavity decomposition, and adaptive bounding-box refinement to obtain a compact representation of indoor scenes directly on resource-constrained mobile hardware. The approach was evaluated on indoor data acquired with the HTC Vive XR Elite. Runtime measurements show a total processing time of $4009 \pm 239$ ms for the evaluated mesh, while 3D Intersection over Union (IoU) is used to evaluate the influence of the segmentation threshold on bounding-box quality. The results show that a normal-similarity threshold of $0.99$ provides the best trade-off between spatial coverage and over-segmentation in the evaluated scene. The proposed pipeline provides a practical basis for asynchronous scene preprocessing, object replacement, collision handling, and spatial interaction in MR applications.
  • P7081 VME-as-Interaction: Designing Voluntary Muscle Effort for Embodied VR Experiences Zhengchun Jia, Jeffrey C. F. Ho
    We introduce the concept of Voluntary Muscle Effort (VME) and further proposes the design perspective of “VME-as-Interaction”, exploring whether users can construct an experience of effort solely through the active engagement of bodily resources in the absence of real physical resistance. We used surface electromyography (sEMG) as a technical medium for detecting active muscle activation, designed a VME-based VR interaction system for a construction safety training scenario, and evaluated it through a within-group experiment (N = 15). The results indicate that VME significantly enhanced immersion whilst increasing self-reported workload; however, participants generally perceived this additional effort as a genuine, natural and valuable experience. This study demonstrates that the experience of effort does not necessarily require actual physical force or extensive bodily exertion that reproduces real-world fatigue, but may emerge through users' active bodily engagement, bodily performance, and bodily imagination.
  • P7117 Reconfiguring Reach: Design Trade-offs of Extended Non-Human Avatar Embodiment in VR Spatial Manipulation Yunhao Yu
    Virtual reality enables avatars that depart from the human body, but it remains unclear how extended non-human limbs trade off reach, control, embodiment, workload, and manipulation behavior. We explore this design space through three deliberately contrasting avatar-control probes: a naturalistic hand, telescopic robotic arm, and octopus tentacle. In a within-subjects study (N=21), participants completed an obstacle-based sorting task while we logged trajectories, performance, VEQ, and NASA-TLX. Results show that reach extension reduces controller movement while expanding tip coverage, workload follows mapping complexity, and subjective embodiment remains comparable despite divergent behavior.
  • P8346 Coherent and Reversible AI-assisted XR Scene Authoring by Spatial Dependency Graphs Mine Dastan, Dongyu Qiu, Haiyan Jiang, Michele Fiorentino, Frank Guan
    The adoption of mixed extended technologies as authoring tools is increasing, as XR and AI-assisted systems simplify the creation of immersive content and narratives and make the process accessible to users without advanced technical expertise or the need for a traditional desktop interface. Existing tools have the ability to capture user intent and rapidly translate it into immersive creative outputs. However, their editing capabilities remain limited, as they often treat objects as isolated entities, failing to preserve semantic and contextual coherence across the generated scene. To preserve scene coherence during actions, we designed a novel Spatial Dependency Graph (SDG) structure for managing inter-object relationships (support and dependency) inferred from bounding-volumes and object-type semantics. SDG is traversed during action to preserve scene coherence through novel cascaded operations on dependent objects: (i) co-motion in case of movements, (ii) re-support find a compatible support in scene, and (iii) remove dependents when no coherent support exist and (iv) reverse which rollback to any prior scene states. We implemented SDG in a MR demonstrator integrating Unity3D and Meta Quest 3 with a text LLM-based command parser. A preliminary study with two participants achieved an adapted SUS score of 82.5/100, indicating promising initial usability and supporting broader dependency-aware XR authoring.
  • P8446 UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction Mikhail Kiselev, Aleksandr Marukhin, Ivan Snegirev, Elizaveta Semenyakina, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
    UltraArUco - a lightweight multilingual library and framework for low-latency, real-time marker-based tracking in mobile augmented reality. Unlike standard OpenCV-based implementations, UltraArUco introduces an optimized multilingual wrapper that reduces per-frame latency in four times, while maintaining high accuracy. Distributed Wi-Fi architecture provides portability, connects mobile device (camera input and detection) with a PC-based visual application, enabling responsive interactions. The framework is validated through an interactive piano simulation, where static ArUco markers on keys enable occlusion-based note triggering, and hand-mounted markers provide spatial gesture recognition. UltraArUco’s system requirements makes it perfect for resource-constrained mobile AR applications, demonstrating viable AR music application without specialized equipment.
  • P9198 Designing Embodied Data Storytelling for Cultural Heritage in Virtual Reality Yuxuan Guo
    Virtual reality (VR) has become a powerful medium for cultural heritage, yet many systems emphasize spatial reconstruction while offering limited support for interpreting the diverse evidence on which heritage narratives depend. We propose a mapping method for designing embodied data storytelling in VR, transforming heterogeneous heritage evidence according to its narrative role rather than its media type: temporal data become embodied narrative checkpoints, event data become interpretive data islands, and geographic data become situated scaled environments. We instantiate the method in a prototype based on the Silk Road and conduct a preliminary study with six participants. Results indicate that the approach connects interaction, spatial experience, and data interpretation, yet embodiment alone is insufficient, users still rely on explanatory anchors such as narration, captions, and mini-maps to understand what each action or spatial representation means.
  • P9219 AR-semble: Preliminary Field Evidence for Augmented Reality-Guided Robot Assembly in Middle School STEM Education Naomi Unkelos Shpigel, Bari Cohen, Shalev Cohen, Doaa Saad
    Assembling physical robots from two-dimensional printed instructions is a persistent challenge in STEM education: novice learners must mentally translate flat, static diagrams into three-dimensional spatial actions, a process that increases cognitive load, generates assembly errors, and creates high dependence on instructor support. We report preliminary field evidence from a classroom deployment of AR-semble, a mobile Augmented Reality application supporting robot assembly in middle school STEM education. AR-semble overlays step-by-step 3D animated guidance onto printed assembly manuals using marker-based tracking (Unity/Vuforia), enabling learner-controlled progression. Sixteen students completed a full six-step robot assembly session (M = 42:28 min), generating detailed navigation logs. A post-session survey (n = 8) assessed perceived usability, learning support, and engagement. Results indicate strong usability (easy to follow: M = 4.12/5; easy to learn: M = 4.00/5) and positive perceived usability and learning support (improved understanding: M = 3.75/5; learning efficiency: M = 3.75/5). Log analysis identifies pages 3 and 4 as navigationally challenging, informing targeted content redesign. These findings bridge the gap between design validation and a planned controlled comparative experiment.
  • P9626 Mapping the Design Space of XR Interruption Delivery: Initial Framework Aditya Sudesh Chandanshive, Jung Who Nam
    As XR integrates into everyday workflows, users face interruptions causing spatial cognitive disorientation. Mitigating this requires a structured understanding of interruption management. We introduce a framework to map this design space, categorizing interruptions by their source location (collocated vs. remote) and delivery timing (immediate vs. deferred). We illustrate this matrix using select examples to demonstrate its structure, discuss initial design strategies for mediating these disruptions, and provide a starting point for future exploration.
  • P9914 Efficiency and Usability of Static and Augmented Reality Head-Up Displays Margaréta Baumgärtner, Boglárka Sándor, Ábel Sulyok, Pál Maák
    Augmented Reality (AR) Head-Up Displays (HUDs) advance driver assistance by directly displaying critical information into the dynamic road environment. Current simulator study (N=56) evaluated driver performance by comparing a static HUD and an eye-tracked AR HUD along a route with sudden hazardous events, analyzing safety, navigation, cognitive workload and user experience. Although high data variance limited statistical significance, distinct descriptive trends emerged. The AR HUD improved navigation efficiency and hedonic acceptance without altering reaction times or cognitive load, though it reduced pragmatic clarity. Consequently, while safety remains equivalent, AR enhances navigation but demands careful interface design to preserve simplicity.

For questions, contact: posters2026@ieeeismar.net

ISMAR 2026 Poster Chairs: Mohammed Safayet Arefin, Andrea Boensch, Francesco Ferrise, Cassidy Nelson