User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interaction space, multi-modal feedback, and high resolution & frame-rate contents. To stimulate the reconstruction of immersive volumetric videos, we introduce ImViD, a multi-view, multi-modal dataset featuring complete space-oriented data capture and various indoor/outdoor scenarios.
Our capture rig supports multi-view video-audio capture while on the move, a capability absent in existing datasets, significantly enhancing the completeness, flexibility, and efficiency of data capture.
@misc{yang2025imvidimmersivevolumetricvideos,
title={ImViD: Immersive Volumetric Videos for Enhanced VR Engagement},
author={Zhengxian Yang and Shi Pan and Shengqi Wang and Haoxiang Wang and Li Lin and Guanjun Li and Zhengqi Wen and Borong Lin and Jianhua Tao and Tao Yu},
year={2025},
eprint={2503.14359},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.14359},
}