Skip to content
GuoHeyuPublic

About

[ICRA26] OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

Topics

Resources

Code of conduct

Contributing

Stars

16 stars

Watchers

1 watching

Forks

Repository files navigation

[ICRA26] OmniVLA:
Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

arXiv License

OmniVLA Teaser

OmniVLA is an novel omni-modality VLA model that integrates novel sensing modalities to enable beyond-RGB robotic perception and manipulation. The core approach is the sensor-masked image, a unified representation that overlays physically meaningful, spatially grounded masks onto the RGB images. These masks are derived from sensors including an infrared camera, a mmWave radar, and a microphone array. This image-native unification keeps sensor input close to RGB statistics to facilitate training, provides a uniform interface across sensor hardware, and enables data-efficient learning with lightweight per-sensor projectors. Building on this, we design a multimodal vision-language-action model architecture and train OmniVLA by extending an RGB-pretrained VLA backbone. We evaluate OmniVLA on challenging real-world tasks that require sensor-modality perception to guide the manipulation. It achieves an average task success rate of 84%, significantly outperforms both RGB-only and raw-sensor-input baseline models by 59% and 28% respectively, meanwhile showing higher learning efficiency and stronger generalization capability.


Demo Video


📑 Table of Contents


1. Installation

1.1 Install Lerobot

Download our source code, then enter the lerobot folder

cd lerobot

Create a virtual environment with Python 3.10 and activate it, e.g. with miniconda:

conda create -y -n lerobot python=3.10
conda activate lerobot

When using miniconda, install ffmpeg in your environment:

conda install ffmpeg=7.1.1 -c conda-forge

Install 🤗 LeRobot, motor control and all other packages:

pip install -e .

NOTE: If you encounter build errors, you may need to install additional dependencies (cmake, build-essential, and ffmpeg libs). On Linux, run: sudo apt-get install cmake build-essential python3-dev pkg-config libavformat-dev libavcodec-dev libavdevice-dev libavutil-dev libswscale-dev libswresample-dev libavfilter-dev pkg-config. For other systems, see: Compiling PyAV

To use huggingface for dataset and model storage, log in with

huggingface-cli login --token <Your_token> --add-to-git-credential
HF_USER=$(huggingface-cli whoami | head -n 1)
echo $HF_USER

To use Weights and Biases for experiment tracking, log in with

wandb login

(note: you will also need to enable WandB in the configuration. See below.)

If you encounter any problem when installing Lerobot, you can refer their official document.

1.2 Install Grounded-SAM-2

Enter the Grounded-SAM-2 folder

cd Grounded-SAM-2

Download the pretrained SAM 2 checkpoints:

cd checkpoints
bash download_ckpts.sh

Download the pretrained Grounding DINO checkpoints:

cd gdino_checkpoints
bash download_ckpts.sh

Install PyTorch environment first. We use cuda-12.1 in our environment to run this demo. Since we need the CUDA compilation environment to compile the Deformable Attention operator used in Grounding DINO, we need to check whether the CUDA environment variables have been set correctly (which you can refer to Grounding DINO Installation for more details). You can set the environment variable manually as follows if you want to build a local GPU environment for Grounding DINO to run Grounded SAM 2:

export CUDA_HOME=/path/to/cuda-12.1/

Install Segment Anything 2:

pip install -e .

Install Grounding DINO:

pip install --no-build-isolation -e grounding_dino

If you encounter any problem when installing Grounded-SAM-2, you can refer their official document.

2. Inference with Existing Model

2.1 Modify Configuration

If you want to infer with existing model, you need to change the configuration first. The configuration path for corresponding model is

lerobot/common/policies/<model name>/configuration_<model name>.py

2.2 Run Scripts

For inference on the local machine, we need to download the model first.

huggingface-cli download ${HF_USER}/<model name> \
--local-dir /home/${PC_USER}/.cache/huggingface/lerobot/${HF_USER}/<model path>

Then we fill the parameters for the python command below and run.

python lerobot/scripts/control_robot.py \
--robot.type=so101 \
--control.type=record \
--control.fps=<fps> \
--control.single_task=<task string> \
--control.repo_id=${HF_USER}/<store path> \
--control.tags='["tutorial"]' \
--control.warmup_time_s=0 \
--control.episode_time_s=<episode time> \
--control.reset_time_s=0 \
--control.num_episodes=1 \
--control.push_to_hub=false \
--control.policy.path=/home/${PC_USER}/.cache/huggingface/lerobot/${HF_USER}/<policy path> \
> <output path>

If local machine doesn't have powerful GPU, we support server-client based inference. On the server, fill the parameters for the python command below and run.

python lerobot/server/so101_eval_fast.py \
--robot.type=so101 \
--control.type=record \
--control.fps=<fps> \
--control.single_task=<task string>  \
--control.repo_id=${HF_USER}/<store path> \ 
--control.tags='["tutorial"]' \
--control.warmup_time_s=0 \
--control.episode_time_s=<episode time> \
--control.reset_time_s=0 \
--control.num_episodes=1 \
--control.push_to_hub=false \
--control.policy.path=/home/${PC_USER}/.cache/huggingface/lerobot/${HF_USER}/<policy path>

On the client connected to local robot arm, fill the parameters for the python command below and run.

python lerobot/server/so101_client.py \
--robot.type=so101 \
--control.type=record \
--control.fps=<fps> \
--control.single_task=<task string> \
--control.repo_id=${HF_USER}/<store path> \
--control.tags='["tutorial"]' \
--control.warmup_time_s=0 \
--control.episode_time_s=<episode time> \
--control.reset_time_s=0 \
--control.num_episodes=1 \
--control.push_to_hub=false \
--control.synthesize=<option> \
--control.crop=<option> \
--control.overlay=<option> \
--control.prompt=<segment prompt> \
> <output path>

If you want to replay the stored inference result, fill the parameters for the python command below and run.

python lerobot/lerobot/scripts/control_robot.py \
--robot.type=so101 \
--control.type=replay \
--control.fps=<fps> \
--control.repo_id=${HF_USER}/<store path> \
--control.episode=0

3. Fine-tuning with a Custom Robot Dataset

3.1 Data Collection

To collect data, we need to calibrate the robot first

python lerobot/scripts/control_robot.py \
--robot.type=so101 \
--robot.cameras="{}" \
--control.type=calibrate

Then we test the camera

python lerobot/common/robot_devices/cameras/opencv.py --images-dir outputs/images_from_opencv_cameras

After that, we teleoperate the robot arm for testing

python lerobot/scripts/control_robot.py \
--robot.type=so101 \
--robot.cameras='{}' \
--control.type=teleoperate

Then we reset the robot arm position and come to data collection after testing all setup.

python lerobot/scripts/control_robot.py \
--robot.type=so101 \
--robot.cameras='{}' \
--control.type=reset_position

To start with dataset collection, fill the parameters for the python command below and run.

python lerobot/scripts/control_robot.py \
--robot.type=so101 \
--control.type=record \
--control.fps=<fps> \
--control.single_task=<task string> \
--control.repo_id=${HF_USER}/<dataset name> \
--control.tags='["tutorial"]' \
--control.warmup_time_s=0 \
--control.episode_time_s=<collection time> \
--control.reset_time_s=0 \
--control.num_episodes=<collection number> \
--control.display_data=false \
--control.push_to_hub=false \
--control.re_thermal=true \
> <Output file name>

After collecting data from real robot by teleoperation, fill the parameters for the python command below and run to post-process the data.

python lerobot/scripts/control_robot.py \
--robot.type=so101 \
--control.type=record \
--control.fps=<fps> \
--control.single_task=<task string> \
--control.repo_id=${HF_USER}/<repo id> \
--control.tags='["tutorial"]' \
--control.warmup_time_s=0 \
--control.episode_time_s=<collection time> \
--control.reset_time_s=0 \
--control.num_episodes=<collection number> \
--control.display_data=false \
--control.push_to_hub=false \
--control.synthesize=<option> \
--control.crop=<option> \
--control.overlay=<option> \
--control.prompt=<segmentation prompt> \
> <Output file name>

Finally, we upload the collected dataset to huggingface.

huggingface-cli upload ${HF_USER}/<dataset name> \
../.cache/huggingface/lerobot/${HF_USER}/<dataset name> \
--repo-type dataset

3.2 Model Fine-tuning

To finetune the model, we need to first download the dataset from huggingface.

huggingface-cli download ${HF_USER}/<dataset name> \
--repo-type dataset \
--local-dir /home/${PC_USER}/.cache/huggingface/lerobot/${HF_USER}/<dataset name>

To fine-tune the policy on single GPU, fill the parameters for the python command below and run

nohup python lerobot/scripts/train.py \
--policy.path=<policy path> \
--dataset.repo_id=${HF_USER}/<dataset name> \
--batch_size=<batch size> \
--steps=<step size> \
--output_dir=outputs/train/<output dir> \
--job_name=<job name> \
--policy.device=cuda \
--wandb.enable=true \
> <Output file name> &

To use wandb for logging training and evaluation curves, make sure you've run wandb login as a one-time setup step. Then, when running the training command above, enable WandB in the configuration by adding --wandb.enable=true.

To fine-tune the policy on multiple GPU in distributed way, fill the parameters for the python command below and run

python lerobot/scripts/train_lightning.py \
--policy.path=<policy path> \
--dataset.repo_id=$${HF_USER}/<dataset name> \
--batch_size=<batch size> \
--output_dir=outputs/train/<output dir> \
--job_name=<job name> \
--policy.device=cuda \
--wandb.enable=true \
--gpus <gpu numnber> \
--nodes <node number> \
--steps <step size> \
--save_freq <save freq>

Finally, we upload the finetuned model to huggingface.

huggingface-cli upload ${HF_USER}/<model name>  outputs/train/<model path>

4. Other Details

4.1 Code Structure

.
├── featureMapCreator # contains code about creating and overlaying sensor images
├── Grounded-SAM-2 # contains code about image segmentation with prompt
├── lerobot
|   ├── configs          # contains config classes with all options that you can override in the command line
|   ├── common           # contains classes and utilities
|   |   ├── datasets       # various datasets of human demonstrations
|   |   ├── policies       # various policies: smolVLA, pi0
|   |   ├── robot_devices  # various real devices: motors, cameras, robots
|   |   └── utils          # various utilities
|   ├── server           # contains server-client to run model on high-end server
|   └── scripts          # contains functions to execute via command line
|       ├── eval.py                 # load policy and evaluate it on an environment
|       ├── train.py                # train a policy via imitation learning and/or reinforcement learning
|       ├── control_robot.py        # teleoperate a real robot, record data, run a policy
|       └── train_lightning.py      # train a policy in a distributed way
├── library # contains code about controling sensors for data collection
|   ├── acoustic          # acoustic sensor
|   ├── Camera            # camera
|   ├── depthCamera_SDK   # depth camera sensor
|   ├── mmWaveRadar       # mmWave radar sensor
|   ├── pressureRadar     # pressure radar sensor
|   └── thermalCamera     # thermal camera sensor
└── outputs               # contains results of scripts execution: logs, videos, model checkpoints

4.2 Hardware Information

RGB Camera: Hikvision 2K USB Camera

Thermal Camera: Xtherm II T2S+

Acoustic Array: Sipeed 6+1 Mic Array

Depth Camera: Orbbec 3D Camera Gemini Pro

mmWave Radar: Calterah 60GAIP4s4s-1, Infineon BGT60ATR24C

pressure Sensor: Leanstar MD30-60

4.3 The LeRobotDataset format

A dataset in LeRobotDataset format is very simple to use. It can be loaded from a repository on the Hugging Face hub or a local folder simply with e.g. dataset = LeRobotDataset("lerobot/aloha_static_coffee") and can be indexed into like any Hugging Face and PyTorch dataset. For instance dataset[0] will retrieve a single temporal frame from the dataset containing observation(s) and an action as PyTorch tensors ready to be fed to a model.

A specificity of LeRobotDataset is that, rather than retrieving a single frame by its index, we can retrieve several frames based on their temporal relationship with the indexed frame, by setting delta_timestamps to a list of relative times with respect to the indexed frame. For example, with delta_timestamps = {"observation.image": [-1, -0.5, -0.2, 0]} one can retrieve, for a given index, 4 frames: 3 "previous" frames 1 second, 0.5 seconds, and 0.2 seconds before the indexed frame, and the indexed frame itself (corresponding to the 0 entry). See example 1_load_lerobot_dataset.py for more details on delta_timestamps.

Under the hood, the LeRobotDataset format makes use of several ways to serialize data which can be useful to understand if you plan to work more closely with this format. We tried to make a flexible yet simple dataset format that would cover most type of features and specificities present in reinforcement learning and robotics, in simulation and in real-world, with a focus on cameras and robot states but easily extended to other types of sensory inputs as long as they can be represented by a tensor.

Here are the important details and internal structure organization of a typical LeRobotDataset instantiated with dataset = LeRobotDataset("lerobot/aloha_static_coffee"). The exact features will change from dataset to dataset but not the main aspects:

dataset attributes:
  ├ hf_dataset: a Hugging Face dataset (backed by Arrow/parquet). Typical features example:
  │  ├ observation.images.cam_high (VideoFrame):
  │  │   VideoFrame = {'path': path to a mp4 video, 'timestamp' (float32): timestamp in the video}
  │  ├ observation.state (list of float32): position of an arm joints (for instance)
  │  ... (more observations)
  │  ├ action (list of float32): goal position of an arm joints (for instance)
  │  ├ episode_index (int64): index of the episode for this sample
  │  ├ frame_index (int64): index of the frame for this sample in the episode ; starts at 0 for each episode
  │  ├ timestamp (float32): timestamp in the episode
  │  ├ next.done (bool): indicates the end of an episode ; True for the last frame in each episode
  │  └ index (int64): general index in the whole dataset
  ├ episode_data_index: contains 2 tensors with the start and end indices of each episode
  │  ├ from (1D int64 tensor): first frame index for each episode — shape (num episodes,) starts with 0
  │  └ to: (1D int64 tensor): last frame index for each episode — shape (num episodes,)
  ├ stats: a dictionary of statistics (max, mean, min, std) for each feature in the dataset, for instance
  │  ├ observation.images.cam_high: {'max': tensor with same number of dimensions (e.g. `(c, 1, 1)` for images, `(c,)` for states), etc.}
  │  ...
  ├ info: a dictionary of metadata on the dataset
  │  ├ codebase_version (str): this is to keep track of the codebase version the dataset was created with
  │  ├ fps (float): frame per second the dataset is recorded/synchronized to
  │  ├ video (bool): indicates if frames are encoded in mp4 video files to save space or stored as png files
  │  └ encoding (dict): if video, this documents the main options that were used with ffmpeg to encode the videos
  ├ videos_dir (Path): where the mp4 videos or png images are stored/accessed
  └ camera_keys (list of string): the keys to access camera features in the item returned by the dataset (e.g. `["observation.images.cam_high", ...]`)

A LeRobotDataset is serialised using several widespread file formats for each of its parts, namely:

  • hf_dataset stored using Hugging Face datasets library serialization to parquet
  • videos are stored in mp4 format to save space
  • metadata are stored in plain json/jsonl files

Dataset can be uploaded/downloaded from the HuggingFace hub seamlessly. To work on a local dataset, you can specify its location with the root argument if it's not in the default ~/.cache/huggingface/lerobot location.

Citation

If you find our work useful in your research, please cite:

@article{guo2025omnivla,
  title={OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation},
  author={Guo, Heyu and Wang, Shanmu and Ma, Ruichun and Jiang, Shiqi and Ghasempour, Yasaman and Abari, Omid and Guo, Baining and Qiu, Lili},
  journal={arXiv preprint arXiv:2511.01210},
  year={2025}
}

About

[ICRA26] OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

Topics

Resources

Code of conduct

Contributing

Stars

16 stars

Watchers

1 watching

Forks

Releases

Contributors

Languages