[NeurIPS 2025] LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
δΈζη README | English README
- NVIDIA GPU with CUDA support (RTX series recommended, Isaac Sim does not support A100/A800)
- Ubuntu 24.04 (tested version)
- conda
- Python 3.11
- Isaac Sim 5.1
Download the code and pull scene assets:
git clone https://github.com/Rui-li023/LabUtopia.git
sudo apt install git-lfs
git lfs pullCreate and activate a new conda environment:
conda create -n labutopia python=3.11 -y
conda activate labutopiaInstall required packages:
# Install PyTorch
pip install torch==2.9.0 torchvision==0.24.0 torchaudio==2.9.0 --index-url https://download.pytorch.org/whl/cu126
# Install Isaac Sim
pip install isaacsim[all,extscache]==5.1.0 --extra-index-url https://pypi.nvidia.com
# Install other dependencies
pip install -r requirements.txt
# Run script to setup .vscode/settings.json
python -m isaacsim --generate-vscode-settingsLabSim/
βββ assets/ # Resource files directory
β βββ chemistry_lab/ # Chemistry lab scene resources
β βββ fetch/ # Fetch robot resources
β βββ navigation/ # Navigation task resources
β βββ robots/ # Robot model resources
βββ config/ # Configuration files directory
β βββ level1_*.yaml # Level 1 basic task configs
β βββ level2_*.yaml # Level 2 combined task configs
β βββ level3_*.yaml # Level 3 generalization task configs
β βββ level4_*.yaml # Level 4 long sequence task configs
β βββ level5_*.yaml # Level 5 navigation task configs
β βββ navigation/ # Navigation specific configs
βββ controllers/ # Controller implementations
β βββ atomic_actions/ # Basic action controllers
β βββ inference_engines/ # Inference engine implementations
β βββ robot_controllers/ # Robot controllers
βββ data_collectors/ # Data collector implementations
βββ factories/ # Factory class implementations
βββ policy/ # Policy model implementations
βββ tasks/ # Task definition implementations
βββ tests/ # Test code
βββ utils/ # Utility functions
βββ scripts/ # Data processing scripts
βββ generate_pointclouds.py # Point cloud generation
βββ merge_dataset.py # Dataset merging
βββ convert_labsim_data_to_lerobot.py # Format conversion
- Modular structure for better code organization and maintainability
- Scene state and observation data acquisition (including camera images, robot states, and scene object states) are handled in tasks
- Robot control and task success condition checking are handled in controllers
This section provides a comprehensive guide for navigation tasks, including data collection, point cloud generation, and obstacle map generation.
Data collection is the foundation for training navigation models. LabUtopia supports comprehensive data collection with Parquet format storage.
Navigation task configurations are located in config/ directory with the prefix level5_Navigation_. Key configuration files include:
Recommended Configuration:
level5_Navigation_parquet.yaml- Full-featured Parquet format data collectionlevel5_Navigation_smooth.yaml- Smooth trajectory collectionlevel5_Navigation_velocity_demo.yaml- Velocity-based control demo
Key Configuration Parameters:
# Basic configuration
name: level5_Navigation_parquet # Task name
task_type: "navigation_task_test_weizi" # Task type
controller_type: "navigation_parquet" # Controller type
mode: "collect" # Collection mode
# Scene configuration
usd_path: "assets/navigation_lab/navigation_lab_01/lab.usd" # Scene file path
# Task parameters
task:
max_steps: 4000 # Maximum steps per episode
navigation_config_path: "config/navigation/navigation_assets_1_4.yaml"
# Navigation goal pairs (start β end)
goal_pairs:
- start: [2.2, 7.11, 0.0] # Start position [x, y, yaw]
end: [5.5, 6.26, 0.0] # Goal position [x, y, yaw]
- start: [5.5, 6.26, 0.0]
end: [6.485, 4.5, 4.71238898038469]
# Motion constraints
max_linear_speed: 0.6 # Maximum linear speed (m/s)
max_angular_speed: 1.5 # Maximum angular speed (rad/s)
position_threshold: 0.08 # Position success threshold (m)
angle_threshold: 0.1 # Angle success threshold (rad)
# Camera configuration
cameras:
- prim_path: "/World/Ridgebase/base_link/Camera_02"
name: "observation"
resolution: [256, 256] # Image resolution
image_type: "rgb+depth" # Collect both RGB and depth
clipping_range: [0.1, 10.0] # Near/far clipping planes
# Robot configuration
robot:
type: "ridgebase" # Robot type
position: [0.0, 0.0, 0.0] # Initial robot position
# Dataset configuration (Parquet format)
dataset:
type: "parquet_format" # Use Parquet format
chunk_size: 1000 # Episodes per chunk
# Video and image saving
save_videos: true # Save MP4 videos
save_images: true # Save single frame images
video:
rgb_codec: "mp4v"
rgb_fps: 30
depth_codec: "mp4v"
depth_fps: 30
depth_colormap: "jet" # jet/viridis/gray
image:
rgb_format: "jpg" # jpg or png
rgb_quality: 95 # JPEG quality (1-100)
depth_format: "png" # PNG format for depth
depth_normalize: true # Normalize depth to 0-255
depth_colormap: "jet" # Depth visualization colormap
# Data collector configuration
collector:
type: "parquet_format" # Parquet format collector
# Collection limits
max_episodes: 100 # Maximum episodes to collectCreate or modify navigation assets configuration in config/navigation/:
# config/navigation/navigation_assets_custom.yaml
assets:
- name: "lab_3"
barrier_image_path: "assets/navigation_lab/navigation_lab_01/barrier.png"
scene_asset_path: "assets/navigation_lab/navigation_lab_01/lab.usd"
scene_scale: [1.0, 1.0, 1.0]
offset_radius: 0.6 # Obstacle offset radius
x_bounds: [0, 9] # Navigation area X bounds
y_bounds: [0, 9] # Navigation area Y boundsBasic Usage:
# Activate conda environment
conda activate labutopia
# Run data collection with default configuration
python main.py --config-name level5_Navigation_parquet
# Run in headless mode (no GUI)
python main.py --config-name level5_Navigation_parquet --headless
# Run without video generation (faster)
python main.py --config-name level5_Navigation_parquet --no-videoAdvanced Usage:
# Custom configuration file
python main.py --config-name level5_Navigation_parquet \
--config-dir config
# Override specific parameters
python main.py --config-name level5_Navigation_parquet \
max_episodes=50 \
task.max_steps=2000Collected data will be organized as follows:
outputs/collect/
βββ 2026.04.27/
βββ 22.30.25_level5_Navigation_parquet/
βββ config.yaml # Configuration snapshot
βββ trajectory_000000/ # Episode 0
β βββ data/
β β βββ chunk-000/
β β βββ episode_000000.parquet # Episode data
β β βββ episode_000000_waypoints.json # Waypoints
β β βββ episode_000000.ply # Point cloud (if generated)
β β βββ rgb/ # RGB images
β β βββ depth/ # Depth images
β βββ videos/
β βββ rgb.mp4 # RGB video
β βββ depth.mp4 # Depth video
βββ trajectory_000001/ # Episode 1
βββ ...
Data Format:
- Parquet files: Contain structured data including observations, actions, robot states
- Waypoints JSON: Navigation waypoints and metadata
- Images/Video: Visual observations from cameras
- PLY files: 3D point cloud scenes (if generated)
Generate colored 3D point clouds from USD scene files with trajectory visualization.
The generate_pointclouds.py script converts USD scene files to PLY format point clouds with:
- Scene geometry: Colored point cloud from USD models
- Planned trajectory: Waypoint path visualization
- Actual trajectory: Robot's executed path from collected data
- Floor ESDF: 2D Euclidean Signed Distance Field for obstacles
Batch Mode (Process Entire Collection):
# Activate environment
conda activate labutopia
# Generate point clouds for all episodes in a collection
python scripts/generate_pointclouds.py \
--config_path /path/to/collection/config.yaml \
--base_url /path/to/LabUtopia \
--overwriteExample:
python scripts/generate_pointclouds.py \
--config_path outputs/collect/2026.04.27/22.30.25_level5_Navigation_parquet/config.yaml \
--base_url /home/pjlab/fbh/LabUtopia \
--overwriteSingle USD Mode:
python scripts/generate_pointclouds.py \
--usd_file /path/to/scene.usd \
--output_ply /path/to/output.ply \
--parquet_path /path/to/episode.parquet \
--waypoints_json /path/to/waypoints.jsonTrajectory Visualization:
# Show both planned and actual trajectories (default)
python scripts/generate_pointclouds.py \
--config_path outputs/collect/xxx/config.yaml \
--base_url /path/to/LabUtopia \
--show_planned_trajectory \
--show_actual_trajectory \
--path_height 1.0 # Trajectory height in metersPoint Cloud Resolution:
# High resolution (smaller voxel size)
python scripts/generate_pointclouds.py \
--config_path outputs/collect/xxx/config.yaml \
--base_url /path/to/LabUtopia \
--resolution 0.01 # 1cm resolution
# Low resolution (larger voxel size, faster)
python scripts/generate_pointclouds.py \
--config_path outputs/collect/xxx/config.yaml \
--base_url /path/to/LabUtopia \
--resolution 0.05 # 5cm resolutionColormap Selection:
# Use different colormaps for depth visualization
python scripts/generate_pointclouds.py \
--config_path outputs/collect/xxx/config.yaml \
--base_url /path/to/LabUtopia \
--colormap jet # Options: jet, viridis, grayGenerate 2D Euclidean Signed Distance Field maps for obstacle detection:
# Generate point clouds with floor ESDF
python scripts/generate_pointclouds.py \
--config_path outputs/collect/xxx/config.yaml \
--base_url /path/to/LabUtopia \
--floor_esdf_2d \
--floor_z 0.0 # Floor Z height
--floor_eps 0.02 # Floor tolerance
--obstacle_height 0.1 # Obstacle height threshold
--max_dist 1.0 # Maximum ESDF distanceESDF Parameters:
--floor_z: Z coordinate of the floor plane (default: 0.0)--floor_eps: Tolerance for floor point detection (default: 0.02m)--obstacle_height: Minimum height for obstacle detection (default: 0.1m)--max_dist: Maximum distance for ESDF computation (default: 1.0m)
Generated PLY Files:
trajectory_000000/data/chunk-000/
βββ episode_000000.parquet # Original episode data
βββ episode_000000_waypoints.json # Waypoint data
βββ episode_000000.ply # Generated point cloud
βββ ...
PLY File Contents:
- Scene points: Colored 3D points from USD geometry (~300K points)
- Planned trajectory: Black colored points showing waypoint path
- Actual trajectory: Black colored points showing executed robot path
- ESDF overlay: Distance-based coloring if floor_esdf_2d enabled
Point Cloud Statistics:
- Average points per episode: ~300,000
- File size: ~5-10 MB per episode
- Format: PLY (Polygon File Format)
- Visualization: CloudCompare, MeshLab, or custom tools
LabUtopia supports automatic generation of obstacle maps through the point cloud pipeline.
ESDF maps provide distance information to obstacles for path planning and collision avoidance.
Generation Methods:
Method 1: During Point Cloud Generation
python scripts/generate_pointclouds.py \
--config_path outputs/collect/xxx/config.yaml \
--base_url /path/to/LabUtopia \
--floor_esdf_2d \
--floor_z 0.0 \
--obstacle_height 0.1 \
--max_dist 2.0 # 2m ESDF rangeMethod 2: Using Existing Point Clouds
# Process existing PLY files to extract ESDF
python scripts/generate_pointclouds.py \
--usd_file /path/to/scene.usd \
--output_ply /path/to/output_with_esdf.ply \
--floor_esdf_2d \
--floor_z 0.0 \
--floor_eps 0.02 \
--obstacle_height 0.15 \
--max_dist 1.5| Parameter | Description | Default | Recommended Range |
|---|---|---|---|
--floor_z |
Floor plane Z coordinate | 0.0 | -0.1 to 0.1 |
--floor_eps |
Floor point detection tolerance | 0.02 | 0.01 to 0.05 |
--obstacle_height |
Minimum obstacle height | 0.1 | 0.05 to 0.3 |
--max_dist |
Maximum ESDF distance | 1.0 | 0.5 to 5.0 |
Parameter Tuning Guidelines:
- Floor epsilon: Larger values include more points as floor, smaller values are more precise
- Obstacle height: Should be above small floor irregularities but below actual obstacles
- Max distance: Larger values provide more navigation information but increase computation time
Load and Visualize:
import numpy as np
import open3d as o3d
# Load PLY file with ESDF
pcd = o3d.io.read_point_cloud("episode_000000.ply")
# Extract ESDF information
points = np.asarray(pcd.points)
colors = np.asarray(pcd.colors)
# ESDF values are encoded in colors
# You can process them for path planningPath Planning Integration:
from your_planner import AStarPlanner
# Load ESDF map
esdf_map = load_esdf_from_ply("episode_000000.ply")
# Create planner
planner = AStarPlanner(esdf_map)
# Plan path
start = (2.2, 7.11)
goal = (5.5, 6.26)
path = planner.plan(start, goal)For custom obstacle map generation, you can extend the pipeline:
# scripts/generate_custom_obstacle_map.py
from utils.usd_to_pointcloud import convert_usd_to_colored_ply_with_trajectory
# Generate custom obstacle map
convert_usd_to_colored_ply_with_trajectory(
usd_file="scene.usd",
output_ply="custom_map.ply",
resolution=0.02,
colormap="viridis",
floor_esdf_2d=True,
floor_z=0.0,
floor_eps=0.015,
obstacle_height=0.12,
max_dist=2.5,
)Here's a complete example from data collection to map generation:
# Step 1: Collect navigation data
conda activate labutopia
python main.py --config-name level5_Navigation_parquet --headless
# Step 2: Generate point clouds with ESDF
python scripts/generate_pointclouds.py \
--config_path outputs/collect/2026.04.27/22.30.25_level5_Navigation_parquet/config.yaml \
--base_url /home/pjlab/fbh/LabUtopia \
--floor_esdf_2d \
--resolution 0.025 \
--colormap jet \
--overwrite
# Step 3: Verify generated data
ls -lh outputs/collect/2026.04.27/22.30.25_level5_Navigation_parquet/*/data/chunk-000/*.ply
# Step 4: Visualize point clouds (optional)
# Using CloudCompare, MeshLab, or Python:
python -c "
import open3d as o3d
pcd = o3d.io.read_point_cloud('outputs/collect/2026.04.27/22.30.25_level5_Navigation_parquet/trajectory_000000/data/chunk-000/episode_000000.ply')
o3d.visualization.draw_geometries([pcd])
"Issue 1: USD File Not Found
FileNotFoundError: USD file not found: /path/to/scene.usd
Solution: Update usd_path in config.yaml to point to the correct USD file location.
Issue 2: Low Point Cloud Quality
Solution: Adjust --resolution parameter. Smaller values (0.01-0.02) give higher quality but larger files.
Issue 3: Missing Trajectories in Point Cloud Solution: Ensure parquet files and waypoints JSON files exist in the trajectory directories.
Issue 4: ESDF Map Generation Fails Solution:
- Check
--floor_zparameter matches actual floor height - Adjust
--floor_epsif floor detection is incorrect - Verify
--obstacle_heightis appropriate for your scene
Issue 5: Memory Errors During Processing Solution:
- Reduce
--resolutionto decrease point count - Process trajectories individually instead of batch mode
- Close other applications to free up memory
The training process uses collected data to train robot policy models.
There are multiple training configurations in the policy/config/ folder:
train_diffusion_unet_image_workspace.yaml- Diffusion model training (recommended)train_act_image_workspace.yaml- ACT model training
Main parameters that need to be adjusted:
# Model configuration
policy:
_target_: policy.policy.diffusion_unet_image_policy.DiffusionUnetImagePolicy
shape_meta: ${shape_meta} # Data shape metadata
# Noise scheduler configuration
noise_scheduler:
num_train_timesteps: 100 # Training timesteps
beta_start: 0.0001 # Beta start value
beta_end: 0.02 # Beta end value
beta_schedule: squaredcos_cap_v2 # Beta schedule strategy
# Observation encoder configuration
obs_encoder:
_target_: policy.model.vision.multi_image_obs_encoder.MultiImageObsEncoder
rgb_model:
_target_: policy.model.vision.model_getter.get_resnet
name: resnet18 # Backbone network
resize_shape: [256, 256] # Resize shape
random_crop: False # Random crop
# Training parameters
training:
device: "cuda:0" # Training device
seed: 42 # Random seed
num_epochs: 8000 # Training epochs
lr: 1.0e-4 # Learning rate
batch_size: 64 # Batch size
gradient_accumulate_every: 1 # Gradient accumulation steps
# Checkpoint saving
checkpoint_every: 30 # Save every 30 epochs
val_every: 10 # Validate every 10 epochs
# Data loader configuration
dataloader:
batch_size: 64 # Batch size
num_workers: 4 # Number of workers
shuffle: True # Whether to shuffle data
# Optimizer configuration
optimizer:
_target_: torch.optim.AdamW
lr: 1.0e-4 # Learning rate
betas: [0.95, 0.999] # Adam parameters
weight_decay: 1.0e-6 # Weight decayModify the corresponding configuration file in the policy/config/task folder, change the dataset_path parameter to your dataset folder location.
# Use diffusion model training
python train.py --config-name=train_diffusion_unet_image_workspace
# Use ACT model training
python train.py --config-name=train_act_image_workspaceTraining logs and models will be saved in the outputs/train/date/time_modelname_taskname/ directory.
Use trained models for inference testing.
Change the mode from collect to infer in the configuration file and add inference-related configurations:
# Basic configuration
mode: "infer" # Change to inference mode
# Inference configuration
infer:
obs_names: {"camera_1_rgb": 'camera_1_rgb', "camera_2_rgb": 'camera_2_rgb'}
# Local inference configuration
policy_model_path: "outputs/train/2025.03.25/12.43.59_train_act_image_pick_pick_data/checkpoints/latest.ckpt"
policy_config_path: "outputs/train/2025.03.25/12.43.59_train_act_image_pick_pick_data/.hydra/config.yaml"
normalizer_path: "outputs/train/2025.03.25/12.43.59_train_act_image_pick_pick_data/checkpoints/normalize.ckpt"
# Remote inference configuration (optional)
type: "remote" # Use remote inference
host: "101.126.156.90" # Server address
port: 56434 # Server port
n_obs_steps: 1 # Observation steps
timeout: 30 # Timeout
max_retries: 3 # Maximum retries
max_episodes: 50 # Inference episodes# Use local model inference
python main.py --config-name level1_pick
# Use remote inference
python main.py --config-name level3_PourLiquidInference results will be saved in the outputs/infer/date/time_taskname/ directory.
Download our modified OpenPI code:
git clone https://github.com/Rui-li023/openpi.git
Convert LabUtopia format data to LeRobot format dataset:
python scripts/convert_labsim_data_to_lerobot.py --data_dir outputs/collect/xxx/xxx/dataset --num_processes 8 --fps 60 --repo_name labutopia/level3-pick
LabUtopia supports using remote servers for model inference.
cd openpi/packages/openpi-client
pip install -e .
Configure the remote inference engine in your config file:
infer:
engine: remote # Use remote inference engine
host: "0.0.0.0" # OpenPI server host
port: 8080 # OpenPI server port (optional)
n_obs_steps: 3 # Observation stepsThe OpenPI client provides simplified WebSocket communication with the remote server:
- Initialize: The client automatically connects to the OpenPI server using WebSocket
- Inference: Sends observation data (images, poses) to the server and receives action predictions
- Data Format: Automatically handles image format conversion and pose data serialization
- Error Handling: Includes fallback mechanisms for failed predictions
The OpenPI server should return actions in one of these formats:
{"action": [action_array]}{"actions": [action_array]}- Any dictionary with a key containing "action"
We welcome contributions from the community! If you have any questions, suggestions, or ideas for improvements, please feel free to:
- Open an Issue: Report bugs, request features, or discuss ideas
- Submit a Pull Request: Contribute code improvements, documentation fixes, or new features
Before submitting a PR, please ensure your code follows the project's coding style and passes relevant tests.
Thank you to all contributors for supporting this project! π
@article{li2025labutopia,
author = {Li, Rui and Hu, Zixuan and Qu, Wenxi and Zhang, Jinouwen and Yin, Zhenfei and Zhang, Sha and Huang, Xuantuo and Wang, Hanqing and Wang, Tai and Pang, Jiangmiao and Ouyang, Wanli and Bai, Lei and Zuo, Wangmeng and Duan, Ling-Yu and Zhou, Dongzhan and Tang, Shixiang},
title = {LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents},
journal = {arXiv preprint arXiv:2505.22634},
year = {2025},
}This repository contains both source code and data assets:
-
Code
Released under the MIT License. -
Data Assets
Released under the CC BY-NC 4.0 License.
Free to use and modify for research and educational purposes only.