This repository is a research fork of EmbodiedScan for the paper:
Exploring image backbone effects for EmbodiedScan semantic occupancy prediction
The code focuses on one controlled question: how much does the 2D image backbone affect RGB-D semantic occupancy prediction when the downstream occupancy pipeline is kept fixed? We keep the EmbodiedScan-style RGB-D projection, depth branch, fusion neck, occupancy head, losses, optimizer, and training schedule unchanged, and replace only the image encoder.
The original EmbodiedScan project provides a broad multi-modal 3D perception suite. This fork narrows the code path to semantic occupancy prediction and adds alternative image backbones:
- CLIP-ResNet + FPN
- CLIP-ViT + FPN-compatible outputs
- BLIP2 visual backbone + FPN-compatible outputs
- DINOv2 ViT + FPN-compatible outputs
All variants use the same occupancy setting:
- RGB-D input from EmbodiedScan.
- Voxel grid:
40 x 40 x 16. - Output channels: empty space plus 80 semantic classes.
- 3D depth branch:
MinkResNet-34. - Fusion and occupancy head: the same dense RGB-D occupancy pipeline used by the baseline configs.
EmbodiedScan/
|-- configs/occupancy/
| |-- mv-occ_8xb1_embodiedscan-occ-80class.py # reproduced ResNet baseline
| |-- mv-occ_clip.py # CLIP-ResNet
| |-- mv-occ_clip_vit.py # CLIP-ViT
| |-- mv-occ_blip2.py # BLIP2
| `-- mv-occ_dinov2.py # DINOv2
|-- embodiedscan/models/backbones/
| |-- mink_resnet.py
| |-- clip_resnet.py
| |-- clip_fpn.py
| |-- bilp2_fpn.py
| `-- dinov2_fpn.py
|-- tools/
| |-- train.py
| |-- test.py
| `-- profile_flops.py
`-- run.sh
The code follows the original EmbodiedScan/OpenMMLab environment. The experiments were developed with:
- Ubuntu 20.04
- CUDA-capable NVIDIA GPU
- Python 3.8
- PyTorch 1.11.0
- MMEngine / MMDetection3D stack used by EmbodiedScan
- MinkowskiEngine
- PyTorch3D
- Transformers and OpenCLIP dependencies for the added image backbones
Install the base project from this directory:
cd EmbodiedScan
conda create -n embodiedscan python=3.8 -y
conda activate embodiedscan
conda install pytorch==1.11.0 torchvision==0.12.0 torchaudio==0.11.0 cudatoolkit=11.3 -c pytorch
python install.py allIf a backbone import fails, check that the extra pretrained-model packages are installed in the active environment. DINOv2 and BLIP2 rely on transformers; CLIP variants may require open_clip_torch depending on the checkpoint format.
Prepare the EmbodiedScan data following the original data guide:
data/README.md
The occupancy configs currently use:
data_root = '/data/embodied_data/data/'Change data_root in the config files if your dataset is stored elsewhere.
Expected local weights:
weights/resnet50-0676ba61.pth
weights/open_clip_pytorch_model.bin
weights/clip_vit_b16_openai.pt
weights/blip2/
weights/dinov2-large/
The repository does not include large pretrained weights or dataset files.
| Variant | Config | Backbone class | Notes |
|---|---|---|---|
| Reproduced baseline | configs/occupancy/mv-occ_8xb1_embodiedscan-occ-80class.py |
mmdet.ResNet |
ResNet-50 + FPN baseline path |
| CLIP-ResNet | configs/occupancy/mv-occ_clip.py |
CLIPResNet |
CLIP-pretrained ResNet-style hierarchy |
| CLIP-ViT | configs/occupancy/mv-occ_clip_vit.py |
CLIPViTBackbone |
ViT-B/16 tokens adapted to FPN outputs |
| BLIP2 | configs/occupancy/mv-occ_blip2.py |
BLIP2Backbone |
BLIP2 visual encoder adapted to FPN outputs |
| DINOv2 | configs/occupancy/mv-occ_dinov2.py |
DINOv2Backbone |
DINOv2 visual encoder adapted to FPN outputs |
The added backbones are registered in:
embodiedscan/models/backbones/__init__.py
Run from the EmbodiedScan directory.
Single-GPU training:
python tools/train.py configs/occupancy/mv-occ_dinov2.py --work-dir=work_dirs/occ_dinov2Distributed training:
python -m torch.distributed.launch \
--nproc_per_node=4 \
--master_port=29500 \
tools/train.py \
configs/occupancy/mv-occ_dinov2.py \
--work-dir=work_dirs/occ_dinov2 \
--launcher=pytorchEvaluate a checkpoint:
python tools/test.py \
configs/occupancy/mv-occ_dinov2.py \
work_dirs/occ_dinov2/epoch_*.pthExample commands for all variants are collected in:
run.sh
Main mIoU comparison:
| Method | mIoU |
|---|---|
| EmbodiedScan reported baseline | 19.97 |
| EmbodiedScan reproduced baseline | 21.20 |
| DROcc(SwinU) | 22.14 |
| DROcc | 23.38 |
| CLIP-ResNet | 17.41 |
| CLIP-ViT | 24.33 |
| BLIP2 | 29.49 |
| DINOv2 | 30.55 |
Resource comparison:
| Method | Train time | Train memory | Test time / instance | Test memory | Parameters |
|---|---|---|---|---|---|
| EmbodiedScan baseline | 38.03 h | 19.86 GiB | 2.65 s | 7.15 GiB | 751.28 M |
| DROcc | 33.83 h | 6.74 GiB | 2.59 s | 2.73 GiB | 114.38 M |
| CLIP-ResNet | 40.06 h | 19.18 GiB | 2.82 s | 7.39 GiB | 751.28 M |
| CLIP-ViT | 35.98 h | 19.73 GiB | 2.70 s | 7.83 GiB | 787.98 M |
| BLIP2 | 49.90 h | 21.71 GiB | 3.75 s | 13.52 GiB | 1718.77 M |
| DINOv2 | 39.02 h | 18.84 GiB | 3.11 s | 10.63 GiB | 1034.62 M |
The main observation is that changing the image backbone can improve occupancy accuracy more than several occupancy-specific architectural changes on the same benchmark. DINOv2 gives the best mIoU, BLIP2 is close but more expensive, CLIP-ViT is a moderate improvement, and CLIP-ResNet is a negative case showing that image pretraining alone does not guarantee better dense occupancy features.
This repository is derived from EmbodiedScan. If you use this code, cite the original EmbodiedScan work and any pretrained backbone used in your experiments.
@inproceedings{embodiedscan,
title={EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI},
author={Wang, Tai and Mao, Xiaohan and Zhu, Chenming and Xu, Runsen and Lyu, Ruiyuan and Li, Peisen and Chen, Xiao and Zhang, Wenwei and Chen, Kai and Xue, Tianfan and Liu, Xihui and Lu, Cewu and Lin, Dahua and Pang, Jiangmiao},
year={2024},
booktitle={IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}
}Follow the original EmbodiedScan license and the licenses of all pretrained model weights, datasets, and third-party dependencies.