Skip to content

Repository files navigation

Exploring Image Backbone Effects for EmbodiedScan Occupancy

This repository is a research fork of EmbodiedScan for the paper:

Exploring image backbone effects for EmbodiedScan semantic occupancy prediction

The code focuses on one controlled question: how much does the 2D image backbone affect RGB-D semantic occupancy prediction when the downstream occupancy pipeline is kept fixed? We keep the EmbodiedScan-style RGB-D projection, depth branch, fusion neck, occupancy head, losses, optimizer, and training schedule unchanged, and replace only the image encoder.

Contents

Project Scope

The original EmbodiedScan project provides a broad multi-modal 3D perception suite. This fork narrows the code path to semantic occupancy prediction and adds alternative image backbones:

  • CLIP-ResNet + FPN
  • CLIP-ViT + FPN-compatible outputs
  • BLIP2 visual backbone + FPN-compatible outputs
  • DINOv2 ViT + FPN-compatible outputs

All variants use the same occupancy setting:

  • RGB-D input from EmbodiedScan.
  • Voxel grid: 40 x 40 x 16.
  • Output channels: empty space plus 80 semantic classes.
  • 3D depth branch: MinkResNet-34.
  • Fusion and occupancy head: the same dense RGB-D occupancy pipeline used by the baseline configs.

Repository Layout

EmbodiedScan/
|-- configs/occupancy/
|   |-- mv-occ_8xb1_embodiedscan-occ-80class.py  # reproduced ResNet baseline
|   |-- mv-occ_clip.py                           # CLIP-ResNet
|   |-- mv-occ_clip_vit.py                       # CLIP-ViT
|   |-- mv-occ_blip2.py                          # BLIP2
|   `-- mv-occ_dinov2.py                         # DINOv2
|-- embodiedscan/models/backbones/
|   |-- mink_resnet.py
|   |-- clip_resnet.py
|   |-- clip_fpn.py
|   |-- bilp2_fpn.py
|   `-- dinov2_fpn.py
|-- tools/
|   |-- train.py
|   |-- test.py
|   `-- profile_flops.py
`-- run.sh

Environment

The code follows the original EmbodiedScan/OpenMMLab environment. The experiments were developed with:

  • Ubuntu 20.04
  • CUDA-capable NVIDIA GPU
  • Python 3.8
  • PyTorch 1.11.0
  • MMEngine / MMDetection3D stack used by EmbodiedScan
  • MinkowskiEngine
  • PyTorch3D
  • Transformers and OpenCLIP dependencies for the added image backbones

Install the base project from this directory:

cd EmbodiedScan
conda create -n embodiedscan python=3.8 -y
conda activate embodiedscan
conda install pytorch==1.11.0 torchvision==0.12.0 torchaudio==0.11.0 cudatoolkit=11.3 -c pytorch
python install.py all

If a backbone import fails, check that the extra pretrained-model packages are installed in the active environment. DINOv2 and BLIP2 rely on transformers; CLIP variants may require open_clip_torch depending on the checkpoint format.

Data and Weights

Prepare the EmbodiedScan data following the original data guide:

data/README.md

The occupancy configs currently use:

data_root = '/data/embodied_data/data/'

Change data_root in the config files if your dataset is stored elsewhere.

Expected local weights:

weights/resnet50-0676ba61.pth
weights/open_clip_pytorch_model.bin
weights/clip_vit_b16_openai.pt
weights/blip2/
weights/dinov2-large/

The repository does not include large pretrained weights or dataset files.

Backbone Variants

Variant Config Backbone class Notes
Reproduced baseline configs/occupancy/mv-occ_8xb1_embodiedscan-occ-80class.py mmdet.ResNet ResNet-50 + FPN baseline path
CLIP-ResNet configs/occupancy/mv-occ_clip.py CLIPResNet CLIP-pretrained ResNet-style hierarchy
CLIP-ViT configs/occupancy/mv-occ_clip_vit.py CLIPViTBackbone ViT-B/16 tokens adapted to FPN outputs
BLIP2 configs/occupancy/mv-occ_blip2.py BLIP2Backbone BLIP2 visual encoder adapted to FPN outputs
DINOv2 configs/occupancy/mv-occ_dinov2.py DINOv2Backbone DINOv2 visual encoder adapted to FPN outputs

The added backbones are registered in:

embodiedscan/models/backbones/__init__.py

Training and Evaluation

Run from the EmbodiedScan directory.

Single-GPU training:

python tools/train.py configs/occupancy/mv-occ_dinov2.py --work-dir=work_dirs/occ_dinov2

Distributed training:

python -m torch.distributed.launch \
  --nproc_per_node=4 \
  --master_port=29500 \
  tools/train.py \
  configs/occupancy/mv-occ_dinov2.py \
  --work-dir=work_dirs/occ_dinov2 \
  --launcher=pytorch

Evaluate a checkpoint:

python tools/test.py \
  configs/occupancy/mv-occ_dinov2.py \
  work_dirs/occ_dinov2/epoch_*.pth

Example commands for all variants are collected in:

run.sh

Results

Main mIoU comparison:

Method mIoU
EmbodiedScan reported baseline 19.97
EmbodiedScan reproduced baseline 21.20
DROcc(SwinU) 22.14
DROcc 23.38
CLIP-ResNet 17.41
CLIP-ViT 24.33
BLIP2 29.49
DINOv2 30.55

Resource comparison:

Method Train time Train memory Test time / instance Test memory Parameters
EmbodiedScan baseline 38.03 h 19.86 GiB 2.65 s 7.15 GiB 751.28 M
DROcc 33.83 h 6.74 GiB 2.59 s 2.73 GiB 114.38 M
CLIP-ResNet 40.06 h 19.18 GiB 2.82 s 7.39 GiB 751.28 M
CLIP-ViT 35.98 h 19.73 GiB 2.70 s 7.83 GiB 787.98 M
BLIP2 49.90 h 21.71 GiB 3.75 s 13.52 GiB 1718.77 M
DINOv2 39.02 h 18.84 GiB 3.11 s 10.63 GiB 1034.62 M

The main observation is that changing the image backbone can improve occupancy accuracy more than several occupancy-specific architectural changes on the same benchmark. DINOv2 gives the best mIoU, BLIP2 is close but more expensive, CLIP-ViT is a moderate improvement, and CLIP-ResNet is a negative case showing that image pretraining alone does not guarantee better dense occupancy features.

Citation and License

This repository is derived from EmbodiedScan. If you use this code, cite the original EmbodiedScan work and any pretrained backbone used in your experiments.

@inproceedings{embodiedscan,
    title={EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI},
    author={Wang, Tai and Mao, Xiaohan and Zhu, Chenming and Xu, Runsen and Lyu, Ruiyuan and Li, Peisen and Chen, Xiao and Zhang, Wenwei and Chen, Kai and Xue, Tianfan and Liu, Xihui and Lu, Cewu and Lin, Dahua and Pang, Jiangmiao},
    year={2024},
    booktitle={IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}
}

Follow the original EmbodiedScan license and the licenses of all pretrained model weights, datasets, and third-party dependencies.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages