Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

HATCH

This is the official code for our ICML 2026 paper, "From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models".

💨 TL;DR

We introduce HATCH, a human-inspired training framework for multimodal large language models that improves multi-image spatial reasoning.

📖 Introduction

While multimodal large language models have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning remains challenging because it requires integrating information across multiple viewpoints. HATCH addresses this challenge with two human-aware training objectives: Patch-Level Spatial Alignment for cross-view correspondence, and Action-then-Answer Reasoning for explicit stepwise viewpoint transformation.

We apply HATCH to Qwen2.5-VL-3B and compare it with baseline training methods and existing models. Our experiments show that:

  • HATCH significantly improves Qwen2.5-VL-3B and outperforms baseline training methods.
  • HATCH outperforms larger open-weight models, including Qwen2.5-VL-72B.
  • HATCH achieves performance closer to, or comparable with, proprietary models such as GPT-5.2.

📝 Changelog

  • [2026.07.02] Initial release of HATCH training and SPAR-Bench evaluation code.

🚀 Quick Start

git clone git@github.com:stjohn2007/HATCH.git
cd HATCH

⚙️ Prepare Training Data

Setup

cd prepare_data
uv venv --python 3.10
source .venv/bin/activate
uv pip install datasets pillow tqdm

Run scripts

HATCH annotations are distributed through Hugging Face Datasets. Images derived from ScanNet and ScanNet++ are not redistributed. Users must prepare the original image assets locally. The preparation script downloads annotations from stjohn2007/HATCH-10k.

# Edit local ScanNet / ScanNet++ image roots in prepare_hatch_jsonl.py.
bash run.sh

This writes JSONL files under:

VLM-R1/data/hatch/

🧠 Training

Setup

cd ../VLM-R1

uv venv --python 3.10
source .venv/bin/activate

# Modify the PyTorch and CUDA versions to match your environment if needed.
uv pip install torch==2.7.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
# If flash-attn installation fails, install a version compatible with your CUDA and PyTorch environment.
bash setup_uv.sh

Run scripts

source .venv/bin/activate

bash scripts/train.sh

The full recipe runs:

bash scripts/pasta.sh
bash scripts/sft.sh
bash scripts/actor_sft.sh
bash scripts/actor_grpo.sh

We tested pasta.sh, sft.sh, and actor_sft.sh on one H100 GPU, and actor_grpo.sh on four H100 GPUs.

Checkpoints are written under:

VLM-R1/exp/{exp_name}/train/checkpoints

📊 Evaluation

Setup

cd ../eval_scripts
uv venv --python 3.10
source .venv/bin/activate

# Modify the PyTorch and CUDA versions to match your environment if needed.
uv pip install torch==2.7.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
uv pip install -r requirements.txt

Run Scripts

Download SPAR-Bench separately and place it under:

hf download jasonzhango/SPAR-Bench --repo-type dataset --local-dir datasets/SPAR-Bench

Then run:

source .venv/bin/activate

python evaluate.py \
  --model_name_or_path ../VLM-R1/exp/hatch_qwen2_5_vl_3b_actor_grpo/train/checkpoints/checkpoint-XXX \
  --data_dir datasets/SPAR-Bench \
  --output_dir results/hatch_actor_grpo_checkpoint_XXX \
  --prompt_type action_all \
  --batch_size 1 \
  --gpu_ids 0

For multi-GPU evaluation:

python evaluate.py \
  --model_name_or_path ../VLM-R1/exp/hatch_qwen2_5_vl_3b_actor_grpo/train/checkpoints/checkpoint-XXX \
  --data_dir datasets/SPAR-Bench \
  --output_dir results/hatch_actor_grpo_checkpoint_XXX \
  --prompt_type action_all \
  --batch_size 1 \
  --gpu_ids 0,1,2,3 \
  --num_processes 4

🤖 Model

The model evaluated in the paper is available at HERE.

⚠️ Note on Reproducibility

Training and evaluation scores may vary across runs due to stochastic sampling, GPU kernels, hardware differences, checkpoint selection, and local rendering or data preparation details.

🙏 Acknowledgement

Our implementation builds on VLM-R1 and SpatialLadder. We thank the authors for releasing their code.

HATCH annotations are derived in part from SPAR-7M. Please also follow the license and terms of the original datasets, including ScanNet and ScanNet++.

📚 Citation

@inproceedings{oi2026hatch,
    title={From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models},
    author={Masanari Oi and Koki Maeda and Ryuto Koike and Daisuke Oba and Nakamasa Inoue and Naoaki Okazaki},
    booktitle={Forty-third International Conference on Machine Learning},
    year={2026},
}

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages