This is the official code for our ICML 2026 paper, "From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models".
We introduce HATCH, a human-inspired training framework for multimodal large language models that improves multi-image spatial reasoning.
While multimodal large language models have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning remains challenging because it requires integrating information across multiple viewpoints. HATCH addresses this challenge with two human-aware training objectives: Patch-Level Spatial Alignment for cross-view correspondence, and Action-then-Answer Reasoning for explicit stepwise viewpoint transformation.
We apply HATCH to Qwen2.5-VL-3B and compare it with baseline training methods and existing models. Our experiments show that:
- HATCH significantly improves Qwen2.5-VL-3B and outperforms baseline training methods.
- HATCH outperforms larger open-weight models, including Qwen2.5-VL-72B.
- HATCH achieves performance closer to, or comparable with, proprietary models such as GPT-5.2.
- [2026.07.02] Initial release of HATCH training and SPAR-Bench evaluation code.
git clone git@github.com:stjohn2007/HATCH.git
cd HATCHSetup
cd prepare_data
uv venv --python 3.10
source .venv/bin/activate
uv pip install datasets pillow tqdmRun scripts
HATCH annotations are distributed through Hugging Face Datasets. Images derived
from ScanNet and ScanNet++ are not redistributed. Users must prepare the
original image assets locally. The preparation script downloads annotations
from stjohn2007/HATCH-10k.
# Edit local ScanNet / ScanNet++ image roots in prepare_hatch_jsonl.py.
bash run.shThis writes JSONL files under:
VLM-R1/data/hatch/
Setup
cd ../VLM-R1
uv venv --python 3.10
source .venv/bin/activate
# Modify the PyTorch and CUDA versions to match your environment if needed.
uv pip install torch==2.7.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
# If flash-attn installation fails, install a version compatible with your CUDA and PyTorch environment.
bash setup_uv.shRun scripts
source .venv/bin/activate
bash scripts/train.shThe full recipe runs:
bash scripts/pasta.sh
bash scripts/sft.sh
bash scripts/actor_sft.sh
bash scripts/actor_grpo.shWe tested pasta.sh, sft.sh, and actor_sft.sh on one H100 GPU, and
actor_grpo.sh on four H100 GPUs.
Checkpoints are written under:
VLM-R1/exp/{exp_name}/train/checkpoints
Setup
cd ../eval_scripts
uv venv --python 3.10
source .venv/bin/activate
# Modify the PyTorch and CUDA versions to match your environment if needed.
uv pip install torch==2.7.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
uv pip install -r requirements.txtRun Scripts
Download SPAR-Bench separately and place it under:
hf download jasonzhango/SPAR-Bench --repo-type dataset --local-dir datasets/SPAR-BenchThen run:
source .venv/bin/activate
python evaluate.py \
--model_name_or_path ../VLM-R1/exp/hatch_qwen2_5_vl_3b_actor_grpo/train/checkpoints/checkpoint-XXX \
--data_dir datasets/SPAR-Bench \
--output_dir results/hatch_actor_grpo_checkpoint_XXX \
--prompt_type action_all \
--batch_size 1 \
--gpu_ids 0For multi-GPU evaluation:
python evaluate.py \
--model_name_or_path ../VLM-R1/exp/hatch_qwen2_5_vl_3b_actor_grpo/train/checkpoints/checkpoint-XXX \
--data_dir datasets/SPAR-Bench \
--output_dir results/hatch_actor_grpo_checkpoint_XXX \
--prompt_type action_all \
--batch_size 1 \
--gpu_ids 0,1,2,3 \
--num_processes 4The model evaluated in the paper is available at HERE.
Training and evaluation scores may vary across runs due to stochastic sampling, GPU kernels, hardware differences, checkpoint selection, and local rendering or data preparation details.
Our implementation builds on VLM-R1 and SpatialLadder. We thank the authors for releasing their code.
HATCH annotations are derived in part from SPAR-7M. Please also follow the license and terms of the original datasets, including ScanNet and ScanNet++.
@inproceedings{oi2026hatch,
title={From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models},
author={Masanari Oi and Koki Maeda and Ryuto Koike and Daisuke Oba and Nakamasa Inoue and Naoaki Okazaki},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
}
