Skip to content

About

Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned perception to its full-image policy, enabling fine-grained visual understanding in a single forward pass without external teachers, labels, or verifiers.

Resources

Stars

327 stars

Watchers

1 watching

Forks

Latest commit

 

History

19 Commits

Folders and files

Repository files navigation

[NeurIPS 2026] Vision-OPD

Vision-OPD: Learning to See Fine-Grained Details for Multimodal LLMs via On-Policy Self-Distillation

📃 Paper | 🤗 Hugging Face

Overview

Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged regional perception to its full-image policy, enabling fine-grained visual understanding in a single forward pass — without external teachers, ground-truth labels, reward verifiers, or inference-time tool use.

Vision-OPD Average Scores

Average scores across fine-grained visual understanding benchmarks, including V* Bench, ZoomBench, HR-Bench 4K, HR-Bench 8K, MME-RealWorld-EN and MME-RealWorld-CN.

Environment Setup

conda create -n vision-opd python=3.12 -y
conda activate vision-opd
pip install --upgrade pip
pip install --no-deps -r requirements.txt
pip install -e . --no-deps
pip install flash-attn --no-build-isolation
pip install causal-conv1d==1.6.1 --no-build-isolation

Quick Start

1. Deployment

Serve the merged checkpoint with vLLM, for example:

vllm serve yuanqianhao/Vision-OPD-9B \
    --gpu-memory-utilization 0.85 \
    --tensor-parallel-size 8 \
    --served-model-name Vision-OPD-9B \
    --trust-remote-code

The server listens on port 8000 by default. You can then query the model via the OpenAI-compatible API at http://localhost:8000/v1/chat/completions.

Please check out our Hugging Face Collection for all public Vision-OPD checkpoints.

2. Evaluation

Evaluate the deployed model on fine-grained visual benchmarks:

API_BASE="http://localhost:8000/v1/" \
OPENAI_MODEL_ID="Vision-OPD-9B" \
JUDGE_API_BASE="YOUR_JUDGE_API_BASE" \
JUDGE_MODEL="YOUR_JUDGE_MODEL_NAME" \
BENCHMARK="vstar,zoombench,hrbench-4k,hrbench-8k,mme-realworld,mme-realworld-cn" \
bash eval/run_eval.sh

Supported benchmarks: vstar, zoombench, hrbench-4k, hrbench-8k, mme-realworld, mme-realworld-cn, mme-realworld-lite, visualprobe, mmvp, cv-bench, mmstar, pope.

The evaluation script runs inference via the OpenAI-compatible API. Judge configuration can be set via JUDGE_API_BASE / JUDGE_MODEL_PATH environment variables. We use openai/gpt-oss-120b as the judge model. Other powerful models like Qwen3.5 or closed-source models are also recommended.

To evaluate Qwen3.5 models as baselines, set ENABLE_THINKING=False to run in non-thinking mode, for example:

API_BASE="http://localhost:8000/v1/" \
OPENAI_MODEL_ID="Qwen3.5-9B" \
ENABLE_THINKING=False \
JUDGE_API_BASE="YOUR_JUDGE_API_BASE" \
JUDGE_MODEL="YOUR_JUDGE_MODEL_NAME" \
BENCHMARK="vstar,zoombench,hrbench-4k,hrbench-8k,mme-realworld,mme-realworld-cn" \
bash eval/run_eval.sh

Training

1. Prepare Training Data

Download and preprocess the Vision-OPD-6K dataset:

python scripts/prepare_data.py --data-dir ./data

This downloads images and metadata from HuggingFace, extracts archives, and converts train.jsonl to the parquet format expected by the training pipeline.

2. Training

Launch Vision-OPD training:

bash scripts/run_vision_opd.sh

Key hyperparameters can be edited at the top of the script. See the script for the full configuration.

3. Merge Checkpoints

After training, merge the FSDP-sharded checkpoint into a standard HuggingFace model:

bash scripts/merge_checkpoint.sh <path_to_checkpoint>

For example:

bash scripts/merge_checkpoint.sh ./checkpoints/Vision-OPD-Qwen3.5-4B/global_step_65/

This merges the FSDP actor shards, saves the model weights, config, tokenizer, and processor into the specified directory. The merged checkpoint can then be loaded directly with transformers or served with vLLM.

Citation

If you find Vision-OPD useful for your research, please consider citing:

@article{yuan2026vision,
  title={Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation},
  author={Yuan, Qianhao and Lou, Jie and Yu, Xing and Lin, Hongyu and Sun, Le and Han, Xianpei and Lu, Yaojie},
  journal={arXiv preprint arXiv:2605.18740},
  year={2026}
}

License

Apache-2.0 License

About

Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned perception to its full-image policy, enabling fine-grained visual understanding in a single forward pass without external teachers, labels, or verifiers.

Resources

Stars

327 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages