Skip to content

Discrepancies in the results for IN1K and Ade20K #6

Description

@kumuji

Dear @YanFangCS @speedinghzl @mc-lan ,

Thank you for sharing the code and checkpoints!

I have been trying to reproduce the frozen visual representation results from the paper, in particular the ImageNet-1K and ADE20K numbers for GenLIP L16 in the table 8.

Differently to your paper, where you use dinov2:

..., we adopt the frozen-backbone evaluation protocol from DINOv2

I used DINOv3 repository as they don't seem to change the protocol for ADE20K or ImageNet.

And since you write:

we use attentive probing on patch features for classification

But you don't explicitely state which exact evaluation protocol you use, I used attention probing from AIM.

In Table 8 of the paper, you report for GenLIP L/16:

  • ImageNet-1K attention probing: 83.9
  • ADE20K linear probing: 41.0 mIoU

Using the GenLIP-L16-NaViT checkpoint with 0.5 normalization, I achieved 42.94 mIoU on ADE20K.
With Dinov3 stats-based normalization, I got 40.7 mIoU.
So, the ADE20K results are close to or slightly higher than the reported number.

However, I can't get close results for IN1K.
Using the GenLIP-L16-NaViT checkpoint on the Dinov3 codebase with 0.5 normalization and an attention probe from the AIM, I can only get 82.8 on IN1K.
However, the result is lower than the reported 83.9.

To verify that the main issue is not my evaluation setup, I reproduced DINOv3 results in the same environment.

  • DINOv3 ADE20K: 54.92 vs. 54.9 reported in the dino paper.
  • DINOv3 ImageNet-1K: 87.1 linear probing and 86.4 attention probing using the AIM protocol. Unfortunately, these number are not in the original DINOv3 paper. However, since DINOv3 ViT-H+ gets 87.9 and ViT-7B gets 88.4, the numbers appear correct.

Could you please share your evaluation protocol for IN1K or provide guidance on how to reproduce the results?

Could you also share the code to reproduce the results for the LLaVA-Next finetunes?

Thank you for the great work and for releasing the project.

Activity

  1. YanFangCS commented on Jun 10, 2026

    @YanFangCS
    Owner

    Hi @kumuji,

    Thanks a lot for your interest in GenLIP and for sharing your detailed reproduction results.

    For the ImageNet-1K result reported in the paper, we used an adapted implementation of attentive probing based on AIM. For the ADE20K evaluation, we similarly made some modifications and adaptations based on the DINOv2 evaluation scripts.

    Unfortunately, due to subsequent changes in my working environment and personnel transitions from bytedance, the original experimental scripts and detailed logs are no longer fully available on my side. I am now working on rebuilding the evaluation pipeline, re-running the evaluation, and carefully verifying the results.

    Once this is completed, I will update this repository with the latest evaluation scripts and reproduced results, and I will also notify you here.

    Thanks again for your careful reproduction effort and for pointing this out.

  2. YanFangCS commented on Jul 10, 2026

    @YanFangCS
    Owner

    Hi @kumuji,

    We have rechecked the ImageNet-1K and ADE20K results recently.

    For ADE20K linear probing, our latest verification shows that the input should be normalized using mean = [127.5, 127.5, 127.5] and std = [127.5, 127.5, 127.5]. This is equivalent to using mean = [0.5, 0.5, 0.5] and std = [0.5, 0.5, 0.5] for inputs scaled to [0, 1]. With this preprocessing, we obtain the following results:

    Evaluation setting GenLIP-L16-NaViT GenLIP-So16-NaViT GenLIP-g16-NaViT
    ADE20K linear probing, 512 px (mIoU) 42.7 43.6 44.8

    For ImageNet-1K attentive probing, we found that the result reported in Table 8 of our technical report was obtained using the first-stage, fixed-resolution checkpoint (e.g., GenLIP-L16-224), rather than the final NaViT-adapted checkpoint. The complete top-1 accuracy results for both checkpoint variants are shown below. All evaluations were conducted at a resolution of 224 × 224.

    Checkpoint variant GenLIP-L16 GenLIP-So16 GenLIP-g16
    First-stage (-224) 83.8 84.3 85.2
    NaViT-adapted (-NaViT) 82.8 83.1 83.0

    The second-stage NaViT adaptation, which includes training on the lengthy caption corpus with high resolution images, results in a slight reduction in discriminative performance under this probing setup. Increasing evaluation resolution may help -navit models to get better results.

    We will correct the checkpoint attribution and update these results in our report. Thank you again for your careful reproduction and for helping us identify this discrepancy.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions