Skip to content

Question about potential inconsistency in latent_action selection between training and evaluation #54

Description

@shiyouheng

Hello,

I've been looking through the codebase and noticed a difference in how latent_action is extracted during training versus evaluation. I wanted to ask about this to better understand the implementation, as it might be a deliberate design choice that I'm not fully grasping.

In vla-scripts/finetune_libero.py, during training, latent_action_tokens are selected using a mask based on action_gt > 32000:

def action_decoder_forward(self, batch, vla_output):
    # ... (other code)
    action_gt = batch["labels"].to(latent_tokens.device)
    mask = action_gt > 32000

    latent_action_tokens = []
    for idx, per_sample_latent_tokens in enumerate(latent_tokens):
        per_sample_latent_action_tokens = per_sample_latent_tokens[mask[idx], :]
        latent_action_tokens.append(per_sample_latent_action_tokens)
    # ... (uses latent_action_tokens for prediction)

However, in the evaluation code (experiments/robot/libero/run_libero_eval.py), latent_action is selected using the positions of the last four tokens, which is equivalent to using generated_ids > 32000 to create a mask for extracting latent_action:

def predict_latent_action(...):
    # ... (other code)
    latent_action = latent_tokens[:, -4:]
    # ...

Since the LlamaForCausalLMbackbone uses next token prediction, I'm wondering if these two approaches might be extracting latent_action from different positions, creating a potential mismatch between training and evaluation.

I also noticed another section of code that accounts for the next token prediction behavior by using the mask with action_gt, which aligns with my understanding of how this should work:

# Compute Accuracy and L1 Loss for Logging
action_logits = output.logits[:, wrapped_model.module.vla.vision_backbone.featurizer.patch_embed.num_patches : -1]
action_preds = action_logits.argmax(dim=2)
action_gt = batch["labels"][:, 1:].to(action_preds.device)
mask = action_gt > 32000

# Compute Accuracy
correct_preds = (action_preds == action_gt) & mask
action_accuracy = correct_preds.sum().float() / mask.sum().float()

Would you be able to clarify whether the difference in latent_action selection between training and evaluation is a deliberate design choice, a potential oversight, or perhaps something I'm misunderstanding about the code flow?

Thank you very much for your time and help—I really appreciate it!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions