Hello,
I've been looking through the codebase and noticed a difference in how latent_action is extracted during training versus evaluation. I wanted to ask about this to better understand the implementation, as it might be a deliberate design choice that I'm not fully grasping.
In vla-scripts/finetune_libero.py, during training, latent_action_tokens are selected using a mask based on action_gt > 32000:
def action_decoder_forward(self, batch, vla_output):
# ... (other code)
action_gt = batch["labels"].to(latent_tokens.device)
mask = action_gt > 32000
latent_action_tokens = []
for idx, per_sample_latent_tokens in enumerate(latent_tokens):
per_sample_latent_action_tokens = per_sample_latent_tokens[mask[idx], :]
latent_action_tokens.append(per_sample_latent_action_tokens)
# ... (uses latent_action_tokens for prediction)
However, in the evaluation code (experiments/robot/libero/run_libero_eval.py), latent_action is selected using the positions of the last four tokens, which is equivalent to using generated_ids > 32000 to create a mask for extracting latent_action:
def predict_latent_action(...):
# ... (other code)
latent_action = latent_tokens[:, -4:]
# ...
Since the LlamaForCausalLMbackbone uses next token prediction, I'm wondering if these two approaches might be extracting latent_action from different positions, creating a potential mismatch between training and evaluation.
I also noticed another section of code that accounts for the next token prediction behavior by using the mask with action_gt, which aligns with my understanding of how this should work:
# Compute Accuracy and L1 Loss for Logging
action_logits = output.logits[:, wrapped_model.module.vla.vision_backbone.featurizer.patch_embed.num_patches : -1]
action_preds = action_logits.argmax(dim=2)
action_gt = batch["labels"][:, 1:].to(action_preds.device)
mask = action_gt > 32000
# Compute Accuracy
correct_preds = (action_preds == action_gt) & mask
action_accuracy = correct_preds.sum().float() / mask.sum().float()
Would you be able to clarify whether the difference in latent_action selection between training and evaluation is a deliberate design choice, a potential oversight, or perhaps something I'm misunderstanding about the code flow?
Thank you very much for your time and help—I really appreciate it!
Hello,
I've been looking through the codebase and noticed a difference in how
latent_actionis extracted during training versus evaluation. I wanted to ask about this to better understand the implementation, as it might be a deliberate design choice that I'm not fully grasping.In
vla-scripts/finetune_libero.py, during training,latent_action_tokensare selected using a mask based onaction_gt > 32000:However, in the evaluation code (
experiments/robot/libero/run_libero_eval.py), latent_action is selected using the positions of the last four tokens, which is equivalent to using generated_ids > 32000 to create a mask for extracting latent_action:Since the
LlamaForCausalLMbackboneuses next token prediction, I'm wondering if these two approaches might be extractinglatent_actionfrom different positions, creating a potential mismatch between training and evaluation.I also noticed another section of code that accounts for the next token prediction behavior by using the mask with
action_gt, which aligns with my understanding of how this should work:Would you be able to clarify whether the difference in
latent_actionselection between training and evaluation is a deliberate design choice, a potential oversight, or perhaps something I'm misunderstanding about the code flow?Thank you very much for your time and help—I really appreciate it!