Thanks for your great work!
The paper mentions that the model is fine-tuned using 180 episodes for Cover Blocks and 308 episodes for Put Back Block. I’m wondering what the starting point of this fine-tuning is.
Is the Qwen3.5-4B + Action Expert model trained directly on these 180/308 episodes, with the Action Expert initialized from scratch? Or has the Qwen3.5-4B + Action Expert model already been pretrained on a large-scale robotics dataset before being fine-tuned on these task-specific demonstrations?
Thanks for your great work!
The paper mentions that the model is fine-tuned using 180 episodes for Cover Blocks and 308 episodes for Put Back Block. I’m wondering what the starting point of this fine-tuning is.
Is the Qwen3.5-4B + Action Expert model trained directly on these 180/308 episodes, with the Action Expert initialized from scratch? Or has the Qwen3.5-4B + Action Expert model already been pretrained on a large-scale robotics dataset before being fine-tuned on these task-specific demonstrations?