Hi, thank you for releasing AutoGaze and NVILA-HD.
I would like to better understand the training procedure used for NVILA-HD. In particular, after integrating AutoGaze into NVILA, did you:
- retrain the model through all five training stages using the full set of training datasets, or
- initialize from an existing NVILA checkpoint and only perform the final adaptation stage?
Could you also provide more details about the NVILA-HD training setup?
These details would be very helpful for reproducing NVILA-HD and for understanding how much additional training is required when integrating AutoGaze into an existing MLLM.
Thank you!
Hi, thank you for releasing AutoGaze and NVILA-HD.
I would like to better understand the training procedure used for NVILA-HD. In particular, after integrating AutoGaze into NVILA, did you:
Could you also provide more details about the NVILA-HD training setup?
These details would be very helpful for reproducing NVILA-HD and for understanding how much additional training is required when integrating AutoGaze into an existing MLLM.
Thank you!