Hi LiteResearcher team,
I’m very interested in reproducing and learning from the SFT → RL training pipeline described in your paper.
I found the Stage-1/Stage-2 RL prompts and the final LiteResearcher-4B-RL checkpoint on Hugging Face, but I couldn’t find the SFT trajectory data used for the cold-start stage.
Do you have any plan to release the SFT trajectories? If releasing the full data is not convenient, would it be possible to provide a small subset or a few examples in the training format? This would be very helpful for understanding and reproducing the cold-start SFT stage.
Thank you!
Hi LiteResearcher team,
I’m very interested in reproducing and learning from the SFT → RL training pipeline described in your paper.
I found the Stage-1/Stage-2 RL prompts and the final LiteResearcher-4B-RL checkpoint on Hugging Face, but I couldn’t find the SFT trajectory data used for the cold-start stage.
Do you have any plan to release the SFT trajectories? If releasing the full data is not convenient, would it be possible to provide a small subset or a few examples in the training format? This would be very helpful for understanding and reproducing the cold-start SFT stage.
Thank you!