Recently, I have been testing the recall agent and the ranking agent with the DeepSeek backend. However, I found that the current mock data corpus is indeed too simple -- Our framework here will be a total overkill. Therefore, I suggest that we can involve industrial-level data for better validation of the framework's real-world effectiveness and further demonstration.
Kuaishou's https://huggingface.co/datasets/OpenOneRec/OpenOneRec-RecIF will be a good start.
onerec_bench_release.parquet contains users' video/ad/goods history, engagement labels, user profiles, CoT reasoning
video_ad_pid2sid.parquet, product_pid2sid.parquet, pid2caption.parquet contains ID mappings
I can go on to develop this feature. Please kindly assign this to me if it looks ok to you, @guoxun , thank you!
Recently, I have been testing the recall agent and the ranking agent with the DeepSeek backend. However, I found that the current mock data corpus is indeed too simple -- Our framework here will be a total overkill. Therefore, I suggest that we can involve industrial-level data for better validation of the framework's real-world effectiveness and further demonstration.
Kuaishou's https://huggingface.co/datasets/OpenOneRec/OpenOneRec-RecIF will be a good start.
onerec_bench_release.parquetcontains users' video/ad/goods history, engagement labels, user profiles, CoT reasoningvideo_ad_pid2sid.parquet,product_pid2sid.parquet,pid2caption.parquetcontains ID mappingsI can go on to develop this feature. Please kindly assign this to me if it looks ok to you, @guoxun , thank you!