1 code implementation • 20 Feb 2024 • Xueyang Feng, Zhi-Yuan Chen, Yujia Qin, Yankai Lin, Xu Chen, Zhiyuan Liu, Ji-Rong Wen
We construct a human-agent collaboration dataset to train this policy model in an offline reinforcement learning environment.