[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.
暂无评论,来聊聊你的看法吧