[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.
[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.
No open issues yet, or sync has not completed.