#707·open-r1

Forward reward always 0

Author: ZhaoyangLi-1Created Oct 24, 2025Updated Oct 24, 2025

I run grpo_classification.py with my own dataset, while my code is always 0. How ot fix.

I use Qwen3-VL-8B-Instruct as the base model