Forward reward always 0
Author: ZhaoyangLi-1Created Oct 24, 2025Updated Oct 24, 2025
I run grpo_classification.py with my own dataset, while my code is always 0. How ot fix.
I use Qwen3-VL-8B-Instruct as the base model
Source: huggingface/open-r1
I run grpo_classification.py with my own dataset, while my code is always 0. How ot fix.
I use Qwen3-VL-8B-Instruct as the base model
Source: huggingface/open-r1