Full-Scale Experiment with exact_accuracy 0
Author: andreyferriyanCreated Sep 5, 2025Updated May 2, 2026
Hi, we have finished training for over a week with a single NVIDIA RTX A6000.
This is how we run the experiment OMP_NUM_THREADS=8 torchrun --nproc-per-node 1 pretrain.py global_batch_size=16.
Compared to the small experiment, this full experiment was worse. Is there anything wrong with my configuration?.
Source: sapientinc/HRM