OpenRLHF · Issues· 381 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #1358
Suboptimal forward dispatch order in make_experience causes unnecessary serialization
Updated Sep 18, 2026 - #1349
LoRA + colocated vLLM: policy→vLLM weight sync fails end-to-end (dtype, then unmapped adapter names)
Updated Sep 14, 2026 - #1317
reward model HTTP service has no authentication, defaults to binding all interfaces, and the training client consumes whatever the endpoint returns
Updated Sep 9, 2026 - #1322
[Bug] Async oversampling checkpoints drop buffered rollouts on resume
Updated Sep 4, 2026 - #1311
[Bug] RewardDataset silently keeps preference pairs that become identical after truncation
Updated Aug 21, 2026 - #1303
[Roadmap] OpenRLHF on Intel XPU
Updated Aug 12, 2026 - #1295
[Bug] seq-mask-tis skips token-level truncation and can produce non-finite PPO updates
Updated Aug 8, 2026 - #1011
assert "expandable_segments:True" not in conf报错
Updated Jul 27, 2026 - #1273
Molt brings an Automodel-powered backend (AutoTP/EP/CP) to OpenRLHF
Updated Jul 27, 2026 - #1270
Droping experiences should also consider standard deviation when using non-binary rewards
Updated Jul 25, 2026 - #1263
Feature request: reward-hacking onset monitoring hooks during PPO/GRPO training
Updated Jul 12, 2026 - #1164
Fix SignalActor concurrency issues and refactor state management based on condition lock
Updated Jul 9, 2026 - #1243
overlong_penalty should exclude tool response tokens from length calculation (use action_mask/action_ranges)
Updated Jun 18, 2026 - #1244
Consider using pre-built Flash Attention kernels via `kernels`
Updated May 28, 2026 - #1236
ray error for multi-nodes training
Updated May 13, 2026