Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

OpenRLHF · Issues· 381 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #1358

    Suboptimal forward dispatch order in make_experience causes unnecessary serialization

    Updated Sep 18, 2026
  • #1349

    LoRA + colocated vLLM: policy→vLLM weight sync fails end-to-end (dtype, then unmapped adapter names)

    Updated Sep 14, 2026
  • #1317

    reward model HTTP service has no authentication, defaults to binding all interfaces, and the training client consumes whatever the endpoint returns

    Updated Sep 9, 2026
  • #1322

    [Bug] Async oversampling checkpoints drop buffered rollouts on resume

    Updated Sep 4, 2026
  • #1311

    [Bug] RewardDataset silently keeps preference pairs that become identical after truncation

    Updated Aug 21, 2026
  • #1303

    [Roadmap] OpenRLHF on Intel XPU

    Updated Aug 12, 2026
  • #1295

    [Bug] seq-mask-tis skips token-level truncation and can produce non-finite PPO updates

    Updated Aug 8, 2026
  • #1011

    assert "expandable_segments:True" not in conf报错

    Updated Jul 27, 2026
  • #1273

    Molt brings an Automodel-powered backend (AutoTP/EP/CP) to OpenRLHF

    Updated Jul 27, 2026
  • #1270

    Droping experiences should also consider standard deviation when using non-binary rewards

    Updated Jul 25, 2026
  • #1263

    Feature request: reward-hacking onset monitoring hooks during PPO/GRPO training

    Updated Jul 12, 2026
  • #1164

    Fix SignalActor concurrency issues and refactor state management based on condition lock

    Updated Jul 9, 2026
  • #1243

    overlong_penalty should exclude tool response tokens from length calculation (use action_mask/action_ranges)

    Updated Jun 18, 2026
  • #1244

    Consider using pre-built Flash Attention kernels via `kernels`

    Updated May 28, 2026
  • #1236

    ray error for multi-nodes training

    Updated May 13, 2026