verl · Issues· 1228 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #5192
Google TPU support in verl with Ray
Updated Sep 18, 2026 - #7914
[megatron] seq-mean-token-mean guard rejects CP=1 and appears overly restrictive for reconstructed THD outputs
Updated Sep 18, 2026 - #7913
[Bug] GMPO geo_mean loss changes when a minibatch is split into more microbatches
bugUpdated Sep 18, 2026 - #6280
`log_prob` and `old_log_prob` differ in on-policy setting when rollout and actor micro batch per gpu differ
Updated Sep 18, 2026 - #7909
Multi-LoRA training: resident adapters and mixed-adapter steps on one base model
Updated Sep 17, 2026 - #7904
[rollout][vllm] Garbled multi-language output after weight sync when `free_cache_engine=true` (sleep/resume corrupts rollout weights)
bugUpdated Sep 17, 2026 - #7899
[bug] equal-length batches of 3D mRoPE position_ids produce an inconsistent jagged layout → split_with_sizes crash in V1 trainer
Updated Sep 17, 2026 - #7060
[Tracking] Sharded delta weight sync (delta_sharded): roadmap & known issues
Updated Sep 15, 2026 - #6252
Qwen3.5/Qwen3.6 35B-A3B 多轮工具调用 Agent RL 训练中出现工具调用格式异常并导致崩溃
Updated Sep 14, 2026 - #7798
[Bug] Qwen3.5 system-only prompt still fails through Continuous Token
Updated Sep 14, 2026 - #7856
[RFC] Tail-aware top-k KL for on-policy distillation
Updated Sep 13, 2026 - #7839
[RFC] Enhanced `verl.single_controller`: a backend-neutral WorkerGroup runtime
Updated Sep 11, 2026 - #7834
[fsdp] Tied embeddings disable rank0-only weight loading, so host RAM scales with rank count; the naive fix unties the model silently
Updated Sep 10, 2026 - #7829
[Bug] KL-Cov token selection changes with the PPO microbatch size
bugUpdated Sep 10, 2026 - #7824
[data] The val_max_samples subsample keys off data.shuffle, so data.validation_shuffle never reaches it
Updated Sep 10, 2026