slime · Issues· 498 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #2214
Your project is on StackMap — a curated map of the AI stack
Updated Sep 19, 2026 - #2388
[Bug] Qwen-VL patch bypasses SGLang token-ID preservation and breaks multi-turn TITO
bugUpdated Sep 16, 2026 - #2387
[Question] Would an adaptive in-reward KL controller fit Slime's PPO/RLHF scope?
questionUpdated Sep 15, 2026 - #1761
[Bug] Resume advances Megatron scheduler with rollout_id instead of train step
bugUpdated Sep 15, 2026 - #2386
Verify evals on Papers with Code
Updated Sep 14, 2026 - #2384
[Bug] Broadcasting full rollout_routed_experts data to every Megatron training rank will cause host OOM
bugUpdated Sep 14, 2026 - #2339
[Discussion] Support background evaluation while fully-async training rollout is running
questionUpdated Sep 14, 2026 - #2370
[Bug] save_model fires unconditionally at the final step regardless of --save-interval; final optimizer-state write can kill the job
Updated Sep 13, 2026 - #2374
[Bug] shipped dynamic-sampling filters crash on fan-out groups (list[list[Sample]]) from multi-turn agent rollouts
Updated Sep 9, 2026 - #1136
[BUG] 稳定训一半OOM,已开cp
Updated Sep 8, 2026 - #1487
Training hangs indefinitely after rollout phase completes on 8 B200 GPUs with CP=4, TP=2
Updated Sep 6, 2026 - #200
UX/Unintended bug: sample.reward is None on aborted sample
Updated Sep 5, 2026 - #2338
[Question] realign 使用本轮 response 长度判断
questionUpdated Sep 5, 2026 - #397
if we use GRPO and args.kl_coef is non-zero, is the KL computation incorrect?
Updated Sep 5, 2026 - #1462
cannot import name 'GPU_MEMORY_TYPE_CUDA_GRAPH' from 'sglang.srt.constants'
Updated Sep 5, 2026