ART · Issues· 143 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #902
Investigate output and gradient variability between matched same-source Qwen TrainerRank actors
Updated Sep 17, 2026 - #870
TrainerRank reusable-cache admission can leave insufficient physical CUDA memory for backward library allocations
Updated Sep 17, 2026 - #848
TrainerRank MoE memory admission: backward-recompute OOM and profile calibration failures
Updated Sep 17, 2026 - #917
Support combined expert and data parallelism in TrainerRank
Updated Sep 16, 2026 - #916
Support pipeline parallelism in TrainerRank without exposing stage scheduling to callers
Updated Sep 16, 2026 - #901
Investigate large hidden-state differences when independent requests are co-packed (Qwen3.6, no-grad)
Updated Sep 15, 2026 - #883
Retained sampled suffix can carry logprobs under a different causal prefix
Updated Sep 11, 2026 - #869
Determine whether a large dynamics RL batch refusal is necessary or avoidable by execution layout
Updated Sep 11, 2026 - #364
Feature: Dr. GRPO Support
enhancementUpdated Sep 4, 2026 - #851
Qwen3.5-35B-A3B at EP1 with CP>1 segfaults in Transformer Engine's grouped GEMM on real-data routing (likely empty local experts)
Updated Sep 4, 2026 - #678
Dedicated LocalBackend vLLM server can become unreachable during LoRA adapter reload
Updated Jul 27, 2026 - #672
Regression: forked GRPO run can produce malformed checkpoint outputs on newer ART ref
bugUpdated May 8, 2026 - #661
LoRA adapters are not properly syncing when using LocalBackend - stale adapters are being used in rollouts
Updated Apr 24, 2026 - #651
LocalBackend fork_checkpoint doesn't update vLLM's initial LoRA
Updated Apr 13, 2026 - #649
Add from_entity parameter to _experimental_fork_checkpoint
Updated Apr 10, 2026