Megatron-LM · Issues· 1371 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #7171
[Feature request] Add Engram memory to HybridModel training
enhancementcommunity-requestwaiting-on-customerUpdated Sep 18, 2026 - #6757
[ROADMAP][2026 Q3] Megatron Core MoE Roadmap
call for contributionUpdated Sep 18, 2026 - #7484
[Inference] Expert tensor parallelism is rejected for inference-optimized MoE and has no regression coverage
community-requestUpdated Sep 18, 2026 - #7452
Silent numerical-correctness bugs in Megatron Core parallelism composition
community-requestUpdated Sep 18, 2026 - #7174
[BUG] Full validation fails with a single dataset
community-requestUpdated Sep 18, 2026 - #7173
[BUG] Scheduler override overwrites per-group LR bounds
community-requestUpdated Sep 18, 2026 - #7172
[BUG] Hybrid MTP1 with MoE fails during loss logging
community-requestUpdated Sep 18, 2026 - #6872
Kimi-K3 training support
enhancementUpdated Sep 18, 2026 - #5676
[ROADMAP][2026 Q3] Megatron Core Roadmap
call for contributionUpdated Sep 18, 2026 - #4167
Feature Request: ScatterMoE (Triton-based Sparse MoE with fused scatter/gather GEMMs)
enhancementUpdated Sep 18, 2026 - #7475
`current_max_attn_logits` breaks CUDA Graph replay because its Python-level assignment is not re-executed
bugcommunity-requestUpdated Sep 18, 2026 - #4468
DeepSeek-V4 training support
enhancementUpdated Sep 18, 2026 - #7338
[BUG] get_grad_norm_fp32 / clip_grad_by_total_norm_fp32 pass a mixed bf16/fp32 gradient list to one TE multi_tensor launch -> illegal memory access (precision-aware optimizer + natively fp32 params)
community-requestUpdated Sep 18, 2026 - #7389
[BUG] Shared expert overlap delays independent expert wgrad until input-gradient merge
community-requestwaiting-on-maintainersUpdated Sep 18, 2026 - #7464
[Bug] Vocab-parallel label smoothing uses the local vocabulary size and the local mean log-probability
bugcommunity-requestUpdated Sep 18, 2026