Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

Megatron-LM · Issues· 1371 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #7171

    [Feature request] Add Engram memory to HybridModel training

    enhancementcommunity-requestwaiting-on-customerUpdated Sep 18, 2026
  • #6757

    [ROADMAP][2026 Q3] Megatron Core MoE Roadmap

    call for contributionUpdated Sep 18, 2026
  • #7484

    [Inference] Expert tensor parallelism is rejected for inference-optimized MoE and has no regression coverage

    community-requestUpdated Sep 18, 2026
  • #7452

    Silent numerical-correctness bugs in Megatron Core parallelism composition

    community-requestUpdated Sep 18, 2026
  • #7174

    [BUG] Full validation fails with a single dataset

    community-requestUpdated Sep 18, 2026
  • #7173

    [BUG] Scheduler override overwrites per-group LR bounds

    community-requestUpdated Sep 18, 2026
  • #7172

    [BUG] Hybrid MTP1 with MoE fails during loss logging

    community-requestUpdated Sep 18, 2026
  • #6872

    Kimi-K3 training support

    enhancementUpdated Sep 18, 2026
  • #5676

    [ROADMAP][2026 Q3] Megatron Core Roadmap

    call for contributionUpdated Sep 18, 2026
  • #4167

    Feature Request: ScatterMoE (Triton-based Sparse MoE with fused scatter/gather GEMMs)

    enhancementUpdated Sep 18, 2026
  • #7475

    `current_max_attn_logits` breaks CUDA Graph replay because its Python-level assignment is not re-executed

    bugcommunity-requestUpdated Sep 18, 2026
  • #4468

    DeepSeek-V4 training support

    enhancementUpdated Sep 18, 2026
  • #7338

    [BUG] get_grad_norm_fp32 / clip_grad_by_total_norm_fp32 pass a mixed bf16/fp32 gradient list to one TE multi_tensor launch -> illegal memory access (precision-aware optimizer + natively fp32 params)

    community-requestUpdated Sep 18, 2026
  • #7389

    [BUG] Shared expert overlap delays independent expert wgrad until input-gradient merge

    community-requestwaiting-on-maintainersUpdated Sep 18, 2026
  • #7464

    [Bug] Vocab-parallel label smoothing uses the local vocabulary size and the local mean log-probability

    bugcommunity-requestUpdated Sep 18, 2026