DeepSpeed · Issues· 1435 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #8572
[BUG] Runtime configuration validation is bypassed under python -O
Updated Sep 17, 2026 - #5537
[BUG] FlopsProfiler upsample flops compute bug
bugtrainingUpdated Sep 17, 2026 - #8443
ZeRO-3 applies Muon's Newton-Schulz once per micro-batch, so gradient accumulation changes the optimizer
Updated Sep 17, 2026 - #8458
[BUG] Do not silently ignore use_shared_prefill with continuous batching
bugtrainingUpdated Sep 17, 2026 - #8489
Deprecate unused DeepSpeed features
Updated Sep 17, 2026 - #8173
[RFC] Support sharding LM heads and adopting Online Softmax
enhancementUpdated Sep 17, 2026 - #5410
[BUG] Gradient Accumulation Steps Initialization Bug in Pipeline Parallel Mode
bugtrainingUpdated Sep 17, 2026 - #1101
zero_to_fp32.py drops buffers
Updated Sep 17, 2026 - #8554
2 tests fail: assert isinstance(model.llm[0], DistributedAttention)
Updated Sep 16, 2026 - #8531
[REQUEST] Register native pinned host memory with device runtime for NPU (follow-up to #8283 and #8315)
enhancementUpdated Sep 16, 2026 - #8197
[REQUEST] OPSD Profile and improve HybridEngine rollout performance
enhancementUpdated Sep 15, 2026 - #8290
HF transformers main injects 'embedding_rowwise' into tp_plan for tied-embedding models; AutoTP rejects the whole plan
Updated Sep 15, 2026 - #7819
[BUG] DeepSpeed ZeRO Stage-3 + CPU offloaded optimizer (CPUAdam) inconsistency metadata between subgroup
bugtrainingUpdated Sep 15, 2026 - #8514
[BUG] AutoTP checkpoint load raises AttributeError on a model containing nn.InstanceNorm1d
Updated Sep 14, 2026 - #5446
[BUG] Can't pickle local object 'instrument_w_nvtx.<locals>.wrapped_fn'
bugtrainingUpdated Sep 14, 2026