DeepSpeedExamples · Issues· 330 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #996
Compute_rewards in PPO:rewards[j, start:ends[j]][-1] += reward_clip[j] is wrong
Updated Dec 22, 2025 - #995
Step1 failed when I run "bash training_scripts/opt/single_gpu/run_1.3b.sh"
Updated Dec 16, 2025 - #989
One example, multiple config files
Updated Sep 8, 2025 - #703
how to understand the code for calculating rewards
Updated Aug 28, 2025 - #984
moe example 404
Updated Jul 25, 2025 - #186
OOM despite ZeRO stage 3
Updated Jul 23, 2025 - #979
Possible to include an example of DeepNVMe + state dict
Updated Jun 20, 2025 - #943
KV_cache offload
Updated Jun 12, 2025 - #172
My deepspeed code is very slow
Updated Jun 2, 2025 - #956
Why Does vf_loss Take the Maximum Value, Rendering Clamp Meaningless?
Updated May 17, 2025 - #969
Apply Zero-3 and LoRA appears empty lora weight [0]
Updated May 11, 2025 - #845
torch.distributed.DistBackendError: NCCL error in: ../torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1333, remote process exited or there was a network error, NCCL version 2.18.6
Updated Apr 24, 2025 - #960
DeepSpeed-FastGen support ascend npu?
Updated Apr 7, 2025 - #892
Does Zero-Inference support TP?
Updated Mar 21, 2025 - #946
Assertion `srcIndex < srcSelectDimSize` failed
Updated Jan 24, 2025