minimind · Issues· 68 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #823
[学习]零基础小白缓慢学习记录
Updated Sep 17, 2026 - #858
[Bug] num_experts_per_tok=1 时 MoE 路由器收不到梯度(topk 权重归一化后退化为常数)
Updated Sep 17, 2026 - #856
在 Mac 上如果不做任何修改直接运行 python trainer/train_pretrain.py,它会完全运行在 CPU 上,无法调用 Mac 的 GPU。
Updated Sep 16, 2026 - #504
【推荐内容】合集
documentationenhancementgood first issueUpdated Sep 15, 2026 - #835
MoE feed-forward performs ~3×num_experts device→host syncs per layer, per forward
Updated Sep 7, 2026 - #824
MiniMind可视化扩展,模型结构、注意力机制和训练流程展示
Updated Aug 31, 2026 - #801
[Feature] 独立可插拔 MoE V2 模块 + moe_type 切换机制
Updated Aug 26, 2026 - #670
感谢大佬开源,基于 HuggingFace 框架重写了核心功能
Updated Aug 20, 2026 - #795
【学习】MiniMind 分布式并行训练扩展:TP / SP / VP / PP / CP
Updated Aug 4, 2026 - #804
[Feature Request] add On-Policy Distillation (OPD)
Updated Jul 24, 2026 - #800
分享自己学习 MiniMind 源码写的精读笔记
Updated Jul 21, 2026 - #810
在训练PPO时遇到的问题
Updated Jul 14, 2026 - #764
MiniMind+Twinkle 线上训练服务
Updated Jun 26, 2026 - #799
有没有出更多数据集的打算
Updated Jun 20, 2026 - #620
如何让模型学会100内的加减法?
help wantedquestionUpdated Jun 12, 2026