nanochat · Issues· 120 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #848
SFT inherits pre-batch-scaling learning rates from base checkpoint
Updated Sep 9, 2026 - #844
Multiple causal-leakage architectures can collapse NanoChat val_bpb to 0.0018
Updated Sep 1, 2026 - #840
TaskSequence in tasks/common.py is unused
Updated Aug 28, 2026 - #838
"Peak memory usage" reports 0.00MiB on CPU and MPS
Updated Aug 28, 2026 - #590
[Bug] scripts/chat_sft.py produces loss: nan from step 00001 on small device-batch-size (≤8) due to fully-masked micro-batches
code robustnessUpdated Aug 26, 2026 - #810
chat_rl: the sampled assistant_end token is masked out, so stopping is never reinforced
potential_bugUpdated Aug 24, 2026 - #825
SDPA fallback: sliding window doesn't reduce memory (full mask built regardless of window size)
Updated Aug 9, 2026 - #698
RTX 3050TI Not detecting m.get_device()
Updated Aug 4, 2026 - #737
Use Official PyTorch FA3 Builds
Updated Aug 2, 2026 - #820
FA3 loader reports success on Blackwell (sm_120) then dies at first kernel launch
Updated Aug 2, 2026 - #756
Token smearing does not correctly handle chunked prefill / chunk inference (first token of a chunk misses its cross-chunk predecessor
potential_bugUpdated Jul 31, 2026 - #284
Add support for TPU?
featureUpdated Jul 26, 2026 - #804
Add support for OpenCL (e.g., for the integrated AMD GPU)?
Updated Jul 10, 2026 - #791
TODO: hash for d34 model
docsUpdated Jun 26, 2026 - #427
base_eval.py: hellaswag gets progressively slower and leaks memory on small models (Mac Studio)
performanceUpdated Jun 3, 2026