Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

nanochat · Issues· 120 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #848

    SFT inherits pre-batch-scaling learning rates from base checkpoint

    Updated Sep 9, 2026
  • #844

    Multiple causal-leakage architectures can collapse NanoChat val_bpb to 0.0018

    Updated Sep 1, 2026
  • #840

    TaskSequence in tasks/common.py is unused

    Updated Aug 28, 2026
  • #838

    "Peak memory usage" reports 0.00MiB on CPU and MPS

    Updated Aug 28, 2026
  • #590

    [Bug] scripts/chat_sft.py produces loss: nan from step 00001 on small device-batch-size (≤8) due to fully-masked micro-batches

    code robustnessUpdated Aug 26, 2026
  • #810

    chat_rl: the sampled assistant_end token is masked out, so stopping is never reinforced

    potential_bugUpdated Aug 24, 2026
  • #825

    SDPA fallback: sliding window doesn't reduce memory (full mask built regardless of window size)

    Updated Aug 9, 2026
  • #698

    RTX 3050TI Not detecting m.get_device()

    Updated Aug 4, 2026
  • #737

    Use Official PyTorch FA3 Builds

    Updated Aug 2, 2026
  • #820

    FA3 loader reports success on Blackwell (sm_120) then dies at first kernel launch

    Updated Aug 2, 2026
  • #756

    Token smearing does not correctly handle chunked prefill / chunk inference (first token of a chunk misses its cross-chunk predecessor

    potential_bugUpdated Jul 31, 2026
  • #284

    Add support for TPU?

    featureUpdated Jul 26, 2026
  • #804

    Add support for OpenCL (e.g., for the integrated AMD GPU)?

    Updated Jul 10, 2026
  • #791

    TODO: hash for d34 model

    docsUpdated Jun 26, 2026
  • #427

    base_eval.py: hellaswag gets progressively slower and leaks memory on small models (Mac Studio)

    performanceUpdated Jun 3, 2026