BitNet · Issues· 332 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #629
Latent int16 overflow in ggml_vec_dot_i2_i8_s_1x1: 128 maddubs results per lane against a 32767 ceiling
Updated Sep 17, 2026 - #631
build/bin/llama-cli throws SIGSEGV, Segmentation fault when using a TL2 GGUF model
Updated Sep 17, 2026 - #630
Verify evals on Papers with Code
Updated Sep 14, 2026 - #628
i2_s AVX2 GEMM folds the int16 accumulator 10x more often than its own bound requires (7.5% measured)
Updated Sep 13, 2026 - #588
[Bug]: BitNet FFN uses SILU instead of ReLU² — perplexity 99.8 vs 17.1 on every CPU backend
Updated Sep 13, 2026 - #600
BitNet-b1.58-2B-4T produces garbage output on ARM64/NEON (no AVX2) — scalar fallback uses wrong I2_S unpacking scheme
Updated Sep 12, 2026 - #618
ggml-cpu.c: `src1_cont` undeclared → build fails on ARM CPUs with i8mm (Apple M2/M3/M4, Graviton3, ...)
Updated Sep 9, 2026 - #602
bitnet-b1.58-2B-4T runs FFN with SiLU instead of relu2 - wrong logits on every backend
Updated Sep 7, 2026 - #622
Regression: Falcon-E models cannot be loaded — "unknown pre-tokenizer type: 'falcon_e'" (support added in #268 was lost in the llama.cpp submodule update)
Updated Sep 6, 2026 - #621
convert-hf-to-gguf-bitnet.py --outtype i2_s silently writes F16 (not I2_S) for LlamaForCausalLM BitNet checkpoints (Falcon3 / Falcon-E 1.58bit)
Updated Sep 6, 2026 - #619
README "Convert from .safetensors" flow is broken: convert-ms-to-gguf-bitnet.py KeyError MODEL_ARCH.BITNET_25, and llama-quantize has no I2_S ftype (setup_env.py i2_s conversion path also affected)
Updated Sep 5, 2026 - #611
setup_env.py on MacBook M2 failed
Updated Sep 5, 2026 - #617
[Bug]: I2_S GEMM fast path produces garbage for multi-token prompts on AVX-only CPUs
Updated Aug 26, 2026 - #547
[Bug]: i2_s quantization produces garbage output on x86 CPUs without AVX2 (Missing fallback)
Updated Aug 26, 2026 - #560
Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead.
Updated Aug 20, 2026