llama-cpp-python · Issues· 678 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #2231
Request for Gemma 3 4B & Gemma 4 Integration Support with llama-cpp-python
Updated Sep 6, 2026 - #2362
Long whitespace tokens detokenize to empty bytes, breaking tokenize/detokenize round-trip
Updated Sep 4, 2026 - #2361
[Bug] DFlash 2 speculative decoding crashes and fails in Python due to missing ctx_other and non-autoregressive graph execution
Updated Aug 20, 2026 - #2013
Can't install with GPU support with Cuda toolkit 12.9 and Cuda 12.9
Updated Aug 16, 2026 - #2352
uv add llama-cpp-python wheels fails for versions above 0.3.30
Updated Jul 31, 2026 - #1169
Docker llama-cpp libcuda.so.1: cannot open shared object file: No such file or directory
Updated Jul 22, 2026 - #2341
Support sm100 and sm120 in CUDA pre‑built wheels
Updated Jul 16, 2026 - #2029
Access Violation issue facing for exe created using pyinstaller
Updated Jul 10, 2026 - #2211
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_array
Updated Jul 7, 2026 - #635
Documentation of server command line parameters.
documentationUpdated Jun 14, 2026 - #2175
How to use "Gemma-4-E4B-it-heretic-GGUF" ?
Updated Jun 11, 2026 - #463
CUBLAS and CLBLAST builds on Windows
documentationduplicatewindowsUpdated May 24, 2026 - #450
OpenBLAS blas_thread_ init Error - Linux Centos 7 with llama-cpp-python == 0.1.67
wontfixUpdated May 24, 2026 - #409
Could not find nvcc, please set CUDAToolkit_ROOT
duplicatebuildhardwarellama.cppwindowsUpdated May 24, 2026 - #383
Prompt for instruction models to be used in server
modelqualityUpdated May 24, 2026