Baike.dev
All toolsTrendingOpen sourceNewsSubmit
Log in
< 返回工具列表
B

BitNet

> 编程语言
开源

Official inference framework for 1-bit LLMs

39.8K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

Official inference framework for 1-bit LLMs

## Overview bitnet.cpp is the official inference framework for 1-bit LLMs (e.g., BitNet b1.58). It offers a suite of optimized kernels that support **fast** and **lossless** inference of 1.58-bit models on **CPU** and **GPU** (NPU support coming next). Try it out via this [online demo](https://demo-bitnet-h0h8hcfqeqhrf5gf.canadacentral-01.azurewebsites.net/), or build and run it on your own [CPU](https://github.com/microsoft/BitNet?tab=readme-ov-file#build-from-source) or [GPU](https://github.com/microsoft/BitNet/blob/main/gpu/README.md). bitnet.cpp achieves speedups of **1.37x** to **5.07x** on ARM CPUs, with larger models experiencing greater performance gains. Additionally, it reduces energy consumption by **55.4%** to **70.0%**, further boosting overall efficiency. On x86 CPUs, speedups range from **2.37x** to **6.17x** with energy reductions between **71.9%** to **82.2%**. Furthermore, bitnet.cpp can run a 100B BitNet b1.58 model on a single CPU, achieving speeds comparable to human reading (5-7 tokens per second), significantly enhancing the potential for running LLMs on local devices. Please refer to the [technical report](https://arxiv.org/abs/2410.16144) for more details. ## Model Releases ### 1. [BitNet-b1.58-2B-4T](https://huggingface.co/microsoft/BitNet-b1.58-2B-4T) - 1-bit Large Language Model **BitNet-b1.58-2B-4T** is the first official BitNet b1.58 model with **2.4B parameters**, trained on **4 trillion tokens**. It is a ternary (1.58-bit) language model that delivers competitive performance with full-precision models of similar size while enabling significantly faster and more energy-efficient inference. - **Fast CPU Inference**: Achieves up to **6.17x speedup** on x86 CPUs and **5.07x** on ARM CPUs compared to full-precision models. - **Energy Efficient**: Reduces energy consumption by up to **82.2%** on x86 and **70.0%** on ARM. - **GPU Support**: Official GPU inference kernel available for accelerated deployment. - **Chat-Ready**: Supports conversational mode for interactive use. [🤗 Hugging Face](https://huggingface.co/microsoft/BitNet-b1.58-2B-4T) | [🔗 Online Demo](https://demo-bitnet-h0h8hcfqeqhrf5gf.canadacentral-01.azurewebsites.net/) | [📄 Technical Report](https://arxiv.org/abs/2410.16144) ### 2. [BitNet-embedding-0.6B](https://huggingface.co/microsoft/BitNet-embedding-0.6B) - 1-bit Embedding Model **BitNet-embedding-0.6B** is a **0.6B-parameter** 1-bit embedding model that achieves competitive embedding quality with significantly faster CPU inference. It is the first model to demonstrate that ternary weights can deliver strong performance on embedding tasks. - **1.42x to 2.28x speedup** over F16 on prefill (8 threads, x86) - **Lossless Quality**: Competitive embedding quality with 2 bits per weight - **I2_S Kernel**: Supports optimized I2_S conversion on x86 CPUs [🤗 Hugging Face](https://huggingface.co/microsoft/BitNet-embedding-0.6B) | [📄 I2_S Guide](docs/bitnet-embeddings-i2s-guide.md) ### 3. [BitNet-embedding-270M](https://huggingface.co/microsoft/BitNet-embedding-270M) - Lightweight 1-bit Embedding Model **BitNet-embedding-270M** is a compact **270M-parameter** 1-bit embedding model designed for resource-constrained environments, offering fast inference with minimal memory footprint. - **1.32x to 1.74x speedup** over F16 on prefill (8 threads, x86) - **Lossless Quality**: Competitive embedding quality with 2 bits per weight - **Lightweight**: Only 270M parameters for edge deployment scenarios [🤗 Hugging Face](https://huggingface.co/microsoft/BitNet-embedding-270M) | [📄 I2_S Guide](docs/bitnet-embeddings-i2s-guide.md) ## Supported Models Model Parameters CPU Kernel I2_S TL1 TL2 Official Models BitNet-b1.58-2B-4T 2.4B x86 ✅ ❌ ✅ ARM ✅ ✅ ❌ BitNet-embedding-0.6B 0.6B x86 ✅ ❌ ❌ ARM ❌ ❌ ❌ BitNet-embedding-270M 270M x86 ✅ ❌ ❌ ARM ❌ ❌ ❌ Community Models bitnet_b1_58-large 0.7B x86 ✅ ❌ ✅ ARM ✅ ✅ ❌ bitnet_b1_58-3B 3.3B x86 ❌ ❌ ✅ ARM ❌ ✅ ❌ Llama3-8B-1.58-100B-tokens 8.0B x86 ✅ ❌ ✅ ARM ✅ ✅ ❌ Falcon3 Family 1B-10B x86 ✅ ❌ ✅ ARM ✅ ✅ ❌ Falcon-E Family 1B-3B x86 ✅ ❌ ✅ ARM ✅ ✅ ❌ ❗️**We use existing 1-bit LLMs available on [Hugging Face](https://huggingface.co/) to demonstrate the inference capabilities of bitnet.cpp. We hope the release of bitnet.cpp will inspire the development of 1-bit LLMs in large-scale settings in terms of model size and training tokens.** ## Installation ### Requirements - python>=3.10 - cmake>=3.22 - clang>=18 - For Windows users, install [Visual Studio 2022](https://visualstudio.microsoft.com/downloads/). In the installer, toggle on at least the following options(this also automatically installs the required additional tools like CMake): - Desktop-development with C++ - C++-CMake Tools for Windows - Git for Windows - C++-Clang Compiler for Windows - MS-Build Support for LLVM-Toolset (clang) - For Debian/Ubuntu users, you can download with [Automatic installation script](https://apt.llvm.org/) `bash -c "$(wget -O - https://apt.llvm.org/llvm.sh)"` - conda (highly recommend) ### Build from source > [!IMPORTANT] > If you are using Windows, please remember to always use a Developer Command Prompt / PowerShell for VS2022 for the following commands. Please refer to the FAQs below if you see any issues. 1. Clone the repo ```bash git clone --recursive https://github.com/microsoft/BitNet.git cd BitNet ``` 2. Install the dependencies ```bash # (Recommended) Create a new conda environment conda create -n bitnet-cpp python=3.10 conda activate bitnet-cpp pip install -r requirements.txt ``` 3. Build the project ```bash # Manually download the model and run with local path huggingface-cli download microsoft/BitNet-b1.58-2B-4T-gguf --local-dir models/BitNet-b1.58-2B-4T python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s ```
usage: setup_env.py [-h] [--hf-repo {1bitLLM/bitnet_b1_58-large,1bitLLM/bitnet_b1_58-3B,HF1BitLLM/Llama3-8B-1.58-100B-tokens,tiiuae/Falcon3-1B-Instruct-1.58bit,tiiuae/Falcon3-3B-Instruct-1.58bit,tiiuae/Falcon3-7B-Instruct-1.58bit,tiiuae/Falcon3-10B-Instruct-1.58bit}] [--model-dir MODEL_DIR] [--log-dir LOG_DIR] [--quant-type {i2_s,tl1}] [--quant-embd]
                    [--use-pretuned]

Setup the environment for running inference

optional arguments:
  -h, --help            show this help message and exit
  --hf-repo {1bitLLM/bitnet_b1_58-large,1bitLLM/bitnet_b1_58-3B,HF1BitLLM/Llama3-8B-1.58-100B-tokens,tiiuae/Falcon3-1B-Instruct-1.58bit,tiiuae/Falcon3-3B-Instruct-1.58bit,tiiuae/Falcon3-7B-Instruct-1.58bit,tiiuae/Falcon3-10B-Instruct-1.58bit}, -hr {1bitLLM/bitnet_b1_58-large,1bitLLM/bitnet_b1_58-3B,HF1BitLLM/Llama3-8B-1.58-100B-tokens,tiiuae/Falcon3-1B-Instruct-1.58bit,tiiuae/Falcon3-3B-Instruct-1.58bit,tiiuae/Falcon3-7B-Instruct-1.58bit,tiiuae/Falcon3-10B-Instruct-1.58bit}
                        Model used for inference
  --model-dir MODEL_DIR, -md MODEL_DIR
                        Directory to save/load the model
  --log-dir LOG_DIR, -ld LOG_DIR
                        Directory to save the logging info
  --quant-type {i2_s,tl1}, -q {i2_s,tl1}
                        Quantization type
  --quant-embd          Quantize the embeddings to f16
  --use-pretuned, -p    Use the pretuned kernel parameters
## Usage ### Basic usage ```bash # Run inference with the quantized model python run_inference.py -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf -p "You are a helpful assistant" -cnv ```
usage: run_inference.py [-h] [-m MODEL] [-n N_PREDICT] -p PROMPT [-t THREADS] [-c CTX_SIZE] [-temp TEMPERATURE] [-cnv]

Run inference

optional arguments:
  -h, --help            show this help message and exit
  -m MODEL, --model MODEL
                        Path to model file
  -n N_PREDICT, --n-predict N_PREDICT
                        Number of tokens to predict when generating text
  -p PROMPT, --prompt PROMPT
                        Prompt to generate text from
  -t THREADS, --threads THREADS
                        Number of threads to use
  -c CTX_SIZE, --ctx-size CTX_SIZE
                        Size of the prompt context
  -temp TEMPERATURE, --temperature TEMPERATURE
                        Temperature, a hyperparameter that controls the randomness of the generated text
  -cnv, --conversation  Whether to enable chat mode or not (for instruct models.)
                        (When this option is turned on, the prompt specified by -p will be used as the system prompt.)
### Demo A demo of bitnet.cpp running a BitNet b1.58 3B model on Apple M2: https://github.com/user-attachments/assets/7f46b736-edec-4828-b809-4be780a3e5b1 ### Benchmark We provide scripts to run the inference benchmark providing a model. ``` … ``` Here's a brief explanation of each argument: - `-m`, `--model`: The path to the model file. This is a required argument

核心特点

  • •Fast CPU Inference: Achieves up to 6.17x speedup on x86 CPUs and 5.07x on ARM CPUs compared to full-precision models.
  • •Energy Efficient: Reduces energy consumption by up to 82.2% on x86 and 70.0% on ARM.
  • •GPU Support: Official GPU inference kernel available for accelerated deployment.
  • •Chat-Ready: Supports conversational mode for interactive use.
  • •1.42x to 2.28x speedup over F16 on prefill (8 threads, x86)
  • •Lossless Quality: Competitive embedding quality with 2 bits per weight
  • •I2_S Kernel: Supports optimized I2_S conversion on x86 CPUs
  • •1.32x to 1.74x speedup over F16 on prefill (8 threads, x86)
  • •Lossless Quality: Competitive embedding quality with 2 bits per weight
  • •Lightweight: Only 270M parameters for edge deployment scenarios

> 标签

C++

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月9日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言
Baike.dev

baike.dev helps you discover great languages, frameworks, databases, DevOps and cloud-native tools.

Quick links

  • Home
  • All tools
  • Trending
  • Open source

About

  • About us
  • Community
  • News

Contribute

Found a great developer tool? Share it with the community.

Submit a tool
© 2026 baike.dev Developer EncyclopediaUpdated daily · Discover great developer tools