Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
< Back to tools
N

neural-speed

> 数据库
Open source

An innovative library for efficient LLM inference via low-bit quantization

352 stars0 likes0 views
WebsiteGitHub

About

An innovative library for efficient LLM inference via low-bit quantization

PROJECT NOT UNDER ACTIVE MANAGEMENT

This project will no longer be maintained by Intel.

Intel has ceased development and contributions including, but not limited to, maintenance, bug fixes, new releases, or updates, to this project.

Intel no longer accepts patches to this project.

Please refer to https://github.com/intel/intel-extension-for-pytorch as an alternative

Neural Speed

Neural Speed is an innovative library designed to support the efficient inference of large language models (LLMs) on Intel platforms through the state-of-the-art (SOTA) low-bit quantization powered by Intel Neural Compressor. The work is inspired by llama.cpp and further optimized for Intel platforms with our innovations in NeurIPS' 2023

Key Features

  • Highly optimized kernels on CPUs with ISAs (AMX, VNNI, AVX512F, AVX_VNNI and AVX2) for N-bit weight (int1, int2, int3, int4, int5, int6, int7 and int8). See details
  • Up to 40x performance speedup on popular LLMs compared with llama.cpp. See details
  • Tensor parallelism across sockets/nodes on CPUs. See details

Neural Speed is under active development so APIs are subject to change.

Supported Hardware

Hardware Supported
Intel Xeon Scalable Processors ✔
Intel Xeon CPU Max Series ✔
Intel Core Processors ✔

Supported Models

Support almost all the LLMs in PyTorch format from Hugging Face such as Llama2, ChatGLM2, Baichuan2, Qwen, Mistral, Whisper, etc. File an issue if your favorite LLM does not work.

Support typical LLMs in GGUF format such as Llama2, Falcon, MPT, Bloom etc. More are coming. Check out the details.

Installation

Install from binary

bash
pip install -r requirements.txt
pip install neural-speed

Build from Source

bash
pip install .

Note: GCC requires version 10+

Quick Start (Transformer-like usage)

Install Intel Extension for Transformers to use Transformer-like APIs.

PyTorch Model from Hugging Face

…

GGUF Model from Hugging Face

…

PyTorch Model from Modelscope

…

As an Inference Backend in Neural Chat Server

Neural Speed can be used in Neural Chat Server of Intel Extension for Transformers. You can choose to enable it by adding use_neural_speed: true in config.yaml.

  • add optimization key section to use Neural Speed and its RTN quantization (example).
yaml
device: "cpu"

# itrex int4 llm runtime optimization
optimization:
    use_neural_speed: true
    optimization_type: "weight_only"
    compute_dtype: "fp32"
    weight_dtype: "int4"
  • add key use_neural_speed and key use_gptq to use Neural Speed and load GPT-Q model (example).
yaml
device: "cpu"
use_neural_speed: true
use_gptq: true

More details please refer to Neural Chat.

Quick Start (llama.cpp-like usage)

Single (One-click) Step

bash
python scripts/run.py model-path --weight_dtype int4 -p "She opened the door and see"

Multiple Steps

Convert and Quantize

bash
# skip the step if GGUF model is from Hugging Face or generated by llama.cpp
python scripts/convert.py --outtype f32 --outfile ne-f32.bin EleutherAI/gpt-j-6b
# Using the quantize script requires a binary installation of Neural Speed
mkdir build&&cd build
cmake ..&&make -j
cd ..
python scripts/quantize.py  --model_name gptj --model_file ne-f32.bin  --out_file ne-q4_j.bin  --build_dir ./build --weight_dtype int4 --alg sym

Inference

bash
# Linux and WSL
OMP_NUM_THREADS= numactl -m 0 -C 0- python scripts/inference.py --model_name llama -m ne-q4_j.bin -c 512 -b 1024 -n 256 -t  --color -p "She opened the door and see"
bash
# Windows
python scripts/inference.py --model_name llama -m ne-q4_j.bin -c 512 -b 1024 -n 256 -t  --color -p "She opened the door and see"

Please refer to Advanced Usage for more details.

Advanced Topics

New model enabling

You can consider adding your own models, please follow the document: graph developer document.

Performance profiling

Enable NEURAL_SPEED_VERBOSE environment variable for performance profiling.

Available modes:

  • 0: Print full information: evaluation time and operator profiling. Need to set NS_PROFILING to ON and recompile.
  • 1: Print evaluation time. Time taken for each evaluation.
  • 2: Profile individual operator. Identify performance bottleneck within the model. Need to set NS_PROFILING to ON and recompile.

Issues· 0 open

View all issuesOpen on GitHub

No open issues yet, or sync has not completed.

> Tags

C++cpufp4fp8gaudi2

No comments yet. Be the first to share.

> Details

PublishedAug 1, 2026
UpdatedSep 17, 2026
Category数据库
PricingOpen source

> Related tools

P
PostgreSQL
功能强大的开源关系型数据库
R
Redis
内存数据结构存储,常用作缓存与队列
M
MySQL
广泛使用的开源关系型数据库