microsoft/BitNetPublic

Official inference framework for 1-bit LLMs

AI summary: The official C++ inference framework designed specifically to run ultra-efficient 1.58-bit and 1-bit Large Language Models natively.

Stars
40.4K
+4 today
Forks
3.7K
Watchers
367
Open issues
212
Open PRs
117
Contributors
~23
Commits
110
Branches
7

C++MITCreated Aug 5, 2024Last push 2mo ago+15 stars this week+133 this month

Quick answers

What is BitNet?
The official C++ inference framework designed specifically to run ultra-efficient 1.58-bit and 1-bit Large Language Models natively.
What does BitNet do?
BitNet is an extremely specialized, hyper-efficient inference framework developed by Microsoft to natively execute radically quantized 1-bit and 1.58-bit Large Language Models. Built fundamentally on the architecture of bitnet.cpp, it replaces the floating-point matrix multiplications required by traditional LLMs with integer addition operations, massively dropping hardware requirements. This framework allows massive models, up to 100 billion parameters, to execute smoothly on standard x86 and ARM CPU hardware, achieving up to 6x speedups and slashing energy consumption by over 70%. By making sub-2-bit quantization practical for production inference, BitNet paves the way for running highly capable LLMs directly on resource-constrained edge devices and mobile processors.
Who is BitNet for?
AI researchers, hardware engineers, and systems developers seeking to deploy massive LLMs locally onto edge devices, laptops, or energy-constrained CPU infrastructure.
How do I get started with BitNet?
git clone https://github.com/microsoft/BitNet && cd BitNet && mkdir build && cd build && cmake .. && make
How popular is BitNet on GitHub?
microsoft/BitNet has 40,355 stars and 3,738 forks on GitHub, and gained 15 stars in the last 7 days.
What license does BitNet use?
microsoft/BitNet is released under the MIT license.

Star history

since Oct 13, 2024
020K40KOct 2024Jun 2025Jan 2026Oct 2026
40.4K stars as of Oct 2, 2026. Before Jul 29, 2026, reconstructed from public GitHub event archives (checked against the repository's real star total); since then measured daily.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 1 commit2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 2 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 2 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 2 commits2026-01-21: 0 commits2026-01-22: 2 commits2026-01-23: 0 commits2026-01-24: 1 commit2026-01-25: 1 commit2026-01-26: 0 commits2026-01-27: 2 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 1 commit2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 1 commit2026-03-10: 1 commit2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 1 commit2026-05-25: 1 commit2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 1 commit2026-07-14: 0 commits2026-07-15: 3 commits2026-07-16: 3 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 1 commit2026-07-21: 1 commit2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 1 commit2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 0 commits2026-08-13: 0 commits2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits
28 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    40,355 stars

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    8 trending appearances

What BitNet does

BitNet is an extremely specialized, hyper-efficient inference framework developed by Microsoft to natively execute radically quantized 1-bit and 1.58-bit Large Language Models. Built fundamentally on the architecture of bitnet.cpp, it replaces the floating-point matrix multiplications required by traditional LLMs with integer addition operations, massively dropping hardware requirements. This framework allows massive models, up to 100 billion parameters, to execute smoothly on standard x86 and ARM CPU hardware, achieving up to 6x speedups and slashing energy consumption by over 70%. By making sub-2-bit quantization practical for production inference, BitNet paves the way for running highly capable LLMs directly on resource-constrained edge devices and mobile processors.

AI researchers, hardware engineers, and systems developers seeking to deploy massive LLMs locally onto edge devices, laptops, or energy-constrained CPU infrastructure.

  • Sub-2-Bit Inference Execution: Natively processes 1.58-bit ternary weights flawlessly, replacing costly matrix multiplications with ultra-fast integer additions.
  • Massive CPU Acceleration: Achieves massive multi-factor inference speedups directly on x86 and ARM processors without requiring discrete graphics cards.
  • Extreme Energy Efficiency: Fundamentally alters processing demands to reduce CPU energy consumption by up to 82% compared to standard full-precision inference.
  • Multi-Hardware Targeting: Ships with optimized kernels natively targeting Apple NEON, x86 AVX, and a specialized GPU pipeline for maximum hardware coverage.
  • Lossless Embedding Support: Executes 1-bit embedding models smoothly, offering highly competitive semantic search quality with significantly faster CPU throughput.

Where teams use it

Edge Device AI Deployment

Mobile engineers utilize the framework to run massive, highly capable LLMs completely offline directly on constrained smartphone CPUs.

Energy-Constrained Datacenters

Cloud providers deploy the framework to drastically slash the electrical and cooling overhead required to serve automated LLM API requests at scale.

Rapid Local Embedding

Data scientists leverage the specialized 1-bit embedding models to perform massive semantic vectorizations locally on their laptops at high speed.

Low-Latency Local Chatbots

Developers embed the framework into desktop applications to guarantee near-instantaneous response times for local AI assistants.

Getting started: git clone https://github.com/microsoft/BitNet && cd BitNet && mkdir build && cd build && cmake .. && make

README

main branch

bitnet.cpp

License: MIT version Hugging Face Technical Report Demo GPU Kernel

📰 News

07/23/2026: 📣 We released VibeASR.cpp — a real-time multilingual ASR inference engine on CPU using BitNet I2_S quantization, achieving RTF < 1 with very few threads on x86 (AVX2) and ARM (NEON) platforms. [Code] [Models] [Report] NEW

07/20/2026: 📣 We released BitNet-embedding-0.6B and BitNet-embedding-270M on Hugging Face — the first 1-bit embedding models that deliver competitive embedding quality with significantly faster inference on CPUs.

  • 1.42x to 2.28x speedup over F16 on BitNet-embedding-0.6B prefill (8 threads)
  • 1.32x to 1.74x speedup over F16 on BitNet-embedding-270M prefill (8 threads)
  • Supports I2_S conversion with optimized kernels on x86 CPUs
  • Lossless inference with 2 bits per weight

07/16/2026: 📣 Released BitNet Embeddings 0.6B/270M: I2_S Conversion and Inference Optimization — detailed guide for converting and running BitNet embedding models with optimized I2_S kernels.

01/15/2026: 📣 Released BitNet CPU Inference Optimization — parallel kernel implementations with configurable tiling and embedding quantization support, achieving 1.15x to 2.1x additional speedup over the original implementation.

05/20/2025: 📣 Released BitNet Official GPU inference kernel — extending 1-bit inference beyond CPUs.

04/14/2025: 📣 Released BitNet Official 2B Parameter Model on Hugging Face — the first official BitNet b1.58 model trained with 4T tokens.

02/18/2025: 📑 Bitnet.cpp: Efficient Edge Inference for Ternary LLMs — system-level paper on bitnet.cpp's architecture and design.

11/08/2024: 📑 BitNet a4.8: 4-bit Activations for 1-bit LLMs — enabling 4-bit activations for further efficiency gains.

10/21/2024: 📑 1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs — the technical report behind bitnet.cpp.

10/17/2024: 📣 bitnet.cpp 1.0 released.

03/21/2024: 📑 The-Era-of-1-bit-LLMs: Training Tips, Code, FAQ

02/27/2024: 📑 The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits — the foundational paper introducing BitNet b1.58.

10/17/2023: 📑 BitNet: Scaling 1-bit Transformers for Large Language Models — the original BitNet paper.

Overview

bitnet.cpp is the official inference framework for 1-bit LLMs (e.g., BitNet b1.58). It offers a suite of optimized kernels that support fast and lossless inference of 1.58-bit models on CPU and GPU (NPU support coming next).

Try it out via this online demo, or build and run it on your own CPU or GPU.

bitnet.cpp achieves speedups of 1.37x to 5.07x on ARM CPUs, with larger models experiencing greater performance gains. Additionally, it reduces energy consumption by 55.4% to 70.0%, further boosting overall efficiency. On x86 CPUs, speedups range from 2.37x to 6.17x with energy reductions between 71.9% to 82.2%. Furthermore, bitnet.cpp can run a 100B BitNet b1.58 model on a single CPU, achieving speeds comparable to human reading (5-7 tokens per second), significantly enhancing the potential for running LLMs on local devices. Please refer to the technical report for more details.

performance_comparison

Model Releases

1. BitNet-b1.58-2B-4T - 1-bit Large Language Model

BitNet-b1.58-2B-4T is the first official BitNet b1.58 model with 2.4B parameters, trained on 4 trillion tokens. It is a ternary (1.58-bit) language model that delivers competitive performance with full-precision models of similar size while enabling significantly faster and more energy-efficient inference.

  • Fast CPU Inference: Achieves up to 6.17x speedup on x86 CPUs and 5.07x on ARM CPUs compared to full-precision models.
  • Energy Efficient: Reduces energy consumption by up to 82.2% on x86 and 70.0% on ARM.
  • GPU Support: Official GPU inference kernel available for accelerated deployment.
  • Chat-Ready: Supports conversational mode for interactive use.

🤗 Hugging Face | 🔗 Online Demo | 📄 Technical Report

BitNet b1.58 2B Benchmark

2. BitNet-embedding-0.6B - 1-bit Embedding Model

BitNet-embedding-0.6B is a 0.6B-parameter 1-bit embedding model that achieves competitive embedding quality with significantly faster CPU inference. It is the first model to demonstrate that ternary weights can deliver strong performance on embedding tasks.

  • 1.42x to 2.28x speedup over F16 on prefill (8 threads, x86)
  • Lossless Quality: Competitive embedding quality with 2 bits per weight
  • I2_S Kernel: Supports optimized I2_S conversion on x86 CPUs

🤗 Hugging Face | 📄 I2_S Guide

BitNet Embedding 0.6B Prefill Performance

3. BitNet-embedding-270M - Lightweight 1-bit Embedding Model

BitNet-embedding-270M is a compact 270M-parameter 1-bit embedding model designed for resource-constrained environments, offering fast inference with minimal memory footprint.

  • 1.32x to 1.74x speedup over F16 on prefill (8 threads, x86)
  • Lossless Quality: Competitive embedding quality with 2 bits per weight
  • Lightweight: Only 270M parameters for edge deployment scenarios

🤗 Hugging Face | 📄 I2_S Guide

BitNet Embedding 270M Prefill Performance

Supported Models

Model Parameters CPU Kernel
I2_S TL1 TL2
Official Models
BitNet-b1.58-2B-4T 2.4B x86 ✅ ❌ ✅
ARM ✅ ✅ ❌
BitNet-embedding-0.6B 0.6B x86 ✅ ❌ ❌
ARM ❌ ❌ ❌
BitNet-embedding-270M 270M x86 ✅ ❌ ❌
ARM ❌ ❌ ❌
Community Models
bitnet_b1_58-large 0.7B x86 ✅ ❌ ✅
ARM ✅ ✅ ❌
bitnet_b1_58-3B 3.3B x86 ❌ ❌ ✅
ARM ❌ ✅ ❌
Llama3-8B-1.58-100B-tokens 8.0B x86 ✅ ❌ ✅
ARM ✅ ✅ ❌
Falcon3 Family 1B-10B x86 ✅ ❌ ✅
ARM ✅ ✅ ❌
Falcon-E Family 1B-3B x86 ✅ ❌ ✅
ARM ✅ ✅ ❌

❗️We use existing 1-bit LLMs available on Hugging Face to demonstrate the inference capabilities of bitnet.cpp. We hope the release of bitnet.cpp will inspire the development of 1-bit LLMs in large-scale settings in terms of model size and training tokens.

Installation

Requirements

  • python>=3.10
  • cmake>=3.22
  • clang>=18
    • For Windows users, install Visual Studio 2022. In the installer, toggle on at least the following options(this also automatically installs the required additional tools like CMake):

      • Desktop-development with C++
      • C++-CMake Tools for Windows
      • Git for Windows
      • C++-Clang Compiler for Windows
      • MS-Build Support for LLVM-Toolset (clang)
    • For Debian/Ubuntu users, you can download with Automatic installation script

      bash -c "$(wget -O - https://apt.llvm.org/llvm.sh)"

  • conda (highly recommend)

Build from source

Important

If you are using Windows, please remember to always use a Developer Command Prompt / PowerShell for VS2022 for the following commands. Please refer to the FAQs below if you see any issues.

  1. Clone the repo
git clone --recursive https://github.com/microsoft/BitNet.git
cd BitNet
  1. Install the dependencies
# (Recommended) Create a new conda environment
conda create -n bitnet-cpp python=3.10
conda activate bitnet-cpp

pip install -r requirements.txt
  1. Build the project
# Manually download the model and run with local path
huggingface-cli download microsoft/BitNet-b1.58-2B-4T-gguf --local-dir models/BitNet-b1.58-2B-4T
python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s
usage: setup_env.py [-h] [--hf-repo {1bitLLM/bitnet_b1_58-large,1bitLLM/bitnet_b1_58-3B,HF1BitLLM/Llama3-8B-1.58-100B-tokens,tiiuae/Falcon3-1B-Instruct-1.58bit,tiiuae/Falcon3-3B-Instruct-1.58bit,tiiuae/Falcon3-7B-Instruct-1.58bit,tiiuae/Falcon3-10B-Instruct-1.58bit}] [--model-dir MODEL_DIR] [--log-dir LOG_DIR] [--quant-type {i2_s,tl1}] [--quant-embd]
                    [--use-pretuned]

Setup the environment for running inference

optional arguments:
  -h, --help            show this help message and exit
  --hf-repo {1bitLLM/bitnet_b1_58-large,1bitLLM/bitnet_b1_58-3B,HF1BitLLM/Llama3-8B-1.58-100B-tokens,tiiuae/Falcon3-1B-Instruct-1.58bit,tiiuae/Falcon3-3B-Instruct-1.58bit,tiiuae/Falcon3-7B-Instruct-1.58bit,tiiuae/Falcon3-10B-Instruct-1.58bit}, -hr {1bitLLM/bitnet_b1_58-large,1bitLLM/bitnet_b1_58-3B,HF1BitLLM/Llama3-8B-1.58-100B-tokens,tiiuae/Falcon3-1B-Instruct-1.58bit,tiiuae/Falcon3-3B-Instruct-1.58bit,tiiuae/Falcon3-7B-Instruct-1.58bit,tiiuae/Falcon3-10B-Instruct-1.58bit}
                        Model used for inference
  --model-dir MODEL_DIR, -md MODEL_DIR
                        Directory to save/load the model
  --log-dir LOG_DIR, -ld LOG_DIR
                        Directory to save the logging info
  --quant-type {i2_s,tl1}, -q {i2_s,tl1}
                        Quantization type
  --quant-embd          Quantize the embeddings to f16
  --use-pretuned, -p    Use the pretuned kernel parameters

Usage

Basic usage

# Run inference with the quantized model
python run_inference.py -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf -p "You are a helpful assistant" -cnv
usage: run_inference.py [-h] [-m MODEL] [-n N_PREDICT] -p PROMPT [-t THREADS] [-c CTX_SIZE] [-temp TEMPERATURE] [-cnv]

Run inference

optional arguments:
  -h, --help            show this help message and exit
  -m MODEL, --model MODEL
                        Path to model file
  -n N_PREDICT, --n-predict N_PREDICT
                        Number of tokens to predict when generating text
  -p PROMPT, --prompt PROMPT
                        Prompt to generate text from
  -t THREADS, --threads THREADS
                        Number of threads to use
  -c CTX_SIZE, --ctx-size CTX_SIZE
                        Size of the prompt context
  -temp TEMPERATURE, --temperature TEMPERATURE
                        Temperature, a hyperparameter that controls the randomness of the generated text
  -cnv, --conversation  Whether to enable chat mode or not (for instruct models.)
                        (When this option is turned on, the prompt specified by -p will be used as the system prompt.)

Demo

A demo of bitnet.cpp running a BitNet b1.58 3B model on Apple M2:

demo.mp4

Benchmark

We provide scripts to run the inference benchmark providing a model.

usage: e2e_benchmark.py -m MODEL [-n N_TOKEN] [-p N_PROMPT] [-t THREADS]  
   
Setup the environment for running the inference  
   
required arguments:  
  -m MODEL, --model MODEL  
                        Path to the model file. 
   
optional arguments:  
  -h, --help  
                        Show this help message and exit. 
  -n N_TOKEN, --n-token N_TOKEN  
                        Number of generated tokens. 
  -p N_PROMPT, --n-prompt N_PROMPT  
                        Prompt to generate text from. 
  -t THREADS, --threads THREADS  
                        Number of threads to use. 

Here's a brief explanation of each argument:

  • -m, --model: The path to the model file. This is a required argument that must be provided when running the script.
  • -n, --n-token: The number of tokens to generate during the inference. It is an optional argument with a default value of 128.
  • -p, --n-prompt: The number of prompt tokens to use for generating text. This is an optional argument with a default value of 512.
  • -t, --threads: The number of threads to use for running the inference. It is an optional argument with a default value of 2.
  • -h, --help: Show the help message and exit. Use this argument to display usage information.

For example:

python utils/e2e_benchmark.py -m /path/to/model -n 200 -p 256 -t 4  

This command would run the inference benchmark using the model located at /path/to/model, generating 200 tokens from a 256 token prompt, utilizing 4 threads.

For the model layout that do not supported by any public model, we provide scripts to generate a dummy model with the given model layout, and run the benchmark on your machine:

python utils/generate-dummy-bitnet-model.py models/bitnet_b1_58-large --outfile models/dummy-bitnet-125m.tl1.gguf --outtype tl1 --model-size 125M

# Run benchmark with the generated model, use -m to specify the model path, -p to specify the prompt processed, -n to specify the number of token to generate
python utils/e2e_benchmark.py -m models/dummy-bitnet-125m.tl1.gguf -p 512 -n 128

Convert from .safetensors Checkpoints

# Prepare the .safetensors model file
huggingface-cli download microsoft/bitnet-b1.58-2B-4T-bf16 --local-dir ./models/bitnet-b1.58-2B-4T-bf16

# Convert to gguf model
python ./utils/convert-helper-bitnet.py ./models/bitnet-b1.58-2B-4T-bf16

Acknowledgements

This project is based on the llama.cpp framework. We would like to thank all the authors for their contributions to the open-source community. Also, bitnet.cpp's kernels are built on top of the Lookup Table methodologies pioneered in T-MAC. For inference of general low-bit LLMs beyond ternary models, we recommend using T-MAC.

FAQ (Frequently Asked Questions)📌

Q1: The build dies with errors building llama.cpp due to issues with std::chrono in log.cpp?

A: This is an issue introduced in recent version of llama.cpp. Please refer to this commit in the discussion to fix this issue.

Q2: How to build with clang in conda environment on windows?

A: Before building the project, verify your clang installation and access to Visual Studio tools by running:

clang -v

This command checks that you are using the correct version of clang and that the Visual Studio tools are available. If you see an error message such as:

'clang' is not recognized as an internal or external command, operable program or batch file.

It indicates that your command line window is not properly initialized for Visual Studio tools.

• If you are using Command Prompt, run:

"C:\Program Files\Microsoft Visual Studio\2022\Professional\Common7\Tools\VsDevCmd.bat" -startdir=none -arch=x64 -host_arch=x64

• If you are using Windows PowerShell, run the following commands:

Import-Module "C:\Program Files\Microsoft Visual Studio\2022\Professional\Common7\Tools\Microsoft.VisualStudio.DevShell.dll" Enter-VsDevShell 3f0e31ad -SkipAutomaticLocation -DevCmdArguments "-arch=x64 -host_arch=x64"

These steps will initialize your environment and allow you to use the correct Visual Studio tools.

View on GitHub

Recent activity

commits and pull requests

Code frequency

additions and deletions
+3K-3KWeek of 2025-09-28: +0 linesWeek of 2025-09-28: -0 linesWeek of 2025-10-05: +0 linesWeek of 2025-10-05: -0 linesWeek of 2025-10-12: +0 linesWeek of 2025-10-12: -0 linesWeek of 2025-10-19: +0 linesWeek of 2025-10-19: -0 linesWeek of 2025-10-26: +0 linesWeek of 2025-10-26: -0 linesWeek of 2025-11-02: +0 linesWeek of 2025-11-02: -0 linesWeek of 2025-11-09: +0 linesWeek of 2025-11-09: -0 linesWeek of 2025-11-16: +940 linesWeek of 2025-11-16: -224 linesWeek of 2025-11-23: +0 linesWeek of 2025-11-23: -0 linesWeek of 2025-11-30: +0 linesWeek of 2025-11-30: -0 linesWeek of 2025-12-07: +0 linesWeek of 2025-12-07: -0 linesWeek of 2025-12-14: +0 linesWeek of 2025-12-14: -0 linesWeek of 2025-12-21: +1,044 linesWeek of 2025-12-21: -34 linesWeek of 2025-12-28: +0 linesWeek of 2025-12-28: -0 linesWeek of 2026-01-04: +0 linesWeek of 2026-01-04: -0 linesWeek of 2026-01-11: +219 linesWeek of 2026-01-11: -3 linesWeek of 2026-01-18: +2,986 linesWeek of 2026-01-18: -1,978 linesWeek of 2026-01-25: +163 linesWeek of 2026-01-25: -12 linesWeek of 2026-02-01: +20 linesWeek of 2026-02-01: -221 linesWeek of 2026-02-08: +0 linesWeek of 2026-02-08: -0 linesWeek of 2026-02-15: +0 linesWeek of 2026-02-15: -0 linesWeek of 2026-02-22: +0 linesWeek of 2026-02-22: -0 linesWeek of 2026-03-01: +0 linesWeek of 2026-03-01: -0 linesWeek of 2026-03-08: +4 linesWeek of 2026-03-08: -4 linesWeek of 2026-03-15: +0 linesWeek of 2026-03-15: -0 linesWeek of 2026-03-22: +0 linesWeek of 2026-03-22: -0 linesWeek of 2026-03-29: +0 linesWeek of 2026-03-29: -0 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +0 linesWeek of 2026-04-19: -0 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +1,412 linesWeek of 2026-05-24: -349 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +0 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +0 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +1,948 linesWeek of 2026-07-12: -476 linesWeek of 2026-07-19: +136 linesWeek of 2026-07-19: -51 linesWeek of 2026-07-26: +3 linesWeek of 2026-07-26: -1 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesWeek of 2026-08-09: +0 linesWeek of 2026-08-09: -0 linesWeek of 2026-08-16: +0 linesWeek of 2026-08-16: -0 linesWeek of 2026-08-23: +0 linesWeek of 2026-08-23: -0 linesWeek of 2026-08-30: +0 linesWeek of 2026-08-30: -0 linesWeek of 2026-09-06: +0 linesWeek of 2026-09-06: -0 linesWeek of 2026-09-13: +0 linesWeek of 2026-09-13: -0 linesWeek of 2026-09-20: +0 linesWeek of 2026-09-20: -0 linesSep 28, 2025Sep 20, 2026
+8.9K lines added, -3.4K removed over the last year.

Commits per week

last 52 weeks
70Week of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 1 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 2 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 2 commitsWeek of 2026-01-18: 5 commitsWeek of 2026-01-25: 3 commitsWeek of 2026-02-01: 1 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 2 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 2 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 7 commitsWeek of 2026-07-19: 2 commitsWeek of 2026-07-26: 1 commitsWeek of 2026-08-02: 0 commitsWeek of 2026-08-09: 0 commitsWeek of 2026-08-16: 0 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 0 commitsWeek of 2026-09-06: 0 commitsWeek of 2026-09-13: 0 commitsWeek of 2026-09-20: 0 commitsSep 28, 2025Sep 20, 2026
28 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 1 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 1 commitsSun 22:00 — 0 commitsSun 23:00 — 1 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 1 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 1 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 0 commitsMon 11:00 — 0 commitsMon 12:00 — 2 commitsMon 13:00 — 1 commitsMon 14:00 — 1 commitsMon 15:00 — 2 commitsMon 16:00 — 2 commitsMon 17:00 — 0 commitsMon 18:00 — 0 commitsMon 19:00 — 0 commitsMon 20:00 — 0 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 0 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 2 commitsTue 4:00 — 1 commitsTue 5:00 — 1 commitsTue 6:00 — 1 commitsTue 7:00 — 3 commitsTue 8:00 — 1 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 1 commitsTue 12:00 — 1 commitsTue 13:00 — 0 commitsTue 14:00 — 2 commitsTue 15:00 — 2 commitsTue 16:00 — 0 commitsTue 17:00 — 3 commitsTue 18:00 — 1 commitsTue 19:00 — 0 commitsTue 20:00 — 0 commitsTue 21:00 — 2 commitsTue 22:00 — 0 commitsTue 23:00 — 0 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 2 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 1 commitsWed 8:00 — 1 commitsWed 9:00 — 1 commitsWed 10:00 — 0 commitsWed 11:00 — 2 commitsWed 12:00 — 0 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 0 commitsWed 17:00 — 2 commitsWed 18:00 — 2 commitsWed 19:00 — 0 commitsWed 20:00 — 1 commitsWed 21:00 — 2 commitsWed 22:00 — 0 commitsWed 23:00 — 0 commitsThu 0:00 — 0 commitsThu 1:00 — 1 commitsThu 2:00 — 0 commitsThu 3:00 — 2 commitsThu 4:00 — 0 commitsThu 5:00 — 2 commitsThu 6:00 — 1 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 1 commitsThu 10:00 — 2 commitsThu 11:00 — 1 commitsThu 12:00 — 1 commitsThu 13:00 — 0 commitsThu 14:00 — 2 commitsThu 15:00 — 1 commitsThu 16:00 — 1 commitsThu 17:00 — 1 commitsThu 18:00 — 4 commitsThu 19:00 — 0 commitsThu 20:00 — 1 commitsThu 21:00 — 2 commitsThu 22:00 — 0 commitsThu 23:00 — 1 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 0 commitsFri 10:00 — 1 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 1 commitsFri 14:00 — 2 commitsFri 15:00 — 0 commitsFri 16:00 — 2 commitsFri 17:00 — 1 commitsFri 18:00 — 0 commitsFri 19:00 — 0 commitsFri 20:00 — 0 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 1 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 1 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 1 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 1 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Mar 14, 2026daily#17+233
Mar 13, 2026daily#11+292
Mar 12, 2026daily#9+337
Mar 11, 2026daily#14+183
Feb 1, 2026daily#17+141
Jan 31, 2026daily#23+134
Jan 30, 2026daily#20+137
Jan 29, 2026daily#18+143