lyogavin/airllmPublic

AirLLM 70B inference with single 4GB GPU

AI summary: Run massive language models like 70B on a single 4GB GPU using layer-by-layer inference.

Stars
29.8K
+778 today
Forks
3.2K
Watchers
269
Open issues
98
Open PRs
35
Contributors
~10
Commits
311
Branches
4

Jupyter NotebookApache-2.0Created Jun 12, 2023Last push 1d agoLatest release v3.1.0+5.5K stars this week+5.6K this month

Star history

since Jul 28, 2024
010K20KJul 2024Mar 2025Nov 2025Aug 2026
29.8K stars as of Aug 7, 2026, tracked back to Jul 28, 2024. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 2 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 1 commit2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 5 commits2026-06-19: 5 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 3 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 3 commits2026-07-04: 2 commits2026-07-05: 1 commit2026-07-06: 1 commit2026-07-07: 1 commit2026-07-08: 1 commit2026-07-09: 1 commit2026-07-10: 1 commit2026-07-11: 1 commit2026-07-12: 1 commit2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 1 commit2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 2 commits2026-07-21: 1 commit2026-07-22: 1 commit2026-07-23: 1 commit2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 4 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 2 commits2026-08-06: 1 commit2026-08-07: 0 commits2026-08-08: 0 commits
42 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    29,821 stars

  • Rising fast

    +5,514 stars this week

  • Actively maintained

    Pushed within 48 hours

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

What airllm does

AirLLM is an inference engine designed to drastically reduce the VRAM requirements for running large language models locally. It achieves this by loading and executing the model one layer at a time, avoiding the need to hold the entire model in memory simultaneously. This allows a 70B model to run on a 4GB GPU without relying on quantization, distillation, or pruning. It uniquely supports massive sparse MoE models like DeepSeek-V3 by streaming individual experts dynamically.

AirLLM is for AI researchers, hobbyists, and developers who lack access to high-end enterprise GPUs but want to run the largest open-source models. Familiarity with Python and basic machine learning concepts is beneficial.

  • Layer-wise execution: Processes neural network layers sequentially to keep VRAM usage incredibly low.
  • Zero quantization required: Maintains full model precision without needing to compress or alter weights.
  • Massive model support: Capable of running 405B Llama 3.1 on 8GB and Kimi K3 on under 4GB.
  • Sparse MoE streaming: Efficiently handles Mixture of Experts architectures by only loading the necessary experts.
  • Mac and PC compatibility: Runs on standard consumer GPUs and Apple Silicon with minimal configuration.

Where teams use it

Local inference on budget hardware

Enables students and hobbyists to run state-of-the-art 70B+ models on older or low-end gaming GPUs.

Evaluating large open models

Allows researchers to test the full-precision outputs of massive models without renting expensive cloud compute.

Privacy-preserving local AI

Facilitates running powerful models completely offline to ensure sensitive data never leaves the local machine.

Prototyping MoE architectures

Provides a low-cost environment for developers to experiment with and deploy Mixture of Experts models.

Getting started: pip install airllm

README

main branch

airllm_logo

Quickstart | Configurations | MacOS | Example notebooks | FAQ

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T) — the largest open-source model released to date — on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.

GitHub Repo stars Downloads

Code License Generic badge Discord PyPI - AirLLM Website Website Support me on Patreon GitHub Sponsors

AI Agents Recommendation:

Updates

[2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on one RTX 6000 Ada. Per-expert streaming loads only the experts a token actually routes to. K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformers 4.56.x, as its remote code does not load on 5.x.

[2026/06] v3.0: FP8 model support + the latest models. Run DeepSeek-V3 (671B) on ~12GB and Qwen3-235B on ~3GB, plus Qwen3, Llama 3.x/4, DeepSeek V2/V3, Phi-4, Gemma and more — all through a single AutoModel.

[2024/08/20] v2.11.0: Support Qwen2.5

[2024/08/18] v2.10.1 Support CPU inference. Support non sharded models. Thanks @NavodPeiris for the great work!

[2024/07/30] Support Llama3.1 405B (example notebook). Support 8bit/4bit quantization.

[2024/04/20] AirLLM supports Llama3 natively already. Run Llama3 70B on 4GB single GPU.

[2023/12/25] v2.8.2: Support MacOS running 70B large language models.

[2023/12/20] v2.7: Support AirLLMMixtral.

[2023/12/20] v2.6: Added AutoModel, automatically detect model type, no need to provide model class to initialize model.

[2023/12/18] v2.5: added prefetching to overlap the model loading and compute. 10% speed improvement.

[2023/12/03] added support of ChatGLM, QWen, Baichuan, Mistral, InternLM!

[2023/12/02] added support for safetensors. Now support all top 10 models in open llm leaderboard.

[2023/12/01] airllm 2.0. Support compressions: 3x run time speed up!

[2023/11/20] airllm Initial version!

Star History

Star History Chart

Table of Contents

Quickstart

1. Install package

First, install the airllm pip package.

pip install airllm

2. Inference

Then, initialize AirLLMLlama2, pass in the huggingface repo ID of the model being used, or the local path, and inference can be performed similar to a regular transformer model.

(You can also specify the path to save the splitted layered model through layer_shards_saving_path when init AirLLMLlama2.

from airllm import AutoModel

MAX_LENGTH = 128
# just pass a hugging face repo id — works with almost any popular model:
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")

# go bigger with the exact same one line:
#model = AutoModel.from_pretrained("Qwen/Qwen3-235B-A22B")     # 235B, runs in ~3GB
#model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-V3")  # 671B, runs in ~12GB

# or use a model's local path...
#model = AutoModel.from_pretrained("/home/ubuntu/.cache/huggingface/hub/models--Qwen--Qwen3-32B/snapshots/...")

input_text = [
        'What is the capital of United States?',
        #'I like',
    ]

input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=False)
           
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=20,
    use_cache=True,
    return_dict_in_generate=True)

output = model.tokenizer.decode(generation_output.sequences[0])

print(output)

Note: During inference, the original model will first be decomposed and saved layer-wise. Please ensure there is sufficient disk space in the huggingface cache directory.

Model Compression - 3x Inference Speed Up!

We just added model compression based on block-wise quantization-based model compression. Which can further speed up the inference speed for up to 3x , with almost ignorable accuracy loss! (see more performance evaluation and why we use block-wise quantization in this paper)

speed_improvement

How to enable model compression speed up:

  • Step 1. make sure you have bitsandbytes installed by pip install -U bitsandbytes
  • Step 2. make sure airllm verion later than 2.0.0: pip install -U airllm
  • Step 3. when initialize the model, passing the argument compression ('4bit' or '8bit'):
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
                     compression='4bit' # specify '8bit' for 8-bit block-wise quantization 
                    )

What are the differences between model compression and quantization?

Quantization normally needs to quantize both weights and activations to really speed things up. Which makes it harder to maintain accuracy and avoid the impact of outliers in all kinds of inputs.

While in our case the bottleneck is mainly at the disk loading, we only need to make the model loading size smaller. So, we get to only quantize the weights' part, which is easier to ensure the accuracy.

Configurations

When initialize the model, we support the following configurations:

  • compression: supported options: 4bit, 8bit for 4-bit or 8-bit block-wise quantization, or by default None for no compression
  • profiling_mode: supported options: True to output time consumptions or by default False
  • layer_shards_saving_path: optionally another path to save the splitted model
  • hf_token: huggingface token can be provided here if downloading gated models like: meta-llama/Llama-2-7b-hf
  • prefetching: prefetching to overlap the model loading and compute. By default, turned on. For now, only AirLLMLlama2 supports this.
  • delete_original: if you don't have too much disk space, you can set delete_original to true to delete the original downloaded hugging face model, only keep the transformed one to save half of the disk space.

MacOS

Just install airllm and run the code the same as on linux. See more in Quick Start.

  • make sure you installed mlx and torch
  • you probably need to install python native see more here
  • only Apple silicon is supported

Example [python notebook] (https://github.com/lyogavin/airllm/blob/main/air_llm/examples/run_on_macos.ipynb)

Example Python Notebook

Example colabs here:

Open In Colab

example of other models (ChatGLM, QWen, Baichuan, Mistral, etc):

Details
  • ChatGLM:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("THUDM/chatglm3-6b-base")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=True)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache= True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
  • QWen:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen-7B")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache=True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
  • Baichuan, InternLM, Mistral, etc:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("baichuan-inc/Baichuan2-7B-Base")
#model = AutoModel.from_pretrained("internlm/internlm-20b")
#model = AutoModel.from_pretrained("mistralai/Mistral-7B-Instruct-v0.1")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache=True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])

To request other model support: here

Supported Models

AirLLM works out of the box with virtually every popular open LLM — just pass its Hugging Face ID to AutoModel.from_pretrained(...). That covers all the major families:

Llama (2 / 3 / 3.1 / 3.3 / 4) · Qwen (1 / 2 / 2.5 / 3, including MoE and FP8) · DeepSeek (V2 / V3 / R1) · Mistral & Mixtral · Phi · Gemma · ChatGLM · Baichuan · InternLM · Yi — and most new models the day they're released.

Tiny GPU, huge models

The trick: AirLLM only ever keeps one layer on the GPU at a time, so the VRAM you need depends on the model's layer size — not its total size. That's how a 671B model fits on a hobbyist card:

Model Size GPU VRAM
Qwen3 / Mistral / Phi (≈8B) 8B ~1–2 GB
Qwen3-30B / Mixtral (MoE) 30–47B ~1–3 GB
Qwen3-235B (MoE) 235B ~3 GB
Llama 3.x 70B (full precision) 70B ~4 GB
Llama 3.1 405B 405B ~8 GB
DeepSeek-V3 671B ~12 GB

Same one line of code for all of them — no special setup.

Acknowledgement

A lot of the code are based on SimJeg's great work in the Kaggle exam competition. Big shoutout to SimJeg:

GitHub account @SimJeg, the code on Kaggle, the associated discussion.

FAQ

1. MetadataIncompleteBuffer

safetensors_rust.SafetensorError: Error while deserializing header: MetadataIncompleteBuffer

If you run into this error, most possible cause is you run out of disk space. The process of splitting model is very disk-consuming. See this. You may need to extend your disk space, clear huggingface .cache and rerun.

2. ValueError: max() arg is an empty sequence

Most likely you are loading QWen or ChatGLM model with Llama2 class. Try the following:

For QWen model:

from airllm import AutoModel #<----- instead of AirLLMLlama2
AutoModel.from_pretrained(...)

For ChatGLM model:

from airllm import AutoModel #<----- instead of AirLLMLlama2
AutoModel.from_pretrained(...)

3. 401 Client Error....Repo model ... is gated.

Some models are gated models, needs huggingface api token. You can provide hf_token:

model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-hf", #hf_token='HF_API_TOKEN')

4. ValueError: Asking to pad but the tokenizer does not have a padding token.

Some model's tokenizer doesn't have padding token, so you can set a padding token or simply turn the padding config off:

input_tokens = model.tokenizer(input_text,
   return_tensors="pt", 
   return_attention_mask=False, 
   truncation=True, 
   max_length=MAX_LENGTH, 
   padding=False  #<-----------   turn off padding 
)

Citing AirLLM

If you find AirLLM useful in your research and wish to cite it, please use the following BibTex entry:

@software{airllm2023,
  author = {Gavin Li},
  title = {AirLLM: scaling large language models on low-end commodity computers},
  url = {https://github.com/lyogavin/airllm/},
  version = {0.0},
  year = {2023},
}

Sponsors

Bloome — Run AI Agent Teams in the Cloud

Run AI Agent Teams in the Cloud — Bloome

Bloome is an AI-agent IM platform: build and run AI agent teams in the cloud with zero setup. Add a skill as an agent in a group chat, run it in one click from web or mobile, and share it with your team — think of it as a group chat where your AI assistants are teammates you can @mention and assign tasks to.

👉 Try Bloome

Contribution

Welcomed contributions, ideas and discussions!

If you find it useful, please ⭐ or buy me a coffee! 🙏

"Buy Me A Coffee"

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

3 total
  1. ## Kimi K3 (2.8T) runs on a single card in 3.72GB Kimi K3 is the largest open-source model released to date. AirLLM runs it on one consumer-class GPU. Measured end to end on a single **RTX 6000 Ada (48GB)** against the full 1.56TB checkpoint, generating real tokens: | | | |---|---| | **Peak VRAM during generation** | **3.72 GB** | | Peak VRAM after init | 0.83 GB | | Init (one-time per process) | 900 s | | Generation | 292 s/token, disk-bound | The reason a 2.8T model needs less VRAM than a 671B one is that sparse MoE checkpoints stream **one expert at a time** rather than a whole layer. K3 holds 896 experts per layer and routes each token to 16 of them — expanded, a layer's experts are ~55GB, but a token only needs ~1GB. AirLLM loads just those. MXFP4 weights also cross PCIe packed and expand on the GPU, moving 4x less data. Fitting the checkpoint on disk needed the same kind of trick: a naive split would want 3.12TB for a 1.56TB model. K3's shards turn out to be pure, one module each, so split layers are hardlinked to the originals instead of copied. ### Before you run K3 K3 brings three requirements of its own, none of them optional: ```bash pip install airllm compressed

  2. AirLLM v3.0.1v3.0.1Jun 30, 2026

    ## AirLLM v3.0.1 Patch release. - **Fix:** `import airllm` failed on a clean install with `ModuleNotFoundError: No module named 'sentencepiece'`. `sentencepiece` is now a declared dependency, so `pip install airllm` works out of the box. - Model-family imports are now defensive: a missing optional dependency for one niche model family no longer breaks the whole package — the generic streaming path always loads. ### Upgrade ```bash pip install --upgrade airllm ```

  3. v3.0.0v3.0.0Jun 30, 2026

    ## AirLLM v3.0.0 Big update: AirLLM now runs today's largest open models on tiny GPUs, with full support for the latest model families and Hugging Face versions — still no quantization, distillation, or pruning required. ### Highlights - **Run the biggest open models on a single small GPU.** Stream 70B models on 4GB, 405B Llama 3.1 on 8GB, and even **DeepSeek-V3 (671B) on ~12GB**. - **Native FP8 support.** Pre-quantized FP8 (block-FP8) checkpoints now load and run correctly — including DeepSeek-V3 and the Qwen3-FP8 family. - **Latest models supported**, including Qwen3 (dense + MoE, e.g. Qwen3-32B, Qwen3-30B-A3B, Qwen3-235B-A22B-FP8), DeepSeek-V3, Phi-4, Mixtral-8x7B, and DeepSeek-V2-Lite. - **Up to date with modern Hugging Face.** Works with current `transformers` / `accelerate` releases, so a plain `pip install airllm` just works — no manual dependency juggling. ### Improvements & fixes - Reworked layer streaming to build on the standard Transformers model path for better model compatibility and `generate()` behavior. - Runtime precision now follows each model's native dtype (e.g. bfloat16) instead of being forced to fp16, fixing garbled output on very deep models.

Code frequency

additions and deletions
+767-767Week of 2025-08-10: +0 linesWeek of 2025-08-10: -0 linesWeek of 2025-08-17: +0 linesWeek of 2025-08-17: -0 linesWeek of 2025-08-24: +0 linesWeek of 2025-08-24: -0 linesWeek of 2025-08-31: +7 linesWeek of 2025-08-31: -2 linesWeek of 2025-09-07: +0 linesWeek of 2025-09-07: -0 linesWeek of 2025-09-14: +0 linesWeek of 2025-09-14: -0 linesWeek of 2025-09-21: +0 linesWeek of 2025-09-21: -0 linesWeek of 2025-09-28: +0 linesWeek of 2025-09-28: -0 linesWeek of 2025-10-05: +0 linesWeek of 2025-10-05: -0 linesWeek of 2025-10-12: +0 linesWeek of 2025-10-12: -0 linesWeek of 2025-10-19: +0 linesWeek of 2025-10-19: -0 linesWeek of 2025-10-26: +0 linesWeek of 2025-10-26: -0 linesWeek of 2025-11-02: +0 linesWeek of 2025-11-02: -0 linesWeek of 2025-11-09: +0 linesWeek of 2025-11-09: -0 linesWeek of 2025-11-16: +0 linesWeek of 2025-11-16: -0 linesWeek of 2025-11-23: +0 linesWeek of 2025-11-23: -0 linesWeek of 2025-11-30: +0 linesWeek of 2025-11-30: -0 linesWeek of 2025-12-07: +0 linesWeek of 2025-12-07: -0 linesWeek of 2025-12-14: +0 linesWeek of 2025-12-14: -0 linesWeek of 2025-12-21: +0 linesWeek of 2025-12-21: -0 linesWeek of 2025-12-28: +0 linesWeek of 2025-12-28: -0 linesWeek of 2026-01-04: +0 linesWeek of 2026-01-04: -0 linesWeek of 2026-01-11: +0 linesWeek of 2026-01-11: -0 linesWeek of 2026-01-18: +0 linesWeek of 2026-01-18: -0 linesWeek of 2026-01-25: +0 linesWeek of 2026-01-25: -0 linesWeek of 2026-02-01: +0 linesWeek of 2026-02-01: -0 linesWeek of 2026-02-08: +0 linesWeek of 2026-02-08: -0 linesWeek of 2026-02-15: +0 linesWeek of 2026-02-15: -0 linesWeek of 2026-02-22: +0 linesWeek of 2026-02-22: -0 linesWeek of 2026-03-01: +0 linesWeek of 2026-03-01: -0 linesWeek of 2026-03-08: +74 linesWeek of 2026-03-08: -0 linesWeek of 2026-03-15: +0 linesWeek of 2026-03-15: -0 linesWeek of 2026-03-22: +0 linesWeek of 2026-03-22: -0 linesWeek of 2026-03-29: +0 linesWeek of 2026-03-29: -0 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +0 linesWeek of 2026-04-19: -0 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +515 linesWeek of 2026-06-14: -588 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +335 linesWeek of 2026-06-28: -384 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +3 linesWeek of 2026-07-12: -3 linesWeek of 2026-07-19: +31 linesWeek of 2026-07-19: -3 linesWeek of 2026-07-26: +767 linesWeek of 2026-07-26: -24 linesWeek of 2026-08-02: +128 linesWeek of 2026-08-02: -50 linesAug 10, 2025Aug 2, 2026
+1.9K lines added, -1.1K removed over the last year.

Commits per week

last 52 weeks
100Week of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 2 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 1 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 10 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 8 commitsWeek of 2026-07-05: 7 commitsWeek of 2026-07-12: 2 commitsWeek of 2026-07-19: 5 commitsWeek of 2026-07-26: 4 commitsWeek of 2026-08-02: 3 commitsAug 10, 2025Aug 2, 2026
42 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 1 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 2 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 5 commitsSun 12:00 — 1 commitsSun 13:00 — 0 commitsSun 14:00 — 2 commitsSun 15:00 — 1 commitsSun 16:00 — 0 commitsSun 17:00 — 2 commitsSun 18:00 — 2 commitsSun 19:00 — 0 commitsSun 20:00 — 2 commitsSun 21:00 — 1 commitsSun 22:00 — 0 commitsSun 23:00 — 1 commitsMon 0:00 — 3 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 1 commitsMon 8:00 — 1 commitsMon 9:00 — 0 commitsMon 10:00 — 1 commitsMon 11:00 — 5 commitsMon 12:00 — 1 commitsMon 13:00 — 3 commitsMon 14:00 — 0 commitsMon 15:00 — 7 commitsMon 16:00 — 13 commitsMon 17:00 — 10 commitsMon 18:00 — 0 commitsMon 19:00 — 1 commitsMon 20:00 — 2 commitsMon 21:00 — 2 commitsMon 22:00 — 2 commitsMon 23:00 — 2 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 1 commitsTue 7:00 — 0 commitsTue 8:00 — 2 commitsTue 9:00 — 2 commitsTue 10:00 — 0 commitsTue 11:00 — 8 commitsTue 12:00 — 5 commitsTue 13:00 — 2 commitsTue 14:00 — 2 commitsTue 15:00 — 0 commitsTue 16:00 — 2 commitsTue 17:00 — 0 commitsTue 18:00 — 3 commitsTue 19:00 — 0 commitsTue 20:00 — 1 commitsTue 21:00 — 0 commitsTue 22:00 — 8 commitsTue 23:00 — 2 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 2 commitsWed 4:00 — 0 commitsWed 5:00 — 1 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 7 commitsWed 9:00 — 4 commitsWed 10:00 — 6 commitsWed 11:00 — 3 commitsWed 12:00 — 0 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 0 commitsWed 17:00 — 2 commitsWed 18:00 — 0 commitsWed 19:00 — 0 commitsWed 20:00 — 1 commitsWed 21:00 — 0 commitsWed 22:00 — 5 commitsWed 23:00 — 0 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 3 commitsThu 9:00 — 6 commitsThu 10:00 — 0 commitsThu 11:00 — 1 commitsThu 12:00 — 0 commitsThu 13:00 — 0 commitsThu 14:00 — 1 commitsThu 15:00 — 5 commitsThu 16:00 — 1 commitsThu 17:00 — 1 commitsThu 18:00 — 5 commitsThu 19:00 — 2 commitsThu 20:00 — 5 commitsThu 21:00 — 4 commitsThu 22:00 — 1 commitsThu 23:00 — 1 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 2 commitsFri 8:00 — 4 commitsFri 9:00 — 3 commitsFri 10:00 — 1 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 3 commitsFri 15:00 — 4 commitsFri 16:00 — 13 commitsFri 17:00 — 11 commitsFri 18:00 — 5 commitsFri 19:00 — 0 commitsFri 20:00 — 0 commitsFri 21:00 — 16 commitsFri 22:00 — 9 commitsFri 23:00 — 0 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 1 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 1 commitsSat 8:00 — 7 commitsSat 9:00 — 2 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 1 commitsSat 17:00 — 4 commitsSat 18:00 — 4 commitsSat 19:00 — 2 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 5 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits28 (64%)
Community commits16 (36%)

44 commits in total over the last year.

DateListRankStars gained
Jan 22, 2026daily#24+147
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    238.5K stars · JavaScript

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    234.7K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    227K stars · Python

  • Significant-Gravitas/AutoGPT

    AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

    186.2K stars · Python

  • f/prompts.chat

    f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.

    166.9K stars · HTML