JustVugg/colibriPublic

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

AI summary: A minimal C engine for running frontier Mixture of Experts models locally.

Stars
23K
+333 today
Forks
2.5K
Watchers
198
Open issues
51
Open PRs
37
Contributors
~87
Commits
1K
Branches
12

CApache-2.0Created Jul 1, 2026Last push 1d agoLatest release v1.4.0+1.6K stars this week+2.5K this month

Star history

since Jul 5, 2026
010K20KJul 2026Jul 2026Jul 2026Aug 2026
23K stars as of Aug 7, 2026, tracked back to Jul 5, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 1 commit2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 6 commits2026-07-06: 16 commits2026-07-07: 4 commits2026-07-08: 2 commits2026-07-09: 6 commits2026-07-10: 16 commits2026-07-11: 7 commits2026-07-12: 24 commits2026-07-13: 16 commits2026-07-14: 60 commits2026-07-15: 58 commits2026-07-16: 52 commits2026-07-17: 32 commits2026-07-18: 32 commits2026-07-19: 61 commits2026-07-20: 35 commits2026-07-21: 34 commits2026-07-22: 47 commits2026-07-23: 26 commits2026-07-24: 16 commits2026-07-25: 18 commits2026-07-26: 15 commits2026-07-27: 17 commits2026-07-28: 24 commits2026-07-29: 24 commits2026-07-30: 17 commits2026-07-31: 9 commits2026-08-01: 12 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
687 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    23,027 stars

  • Breakout launch

    23,027 stars in 37 days

  • Rising fast

    +1,571 stars this week

  • Very active

    687 commits in 52 weeks

  • Outside contributions

    73% of recent commits from the community

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    9 trending appearances

What colibri does

Colibrì is an ultra-lightweight inference engine written in pure C with zero external dependencies, designed specifically for running massive Mixture of Experts (MoE) models on consumer hardware. It achieves this by aggressively streaming experts directly from disk rather than loading the entire model into RAM. This approach drastically reduces the memory footprint, allowing models that would normally require data center hardware to run on everyday machines. It prioritizes simplicity and broad compatibility.

Developers and AI enthusiasts who want to run state-of-the-art MoE models locally on limited hardware. It appeals to those favoring minimal, dependency-free implementations.

  • Pure C implementation: Built entirely in C with zero external dependencies for maximum portability.
  • Disk-streamed experts: Loads only necessary parts of the MoE model into RAM, minimizing memory usage.
  • Minimal footprint: A tiny execution engine that doesn't bloat your system.
  • Hardware accessibility: Enables running massive models on consumer-grade hardware you already own.
  • High performance: Optimized for efficient execution despite relying on disk streaming.

Where teams use it

Local MoE inference

Run advanced Mixture of Experts models locally without requiring specialized AI hardware.

Edge computing

Deploy complex models on resource-constrained devices at the edge.

Offline AI access

Utilize powerful frontier models securely in completely offline environments.

Embedded systems integration

Embed advanced AI capabilities into systems with strict memory limitations.

README

main branch

colibrì — tiny engine, immense model

Website Latest release

Website · Discord · English · 简体中文 · 繁體中文 · Italiano

Tiny engine, immense model. Run frontier MoE models — 744B to 2.8T parameters — on consumer and heterogeneous hardware, in pure C with zero engine dependencies, by treating storage, RAM, and VRAM as one inference hierarchy.

Four families run today: GLM-5.2 (744B), Inkling (975B), Kimi K3 (2.8T) and OLMoE (7B) — one C file each, the same coli chat / coli serve / coli web front end. Full roster ↓

Colibrì is an inference engine you can run today, and an open research platform. Its primary goal is to pursue inference-side performance across the entire software/hardware boundary — model formats, memory hierarchy, storage I/O, placement, scheduling, kernels, speculation, and CPU/GPU overlap — so large models depend less on scarce hardware and cost less to run.

Colibrì treats VRAM, RAM, and storage as one managed memory hierarchy, and it is deliberately a place to test aggressive systems ideas — so there is no SLA on speed, and a hard guarantee on semantics: experiments must earn their place through reproducible end-to-end measurements, and the default policy never silently changes model precision or router semantics. Insufficient fast memory may reduce speed; it must not quietly redefine the model.

$ ./coli chat
  🐦 colibri v1.4.0 — GLM-5.2 · 744B MoE · int4 · streaming CPU
  ✓ ready in 32s · resident 9.9 GB
  › ciao!
  ◆ Ciao! 😊 Come posso aiutarti oggi?

See it running

colibrì web dashboard — live metrics, hardware panel, expert tiers

The web dashboard (./coli web): a 744B model at 4 tok/s, TTFT 1.6 s, disk 0 — full expert residency on 6× RTX 5090, with live token metrics, the per-turn time breakdown, the VRAM/RAM/disk tier bar and the live mini-brain in the corner.

the Brain page — 19,456 experts as a live cortex

The Brain page: all 19,456 experts as a living cortex — colour is the storage tier, brightness is routing heat, and every expert routed in a turn flashes white. Hovering shows the expert's measured topic affinity.

the Atlas page — the measured expert atlas as a 3-D galaxy

The Atlas page: the measured expert atlas as a 3-D galaxy — 13,260 characterised experts, 1,041 replicated specialists clustering by topic (poetry, law, Chinese, SQL…). Position is measured routing affinity, not a learned embedding. Drag to spin.

The research mission

Frontier inference should not require datacenter-class hardware by default. Colibrì's research target is simple: reduce the hardware dependency and total cost of inference by optimizing every part of the inference path that evidence shows is limiting it.

That includes changing how weights are represented and moved, deciding what lives in VRAM, RAM, or storage, overlapping heterogeneous compute, reducing launch and synchronization overhead, exploiting sparsity and reuse, and testing new decoding algorithms. Nothing is protected merely because it is conventional; nothing is adopted merely because a microbenchmark looks fast. The deciding result is end-to-end inference on real machines, with correctness and quality measured alongside throughput, latency, memory, and cost.

The practical consequence is accessibility: run a 744B-parameter model on hardware you already own, watch every expert fire in real time, and change the code that does it. Not renting intelligence behind an API — holding it: probing it, measuring it, improving it. The engine is deliberately small enough that the next useful optimization can come from anyone willing to measure it.

Core techniques and measured findings

  • One hierarchy, not one memory threshold. VRAM, RAM, and NVMe are placement tiers for the same weights; limited fast memory changes speed, not model semantics.
  • A JIT for weights. Measured routing heat drives a per-layer LRU, a learned pinned hot-store, and one-layer-ahead prefetch instead of loading every expert. It wins on repeatable workloads; history can overfit, and lookahead can lose on some hosts, so both remain measurable policies rather than promises.
  • I/O is part of the engine. Batched expert unions, overlapped reads and compute, O_DIRECT, and weighted dual-SSD striping attack the streaming path rather than pretending storage latency is free. O_DIRECT is drive-dependent, and dual-SSD still needs broader end-to-end community A/Bs.
  • Heterogeneous execution. CPU, CUDA, Metal, NUMA memory, and partial or full expert residency share one runtime and can be combined according to the machine; the profitable combination depends on compute, bandwidth, residency, and workload.
  • Compressed state without a different model. Token-exact forward validation, 57× smaller MLA KV state, persistent warm conversations, and faithful DSA keep optimization tied to correctness. These are memory, latency, and correctness properties — not a blanket throughput claim.
  • Speculation that must earn its keep. Native MTP and grammar-forced drafts are measured end to end and can be disabled when acceptance does not repay verification.

Open hypotheses, experiments, and how to help

Colibrì treats an optimization as a hypothesis until a controlled end-to-end A/B shows otherwise. These are the main questions now:

hypothesis evidence so far experiment still needed
Routing history can place experts better than plain LRU learned pins improve repeated workloads, but can overfit a prompt held-out, cross-session A/Bs across coding, chat, multilingual, and long-context workloads
Multiple SSDs can turn independent bandwidth into decode speed weighted mirror/split routing is implemented and validated; the bandwidth model is sound cold-cache one-drive vs two-drive GLM-5.2 runs on real, independent controllers
A hardware-aware planner can approach each machine's best configuration automatically RAM/VRAM budgets and several backends are detected today compare the generated plan with a controlled parameter sweep across laptops, workstations, NUMA hosts, and multi-GPU systems
Lossless or quality-bounded representations can reduce weight movement enough to matter format and quantization ablations exist, with correctness/quality gates reproduce quality, bytes moved, latency, and cost per useful token together — not compression ratio alone
Routing-aware speculation can pay before near-full residency MTP and grammar drafts work, but MTP has also measured a 32% loss around 85% expert hit map the break-even surface across acceptance, expert hit rate, batch union, and draft depth
CPU/GPU overlap can hide transfer and synchronization rather than merely move the bottleneck CUDA and Metal wins exist, but fast CPUs and low residency can erase them per-stage profiles and one-variable A/Bs across PCIe, unified-memory, and full-resident machines

Want to help? Pick one row and publish the negative results too. Record the hardware, commit, model/container, exact command, prompt, cache state, throughput, TTFT, expert hit rate, bytes read, and quality check; change one variable, repeat the run, and attach raw logs. Start with CONTRIBUTING.md, compare against the benchmark protocol, then open an experiment issue. A well-controlled failure is more valuable here than an unexplained fast number.

The idea

A 744B Mixture-of-Experts model activates only ~40B parameters per token — and only ~11 GB of those change from token to token (the routed experts):

only ~5.4% of parameters are active per token

So the model doesn't need to fit in fast memory — it needs to be placed:

  • the dense part (attention, shared experts, embeddings — ~17B params) stays resident in RAM at int4 (~9.9 GB);
  • the 19,456 routed experts (75 MoE layers × 256 + the MTP head, ~19 MB each at int4) live on disk (~370 GB) and are streamed on demand, with a per-layer LRU cache, a learned pinned hot-store, and an optional VRAM tier.

Think of the core algorithm as a JIT, but for weights. A compiler JIT never compiles the whole program — it watches what actually runs and compiles the hot paths, just in time. colibrì makes the same bet about a 744B parameter space: parameters are not resident state to be held, they are data to be staged across a heterogeneous storage hierarchy (VRAM / RAM / NVMe), exactly when the router proves they are needed. Measured routing heat decides which experts earn which tier, the router runs a layer ahead so prefetch hides the staging latency, and — like a JIT — the engine learns your workload: the more you run, the hotter the right experts get. It works because routing has measurable structure (see the expert atlas) — and structure is cacheable.

The engine is a single C file (c/glm.c) plus small headers. No BLAS, no Python at runtime, no GPU required.

How it works

The per-token path

route → union → place → overlap → learn

Every layer of every token walks the same five steps. The design goal is that placement only ever decides speed — the router's decisions and the weights' precision are the same whether an expert answered from VRAM or from disk.

One memory hierarchy instead of one memory requirement

VRAM / RAM / NVMe three-tier expert residency

Dual-SSD: two copies of the model, twice the read bandwidth

Decode is disk-bound on most machines, and expert reads are read-only — so if you have a second SSD, put a full copy of the model on it and let the engine stream from both drives at once:

COLI_MODEL=/fast/glm52_i4 COLI_MODEL_MIRROR=/second/glm52_i4 ./coli chat
COLI_DISK_WEIGHTS=9,3 ...   # optional: primary,mirror bandwidth ratio (else measured at startup)

Each expert is routed to one drive by a deterministic hash, weighted by the two drives' measured (or declared) bandwidth, so readahead/PILOT prefetch and the demand read always hit the same drive and nothing is cached twice. The aggregate bandwidth is the sum of both drives — a 9 GB/s + 3 GB/s pair reads experts ~33% faster than the fast drive alone, and the OMP-parallel pin/warmup load streams from both. Details worth knowing:

  • the mirror is validated at startup (per-file size + safetensors header must be byte-identical to the primary); divergent or missing files silently stay on the primary, so a partial mirror is fine — a smaller second SSD holding only some shards still helps;
  • the mirror is never written: .coli_usage, .coli_kv and all sidecars stay on the primary;
  • a read error on the mirror falls back to the primary (one warning, no crash), so unplugging the second drive mid-run degrades instead of killing the server;
  • routing never changes tokens — both copies are byte-identical, and the per-run MIRROR: stats line shows GB served per drive.

The same engine spans the whole range: on a 25 GB laptop everything streams from disk (slow but correct); on a large host the entire expert set becomes resident (CUDA_EXPERT_GB=auto PIN_GB=all) and disk drops out of the decode path entirely. Between the tiers sits a learning cache: the engine records which experts your workload routes to (.coli_usage, updated every turn) and pins the hottest ones automatically — colibrì literally gets faster the more you use it. On multi-socket hosts, COLI_NUMA=1 interleaves the resident weights across memory controllers (#82).

For a second drive that cannot hold the whole model, Colibri can rank a partial mirror from the expert history it already learns. Run a few representative prompts first so .coli_usage reflects the workload, then plan, stage, and verify the mirror:

./c/coli mirror plan  --model /fast/glm52_i4 --mirror /second/glm52_i4 \
  --budget-gib 200 --reserve-gib 20
./c/coli mirror stage --model /fast/glm52_i4 --mirror /second/glm52_i4 \
  --budget-gib 200 --reserve-gib 20
./c/coli mirror verify --model /fast/glm52_i4 --mirror /second/glm52_i4

The planner reads safetensors headers directly, follows split-model directories from COLI_MODEL_DIRS, and prioritizes shards that can serve the hottest routed experts. Staging never changes the primary model: it copies through temporary files, preserves the requested free-space reserve, verifies every shard with SHA-256, never deletes an existing mirror shard, and atomically publishes a receipt only after the selected mirror is ready.

Never wait for the disk twice

Misses are expensive, so the engine spends most of its cleverness avoiding and overlapping them: each expert's three matrices are stored adjacent and read in one pread; a bounded async I/O pool (PIPE=1, default) loads missing experts while resident ones compute; batched positions read each unique expert once (batch-union); and a router-lookahead thread (PILOT=1) prefetches the next layer's experts — routing is measurably 71.6% predictable one layer ahead. On GPUs, the resident pipeline (COLI_CUDA_PIPE=2) keeps the residual stream on-device across layers so the CPU expert loop runs uninterrupted; on Apple Silicon an experimental Metal backend does the batched expert math on the unified-memory GPU; and a Vulkan backend brings the expert tier, dense projections, and the MLA attention core to any GPU with a Vulkan 1.2 driver — including AMD cards via Mesa/RADV (the only backend for cards the vendor stacks no longer support, like the RX 580, and competitive with ROCm on RDNA4 — see the benchmarking notes).

On real NVMe, measure DIRECT=1. O_DIRECT bypasses the page cache and is often a large win on drives with DRAM cache and bandwidth headroom (+34% decode measured with PIPE=1 on a Blackwell/Windows box; 4.25→9.69 GB/s in iobench on a GB10) — but it is drive-dependent: QLC/DRAM-less or virtualised disks can be neutral to negative. Try it first; keep what your hardware rewards.

Faithful model, compressed state

The forward pass is validated against a transformers oracle (teacher-forcing typically 30-32/32; two tiny-oracle positions are floating-point near-ties and toolchain-dependent). MLA attention stores a compressed KV state — 576 floats/token instead of 32,768 (57× smaller) — and persists it across restarts (.coli_kv): conversations reopen warm with zero re-prefill, byte-identical to an uninterrupted session. DSA sparse attention (GLM-5.2's lightning indexer) is implemented faithfully and validated by forcing full-key selection to reproduce dense attention exactly.

Speculative decoding, honestly

GLM-5.2's native MTP head drafts tokens that the main model verifies in one batched forward — 2.2–2.8 tokens/forward when it pays. Two hard-won rules ship as defaults: the MTP head must be int8 (int4 heads collapse to 0–4% acceptance, #8), and draft and verify must compute the same functionSPEC_PIN=1 pins both to one kernel family (#163 is the full forensic story). Grammar-forced drafts (GRAMMAR=file.gbnf) add ~free acceptance on constrained JSON output. Whether speculation is a net win depends on your cache temperature — measure, and use DRAFT=0 when it doesn't pay.

What it achieves

measured decode speed by hardware class

Same engine, same int4 container — the hardware only changes where the experts live. Highlights from the full benchmark tables:

  • 6× RTX 5090, full residency: 5.8–6.8 tok/s decode, TTFT ~13 s (experiment log);
  • 128 GB CPU-only desktop: ~1.8 tok/s warm (#200);
  • single RTX 5070 Ti laptop-class box: 1.07 tok/s via the GPU-resident pipeline (#273);
  • 25 GB dev box: 0.05–0.1 tok/s cold — the proven floor where this project started, and still the honest baseline.

Quality is measured, not assumed: the int4 container's quantization cost and the scale-granularity/rotation ablations live in docs/benchmarks.md and #108/#81.

Get started

You need two things: the program (a few hundred KB) and the model (372 GB). Step-by-step for every platform in the Quick Start guide.

1. Get colibri

Download a prebuilt release — Linux, macOS and Windows, no compiler needed. Take the archive for your platform from Releases and unpack it:

mkdir colibri && tar xzf colibri-v1.1.0-linux-x86_64.tar.gz -C colibri && cd colibri
python3 coli info                         # engine ready ✓

Inside you get the engine (colibri, colibri.exe on Windows), the coli launcher and its Python helpers. Nothing to rename or configure — coli finds the engine next to itself. You only need Python 3 installed: the launcher and the API gateway are Python scripts, while the engine itself is pure C with zero dependencies.

Or build from source — needs gcc (or clang) with OpenMP:

git clone https://github.com/JustVugg/colibri && cd colibri/c
./setup.sh                                # checks gcc/OpenMP, builds, self-tests

Want coli on your PATH? From a checkout, pip install -e . registers it (the engine still lives in c/ — an editable install from the clone, not a wheel).

2. Get the model

A pre-converted GLM-5.2 int4 container is on Hugging Face — use the group-scaled (gs64) build with the int8 MTP head. It is about 372 GB, so put it on a disk with the room, ideally a fast one:

https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp

⚠️ Use the gs64 container above, not the older per-row int4 mirrors (mateogrgic/…, jlnsrk/…): those measure ~9pp worse on quality and are the root cause of the original think-mode loops and never-terminating generations in #455. The gs64 container fixed those controlled per-row A/Bs, but it is not a general repetition or EOS-starvation guard. The MTP head must also be int8, not int4 (int4 → 0% draft acceptance, #8): ls -l <model>/out-mtp-* — int8 (correct) is 3527131672 / 5366238584 / 1065950496.

Or convert from the FP8 source yourself — one resumable command that never needs the full 756 GB on disk at once:

./coli convert --model /nvme/glm52_i4     # download+convert shard by shard (python, one-time)

Other supported models

GLM-5.2 is the reference model, but the same streaming approach runs three more families. Each is a sibling engine — one C file, its own architecture, the same coli chat / coli serve / coli web front end (the launcher picks the binary from the model's config.json):

Family Total / active Weights Build Docs
GLM-5.2 744B / 40B mastouri/…-int4-g64-with-int8-mtp (372 GB) make -C c glm this page
Inkling (Thinking Machines) 975B / 41B nbeerbower/Inkling-colibri-int4 (469 GB) make -C c inkling inkling.md
Kimi K3 (Moonshot) 2.8T / 104B moonshotai/Kimi-K3 — original checkpoint, routed experts stay native MXFP4 make -C c kimi_k3 kimi_k3.md
OLMoE (AI2) 7B / 1B converted with c/tools/convert_olmoe_merged.py make -C c olmoe

Kimi K3 needs no conversion: its QAT-trained MXFP4 experts are streamed straight from the original Hugging Face shards, and the bf16 dense set is quantized at load time. Inkling ships int4 experts but bf16 dense weights (49.4 GB resident); on a host that cannot hold those, inkling.md has a one-pass tool that brings the dense set to 15.3 GB and lets the 975B run on a 25 GB box — with the honest trade-off written down.

3. Run it

COLI_MODEL=/nvme/glm52_i4 ./coli chat     # RAM budget, cache and MTP auto-detected
COLI_MODEL=/nvme/glm52_i4 ./coli plan     # inspect the planned VRAM/RAM/disk placement
COLI_MODEL=/nvme/glm52_i4 ./coli doctor   # read-only readiness check
COLI_MODEL=/nvme/glm52_i4 ./coli doctor --deep  # strict tensors/shards/index/mirror preflight
COLI_MODEL=/nvme/glm52_i4 ./coli tune     # measure and save this machine's fastest safe execution profile
./coli web  --model /nvme/glm52_i4        # API + web dashboard on one port
./coli serve --model /nvme/glm52_i4       # OpenAI-compatible API only

On Windows the same commands work with python coli chat --model D:\glm52_i4. The engine at runtime is pure C — python is only used by the one-time converter and the optional API gateway.

The same commands run any of the models

coli reads the model's config.json, picks the matching engine binary, and renders that family's chat template — so nothing about the command line changes between models. Build the engine you want once, then just point COLI_MODEL at the right directory:

make -C c glm                                     # GLM-5.2
make -C c inkling                                 # Inkling
make -C c kimi_k3                                 # Kimi K3

COLI_MODEL=/nvme/glm52_i4      ./coli chat        # TUI
COLI_MODEL=/nvme/inkling_i4    ./coli chat
COLI_MODEL=/nvme/kimi_k3       ./coli chat

./coli web --model /nvme/inkling_i4               # API + dashboard, same port
./coli web --model /nvme/kimi_k3
./coli serve --model /nvme/inkling_i4             # API only

For the non-GLM engines coli chat starts the gateway locally and attaches the TUI to it, so the TUI, the API and the dashboard all go through the same arch-aware chat template — you never have to pass the template yourself.

Two things that differ per model, both documented in the per-model page:

  • Inkling on a RAM-tight host needs the int4 dense container and a small expert cache: ./coli chat --model /nvme/inkling_i4 --cap 2 (see inkling.md — the default --cap 8 wants ~14 GB of cache on top of the resident set).
  • Kimi K3 streams its MXFP4 experts from the original checkpoint, so there is nothing to convert — but the snapshot is ~1.6 TB (see kimi_k3.md).

4. Go deeper

topic doc
Benchmarks, community datapoints, quality measurements docs/benchmarks.md
Tuning knobs, policies, the learning cache, prefetch docs/tuning.md
Windows 11 native build (+ CUDA DLL) docs/windows.md
CUDA backend, VRAM expert tier, full residency docs/cuda.md
Vulkan backend (any GPU: AMD via RADV, incl. cards ROCm dropped) docs/vulkan.md
Apple Silicon Metal backend docs/metal.md
OpenAI-compatible API, KV slots, web dashboard docs/api.md
Grammar-forced drafts (structured output) docs/grammar-draft.md
Environment variable inventory docs/ENVIRONMENT.md

What's next

  • Inference-systems research is the product. The current hierarchy is LRU + a learned pin set; active work spans model formats, compression, placement, scheduling, I/O, CPU/GPU kernels, heterogeneous overlap, KV state, and routing-aware speculation. The objective is lower hardware requirements and lower cost per useful token. Everything lands the way this project works: measured end to end, reviewed, and developed in the open.
  • More open models. The tiering algorithm is model-agnostic: any MoE with routed experts can be staged the same way. GLM-5.2 and OLMoE run today; support for more open-weight families — Kimi K2 (Moonshot AI), Qwen3 MoE (Alibaba), MiniMax — is on the roadmap.

Supporting the project

colibrì started as a one-person project on a 12-core laptop with 25 GB of RAM; today its numbers come from a community of real machines. If it's useful to you:

  • ⭐ star the repo and share it;
  • 🐛 open issues with benchmark numbers from your hardware — datapoints move this project more than anything else;
  • 💬 join the Discord community to discuss experiments, hardware results, and research directions;
  • 💬 reach out via GitHub issues to sponsor development or donate hardware.

Repo layout

Makefile                  root build/check entry point
c/
├── glm.c                 single-file GLM engine
├── st.h, tok.h, json.h   runtime headers
├── backend_cuda.*        optional CUDA tier
├── Makefile              build and local checks
├── coli                  user-facing CLI
├── openai_server.py      OpenAI-compatible HTTP gateway
├── setup.sh              one-command local setup
├── tools/                offline conversion, fixtures and benchmarks
├── scripts/              long-running conversion helpers
└── tests/                dependency-free C and Python tests
web/                      browser UI (pure OpenAI-API client)
desktop/                  Tauri v2 desktop shell wrapping the web UI
docs/                     reference docs, experiments, media

The runtime path intentionally stays flat and readable: glm.c plus its small headers. From the repository root, make, make check, and make clean delegate to the engine Makefile.

Why "colibrì"

The hummingbird weighs a few grams, hovers in place, and visits a thousand flowers a day. This engine keeps a 744-billion-parameter giant alive on hummingbird rations: 25 GB of RAM, twelve CPU cores, and a lot of disk patience.

Acknowledgements

colibrì is an engine; the minds it runs are a gift. Thank you to the teams releasing frontier-class weights in the open — Z.ai (GLM), Moonshot AI (Kimi), Alibaba Qwen, MiniMax, and Allen AI (OLMoE) — and to every contributor who benchmarked, bisected, replicated an atlas run, or sent a patch. This project is proof of what open weights make possible.

License

Apache 2.0. GLM-5.2 weights are released by Z.ai under MIT.

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Discussions

all 37

Releases and announcements

6 total
  1. **Two things change what a colibrì release is: a third GPU backend, and — for the first time — archives that actually contain the models the front page promises.** 162 commits, 29 pull requests, 25 of them from contributors. --- ## 🔴 If you downloaded v1.3.0 to run Kimi K3 or Inkling, it could not work Every archive until now contained `c/colibri` alone. Point it at Kimi K3 and there was no engine — while the README front page promised four model families. Nothing you could have typed would have helped (#720). The cause was worse than an oversight in packaging: **`inkling` and `kimi_k3` did not compile on Windows at all.** Their serve loops polled stdin with `fd_set`/`select`, which msys2/UCRT64 does not provide in that form. Measured across the release matrix: Linux ✅, macOS ✅, Windows ❌. Three fixes, in order: - **#736** — the portable version now lives once in `compat.h`, carrying the two bug fixes `colibri.c` had already absorbed: #139 (`select()` on a pipe handle routes to winsock and always returns `SOCKET_ERROR`) and #195 (anonymous pipes are not waitable objects; `PeekNamedPipe` fails on file/console handles). Copying that a third and fourth time would have reintrod

  2. colibri v1.3.0v1.3.0Jul 29, 20261.1K downloads

    **Three MoE families now run on the same engine, from 744B to 2.8T parameters — and a 975B model answers on a 25 GB machine.** Pure C, zero engine dependencies, storage/RAM/VRAM treated as one inference hierarchy. --- ## New: two more model families ### Kimi K3 — 2.8T total / 104B active (#676) A sibling engine (`c/kimi_k3.c`) that streams Moonshot's **QAT-trained MXFP4 routed experts straight from the original Hugging Face shards** — never re-encoded, never converted — and quantizes the bf16 dense set at load time. Includes the KDA + gated-NoPE-MLA + AttnRes + LatentMoE architecture, an MXFP4 matmul kernel (scalar, AVX2, and int8-activation), chunked prefill (bit-identical, 2.6×), parallel `O_DIRECT` expert reads, and `--chat` for K3's XTML format. > **Preview status, stated plainly:** the engine and tokenizer are validated — the multi-turn wire matches Moonshot's official `encoding_k3.py` at **77/77 token ids** — and the regression suite passes. Full-model generation is still gated on a host with ~1.6 TB of storage for the snapshot; **nobody has run it end to end yet.** If you have that hardware, this one is yours to try. ### Inkling — 975B total / 41B active, now on a sma

  3. colibri v1.2.0v1.2.0Jul 28, 2026871 downloads

    Pure-C MoE engine that streams experts from disk — run a 744B model on a ~25 GB box. Highlights since v1.1.1: ### Hardware & correctness - Fix OOM on integrated / unified-memory GPUs — **GB10 / DGX Spark** (#653) - Recognize **AMD/ROCm** in `doctor` and `resource_plan` — a HIP build is no longer reported CPU-only (#662, #663) - Chat stop-set fix (#633 / #381); async packed-int4 parity (#632); stable KV slot per conversation (#634, #639) ### Performance - **AVX2 `matmul_e8`** — fmt=6 (E8/IQ3) was 92% of decode on the scalar kernel (#654) - Native **SIMD fmt=6 encoder**, 15× over numpy, byte-identical (#655) - **AVX-512 `matmul_i3`** for fmt=5 int3-g64 (#661) ### Web & docs - Web **dashboard** with the live Expert Atlas (#641); new landing page (#645) - README / quickstart point to the **gs64** model container (#642) Verify your download against `SHA256SUMS.txt`.

  4. colibri v1.1.1v1.1.1Jul 22, 20263.7K downloads

    A same-day patch release. **Windows users on v1.1.0 should upgrade**: Microsoft Defender flags the v1.1.0 Windows binary, and the cause was ours. ### Fixed - **107 KB of zeros were shipped inside every binary** (#527, #532) — and that is what antivirus ML heuristics were reacting to. `static GrDraft g_grd={.max=24};` looks harmless, but `GrDraft` is ~107 KB (the grammar's 1024 static rules plus the PDA walker) and **any** initializer moves the whole struct out of `.bss` and into `.data`, writing 106,848 bytes of near-zero-entropy data into the file — in a *writable* section, which is the classic shape of an unpacking buffer for a packed payload. Section forensics against v1.0.0 (clean on the same Defender definitions) isolated it: identical toolchain, identical PE layout, `.data` 1,840 → 108,752 bytes. A Windows build with the fix scans clean where v1.1.0 does not. Every platform's binary also gets smaller: the Linux engine drops 474,904 → 368,016 bytes, **-22.5%**. - **`python3 openai_server.py` was broken on a clean checkout** (#526) — the gateway still looked for an engine named `glm` after the #391 rename. It resolves `colibri`/`colibri.exe` first now

  5. colibri v1.1.0v1.1.0Jul 22, 2026640 downloads

    A community release. 27 pull requests from more than 20 contributors, 216 commits since v1.0.0. Most of what follows was found, measured, or fixed by people who do not work on this project and had nothing to gain from it. ### Added - **AMD GPU support (HIP/ROCm)** (#339) — single-source `backend_gpu_compat.h` with a WMMA dispatch gate, so one codebase builds for CUDA and HIP. Validated on an RX 9070 XT (RDNA4, ROCm 7.2): token-exact against CPU on a real fmt=4 gs64 container, with resident dense *and* with routed experts in VRAM, plus a fail-injection control proving the GPU actually executed the work. - **Dual-SSD streaming** (`COLI_MODEL_MIRROR`, #421) — read the model from two drives at once, roughly doubling streaming bandwidth on a disk-bound host. - **N-drive shard split** (`COLI_MODEL_DIRS`, #469) — capacity aggregation: run a container no single drive can hold, spread across several with no duplication. - **fmt=5 (int3-g64)** (#168) — 3-bit weights with per-64 group scales: measured 3.3x lower outlier-row error than per-row int4 at 25% fewer bytes. - **fmt=6 (E8/IQ3 lattice)** (#465) — CPU decode kernel and dispatch; index codec tooling (#458). - **`tools/t

Code frequency

additions and deletions
+36.5K-36.5KWeek of 2026-06-28: +202 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +13,871 linesWeek of 2026-07-05: -2,463 linesWeek of 2026-07-12: +36,477 linesWeek of 2026-07-12: -3,912 linesWeek of 2026-07-19: +31,362 linesWeek of 2026-07-19: -9,676 linesWeek of 2026-07-26: +17,232 linesWeek of 2026-07-26: -2,891 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesJun 28, 2026Aug 2, 2026
+99.1K lines added, -18.9K removed over the last year.

Commits per week

last 52 weeks
2740Week of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 1 commitsWeek of 2026-07-05: 57 commitsWeek of 2026-07-12: 274 commitsWeek of 2026-07-19: 237 commitsWeek of 2026-07-26: 118 commitsWeek of 2026-08-02: 0 commitsAug 10, 2025Aug 2, 2026
687 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 5 commitsSun 1:00 — 3 commitsSun 2:00 — 3 commitsSun 3:00 — 2 commitsSun 4:00 — 4 commitsSun 5:00 — 1 commitsSun 6:00 — 1 commitsSun 7:00 — 3 commitsSun 8:00 — 1 commitsSun 9:00 — 2 commitsSun 10:00 — 10 commitsSun 11:00 — 3 commitsSun 12:00 — 6 commitsSun 13:00 — 9 commitsSun 14:00 — 7 commitsSun 15:00 — 4 commitsSun 16:00 — 3 commitsSun 17:00 — 3 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 5 commitsSun 21:00 — 7 commitsSun 22:00 — 3 commitsSun 23:00 — 21 commitsMon 0:00 — 12 commitsMon 1:00 — 2 commitsMon 2:00 — 1 commitsMon 3:00 — 2 commitsMon 4:00 — 2 commitsMon 5:00 — 1 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 1 commitsMon 9:00 — 1 commitsMon 10:00 — 0 commitsMon 11:00 — 1 commitsMon 12:00 — 3 commitsMon 13:00 — 2 commitsMon 14:00 — 2 commitsMon 15:00 — 3 commitsMon 16:00 — 2 commitsMon 17:00 — 10 commitsMon 18:00 — 11 commitsMon 19:00 — 7 commitsMon 20:00 — 5 commitsMon 21:00 — 5 commitsMon 22:00 — 2 commitsMon 23:00 — 9 commitsTue 0:00 — 8 commitsTue 1:00 — 12 commitsTue 2:00 — 4 commitsTue 3:00 — 4 commitsTue 4:00 — 2 commitsTue 5:00 — 0 commitsTue 6:00 — 1 commitsTue 7:00 — 5 commitsTue 8:00 — 5 commitsTue 9:00 — 6 commitsTue 10:00 — 3 commitsTue 11:00 — 2 commitsTue 12:00 — 5 commitsTue 13:00 — 7 commitsTue 14:00 — 7 commitsTue 15:00 — 8 commitsTue 16:00 — 1 commitsTue 17:00 — 2 commitsTue 18:00 — 7 commitsTue 19:00 — 5 commitsTue 20:00 — 6 commitsTue 21:00 — 2 commitsTue 22:00 — 11 commitsTue 23:00 — 9 commitsWed 0:00 — 8 commitsWed 1:00 — 12 commitsWed 2:00 — 10 commitsWed 3:00 — 3 commitsWed 4:00 — 2 commitsWed 5:00 — 1 commitsWed 6:00 — 2 commitsWed 7:00 — 3 commitsWed 8:00 — 6 commitsWed 9:00 — 4 commitsWed 10:00 — 1 commitsWed 11:00 — 1 commitsWed 12:00 — 5 commitsWed 13:00 — 6 commitsWed 14:00 — 7 commitsWed 15:00 — 6 commitsWed 16:00 — 6 commitsWed 17:00 — 2 commitsWed 18:00 — 4 commitsWed 19:00 — 8 commitsWed 20:00 — 7 commitsWed 21:00 — 4 commitsWed 22:00 — 11 commitsWed 23:00 — 13 commitsThu 0:00 — 10 commitsThu 1:00 — 8 commitsThu 2:00 — 2 commitsThu 3:00 — 1 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 2 commitsThu 7:00 — 0 commitsThu 8:00 — 3 commitsThu 9:00 — 2 commitsThu 10:00 — 7 commitsThu 11:00 — 12 commitsThu 12:00 — 7 commitsThu 13:00 — 4 commitsThu 14:00 — 5 commitsThu 15:00 — 12 commitsThu 16:00 — 7 commitsThu 17:00 — 2 commitsThu 18:00 — 1 commitsThu 19:00 — 4 commitsThu 20:00 — 6 commitsThu 21:00 — 3 commitsThu 22:00 — 1 commitsThu 23:00 — 2 commitsFri 0:00 — 6 commitsFri 1:00 — 2 commitsFri 2:00 — 2 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 4 commitsFri 8:00 — 5 commitsFri 9:00 — 5 commitsFri 10:00 — 7 commitsFri 11:00 — 3 commitsFri 12:00 — 3 commitsFri 13:00 — 3 commitsFri 14:00 — 6 commitsFri 15:00 — 3 commitsFri 16:00 — 4 commitsFri 17:00 — 4 commitsFri 18:00 — 8 commitsFri 19:00 — 0 commitsFri 20:00 — 2 commitsFri 21:00 — 1 commitsFri 22:00 — 3 commitsFri 23:00 — 2 commitsSat 0:00 — 3 commitsSat 1:00 — 8 commitsSat 2:00 — 5 commitsSat 3:00 — 6 commitsSat 4:00 — 5 commitsSat 5:00 — 1 commitsSat 6:00 — 0 commitsSat 7:00 — 1 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 6 commitsSat 13:00 — 10 commitsSat 14:00 — 3 commitsSat 15:00 — 3 commitsSat 16:00 — 2 commitsSat 17:00 — 1 commitsSat 18:00 — 2 commitsSat 19:00 — 2 commitsSat 20:00 — 1 commitsSat 21:00 — 2 commitsSat 22:00 — 2 commitsSat 23:00 — 6 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits278 (27%)
Community commits740 (73%)

1,018 commits in total over the last year.

DateListRankStars gained
Jul 27, 2026daily#13+4
Jul 25, 2026daily#21+3
Jul 16, 2026daily#11+7
Jul 15, 2026daily#7+11
Jul 14, 2026daily#1+11
Jul 13, 2026daily#1+14
Jul 12, 2026daily#2+9
Jul 11, 2026daily#1+7
Jul 10, 2026daily#1+9
  • Genymobile/scrcpy

    Display and control your Android device

    147.1K stars · C

  • microsoft/PowerToys

    Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

    137.6K stars · C

  • colbymchenry/codegraph

    Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local

    65.2K stars · C

  • DeusData/codebase-memory-mcp

    High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

    37.8K stars · C

  • antirez/ds4

    DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

    20.8K stars · C

  • erincatto/box3d

    Box3D is a 3D physics engine for games

    5.9K stars · C