RyanCodrai/turbovecPublic

A vector index built on TurboQuant, written in Rust with Python bindings

AI summary: A memory-efficient Rust vector index with Python bindings, built on Google's TurboQuant algorithm.

Stars
17.3K
+8 today
Forks
1.5K
Watchers
75
Open issues
19
Open PRs
1
Contributors
~8
Commits
366
Branches
4

RustMITCreated Mar 26, 2026Last push today+36 stars this week+613 this month

Quick answers

What is turbovec?
A memory-efficient Rust vector index with Python bindings, built on Google's TurboQuant algorithm.
What does turbovec do?
turbovec is a high-performance vector index implemented in Rust that utilizes Google Research's TurboQuant algorithm to achieve extreme memory compression. It allows a 10 million document corpus to fit into 4 GB of RAM instead of the typical 31 GB required for float32. The system maps coordinates to a canonical Beta distribution using Hadamard transforms and applies Lloyd-Max scalar quantization without requiring a separate training phase. It features fast, hand-written SIMD search kernels for both ARM and x86 architectures, often outperforming traditional FAISS indexes in speed. It supports online ingestion, incremental saves, and search-time filtering.
Who is turbovec for?
Data engineers, ML practitioners, and backend developers needing fast, memory-efficient vector search for RAG or similarity applications. Familiarity with Python or Rust is required.
How do I get started with turbovec?
pip install turbovec
How popular is turbovec on GitHub?
RyanCodrai/turbovec has 17,277 stars and 1,480 forks on GitHub, and gained 36 stars in the last 7 days.
What license does turbovec use?
RyanCodrai/turbovec is released under the MIT license.

Star history

since Jul 28, 2026
05K10K15KJul 2026Aug 2026Sep 2026Oct 2026
17.3K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepOctMonWedFri2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 7 commits2026-03-27: 10 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 14 commits2026-03-31: 7 commits2026-04-01: 10 commits2026-04-02: 0 commits2026-04-03: 7 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 16 commits2026-04-14: 7 commits2026-04-15: 8 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 9 commits2026-04-19: 6 commits2026-04-20: 8 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 1 commit2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 5 commits2026-05-18: 8 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 4 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 6 commits2026-05-26: 2 commits2026-05-27: 5 commits2026-05-28: 0 commits2026-05-29: 2 commits2026-05-30: 2 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 6 commits2026-06-10: 2 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 3 commits2026-07-24: 11 commits2026-07-25: 15 commits2026-07-26: 15 commits2026-07-27: 25 commits2026-07-28: 6 commits2026-07-29: 57 commits2026-07-30: 20 commits2026-07-31: 17 commits2026-08-01: 0 commits2026-08-02: 6 commits2026-08-03: 0 commits2026-08-04: 1 commit2026-08-05: 2 commits2026-08-06: 1 commit2026-08-07: 1 commit2026-08-08: 5 commits2026-08-09: 1 commit2026-08-10: 13 commits2026-08-11: 0 commits2026-08-12: 0 commits2026-08-13: 0 commits2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 4 commits2026-08-17: 0 commits2026-08-18: 3 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 4 commits2026-10-03: 0 commits2026-10-04: 0 commits2026-10-05: 0 commits2026-10-06: 0 commits2026-10-07: 0 commits2026-10-08: 0 commits2026-10-09: 0 commits2026-10-10: 0 commits
362 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    17,277 stars

  • Actively maintained

    Pushed within 48 hours

  • Well documented

    High community health score

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    9 trending appearances

What turbovec does

turbovec is a high-performance vector index implemented in Rust that utilizes Google Research's TurboQuant algorithm to achieve extreme memory compression. It allows a 10 million document corpus to fit into 4 GB of RAM instead of the typical 31 GB required for float32. The system maps coordinates to a canonical Beta distribution using Hadamard transforms and applies Lloyd-Max scalar quantization without requiring a separate training phase. It features fast, hand-written SIMD search kernels for both ARM and x86 architectures, often outperforming traditional FAISS indexes in speed. It supports online ingestion, incremental saves, and search-time filtering.

Data engineers, ML practitioners, and backend developers needing fast, memory-efficient vector search for RAG or similarity applications. Familiarity with Python or Rust is required.

  • Online Vector Ingestion: Adds vectors dynamically with no training step, parameter tuning, or index rebuilds required.
  • SIMD-Optimized Search: Uses hand-written NEON (ARM) and AVX-512BW (x86) kernels for blazing fast similarity search.
  • Extreme Memory Compression: Applies advanced Hadamard transforms and scalar quantization to provide up to 16x memory savings.
  • Search-Time Filtering: Honors ID allowlists or slot bitmasks directly within the SIMD kernel to prevent over-fetching.
  • Incremental Disk Syncing: Persists only the changes made since the last sync, ensuring crash-safe, fast saves.
  • Seamless Framework Integration: Provides drop-in replacements for vector stores in LangChain, LlamaIndex, Haystack, and Agno.

Where teams use it

High-Scale Vector Search

Running massive vector databases on memory-constrained or cost-sensitive hardware environments.

Low-Latency RAG Pipelines

Improving retrieval speed and reducing memory footprint in retrieval-augmented generation architectures.

On-Device Embeddings

Deploying large vector indices on edge devices, consumer hardware, or air-gapped systems.

Hybrid Search Retrieval

Combining dense vector reranking with candidate sets produced by external systems like SQL or BM25.

Getting started: pip install turbovec

README

main branch

turbovec — Google's TurboQuant for vector search

License PyPI version crates.io version TurboQuant paper


A 10 million document corpus takes 31 GB of RAM as float32. turbovec fits it in 4 GB - and searches it faster than FAISS.

turbovec is a Rust vector index with Python bindings, built on Google Research's TurboQuant algorithm — a data-oblivious quantizer with near-optimal distortion and no separate training phase.

  • Online ingest. Add vectors, they're indexed — no train step, no parameter tuning, no rebuilds as the corpus grows.
  • Fast SIMD search. Hand-written kernels — NEON SDOT/SMMLA on ARM, AVX-512 VNNI and vpermb on x86, with AVX2 and scalar fallbacks — beat FAISS IndexPQFastScan in every measured config, averaging 3.4× at 4-bit and 23% at 2-bit across the eight cells of each width, on both architectures.
  • Incremental saves. sync(path) persists just what changed since the last sync — one fsync per call, crash-safe at any byte, and a removal or a small append costs milliseconds however large the index. write/load stay for whole-file snapshots.
  • Filter at search time. Pass an id allowlist (or a slot bitmask) to search() and the kernel honours it directly. You always get up to k results from the allowed set — no over-fetching, no recall hit on selective filters.
  • Pure local. No managed service, no data leaving your machine or VPC. Pair with any open-source embedding model for a fully air-gapped RAG stack.

Building RAG where privacy, memory, or latency matters? You're in the right place.

Python

pip install turbovec
from turbovec import TurboQuantIndex

index = TurboQuantIndex(dim=1536, bit_width=4)
index.add(vectors)
index.add(more_vectors)

scores, indices = index.search(query, k=10)

index.write("my_index.tv")
loaded = TurboQuantIndex.load("my_index.tv")

index.sync("my_index.tv")   # after more changes: durable incremental save

vectors and query are 2-D float32 arrays of shape (n, dim) — other dtypes are rejected rather than silently converted, so cast with np.asarray(x, dtype=np.float32) first if needed.

Need stable ids that survive deletes? Use IdMapIndex:

import numpy as np
from turbovec import IdMapIndex

index = IdMapIndex(dim=1536, bit_width=4)
index.add_with_ids(vectors, np.array([1001, 1002, 1003], dtype=np.uint64))

scores, ids = index.search(query, k=10)   # ids are your uint64 external ids
index.remove(1002)                         # O(1) by id

index.write("my_index.tvim")
loaded = IdMapIndex.load("my_index.tvim")

index.sync("my_index.tvim")   # durable incremental save, ids included

Hybrid retrieval (filtered search)

Restrict results to a candidate set produced by another system (SQL, BM25, ACL, time window, …):

import numpy as np
from turbovec import IdMapIndex

idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, ids)

# Stage 1: external system narrows to candidate ids.
allowed = np.array(db.execute("SELECT id FROM docs WHERE tenant=?", (t,)).fetchall(),
                   dtype=np.uint64)

# Stage 2: dense rerank within the candidate set.
scores, ids = idx.search(query, k=10, allowlist=allowed)

Filtering happens inside the SIMD kernel at 32-vector block granularity: blocks with no allowed slots are short-circuited before any LUT lookup or scoring work, and individual non-allowed slots inside scored blocks are dropped at heap-insert. Selective allowlists (small fraction of the index allowed) therefore avoid most of the SIMD cost rather than paying it and discarding the result afterwards.

The output length is min(k, n_allowed), where n_allowed counts distinct allowed vectors — when fewer vectors are allowed than k you get exactly that many results rather than padded fallbacks.

See docs/api.md for the full reference.

Framework integrations

Drop-in replacements for the in-tree reference vector / document stores in each framework. Same public surface, same persistence semantics, same retriever and pipeline wiring — swap the import and keep your pipeline.

  • LangChain — pip install turbovec[langchain] · replaces langchain_core.vectorstores.InMemoryVectorStore
  • LlamaIndex — pip install turbovec[llama-index] · replaces llama_index.core.vector_stores.SimpleVectorStore
  • Haystack — pip install turbovec[haystack] · replaces haystack.document_stores.in_memory.InMemoryDocumentStore
  • Agno — pip install turbovec[agno] · replaces agno.vectordb.lancedb.LanceDb

Rust

cargo add turbovec
use turbovec::TurboQuantIndex;

let mut index = TurboQuantIndex::new(1536, 4).unwrap();
index.add(&vectors);
let results = index.search(&queries, 10);
index.write("index.tv").unwrap();
let loaded = TurboQuantIndex::load("index.tv").unwrap();

For stable external ids that survive deletes:

use turbovec::IdMapIndex;

let mut index = IdMapIndex::new(1536, 4).unwrap();
index.add_with_ids(&vectors, &[1001, 1002, 1003]).unwrap();
let (scores, ids) = index.search(&queries, 10);
index.remove(1002);
index.write("index.tvim").unwrap();
let loaded = IdMapIndex::load("index.tvim").unwrap();

Recall

TurboQuant vs FAISS IndexPQ (LUT256, nbits=8) — the paper's Section 4.4 baseline. 100K vectors, k=64. FAISS PQ sub-quantizer counts sized to match TurboQuant's bit rate (m=d/4 at 2-bit, m=d/2 at 4-bit).

Recall GloVe d=200

Recall d=1536

Recall d=3072

The charts plot calibrated TurboQuant (TQ+). Across OpenAI d=1536 and d=3072, TQ+ beats FAISS at R@1 on three of four cells (by 0.9–2.9 points; d=1536 4-bit trails by 0.7), and both reach 1.0 by k=8 (≥0.997 already at k≤4). GloVe d=200 is the harder regime — at low dim the asymptotic Beta assumption is looser. TQ+ lands ahead of FAISS at R@1 at both bit widths (+1.9 at 4-bit, +0.8 at 2-bit), with FAISS keeping a slim edge at 2-bit from k≈8. Uncalibrated numbers are in the JSONs (tq_recalls).

A note on baselines. We compare against FAISS IndexPQ (LUT256, nbits=8, float32 LUT) because it's the default production-grade PQ most users would reach for. This is a stronger baseline than the custom u8-LUT PQ in the TurboQuant paper — FAISS uses a higher-precision LUT at scoring time and k-means++ for codebook training. We reproduce the paper's TurboQuant numbers on OpenAI d=1536 / d=3072 and hit similar numbers to other community reference implementations on low-dim embeddings (see turboquant-py at d=384). On GloVe (d=200) — the low-dim regime where the asymptotic Beta assumption is loosest — TurboQuant lands ahead of FAISS at 4-bit but trails it at 2-bit; TQ+ calibration recovers the 2-bit deficit at R@1 (0.572 vs FAISS's 0.564), with FAISS keeping a slim edge at deeper k.

Full results: d=1536 2-bit, d=1536 4-bit, d=3072 2-bit, d=3072 4-bit, GloVe 2-bit, GloVe 4-bit.

Compression

Compression

Search Speed

All benchmarks: 100K vectors, 1K queries, k=64, median of 5 runs.

ARM (GCP c4a-standard-8, Google Axion, 8 vCPUs)

ARM Speed — Single-threaded

ARM Speed — Multi-threaded

On ARM, TurboQuant beats FAISS FastScan in every config, averaging 3.5× at 4-bit (3.4–3.7× across cells — the SDOT/SMMLA dot-product kernels score the vector-major layout directly) and 26% at 2-bit (22–29%).

x86 (Intel Xeon Platinum 8481C / Sapphire Rapids, 8 vCPUs)

x86 Speed — Single-threaded

x86 Speed — Multi-threaded

On x86, TurboQuant wins every config, averaging 3.4× at 4-bit (3.2–3.5× across cells — the AVX-512 VNNI dot-product kernel on the vector-major layout) and 20% at 2-bit (5–32%), where the vpermb LUT scan carries the short 2-bit accumulate loop.

Insertion & Removal Latency

Same corpus as the search cells: 100K OpenAI vectors, median of 5 runs, timed loops including the Python-call overhead a caller actually pays per op. Insertion measures per-vector add() latency on a warm, populated index (built untimed) at n=1 — a single-vector add() — and n=100 — a 100-vector batch, showing how far batching amortizes the per-call overhead — against add() into the trained, populated FAISS IndexPQFastScan (training untimed). A single add() lands in 6.3–19.7 µs depending on the cell (7.6–13.9× faster than a FAISS single add), and a 100-vector batch amortizes TurboQuant to 4.6–16.3 µs/vector (4.6–15.1× faster than the same batch into FAISS). Removal measures per-op remove-by-id latency at n=1 (the steady per-op rate over 1000 removes) and n=100 (the first 100 removes on a fresh index): IdMapIndex.remove(id) — O(1) swap-and-pop plus the id-map bookkeeping — lands at 0.44–1.22 µs and 0.59–1.37 µs per op across the cells. The FAISS column is the same user-visible operation, remove_ids on an IndexIDMap over IndexPQFastScan, which repacks the stored codes on every call: 0.19–1.02 s per single remove at 100K, with cost doubling alongside code size — which is why the removal charts use a log-scale axis. Charts show the single-threaded cells (RAYON_NUM_THREADS=1); the _mt cells are measured too and match at n=1, since a single add is serial. Scripts: benchmarks/suite/.

ARM (GCP c4a-standard-8, Google Axion, 8 vCPUs)

ARM Online Insert Latency — Single-threaded

ARM Online Remove Latency — Single-threaded

Full results: d=1536 2-bit insert, d=1536 4-bit insert, d=3072 2-bit insert, d=3072 4-bit insert, and the matching speed_remove_* and _mt files.

x86 (Intel Xeon Platinum 8481C / Sapphire Rapids, 8 vCPUs)

x86 Online Insert Latency — Single-threaded

x86 Online Remove Latency — Single-threaded

Full results: d=1536 2-bit insert, d=1536 4-bit insert, d=3072 2-bit insert, d=3072 4-bit insert, and the matching speed_remove_* and _mt files.

Save & Load

Same corpus as the search cells: 100K OpenAI vectors, median of 5 runs. TurboQuant serializes to a single .tv file with an fsync + atomic rename; FAISS is write_index / read_index on the precision-matched IndexPQFastScan (sub-quantizer count matched to TurboQuant's bit rate, as in the search cells). Save (warm) is a write after a search has run, so the blocked layout cache is populated. Load → first search opens a fresh index and times the first query — separating bare deserialization (the page cache is warm throughout, so this is layout work, not cold-storage I/O) from the first-query cost. Round-trip chains the checkpoint/resume cycle an embedding store actually pays — mutate 1K vectors → save → reopen → serve the first query; FAISS has no measured equivalent for this path, so it is shown for TurboQuant only. On the smaller payloads the round-trip can come in below the isolated post-mutation ("dirty") write: the two are timed in separate suite steps, and at small file sizes the standalone fsync in the dirty-write step dominates and inflates it — a measurement artifact of the harness, not a repack win in the combined path. Single-threaded cells pin RAYON_NUM_THREADS=1. Scripts: benchmarks/suite/.

ARM (GCP c4a-standard-8, Google Axion, 8 vCPUs)

ARM Save/Load — Single-threaded

ARM Save/Load — Multi-threaded

Full results: d=1536 2-bit persist ST, MT, d=1536 4-bit persist ST, MT, d=3072 2-bit persist ST, MT, d=3072 4-bit persist ST, MT.

x86 (Intel Xeon Platinum 8481C / Sapphire Rapids, 8 vCPUs)

x86 Save/Load — Single-threaded

x86 Save/Load — Multi-threaded

Full results: d=1536 2-bit persist ST, MT, d=1536 4-bit persist ST, MT, d=3072 2-bit persist ST, MT, d=3072 4-bit persist ST, MT.

How it works

Each vector is a direction on a high-dimensional hypersphere. TurboQuant compresses these directions using a simple insight: after applying a random rotation, every coordinate follows a known distribution -- regardless of the input data.

1. Normalize. Strip the length (norm) from each vector and store it as a single float. Now every vector is a unit direction on the hypersphere.

2. Random rotation. Multiply all vectors by the same random orthogonal matrix. After rotation, each coordinate independently follows a Beta distribution that converges to Gaussian N(0, 1/d) in high dimensions. This holds for any input data -- the rotation makes the coordinate distribution predictable.

3. Per-coordinate calibration (TQ+). The Beta distribution from step 2 is asymptotic — at finite dimensions, individual coordinates drift from the canonical shape (especially low-bit and word-vector-style embeddings). TQ+ fits two scalars per coordinate — a shift and a scale — mapping each coordinate's empirical quantiles onto the codebook's outermost centroids. The probability level comes from the codebook, so it tracks the bit width (~0.933 at 2-bit, ~0.996 at 4-bit) rather than being fixed. The Lloyd-Max codebook then quantizes against the target distribution it was designed for. The fit is explicit: call index.calibrate(sample) once with a random, representative sample of your vectors (~1024 rows is enough — a draw of that size matches fitting on the whole corpus) before adding; afterwards the calibration is committed and reused by every add — no retraining, no rebuilds, no separate train phase. An index you never calibrate is plain TurboQuant. index.calibration_state reports "uncalibrated" or "calibrated". Recall gain: up to +2.2pp at @1 on the cells that drift most (e.g. GloVe at 2-bit).

4. Lloyd-Max scalar quantization. Since the distribution is known, we can precompute the optimal way to bucket each coordinate. For 2-bit, that's 4 buckets; for 4-bit, 16 buckets. The Lloyd-Max algorithm finds bucket boundaries and centroids that minimize mean squared error. These are computed once from the math, not from the data.

5. Bit-pack. Each coordinate is now a small integer (0-3 for 2-bit, 0-15 for 4-bit). Pack these tightly into bytes. A 1536-dim vector goes from 6,144 bytes (FP32) to 384 bytes (2-bit). That's 16x compression.

6. Length-renormalized scoring. Scalar quantization systematically underestimates inner products — the reconstructed unit direction is a little shorter than the original. We compute one scalar per vector at encode time — the inner product of the rotated unit vector with its own centroid reconstruction — and store ||v|| / ⟨u, x̂⟩ alongside each compressed vector. The search kernel multiplies the per-candidate score by this scalar before heap insertion, turning the inner-product estimator from downward-biased into unbiased at zero search-time cost and zero extra storage. The recall gain shows up most at low bit widths, where the quantization shrinkage is largest.

Encoding cost: one extra d-dimensional dot product per vector to compute ⟨u, x̂⟩. On 1M vectors at d=1536 this is sub-second of additional encode time — a one-shot price paid at ingest, not at query.

Search. Instead of decompressing every database vector, we rotate the query once into the same domain and score directly against the codebook values. The scoring kernel uses SIMD intrinsics (NEON on ARM; AVX-512BW on modern x86, falling back to AVX2, then to a scalar path on pre-AVX2 CPUs) with nibble-split lookup tables for maximum throughput.

The Lloyd-Max codebook achieves distortion within a factor of 2.7x of the information-theoretic lower bound (Shannon's distortion-rate limit); the length-renormalization step removes the residual bias the Lloyd-Max codebook introduces on the inner-product estimator itself.

Building

Python (via maturin)

pip install maturin
cd turbovec-python
maturin build --release
pip install target/wheels/*.whl

Rust

cargo build --release

All x86_64 builds target x86-64-v2 (SSE4.2 baseline, Nehalem 2008+) via .cargo/config.toml, so any x86-64-v2 CPU can run the whole crate. The AVX-512 and AVX2 kernels are #[target_feature]-gated and selected at runtime via is_x86_feature_detected!, so they kick in on hardware that supports them regardless of the compile baseline; CPUs with neither run the scalar fallback.

Running benchmarks

Download datasets:

python3 benchmarks/download_data.py all            # all datasets
python3 benchmarks/download_data.py glove          # GloVe d=200
python3 benchmarks/download_data.py openai-1536    # OpenAI DBpedia d=1536
python3 benchmarks/download_data.py openai-3072    # OpenAI DBpedia d=3072

Each benchmark is a self-contained script in benchmarks/suite/. Run any one individually:

python3 benchmarks/suite/speed_d1536_2bit_arm_mt.py
python3 benchmarks/suite/recall_d1536_2bit.py
python3 benchmarks/suite/compression.py

Run all benchmarks for a category:

for f in benchmarks/suite/speed_*arm*.py; do python3 "$f"; done    # all ARM speed
for f in benchmarks/suite/speed_*x86*.py; do python3 "$f"; done    # all x86 speed
for f in benchmarks/suite/recall_*.py; do python3 "$f"; done       # all recall
python3 benchmarks/suite/compression.py                            # compression

Results are saved as JSON to benchmarks/results/. Regenerate charts:

python3 benchmarks/create_diagrams.py

Quick harness for optimization work

The suite above is the source of every published number — real embeddings, FAISS comparator, fixed shapes, run on the two official environments. For the inner loop of an optimization pass there's also a Rust harness that reproduces the four mutation metrics (cold bulk add, warm append, single add, remove) on deterministic synthetic vectors, so a hypothesis can be measured in seconds on any machine with no dataset and no FAISS:

cargo run --release --example insert_bench -- --dim 1536 --bits 2
RAYON_NUM_THREADS=1 cargo run --release --example insert_bench

It is a screening tool, not a source of published numbers.

examples/encode_hash prints a per-stage hash of the encode pipeline for a fixed input; CI runs it on every OS in the matrix and fails if they disagree, which is how cross-platform byte identity of the encode is checked.

References

View on GitHub

Recent activity

commits and pull requests

Code frequency

additions and deletions
+43.4K-43.4KWeek of 2026-03-22: +688 linesWeek of 2026-03-22: -120 linesWeek of 2026-03-29: +10,795 linesWeek of 2026-03-29: -5,454 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +2,648 linesWeek of 2026-04-12: -1,271 linesWeek of 2026-04-19: +4,674 linesWeek of 2026-04-19: -556 linesWeek of 2026-04-26: +29 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +9,152 linesWeek of 2026-05-17: -685 linesWeek of 2026-05-24: +7,229 linesWeek of 2026-05-24: -1,708 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +1,460 linesWeek of 2026-06-07: -2,144 linesWeek of 2026-06-14: +0 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +0 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +0 linesWeek of 2026-07-12: -0 linesWeek of 2026-07-19: +9,547 linesWeek of 2026-07-19: -2,197 linesWeek of 2026-07-26: +43,436 linesWeek of 2026-07-26: -8,154 linesWeek of 2026-08-02: +38,104 linesWeek of 2026-08-02: -6,115 linesWeek of 2026-08-09: +3,671 linesWeek of 2026-08-09: -177 linesWeek of 2026-08-16: +3,541 linesWeek of 2026-08-16: -5,129 linesWeek of 2026-08-23: +0 linesWeek of 2026-08-23: -0 linesWeek of 2026-08-30: +0 linesWeek of 2026-08-30: -0 linesWeek of 2026-09-06: +0 linesWeek of 2026-09-06: -0 linesWeek of 2026-09-13: +0 linesWeek of 2026-09-13: -0 linesWeek of 2026-09-20: +0 linesWeek of 2026-09-20: -0 linesMar 22, 2026Sep 20, 2026
+135K lines added, -33.7K removed over the last year.

Commits per week

last 52 weeks
1400Week of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 17 commitsWeek of 2026-03-29: 38 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 40 commitsWeek of 2026-04-19: 14 commitsWeek of 2026-04-26: 1 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 17 commitsWeek of 2026-05-24: 17 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 8 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 29 commitsWeek of 2026-07-26: 140 commitsWeek of 2026-08-02: 16 commitsWeek of 2026-08-09: 14 commitsWeek of 2026-08-16: 7 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 0 commitsWeek of 2026-09-06: 0 commitsWeek of 2026-09-13: 0 commitsWeek of 2026-09-20: 0 commitsWeek of 2026-09-27: 4 commitsWeek of 2026-10-04: 0 commitsOct 12, 2025Oct 4, 2026
362 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 1 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 4 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 3 commitsSun 10:00 — 2 commitsSun 11:00 — 5 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 1 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 2 commitsSun 19:00 — 3 commitsSun 20:00 — 3 commitsSun 21:00 — 4 commitsSun 22:00 — 0 commitsSun 23:00 — 9 commitsMon 0:00 — 2 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 1 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 4 commitsMon 11:00 — 6 commitsMon 12:00 — 4 commitsMon 13:00 — 4 commitsMon 14:00 — 7 commitsMon 15:00 — 10 commitsMon 16:00 — 7 commitsMon 17:00 — 7 commitsMon 18:00 — 10 commitsMon 19:00 — 7 commitsMon 20:00 — 5 commitsMon 21:00 — 4 commitsMon 22:00 — 7 commitsMon 23:00 — 5 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 1 commitsTue 10:00 — 0 commitsTue 11:00 — 1 commitsTue 12:00 — 1 commitsTue 13:00 — 3 commitsTue 14:00 — 2 commitsTue 15:00 — 4 commitsTue 16:00 — 6 commitsTue 17:00 — 5 commitsTue 18:00 — 3 commitsTue 19:00 — 0 commitsTue 20:00 — 1 commitsTue 21:00 — 1 commitsTue 22:00 — 1 commitsTue 23:00 — 3 commitsWed 0:00 — 6 commitsWed 1:00 — 6 commitsWed 2:00 — 3 commitsWed 3:00 — 3 commitsWed 4:00 — 5 commitsWed 5:00 — 3 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 1 commitsWed 10:00 — 0 commitsWed 11:00 — 9 commitsWed 12:00 — 4 commitsWed 13:00 — 3 commitsWed 14:00 — 2 commitsWed 15:00 — 4 commitsWed 16:00 — 9 commitsWed 17:00 — 4 commitsWed 18:00 — 5 commitsWed 19:00 — 4 commitsWed 20:00 — 5 commitsWed 21:00 — 3 commitsWed 22:00 — 4 commitsWed 23:00 — 1 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 1 commitsThu 6:00 — 1 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 1 commitsThu 10:00 — 1 commitsThu 11:00 — 3 commitsThu 12:00 — 3 commitsThu 13:00 — 1 commitsThu 14:00 — 8 commitsThu 15:00 — 2 commitsThu 16:00 — 1 commitsThu 17:00 — 0 commitsThu 18:00 — 0 commitsThu 19:00 — 1 commitsThu 20:00 — 1 commitsThu 21:00 — 1 commitsThu 22:00 — 7 commitsThu 23:00 — 3 commitsFri 0:00 — 2 commitsFri 1:00 — 2 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 6 commitsFri 9:00 — 4 commitsFri 10:00 — 1 commitsFri 11:00 — 3 commitsFri 12:00 — 0 commitsFri 13:00 — 9 commitsFri 14:00 — 1 commitsFri 15:00 — 2 commitsFri 16:00 — 1 commitsFri 17:00 — 5 commitsFri 18:00 — 4 commitsFri 19:00 — 1 commitsFri 20:00 — 0 commitsFri 21:00 — 1 commitsFri 22:00 — 9 commitsFri 23:00 — 1 commitsSat 0:00 — 2 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 1 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 2 commitsSat 14:00 — 2 commitsSat 15:00 — 1 commitsSat 16:00 — 3 commitsSat 17:00 — 4 commitsSat 18:00 — 2 commitsSat 19:00 — 0 commitsSat 20:00 — 1 commitsSat 21:00 — 1 commitsSat 22:00 — 11 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits354 (97%)
Community commits12 (3%)

366 commits in total over the last year.

DateListRankStars gained
Aug 22, 2026daily#12+230
Aug 21, 2026daily#12+230
Jun 10, 2026daily#22+19
Jun 9, 2026daily#5+34
Jun 8, 2026daily#5+66
Jun 7, 2026daily#3+166
Jun 6, 2026daily#14+65
May 22, 2026daily#22+45
May 21, 2026daily#20+43
  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    373.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    285.8K stars · Python

  • tensorflow/tensorflow

    An Open Source Machine Learning Framework for Everyone

    200.7K stars · C++

  • yt-dlp/yt-dlp

    A feature-rich command-line audio/video downloader

    195.5K stars · Python

  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    195.2K stars · Rust

  • Significant-Gravitas/AutoGPT

    AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

    187.7K stars · Python