xberg-io/xbergPublic

Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.

AI summary: A high-performance document intelligence framework for extracting structured data from over 98 file formats.

Stars
9.4K
+7 today
Forks
588
Watchers
37
Open issues
20
Open PRs
3
Contributors
~62
Commits
9.8K
Branches
31

RustMITCreated Jan 31, 2025Last push 1d agoLatest release v1.2.9+32 stars this week+106 this month

Quick answers

What is xberg?
A high-performance document intelligence framework for extracting structured data from over 98 file formats.
What does xberg do?
Kreuzberg is an advanced, polyglot document intelligence tool built on a robust Rust core that specializes in deep data extraction. It is designed to parse text, metadata, and embedded images from a massive variety of sources, including PDFs, Office documents, and over 98 distinct file formats. The framework exposes this capability through a wide array of interfaces, offering native bindings for Python, Node.js, Java, Go, and C#, alongside REST APIs and CLI tools. By providing high-speed, reliable parsing across virtually any tech stack, it enables seamless ingestion of complex, unstructured documents into modern AI processing pipelines.
Who is xberg for?
Data engineers, backend developers, and researchers requiring a blazing fast, highly reliable framework to extract structured data from diverse document formats.
How do I get started with xberg?
Choose the appropriate binding for your language (e.g., pip install xberg, npm install @xberg-io/xberg) or explore the Rust crate.
How popular is xberg on GitHub?
xberg-io/xberg has 9,366 stars and 588 forks on GitHub, and gained 32 stars in the last 7 days.
What license does xberg use?
xberg-io/xberg is released under the MIT license.

Star history

since Jul 29, 2026
02.5K5K7.5KJul 2026Aug 2026Sep 2026Oct 2026
9.4K stars as of Oct 3, 2026. Measured daily since Jul 29, 2026; GitHub no longer exposes earlier star timestamps.

Update history

1 recorded
  • Oct 4, 2026Previously tracked as kreuzberg-dev/kreuzberg; its 1 daily snapshot and 3 trending appearances were merged into this profile. Stars: 8,712 on 2026-07-29 under the old name, 9,366 on 2026-10-03 (+654).

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-28: 30 commits2025-09-29: 26 commits2025-09-30: 7 commits2025-10-01: 4 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 1 commit2025-10-07: 0 commits2025-10-08: 1 commit2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 13 commits2025-10-12: 1 commit2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 2 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 3 commits2025-11-05: 2 commits2025-11-06: 0 commits2025-11-07: 73 commits2025-11-08: 20 commits2025-11-09: 27 commits2025-11-10: 18 commits2025-11-11: 15 commits2025-11-12: 11 commits2025-11-13: 37 commits2025-11-14: 36 commits2025-11-15: 29 commits2025-11-16: 15 commits2025-11-17: 16 commits2025-11-18: 19 commits2025-11-19: 12 commits2025-11-20: 57 commits2025-11-21: 55 commits2025-11-22: 25 commits2025-11-23: 34 commits2025-11-24: 11 commits2025-11-25: 13 commits2025-11-26: 17 commits2025-11-27: 67 commits2025-11-28: 25 commits2025-11-29: 25 commits2025-11-30: 33 commits2025-12-01: 23 commits2025-12-02: 28 commits2025-12-03: 24 commits2025-12-04: 34 commits2025-12-05: 17 commits2025-12-06: 28 commits2025-12-07: 38 commits2025-12-08: 30 commits2025-12-09: 20 commits2025-12-10: 26 commits2025-12-11: 28 commits2025-12-12: 41 commits2025-12-13: 58 commits2025-12-14: 22 commits2025-12-15: 35 commits2025-12-16: 32 commits2025-12-17: 26 commits2025-12-18: 23 commits2025-12-19: 47 commits2025-12-20: 53 commits2025-12-21: 70 commits2025-12-22: 45 commits2025-12-23: 36 commits2025-12-24: 34 commits2025-12-25: 6 commits2025-12-26: 32 commits2025-12-27: 64 commits2025-12-28: 36 commits2025-12-29: 36 commits2025-12-30: 24 commits2025-12-31: 33 commits2026-01-01: 8 commits2026-01-02: 29 commits2026-01-03: 28 commits2026-01-04: 20 commits2026-01-05: 16 commits2026-01-06: 13 commits2026-01-07: 24 commits2026-01-08: 22 commits2026-01-09: 20 commits2026-01-10: 35 commits2026-01-11: 20 commits2026-01-12: 9 commits2026-01-13: 13 commits2026-01-14: 13 commits2026-01-15: 6 commits2026-01-16: 11 commits2026-01-17: 11 commits2026-01-18: 18 commits2026-01-19: 14 commits2026-01-20: 14 commits2026-01-21: 14 commits2026-01-22: 16 commits2026-01-23: 5 commits2026-01-24: 8 commits2026-01-25: 17 commits2026-01-26: 15 commits2026-01-27: 17 commits2026-01-28: 23 commits2026-01-29: 10 commits2026-01-30: 14 commits2026-01-31: 25 commits2026-02-01: 9 commits2026-02-02: 10 commits2026-02-03: 5 commits2026-02-04: 14 commits2026-02-05: 13 commits2026-02-06: 21 commits2026-02-07: 21 commits2026-02-08: 18 commits2026-02-09: 27 commits2026-02-10: 14 commits2026-02-11: 21 commits2026-02-12: 18 commits2026-02-13: 10 commits2026-02-14: 7 commits2026-02-15: 13 commits2026-02-16: 16 commits2026-02-17: 18 commits2026-02-18: 17 commits2026-02-19: 14 commits2026-02-20: 9 commits2026-02-21: 17 commits2026-02-22: 6 commits2026-02-23: 7 commits2026-02-24: 35 commits2026-02-25: 23 commits2026-02-26: 25 commits2026-02-27: 8 commits2026-02-28: 13 commits2026-03-01: 5 commits2026-03-02: 11 commits2026-03-03: 8 commits2026-03-04: 12 commits2026-03-05: 17 commits2026-03-06: 7 commits2026-03-07: 17 commits2026-03-08: 21 commits2026-03-09: 9 commits2026-03-10: 15 commits2026-03-11: 23 commits2026-03-12: 12 commits2026-03-13: 22 commits2026-03-14: 52 commits2026-03-15: 23 commits2026-03-16: 13 commits2026-03-17: 19 commits2026-03-18: 28 commits2026-03-19: 38 commits2026-03-20: 39 commits2026-03-21: 36 commits2026-03-22: 10 commits2026-03-23: 29 commits2026-03-24: 31 commits2026-03-25: 24 commits2026-03-26: 20 commits2026-03-27: 17 commits2026-03-28: 1 commit2026-03-29: 12 commits2026-03-30: 64 commits2026-03-31: 26 commits2026-04-01: 38 commits2026-04-02: 29 commits2026-04-03: 21 commits2026-04-04: 30 commits2026-04-05: 26 commits2026-04-06: 21 commits2026-04-07: 18 commits2026-04-08: 12 commits2026-04-09: 13 commits2026-04-10: 12 commits2026-04-11: 7 commits2026-04-12: 4 commits2026-04-13: 13 commits2026-04-14: 36 commits2026-04-15: 4 commits2026-04-16: 11 commits2026-04-17: 22 commits2026-04-18: 10 commits2026-04-19: 18 commits2026-04-20: 46 commits2026-04-21: 7 commits2026-04-22: 13 commits2026-04-23: 18 commits2026-04-24: 20 commits2026-04-25: 30 commits2026-04-26: 25 commits2026-04-27: 14 commits2026-04-28: 16 commits2026-04-29: 38 commits2026-04-30: 48 commits2026-05-01: 10 commits2026-05-02: 27 commits2026-05-03: 19 commits2026-05-04: 22 commits2026-05-05: 44 commits2026-05-06: 9 commits2026-05-07: 90 commits2026-05-08: 76 commits2026-05-09: 73 commits2026-05-10: 86 commits2026-05-11: 17 commits2026-05-12: 15 commits2026-05-13: 56 commits2026-05-14: 30 commits2026-05-15: 40 commits2026-05-16: 11 commits2026-05-17: 31 commits2026-05-18: 28 commits2026-05-19: 19 commits2026-05-20: 19 commits2026-05-21: 65 commits2026-05-22: 51 commits2026-05-23: 18 commits2026-05-24: 18 commits2026-05-25: 32 commits2026-05-26: 28 commits2026-05-27: 9 commits2026-05-28: 43 commits2026-05-29: 18 commits2026-05-30: 69 commits2026-05-31: 47 commits2026-06-01: 17 commits2026-06-02: 45 commits2026-06-03: 37 commits2026-06-04: 31 commits2026-06-05: 30 commits2026-06-06: 18 commits2026-06-07: 23 commits2026-06-08: 12 commits2026-06-09: 14 commits2026-06-10: 20 commits2026-06-11: 15 commits2026-06-12: 15 commits2026-06-13: 7 commits2026-06-14: 18 commits2026-06-15: 9 commits2026-06-16: 24 commits2026-06-17: 90 commits2026-06-18: 33 commits2026-06-19: 10 commits2026-06-20: 53 commits2026-06-21: 11 commits2026-06-22: 24 commits2026-06-23: 24 commits2026-06-24: 29 commits2026-06-25: 45 commits2026-06-26: 23 commits2026-06-27: 58 commits2026-06-28: 49 commits2026-06-29: 29 commits2026-06-30: 22 commits2026-07-01: 7 commits2026-07-02: 67 commits2026-07-03: 33 commits2026-07-04: 12 commits2026-07-05: 25 commits2026-07-06: 18 commits2026-07-07: 40 commits2026-07-08: 23 commits2026-07-09: 40 commits2026-07-10: 12 commits2026-07-11: 20 commits2026-07-12: 33 commits2026-07-13: 21 commits2026-07-14: 2 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 17 commits2026-07-18: 5 commits2026-07-19: 30 commits2026-07-20: 54 commits2026-07-21: 47 commits2026-07-22: 58 commits2026-07-23: 56 commits2026-07-24: 50 commits2026-07-25: 39 commits2026-07-26: 41 commits2026-07-27: 39 commits2026-07-28: 42 commits2026-07-29: 34 commits2026-07-30: 33 commits2026-07-31: 71 commits2026-08-01: 60 commits2026-08-02: 41 commits2026-08-03: 21 commits2026-08-04: 40 commits2026-08-05: 87 commits2026-08-06: 108 commits2026-08-07: 85 commits2026-08-08: 61 commits2026-08-09: 30 commits2026-08-10: 21 commits2026-08-11: 5 commits2026-08-12: 23 commits2026-08-13: 21 commits2026-08-14: 21 commits2026-08-15: 4 commits2026-08-16: 14 commits2026-08-17: 6 commits2026-08-18: 51 commits2026-08-19: 46 commits2026-08-20: 32 commits2026-08-21: 36 commits2026-08-22: 99 commits2026-08-23: 47 commits2026-08-24: 47 commits2026-08-25: 26 commits2026-08-26: 12 commits2026-08-27: 37 commits2026-08-28: 55 commits2026-08-29: 57 commits2026-08-30: 35 commits2026-08-31: 20 commits2026-09-01: 47 commits2026-09-02: 30 commits2026-09-03: 15 commits2026-09-04: 14 commits2026-09-05: 19 commits2026-09-06: 16 commits2026-09-07: 28 commits2026-09-08: 24 commits2026-09-09: 19 commits2026-09-10: 18 commits2026-09-11: 20 commits2026-09-12: 19 commits2026-09-13: 5 commits2026-09-14: 26 commits2026-09-15: 20 commits2026-09-16: 26 commits2026-09-17: 5 commits2026-09-18: 16 commits2026-09-19: 46 commits2026-09-20: 31 commits2026-09-21: 16 commits2026-09-22: 97 commits2026-09-23: 35 commits2026-09-24: 24 commits2026-09-25: 2 commits2026-09-26: 0 commits
8,837 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Very active

    8,837 commits in 52 weeks

  • Well documented

    High community health score

  • Permissive license

    MIT

  • Repeat trending

    3 trending appearances

What xberg does

Kreuzberg is an advanced, polyglot document intelligence tool built on a robust Rust core that specializes in deep data extraction. It is designed to parse text, metadata, and embedded images from a massive variety of sources, including PDFs, Office documents, and over 98 distinct file formats. The framework exposes this capability through a wide array of interfaces, offering native bindings for Python, Node.js, Java, Go, and C#, alongside REST APIs and CLI tools. By providing high-speed, reliable parsing across virtually any tech stack, it enables seamless ingestion of complex, unstructured documents into modern AI processing pipelines.

Data engineers, backend developers, and researchers requiring a blazing fast, highly reliable framework to extract structured data from diverse document formats.

  • Extensive format support: Capable of extracting text, metadata, and images from over 98 distinct file types, including complex PDFs and Office formats.
  • High-performance Rust core: Built natively in Rust to ensure exceptional speed, memory safety, and minimal resource overhead during large batch processing.
  • Polyglot bindings: Offers native libraries and integration support for a vast array of languages, including Python, Java, Go, C#, Ruby, and WebAssembly.
  • Flexible execution modes: Can be utilized directly via language bindings, executed through a CLI, or deployed as a REST API or MCP server.
  • Structured output generation: Converts highly unstructured document data into clean, structured formats suitable for downstream AI analysis or database ingestion.

Where teams use it

Large-scale document ingestion

Data engineering teams utilize the tool to rapidly parse thousands of legacy PDF reports into a structured data lake for machine learning training.

Cross-platform data extraction

Enterprise architectures integrate the polyglot bindings to ensure identical document parsing behavior across their Python microservices and Java backends.

Web-based file processing

Frontend developers leverage the WebAssembly build to securely parse sensitive Office documents entirely within the user's browser.

Automated text mining

Researchers use the CLI tool in automated bash scripts to quickly extract metadata and plain text from massive, diverse document archives.

Getting started: Choose the appropriate binding for your language (e.g., pip install xberg, npm install @xberg-io/xberg) or explore the Rust crate.

README

main branch

Xberg

Xberg

The fast, precise document-intelligence engine — for every language.

Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data. One engine handles format detection, reading, OCR, and extraction, so you never stitch a pipeline together from a dozen libraries.

107 formats · 141 file extensions · 371 code languages · 15 language bindings · 6 output formats · OCR · transcription · embeddings

The fastest, most precise open-source document and PDF-to-Markdown engine — see the benchmarks.

Install · What you get · Capabilities · CLI · Docs

Xberg is the next iteration of Kreuzberg. Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line.


What you get

Point Xberg at anything — a PDF, a spreadsheet, a scanned image, an audio file, a URL, an archive, a source tree — and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server.

Capability What you get
107 document formats PDFs, Office, images, HTML, email, e-books, scientific publications, and structured data across 141 file extensions, with intelligent MIME detection and bounded extraction controls.
URLs & the web Point Xberg at an http(s) URL — it fetches and extracts a single document, or crawls and follows links (Auto / Document / Crawl modes via the crawlberg engine). Requires the url-ingestion feature.
Audio & video transcription Speech-to-text from MP3, M4A, WAV, WebM, and MP4 tracks via Whisper ONNX (tiny → large-v3). Requires the transcription feature.
Archives, traversed List and recursively extract nested .zip, .tar, .gz, .7z — documents inside documents — guarded by zip-bomb, compression-ratio, and nesting-depth limits.
OCR on demand Tesseract, PaddleOCR, Candle, or VLM backends — fallback chains, confidence scores, language auto-detection, extensible via plugins.
Layout & tables ML layout models (PP-DocLayout-V3, RT-DETR) and table structure (TATR, SLANet) reconstruct reading order and cell grids for clean Markdown.
Code intelligence Functions, classes, imports, symbols, docstrings from 371 programming languages. Syntax-aware chunking for RAG pipelines.
Embeddings & search Local (ONNX) or provider-hosted embeddings (165 providers via liter-llm), sparse and late-interaction, cross-encoder reranking.
Enrichment NER, keyword extraction (YAKE/RAKE), summarization, translation, redaction, page classification, QR detection, language detection, token reduction (TOON).
Structured extraction Schema-driven JSON straight from any document via local (Ollama, LM Studio, vLLM) or hosted LLMs — no prompt engineering.
6 output formats Plain text, Markdown, Djot, HTML, JSON tree, or Docling DocTags, plus registered custom renderers.
Runs anywhere Library, CLI (14 commands), REST API (xberg serve), MCP server, Docker, Helm — CPU by default, no GPU required. Content-hash caching, parallel batch, per-file timeouts.

Capabilities marked requires a feature are Cargo feature flags on the core crate (url-ingestion, transcription, reranker, layout/ORT). Prebuilt language packages and the Docker image bundle the common set; a from-source build enables only what you select.


Installation

Language Packages

Python
pip install xberg

See Python README for full documentation.

Node.js / TypeScript
npm install @xberg-io/xberg

See Node.js README for full documentation.

Rust
cargo add xberg

See Rust README for full documentation.

Go
go get github.com/xberg-io/xberg/packages/go@latest

⚠️ The repository root is not a Go module — go get github.com/xberg-io/xberg will fail. Always target the /packages/go subdirectory as shown above.

See Go README for full documentation.

Java

Available on Maven Central as io.xberg:xberg. See Java README for the dependency snippet.

C#
dotnet add package XbergIo.Xberg

See C# README for full documentation.

Ruby
gem install xberg

See Ruby README for full documentation.

PHP
composer require xberg-io/xberg

See PHP README for full documentation.

Elixir

Add {:xberg, "~> 1.0"} to your mix.exs dependencies. See Elixir README for full documentation.

WebAssembly
npm install @xberg-io/xberg-wasm

See WebAssembly README for full documentation.

Kotlin (Android)

Available on Maven Central as io.xberg:xberg-android. See Kotlin README for the dependency snippet.

Swift

Add via Swift Package Manager. See Swift README for full documentation.

Dart / Flutter
dart pub add xberg

See Dart README for full documentation.

Zig

Add via zig fetch. See Zig README for full documentation.

C/C++ (FFI)

Build from source as part of this workspace. See C (FFI) README for full documentation.

CLI & Deployment

CLI Tool
brew install xberg-io/tap/xberg

Windows users can install the same binary through Scoop:

scoop bucket add xberg https://github.com/xberg-io/scoop-bucket
scoop install xberg

14 commands: extract, batch, detect, formats, version, cache, tree-sitter, doctor, serve, mcp, api, embed, chunk, and completions.

See CLI usage guide for detailed documentation.

Docker
docker pull ghcr.io/xberg-io/xberg:latest

Run in API, CLI, or MCP modes. See Docker guide for examples.

REST API Server
xberg serve --host 0.0.0.0 --port 8000

One POST endpoint handles all formats. Returns JSON or Markdown. Stream large files. See API server guide.

MCP Server
xberg mcp --transport stdio

9 tools (extract, extract_batch, detect_mime_type, cache_stats, list_formats, cache_clear, get_version, cache_manifest, cache_warm). 3 prompts (extract_document, extract_with_ocr, semantic_search). 4 resources (formats, models, OCR languages, embedding presets).

Add to Claude Desktop or Cursor:

{
  "mcpServers": {
    "xberg": { "command": "xberg", "args": ["mcp"] }
  }
}

See MCP integration guide.

AI Coding Assistants

Install the Xberg plugin from xberg-io/xberg. Ships extraction APIs, OCR backends, configuration, and language conventions.

Claude Code
/plugin marketplace add xberg-io/xberg
/plugin install xberg@xberg
Codex CLI
/plugins add https://github.com/xberg-io/xberg

Search for xberg and select Install Plugin.

Cursor

Settings → Plugins → Add from URL → https://github.com/xberg-io/xberg, then select xberg.

Gemini CLI
gemini extensions install https://github.com/xberg-io/xberg
Factory Droid
droid plugin marketplace add https://github.com/xberg-io/xberg
droid plugin install xberg@xberg
GitHub Copilot CLI
copilot plugin marketplace add https://github.com/xberg-io/xberg
copilot plugin install xberg@xberg
opencode

Add to opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@xberg-io/opencode-xberg"]
}

Quick Start

Extract text from a document:

use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result<()> {
    let config = ExtractionConfig::default();
    let output = extract(
        ExtractInput::from_uri("document.pdf"),
        &config
    ).await?;

    println!("{}", output.results[0].content);
    Ok(())
}

Common use cases — see Quick start guide for language-specific examples, OCR, batch processing, and API configuration.


Capabilities

Full feature list

Supported File Formats (107 formats · 141 file extensions · 56 MIME aliases)

107 formats across 140 unique file extensions, with 56 compatibility MIME aliases, intelligent format detection, and comprehensive metadata extraction.

Office Documents
Category Formats Capabilities
Word Processing .docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6 Full text, tables, images, metadata, styles
Spreadsheets .xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers Sheet data, formulas, cell metadata, charts
Presentations .pptx, .pptm, .ppt, .pps, .ppsx, .potx, .potm, .pot, .odp, .key Slides, speaker notes, images, metadata
PDF .pdf Text, tables, images, metadata, OCR support
eBooks .epub, .fb2 Chapters, metadata, embedded resources
Database .dbf, .sqlite, .sqlite3, .db, .gpkg, .gpkx Bounded table extraction, schema metadata, GeoPackage detection
Hangul .hwp, .hwpx Korean document format, text extraction
Images (OCR-Enabled)
Category Formats Features
Raster .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif OCR, table detection, EXIF metadata, dimensions, color space
Advanced .jp2, .jpg2, .j2c, .j2k, .jpc, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm OCR via pure-Rust JPEG2000 decoder, JBIG2 support, table detection
HEIC family .heic, .heics, .heif, .heifs, .hif, .avif, .avcs EXIF metadata, optional pixel decoding
Vector .svg DOM parsing, embedded text, graphics metadata
Audio & Video
Category Formats Features
Audio .mp3, .mpga, .m4a, .wav, .webm Whisper transcription
MP4 audio track .mp4, .mpg4, .mp4v, .m4v Audio-track transcription only
MPEG audio track .mpeg, .mpg, .mpe, .m1v, .m2v Audio-track transcription only
WebM audio track .webm Audio-track transcription only
Web & Data
Category Formats Features
Markup .html, .htm, .xhtml, .xht, .xml, .kml, .svg DOM parsing, metadata (Open Graph, Twitter Card), link extraction
Structured Data .json, .geojson, .jsonl, .ndjson, .yaml, .yml, .toml, .csv, .tsv Schema detection, nested structures, validation
Text & Markdown .txt, .adoc, .asciidoc, .vtt, .md, .markdown, .commonmark, .qmd, .rmd, .djot, .dj, .mdx, .doctags, .rst, .org, .rtf AsciiDoc, CommonMark, MyST Markdown, Quarto, R Markdown, Djot, MDX, DocTags, reStructuredText, Org Mode
Email & Archives
Category Formats Features
Email .eml, .msg, .pst Headers, body (HTML/plain), attachments, threading
Archives .zip, .tar, .tgz, .gz, .7z File listing, nested archives, metadata, recursive extraction
Academic & Scientific
Category Formats Features
Citations .bib, .ris, .nbib, .enw Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX
Scientific .tex, .latex, .typ, .typst, .jats, .nxml LaTeX, Typst, PubMed JATS
Text notebooks .ipynb, .md, .py, .R, .jl Jupyter, MyST-NB, Jupytext percent/light, saved outputs, cell visibility tags
Publishing .fb2, .docbook, .dbk, .docbook4, .docbook5, .opml FictionBook, DocBook XML, OPML outlines

Code Intelligence (371 Languages)

Extract structure from 371 programming languages via tree-sitter:

Feature Description
Structure Extraction Functions, classes, methods, structs, interfaces, enums
Import/Export Analysis Module dependencies, re-exports, wildcard imports
Symbol Extraction Variables, constants, type aliases, properties
Docstring Parsing Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats
Syntax-Aware Chunking Split code by semantic boundaries for RAG pipelines
Diagnostics Parse errors with line/column positions

Powered by tree-sitter-language-pack.

Output Formats (6)

Format Use case Example
Plain Raw text, no markup "Chapter 1\nIntroduction"
Markdown Readable, structured, RAG-friendly "# Chapter 1\n## Introduction"
Djot Modern lightweight markup Similar to Markdown but stricter
HTML Styled, browser-ready <h1>Chapter 1</h1>
JSON Machine-readable tree structure Hierarchical sections with heading levels
DocTags Docling-compatible tag stream for document elements and tables <text>Chapter 1</text>

Deployment Modes

Mode Command Transport Use case
Library xberg::extract() Async functions Embed in your application
CLI xberg extract document.pdf 14 commands Scripts, batch jobs, CI/CD
REST API xberg serve HTTP POST Microservice, serverless deployment
MCP Server xberg mcp stdio or HTTP Claude, Cursor, IDE agents
Docker docker run ghcr.io/xberg-io/xberg All modes Container deployment

OCR Backends

  • Tesseract — Native C FFI (Linux/macOS/Windows) and WASM (browser)
  • PaddleOCR — ONNX Runtime, mobile-optimized models
  • Candle — Pure Rust, CPU-only, lightweight
  • VLM — GPT-4 Vision, Claude Vision, Gemini Vision, or 165 providers via liter-llm

Fallback chains. Extensible via plugin system.

Embeddings

Local (ONNX Runtime):

  • Preset models: fast, balanced (default), quality, multilingual
  • Dimensions: 384, 768, 1024

Provider-hosted:

  • OpenAI, Anthropic, Google, Hugging Face, Mistral, Cohere, and 165 providers total
  • Via liter-llm integration

Reranking:

  • Local ONNX rerankers (cross-encoder models)
  • Provider-hosted: Cohere Rerank, others

Structured LLM Extraction

Local engines: Ollama, LM Studio, vLLM

Remote: OpenAI, Anthropic, Google, Mistral, Cohere, and 165 providers via liter-llm

Schema validation. Temperature, top-p, frequency penalty tuning.

Enrichment

  • NER — GLiNER or LLM-based entity recognition
  • Redaction — Mask PII (phone, email, SSN, credit card, addresses)
  • Summarization — Document and section summaries via LLM
  • Translation — Multi-language via LLM
  • Page Classification — Tag document pages (cover, toc, content, etc.)
  • QR Code Detection — Extract and decode QR codes from images
  • Keyword Extraction — YAKE or RAKE algorithms
  • Language Detection — Detect document language
  • Layout Detection — RT-DETR + TATR models for document structure
  • Table Extraction — Cell-level structure and content
  • Token Reduction — TOON wire format (~30–50% fewer tokens than JSON)

CLI Reference

All 14 commands
Command Subcommands Purpose
extract — Extract text from a single document (path, URL, or stdin)
batch — Extract from multiple documents in parallel
detect — Identify MIME type of a file
formats — List all supported formats and MIME types
version — Show Xberg version
cache stats, clear, manifest, warm Manage extraction cache and models
tree-sitter download, list, cache-dir, clean Manage code-intelligence grammars
doctor — Diagnose the local installation and runtime dependencies
serve — Start REST API server (default: http://127.0.0.1:8000)
mcp — Start MCP server (stdio or HTTP transport)
api schema Output OpenAPI 3.1 specification
embed — Generate embeddings for text (local or provider-hosted)
chunk — Split text into chunks (text, markdown, YAML, or semantic)
completions — Generate shell completion scripts

Run xberg --help or xberg <command> --help for detailed options.


Documentation

Full guides, API references for every binding, format reference, and configuration docs live at xberg.io.


Built with Xberg

Projects that declare Xberg as a dependency. Xberg was previously published as kreuzberg, and most of these projects declare the package under that name.

Project What it is Stars
basemind AI context and content layer for coding agents over one MCP server: code map, document RAG, shared memory and web crawl Stars
delulu A suite of MCP servers and CLI tools that give your LLM better search and fewer hallucinations Stars
docs-mcp-server Grounded documentation MCP server, an open-source alternative to Context7, Nia and Ref.Tools Stars
erato The open-source AI platform Stars
fastmail-cli CLI and MCP server for Fastmail: email, contacts, masked email, attachments and text extraction Stars
ghfdb-portal Web portal for the Global Heat Flow Database Stars
hawki-toolkit-file-converter Prepares and converts PDF files for the HAWKI toolkit Stars
haystack-core-integrations Integrations that extend Haystack with extra components and document stores Stars
kreuzakt A search engine for humans and computers, aimed at your most boring documents Stars
lilbee The whole local AI stack in one executable, with conversational search and cited answers over your files, code and the web Stars
llm-workflow-engine Power CLI and workflow manager for LLMs Stars
MANSPIDER Spiders entire networks for files sitting on SMB shares, searching filenames or contents with regex Stars
otoroshi-llm-extension Connect, secure and manage LLM models behind one OpenAI-compatible API Stars
sift-kg Turns a collection of documents into a knowledge graph, extracting entities and relationships with an LLM Stars
sirchmunk Turns raw data into a self-evolving, real-time search and intelligence layer Stars
support-chatbot Level-1 support chatbot for the Netherlands Red Cross 510 team Stars

Using Xberg in your project? Open a PR adding it to this list.


Contributing

Contributions are welcome! See CONTRIBUTING.md for guidelines.

Join our Discord community for questions and discussion.


Part of Xberg.io

  • Xberg — the open-source content-intelligence engine: text, tables, and metadata from 107 formats (141 file extensions), with OCR, transcription, and code intelligence. MIT.
  • Xberg Pro — a complete self-hosted content-intelligence backend in a single container. Commercial.
  • Xberg Enterprise — the distributed, governed content-intelligence platform, scaled on Kubernetes with team governance and support. Commercial.
  • crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
  • html-to-markdown — fast, lossless HTML→Markdown engine.
  • liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
  • tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
  • alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

License

MIT License (MIT) — see LICENSE for details.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

86 total
  1. v1.2.9v1.2.9Sep 24, 202635 downloads

    ### Added - **(ocr): the near-empty PDF fallback, embedded-image recognition and the per-page selection gate each have a setting of their own.** All three behaviours were keyed on one condition -- whether `ExtractionConfig.ocr` was present -- so a caller could not have one without the others. On a 731-page mixed native and scanned document that cost 102.8 s with an OCR block against 57.1 s without, while dropping the block lost the near-empty fallback (a scanned page carrying a text stamp fell from 996 to 6 characters) and picture text in DOCX files (250 to 36). `ocr_near_empty_fallback`, `ocr_scanned_page_quality_gate` and `ocr_embedded_images` are each `Option<bool>` whose `None` reproduces exactly the previous derived behaviour, so no existing caller changes; setting one to `Some(true)` without an `ocr` block reads `OcrConfig::default()` thresholds and still requires a registered automatic backend. A third copy of the embedded-image predicate was found and folded into the same helper -- `extraction/image_ocr.rs` returned early on `config.ocr.is_none()` independently of the two known sites, so the new setting would otherwise have been accepted and then silently ignored. `Extract

  2. v1.2.8v1.2.8Sep 23, 2026234 downloads

    ### Fixed - **(ocr): PDF OCR no longer upscales a page raster the layout pass already rendered.** The `source_dpi` hint that tells the image preprocessor how finely a page was sampled was derived only on the route that rendered the pages itself. When layout detection is on, OCR reuses the rasters the layout pass produced, and with the hint absent the preprocessor fell back to assuming 72 dpi. Both passes render at the same resolution -- `effective_pdf_render_dpi`, 150 by default -- so a correctly sampled raster was resampled toward the 300 dpi target by about 4.2x instead of the intended 2x, and clamped at the 4096 px ceiling: a US Letter page went to 3165 x 4096 where the OCR-rendered path produced 2550 x 3300. The extra pixels are interpolated, so they cost memory and recognition time and carry no detail the raster did not already hold. The resolution is now read from the raster's own pixel dimensions against the source page's MediaBox, which needs no extra PDF parse -- the same open already served the `/Rotate` hint on this route. A raster whose two axes disagree on the resolution they imply is not a whole-page MediaBox-oriented render of that page (a display-oriented render of

  3. v1.2.7v1.2.7Sep 22, 2026169 downloads

    ### Added - **(config): `ConcurrencyConfig::max_concurrent_ocr` and the `--max-concurrent-ocr` CLI flag set concurrent Tesseract recognition sessions on their own.** Use it when the host has cores to spare but not the memory to run a recognition session on each of them. The value is applied as given: neither the thread budget nor the host's free memory reduces it, since both of those bound only the automatic limit. The first extraction in a process fixes the session count for the rest of that process, because the admission semaphore and the Tesseract handle pool that enforce it are built once and the pool's capacity is fixed when it is constructed; a later extraction that names a different value keeps the first one and logs a one-time `WARN` naming both numbers. Set the value on the first extraction, or run one process per value. `ConcurrencyConfig` is not `#[non_exhaustive]`, so the added field breaks any Rust caller that builds the struct by literal without `..Default::default()`; such a caller must add the field or the rest pattern. Callers on every other binding are unaffected. (GH#1727) ### Changed - **(pdf): scan detection and fabricated-mapping OCR routing now use the thr

  4. v1.2.6v1.2.6Sep 20, 2026358 downloads

    ### Added - **(pdf/ocr): `OcrQualityThresholds::enable_plausibility_ocr_routing` and `OcrQualityThresholds::min_reliable_language_chunk_ratio` control language/dictionary-plausibility OCR routing, and `PdfMetadata.implausible_text_pages` reports which pages a page's decoded text failed to read as any real language.** (GH#1696) - **(config): `ExtractionConfig::runs_ocr_on_embedded_images` and `ExtractionConfig::wants_own_bytes_in_result` are exposed on every binding**, beside the existing `needs_image_data`. The first is the predicate the pipeline uses to decide whether a container's embedded images are OCR'd; the second is the pre-GH#1662 formula (`extract_images`, captioning or QR codes) that decides whether a standalone image's own bytes are echoed into `images`. (GH#1662) ### Changed - **(ocr): the `ocr`, `ocr-wasm`, and `ocr-pipeline` Cargo features now imply `language-detection`.** The language/dictionary-plausibility OCR-routing signal (GH#1696) is gated the same as the rest of the OCR module and would otherwise be a silent hole under a build that enables only `ocr-pipeline` (the VLM-only, Tesseract-free pipeline). whatlang (behind `language-detection`) depends only on `ha

  5. v1.2.5v1.2.5Sep 18, 2026458 downloads

    1.2.4 was tagged but never reached a package registry: its generated binding files still named 1.2.3, so its publish run was cancelled. 1.2.5 is the first published release carrying the 1.2.4 changes below in addition to its own. ### Fixed - **(python): a config field holding a nested data-carrying enum is no longer dropped by `from_json` / `to_json`.** The Python DTOs marked every field whose type transitively held a data enum as `#[serde(skip)]` -- `CaptioningConfig.llm` and `ChunkingConfig`'s nested options among them -- so a JSON payload round-tripped through such a config silently lost those fields. Regenerated on alef 0.92.1, whose Python emitter no longer treats a serializable enum wrapper as opaque. (alef#394) - **(ppt): a legacy `.ppt`'s embedded OLE objects are extracted.** A Word document or Excel sheet inserted as an object -- the way PowerPoint 97-2003 decks routinely carry a table -- lives in the deck's `ExOleObjStg` records, which the extractor walked over as opaque bytes, so the slide came out as its title and nothing else. The `.pptx` path already recursed into `ppt/embeddings/`; the legacy path now does the same: each embedded object's storage

Code frequency

additions and deletions
+1.6M-1.6MWeek of 2025-09-07: +13,451 linesWeek of 2025-09-07: -13,335 linesWeek of 2025-09-14: +15,875 linesWeek of 2025-09-14: -9,874 linesWeek of 2025-09-21: +7,738 linesWeek of 2025-09-21: -7,558 linesWeek of 2025-09-28: +28,471 linesWeek of 2025-09-28: -17,651 linesWeek of 2025-10-05: +1,493 linesWeek of 2025-10-05: -1,179 linesWeek of 2025-10-12: +4 linesWeek of 2025-10-12: -2 linesWeek of 2025-10-19: +0 linesWeek of 2025-10-19: -0 linesWeek of 2025-10-26: +7 linesWeek of 2025-10-26: -7 linesWeek of 2025-11-02: +252,287 linesWeek of 2025-11-02: -64,214 linesWeek of 2025-11-09: +39,972 linesWeek of 2025-11-09: -36,045 linesWeek of 2025-11-16: +84,133 linesWeek of 2025-11-16: -30,074 linesWeek of 2025-11-23: +111,499 linesWeek of 2025-11-23: -92,131 linesWeek of 2025-11-30: +87,164 linesWeek of 2025-11-30: -41,796 linesWeek of 2025-12-07: +105,979 linesWeek of 2025-12-07: -56,250 linesWeek of 2025-12-14: +71,387 linesWeek of 2025-12-14: -55,173 linesWeek of 2025-12-21: +154,351 linesWeek of 2025-12-21: -49,489 linesWeek of 2025-12-28: +254,388 linesWeek of 2025-12-28: -97,283 linesWeek of 2026-01-04: +72,763 linesWeek of 2026-01-04: -115,831 linesWeek of 2026-01-11: +16,987 linesWeek of 2026-01-11: -30,619 linesWeek of 2026-01-18: +96,614 linesWeek of 2026-01-18: -79,333 linesWeek of 2026-01-25: +250,729 linesWeek of 2026-01-25: -41,892 linesWeek of 2026-02-01: +56,133 linesWeek of 2026-02-01: -28,100 linesWeek of 2026-02-08: +477,705 linesWeek of 2026-02-08: -54,452 linesWeek of 2026-02-15: +123,420 linesWeek of 2026-02-15: -437,885 linesWeek of 2026-02-22: +74,569 linesWeek of 2026-02-22: -26,520 linesWeek of 2026-03-01: +46,076 linesWeek of 2026-03-01: -14,356 linesWeek of 2026-03-08: +89,017 linesWeek of 2026-03-08: -34,522 linesWeek of 2026-03-15: +149,525 linesWeek of 2026-03-15: -230,556 linesWeek of 2026-03-22: +107,456 linesWeek of 2026-03-22: -124,285 linesWeek of 2026-03-29: +255,855 linesWeek of 2026-03-29: -332,199 linesWeek of 2026-04-05: +65,413 linesWeek of 2026-04-05: -35,074 linesWeek of 2026-04-12: +30,476 linesWeek of 2026-04-12: -23,243 linesWeek of 2026-04-19: +482,046 linesWeek of 2026-04-19: -485,716 linesWeek of 2026-04-26: +614,721 linesWeek of 2026-04-26: -566,257 linesWeek of 2026-05-03: +728,855 linesWeek of 2026-05-03: -761,861 linesWeek of 2026-05-10: +756,984 linesWeek of 2026-05-10: -726,049 linesWeek of 2026-05-17: +534,748 linesWeek of 2026-05-17: -484,592 linesWeek of 2026-05-24: +299,416 linesWeek of 2026-05-24: -276,918 linesWeek of 2026-05-31: +324,181 linesWeek of 2026-05-31: -199,096 linesWeek of 2026-06-07: +647,002 linesWeek of 2026-06-07: -551,934 linesWeek of 2026-06-14: +848,363 linesWeek of 2026-06-14: -629,446 linesWeek of 2026-06-21: +1,474,191 linesWeek of 2026-06-21: -1,550,720 linesWeek of 2026-06-28: +1,042,465 linesWeek of 2026-06-28: -960,729 linesWeek of 2026-07-05: +637,984 linesWeek of 2026-07-05: -578,241 linesWeek of 2026-07-12: +97,190 linesWeek of 2026-07-12: -92,290 linesWeek of 2026-07-19: +401,229 linesWeek of 2026-07-19: -275,744 linesWeek of 2026-07-26: +478,527 linesWeek of 2026-07-26: -414,473 linesWeek of 2026-08-02: +591,024 linesWeek of 2026-08-02: -409,278 linesWeek of 2026-08-09: +421,552 linesWeek of 2026-08-09: -412,946 linesWeek of 2026-08-16: +676,805 linesWeek of 2026-08-16: -190,861 linesWeek of 2026-08-23: +556,620 linesWeek of 2026-08-23: -549,690 linesWeek of 2026-08-30: +77,525 linesWeek of 2026-08-30: -32,757 linesSep 7, 2025Aug 30, 2026
+14.8M lines added, -12.3M removed over the last year.

Commits per week

last 52 weeks
4430Week of 2025-09-28: 67 commitsWeek of 2025-10-05: 15 commitsWeek of 2025-10-12: 1 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 2 commitsWeek of 2025-11-02: 98 commitsWeek of 2025-11-09: 173 commitsWeek of 2025-11-16: 199 commitsWeek of 2025-11-23: 192 commitsWeek of 2025-11-30: 187 commitsWeek of 2025-12-07: 241 commitsWeek of 2025-12-14: 238 commitsWeek of 2025-12-21: 287 commitsWeek of 2025-12-28: 194 commitsWeek of 2026-01-04: 150 commitsWeek of 2026-01-11: 83 commitsWeek of 2026-01-18: 89 commitsWeek of 2026-01-25: 121 commitsWeek of 2026-02-01: 93 commitsWeek of 2026-02-08: 115 commitsWeek of 2026-02-15: 104 commitsWeek of 2026-02-22: 117 commitsWeek of 2026-03-01: 77 commitsWeek of 2026-03-08: 154 commitsWeek of 2026-03-15: 196 commitsWeek of 2026-03-22: 132 commitsWeek of 2026-03-29: 220 commitsWeek of 2026-04-05: 109 commitsWeek of 2026-04-12: 100 commitsWeek of 2026-04-19: 152 commitsWeek of 2026-04-26: 178 commitsWeek of 2026-05-03: 333 commitsWeek of 2026-05-10: 255 commitsWeek of 2026-05-17: 231 commitsWeek of 2026-05-24: 217 commitsWeek of 2026-05-31: 225 commitsWeek of 2026-06-07: 106 commitsWeek of 2026-06-14: 237 commitsWeek of 2026-06-21: 214 commitsWeek of 2026-06-28: 219 commitsWeek of 2026-07-05: 178 commitsWeek of 2026-07-12: 78 commitsWeek of 2026-07-19: 334 commitsWeek of 2026-07-26: 320 commitsWeek of 2026-08-02: 443 commitsWeek of 2026-08-09: 125 commitsWeek of 2026-08-16: 284 commitsWeek of 2026-08-23: 281 commitsWeek of 2026-08-30: 180 commitsWeek of 2026-09-06: 144 commitsWeek of 2026-09-13: 144 commitsWeek of 2026-09-20: 205 commitsSep 28, 2025Sep 20, 2026
8.8K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 7 commitsSun 1:00 — 12 commitsSun 2:00 — 8 commitsSun 3:00 — 12 commitsSun 4:00 — 2 commitsSun 5:00 — 1 commitsSun 6:00 — 23 commitsSun 7:00 — 48 commitsSun 8:00 — 94 commitsSun 9:00 — 81 commitsSun 10:00 — 67 commitsSun 11:00 — 72 commitsSun 12:00 — 58 commitsSun 13:00 — 93 commitsSun 14:00 — 99 commitsSun 15:00 — 126 commitsSun 16:00 — 92 commitsSun 17:00 — 75 commitsSun 18:00 — 82 commitsSun 19:00 — 83 commitsSun 20:00 — 75 commitsSun 21:00 — 53 commitsSun 22:00 — 42 commitsSun 23:00 — 10 commitsMon 0:00 — 9 commitsMon 1:00 — 12 commitsMon 2:00 — 2 commitsMon 3:00 — 6 commitsMon 4:00 — 0 commitsMon 5:00 — 10 commitsMon 6:00 — 32 commitsMon 7:00 — 80 commitsMon 8:00 — 78 commitsMon 9:00 — 58 commitsMon 10:00 — 80 commitsMon 11:00 — 62 commitsMon 12:00 — 40 commitsMon 13:00 — 56 commitsMon 14:00 — 61 commitsMon 15:00 — 53 commitsMon 16:00 — 100 commitsMon 17:00 — 85 commitsMon 18:00 — 71 commitsMon 19:00 — 67 commitsMon 20:00 — 75 commitsMon 21:00 — 43 commitsMon 22:00 — 45 commitsMon 23:00 — 20 commitsTue 0:00 — 22 commitsTue 1:00 — 11 commitsTue 2:00 — 12 commitsTue 3:00 — 12 commitsTue 4:00 — 8 commitsTue 5:00 — 24 commitsTue 6:00 — 17 commitsTue 7:00 — 61 commitsTue 8:00 — 91 commitsTue 9:00 — 85 commitsTue 10:00 — 81 commitsTue 11:00 — 65 commitsTue 12:00 — 54 commitsTue 13:00 — 64 commitsTue 14:00 — 68 commitsTue 15:00 — 76 commitsTue 16:00 — 108 commitsTue 17:00 — 80 commitsTue 18:00 — 72 commitsTue 19:00 — 59 commitsTue 20:00 — 85 commitsTue 21:00 — 49 commitsTue 22:00 — 31 commitsTue 23:00 — 20 commitsWed 0:00 — 28 commitsWed 1:00 — 16 commitsWed 2:00 — 7 commitsWed 3:00 — 8 commitsWed 4:00 — 6 commitsWed 5:00 — 13 commitsWed 6:00 — 52 commitsWed 7:00 — 86 commitsWed 8:00 — 100 commitsWed 9:00 — 83 commitsWed 10:00 — 81 commitsWed 11:00 — 62 commitsWed 12:00 — 44 commitsWed 13:00 — 42 commitsWed 14:00 — 75 commitsWed 15:00 — 88 commitsWed 16:00 — 92 commitsWed 17:00 — 67 commitsWed 18:00 — 79 commitsWed 19:00 — 61 commitsWed 20:00 — 84 commitsWed 21:00 — 87 commitsWed 22:00 — 39 commitsWed 23:00 — 17 commitsThu 0:00 — 6 commitsThu 1:00 — 12 commitsThu 2:00 — 5 commitsThu 3:00 — 10 commitsThu 4:00 — 3 commitsThu 5:00 — 11 commitsThu 6:00 — 38 commitsThu 7:00 — 60 commitsThu 8:00 — 63 commitsThu 9:00 — 89 commitsThu 10:00 — 81 commitsThu 11:00 — 98 commitsThu 12:00 — 71 commitsThu 13:00 — 68 commitsThu 14:00 — 88 commitsThu 15:00 — 117 commitsThu 16:00 — 158 commitsThu 17:00 — 87 commitsThu 18:00 — 57 commitsThu 19:00 — 83 commitsThu 20:00 — 103 commitsThu 21:00 — 78 commitsThu 22:00 — 74 commitsThu 23:00 — 37 commitsFri 0:00 — 30 commitsFri 1:00 — 26 commitsFri 2:00 — 13 commitsFri 3:00 — 7 commitsFri 4:00 — 6 commitsFri 5:00 — 7 commitsFri 6:00 — 36 commitsFri 7:00 — 86 commitsFri 8:00 — 68 commitsFri 9:00 — 63 commitsFri 10:00 — 59 commitsFri 11:00 — 75 commitsFri 12:00 — 69 commitsFri 13:00 — 54 commitsFri 14:00 — 94 commitsFri 15:00 — 61 commitsFri 16:00 — 83 commitsFri 17:00 — 123 commitsFri 18:00 — 101 commitsFri 19:00 — 56 commitsFri 20:00 — 98 commitsFri 21:00 — 101 commitsFri 22:00 — 32 commitsFri 23:00 — 23 commitsSat 0:00 — 4 commitsSat 1:00 — 8 commitsSat 2:00 — 10 commitsSat 3:00 — 12 commitsSat 4:00 — 14 commitsSat 5:00 — 14 commitsSat 6:00 — 15 commitsSat 7:00 — 62 commitsSat 8:00 — 128 commitsSat 9:00 — 91 commitsSat 10:00 — 72 commitsSat 11:00 — 89 commitsSat 12:00 — 97 commitsSat 13:00 — 101 commitsSat 14:00 — 110 commitsSat 15:00 — 96 commitsSat 16:00 — 105 commitsSat 17:00 — 91 commitsSat 18:00 — 94 commitsSat 19:00 — 103 commitsSat 20:00 — 100 commitsSat 21:00 — 72 commitsSat 22:00 — 32 commitsSat 23:00 — 25 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Jan 13, 2026daily#13+189
Jan 12, 2026daily#14+214
Jan 11, 2026daily#7+299
  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    373.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    285.8K stars · Python

  • tensorflow/tensorflow

    An Open Source Machine Learning Framework for Everyone

    200.7K stars · C++

  • yt-dlp/yt-dlp

    A feature-rich command-line audio/video downloader

    195.5K stars · Python

  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    195.2K stars · Rust

  • Significant-Gravitas/AutoGPT

    AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

    187.7K stars · Python