xberg-io/xbergPublic

A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.

AI summary: High-performance polyglot document extraction framework built in Rust.

Stars
8.9K
+14 today
Forks
536
Watchers
31
Open issues
8
Open PRs
0
Contributors
~55
Commits
7.8K
Branches
49

RustMITCreated Jan 31, 2025Last push 1d agoLatest release v1.0.3+196 stars this week+196 this month

Star history

since Jul 29, 2026
02.5K5K7.5KJul 2026Jul 2026Aug 2026Aug 2026
8.9K stars as of Aug 6, 2026, tracked back to Jul 29, 2026.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulMonWedFri2025-08-03: 0 commits2025-08-04: 0 commits2025-08-05: 0 commits2025-08-06: 0 commits2025-08-07: 0 commits2025-08-08: 0 commits2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 2 commits2025-08-12: 1 commit2025-08-13: 8 commits2025-08-14: 0 commits2025-08-15: 2 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 2 commits2025-08-24: 13 commits2025-08-25: 3 commits2025-08-26: 1 commit2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 19 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 8 commits2025-09-03: 16 commits2025-09-04: 18 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 7 commits2025-09-10: 4 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 24 commits2025-09-14: 13 commits2025-09-15: 15 commits2025-09-16: 8 commits2025-09-17: 9 commits2025-09-18: 0 commits2025-09-19: 4 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 14 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 6 commits2025-09-28: 30 commits2025-09-29: 26 commits2025-09-30: 7 commits2025-10-01: 4 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 1 commit2025-10-07: 0 commits2025-10-08: 1 commit2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 13 commits2025-10-12: 1 commit2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 2 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 3 commits2025-11-05: 2 commits2025-11-06: 0 commits2025-11-07: 73 commits2025-11-08: 20 commits2025-11-09: 27 commits2025-11-10: 18 commits2025-11-11: 15 commits2025-11-12: 11 commits2025-11-13: 37 commits2025-11-14: 36 commits2025-11-15: 29 commits2025-11-16: 15 commits2025-11-17: 16 commits2025-11-18: 19 commits2025-11-19: 12 commits2025-11-20: 57 commits2025-11-21: 55 commits2025-11-22: 25 commits2025-11-23: 34 commits2025-11-24: 11 commits2025-11-25: 13 commits2025-11-26: 17 commits2025-11-27: 67 commits2025-11-28: 25 commits2025-11-29: 25 commits2025-11-30: 33 commits2025-12-01: 23 commits2025-12-02: 28 commits2025-12-03: 24 commits2025-12-04: 34 commits2025-12-05: 17 commits2025-12-06: 28 commits2025-12-07: 38 commits2025-12-08: 30 commits2025-12-09: 20 commits2025-12-10: 26 commits2025-12-11: 28 commits2025-12-12: 41 commits2025-12-13: 58 commits2025-12-14: 22 commits2025-12-15: 35 commits2025-12-16: 32 commits2025-12-17: 26 commits2025-12-18: 23 commits2025-12-19: 47 commits2025-12-20: 53 commits2025-12-21: 70 commits2025-12-22: 45 commits2025-12-23: 36 commits2025-12-24: 34 commits2025-12-25: 6 commits2025-12-26: 32 commits2025-12-27: 64 commits2025-12-28: 36 commits2025-12-29: 36 commits2025-12-30: 24 commits2025-12-31: 33 commits2026-01-01: 8 commits2026-01-02: 29 commits2026-01-03: 28 commits2026-01-04: 20 commits2026-01-05: 16 commits2026-01-06: 13 commits2026-01-07: 24 commits2026-01-08: 22 commits2026-01-09: 20 commits2026-01-10: 35 commits2026-01-11: 20 commits2026-01-12: 9 commits2026-01-13: 13 commits2026-01-14: 13 commits2026-01-15: 6 commits2026-01-16: 11 commits2026-01-17: 11 commits2026-01-18: 18 commits2026-01-19: 14 commits2026-01-20: 14 commits2026-01-21: 14 commits2026-01-22: 16 commits2026-01-23: 5 commits2026-01-24: 8 commits2026-01-25: 17 commits2026-01-26: 15 commits2026-01-27: 17 commits2026-01-28: 23 commits2026-01-29: 10 commits2026-01-30: 14 commits2026-01-31: 25 commits2026-02-01: 9 commits2026-02-02: 10 commits2026-02-03: 5 commits2026-02-04: 14 commits2026-02-05: 13 commits2026-02-06: 21 commits2026-02-07: 21 commits2026-02-08: 18 commits2026-02-09: 27 commits2026-02-10: 14 commits2026-02-11: 21 commits2026-02-12: 18 commits2026-02-13: 10 commits2026-02-14: 7 commits2026-02-15: 13 commits2026-02-16: 16 commits2026-02-17: 18 commits2026-02-18: 17 commits2026-02-19: 14 commits2026-02-20: 9 commits2026-02-21: 17 commits2026-02-22: 6 commits2026-02-23: 7 commits2026-02-24: 35 commits2026-02-25: 23 commits2026-02-26: 25 commits2026-02-27: 8 commits2026-02-28: 13 commits2026-03-01: 5 commits2026-03-02: 11 commits2026-03-03: 8 commits2026-03-04: 12 commits2026-03-05: 17 commits2026-03-06: 7 commits2026-03-07: 17 commits2026-03-08: 21 commits2026-03-09: 9 commits2026-03-10: 15 commits2026-03-11: 23 commits2026-03-12: 12 commits2026-03-13: 22 commits2026-03-14: 52 commits2026-03-15: 23 commits2026-03-16: 13 commits2026-03-17: 19 commits2026-03-18: 28 commits2026-03-19: 38 commits2026-03-20: 39 commits2026-03-21: 36 commits2026-03-22: 10 commits2026-03-23: 29 commits2026-03-24: 31 commits2026-03-25: 24 commits2026-03-26: 20 commits2026-03-27: 17 commits2026-03-28: 1 commit2026-03-29: 12 commits2026-03-30: 64 commits2026-03-31: 26 commits2026-04-01: 38 commits2026-04-02: 29 commits2026-04-03: 21 commits2026-04-04: 30 commits2026-04-05: 26 commits2026-04-06: 21 commits2026-04-07: 18 commits2026-04-08: 12 commits2026-04-09: 13 commits2026-04-10: 12 commits2026-04-11: 7 commits2026-04-12: 4 commits2026-04-13: 13 commits2026-04-14: 36 commits2026-04-15: 4 commits2026-04-16: 11 commits2026-04-17: 22 commits2026-04-18: 10 commits2026-04-19: 18 commits2026-04-20: 46 commits2026-04-21: 7 commits2026-04-22: 13 commits2026-04-23: 18 commits2026-04-24: 20 commits2026-04-25: 30 commits2026-04-26: 25 commits2026-04-27: 14 commits2026-04-28: 16 commits2026-04-29: 38 commits2026-04-30: 48 commits2026-05-01: 10 commits2026-05-02: 27 commits2026-05-03: 19 commits2026-05-04: 22 commits2026-05-05: 44 commits2026-05-06: 9 commits2026-05-07: 90 commits2026-05-08: 76 commits2026-05-09: 73 commits2026-05-10: 86 commits2026-05-11: 17 commits2026-05-12: 15 commits2026-05-13: 56 commits2026-05-14: 30 commits2026-05-15: 40 commits2026-05-16: 11 commits2026-05-17: 31 commits2026-05-18: 28 commits2026-05-19: 19 commits2026-05-20: 19 commits2026-05-21: 65 commits2026-05-22: 51 commits2026-05-23: 18 commits2026-05-24: 18 commits2026-05-25: 32 commits2026-05-26: 28 commits2026-05-27: 9 commits2026-05-28: 43 commits2026-05-29: 18 commits2026-05-30: 69 commits2026-05-31: 47 commits2026-06-01: 17 commits2026-06-02: 45 commits2026-06-03: 37 commits2026-06-04: 31 commits2026-06-05: 30 commits2026-06-06: 18 commits2026-06-07: 23 commits2026-06-08: 12 commits2026-06-09: 14 commits2026-06-10: 20 commits2026-06-11: 15 commits2026-06-12: 15 commits2026-06-13: 7 commits2026-06-14: 18 commits2026-06-15: 9 commits2026-06-16: 24 commits2026-06-17: 90 commits2026-06-18: 33 commits2026-06-19: 10 commits2026-06-20: 53 commits2026-06-21: 11 commits2026-06-22: 24 commits2026-06-23: 24 commits2026-06-24: 29 commits2026-06-25: 45 commits2026-06-26: 23 commits2026-06-27: 58 commits2026-06-28: 49 commits2026-06-29: 29 commits2026-06-30: 22 commits2026-07-01: 7 commits2026-07-02: 67 commits2026-07-03: 33 commits2026-07-04: 12 commits2026-07-05: 25 commits2026-07-06: 18 commits2026-07-07: 40 commits2026-07-08: 23 commits2026-07-09: 40 commits2026-07-10: 12 commits2026-07-11: 20 commits2026-07-12: 33 commits2026-07-13: 21 commits2026-07-14: 2 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 17 commits2026-07-18: 5 commits2026-07-19: 30 commits2026-07-20: 54 commits2026-07-21: 47 commits2026-07-22: 58 commits2026-07-23: 56 commits2026-07-24: 50 commits2026-07-25: 39 commits2026-07-26: 41 commits2026-07-27: 39 commits2026-07-28: 42 commits2026-07-29: 34 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits
7,064 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Very active

    7,064 commits in 52 weeks

  • Well documented

    High community health score

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

What xberg does

Xberg is a powerful document intelligence framework designed to extract text, metadata, and structured data from over 98 different file formats. Built with a core engine in Rust for safety and speed, it handles complex documents like PDFs and Office files with ease. It goes beyond simple text scraping by maintaining structural integrity, such as recognizing tables and extracting embedded images. Its standout feature is its massive cross-language support, providing native bindings for almost every major programming language and environment.

Data engineers, backend developers, and AI builders constructing RAG systems or document processing pipelines. Suitable for almost any tech stack.

  • Massive format support: Accurately parses over 98 file types including PDFs, DOCX, and images via OCR.
  • Polyglot bindings: Native libraries available for Python, Node, Java, Go, C#, Ruby, PHP, and more.
  • Structural extraction: Intelligently identifies and extracts tables, metadata, and embedded media, not just raw text.
  • Rust core: Ensures blazing fast performance and memory safety, critical for processing untrusted documents.
  • Flexible deployment: Can be used as a CLI tool, a REST API, or an MCP server for AI integration.

Where teams use it

RAG pipeline ingestion

Data engineers use Xberg to reliably chunk and parse massive corporate document troves for vector databases.

Automated data entry

Financial systems extract structured table data from vendor invoices without manual transcription.

Enterprise search

Search indexing services extract deep metadata and text from mixed legacy file formats.

Getting started: Install the package for your specific language (e.g., npm install @xberg-io/xberg or pip install xberg).

README

main branch

Xberg

Extract clean text, tables, and structured data from documents and code — no format detection, no OCR setup, no stitched-together libraries. One engine, 15 language bindings, runs anywhere.

Xberg is the next iteration of Kreuzberg. Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line.

Feed documents → get clean text, tables, metadata, transcripts, code intelligence · Run it library, CLI, REST API, or MCP server · No GPU needed · Stream multi-GB files · Cache results.

Documents · Images · Spreadsheets · Email · Archives · Code · Audio · Video

crates.io npm PyPI License: MIT

Quick start · What you get · Capabilities · CLI · Docs


Feed any document—get structured text. Extract, batch, stream, or crawl.


What you get

Point Xberg at anything — a PDF, a spreadsheet, a scanned image, an audio file, a source tree — and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server.

What it does How
Extract from 98 formats PDFs, Office, images, HTML, email, archives, scientific publications, and code — intelligent MIME detection, streaming for large files.
6 output formats Plain text, Markdown, Djot, HTML, JSON tree structure, or Structured (JSON with OCR metadata and bounding boxes).
Code intelligence Functions, classes, imports, symbols, docstrings from 306 programming languages. Syntax-aware chunking for RAG pipelines.
Crawl & recurse Follow URLs, extract documents from within documents (nested archives, embedded PDFs). Auto/Document/Crawl modes.
OCR on demand Tesseract, PaddleOCR, Candle, or VLM backends — fallback chains, extensible via plugins. Confidence scores. Language auto-detection.
Transcription Whisper ONNX for audio/video tracks (MP3, M4A, WAV, WebM, MP4).
Embeddings & search Local (ONNX models) or provider-hosted (OpenAI, Anthropic, Google, 143 providers via liter-llm). Reranking.
Structured outputs LLM-powered extraction — local (Ollama, LM Studio, vLLM) or remote (OpenAI, Anthropic, Google).
Enrichment NER, redaction, summarization, translation, QR code detection, page classification, keyword extraction (YAKE/RAKE), language detection, layout detection, table extraction, token reduction (TOON).
Batch & parallel Process 100s of documents in parallel. Per-file timeouts. Configurable batch concurrency (max_concurrent_extractions).
Caching Content-hash cache keys — skip re-extraction when the file and config are unchanged.
Deployment Library, CLI (12 commands), REST API (xberg serve), MCP server (9 tools, 3 prompts, 4 resources), Docker.

Installation

Language Packages

Python
pip install xberg

See Python README for full documentation.

Node.js / TypeScript
npm install @xberg-io/xberg

See Node.js README for full documentation.

Rust
cargo add xberg

See Rust README for full documentation.

Go
go get github.com/xberg-io/xberg/packages/go@latest

⚠️ The repository root is not a Go module — go get github.com/xberg-io/xberg will fail. Always target the /packages/go subdirectory as shown above.

See Go README for full documentation.

Java

Available on Maven Central as io.xberg:xberg. See Java README for the dependency snippet.

C#
dotnet add package Xberg

See C# README for full documentation.

Ruby
gem install xberg

See Ruby README for full documentation.

PHP
composer require xberg-io/xberg

See PHP README for full documentation.

Elixir

Add {:xberg, "~> 1.0"} to your mix.exs dependencies. See Elixir README for full documentation.

WebAssembly
npm install @xberg-io/xberg-wasm

See WebAssembly README for full documentation.

Kotlin (Android)

Available on Maven Central as io.xberg:xberg-android. See Kotlin README for the dependency snippet.

Swift

Add via Swift Package Manager. See Swift README for full documentation.

Dart / Flutter
dart pub add xberg

See Dart README for full documentation.

Zig

Add via zig fetch. See Zig README for full documentation.

C/C++ (FFI)

Build from source as part of this workspace. See C (FFI) README for full documentation.

CLI & Deployment

CLI Tool
brew install xberg-io/tap/xberg

12 commands: extract, batch, detect, formats, version, cache (stats/clear/manifest/warm), serve, mcp, api, embed, chunk, completions.

See CLI usage guide for detailed documentation.

Docker
docker pull ghcr.io/xberg-io/xberg:latest

Run in API, CLI, or MCP modes. See Docker guide for examples.

REST API Server
xberg serve --host 0.0.0.0 --port 8000

One POST endpoint handles all formats. Returns JSON or Markdown. Stream large files. See API server guide.

MCP Server
xberg mcp --transport stdio

9 tools (extract, extract_batch, detect_mime_type, cache_stats, list_formats, cache_clear, get_version, cache_manifest, cache_warm). 3 prompts (extract_document, extract_with_ocr, semantic_search). 4 resources (formats, models, OCR languages, embedding presets).

Add to Claude Desktop or Cursor:

{
  "mcpServers": {
    "xberg": { "command": "xberg", "args": ["mcp"] }
  }
}

See MCP integration guide.

AI Coding Assistants

Install the Xberg plugin from xberg-io/xberg. Ships extraction APIs, OCR backends, configuration, and language conventions.

Claude Code
/plugin marketplace add xberg-io/xberg
/plugin install xberg@xberg
Codex CLI
/plugins add https://github.com/xberg-io/xberg

Search for xberg and select Install Plugin.

Cursor

Settings → Plugins → Add from URL → https://github.com/xberg-io/xberg, then select xberg.

Gemini CLI
gemini extensions install https://github.com/xberg-io/xberg
Factory Droid
droid plugin marketplace add https://github.com/xberg-io/xberg
droid plugin install xberg@xberg
GitHub Copilot CLI
copilot plugin marketplace add https://github.com/xberg-io/xberg
copilot plugin install xberg@xberg
opencode

Add to opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@xberg-io/opencode-xberg"]
}

Quick Start

Extract text from a document:

use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result<()> {
    let config = ExtractionConfig::default();
    let output = extract(
        ExtractInput::from_uri("document.pdf"),
        &config
    ).await?;

    println!("{}", output.results[0].content);
    Ok(())
}

Common use cases — see Quick start guide for language-specific examples, OCR, batch processing, and API configuration.


Capabilities

Full feature list

Supported File Formats (98)

98 file formats across 8 major categories with intelligent format detection and comprehensive metadata extraction.

Office Documents

Category Formats Capabilities
Word Processing .docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6 Full text, tables, images, metadata, styles
Spreadsheets .xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers Sheet data, formulas, cell metadata, charts
Presentations .pptx, .pptm, .ppt, .ppsx, .potx, .potm, .pot, .odp, .key Slides, speaker notes, images, metadata
PDF .pdf Text, tables, images, metadata, OCR support
eBooks .epub, .fb2 Chapters, metadata, embedded resources
Database .dbf Table data extraction, field type support
Hangul .hwp, .hwpx Korean document format, text extraction

Images (OCR-Enabled)

Category Formats Features
Raster .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif OCR, table detection, EXIF metadata, dimensions, color space
Advanced .jp2, .jpx, .jpm, .mj2, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm OCR via pure-Rust JPEG2000 decoder, JBIG2 support, table detection
HEIC family .heic, .heics, .heif, .avif, .avcs EXIF metadata, optional pixel decoding
Vector .svg DOM parsing, embedded text, graphics metadata

Audio & Video

Category Formats Features
Audio .mp3, .mpga, .m4a, .wav, .webm Whisper transcription
Video audio track .mp4, .mpeg, .webm Audio-track transcription only

Web & Data

Category Formats Features
Markup .html, .htm, .xhtml, .xml, .svg DOM parsing, metadata (Open Graph, Twitter Card), link extraction
Structured Data .json, .yaml, .yml, .toml, .csv, .tsv Schema detection, nested structures, validation
Text & Markdown .txt, .md, .markdown, .djot, .mdx, .rst, .org, .rtf CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode

Email & Archives

Category Formats Features
Email .eml, .msg, .pst Headers, body (HTML/plain), attachments, threading
Archives .zip, .tar, .tgz, .gz, .7z File listing, nested archives, metadata, recursive extraction

Academic & Scientific

Category Formats Features
Citations .bib, .ris, .nbib, .enw Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX
Scientific .tex, .latex, .typ, .typst, .jats, .ipynb LaTeX, Typst, Jupyter notebooks, PubMed JATS
Publishing .fb2, .docbook, .dbk, .docbook4, .docbook5, .opml FictionBook, DocBook XML, OPML outlines

Code Intelligence (306 Languages)

Extract structure from 306 programming languages via tree-sitter:

Feature Description
Structure Extraction Functions, classes, methods, structs, interfaces, enums
Import/Export Analysis Module dependencies, re-exports, wildcard imports
Symbol Extraction Variables, constants, type aliases, properties
Docstring Parsing Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats
Syntax-Aware Chunking Split code by semantic boundaries for RAG pipelines
Diagnostics Parse errors with line/column positions

Powered by tree-sitter-language-pack.

Output Formats (6)

Format Use case Example
Plain Raw text, no markup "Chapter 1\nIntroduction"
Markdown Readable, structured, RAG-friendly "# Chapter 1\n## Introduction"
Djot Modern lightweight markup Similar to Markdown but stricter
HTML Styled, browser-ready <h1>Chapter 1</h1>
JSON Machine-readable tree structure Hierarchical sections with heading levels
Structured OCR metadata, bounding boxes JSON with elements[] containing {text, bbox, confidence}

Deployment Modes

Mode Command Transport Use case
Library xberg::extract() Async functions Embed in your application
CLI xberg extract document.pdf 12 commands Scripts, batch jobs, CI/CD
REST API xberg serve HTTP POST Microservice, serverless deployment
MCP Server xberg mcp stdio or HTTP Claude, Cursor, IDE agents
Docker docker run ghcr.io/xberg-io/xberg All modes Container deployment

OCR Backends

  • Tesseract — Native C FFI (Linux/macOS/Windows) and WASM (browser)
  • PaddleOCR — ONNX Runtime, mobile-optimized models
  • Candle — Pure Rust, CPU-only, lightweight
  • VLM — GPT-4 Vision, Claude Vision, Gemini Vision, or 143 providers via liter-llm

Fallback chains. Extensible via plugin system.

Embeddings

Local (ONNX Runtime):

  • Preset models: fast, balanced (default), quality, multilingual
  • Dimensions: 384, 768, 1024

Provider-hosted:

  • OpenAI, Anthropic, Google, Hugging Face, Mistral, Cohere, and 143 providers total
  • Via liter-llm integration

Reranking:

  • Local ONNX rerankers (cross-encoder models)
  • Provider-hosted: Cohere Rerank, others

Structured LLM Extraction

Local engines: Ollama, LM Studio, vLLM

Remote: OpenAI, Anthropic, Google, Mistral, Cohere, and 143 providers via liter-llm

Schema validation. Temperature, top-p, frequency penalty tuning.

Enrichment

  • NER — GLiNER or LLM-based entity recognition
  • Redaction — Mask PII (phone, email, SSN, credit card, addresses)
  • Summarization — Document and section summaries via LLM
  • Translation — Multi-language via LLM
  • Page Classification — Tag document pages (cover, toc, content, etc.)
  • QR Code Detection — Extract and decode QR codes from images
  • Keyword Extraction — YAKE or RAKE algorithms
  • Language Detection — Detect document language
  • Layout Detection — RT-DETR + TATR models for document structure
  • Table Extraction — Cell-level structure and content
  • Token Reduction — TOON wire format (~30–50% fewer tokens than JSON)

CLI Reference

All 12 commands
Command Subcommands Purpose
extract Extract text from a single document (path, URL, or stdin)
batch Extract from multiple documents in parallel
detect Identify MIME type of a file
formats List all 98 supported formats and MIME types
version Show Xberg version
cache stats, clear, manifest, warm Manage extraction cache and models
serve Start REST API server (default: http://127.0.0.1:8000)
mcp Start MCP server (stdio or HTTP transport)
api schema Output OpenAPI 3.1 specification
embed Generate embeddings for text (local or provider-hosted)
chunk Split text into chunks (text, markdown, YAML, or semantic)
completions Generate shell completion scripts

Run xberg --help or xberg <command> --help for detailed options.


Documentation

Full guides, API references for every binding, format reference, and configuration docs live at xberg.io.


Contributing

Contributions are welcome! See CONTRIBUTING.md for guidelines.

Join our Discord community for questions and discussion.


Part of Xberg.dev

Xberg is one of six open-source projects from Kreuzberg, Inc.:

  • Xberg — document intelligence: text, tables, metadata from 98+ formats with optional OCR.
  • Xberg Enterprise — managed extraction API with SDKs, dashboards, and observability.
  • crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
  • html-to-markdown — fast, lossless HTML→Markdown engine.
  • liter-llm — universal LLM API client with native bindings for 14 languages and 143 providers.
  • tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
  • alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

License

MIT License (MIT) — see LICENSE for details.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

55 total
  1. v1.0.3v1.0.3Jul 29, 2026135 downloads

    Patch release. Completes the 1.0.2 rollout (core registry packages never published under 1.0.2) and lands a set of OCR/PDF extraction fixes plus the rmcp 3.0 upgrade. ### Fixed - VLM OCR now honors `XBERG_LLM_*` env credentials when a custom `base_url` is set, and `vlm_fallback` `on_low_quality` fires for bare images; per-stage OCR pipeline failures surface as processing warnings (#1339). - Tesseract no longer creates cache directories when caching is disabled (#1336). - Light-text-on-dark-background scans are auto-inverted before OCR, and the `invert_colors` config is honored as an explicit override (#1337). - NER and summarization processors are now compiled into container builds; whole-document text failures route to OCR under `ScannedPages` (#1338, partial — the `Auto`-strategy case still needs a repro). - Sparse continuation rows no longer flatten numeric line-item tables (#1333). - Linux builds without CUDA or TensorRT no longer fail under strict warnings on an unused ONNX Runtime execution-provider trait import. - Label-heavy financial tables are recovered and stitched without merging independent aligned tables. - PDF/OCR Markdown: explicit word boundaries and changelog he

  2. v1.0.2v1.0.2Jul 29, 202619 downloads

    Release v1.0.2

  3. v1.0.1v1.0.1Jul 28, 2026101 downloads

    xberg 1.0.1 is a maintenance release that fixes two extraction bugs and republishes the WordPerfect support crate so command-line builds link cleanly from crates.io. ## Fixed - **#1321 — Borderless tables on mixed pages.** Text-heavy borderless tables are now recovered even when the same page also contains an ML-detected table. The geometric-table fallback runs per region instead of per page, so a single model `Table` hint no longer suppresses grid recovery for the rest of the page, and words already inside a detected table are excluded to avoid double detection. - **#1326 — RTF font charset decoding.** RTF hex byte escapes now decode through the active font's `\fcharsetN` charset (mapped to a Windows codepage), falling back to `\ansicpgNNNN` and then Windows-1252. Documents that declare a Cyrillic or other non-ANSI font in the font table decode as readable text instead of mojibake, with font switches tracked across nested groups. ## Packaging - Republishes `xberg-libwpd` with the static zlib link fix. The 1.0.0 crate was published before the fix and left the librevenge `inflateInit2_`/`inflate`/`inflateEnd` symbols undefined at final link, which broke `xberg-cl

  4. v1.0.0v1.0.0Jul 28, 202637 downloads

    # xberg 1.0.0 The first stable release of **xberg** — the document-intelligence engine previously developed as **Kreuzberg**. xberg 1.0.0 is the direct successor to Kreuzberg v4.9, carrying the same Rust core and extraction-API lineage forward under a new name and an MIT license. The Kreuzberg v4 line continues as LTS at [kreuzberg-dev/kreuzberg-lts](https://github.com/kreuzberg-dev/kreuzberg-lts). This is a large release. Beyond the rename, the PDF stack moved to a pure-Rust backend, OCR grew from a single engine into a family of classical and vision-language models, and whole new capabilities landed: audio/video transcription, named-entity recognition, structured LLM extraction, sparse and late-interaction retrieval, and four new language bindings — 15 in total over one engine. ## Highlights - **Pure-Rust PDF backend.** pdfium is gone; `pdf_oxide` is now the sole PDF backend, with a layout-aware pipeline — ONNX layout detection (PP-DocLayoutV3 / RT-DETR), Docling-style reading-order reconstruction, selective per-page OCR for scanned documents, and AcroForm/XFA form extraction. - **A family of OCR backends.** Alongside Tesseract: a native **PaddleOCR** backend (PP-OCRv6) a

  5. Benchmark Results 2026-07-28 (e07fabb)benchmark-run-30366588761Jul 28, 2026pre-release22 downloads

    Comparative benchmark results from workflow run [30366588761](https://github.com/xberg-io/xberg/actions/runs/30366588761). **Commit:** e07fabba96fd3fb9073be97c751283146619248e **Date:** 2026-07-28

Commits per week

last 52 weeks
3340Week of 2025-08-03: 0 commitsWeek of 2025-08-10: 13 commitsWeek of 2025-08-17: 2 commitsWeek of 2025-08-24: 36 commitsWeek of 2025-08-31: 42 commitsWeek of 2025-09-07: 35 commitsWeek of 2025-09-14: 49 commitsWeek of 2025-09-21: 20 commitsWeek of 2025-09-28: 67 commitsWeek of 2025-10-05: 15 commitsWeek of 2025-10-12: 1 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 2 commitsWeek of 2025-11-02: 98 commitsWeek of 2025-11-09: 173 commitsWeek of 2025-11-16: 199 commitsWeek of 2025-11-23: 192 commitsWeek of 2025-11-30: 187 commitsWeek of 2025-12-07: 241 commitsWeek of 2025-12-14: 238 commitsWeek of 2025-12-21: 287 commitsWeek of 2025-12-28: 194 commitsWeek of 2026-01-04: 150 commitsWeek of 2026-01-11: 83 commitsWeek of 2026-01-18: 89 commitsWeek of 2026-01-25: 121 commitsWeek of 2026-02-01: 93 commitsWeek of 2026-02-08: 115 commitsWeek of 2026-02-15: 104 commitsWeek of 2026-02-22: 117 commitsWeek of 2026-03-01: 77 commitsWeek of 2026-03-08: 154 commitsWeek of 2026-03-15: 196 commitsWeek of 2026-03-22: 132 commitsWeek of 2026-03-29: 220 commitsWeek of 2026-04-05: 109 commitsWeek of 2026-04-12: 100 commitsWeek of 2026-04-19: 152 commitsWeek of 2026-04-26: 178 commitsWeek of 2026-05-03: 333 commitsWeek of 2026-05-10: 255 commitsWeek of 2026-05-17: 231 commitsWeek of 2026-05-24: 217 commitsWeek of 2026-05-31: 225 commitsWeek of 2026-06-07: 106 commitsWeek of 2026-06-14: 237 commitsWeek of 2026-06-21: 214 commitsWeek of 2026-06-28: 219 commitsWeek of 2026-07-05: 178 commitsWeek of 2026-07-12: 78 commitsWeek of 2026-07-19: 334 commitsWeek of 2026-07-26: 156 commitsAug 3, 2025Jul 26, 2026
7.1K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 6 commitsSun 1:00 — 9 commitsSun 2:00 — 8 commitsSun 3:00 — 3 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 16 commitsSun 7:00 — 39 commitsSun 8:00 — 82 commitsSun 9:00 — 66 commitsSun 10:00 — 57 commitsSun 11:00 — 52 commitsSun 12:00 — 42 commitsSun 13:00 — 83 commitsSun 14:00 — 85 commitsSun 15:00 — 112 commitsSun 16:00 — 84 commitsSun 17:00 — 61 commitsSun 18:00 — 67 commitsSun 19:00 — 75 commitsSun 20:00 — 65 commitsSun 21:00 — 43 commitsSun 22:00 — 32 commitsSun 23:00 — 9 commitsMon 0:00 — 8 commitsMon 1:00 — 8 commitsMon 2:00 — 2 commitsMon 3:00 — 6 commitsMon 4:00 — 0 commitsMon 5:00 — 10 commitsMon 6:00 — 19 commitsMon 7:00 — 60 commitsMon 8:00 — 73 commitsMon 9:00 — 50 commitsMon 10:00 — 69 commitsMon 11:00 — 54 commitsMon 12:00 — 32 commitsMon 13:00 — 44 commitsMon 14:00 — 49 commitsMon 15:00 — 50 commitsMon 16:00 — 91 commitsMon 17:00 — 65 commitsMon 18:00 — 47 commitsMon 19:00 — 59 commitsMon 20:00 — 66 commitsMon 21:00 — 36 commitsMon 22:00 — 44 commitsMon 23:00 — 18 commitsTue 0:00 — 17 commitsTue 1:00 — 7 commitsTue 2:00 — 7 commitsTue 3:00 — 6 commitsTue 4:00 — 3 commitsTue 5:00 — 21 commitsTue 6:00 — 14 commitsTue 7:00 — 48 commitsTue 8:00 — 74 commitsTue 9:00 — 79 commitsTue 10:00 — 62 commitsTue 11:00 — 50 commitsTue 12:00 — 33 commitsTue 13:00 — 54 commitsTue 14:00 — 59 commitsTue 15:00 — 62 commitsTue 16:00 — 63 commitsTue 17:00 — 65 commitsTue 18:00 — 36 commitsTue 19:00 — 48 commitsTue 20:00 — 64 commitsTue 21:00 — 38 commitsTue 22:00 — 22 commitsTue 23:00 — 13 commitsWed 0:00 — 27 commitsWed 1:00 — 12 commitsWed 2:00 — 7 commitsWed 3:00 — 5 commitsWed 4:00 — 4 commitsWed 5:00 — 12 commitsWed 6:00 — 46 commitsWed 7:00 — 79 commitsWed 8:00 — 81 commitsWed 9:00 — 76 commitsWed 10:00 — 72 commitsWed 11:00 — 40 commitsWed 12:00 — 29 commitsWed 13:00 — 28 commitsWed 14:00 — 53 commitsWed 15:00 — 61 commitsWed 16:00 — 63 commitsWed 17:00 — 53 commitsWed 18:00 — 54 commitsWed 19:00 — 43 commitsWed 20:00 — 81 commitsWed 21:00 — 63 commitsWed 22:00 — 33 commitsWed 23:00 — 17 commitsThu 0:00 — 6 commitsThu 1:00 — 10 commitsThu 2:00 — 4 commitsThu 3:00 — 9 commitsThu 4:00 — 3 commitsThu 5:00 — 11 commitsThu 6:00 — 19 commitsThu 7:00 — 41 commitsThu 8:00 — 47 commitsThu 9:00 — 67 commitsThu 10:00 — 63 commitsThu 11:00 — 76 commitsThu 12:00 — 66 commitsThu 13:00 — 62 commitsThu 14:00 — 86 commitsThu 15:00 — 98 commitsThu 16:00 — 124 commitsThu 17:00 — 76 commitsThu 18:00 — 53 commitsThu 19:00 — 63 commitsThu 20:00 — 74 commitsThu 21:00 — 58 commitsThu 22:00 — 53 commitsThu 23:00 — 35 commitsFri 0:00 — 21 commitsFri 1:00 — 17 commitsFri 2:00 — 4 commitsFri 3:00 — 5 commitsFri 4:00 — 6 commitsFri 5:00 — 6 commitsFri 6:00 — 25 commitsFri 7:00 — 70 commitsFri 8:00 — 55 commitsFri 9:00 — 61 commitsFri 10:00 — 44 commitsFri 11:00 — 66 commitsFri 12:00 — 56 commitsFri 13:00 — 44 commitsFri 14:00 — 72 commitsFri 15:00 — 36 commitsFri 16:00 — 61 commitsFri 17:00 — 109 commitsFri 18:00 — 81 commitsFri 19:00 — 38 commitsFri 20:00 — 75 commitsFri 21:00 — 67 commitsFri 22:00 — 23 commitsFri 23:00 — 9 commitsSat 0:00 — 1 commitsSat 1:00 — 4 commitsSat 2:00 — 3 commitsSat 3:00 — 3 commitsSat 4:00 — 4 commitsSat 5:00 — 6 commitsSat 6:00 — 6 commitsSat 7:00 — 44 commitsSat 8:00 — 106 commitsSat 9:00 — 78 commitsSat 10:00 — 65 commitsSat 11:00 — 59 commitsSat 12:00 — 90 commitsSat 13:00 — 78 commitsSat 14:00 — 80 commitsSat 15:00 — 68 commitsSat 16:00 — 85 commitsSat 17:00 — 81 commitsSat 18:00 — 79 commitsSat 19:00 — 84 commitsSat 20:00 — 81 commitsSat 21:00 — 42 commitsSat 22:00 — 21 commitsSat 23:00 — 12 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    277.2K stars · Python

  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    195K stars · Rust

  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    194.9K stars · Rust

  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    194.9K stars · Rust

  • Significant-Gravitas/AutoGPT

    AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

    186.3K stars · Python