run-llama/liteparsePublic

A fast, helpful, and open-source document parser

AI summary: A lightweight parser for extracting and structuring text from complex document formats.

Stars
12.8K
+36 today
Forks
870
Watchers
40
Open issues
24
Open PRs
13
Contributors
~36
Commits
1.1K
Branches
31

RustApache-2.0Created Feb 9, 2026Last push 2d agoLatest release node-v2.14.7+175 stars this week+532 this month

Quick answers

What is liteparse?
A lightweight parser for extracting and structuring text from complex document formats.
What does liteparse do?
LiteParse is a utility library focused on efficiently extracting clean, structured text from complex document formats like PDFs, DOCX, and HTML. It addresses the common challenge in LLM (Large Language Model) application development where raw document ingestion produces noisy, unstructured data. The library utilizes heuristic-based parsing to identify and preserve the structural hierarchy of documents, such as headers, paragraphs, lists, and tables. By converting messy source files into clean, chunkable text formats, it significantly improves the performance and accuracy of Retrieval-Augmented Generation (RAG) pipelines.
Who is liteparse for?
AI engineers, data scientists, and developers building LLM applications, particularly those focused on Retrieval-Augmented Generation (RAG) and document processing.
How do I get started with liteparse?
pip install liteparse
How popular is liteparse on GitHub?
run-llama/liteparse has 12,769 stars and 870 forks on GitHub, and gained 175 stars in the last 7 days.
What license does liteparse use?
run-llama/liteparse is released under the Apache-2.0 license.

Star history

since Jul 29, 2026
05K10KJul 2026Aug 2026Sep 2026Oct 2026
12.8K stars as of Oct 3, 2026. Measured daily since Jul 29, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 2 commits2026-02-10: 2 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 6 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 7 commits2026-02-17: 14 commits2026-02-18: 9 commits2026-02-19: 22 commits2026-02-20: 3 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 2 commits2026-02-24: 2 commits2026-02-25: 0 commits2026-02-26: 2 commits2026-02-27: 0 commits2026-02-28: 3 commits2026-03-01: 0 commits2026-03-02: 1 commit2026-03-03: 0 commits2026-03-04: 3 commits2026-03-05: 4 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 3 commits2026-03-09: 1 commit2026-03-10: 1 commit2026-03-11: 5 commits2026-03-12: 2 commits2026-03-13: 2 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 5 commits2026-03-17: 1 commit2026-03-18: 10 commits2026-03-19: 7 commits2026-03-20: 9 commits2026-03-21: 2 commits2026-03-22: 1 commit2026-03-23: 10 commits2026-03-24: 5 commits2026-03-25: 15 commits2026-03-26: 9 commits2026-03-27: 15 commits2026-03-28: 15 commits2026-03-29: 7 commits2026-03-30: 1 commit2026-03-31: 2 commits2026-04-01: 2 commits2026-04-02: 9 commits2026-04-03: 1 commit2026-04-04: 1 commit2026-04-05: 1 commit2026-04-06: 4 commits2026-04-07: 5 commits2026-04-08: 0 commits2026-04-09: 1 commit2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 1 commit2026-04-13: 11 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 1 commit2026-04-17: 3 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 8 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 4 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 2 commits2026-04-29: 1 commit2026-04-30: 0 commits2026-05-01: 2 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 1 commit2026-05-05: 0 commits2026-05-06: 2 commits2026-05-07: 0 commits2026-05-08: 5 commits2026-05-09: 2 commits2026-05-10: 6 commits2026-05-11: 9 commits2026-05-12: 6 commits2026-05-13: 5 commits2026-05-14: 18 commits2026-05-15: 9 commits2026-05-16: 2 commits2026-05-17: 3 commits2026-05-18: 4 commits2026-05-19: 15 commits2026-05-20: 23 commits2026-05-21: 12 commits2026-05-22: 4 commits2026-05-23: 4 commits2026-05-24: 8 commits2026-05-25: 17 commits2026-05-26: 7 commits2026-05-27: 10 commits2026-05-28: 4 commits2026-05-29: 23 commits2026-05-30: 0 commits2026-05-31: 5 commits2026-06-01: 4 commits2026-06-02: 19 commits2026-06-03: 0 commits2026-06-04: 7 commits2026-06-05: 5 commits2026-06-06: 0 commits2026-06-07: 1 commit2026-06-08: 10 commits2026-06-09: 6 commits2026-06-10: 2 commits2026-06-11: 12 commits2026-06-12: 4 commits2026-06-13: 1 commit2026-06-14: 3 commits2026-06-15: 3 commits2026-06-16: 0 commits2026-06-17: 14 commits2026-06-18: 3 commits2026-06-19: 2 commits2026-06-20: 1 commit2026-06-21: 0 commits2026-06-22: 1 commit2026-06-23: 3 commits2026-06-24: 7 commits2026-06-25: 5 commits2026-06-26: 2 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 9 commits2026-06-30: 3 commits2026-07-01: 2 commits2026-07-02: 1 commit2026-07-03: 1 commit2026-07-04: 1 commit2026-07-05: 0 commits2026-07-06: 2 commits2026-07-07: 3 commits2026-07-08: 3 commits2026-07-09: 0 commits2026-07-10: 3 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 9 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 1 commit2026-07-19: 8 commits2026-07-20: 4 commits2026-07-21: 12 commits2026-07-22: 2 commits2026-07-23: 7 commits2026-07-24: 1 commit2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 3 commits2026-07-28: 5 commits2026-07-29: 3 commits2026-07-30: 0 commits2026-07-31: 3 commits2026-08-01: 0 commits2026-08-02: 16 commits2026-08-03: 2 commits2026-08-04: 2 commits2026-08-05: 4 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 5 commits2026-08-11: 10 commits2026-08-12: 3 commits2026-08-13: 3 commits2026-08-14: 3 commits2026-08-15: 5 commits2026-08-16: 4 commits2026-08-17: 3 commits2026-08-18: 2 commits2026-08-19: 14 commits2026-08-20: 9 commits2026-08-21: 2 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 10 commits2026-08-25: 14 commits2026-08-26: 2 commits2026-08-27: 10 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 4 commits2026-09-01: 6 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 2 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 2 commits2026-09-08: 2 commits2026-09-09: 11 commits2026-09-10: 3 commits2026-09-11: 3 commits2026-09-12: 1 commit2026-09-13: 0 commits2026-09-14: 6 commits2026-09-15: 4 commits2026-09-16: 0 commits2026-09-17: 2 commits2026-09-18: 1 commit2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 4 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits
846 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    12,769 stars

  • Very active

    846 commits in 52 weeks

  • Permissive license

    Apache-2.0

  • Repeat trending

    4 trending appearances

What liteparse does

LiteParse is a utility library focused on efficiently extracting clean, structured text from complex document formats like PDFs, DOCX, and HTML. It addresses the common challenge in LLM (Large Language Model) application development where raw document ingestion produces noisy, unstructured data. The library utilizes heuristic-based parsing to identify and preserve the structural hierarchy of documents, such as headers, paragraphs, lists, and tables. By converting messy source files into clean, chunkable text formats, it significantly improves the performance and accuracy of Retrieval-Augmented Generation (RAG) pipelines.

AI engineers, data scientists, and developers building LLM applications, particularly those focused on Retrieval-Augmented Generation (RAG) and document processing.

  • Multi-Format Extraction: Extracts text seamlessly from PDFs, Word documents (DOCX), and HTML files.
  • Structural Preservation: Intelligently identifies and preserves document hierarchy (headers, paragraphs, lists) during extraction.
  • Lightweight Architecture: Designed for fast execution with minimal dependencies, suitable for high-throughput ingestion pipelines.
  • Table Processing: Features specialized algorithms for extracting tabular data into structured, machine-readable formats.
  • LLM Readiness: Outputs clean, sanitized text optimized for chunking and embedding in RAG applications.

Where teams use it

RAG Pipeline Ingestion

Parsing thousands of corporate PDFs into clean text chunks for a company-wide semantic search engine.

Data Extraction

Automating the extraction of structured tabular data from weekly financial reports.

Document Sanitization

Cleaning up messy HTML scraped from the web into standardized markdown before feeding it to an LLM.

Knowledge Base Generation

Converting legacy DOCX manuals into structured text to train specialized AI models.

Getting started: pip install liteparse

README

main branch

LiteParse

CI | Crates.io version | npm version | wasm version | PyPI version | License | Docs

English | 简体中文

out

Looking for LiteParse V1? Follow this link to the old code

LiteParse is a standalone OSS PDF parsing tool focused exclusively on fast and light parsing. It provides high-quality spatial text parsing with bounding boxes, without proprietary LLM features or cloud dependencies. Everything runs locally on your machine.

Hitting the limits of local parsing? For complex documents (dense tables, multi-column layouts, charts, handwritten text, or scanned PDFs), you'll get significantly better results with LlamaParse, our cloud-based document parser built for production document pipelines. LlamaParse handles the hard stuff so your models see clean, structured data and markdown.

Sign up for LlamaParse free

Overview

  • Fast Text Parsing: Spatial text parsing using PDFium, ~2-5ms per page
  • Flexible OCR System:
    • Built-in: Tesseract (zero setup, bundled with the library)
    • HTTP Servers: Plug in any OCR server (EasyOCR, PaddleOCR, custom)
    • Standard API: Simple, well-defined OCR API specification
  • Complexity Detection: Cheaply check whether a document needs OCR or heavier parsing — route, reject, or estimate cost before a full parse
  • Screenshot Generation: Generate high-quality page screenshots for LLM agents
  • Multiple Output Formats: Markdown, JSON, and Text
  • Markdown Output: Structured Markdown with headings, tables, lists, images, and links — great for feeding LLMs and RAG pipelines
  • Bounding Boxes: Precise text positioning information
  • Multi-language: Use from Rust, Node.js/TypeScript, Python, or the browser (WASM)
  • Worker Pool Mode (Python & Node.js): Parse in persistent worker processes for true parallelism (PDFium otherwise serializes concurrent parses) and hard per-parse timeouts — rogue documents are killed, identified by name, and never stall the pipeline
  • Multi-platform: Linux, macOS (Intel/ARM), Windows
flowchart LR
      subgraph Input["Input Formats"]
          direction TB
          PDF["PDF"]
          DOCX["DOCX"]
          XLSX["XLSX"]
          PPTX["PPTX"]
          IMG["Images"]
      end

      subgraph Core["Rust Core"]
          direction TB
          CONV["Format Conversion\nLibreOffice / Rust image + resvg + usvg crates"]
          EXTRACT["Text Extraction\nPDFium C library"]
          OCR["Selective OCR\nTesseract / HTTP / Custom"]
          MERGE["OCR Merge\nNative text + OCR results"]
          PROJ["Grid Projection\nSpatial layout reconstruction"]
          CONV --> EXTRACT
          EXTRACT --> OCR --> MERGE --> PROJ
          EXTRACT --> MERGE
      end

      subgraph Output[" Output "]
          direction TB
          JSON["Structured JSON\ntext + bounding boxes"]
          TEXT["Plain Text\nlayout-preserved"]
          SCREEN["Screenshots\nPNG rendering"]
      end

      subgraph Bindings["Language Bindings"]
          direction TB
          NAPI["Node.js / TypeScript\nnapi-rs"]
          PYO3["Python\nPyO3"]
          WASM["Browser / WASM\nwasm-bindgen"]
          CLI["CLI\ncargo / npm / pip"]
          NAPI ~~~ PYO3 ~~~ WASM ~~~ CLI
      end

      PDF --> EXTRACT
      DOCX & XLSX & PPTX & IMG --> CONV
      PROJ --> JSON & TEXT & SCREEN
      JSON & TEXT & SCREEN --> Bindings

      style Input fill:#F5F5F5,color:#000000,stroke:#37D7FA,stroke-width:2px
      style Core fill:#F5F5F5,color:#000000,stroke:#3E18F9,stroke-width:2px
      style Output fill:#F5F5F5,color:#000000,stroke:#FF8705,stroke-width:2px
      style Bindings fill:#F5F5F5,color:#000000,stroke:#FF8DF2,stroke-width:2px

      style PDF fill:#96E7F9,color:#000000,stroke:#37D7FA,stroke-width:1px
      style DOCX fill:#96E7F9,color:#000000,stroke:#37D7FA,stroke-width:1px
      style XLSX fill:#96E7F9,color:#000000,stroke:#37D7FA,stroke-width:1px
      style PPTX fill:#96E7F9,color:#000000,stroke:#37D7FA,stroke-width:1px
      style IMG fill:#96E7F9,color:#000000,stroke:#37D7FA,stroke-width:1px

      style CONV fill:#92AEFF,color:#000000,stroke:#4B72FE,stroke-width:1px
      style EXTRACT fill:#92AEFF,color:#000000,stroke:#4B72FE,stroke-width:1px
      style OCR fill:#92AEFF,color:#000000,stroke:#4B72FE,stroke-width:1px
      style MERGE fill:#92AEFF,color:#000000,stroke:#4B72FE,stroke-width:1px
      style PROJ fill:#4B72FE,color:#FFFFFF,stroke:#3E18F9,stroke-width:2px

      style JSON fill:#FFBD74,color:#000000,stroke:#FF8705,stroke-width:1px
      style TEXT fill:#FFBD74,color:#000000,stroke:#FF8705,stroke-width:1px
      style SCREEN fill:#FFBD74,color:#000000,stroke:#FF8705,stroke-width:1px

      style NAPI fill:#FFBFF8,color:#000000,stroke:#FF8DF2,stroke-width:1px
      style PYO3 fill:#FFBFF8,color:#000000,stroke:#FF8DF2,stroke-width:1px
      style WASM fill:#FFBFF8,color:#000000,stroke:#FF8DF2,stroke-width:1px
      style CLI fill:#FFBFF8,color:#000000,stroke:#FF8DF2,stroke-width:1px
Loading

Benchmarks

LiteParse is measured on multiple public Doc→Markdown benchmarks. All numbers below were produced on this machine from one command (see Reproducing), with every tool at its latest release as of 2026-09-09. LiteParse and the "model-free" competitors use no ML model at all (no LLM, no layout model, no GPU). All tested methods have permissive licenses and can be run locally with minimal dependencies.

Benchmark Metric LiteParse + Tesseract OCR + PaddleOCR Best other model-free tool
ParseBench (2,049 docs) Overall (mean of 5 categories) 0.364 0.380 0.389 pdf-inspector 0.283
opendataloader-bench (200 docs) Overall (NID + TEDS + MHS) 0.886 0.896 0.901 opendataloader 0.842
olmOCR-bench (1,403 pages) % tests passed 39.6 41.1 42.2 pdf-inspector 33.7
ParseBench

Rule-based scoring, no LLM judge. Each column is ParseBench's canonical per-category metric (Tables = GriTS/TRM composite; the others are rule pass-rates). Overall is the mean of the five, as on the ParseBench leaderboard.

Pipeline Overall Tables Charts Content Faithfulness Semantic Formatting Visual Grounding
LiteParse + PaddleOCR 0.389 0.430 0.013 0.787 0.402 0.314
LiteParse + Tesseract 0.380 0.428 0.012 0.751 0.399 0.307
LiteParse (no OCR) 0.364 0.424 0.013 0.700 0.385 0.297
pdf-inspector 1.19 0.283 0.277 0.017 0.598 0.426 0.099
opendataloader 2.5.7 0.277 0.349 0.006 0.663 0.258 0.110
markitdown 0.1.7 0.185 0.158 0.020 0.652 0.001 0.110

Notes:

  • Visual Grounding scores layout blocks with bounding boxes (lit parse --extract-blocks).
  • Charts is near zero for every tool here — none reconstruct chart data.
opendataloader-bench

NID = reading-order similarity, TEDS = table structure, MHS = heading hierarchy. Overall is the harness's own mean.

Engine Overall NID TEDS MHS
LiteParse + PaddleOCR 0.901 0.932 0.832 0.840
LiteParse + Tesseract 0.896 0.928 0.829 0.828
LiteParse (no OCR) 0.886 0.917 0.818 0.821
nutrient (commercial) 0.885 0.925 0.708 0.819
opendataloader 2.5.7 0.842 0.912 0.483 0.757
markitdown 0.1.7 0.589 0.844 0.273 0.000

Notes:

  • nutrient has no runnable parser in this harness (commercial)
  • The corpus is native-text PDFs, so OCR gains come from embedded figures and a handful of scanned pages.
olmOCR-bench

Score = average of per-category pass rates. The two math categories require LaTeX output and are 0% for every tool here; they still count in the average.

Engine Overall baseline headers_footers multi_column table_tests long_tiny_text old_scans arxiv_math old_scans_math
LiteParse + PaddleOCR 42.2 99.9 48.7 69.1 54.0 46.4 19.4 0.0 0.0
LiteParse + Tesseract 41.1 99.9 52.1 66.2 54.1 42.5 13.9 0.0 0.0
LiteParse (no OCR) 39.6 99.9 55.8 66.3 52.5 29.2 13.3 0.0 0.0
pdf-inspector 1.19 33.7 82.9 62.1 49.7 43.6 17.6 13.3 0.0 0.0
opendataloader 2.5.7 32.5 86.9 36.6 63.7 24.9 34.8 13.3 0.0 0.0
markitdown 0.1.7 28.7 86.8 38.8 39.3 19.9 31.2 13.3 0.0 0.0

Notes:

  • headers_footers expects letterhead and footer text to be absent. OCR recovers that text from logo and address images on single-page documents, where the repeated-header filter cannot fire, so the OCR rows score lower there by design. We chose not to drop text by page position. pdf-inspector 1.19 refuses scanned and image-based pages outright (192 of the 1,403), which is why it tops this category while scoring nothing on those pages elsewhere.
  • old_scans is largely cursive handwriting; Tesseract cannot read it, PaddleOCR partially can.
How the runs are configured
  • Same scorer for every tool. Each benchmark's own evaluator, unmodified. The ground truth in all three corpora is plain text, so every tool including LiteParse runs with hyperlink syntax off (--no-links); everything else is the default lit parse --format markdown.
  • OCR modes. No OCR is --no-ocr. Tesseract is the built-in engine with no setup. PaddleOCR is the PP-OCRv5 mobile detector and English recognizer served over the OCR HTTP API by ocr/rapidocr, which runs the PaddleOCR models on ONNX Runtime — the same models as ocr/paddleocr, just 20–40× faster per page on CPU-only machines.
  • Competitors are the free converters we could run locally, each at its latest release on 2026-09-09: markitdown 0.1.7, opendataloader-pdf 2.5.7 (its non-hybrid mode), pdf-inspector 1.19.0. Each benchmark's own runner for the tool is used as-is. Numbers from public leaderboards or LLM-assisted modes are not mixed in.
  • Machine: Apple M2 Max, 12 cores, 32 GB. LiteParse rows: v2.14.4 (2026-09-09).

Reproducing the benchmarks

./run_benchmarks.sh --liteparse-only                # all three benches, no OCR
./run_benchmarks.sh --liteparse-only --ocr=tesseract
( cd ocr/rapidocr && uv run server.py ) &           # then:
./run_benchmarks.sh --liteparse-only --ocr=paddle
./run_benchmarks.sh --competitors-only              # re-run the free competitors only
./run_benchmarks.sh                                 # everything

Results and a SUMMARY.md land in bench_results/latest/.

Installation

Install via your preferred package manager. All versions (except WASM) ship with the same lit CLI.

Language Install Library Docs
Node.js / TypeScript npm i -g @llamaindex/liteparse Node.js README
Python pip install liteparse Python README
Rust cargo install liteparse (CLI) / cargo add liteparse (lib) Rust README (crates.io)
Browser (WASM) npm i @llamaindex/liteparse-wasm WASM README

Agent Skill

You can use liteparse as an agent skill, downloading it with the skills CLI tool:

npx skills add run-llama/llamaparse-agent-skills --skill liteparse

Or copy-pasting the SKILL.md file to your own skills setup.

See the Agent Skill guide for requirements and usage patterns.

CLI Usage

The CLI is the same across all installations (npm, pip, cargo install).

Parse Files

# Basic parsing
lit parse document.pdf

# Parse to Markdown — headings, tables, lists, images, links
lit parse document.pdf --format markdown -o output.md

# Parse with specific format
lit parse document.pdf --format json -o output.json

# Parse specific pages
lit parse document.pdf --target-pages "1-5,10,15-20"

# Parse without OCR
lit parse document.pdf --no-ocr

# Include page-scoped vector path data in JSON
lit parse document.pdf --format json --extract-vector-graphics

# Include rich per-item PDF text metadata
lit parse document.pdf --format json --extract-text-metadata

# Include page annotations in structured JSON
lit parse document.pdf --format json --extract-annotations

# Include AcroForm widget fields and values (repairs orphaned widgets in memory)
lit parse document.pdf --format json --extract-form-fields

# Parse a remote PDF
curl -sL https://example.com/report.pdf | lit parse -

Markdown Output

LiteParse can render documents directly to Markdown. This means reconstructing headings, tables, lists, images, and links from the spatial layout. This is ideal for feeding documents to LLMs and RAG pipelines. This mode is purely heuristics and rule-based, so complex documents may not render perfectly, but it will be fast.

# Render to Markdown
lit parse document.pdf --format markdown -o output.md

# Strip images instead of emitting placeholders
lit parse document.pdf --format markdown --image-mode off

# Extract embedded images to disk and reference them from the markdown
lit parse document.pdf --format markdown --image-mode embed --extract-images --image-output-dir ./images

# Extract image bytes and metadata without changing Markdown image handling
lit parse document.pdf --format json --extract-images

# Emit link text as plain text (no [text](url) syntax)
lit parse document.pdf --format markdown --no-links

# Include tagged-PDF logical structure in JSON
lit parse document.pdf --format json --extract-structure-tree

# Include the classified layout blocks (with bounding boxes) in JSON
lit parse document.pdf --format json --extract-blocks

Image handling is controlled by --image-mode:

Mode Behavior
placeholder (default) Emits ![](img_pN_K.png) references in reading order
off Strips images entirely
embed Emits the same image references as placeholder

--extract-images is the only option that enables embedded-image extraction. --image-output-dir requires it and writes the extracted bytes to disk. JSON output contains each image's name, path, page bbox, intrinsic pixel dimensions, rotation, format, and duplicate relationship; pixel bytes are never embedded in JSON. Identical image resources reuse the same output file.

Library callers can opt in with extract_images: true (Rust), extractImages: true (Node/WASM), or extract_images=True (Python). It defaults to false. Markdown image mode controls presentation only; placeholder refs are still discovered without bytes.

Markdown reconstruction quality varies with document complexity. For the hardest documents (dense tables, multi-column layouts, scans), LlamaParse remains the most accurate option.

Vector Graphics

Vector path output is opt-in because path-heavy PDFs can produce large payloads. Enable it with --extract-vector-graphics, Rust/Python extract_vector_graphics = true, or JavaScript/WASM extractVectorGraphics: true. Each page then includes vector_graphics (vectorGraphics in JavaScript) with:

  • shapes: path bounding box, stroke/fill paint state and ARGB colors, and whether the path contains a Bezier curve.
  • lines: compatible horizontal/vertical segments merged using stroke width and paint colors, with top-left 72-DPI viewport coordinates.

The representation follows LlamaParse PDFium path extraction; LiteParse calls the shape rectangle bbox rather than PDFium's coords, and uses width / height rather than w / h. The field is absent (or None/undefined) by default. Diagonal and curved segments are represented by their parent shape but are not emitted as lines.

Tagged PDF structure tree

Enable --extract-structure-tree (Rust/Python extract_structure_tree, JavaScript/WASM extractStructureTree) to add a page-scoped structure_tree. It preserves every root and recursively exposes element type, ID, actual/alternate text, title, typed scalar attributes, marked-content IDs, children, and referenced link annotations. The field is absent by default; enabled untagged pages contain roots: [].

Layout blocks

The Markdown renderer works by classifying each page into blocks — headings, paragraphs, list items, tables, code, rules, figures — and then rendering them. Enable --extract-blocks (Rust/Python extract_blocks, JavaScript/WASM extractBlocks) to get that decomposition as data instead of only as rendered text, with the coordinates the classifier used.

Each page gains a blocks array in reading order — the same order, and the same blocks, the Markdown output is built from. Every block carries:

  • kind: one of heading, paragraph, list_item, code, table, grid_fallback, rule, figure.
  • bbox: the region it occupies, in the same top-left 72-DPI viewport space as text_items. This is the union of every source line that fed the block, so a wrapped heading or a multi-line paragraph reports its whole band.
  • Kind-specific fields, omitted when they don't apply: text and level for headings; ordered / marker for list items; lines and lang for code; header and rows for tables; id / format for figures.

Table cells are objects, not bare strings — each has text and its own bbox, so a cell can be mapped back to the region of the page it was read from. For ruled tables that box is the drawn grid cell; for borderless tables it is the extent of the spans the cell was built from. Cells that exist only to square off a ragged grid carry no bbox, since they have no ink behind them.

{
  "kind": "table",
  "bbox": { "x": 72.0, "y": 310.5, "width": 468.0, "height": 96.0 },
  "header": [
    { "text": "Territory Code", "bbox": { "x": 72.0, "y": 310.5, "width": 120.0, "height": 24.0 } },
    { "text": "Factor",         "bbox": { "x": 192.0, "y": 310.5, "width": 96.0, "height": 24.0 } }
  ],
  "rows": [
    [
      { "text": "001", "bbox": { "x": 72.0, "y": 334.5, "width": 120.0, "height": 24.0 } },
      { "text": "1.25", "bbox": { "x": 192.0, "y": 334.5, "width": 96.0, "height": 24.0 } }
    ]
  ]
}

Document metadata, content bounds, and XFA packets

Parse results (Rust/Node/Python APIs) carry the document's /Info creator and producer entries when present; these are API-only and never appear in CLI JSON output. Enable extract_document_metadata (JavaScript/WASM extractDocumentMetadata) to add doc_meta/docMeta, a provenance object with the /Info creation/modification dates, PDF version and encryption permissions, signature state, incremental-save markers, trailer ID comparison, the document catalog's XMP packet (capped at 64 KiB, with xmp_truncated when it was cut), and source file size. It is off by default because it streams the whole source file once; it is absent for inputs converted from a non-PDF format, where the facts would describe the intermediate PDF rather than your file. xmp needs a structural parse of the document, so it is skipped (left absent) for sources over 16 MiB and in WASM builds — the other fields are unaffected. Enable --extract-content-bounds (Rust/Python extract_content_bounds, JavaScript/WASM extractContentBounds) to add a per-page content_bounds: the union bbox of the page's top-level content objects in viewport coords (absent for empty pages). Enable --extract-xfa-packets (Rust/Python extract_xfa_packets, JavaScript/WASM extractXfaPackets) to add xfa_packets with each raw XFA packet's index, name, byte length, and XML content; non-XFA documents yield an empty list. All of these are off by default, so default JSON output is unchanged.

Screenshot raster signals

Screenshots draw AcroForm field appearances (filled values, checkbox states) on top of the page raster, so form data is visible in the render and to OCR. Each screenshot result reports is_solid_fill (blank page after render), and with detect_screenshot_rects (Node detectScreenshotRects) also rects: solid same-color rectangles and lines found in the raster in viewport coords, which covers scanned/flattened pages that carry no vector paths.

Check Complexity

Before committing to a full parse, check whether a document actually needs OCR or heavier processing. This is a cheap, text-layer-only pass — useful for routing documents to different pipelines, rejecting ones you can't handle, or estimating cost.

# Print the complexity verdict and per-page JSON
lit is-complex document.pdf

# Use as a shell predicate — only parse with --no-ocr when the document is simple
lit is-complex document.pdf --quiet && lit parse document.pdf --no-ocr

# List the pages that need OCR
lit is-complex document.pdf --compact | jq '[.[] | select(.needs_ocr) | .page_number]'

It always prints per-page JSON to stdout, a human-readable verdict to stderr, and exits non-zero when any page needs OCR. Each page carries a needs_ocr verdict and a list of reasons (scanned, no-text, sparse-text, embedded-images, garbled, vector-text, annotation-text).

Batch Parsing

Parse an entire directory of documents:

lit batch-parse ./input-directory ./output-directory

Generate Screenshots

Screenshots are essential for LLM agents to extract visual information that text alone cannot capture.

# Screenshot all pages
lit screenshot document.pdf -o ./screenshots

# Screenshot specific pages
lit screenshot document.pdf --target-pages "1,3,5" -o ./screenshots

# Custom DPI
lit screenshot document.pdf --dpi 300 -o ./screenshots

CLI Reference

Parse Command
lit parse [OPTIONS] <file>

Options:
  -o, --output <file>          Output file path
      --format <format>        Output format: json|text|markdown [default: text]
      --no-ocr                 Disable OCR
      --ocr-language <lang>    OCR language, Tesseract format [default: eng]
      --ocr-server-url <url>   HTTP OCR server URL (uses Tesseract if not provided)
      --tessdata-path <path>   Path to tessdata directory
      --max-pages <n>          Max pages to parse [default: 1000]
      --target-pages <pages>   Pages to parse (e.g., "1-5,10,15-20")
      --dpi <dpi>              Rendering DPI [default: 150]
      --image-mode <mode>      Markdown image handling: off|placeholder|embed [default: placeholder]
      --extract-images         Extract embedded image bytes and metadata
      --image-output-dir <dir> Write extracted images; requires --extract-images
      --extract-text-metadata  Include rich PDF text metadata in text items
      --extract-vector-graphics Include page vector shapes and merged H/V lines
      --no-links               Emit link anchor text as plain text (no [text](url)) in markdown
      --keep-headers-footers   Keep running headers/footers in markdown (skip repeated-line stripping)
      --extract-annotations    Include PDF annotations in page output
      --extract-form-fields    Include AcroForm widget fields and values
      --preserve-small-text    Keep very small text
      --password <password>    Password for encrypted documents
      --num-workers <n>        Concurrent OCR workers [default: CPU cores - 1]
  -q, --quiet                  Suppress progress output
  -h, --help                   Print help
Batch Parse Command
lit batch-parse [OPTIONS] <input-dir> <output-dir>

Options:
      --format <format>        Output format: json|text|markdown [default: text]
      --no-ocr                 Disable OCR
      --ocr-language <lang>    OCR language [default: eng]
      --ocr-server-url <url>   HTTP OCR server URL
      --tessdata-path <path>   Path to tessdata directory
      --max-pages <n>          Max pages per file [default: 1000]
      --dpi <dpi>              Rendering DPI [default: 150]
      --recursive              Recursively search input directory
      --extension <ext>        Only process files with this extension (e.g., ".pdf")
      --password <password>    Password for encrypted documents
      --num-workers <n>        Concurrent OCR workers
  -q, --quiet                  Suppress progress output
  -h, --help                   Print help
Screenshot Command
lit screenshot [OPTIONS] <file>

Options:
  -o, --output-dir <dir>       Output directory [default: ./screenshots]
      --target-pages <pages>   Pages to screenshot (e.g., "1,3,5" or "1-5")
      --dpi <dpi>              Rendering DPI [default: 150]
      --password <password>    Password for encrypted documents
  -q, --quiet                  Suppress progress output
  -h, --help                   Print help
Is-Complex Command
lit is-complex [OPTIONS] <file>

Options:
      --compact                Emit dense, whitespace-free JSON instead of pretty-printed
      --max-pages <n>          Max pages to check [default: 1000]
      --target-pages <pages>   Pages to check (e.g., "1-5,10,15-20")
      --password <password>    Password for encrypted documents
  -q, --quiet                  Suppress the stderr verdict
  -h, --help                   Print help

Prints per-page JSON to stdout and a COMPLEX/SIMPLE verdict to stderr; exits non-zero when any page needs OCR, so it composes as a shell predicate.

OCR Setup

Default: Tesseract

Tesseract is bundled and works out of the box:

lit parse document.pdf                    # OCR enabled by default
lit parse document.pdf --ocr-language fra # Specify language
lit parse document.pdf --no-ocr           # Disable OCR

For offline or air-gapped environments, set TESSDATA_PREFIX to a directory containing pre-downloaded .traineddata files:

export TESSDATA_PREFIX=/path/to/tessdata
lit parse document.pdf --ocr-language eng

Or pass the path directly:

lit parse document.pdf --tessdata-path /path/to/tessdata

Optional: HTTP OCR Servers

For higher accuracy or better performance, you can use an HTTP OCR server. We provide ready-to-use example wrappers for popular OCR engines:

You can integrate any OCR service by implementing the simple LiteParse OCR API specification (see OCR_API_SPEC.md).

The API requires:

  • POST /ocr endpoint
  • Accepts file and language parameters
  • Returns JSON: { results: [{ text, bbox: [x1,y1,x2,y2], confidence }] }

Multi-Format Input Support

LiteParse supports automatic conversion of various document formats to PDF before parsing.

Supported Input Formats

Office Documents (via LibreOffice)
  • Word: .doc, .docx, .docm, .odt, .rtf, .pages
  • PowerPoint: .ppt, .pptx, .pptm, .odp, .key
  • Spreadsheets: .xls, .xlsx, .xlsm, .ods, .csv, .tsv, .numbers

Install LibreOffice for automatic conversion:

# macOS
brew install --cask libreoffice

# Ubuntu/Debian
apt-get install libreoffice

# Windows
choco install libreoffice-fresh

On Windows, you may need to add LibreOffice's program directory (usually C:\Program Files\LibreOffice\program) to your PATH.

Images (native support)
  • Formats: .jpg, .jpeg, .png, .gif, .bmp, .tiff, .webp, .svg

![NOTE]

As of v2.8.0, imagemagick is no longer required to convert images to PDF. Conversion is natively handled by the rust code.

Environment Variables

Variable Description
TESSDATA_PREFIX Path to a directory containing Tesseract .traineddata files. Used for offline/air-gapped environments.

Development

The project is a Rust workspace with the core library and language-specific binding crates.

crates/
├── liteparse/          # Core library + CLI binary
├── liteparse-napi/     # Node.js bindings (napi-rs)
├── liteparse-python/   # Python bindings (PyO3)
├── liteparse-wasm/     # WASM bindings (wasm-bindgen)
├── pdfium/             # PDFium Rust wrapper
└── pdfium-sys/         # PDFium FFI bindings
packages/
├── node/               # npm package (TS wrapper + native binary)
├── python/             # PyPI package (Python wrapper + native binary)
└── wasm/               # WASM npm package

Building

# Build the CLI
cargo build --release -p liteparse

# Build Node.js bindings
cd packages/node && npm run build

# Build Python bindings
cd packages/python && maturin develop --release

# Build WASM
cd packages/wasm && npm run build

We provide a fairly rich AGENTS.md/CLAUDE.md that we recommend using to help with development + coding agents.

License

Apache 2.0

Credits

Built on top of:

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

195 total
  1. WASM v2.14.7wasm-v2.14.7Sep 22, 2026

    **Full Changelog**: https://github.com/run-llama/liteparse/compare/crates-v2.14.7...wasm-v2.14.7

  2. Python v2.14.7python-v2.14.7Sep 22, 2026101 downloads

    ## What's Changed * feat: drastically reduce python import times by @logan-markewich in https://github.com/run-llama/liteparse/pull/462 * chore: update bench provider code by @logan-markewich in https://github.com/run-llama/liteparse/pull/464 * feat: expose the `parse()` pipeline as public stage functions by @logan-markewich in https://github.com/run-llama/liteparse/pull/467 **Full Changelog**: https://github.com/run-llama/liteparse/compare/node-v2.14.6...python-v2.14.7

  3. Node.js v2.14.7node-v2.14.7Sep 22, 2026138 downloads

    **Full Changelog**: https://github.com/run-llama/liteparse/compare/crates-v2.14.7...node-v2.14.7

  4. Crates & CLI v2.14.7crates-v2.14.7Sep 22, 2026189 downloads

    ## What's Changed * feat: drastically reduce python import times by @logan-markewich in https://github.com/run-llama/liteparse/pull/462 * chore: update bench provider code by @logan-markewich in https://github.com/run-llama/liteparse/pull/464 * feat: expose the `parse()` pipeline as public stage functions by @logan-markewich in https://github.com/run-llama/liteparse/pull/467 **Full Changelog**: https://github.com/run-llama/liteparse/compare/node-v2.14.6...crates-v2.14.7

  5. WASM v2.14.6wasm-v2.14.6Sep 15, 2026

    **Full Changelog**: https://github.com/run-llama/liteparse/compare/crates-v2.14.6...wasm-v2.14.6

Code frequency

additions and deletions
+133.8K-133.8KWeek of 2026-02-08: +82,176 linesWeek of 2026-02-08: -363 linesWeek of 2026-02-15: +13,473 linesWeek of 2026-02-15: -3,436 linesWeek of 2026-02-22: +2,077 linesWeek of 2026-02-22: -539 linesWeek of 2026-03-01: +2,115 linesWeek of 2026-03-01: -790 linesWeek of 2026-03-08: +2,835 linesWeek of 2026-03-08: -125 linesWeek of 2026-03-15: +2,544 linesWeek of 2026-03-15: -776 linesWeek of 2026-03-22: +4,489 linesWeek of 2026-03-22: -911 linesWeek of 2026-03-29: +1,473 linesWeek of 2026-03-29: -977 linesWeek of 2026-04-05: +1,309 linesWeek of 2026-04-05: -311 linesWeek of 2026-04-12: +98,197 linesWeek of 2026-04-12: -64,737 linesWeek of 2026-04-19: +692 linesWeek of 2026-04-19: -102 linesWeek of 2026-04-26: +19 linesWeek of 2026-04-26: -13 linesWeek of 2026-05-03: +7,382 linesWeek of 2026-05-03: -2,991 linesWeek of 2026-05-10: +24,440 linesWeek of 2026-05-10: -133,770 linesWeek of 2026-05-17: +4,854 linesWeek of 2026-05-17: -8,972 linesWeek of 2026-05-24: +6,045 linesWeek of 2026-05-24: -1,783 linesWeek of 2026-05-31: +16,443 linesWeek of 2026-05-31: -4,452 linesWeek of 2026-06-07: +8,240 linesWeek of 2026-06-07: -1,816 linesWeek of 2026-06-14: +2,428 linesWeek of 2026-06-14: -1,718 linesWeek of 2026-06-21: +3,815 linesWeek of 2026-06-21: -280 linesWeek of 2026-06-28: +1,892 linesWeek of 2026-06-28: -374 linesWeek of 2026-07-05: +2,249 linesWeek of 2026-07-05: -370 linesWeek of 2026-07-12: +2,847 linesWeek of 2026-07-12: -237 linesWeek of 2026-07-19: +10,204 linesWeek of 2026-07-19: -5,350 linesWeek of 2026-07-26: +2,551 linesWeek of 2026-07-26: -450 linesWeek of 2026-08-02: +42,713 linesWeek of 2026-08-02: -1,161 linesWeek of 2026-08-09: +121,328 linesWeek of 2026-08-09: -76,376 linesWeek of 2026-08-16: +16,596 linesWeek of 2026-08-16: -2,402 linesWeek of 2026-08-23: +19,999 linesWeek of 2026-08-23: -5,377 linesWeek of 2026-08-30: +2,123 linesWeek of 2026-08-30: -104,498 linesWeek of 2026-09-06: +2,871 linesWeek of 2026-09-06: -292 linesWeek of 2026-09-13: +2,441 linesWeek of 2026-09-13: -613 linesWeek of 2026-09-20: +647 linesWeek of 2026-09-20: -619 linesWeek of 2026-09-27: +0 linesWeek of 2026-09-27: -0 linesFeb 8, 2026Sep 27, 2026
+513.5K lines added, -427K removed over the last year.

Commits per week

last 52 weeks
700Week of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 10 commitsWeek of 2026-02-15: 55 commitsWeek of 2026-02-22: 9 commitsWeek of 2026-03-01: 8 commitsWeek of 2026-03-08: 14 commitsWeek of 2026-03-15: 34 commitsWeek of 2026-03-22: 70 commitsWeek of 2026-03-29: 23 commitsWeek of 2026-04-05: 11 commitsWeek of 2026-04-12: 16 commitsWeek of 2026-04-19: 12 commitsWeek of 2026-04-26: 5 commitsWeek of 2026-05-03: 10 commitsWeek of 2026-05-10: 55 commitsWeek of 2026-05-17: 65 commitsWeek of 2026-05-24: 69 commitsWeek of 2026-05-31: 40 commitsWeek of 2026-06-07: 36 commitsWeek of 2026-06-14: 26 commitsWeek of 2026-06-21: 18 commitsWeek of 2026-06-28: 17 commitsWeek of 2026-07-05: 11 commitsWeek of 2026-07-12: 10 commitsWeek of 2026-07-19: 34 commitsWeek of 2026-07-26: 14 commitsWeek of 2026-08-02: 24 commitsWeek of 2026-08-09: 29 commitsWeek of 2026-08-16: 34 commitsWeek of 2026-08-23: 36 commitsWeek of 2026-08-30: 12 commitsWeek of 2026-09-06: 22 commitsWeek of 2026-09-13: 13 commitsWeek of 2026-09-20: 4 commitsSep 27, 2025Sep 20, 2026
846 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 1 commitsSun 5:00 — 1 commitsSun 6:00 — 1 commitsSun 7:00 — 0 commitsSun 8:00 — 1 commitsSun 9:00 — 6 commitsSun 10:00 — 2 commitsSun 11:00 — 0 commitsSun 12:00 — 1 commitsSun 13:00 — 3 commitsSun 14:00 — 2 commitsSun 15:00 — 11 commitsSun 16:00 — 4 commitsSun 17:00 — 4 commitsSun 18:00 — 0 commitsSun 19:00 — 1 commitsSun 20:00 — 13 commitsSun 21:00 — 10 commitsSun 22:00 — 4 commitsSun 23:00 — 2 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 1 commitsMon 3:00 — 1 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 4 commitsMon 9:00 — 11 commitsMon 10:00 — 23 commitsMon 11:00 — 7 commitsMon 12:00 — 14 commitsMon 13:00 — 11 commitsMon 14:00 — 9 commitsMon 15:00 — 11 commitsMon 16:00 — 14 commitsMon 17:00 — 2 commitsMon 18:00 — 4 commitsMon 19:00 — 5 commitsMon 20:00 — 12 commitsMon 21:00 — 19 commitsMon 22:00 — 3 commitsMon 23:00 — 0 commitsTue 0:00 — 2 commitsTue 1:00 — 2 commitsTue 2:00 — 5 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 2 commitsTue 9:00 — 12 commitsTue 10:00 — 6 commitsTue 11:00 — 15 commitsTue 12:00 — 13 commitsTue 13:00 — 11 commitsTue 14:00 — 11 commitsTue 15:00 — 20 commitsTue 16:00 — 26 commitsTue 17:00 — 19 commitsTue 18:00 — 13 commitsTue 19:00 — 1 commitsTue 20:00 — 1 commitsTue 21:00 — 3 commitsTue 22:00 — 2 commitsTue 23:00 — 2 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 5 commitsWed 10:00 — 10 commitsWed 11:00 — 12 commitsWed 12:00 — 10 commitsWed 13:00 — 8 commitsWed 14:00 — 11 commitsWed 15:00 — 15 commitsWed 16:00 — 19 commitsWed 17:00 — 13 commitsWed 18:00 — 13 commitsWed 19:00 — 5 commitsWed 20:00 — 7 commitsWed 21:00 — 14 commitsWed 22:00 — 6 commitsWed 23:00 — 3 commitsThu 0:00 — 3 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 1 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 13 commitsThu 10:00 — 9 commitsThu 11:00 — 6 commitsThu 12:00 — 11 commitsThu 13:00 — 16 commitsThu 14:00 — 14 commitsThu 15:00 — 26 commitsThu 16:00 — 23 commitsThu 17:00 — 8 commitsThu 18:00 — 4 commitsThu 19:00 — 5 commitsThu 20:00 — 5 commitsThu 21:00 — 6 commitsThu 22:00 — 2 commitsThu 23:00 — 0 commitsFri 0:00 — 1 commitsFri 1:00 — 0 commitsFri 2:00 — 2 commitsFri 3:00 — 2 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 4 commitsFri 9:00 — 8 commitsFri 10:00 — 7 commitsFri 11:00 — 6 commitsFri 12:00 — 11 commitsFri 13:00 — 8 commitsFri 14:00 — 6 commitsFri 15:00 — 17 commitsFri 16:00 — 10 commitsFri 17:00 — 0 commitsFri 18:00 — 7 commitsFri 19:00 — 6 commitsFri 20:00 — 3 commitsFri 21:00 — 6 commitsFri 22:00 — 6 commitsFri 23:00 — 4 commitsSat 0:00 — 1 commitsSat 1:00 — 1 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 4 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 3 commitsSat 12:00 — 5 commitsSat 13:00 — 7 commitsSat 14:00 — 1 commitsSat 15:00 — 2 commitsSat 16:00 — 3 commitsSat 17:00 — 6 commitsSat 18:00 — 1 commitsSat 19:00 — 4 commitsSat 20:00 — 0 commitsSat 21:00 — 2 commitsSat 22:00 — 2 commitsSat 23:00 — 1 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
May 30, 2026daily#22+62
May 29, 2026daily#18+44
May 28, 2026daily#19+44
Mar 20, 2026daily#24+126
  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    195.2K stars · Rust

  • microsoft/markitdown

    Python tool for converting files and office documents to Markdown.

    188.4K stars · Python

  • farion1231/cc-switch

    A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

    140K stars · Rust

  • openai/codex

    Lightweight coding agent that runs in your terminal

    127.8K stars · Rust

  • denoland/deno

    A modern runtime for JavaScript and TypeScript.

    108.6K stars · Rust

  • ruvnet/RuView

    π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

    96.4K stars · Rust