opendataloader-project/opendataloader-pdfPublic

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.

Stars
28.3K
+62 today
Forks
2.7K
Watchers
113
Open issues
59
Open PRs
25
Contributors
~28
Commits
849
Branches
20

JavaApache-2.0Created May 13, 2025Last push todayLatest release v2.5.0+220 stars this week+260 this month

Star history

since Aug 24, 2025
010K20KAug 2025Dec 2025Apr 2026Aug 2026
28.3K stars as of Aug 7, 2026, tracked back to Aug 24, 2025. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 1 commit2025-08-13: 1 commit2025-08-14: 1 commit2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 1 commit2025-08-20: 1 commit2025-08-21: 0 commits2025-08-22: 1 commit2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 6 commits2025-08-27: 0 commits2025-08-28: 1 commit2025-08-29: 1 commit2025-08-30: 2 commits2025-08-31: 5 commits2025-09-01: 13 commits2025-09-02: 0 commits2025-09-03: 1 commit2025-09-04: 9 commits2025-09-05: 4 commits2025-09-06: 1 commit2025-09-07: 3 commits2025-09-08: 1 commit2025-09-09: 5 commits2025-09-10: 0 commits2025-09-11: 5 commits2025-09-12: 8 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 10 commits2025-09-16: 3 commits2025-09-17: 6 commits2025-09-18: 0 commits2025-09-19: 2 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 4 commits2025-09-23: 0 commits2025-09-24: 1 commit2025-09-25: 1 commit2025-09-26: 7 commits2025-09-27: 1 commit2025-09-28: 0 commits2025-09-29: 4 commits2025-09-30: 7 commits2025-10-01: 2 commits2025-10-02: 4 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 2 commits2025-10-08: 0 commits2025-10-09: 1 commit2025-10-10: 5 commits2025-10-11: 1 commit2025-10-12: 0 commits2025-10-13: 1 commit2025-10-14: 0 commits2025-10-15: 3 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 2 commits2025-10-21: 2 commits2025-10-22: 1 commit2025-10-23: 1 commit2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 1 commit2025-10-27: 0 commits2025-10-28: 1 commit2025-10-29: 1 commit2025-10-30: 4 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 1 commit2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 2 commits2025-11-07: 1 commit2025-11-08: 0 commits2025-11-09: 2 commits2025-11-10: 2 commits2025-11-11: 3 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 2 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 3 commits2025-11-19: 1 commit2025-11-20: 1 commit2025-11-21: 1 commit2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 2 commits2025-11-25: 2 commits2025-11-26: 0 commits2025-11-27: 1 commit2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 3 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 1 commit2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 2 commits2025-12-09: 0 commits2025-12-10: 3 commits2025-12-11: 0 commits2025-12-12: 1 commit2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 7 commits2025-12-16: 8 commits2025-12-17: 9 commits2025-12-18: 5 commits2025-12-19: 7 commits2025-12-20: 10 commits2025-12-21: 1 commit2025-12-22: 7 commits2025-12-23: 1 commit2025-12-24: 1 commit2025-12-25: 0 commits2025-12-26: 2 commits2025-12-27: 1 commit2025-12-28: 0 commits2025-12-29: 1 commit2025-12-30: 2 commits2025-12-31: 6 commits2026-01-01: 0 commits2026-01-02: 27 commits2026-01-03: 32 commits2026-01-04: 5 commits2026-01-05: 11 commits2026-01-06: 3 commits2026-01-07: 2 commits2026-01-08: 5 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 2 commits2026-01-13: 7 commits2026-01-14: 0 commits2026-01-15: 2 commits2026-01-16: 1 commit2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 5 commits2026-01-21: 1 commit2026-01-22: 6 commits2026-01-23: 2 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 1 commit2026-01-27: 0 commits2026-01-28: 1 commit2026-01-29: 2 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 1 commit2026-02-02: 1 commit2026-02-03: 7 commits2026-02-04: 2 commits2026-02-05: 7 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 1 commit2026-02-10: 1 commit2026-02-11: 1 commit2026-02-12: 3 commits2026-02-13: 1 commit2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 1 commit2026-02-19: 1 commit2026-02-20: 7 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 11 commits2026-02-24: 2 commits2026-02-25: 2 commits2026-02-26: 8 commits2026-02-27: 2 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 2 commits2026-03-03: 2 commits2026-03-04: 6 commits2026-03-05: 1 commit2026-03-06: 3 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 5 commits2026-03-10: 12 commits2026-03-11: 15 commits2026-03-12: 1 commit2026-03-13: 6 commits2026-03-14: 0 commits2026-03-15: 1 commit2026-03-16: 13 commits2026-03-17: 1 commit2026-03-18: 6 commits2026-03-19: 6 commits2026-03-20: 11 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 9 commits2026-03-24: 17 commits2026-03-25: 9 commits2026-03-26: 5 commits2026-03-27: 4 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 5 commits2026-03-31: 12 commits2026-04-01: 8 commits2026-04-02: 2 commits2026-04-03: 3 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 8 commits2026-04-07: 1 commit2026-04-08: 1 commit2026-04-09: 0 commits2026-04-10: 5 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 1 commit2026-04-14: 19 commits2026-04-15: 6 commits2026-04-16: 10 commits2026-04-17: 18 commits2026-04-18: 11 commits2026-04-19: 0 commits2026-04-20: 7 commits2026-04-21: 0 commits2026-04-22: 1 commit2026-04-23: 0 commits2026-04-24: 2 commits2026-04-25: 0 commits2026-04-26: 1 commit2026-04-27: 0 commits2026-04-28: 2 commits2026-04-29: 14 commits2026-04-30: 13 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 2 commits2026-05-05: 1 commit2026-05-06: 4 commits2026-05-07: 1 commit2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 2 commits2026-05-13: 5 commits2026-05-14: 10 commits2026-05-15: 6 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 2 commits2026-05-19: 3 commits2026-05-20: 4 commits2026-05-21: 3 commits2026-05-22: 2 commits2026-05-23: 0 commits2026-05-24: 1 commit2026-05-25: 1 commit2026-05-26: 0 commits2026-05-27: 2 commits2026-05-28: 0 commits2026-05-29: 1 commit2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 1 commit2026-06-02: 0 commits2026-06-03: 2 commits2026-06-04: 1 commit2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 1 commit2026-06-08: 1 commit2026-06-09: 3 commits2026-06-10: 0 commits2026-06-11: 3 commits2026-06-12: 3 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 1 commit2026-06-16: 1 commit2026-06-17: 0 commits2026-06-18: 17 commits2026-06-19: 2 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 4 commits2026-06-23: 3 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 1 commit2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 1 commit2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 1 commit2026-07-09: 3 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 2 commits2026-07-14: 4 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 3 commits2026-07-21: 15 commits2026-07-22: 2 commits2026-07-23: 1 commit2026-07-24: 2 commits2026-07-25: 1 commit2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 2 commits2026-07-29: 1 commit2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 1 commit2026-08-05: 0 commits2026-08-06: 1 commit2026-08-07: 0 commits2026-08-08: 0 commits
844 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    28,250 stars

  • Very active

    844 commits in 52 weeks

  • Well documented

    High community health score

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    6 trending appearances

What opendataloader-pdf does

OpenDataLoader PDF solves the notoriously difficult problem of extracting structured, accurate data from complex PDF documents for AI consumption. It achieves top-tier benchmark scores (0.907 overall) by accurately interpreting multi-column layouts, complex tables, and scientific formatting. The technical approach relies on a fast, deterministic local parsing engine that extracts bounding boxes and text, coupled with an AI hybrid mode that activates only for highly complex pages that defy standard rules. What makes it distinctive is its focus on producing truly 'AI-ready' formats like Markdown and JSON with precise spatial coordinates, bridging the gap between visual documents and language models. It eliminates the need for expensive, cloud-based OCR APIs for the vast majority of document processing tasks.

This project is built for data engineers, AI developers, and researchers building document-heavy data pipelines. It requires basic familiarity with Python or CLI tools and an understanding of data formats like JSON and Markdown.

  • Deterministic Local Extraction: processes standard PDFs rapidly on local hardware without requiring external AI or cloud API calls.
  • AI Hybrid Fallback: intelligently routes highly complex or malformed pages to an AI engine to ensure accurate extraction.
  • Spatial Coordinate Mapping: outputs JSON that includes precise bounding boxes, allowing downstream applications to understand document layout.
  • Advanced Table Parsing: achieves a 0.928 accuracy score on tables, correctly identifying rows and columns even in borderless formats.
  • Markdown Generation: seamlessly converts complex PDF formatting into clean, semantic Markdown optimized for RAG pipelines.

Where teams use it

RAG Pipeline Ingestion

Data engineers use the tool to convert massive archives of corporate PDFs into clean Markdown for indexing in vector databases.

Scientific Paper Analysis

Researchers extract tables and multi-column text from dense academic papers to train specialized domain models.

Accessibility Automation

Compliance teams leverage the precise bounding boxes and HTML output to automatically generate tagged, accessible versions of legacy documents.

Financial Document Scraping

Financial analysts parse quarterly reports to reliably extract tabular data into structured JSON for automated modeling.

Getting started: pip install opendataloader-pdf && odl-pdf extract sample.pdf

README

main branch

OpenDataLoader PDF

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

License PyPI version npm version Maven Central Java

opendataloader-project%2Fopendataloader-pdf | Trendshift

🔍 PDF parser for AI data extraction — Extract Markdown, JSON (with bounding boxes), and HTML from any PDF. #1 in benchmarks (0.907 overall). Deterministic local mode + AI hybrid mode for complex pages.

  • How accurate is it? — #1 in benchmarks: 0.907 overall, 0.928 table accuracy across 200 real-world PDFs including multi-column and scientific papers. Deterministic local mode + AI hybrid mode for complex pages (benchmarks)
  • Scanned PDFs and OCR? — Yes. Built-in OCR (80+ languages) in hybrid mode. Works with poor-quality scans at 300 DPI+ (hybrid mode)
  • Tables, formulas, images, charts? — Yes. Complex/borderless tables, LaTeX formulas, and AI-generated picture/chart descriptions all via hybrid mode (hybrid mode)
  • How do I use this for RAG?pip install opendataloader-pdf, convert in 3 lines. Outputs structured Markdown for chunking, JSON with bounding boxes for source citations, and HTML. LangChain integration available. Python, Node.js, Java SDKs (quick start | LangChain)

PDF accessibility automation — Auto-tag untagged PDFs into screen-reader-ready Tagged PDFs at scale. First open-source tool to generate Tagged PDFs end-to-end.

  • What's the problem? — Accessibility regulations are now enforced worldwide. Manual PDF remediation costs $50–200 per document and doesn't scale (regulations)
  • What's free? — Layout analysis + auto-tagging (Apache 2.0). Untagged PDF in → Tagged PDF out. No proprietary SDK dependency (auto-tagging)
  • What about PDF/UA compliance? — Converting Tagged PDF to PDF/UA-1 or PDF/UA-2 is an enterprise add-on. Auto-tagging generates the Tagged PDF; PDF/UA export is the final step (pipeline)
  • Why trust this? — Built in collaboration with Dual Lab (veraPDF developers) based on PDF Association specifications, best practice guides and expertise of the PDF Community. Auto-tagging follows the Well-Tagged PDF specification, validated with veraPDF (collaboration)

Get Started in 30 Seconds

Requires: Java 11+ and Python 3.10+ (Node.js | Java also available)

Before you start: run java -version. If not found, install JDK 11+ from Adoptium.

pip install -U opendataloader-pdf
import opendataloader_pdf

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="markdown,json"
)

OpenDataLoader PDF layout analysis — headings, tables, images detected with bounding boxes

Annotated PDF output — each element (heading, paragraph, table, image) detected with bounding boxes and semantic type.

What Problems Does This Solve?

Problem Solution Status
PDF structure lost during parsing — wrong reading order, broken tables, no element coordinates Deterministic local PDF to Markdown/JSON with bounding boxes, XY-Cut++ reading order Shipped
Complex tables, scanned PDFs, formulas, charts need AI-level understanding Hybrid mode routes complex pages to AI backend (#1 in benchmarks) Shipped
Manual PDF remediation cost — Accessibility regulations (EAA, ADA, Section 508) demand Tagged PDFs. Manual remediation costs $50–200/doc Auto-tag untagged PDFs into Tagged PDFs (free, Apache 2.0). Foundation for PDF/UA workflows; full PDF/UA-1/2 export is an enterprise add-on Auto-tag: Shipped. PDF/UA export: Enterprise

Capability Matrix

Capability Supported Tier
Data extraction
Extract text with correct reading order Yes Free
Bounding boxes for every element Yes Free
Table extraction (simple borders) Yes Free
Table extraction (complex/borderless) Yes Free (Hybrid)
Heading hierarchy detection Yes Free
List detection (numbered, bulleted, nested) Yes Free
Image extraction with coordinates Yes Free
AI chart/image description Yes Free (Hybrid)
OCR for scanned PDFs Yes Free (Hybrid)
Formula extraction (LaTeX) Yes Free (Hybrid)
Tagged PDF structure extraction Yes Free
AI safety (prompt injection filtering) Yes Free
Header/footer/watermark filtering Yes Free
Accessibility
Auto-tagging → Tagged PDF for untagged PDFs Yes Free (Apache 2.0)
PDF/UA-1, PDF/UA-2 export 💼 Available Enterprise
Accessibility studio (visual editor) 💼 Available Enterprise
Limitations
Process Word/Excel/PPT No
GPU required No

Extraction Benchmarks

opendataloader-pdf [hybrid] ranks #1 overall (0.907) across reading order, table, and heading extraction accuracy.

Engine Overall Reading Order Table Heading Speed (s/page) License
opendataloader [hybrid] 0.907 0.934 0.928 0.821 0.463 Apache-2.0
nutrient 0.885 0.925 0.708 0.819 0.008 Commercial
docling 0.882 0.898 0.887 0.824 0.762 MIT
marker 0.861 0.890 0.808 0.796 53.932 GPL-3.0
unstructured [hi_res] 0.841 0.904 0.588 0.749 3.008 Apache-2.0
edgeparse 0.837 0.894 0.717 0.706 0.036 Apache-2.0
opendataloader 0.831 0.902 0.489 0.739 0.015 Apache-2.0
mineru 0.831 0.857 0.873 0.743 5.962 AGPL-3.0
pymupdf4llm 0.732 0.885 0.401 0.412 0.091 AGPL-3.0
unstructured 0.686 0.882 0.000 0.388 0.077 Apache-2.0
markitdown 0.589 0.844 0.273 0.000 0.114 MIT
liteparse 0.576 0.866 0.000 0.000 1.061 Apache-2.0

Scores normalized to [0, 1]. Higher is better for accuracy; lower is better for speed. Bold = best. Full benchmark details

Benchmark

Quality Breakdown

Which Mode Should I Use?

Your Document Mode Install Server Command Client Command
Standard digital PDF Fast (default) pip install opendataloader-pdf None needed opendataloader-pdf file1.pdf file2.pdf folder/
Complex or nested tables Hybrid pip install "opendataloader-pdf[hybrid]" opendataloader-pdf-hybrid --port 5002 opendataloader-pdf --hybrid docling-fast file1.pdf file2.pdf folder/
Scanned / image-based PDF Hybrid + OCR pip install "opendataloader-pdf[hybrid]" opendataloader-pdf-hybrid --port 5002 --force-ocr opendataloader-pdf --hybrid docling-fast file1.pdf file2.pdf folder/
Non-English scanned PDF Hybrid + OCR pip install "opendataloader-pdf[hybrid]" opendataloader-pdf-hybrid --port 5002 --force-ocr --ocr-lang "ko,en" opendataloader-pdf --hybrid docling-fast file1.pdf file2.pdf folder/
Mathematical formulas Hybrid + formula pip install "opendataloader-pdf[hybrid]" opendataloader-pdf-hybrid --enrich-formula opendataloader-pdf --hybrid docling-fast --hybrid-mode full file1.pdf file2.pdf folder/
Charts needing description Hybrid + picture pip install "opendataloader-pdf[hybrid]" opendataloader-pdf-hybrid --enrich-picture-description opendataloader-pdf --hybrid docling-fast --hybrid-mode full file1.pdf file2.pdf folder/
Untagged PDFs needing accessibility Auto-tagging → Tagged PDF pip install opendataloader-pdf None needed opendataloader-pdf --format tagged-pdf file1.pdf file2.pdf folder/

Quick Start

Python

pip install -U opendataloader-pdf
import opendataloader_pdf

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="markdown,json"
)

Node.js

npm install @opendataloader/pdf
import { convert } from '@opendataloader/pdf';

await convert(['file1.pdf', 'file2.pdf', 'folder/'], {
  outputDir: 'output/',
  format: 'markdown,json'
});

Java

<dependency>
  <groupId>org.opendataloader</groupId>
  <artifactId>opendataloader-pdf-core</artifactId>
</dependency>

Python Quick Start | Node.js Quick Start | Java Quick Start

Hybrid Mode: #1 Accuracy for Complex PDFs

Hybrid mode combines fast local Java processing with AI backends. Simple pages stay local (0.02s); complex pages route to AI for +90% table accuracy.

Don't combine with --use-struct-tree on tagged PDFs. --use-struct-tree takes precedence, so the hybrid backend is not called (a warning is logged). If you want the hybrid backend, drop --use-struct-tree.

pip install -U "opendataloader-pdf[hybrid]"

Terminal 1 — Start the backend server:

opendataloader-pdf-hybrid --port 5002

Terminal 2 — Process PDFs:

# Batch all files in one call — each invocation spawns a JVM process, so repeated calls are slow
opendataloader-pdf --hybrid docling-fast file1.pdf file2.pdf folder/

Python:

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    hybrid="docling-fast"
)

OCR for Scanned PDFs

Start the backend with --force-ocr for image-based PDFs with no selectable text:

opendataloader-pdf-hybrid --port 5002 --force-ocr

For non-English documents, specify the language:

opendataloader-pdf-hybrid --port 5002 --force-ocr --ocr-lang "ko,en"

Supported languages: en, ko, ja, ch_sim, ch_tra, de, fr, ar, and more.

Formula Extraction (LaTeX)

Extract mathematical formulas as LaTeX from scientific PDFs:

# Server: enable formula enrichment
opendataloader-pdf-hybrid --enrich-formula

# Batch all files in one call — each invocation spawns a JVM process, so repeated calls are slow
opendataloader-pdf --hybrid docling-fast --hybrid-mode full file1.pdf file2.pdf folder/

Output in JSON:

{
  "type": "formula",
  "page number": 1,
  "bounding box": [226.2, 144.7, 377.1, 168.7],
  "content": "\\frac{f(x+h) - f(x)}{h}"
}

Note: Formula and picture description enrichments require --hybrid-mode full on the client side.

Chart & Image Description

Generate AI descriptions for charts and images — useful for RAG search and accessibility alt text:

# Server
opendataloader-pdf-hybrid --enrich-picture-description

# Batch all files in one call — each invocation spawns a JVM process, so repeated calls are slow
opendataloader-pdf --hybrid docling-fast --hybrid-mode full file1.pdf file2.pdf folder/

Output in JSON:

{
  "type": "picture",
  "page number": 1,
  "bounding box": [72.0, 400.0, 540.0, 650.0],
  "description": "A bar chart showing waste generation by region from 2016 to 2030..."
}

Uses SmolVLM (256M), a lightweight vision model. Custom prompts supported via --picture-description-prompt.

Hancom Data Loader Integration — Coming Soon

Enterprise-grade AI document analysis via Hancom Data Loader — customer-customized models trained on your domain-specific documents. 30+ element types (tables, charts, formulas, captions, footnotes, etc.), VLM-based image/chart understanding, complex table extraction (merged cells, nested tables), SLA-backed OCR for scanned documents, and native HWP/HWPX support. Supports PDF, DOCX, XLSX, PPTX, HWP, PNG, JPG. Live demo

Hybrid Mode Guide

Output Formats

Format Use Case
JSON Structured data with bounding boxes, semantic types
Markdown Clean text for LLM context, RAG chunks
HTML Web display with styling
Annotated PDF Visual debugging — see detected structures (sample)
Text Plain text extraction

Combine formats: format="json,markdown"

JSON Output Example

{
  "type": "heading",
  "id": 42,
  "level": "Title",
  "page number": 1,
  "bounding box": [72.0, 700.0, 540.0, 730.0],
  "heading level": 1,
  "font": "Helvetica-Bold",
  "font size": 24.0,
  "text color": "[0.0]",
  "content": "Introduction"
}
Field Description
type Element type: heading, paragraph, table, list, image, caption, formula
id Unique identifier for cross-referencing
page number 1-indexed page reference
bounding box [left, bottom, right, top] in PDF points (72pt = 1 inch)
heading level Heading depth (1+)
content Extracted text

Full JSON Schema

Advanced Features

Tagged PDF Support

When a PDF has structure tags, OpenDataLoader extracts the exact layout the author intended — no guessing, no heuristics. Headings, lists, tables, and reading order are preserved from the source.

Output quality depends on tag quality. Not all tagged PDFs are well-tagged. For PDFs with sparse or incorrect tags, the default heuristic mode or --hybrid docling-fast often produces better results.

--use-struct-tree takes precedence over --hybrid. If both are set on a tagged PDF, the structure tree is used and the hybrid backend is not called (a well-tagged PDF already carries reading order and structure). Drop --use-struct-tree if you want the hybrid backend instead.

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    use_struct_tree=True           # Use native PDF structure tags
)

Most PDF parsers ignore structure tags entirely. Learn more

AI Safety: Prompt Injection Protection

PDFs can contain hidden prompt injection attacks. OpenDataLoader automatically filters:

  • Hidden text (transparent, zero-size fonts)
  • Off-page content
  • Suspicious invisible layers

To sanitize sensitive data (emails, URLs, phone numbers → placeholders), enable it explicitly:

# Batch all files in one call — each invocation spawns a JVM process, so repeated calls are slow
opendataloader-pdf file1.pdf file2.pdf folder/ --sanitize

AI Safety Guide

LangChain Integration

pip install -U langchain-opendataloader-pdf
from langchain_opendataloader_pdf import OpenDataLoaderPDFLoader

loader = OpenDataLoaderPDFLoader(
    file_path=["file1.pdf", "file2.pdf", "folder/"],
    format="text"
)
documents = loader.load()

LangChain Docs | GitHub | PyPI

Advanced Options

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="json,markdown,pdf",
    image_output="embedded",        # "off", "embedded" (Base64), or "external" (default)
    image_format="jpeg",            # "png" or "jpeg"
    use_struct_tree=True,           # Use native PDF structure
)

Full CLI Options Reference

PDF Accessibility & PDF/UA Conversion

Problem: Millions of existing PDFs lack structure tags, failing accessibility regulations (EAA, ADA/Section 508, Korea Digital Inclusion Act). Manual remediation costs $50–200 per document and doesn't scale.

OpenDataLoader's approach: Built in collaboration with PDF Association and Dual Lab (developers of veraPDF, the industry-reference open-source PDF/A and PDF/UA validator). Auto-tagging follows the Well-Tagged PDF specification and is validated programmatically using veraPDF — automated conformance checks against PDF accessibility standards, not manual review. No existing open-source tool generates Tagged PDFs end-to-end — most rely on proprietary SDKs for the tag-writing step. OpenDataLoader does it all under Apache 2.0. (collaboration details)

Regulation Deadline Requirement
European Accessibility Act (EAA) June 28, 2025 Accessible digital products across the EU
ADA & Section 508 In effect U.S. federal agencies and public accommodations
Digital Inclusion Act In effect South Korea digital service accessibility

Standards & Validation

Aspect Detail
Specification Well-Tagged PDF by PDF Association
Validation veraPDF — industry-reference open-source PDF/A & PDF/UA validator
Collaboration PDF Association + Dual Lab (veraPDF developers) co-develop tagging and validation
License Auto-tagging → Tagged PDF: Apache 2.0 (free). PDF/UA export: Enterprise

Accessibility Pipeline

Step Feature Status Tier
1. Audit Read existing PDF tags, detect untagged PDFs Shipped Free
2. Auto-tag → Tagged PDF Generate structure tags for untagged PDFs Shipped Free (Apache 2.0)
3. Export PDF/UA Convert to PDF/UA-1 or PDF/UA-2 compliant files 💼 Available Enterprise
4. Visual editing Accessibility studio — review and fix tags 💼 Available Enterprise

💼 Enterprise features are available on request. Contact us to get started.

Auto-Tagging

Generate Tagged PDFs from untagged PDFs — output is a screen-reader-ready PDF with structure tags (headings, paragraphs, lists, tables, reading order).

import opendataloader_pdf

# Untagged PDF in → Tagged PDF out
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="tagged-pdf"
)
# CLI
opendataloader-pdf --format tagged-pdf file1.pdf file2.pdf folder/

Combine with other formats: format="json,tagged-pdf".

End-to-End Compliance Workflow

Existing PDFs (untagged)
    │
    ▼
┌─────────────────┐    ┌──────────────────┐    ┌─────────────────┐    ┌──────────────────┐
│  1. Audit       │───>│  2. Auto-Tag     │───>│  3. Export      │───>│  4. Studio       │
│  (check tags)   │    │  (→ Tagged PDF)  │    │  (PDF/UA)       │    │  (visual editor) │
└─────────────────┘    └──────────────────┘    └─────────────────┘    └──────────────────┘
        │                       │                       │                      │
        ▼                       ▼                       ▼                      ▼
  use_struct_tree      format="tagged-pdf"        PDF/UA export       Accessibility Studio
  (Available now)      (Available, Apache 2.0)    (Enterprise)        (Enterprise)

PDF Accessibility Guide

Roadmap

Feature Timeline Tier
Hancom Data Loader — Enterprise AI document analysis, customer-customized models, VLM-based chart/image understanding, production-grade OCR Q2-Q3 2026 Planned
Structure validation — Verify PDF tag trees Q3 2026 Planned

Full Roadmap

Frequently Asked Questions

What is the best PDF parser for RAG?

For RAG pipelines, you need a parser that preserves document structure, maintains correct reading order, and provides element coordinates for citations. OpenDataLoader is designed specifically for this — it outputs structured JSON with bounding boxes, handles multi-column layouts with XY-Cut++, and runs locally without GPU. In hybrid mode, it ranks #1 overall (0.907) in benchmarks.

What is the best open-source PDF parser?

OpenDataLoader PDF is the only open-source parser that combines: rule-based deterministic extraction (no GPU), bounding boxes for every element, XY-Cut++ reading order, built-in AI safety filters, native Tagged PDF support, and hybrid AI mode for complex documents. It ranks #1 in overall accuracy (0.907) while running locally on CPU.

How do I extract tables from PDF for LLM?

OpenDataLoader detects tables using border analysis and text clustering, preserving row/column structure. For complex tables, enable hybrid mode for +90% accuracy improvement (0.489 to 0.928 TEDS score):

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="json",
    hybrid="docling-fast"           # For complex tables
)

How does it compare to docling, marker, or pymupdf4llm?

OpenDataLoader [hybrid] ranks #1 overall (0.907) across reading order, table, and heading accuracy. Key differences: docling (0.882) is strong but lacks bounding boxes and AI safety filters. marker (0.861) requires GPU and is 1000x slower (53.932s/page). pymupdf4llm (0.732) is fast but has poor table (0.401) and heading (0.412) accuracy. OpenDataLoader is the only parser that combines deterministic local extraction, bounding boxes for every element, and built-in prompt injection protection. See full benchmark.

Can I use this without sending data to the cloud?

Yes. OpenDataLoader runs 100% locally. No API calls, no data transmission — your documents never leave your environment. The hybrid mode backend also runs locally on your machine. Ideal for legal, healthcare, and financial documents.

Does it support OCR for scanned PDFs?

Yes, via hybrid mode. Install with pip install "opendataloader-pdf[hybrid]", start the backend with --force-ocr, then process as usual. Supports multiple languages including Korean, Japanese, Chinese, Arabic, and more via --ocr-lang.

Does it work with Korean, Japanese, or Chinese documents?

Yes. For digital PDFs, text extraction works out of the box. For scanned PDFs, use hybrid mode with --force-ocr --ocr-lang "ko,en" (or ja, ch_sim, ch_tra). Coming soon: Hancom Data Loader integration — enterprise-grade AI document analysis with built-in production-grade OCR and customer-customized models optimized for your specific document types and workflows.

How fast is it?

Local mode processes 60+ pages per second on CPU (0.02s/page). Hybrid mode processes 2+ pages per second (0.46s/page) with significantly higher accuracy for complex documents. No GPU required. Benchmarked on Apple M4. Full benchmark details. With multi-process batch processing, throughput exceeds 100 pages per second on 8+ core machines.

Does it handle multi-column layouts?

Yes. OpenDataLoader uses XY-Cut++ reading order analysis to correctly sequence text across multi-column pages, sidebars, and mixed layouts. This works in both local and hybrid modes without any configuration.

What is hybrid mode?

Hybrid mode combines fast local Java processing with an AI backend. Simple pages are processed locally (0.02s/page); complex pages (tables, scanned content, formulas, charts) are automatically routed to the AI backend for higher accuracy. The backend runs locally on your machine — no cloud required. See Which Mode Should I Use? and Hybrid Mode Guide.

Does it work with LangChain?

Yes. Install langchain-opendataloader-pdf for an official LangChain document loader integration. See LangChain docs.

How do I chunk PDFs for RAG?

OpenDataLoader outputs structured Markdown with headings, tables, and lists preserved — ideal input for semantic chunking. Each element in JSON output includes type, heading level, and page number, so you can split by section or page boundary. For most RAG pipelines: parse with format="markdown" for text chunks, or format="json" when you need element-level control. Pair with LangChain's RecursiveCharacterTextSplitter or your own heading-based splitter for best results.

How do I cite PDF sources in RAG answers?

Every element in JSON output includes a bounding box ([left, bottom, right, top] in PDF points) and page number. When your RAG pipeline returns an answer, map the source chunk back to its bounding box to highlight the exact location in the original PDF. This enables "click to source" UX — users see which paragraph, table, or figure the answer came from. No other open-source parser provides bounding boxes for every element by default.

How do I convert PDF to Markdown for LLM?

import opendataloader_pdf

# Batch all files in one call — each convert() spawns a JVM process, so repeated calls are slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="markdown"
)

OpenDataLoader preserves heading hierarchy, table structure, and reading order in the Markdown output. For complex documents with borderless tables or scanned pages, use hybrid mode (hybrid="docling-fast") for higher accuracy. The output is clean enough to feed directly into LLM context windows or RAG chunking pipelines.

Is there an automated PDF accessibility remediation tool?

Yes. OpenDataLoader is the first open-source tool that automates PDF accessibility end-to-end. Built in collaboration with PDF Association and Dual Lab (veraPDF developers), auto-tagging follows the Well-Tagged PDF specification and is validated programmatically using veraPDF. The layout analysis engine detects document structure (headings, tables, lists, reading order) and generates accessibility tags automatically. Auto-tagging converts untagged PDFs into Tagged PDFs under Apache 2.0 — no proprietary SDK dependency. Use format="tagged-pdf" (Python/Node.js) or --format tagged-pdf (CLI). For organizations needing full PDF/UA compliance, enterprise add-ons provide PDF/UA export and a visual tag editor. This replaces manual remediation workflows that typically cost $50–200+ per document.

Is this really the first open-source PDF auto-tagging tool?

Yes. Existing tools either depend on proprietary SDKs for writing structure tags, only output non-PDF formats (e.g., Docling outputs Markdown/JSON but cannot produce Tagged PDFs), or require manual intervention. OpenDataLoader is the first to do layout analysis → tag generation → Tagged PDF output entirely under an open-source license (Apache 2.0), with no proprietary dependency. Auto-tagging follows the PDF Association's Well-Tagged PDF specification and is validated using veraPDF, the industry-reference open-source PDF/A and PDF/UA validator.

How do I convert existing PDFs to PDF/UA?

OpenDataLoader provides an end-to-end pipeline: audit existing PDFs for tags (use_struct_tree=True), auto-tag untagged PDFs into Tagged PDFs (format="tagged-pdf", free under Apache 2.0), and export as PDF/UA-1 or PDF/UA-2 (enterprise add-on). Auto-tagging follows the PDF Association's Well-Tagged PDF specification and is validated using veraPDF. Auto-tagging generates the Tagged PDF; PDF/UA export is the final step. Contact us for enterprise integration.

How do I make my PDFs accessible for EAA compliance?

The European Accessibility Act requires accessible digital products by June 28, 2025. OpenDataLoader supports the full remediation workflow: audit → auto-tag → Tagged PDF → PDF/UA export. Auto-tagging follows the PDF Association's Well-Tagged PDF specification and is validated using veraPDF, ensuring standards-compliant output. Auto-tagging to Tagged PDF is open-source under Apache 2.0. PDF/UA export and accessibility studio are enterprise add-ons. See our Accessibility Guide.

Is OpenDataLoader PDF free?

The core library is open-source under Apache 2.0 — free for commercial use. This includes all extraction features (text, tables, images, OCR, formulas, charts via hybrid mode), AI safety filters, Tagged PDF support, and auto-tagging to Tagged PDF. We are committed to keeping the core accessibility pipeline (layout analysis → auto-tagging → Tagged PDF) free and open-source. Enterprise add-ons (PDF/UA export, accessibility studio) are available for organizations needing end-to-end regulatory compliance.

Why did the license change from MPL 2.0 to Apache 2.0?

MPL 2.0 requires file-level copyleft, which often triggers legal review before enterprise adoption. Apache 2.0 is fully permissive — no copyleft obligations, easier to integrate into commercial projects. If you are using a pre-2.0 version, it remains under MPL 2.0 and you can continue using it. Upgrading to 2.0+ means your project follows Apache 2.0 terms, which are strictly more permissive — no additional obligations, no action needed on your side.

Documentation

Contributing

We welcome contributions! See CONTRIBUTING.md for guidelines.

License

Apache License 2.0

Note: Versions prior to 2.0 are licensed under the Mozilla Public License 2.0.


Found this useful? Give us a star to help others discover OpenDataLoader.

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Discussions

all 15

Releases and announcements

63 total
  1. Release v2.5.0v2.5.0Jul 14, 2026581 downloads

    ## What's Changed * feat(json): unify image alt schema with alt + alt_source, preserve picture metadata by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/537 * fix(hybrid): route crops/page images to the current document, not the first by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/541 * Add TOC to outputs and TaggedDocumentProcessor by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/552 * Improve location of structure elements for annots by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/557 * Improve table header processing by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/559 * fix: detect strikethroughs from line arts by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/560 * Fix escaping of HTML text output by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/561 * fix(images): release BufferedImage and clarify ContrastRatioConsumer ownership (#458) by @hnc-jglee in https://github.com/opendataloader-project/opendataloader-pdf/pull/543 * Improve

  2. Release v2.4.7v2.4.7May 27, 20261.5K downloads

    ## What's Changed * feat(hybrid): persist full-page DLA renders for evidence overlays by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/529 * feat(json): emit per-node ai_score + pdfua_tag, fix paragraph/table metadata holes by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/530 * feat(hybrid): expose DLA raw object_id on ElementMetadata for downstream consumers by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/532 * Update HTML output with layout attributes by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/524 * Add font-size property to formatted html by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/533 * Improve StrikethroughProcessor by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/534 * Auto-tagging - Improve location of structure elements for annots by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/520 * fix(auto-tagging): assign unique /ID to Note / FENote struct elements (PDF/UA-1 §7.9.1) by @bundolee in https://github.com/opendataload

  3. Release v2.4.6v2.4.6May 21, 2026139 downloads

    ## What's Changed * fix(tagged-pdf): emit "Created" info log after saving (PDFDLOSP-26) by @hnc-jglee in https://github.com/opendataloader-project/opendataloader-pdf/pull/519 * chore(deps): bump idna to 3.15 (GHSA-65pc-fj4g-8rjx) by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/522 * fix(json): apply --include-header-footer to JSON output by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/523 * fix(image-output): make embedded mode self-contained by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/525 * feat(hybrid): expose raw response + reference environment + client wall-clock for downstream tooling by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/526 * fix(processors): surface clean error for corrupted or truncated PDFs by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/527 **Full Changelog**: https://github.com/opendataloader-project/opendataloader-pdf/compare/v2.4.4...v2.4.6

  4. Release v2.4.4v2.4.4May 19, 202642 downloads

    ## What's Changed * fix(python): deduplicate CLI failure output by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/495 * Add check for PUA in Alt entry by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/491 * chore(deps): bump urllib3 to 2.7.0 and python-multipart to 0.0.28 for security advisories by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/498 * fix(cli): emit error for non-PDF top-level input by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/496 * Tagged PDF - Add id to content items by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/501 * fix(hybrid)!: fail fast when backend left pages unprocessed with fallback disabled by @hyunhee-jo in https://github.com/opendataloader-project/opendataloader-pdf/pull/499 * fix(hybrid): forward --picture-description-prompt to docling VLM by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/503 * fix(cli): announce when a folder contains zero processable PDFs (PDFDLOSP-15) by @bundolee in https://github.com/opendataloader-project/opendatal

  5. Release v2.4.3v2.4.3May 7, 2026340 downloads

    ## What's Changed * Auto-tagging. Fix PDFStreamWriter by @MaximPlusov in https://github.com/opendataloader-project/opendataloader-pdf/pull/481 * Add pdf version option (from pdfua) by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/487 * Remove StructParent from annotations that are artifacts by @LonelyMidoriya in https://github.com/opendataloader-project/opendataloader-pdf/pull/488 * fix: pass headless JVM flags so macOS does not show Dock icon by @bundolee in https://github.com/opendataloader-project/opendataloader-pdf/pull/489 **Full Changelog**: https://github.com/opendataloader-project/opendataloader-pdf/compare/v2.4.2...v2.4.3

Code frequency

additions and deletions
+120.7K-120.7KWeek of 2025-08-10: +9,105 linesWeek of 2025-08-10: -6,304 linesWeek of 2025-08-17: +1,838 linesWeek of 2025-08-17: -2,174 linesWeek of 2025-08-24: +446 linesWeek of 2025-08-24: -460 linesWeek of 2025-08-31: +10,268 linesWeek of 2025-08-31: -2,578 linesWeek of 2025-09-07: +13,088 linesWeek of 2025-09-07: -8,085 linesWeek of 2025-09-14: +8,912 linesWeek of 2025-09-14: -6,468 linesWeek of 2025-09-21: +1,232 linesWeek of 2025-09-21: -514 linesWeek of 2025-09-28: +2,363 linesWeek of 2025-09-28: -1,345 linesWeek of 2025-10-05: +282 linesWeek of 2025-10-05: -1,250 linesWeek of 2025-10-12: +79 linesWeek of 2025-10-12: -28 linesWeek of 2025-10-19: +422 linesWeek of 2025-10-19: -28 linesWeek of 2025-10-26: +87 linesWeek of 2025-10-26: -35 linesWeek of 2025-11-02: +123 linesWeek of 2025-11-02: -26 linesWeek of 2025-11-09: +299 linesWeek of 2025-11-09: -317 linesWeek of 2025-11-16: +571 linesWeek of 2025-11-16: -677 linesWeek of 2025-11-23: +60 linesWeek of 2025-11-23: -33 linesWeek of 2025-11-30: +109 linesWeek of 2025-11-30: -385 linesWeek of 2025-12-07: +547 linesWeek of 2025-12-07: -254 linesWeek of 2025-12-14: +14,961 linesWeek of 2025-12-14: -15,387 linesWeek of 2025-12-21: +999 linesWeek of 2025-12-21: -3,505 linesWeek of 2025-12-28: +120,749 linesWeek of 2025-12-28: -16,588 linesWeek of 2026-01-04: +11,642 linesWeek of 2026-01-04: -8,040 linesWeek of 2026-01-11: +650 linesWeek of 2026-01-11: -81 linesWeek of 2026-01-18: +3,645 linesWeek of 2026-01-18: -4,617 linesWeek of 2026-01-25: +234 linesWeek of 2026-01-25: -142 linesWeek of 2026-02-01: +15,829 linesWeek of 2026-02-01: -889 linesWeek of 2026-02-08: +140 linesWeek of 2026-02-08: -13 linesWeek of 2026-02-15: +574 linesWeek of 2026-02-15: -13 linesWeek of 2026-02-22: +2,955 linesWeek of 2026-02-22: -1,707 linesWeek of 2026-03-01: +463 linesWeek of 2026-03-01: -668 linesWeek of 2026-03-08: +3,707 linesWeek of 2026-03-08: -1,988 linesWeek of 2026-03-15: +4,035 linesWeek of 2026-03-15: -898 linesWeek of 2026-03-22: +4,820 linesWeek of 2026-03-22: -99,133 linesWeek of 2026-03-29: +1,798 linesWeek of 2026-03-29: -216 linesWeek of 2026-04-05: +375 linesWeek of 2026-04-05: -223 linesWeek of 2026-04-12: +9,994 linesWeek of 2026-04-12: -4,930 linesWeek of 2026-04-19: +1,358 linesWeek of 2026-04-19: -1,015 linesWeek of 2026-04-26: +6,045 linesWeek of 2026-04-26: -1,856 linesWeek of 2026-05-03: +154 linesWeek of 2026-05-03: -91 linesWeek of 2026-05-10: +1,830 linesWeek of 2026-05-10: -199 linesWeek of 2026-05-17: +1,921 linesWeek of 2026-05-17: -214 linesWeek of 2026-05-24: +432 linesWeek of 2026-05-24: -96 linesWeek of 2026-05-31: +443 linesWeek of 2026-05-31: -118 linesWeek of 2026-06-07: +2,396 linesWeek of 2026-06-07: -126 linesWeek of 2026-06-14: +1,899 linesWeek of 2026-06-14: -1,609 linesWeek of 2026-06-21: +49 linesWeek of 2026-06-21: -38 linesWeek of 2026-06-28: +243 linesWeek of 2026-06-28: -6 linesWeek of 2026-07-05: +113 linesWeek of 2026-07-05: -114 linesWeek of 2026-07-12: +830 linesWeek of 2026-07-12: -302 linesWeek of 2026-07-19: +6,007 linesWeek of 2026-07-19: -2,843 linesWeek of 2026-07-26: +297 linesWeek of 2026-07-26: -31 linesWeek of 2026-08-02: +22 linesWeek of 2026-08-02: -14 linesAug 10, 2025Aug 2, 2026
+271.4K lines added, -198.7K removed over the last year.

Commits per week

last 52 weeks
680Week of 2025-08-10: 3 commitsWeek of 2025-08-17: 3 commitsWeek of 2025-08-24: 10 commitsWeek of 2025-08-31: 33 commitsWeek of 2025-09-07: 22 commitsWeek of 2025-09-14: 21 commitsWeek of 2025-09-21: 14 commitsWeek of 2025-09-28: 17 commitsWeek of 2025-10-05: 9 commitsWeek of 2025-10-12: 4 commitsWeek of 2025-10-19: 6 commitsWeek of 2025-10-26: 7 commitsWeek of 2025-11-02: 4 commitsWeek of 2025-11-09: 9 commitsWeek of 2025-11-16: 6 commitsWeek of 2025-11-23: 5 commitsWeek of 2025-11-30: 4 commitsWeek of 2025-12-07: 6 commitsWeek of 2025-12-14: 46 commitsWeek of 2025-12-21: 13 commitsWeek of 2025-12-28: 68 commitsWeek of 2026-01-04: 26 commitsWeek of 2026-01-11: 12 commitsWeek of 2026-01-18: 14 commitsWeek of 2026-01-25: 4 commitsWeek of 2026-02-01: 18 commitsWeek of 2026-02-08: 7 commitsWeek of 2026-02-15: 9 commitsWeek of 2026-02-22: 25 commitsWeek of 2026-03-01: 14 commitsWeek of 2026-03-08: 39 commitsWeek of 2026-03-15: 38 commitsWeek of 2026-03-22: 44 commitsWeek of 2026-03-29: 30 commitsWeek of 2026-04-05: 15 commitsWeek of 2026-04-12: 65 commitsWeek of 2026-04-19: 10 commitsWeek of 2026-04-26: 30 commitsWeek of 2026-05-03: 8 commitsWeek of 2026-05-10: 23 commitsWeek of 2026-05-17: 14 commitsWeek of 2026-05-24: 5 commitsWeek of 2026-05-31: 4 commitsWeek of 2026-06-07: 11 commitsWeek of 2026-06-14: 21 commitsWeek of 2026-06-21: 7 commitsWeek of 2026-06-28: 2 commitsWeek of 2026-07-05: 4 commitsWeek of 2026-07-12: 6 commitsWeek of 2026-07-19: 24 commitsWeek of 2026-07-26: 3 commitsWeek of 2026-08-02: 2 commitsAug 10, 2025Aug 2, 2026
844 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 1 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 1 commitsSun 10:00 — 3 commitsSun 11:00 — 0 commitsSun 12:00 — 3 commitsSun 13:00 — 2 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 2 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 1 commitsSun 20:00 — 1 commitsSun 21:00 — 3 commitsSun 22:00 — 3 commitsSun 23:00 — 2 commitsMon 0:00 — 3 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 1 commitsMon 9:00 — 6 commitsMon 10:00 — 6 commitsMon 11:00 — 10 commitsMon 12:00 — 4 commitsMon 13:00 — 11 commitsMon 14:00 — 6 commitsMon 15:00 — 7 commitsMon 16:00 — 12 commitsMon 17:00 — 25 commitsMon 18:00 — 28 commitsMon 19:00 — 6 commitsMon 20:00 — 9 commitsMon 21:00 — 5 commitsMon 22:00 — 8 commitsMon 23:00 — 5 commitsTue 0:00 — 4 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 1 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 2 commitsTue 9:00 — 7 commitsTue 10:00 — 7 commitsTue 11:00 — 25 commitsTue 12:00 — 4 commitsTue 13:00 — 24 commitsTue 14:00 — 19 commitsTue 15:00 — 17 commitsTue 16:00 — 5 commitsTue 17:00 — 26 commitsTue 18:00 — 15 commitsTue 19:00 — 2 commitsTue 20:00 — 7 commitsTue 21:00 — 7 commitsTue 22:00 — 4 commitsTue 23:00 — 0 commitsWed 0:00 — 2 commitsWed 1:00 — 6 commitsWed 2:00 — 3 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 2 commitsWed 9:00 — 8 commitsWed 10:00 — 11 commitsWed 11:00 — 8 commitsWed 12:00 — 7 commitsWed 13:00 — 30 commitsWed 14:00 — 12 commitsWed 15:00 — 4 commitsWed 16:00 — 10 commitsWed 17:00 — 12 commitsWed 18:00 — 11 commitsWed 19:00 — 3 commitsWed 20:00 — 1 commitsWed 21:00 — 2 commitsWed 22:00 — 2 commitsWed 23:00 — 3 commitsThu 0:00 — 1 commitsThu 1:00 — 1 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 9 commitsThu 7:00 — 0 commitsThu 8:00 — 1 commitsThu 9:00 — 2 commitsThu 10:00 — 4 commitsThu 11:00 — 14 commitsThu 12:00 — 7 commitsThu 13:00 — 11 commitsThu 14:00 — 17 commitsThu 15:00 — 9 commitsThu 16:00 — 17 commitsThu 17:00 — 14 commitsThu 18:00 — 10 commitsThu 19:00 — 14 commitsThu 20:00 — 7 commitsThu 21:00 — 1 commitsThu 22:00 — 7 commitsThu 23:00 — 3 commitsFri 0:00 — 4 commitsFri 1:00 — 1 commitsFri 2:00 — 0 commitsFri 3:00 — 1 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 1 commitsFri 9:00 — 4 commitsFri 10:00 — 18 commitsFri 11:00 — 14 commitsFri 12:00 — 13 commitsFri 13:00 — 12 commitsFri 14:00 — 12 commitsFri 15:00 — 10 commitsFri 16:00 — 20 commitsFri 17:00 — 12 commitsFri 18:00 — 7 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 8 commitsFri 22:00 — 4 commitsFri 23:00 — 9 commitsSat 0:00 — 2 commitsSat 1:00 — 3 commitsSat 2:00 — 3 commitsSat 3:00 — 2 commitsSat 4:00 — 2 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 2 commitsSat 9:00 — 5 commitsSat 10:00 — 5 commitsSat 11:00 — 1 commitsSat 12:00 — 0 commitsSat 13:00 — 3 commitsSat 14:00 — 3 commitsSat 15:00 — 2 commitsSat 16:00 — 3 commitsSat 17:00 — 3 commitsSat 18:00 — 2 commitsSat 19:00 — 4 commitsSat 20:00 — 8 commitsSat 21:00 — 0 commitsSat 22:00 — 6 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Apr 11, 2026daily#22+239
Apr 8, 2026daily#24+138
Mar 21, 2026daily#15+223
Mar 20, 2026daily#7+300
Mar 19, 2026daily#14+178
Mar 18, 2026daily#17+189
  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    385.5K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    268.6K stars · Shell

  • NousResearch/hermes-agent

    The agent that grows with you

    227K stars · Python