firecrawl/pdf-inspectorPublic

Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

AI summary: A high-performance Rust library for intelligent PDF classification, layout-aware text extraction, and clean Markdown conversion.

Stars
19.5K
+51 today
Forks
1.3K
Watchers
59
Open issues
105
Open PRs
122
Contributors
~12
Commits
553
Branches
239

RustMITCreated Feb 6, 2026Last push 3d agoLatest release v1.25.2+153 stars this week+977 this month

Quick answers

What is pdf-inspector?
A high-performance Rust library for intelligent PDF classification, layout-aware text extraction, and clean Markdown conversion.
What does pdf-inspector do?
PDF-inspector is an ultra-fast Rust library developed by Firecrawl to intelligently parse and extract data from PDF documents. It drastically reduces processing costs by quickly classifying PDFs as either text-based or scanned, completely bypassing expensive OCR services for the majority of files that don't need them. The library performs position-aware extraction that understands reading order, multi-column layouts, and complex tables, converting the output into clean, structured Markdown. It offers broad accessibility through bindings for Python, Node.js, and WebAssembly, making it a versatile tool for data ingestion pipelines.
Who is pdf-inspector for?
Data engineers, AI developers, and backend teams building document processing pipelines or RAG systems that require fast, accurate PDF ingestion.
How do I get started with pdf-inspector?
See docs for Rust, Python, Node.js or WASM bindings
How popular is pdf-inspector on GitHub?
firecrawl/pdf-inspector has 19,462 stars and 1,318 forks on GitHub, and gained 153 stars in the last 7 days.
What license does pdf-inspector use?
firecrawl/pdf-inspector is released under the MIT license.

Star history

since Aug 4, 2026
05K10K15KAug 2026Aug 2026Sep 2026Oct 2026
19.5K stars as of Oct 2, 2026. Measured daily since Aug 4, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 6 commits2026-02-07: 16 commits2026-02-08: 2 commits2026-02-09: 5 commits2026-02-10: 1 commit2026-02-11: 7 commits2026-02-12: 5 commits2026-02-13: 8 commits2026-02-14: 5 commits2026-02-15: 5 commits2026-02-16: 12 commits2026-02-17: 13 commits2026-02-18: 11 commits2026-02-19: 2 commits2026-02-20: 9 commits2026-02-21: 2 commits2026-02-22: 1 commit2026-02-23: 6 commits2026-02-24: 1 commit2026-02-25: 5 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 1 commit2026-03-03: 4 commits2026-03-04: 6 commits2026-03-05: 6 commits2026-03-06: 1 commit2026-03-07: 2 commits2026-03-08: 0 commits2026-03-09: 6 commits2026-03-10: 0 commits2026-03-11: 5 commits2026-03-12: 5 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 4 commits2026-03-18: 8 commits2026-03-19: 5 commits2026-03-20: 5 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 7 commits2026-03-24: 2 commits2026-03-25: 7 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 30 commits2026-04-03: 1 commit2026-04-04: 21 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 5 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 2 commits2026-04-12: 0 commits2026-04-13: 5 commits2026-04-14: 7 commits2026-04-15: 2 commits2026-04-16: 1 commit2026-04-17: 6 commits2026-04-18: 0 commits2026-04-19: 1 commit2026-04-20: 8 commits2026-04-21: 3 commits2026-04-22: 2 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 3 commits2026-04-27: 6 commits2026-04-28: 2 commits2026-04-29: 3 commits2026-04-30: 1 commit2026-05-01: 0 commits2026-05-02: 2 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 1 commit2026-05-07: 2 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 5 commits2026-05-13: 0 commits2026-05-14: 6 commits2026-05-15: 1 commit2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 1 commit2026-05-19: 0 commits2026-05-20: 1 commit2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 1 commit2026-05-28: 1 commit2026-05-29: 1 commit2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 2 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 3 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 1 commit2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 2 commits2026-06-24: 3 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 2 commits2026-07-09: 1 commit2026-07-10: 9 commits2026-07-11: 8 commits2026-07-12: 0 commits2026-07-13: 2 commits2026-07-14: 16 commits2026-07-15: 12 commits2026-07-16: 8 commits2026-07-17: 2 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 6 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 4 commits2026-08-04: 4 commits2026-08-05: 6 commits2026-08-06: 2 commits2026-08-07: 2 commits2026-08-08: 1 commit2026-08-09: 3 commits2026-08-10: 5 commits2026-08-11: 4 commits2026-08-12: 5 commits2026-08-13: 5 commits2026-08-14: 2 commits2026-08-15: 0 commits2026-08-16: 14 commits2026-08-17: 8 commits2026-08-18: 4 commits2026-08-19: 1 commit2026-08-20: 4 commits2026-08-21: 6 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 1 commit2026-09-02: 4 commits2026-09-03: 1 commit2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 3 commits2026-09-08: 3 commits2026-09-09: 4 commits2026-09-10: 3 commits2026-09-11: 2 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 4 commits2026-09-15: 0 commits2026-09-16: 1 commit2026-09-17: 2 commits2026-09-18: 1 commit2026-09-19: 0 commits2026-09-20: 10 commits2026-09-21: 10 commits2026-09-22: 3 commits2026-09-23: 0 commits2026-09-24: 1 commit2026-09-25: 2 commits2026-09-26: 0 commits2026-09-27: 7 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits
546 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    19,462 stars

  • Very active

    546 commits in 52 weeks

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    17 trending appearances

What pdf-inspector does

PDF-inspector is an ultra-fast Rust library developed by Firecrawl to intelligently parse and extract data from PDF documents. It drastically reduces processing costs by quickly classifying PDFs as either text-based or scanned, completely bypassing expensive OCR services for the majority of files that don't need them. The library performs position-aware extraction that understands reading order, multi-column layouts, and complex tables, converting the output into clean, structured Markdown. It offers broad accessibility through bindings for Python, Node.js, and WebAssembly, making it a versatile tool for data ingestion pipelines.

Data engineers, AI developers, and backend teams building document processing pipelines or RAG systems that require fast, accurate PDF ingestion.

  • Smart classification: Analyzes content streams in milliseconds to intelligently detect and route text-based versus scanned documents.
  • Layout-aware extraction: accurately parses multi-column layouts, X/Y coordinates, and reading order without relying on OCR.
  • High-fidelity Markdown: Converts extracted data into rich Markdown, correctly identifying headings, lists, and code blocks via font heuristics.
  • Advanced table detection: Combines rectangle-based drawing ops and text alignment heuristics to accurately reconstruct complex data tables.
  • Cross-platform bindings: Provides native performance across different ecosystems with integrations for Python, Node.js, and browser WebAssembly.

Where teams use it

RAG pipeline ingestion

Rapidly converting thousands of corporate PDF reports into clean Markdown to feed into a vector database for LLM retrieval.

Intelligent document routing

Implementing a preprocessing step that sends only scanned documents to expensive cloud OCR APIs, saving significant costs.

Financial data extraction

Accurately parsing complex, multi-page financial tables from annual reports into structured data formats.

In-browser PDF processing

Using the WebAssembly bindings to securely extract text from sensitive user documents entirely client-side.

Getting started: See docs for Rust, Python, Node.js or WASM bindings

README

main branch

pdf-inspector

Crates.io npm PyPI License: MIT

Fast Rust library for PDF classification and text extraction. By default it detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown without OCR. Native Rust and CLI consumers can opt into selective OCR. Includes bindings for Python, Node.js, and browser WebAssembly.

Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.

Features

  • Smart classification — Detect TextBased, Scanned, ImageBased, or Mixed PDFs in ~10-50ms by sampling content streams. Returns a confidence score (0.0-1.0) and per-page OCR routing.
  • Text extraction — Position-aware extraction with font info, X/Y coordinates, and automatic multi-column reading order. Rotated runs (margin stamps, chart axis titles) keep a true axis-aligned box and report their rotation angle instead of collapsing to zero width.
  • Markdown conversion — Headings (H1-H4 via font size ratios), bullet/numbered/letter lists, code blocks (monospace font detection), tables (rectangle-based and heuristic), bold/italic formatting, URL linking, and page breaks.
  • Table detection — Dual-mode: rectangle-based detection from PDF drawing ops, plus heuristic detection from text alignment. Handles financial tables, footnotes, and continuation tables across pages.
  • CID font support — ToUnicode CMap decoding for Type0/Identity-H fonts, UTF-16BE, UTF-8, and Latin-1 encodings.
  • Multi-column layout — Automatic detection of newspaper-style columns, sequential reading order, and RTL text support.
  • Encoding issue detection — Automatically flags broken font encodings so callers can fall back to OCR.
  • Selective OCR — Rust, CLI, Python, and Node can render only pages that need OCR, run PP-OCRv6 Small locally, and preserve per-page provenance and hosted-fallback recommendations.
  • Single document load — The document is parsed once and shared between detection and extraction, avoiding redundant I/O.
  • Browser WebAssembly — Run the same Rust parser locally in browsers and Web Workers, with embedded CMaps and no server round trip.
  • Lightweight by default — The default Rust and browser builds remain pure extraction. Native Python and Node packages include the OCR integration, but PDFium, ONNX Runtime, and model files remain external and are touched only when a page is routed to OCR.

Benchmark

Evaluated on the opendataloader-bench corpus (200 PDFs). Only local engines without model-based PDF parsing are shown; OCR was disabled. Scores are 0-1, higher is better.

Engine Overall Reading Order (NID) Tables (TEDS) Headings (MHS) Speed (200 docs)
pdf-inspector 0.875 0.915 0.814 0.788 0.470s
liteparse 0.873 0.913 0.693 0.811 0.750s
opendataloader 0.831 0.902 0.489 0.739 2.569s
pymupdf4llm 0.735 0.886 0.401 0.424 17.117s
markitdown 0.589 0.844 0.273 0.000 16.165s

Results were refreshed on July 31, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up run, with each parser processing documents sequentially in a single process.

The complete parser configuration, per-document predictions, evaluator output, and generated charts are available in the reproducible results branch.

Best fit: Native-text PDFs where speed, reading order, and table structure matter. In this comparison, pdf-inspector delivered the higher overall, reading-order, and table scores, along with the fastest complete run. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.

Use the paired benchmark harness to compare two local builds against the exact same corpus and evaluator revision.

Quick start

Python

pip install pdf-inspector
import pdf_inspector

result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type)   # "text_based", "scanned", "image_based", "mixed"
print(result.markdown)   # Markdown string or None

# Selective OCR; clean text PDFs do not load the external OCR runtime.
ocr = pdf_inspector.process_pdf_with_ocr("document.pdf")
print(ocr.pages_routed_to_ocr)

Full API reference: docs/python.md

Node.js

npm install @firecrawl/pdf-inspector
import { readFileSync } from 'fs';
import { processPdf, processPdfWithOcr } from '@firecrawl/pdf-inspector';

const pdf = readFileSync('document.pdf');
const result = processPdf(pdf);
console.log(result.pdfType);   // "TextBased", "Scanned", "ImageBased", "Mixed"
console.log(result.markdown);  // Markdown string or null

const ocr = await processPdfWithOcr(pdf); // selective OCR, off the event loop
console.log(ocr.pagesRoutedToOcr);

Full API reference: napi/README.md

Browser WebAssembly

npm install @firecrawl/pdf-inspector-wasm
import init, { processPdf } from '@firecrawl/pdf-inspector-wasm';

await init();
const response = await fetch('/document.pdf');
const pdf = new Uint8Array(await response.arrayBuffer());
const result = processPdf(pdf);

console.log(result.pdfType);
console.log(result.markdown);

Full API reference: wasm/README.md

Rust

Install from crates.io:

cargo add pdf-inspector

Or add it manually:

[dependencies]
pdf-inspector = "1"
use pdf_inspector::process_pdf;

let result = process_pdf("document.pdf")?;
println!("Type: {:?}", result.pdf_type);
if let Some(markdown) = &result.markdown {
    println!("{}", markdown);
}

Full API reference: docs/rust-api.md

CLI

# Install the CLI tools
cargo install pdf-inspector

# Convert PDF to Markdown
pdf2md document.pdf

# JSON output (for piping), with the document information (title, author, producer, dates, ...)
pdf2md document.pdf --json

# Positioned TextItem JSON (coordinates relative to the visible page box): axis-aligned box, rotation, font, paint (fill and stroke colour, render mode), underline metadata
pdf2md document.pdf --items-json

# Raw markdown only (no headers)
pdf2md document.pdf --raw

# Token-efficient output (collapses long dot leaders and similar source padding)
pdf2md document.pdf --compact

# Insert page break markers (<!-- Page N -->)
pdf2md document.pdf --pages

# Process only specific pages
pdf2md document.pdf --select-pages 1,3,5-10

# Detection only (no extraction)
detect-pdf document.pdf
detect-pdf document.pdf --json

# Detection + layout analysis (tables, columns)
detect-pdf document.pdf --analyze --json

Rust and CLI consumers opt into OCR at build time:

cargo install pdf-inspector --features ocr --bin pdf2md
PDFIUM_LIB_PATH=/path/to/libpdfium ORT_DYLIB_PATH=/path/to/libonnxruntime \
  pdf2md scan.pdf --ocr auto --json

The OCR JSON envelope is versioned and reports routed pages, per-page source and confidence, warnings, and pages recommended for the hosted document pipeline. Native Python and Node packages expose the same pipeline without a source-build feature. All native entry points still require separately installed PDFium and ONNX Runtime libraries only when OCR is routed. See the OCR runtime setup guide for pinned downloads, platform support, model-cache behavior, and hosted-fallback integration. See the Rust API guide for lower-level controls.

From a source checkout, use cargo run --bin pdf2md -- document.pdf or cargo run --bin detect-pdf -- document.pdf instead.

Architecture

PDF bytes
  │
  ├─► detector         → PdfType (TextBased / Scanned / ImageBased / Mixed)
  │
  └─► extractor
        ├─ fonts        → font widths, encodings
        ├─ content_stream → walk PDF operators → TextItems + PdfRects
        ├─ xobjects     → Form XObject text, image placeholders
        ├─ links        → hyperlinks, AcroForm fields
        └─ layout       → column detection → line grouping → reading order
              │
              ├─► tables
              │     ├─ detect_rects      → rectangle-based tables (union-find)
              │     ├─ detect_heuristic  → alignment-based tables
              │     ├─ grid              → column/row assignment → cells
              │     └─ format            → cells → Markdown table
              │
              └─► markdown
                    ├─ analysis     → font stats, heading tiers
                    ├─ preprocess   → merge headings, drop caps
                    ├─ convert      → line loop + table/image insertion
                    ├─ classify     → captions, lists, code
                    └─ postprocess  → cleanup → final Markdown

The document is loaded once via load_document_from_path / load_document_from_mem and shared between the detection and extraction stages, so there's no redundant parsing.

Project structure

src/
  lib.rs                — Public API, PdfOptions builder, convenience functions
  python.rs             — PyO3 Python bindings
  types.rs              — Shared types: TextItem, TextLine, PdfRect, ItemType
  text_utils.rs         — Character/text helpers (CJK, RTL, ligatures, bold/italic)
  process_mode.rs       — ProcessMode enum (DetectOnly, Analyze, Full)
  detector.rs           — Fast PDF type detection without full document load
  glyph_names.rs        — Adobe Glyph List → Unicode mapping
  tounicode.rs          — ToUnicode CMap parsing for CID-encoded text
  extractor/            — Text extraction pipeline
  tables/               — Table detection and formatting
  markdown/             — Markdown conversion and structure detection
  bin/                  — CLI tools (pdf2md, detect_pdf)
napi/                   — Node.js/Bun bindings (napi-rs)
wasm/                   — Browser bindings (wasm-bindgen)

How classification works

  1. Parse the xref table and page tree (no full object load)
  2. Select pages based on ScanStrategy (default: a sample of 8 pages)
  3. Look for Tj/TJ (text operators) and Do (image operators) in content streams
  4. Classify based on text operator presence across sampled pages

This detects 300+ page PDFs in milliseconds. The result includes pages_needing_ocr — a list of specific page numbers that lack text, enabling per-page OCR routing instead of all-or-nothing.

Scan strategies

Strategy Behavior Best for
EarlyExit Scan all pages, stop on first non-text page Pipelines routing TextBased PDFs to fast extraction
Full Scan all pages, no early exit Accurate Mixed vs Scanned classification
Sample(n) (default: Sample(8)) Sample n evenly distributed pages (first, last, middle) Very large PDFs where speed matters more than precision
Pages(vec) Only scan specific 1-indexed page numbers When the caller knows which pages to check

Markdown output

The converter handles:

Element How it's detected
Headings (H1-H4) Font size tiers relative to body text, with 0.5pt clustering
Bold/italic Font name patterns (Bold, Italic, Oblique)
Bullet lists •, -, *, ○, ●, ◦ prefixes
Numbered lists 1., 1), (1) patterns
Letter lists a., a), (a) patterns
Code blocks Monospace fonts (Courier, Consolas, Monaco, Menlo, Fira Code, JetBrains Mono) and keyword detection
Tables Rectangle-based detection from PDF drawing ops + heuristic detection from text alignment
Financial tables Token splitting for consolidated numeric values
Captions "Figure", "Table", "Source:" prefix detection
Sub/superscript Font size and Y-offset relative to baseline
URLs Converted to Markdown links
Hyphenation Rejoins words broken across lines
Page numbers Filtered from output
Drop caps Large initial letters merged with following text
Dot leaders TOC-style dots collapsed to " ... "

Use case: smart PDF routing

pdf-inspector was built for pipelines that process PDFs at scale. Instead of sending every PDF through OCR:

PDF arrives
  → pdf-inspector classifies it (~20ms)
  → TextBased + high confidence?
      YES → extract locally (~150ms), done
      NO  → send to OCR service (2-10s)

This saves cost and latency for the majority of PDFs that are already text-based (reports, papers, invoices, legal docs).

Debugging

See docs/debugging.md for RUST_LOG environment variable usage.

License

MIT

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

14 total
  1. Changes since 1.25.1: - Size a super- or subscript run by its letters and digits: an exponent whose minus sign a symbol font sets a design size above its digits (as TeX does) keeps its `<sup>` markup; a run of signs alone, or with signs more than a quarter larger than its letters and digits, is still sized by its largest glyph (#603). - Skip a subset font's ToUnicode repairs where the font itself shows the CMap is right: the repair through a `/CIDToGIDMap` for a CMap keyed by code, and the renumbering of a subset that kept its glyph indexes (#605). No API changes. Release preparation: #606. [Full changelog](https://github.com/firecrawl/pdf-inspector/compare/v1.25.1...v1.25.2)

  2. Changes since 1.25.0: - Read simple fonts (Type1, TrueType, Type3) one byte per code when their ToUnicode CMap declares a two-byte codespace: text shown as a kerned run of short strings no longer loses every two-byte string (`Income Statement` read as `Iometatent`). The stale-CMap check for Type1 fonts judges such a CMap too, and the page detector counts such a font's text byte by byte (#595). - Read a Type1 font whose `/Encoding` names no base encoding through the built-in encoding of its embedded program (PDF 32000-1:2008, Table 114): TeX's fonts read their ligatures (`efficiency`, was `eciency`), curly quotes, dashes and math symbols. A code the program leaves at `.notdef` reads as nothing, the word space excepted, and the base-14 width fallback measures such fonts as they read (#596). - Take a repaired ToUnicode CMap for a string of one or two glyphs only on evidence beyond short common words: `engagement` no longer reads as `engageaent`, `AND` as `ANa` or `Three` as `ThaTe` (#597). No API changes. Release preparation: #604. [Full changelog](https://github.com/firecrawl/pdf-inspector/compare/v1.25.0...v1.25.1)

  3. Changes since 1.24.0: - Node: `extractTextWithPositionsAsync`, the async variant of `extractTextWithPositions`: the same arguments and items, run on the libuv thread pool so a slow page no longer blocks the caller's event loop (#592). - Underline and strikeout detection no longer takes quadratic time on pages drawn from many thin filled rects or short strokes: a page of about 200,000 thin rects went from about 40 s to under 0.5 s, with the same marks (#592). - The `linux-x64-gnu` Node binary is cross-built for a GLIBC_2.17 floor again, so it loads on Amazon Linux 2023 (the managed AWS Lambda Node runtimes), RHEL 9 and Debian 11 (#586). API: Node `extractTextWithPositionsAsync`. Release preparation: #593. [Full changelog](https://github.com/firecrawl/pdf-inspector/compare/v1.24.0...v1.25.0)

  4. Changes since 1.23.0: - Report the paint each text run was shown with: `TextItem` gains `fill_color` and `stroke_color` (8-bit sRGB, read from DeviceRGB, DeviceGray, DeviceCMYK, ICCBased by component count and Indexed palettes; `None` for Separation, DeviceN, Pattern and CIE-based spaces) and `render_mode` (the `Tr` mode, so invisible text and clipping-only text can be told apart). Extraction and markdown are unchanged (#579). - Read the document information dictionary's `/Author`, `/Subject`, `/Keywords`, `/Creator`, `/Producer`, `/CreationDate` and `/ModDate` beside `/Title`, each decoded as a PDF text string (#579). - Decode the document `/Title` as a PDF text string: titles in PDFDocEncoding no longer read with U+FFFD, titles in UTF-16LE no longer read as mojibake, and titles given by reference are read (#579). - Extract text a page shows with the `"` operator, which the page parser skipped together with the spacing and line move it makes (#579). - Take an ActualText span's weight from the paint of its first painted glyph rather than from the paint in force at the span's end (#579). - Keep two lines of small type apart when their baselines lie less than 5 pt apart: a stacked t

  5. Changes since 1.22.1: - Read a Form XObject whose `/BBox` numerals are too large for any parser: such numerals are saturated in place before the document is read, so the form's text is extracted and the OCR pipeline renders it instead of an empty page (#560). - Read right-to-left runs shown in reading order forwards: a visible run of right-to-left letters painted forwards now votes for visual storage, so pages that came out word-mirrored read correctly; the logical-order reading is kept for invisible text layers (#561). - Read a re-encoded simple font by its `/Differences` names when it kept a stale ToUnicode CMap, and keep one glyph's characters together through the visual-order read-back (#562). - Compose a detached spacing accent with the letter it is painted over instead of leaving a stray accent a word on (#563). - Treat a return from a zero-advance dependent sign placed behind the pen as no word gap, so scripts that attach signs by stepping the pen back keep their words whole (#564). - Classify a page whose text layer is drawn invisibly under a covering image as a scan with the OCR reason `invisible_text_layer`, following the text rendering mode, the transformation, the grap

Code frequency

additions and deletions
+38.7K-38.7KWeek of 2026-02-01: +4,413 linesWeek of 2026-02-01: -369 linesWeek of 2026-02-08: +6,282 linesWeek of 2026-02-08: -985 linesWeek of 2026-02-15: +38,676 linesWeek of 2026-02-15: -10,724 linesWeek of 2026-02-22: +1,251 linesWeek of 2026-02-22: -411 linesWeek of 2026-03-01: +2,278 linesWeek of 2026-03-01: -511 linesWeek of 2026-03-08: +5,829 linesWeek of 2026-03-08: -962 linesWeek of 2026-03-15: +3,826 linesWeek of 2026-03-15: -290 linesWeek of 2026-03-22: +3,599 linesWeek of 2026-03-22: -668 linesWeek of 2026-03-29: +6,109 linesWeek of 2026-03-29: -2,081 linesWeek of 2026-04-05: +734 linesWeek of 2026-04-05: -59 linesWeek of 2026-04-12: +3,944 linesWeek of 2026-04-12: -295 linesWeek of 2026-04-19: +2,247 linesWeek of 2026-04-19: -238 linesWeek of 2026-04-26: +5,904 linesWeek of 2026-04-26: -357 linesWeek of 2026-05-03: +689 linesWeek of 2026-05-03: -32 linesWeek of 2026-05-10: +1,817 linesWeek of 2026-05-10: -169 linesWeek of 2026-05-17: +495 linesWeek of 2026-05-17: -33 linesWeek of 2026-05-24: +1,033 linesWeek of 2026-05-24: -32 linesWeek of 2026-05-31: +939 linesWeek of 2026-05-31: -81 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +148 linesWeek of 2026-06-14: -14 linesWeek of 2026-06-21: +1,403 linesWeek of 2026-06-21: -139 linesWeek of 2026-06-28: +0 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +4,698 linesWeek of 2026-07-05: -652 linesWeek of 2026-07-12: +12,610 linesWeek of 2026-07-12: -855 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +200 linesWeek of 2026-07-26: -55 linesWeek of 2026-08-02: +7,379 linesWeek of 2026-08-02: -350 linesWeek of 2026-08-09: +5,582 linesWeek of 2026-08-09: -407 linesWeek of 2026-08-16: +18,132 linesWeek of 2026-08-16: -1,708 linesWeek of 2026-08-23: +0 linesWeek of 2026-08-23: -0 linesWeek of 2026-08-30: +8,578 linesWeek of 2026-08-30: -1,457 linesWeek of 2026-09-06: +7,940 linesWeek of 2026-09-06: -319 linesWeek of 2026-09-13: +5,249 linesWeek of 2026-09-13: -480 linesWeek of 2026-09-20: +32,463 linesWeek of 2026-09-20: -2,591 linesWeek of 2026-09-27: +2,563 linesWeek of 2026-09-27: -147 linesFeb 1, 2026Sep 27, 2026
+197K lines added, -27.5K removed over the last year.

Commits per week

last 52 weeks
540Week of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 22 commitsWeek of 2026-02-08: 33 commitsWeek of 2026-02-15: 54 commitsWeek of 2026-02-22: 13 commitsWeek of 2026-03-01: 20 commitsWeek of 2026-03-08: 16 commitsWeek of 2026-03-15: 22 commitsWeek of 2026-03-22: 16 commitsWeek of 2026-03-29: 52 commitsWeek of 2026-04-05: 7 commitsWeek of 2026-04-12: 21 commitsWeek of 2026-04-19: 14 commitsWeek of 2026-04-26: 17 commitsWeek of 2026-05-03: 3 commitsWeek of 2026-05-10: 12 commitsWeek of 2026-05-17: 2 commitsWeek of 2026-05-24: 3 commitsWeek of 2026-05-31: 5 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 1 commitsWeek of 2026-06-21: 5 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 20 commitsWeek of 2026-07-12: 40 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 6 commitsWeek of 2026-08-02: 19 commitsWeek of 2026-08-09: 24 commitsWeek of 2026-08-16: 37 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 6 commitsWeek of 2026-09-06: 15 commitsWeek of 2026-09-13: 8 commitsWeek of 2026-09-20: 26 commitsWeek of 2026-09-27: 7 commitsOct 4, 2025Sep 27, 2026
546 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 2 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 3 commitsSun 9:00 — 8 commitsSun 10:00 — 1 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 2 commitsSun 16:00 — 3 commitsSun 17:00 — 2 commitsSun 18:00 — 0 commitsSun 19:00 — 1 commitsSun 20:00 — 0 commitsSun 21:00 — 5 commitsSun 22:00 — 2 commitsSun 23:00 — 17 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 2 commitsMon 8:00 — 7 commitsMon 9:00 — 11 commitsMon 10:00 — 12 commitsMon 11:00 — 4 commitsMon 12:00 — 7 commitsMon 13:00 — 5 commitsMon 14:00 — 8 commitsMon 15:00 — 3 commitsMon 16:00 — 4 commitsMon 17:00 — 7 commitsMon 18:00 — 6 commitsMon 19:00 — 3 commitsMon 20:00 — 5 commitsMon 21:00 — 4 commitsMon 22:00 — 7 commitsMon 23:00 — 0 commitsTue 0:00 — 2 commitsTue 1:00 — 1 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 1 commitsTue 8:00 — 3 commitsTue 9:00 — 4 commitsTue 10:00 — 4 commitsTue 11:00 — 6 commitsTue 12:00 — 5 commitsTue 13:00 — 6 commitsTue 14:00 — 4 commitsTue 15:00 — 2 commitsTue 16:00 — 4 commitsTue 17:00 — 7 commitsTue 18:00 — 9 commitsTue 19:00 — 5 commitsTue 20:00 — 4 commitsTue 21:00 — 8 commitsTue 22:00 — 4 commitsTue 23:00 — 0 commitsWed 0:00 — 1 commitsWed 1:00 — 6 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 3 commitsWed 9:00 — 4 commitsWed 10:00 — 13 commitsWed 11:00 — 9 commitsWed 12:00 — 6 commitsWed 13:00 — 14 commitsWed 14:00 — 7 commitsWed 15:00 — 9 commitsWed 16:00 — 7 commitsWed 17:00 — 7 commitsWed 18:00 — 3 commitsWed 19:00 — 0 commitsWed 20:00 — 1 commitsWed 21:00 — 3 commitsWed 22:00 — 3 commitsWed 23:00 — 5 commitsThu 0:00 — 0 commitsThu 1:00 — 2 commitsThu 2:00 — 1 commitsThu 3:00 — 1 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 2 commitsThu 9:00 — 3 commitsThu 10:00 — 9 commitsThu 11:00 — 11 commitsThu 12:00 — 16 commitsThu 13:00 — 3 commitsThu 14:00 — 10 commitsThu 15:00 — 11 commitsThu 16:00 — 3 commitsThu 17:00 — 10 commitsThu 18:00 — 3 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 4 commitsThu 23:00 — 2 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 1 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 1 commitsFri 8:00 — 3 commitsFri 9:00 — 1 commitsFri 10:00 — 3 commitsFri 11:00 — 7 commitsFri 12:00 — 3 commitsFri 13:00 — 10 commitsFri 14:00 — 4 commitsFri 15:00 — 4 commitsFri 16:00 — 6 commitsFri 17:00 — 4 commitsFri 18:00 — 7 commitsFri 19:00 — 5 commitsFri 20:00 — 1 commitsFri 21:00 — 6 commitsFri 22:00 — 6 commitsFri 23:00 — 1 commitsSat 0:00 — 6 commitsSat 1:00 — 9 commitsSat 2:00 — 4 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 1 commitsSat 11:00 — 4 commitsSat 12:00 — 3 commitsSat 13:00 — 11 commitsSat 14:00 — 1 commitsSat 15:00 — 1 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 2 commitsSat 19:00 — 6 commitsSat 20:00 — 3 commitsSat 21:00 — 3 commitsSat 22:00 — 2 commitsSat 23:00 — 4 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Sep 3, 2026daily#7+541
Sep 2, 2026daily#7+541
Sep 1, 2026daily#8+228
Aug 31, 2026monthly#2+15,332
Aug 31, 2026daily#13+199
Aug 30, 2026monthly#2+15,332
Aug 15, 2026weekly#8+3,251
Aug 14, 2026weekly#8+3,251
Aug 13, 2026weekly#4+4,043
Aug 12, 2026weekly#3+5,367
Aug 11, 2026weekly#2+7,143
Aug 10, 2026weekly#1+8,641
Aug 8, 2026daily#9+1,190
Aug 7, 2026daily#9+1,190
Aug 6, 2026daily#5+1,582
  • freeCodeCamp/freeCodeCamp

    freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

    456.7K stars · TypeScript

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    373.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    285.8K stars · Python

  • tensorflow/tensorflow

    An Open Source Machine Learning Framework for Everyone

    200.7K stars · C++

  • yt-dlp/yt-dlp

    A feature-rich command-line audio/video downloader

    195.5K stars · Python

  • ultraworkers/claw-code

    An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

    195.2K stars · Rust