Trending repositories: pdf-parsing

3 tracked repositories tagged with pdf-parsing, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

3 of 3 repositories

  • opendataloader-project/opendataloader-pdf

    PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

    AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.

    28,250dataJavaApache-2.0
  • firecrawl/pdf-inspector

    Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

    AI summary: A fast Rust library for PDF classification, text extraction, and Markdown conversion without relying on expensive OCR.

    12,958developer-toolsRustMIT
  • xberg-io/xberg

    A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.

    AI summary: High-performance polyglot document extraction framework built in Rust.

    8,910dataRustMIT