Trending repositories: pdf-parsing
3 tracked repositories tagged with pdf-parsing, ordered by stars. Use the topic filters below to narrow further.
3 of 3 repositories
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.
28,250dataJavaApache-2.0firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
AI summary: A fast Rust library for PDF classification, text extraction, and Markdown conversion without relying on expensive OCR.
12,958developer-toolsRustMITxberg-io/xberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.
AI summary: High-performance polyglot document extraction framework built in Rust.
8,910dataRustMIT