Trending repositories: pdf-extraction
5 tracked repositories tagged with pdf-extraction, ordered by stars. Use the topic filters below to narrow further.
5 of 5 repositories
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.
28,188dataJavaApache-2.0firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
AI summary: A fast Rust library for PDF classification, text extraction, and Markdown conversion without relying on expensive OCR.
12,943developer-toolsRustMITxberg-io/xberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.
AI summary: High-performance polyglot document extraction framework built in Rust.
8,910dataRustMITxberg-io/xberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 98+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.
AI summary: A unified document intelligence engine for extracting structured data from 98 file formats.
8,712dataRustMITdwzhu-pku/PaperBanana
PaperBanana: Automating Academic Illustration For AI Scientists
AI summary: A streamlined tool for extracting, organizing, and summarizing academic research papers.
6,896productivityPythonApache-2.0