Trending repositories: text-extraction

4 tracked repositories tagged with text-extraction, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

4 of 4 repositories

  • firecrawl/pdf-inspector

    Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

    AI summary: A fast Rust library for PDF classification, text extraction, and Markdown conversion without relying on expensive OCR.

    12,943developer-toolsRustMIT
  • run-llama/liteparse

    A fast, helpful, and open-source document parser

    AI summary: A fast, lightweight, and open-source Rust tool for high-quality spatial PDF parsing.

    11,931dataRustApache-2.0
  • xberg-io/xberg

    A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.

    AI summary: High-performance polyglot document extraction framework built in Rust.

    8,910dataRustMIT
  • xberg-io/xberg

    A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 98+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.

    AI summary: A unified document intelligence engine for extracting structured data from 98 file formats.

    8,712dataRustMIT