Trending repositories: pdf-parser

4 tracked repositories tagged with pdf-parser, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

4 of 4 repositories

  • opendatalab/MinerU

    Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

    AI summary: High-accuracy parsing engine that converts complex documents into LLM-ready Markdown and JSON.

    77,081dataPythonOther
  • opendataloader-project/opendataloader-pdf

    PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

    AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.

    28,188dataJavaApache-2.0
  • firecrawl/pdf-inspector

    Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

    AI summary: A fast Rust library for PDF classification, text extraction, and Markdown conversion without relying on expensive OCR.

    12,943developer-toolsRustMIT
  • run-llama/liteparse

    A fast, helpful, and open-source document parser

    AI summary: A fast, lightweight, and open-source Rust tool for high-quality spatial PDF parsing.

    11,931dataRustApache-2.0