Trending repositories: document-processing

6 tracked repositories tagged with document-processing, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

6 of 6 repositories

  • opendataloader-project/opendataloader-pdf

    PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

    AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.

    28,250dataJavaApache-2.0
  • eigent-ai/eigent

    Eigent: The Open Source Cowork Desktop - Local and Free Alternative to Claude Cowork and Codex

    AI summary: An intelligent data extraction and transformation pipeline for unstructured enterprise documents.

    14,783dataTypeScriptApache-2.0
  • firecrawl/pdf-inspector

    Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

    AI summary: A fast Rust library for PDF classification, text extraction, and Markdown conversion without relying on expensive OCR.

    12,958developer-toolsRustMIT
  • run-llama/liteparse

    A fast, helpful, and open-source document parser

    AI summary: A fast, lightweight, and open-source Rust tool for high-quality spatial PDF parsing.

    11,931dataRustApache-2.0
  • langflow-ai/openrag

    OpenRAG is a comprehensive, single package Retrieval-Augmented Generation platform built on Langflow, Docling, and Opensearch.

    AI summary: An intelligent, agent-powered document search and RAG platform.

    4,405ai-mlPythonApache-2.0
  • BehemothSpongeSquare/adobe-acrobat

    AI summary: This repository provides scripts and tools for automating tasks within Adobe Acrobat, such as PDF creation, editing, and manipulation. It leverages Acrobat's JavaScript API to enable programmatic interaction with PDF files. The distinctive aspect is its focus on streamlining repetitive PDF workflows through custom scripts tailored to Acrobat's capabilities.

    0productivity