opendatalab/MinerUPublic

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

AI summary: High-accuracy parsing engine that converts complex documents into LLM-ready Markdown and JSON.

Stars
77.1K
+112 today
Forks
6.5K
Watchers
271
Open issues
50
Open PRs
36
Contributors
~98
Commits
5.7K
Branches
80

PythonOtherCreated Feb 29, 2024Last push 1d agoLatest release mineru-3.4.4-released+701 stars this week+901 this month

Star history

since Jun 23, 2024
020K40K60KJun 2024Mar 2025Nov 2025Aug 2026
77.1K stars as of Aug 7, 2026, tracked back to Jun 23, 2024. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 3 commits2025-08-12: 0 commits2025-08-13: 8 commits2025-08-14: 4 commits2025-08-15: 4 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 9 commits2025-08-19: 0 commits2025-08-20: 3 commits2025-08-21: 4 commits2025-08-22: 1 commit2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 5 commits2025-08-26: 7 commits2025-08-27: 7 commits2025-08-28: 4 commits2025-08-29: 7 commits2025-08-30: 1 commit2025-08-31: 0 commits2025-09-01: 1 commit2025-09-02: 1 commit2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 12 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 2 commits2025-09-09: 0 commits2025-09-10: 4 commits2025-09-11: 1 commit2025-09-12: 9 commits2025-09-13: 0 commits2025-09-14: 2 commits2025-09-15: 10 commits2025-09-16: 6 commits2025-09-17: 10 commits2025-09-18: 15 commits2025-09-19: 26 commits2025-09-20: 18 commits2025-09-21: 1 commit2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 1 commit2025-09-26: 8 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 1 commit2025-09-30: 3 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 1 commit2025-10-10: 1 commit2025-10-11: 1 commit2025-10-12: 1 commit2025-10-13: 2 commits2025-10-14: 3 commits2025-10-15: 3 commits2025-10-16: 13 commits2025-10-17: 9 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 6 commits2025-10-21: 3 commits2025-10-22: 12 commits2025-10-23: 7 commits2025-10-24: 20 commits2025-10-25: 1 commit2025-10-26: 0 commits2025-10-27: 1 commit2025-10-28: 11 commits2025-10-29: 8 commits2025-10-30: 15 commits2025-10-31: 14 commits2025-11-01: 1 commit2025-11-02: 0 commits2025-11-03: 17 commits2025-11-04: 8 commits2025-11-05: 1 commit2025-11-06: 2 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 1 commit2025-11-11: 12 commits2025-11-12: 6 commits2025-11-13: 5 commits2025-11-14: 3 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 1 commit2025-11-18: 12 commits2025-11-19: 9 commits2025-11-20: 14 commits2025-11-21: 4 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 1 commit2025-11-25: 17 commits2025-11-26: 16 commits2025-11-27: 7 commits2025-11-28: 1 commit2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 10 commits2025-12-02: 10 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 1 commit2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 3 commits2025-12-09: 2 commits2025-12-10: 0 commits2025-12-11: 2 commits2025-12-12: 14 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 4 commits2025-12-16: 4 commits2025-12-17: 2 commits2025-12-18: 8 commits2025-12-19: 6 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 12 commits2025-12-23: 12 commits2025-12-24: 7 commits2025-12-25: 4 commits2025-12-26: 10 commits2025-12-27: 2 commits2025-12-28: 2 commits2025-12-29: 3 commits2025-12-30: 15 commits2025-12-31: 2 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 6 commits2026-01-05: 5 commits2026-01-06: 13 commits2026-01-07: 5 commits2026-01-08: 0 commits2026-01-09: 4 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 2 commits2026-01-13: 4 commits2026-01-14: 5 commits2026-01-15: 1 commit2026-01-16: 6 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 4 commits2026-01-20: 5 commits2026-01-21: 5 commits2026-01-22: 15 commits2026-01-23: 16 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 8 commits2026-01-27: 0 commits2026-01-28: 4 commits2026-01-29: 13 commits2026-01-30: 20 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 6 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 11 commits2026-02-06: 2 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 7 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 10 commits2026-02-25: 7 commits2026-02-26: 12 commits2026-02-27: 5 commits2026-02-28: 9 commits2026-03-01: 4 commits2026-03-02: 2 commits2026-03-03: 5 commits2026-03-04: 5 commits2026-03-05: 6 commits2026-03-06: 2 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 4 commits2026-03-10: 4 commits2026-03-11: 3 commits2026-03-12: 7 commits2026-03-13: 4 commits2026-03-14: 1 commit2026-03-15: 2 commits2026-03-16: 7 commits2026-03-17: 9 commits2026-03-18: 4 commits2026-03-19: 11 commits2026-03-20: 12 commits2026-03-21: 13 commits2026-03-22: 5 commits2026-03-23: 8 commits2026-03-24: 13 commits2026-03-25: 9 commits2026-03-26: 19 commits2026-03-27: 7 commits2026-03-28: 13 commits2026-03-29: 18 commits2026-03-30: 19 commits2026-03-31: 7 commits2026-04-01: 12 commits2026-04-02: 3 commits2026-04-03: 4 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 6 commits2026-04-08: 0 commits2026-04-09: 3 commits2026-04-10: 4 commits2026-04-11: 7 commits2026-04-12: 0 commits2026-04-13: 6 commits2026-04-14: 13 commits2026-04-15: 13 commits2026-04-16: 12 commits2026-04-17: 11 commits2026-04-18: 4 commits2026-04-19: 0 commits2026-04-20: 8 commits2026-04-21: 8 commits2026-04-22: 16 commits2026-04-23: 10 commits2026-04-24: 3 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 6 commits2026-04-28: 11 commits2026-04-29: 1 commit2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 4 commits2026-05-05: 0 commits2026-05-06: 5 commits2026-05-07: 7 commits2026-05-08: 5 commits2026-05-09: 8 commits2026-05-10: 5 commits2026-05-11: 7 commits2026-05-12: 7 commits2026-05-13: 13 commits2026-05-14: 13 commits2026-05-15: 7 commits2026-05-16: 3 commits2026-05-17: 13 commits2026-05-18: 9 commits2026-05-19: 6 commits2026-05-20: 5 commits2026-05-21: 5 commits2026-05-22: 5 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 10 commits2026-05-26: 9 commits2026-05-27: 21 commits2026-05-28: 3 commits2026-05-29: 3 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 4 commits2026-06-02: 15 commits2026-06-03: 0 commits2026-06-04: 6 commits2026-06-05: 13 commits2026-06-06: 11 commits2026-06-07: 0 commits2026-06-08: 3 commits2026-06-09: 12 commits2026-06-10: 10 commits2026-06-11: 17 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 1 commit2026-06-16: 5 commits2026-06-17: 8 commits2026-06-18: 9 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 7 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 2 commits2026-07-09: 0 commits2026-07-10: 3 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
1,496 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Landmark project

    77,081 stars

  • Very active

    1,496 commits in 52 weeks

  • Repeat trending

    4 trending appearances

  • Top 10% tracked

    Rank 91 of 1058

What MinerU does

MinerU solves the problem of extracting clean data from complex, unstructured documents for AI applications. It acts as a high-precision parsing engine that transforms PDFs, DOCX, PPTX, and XLSX files—including those with scanned images, complex tables, and multi-column layouts—into structured Markdown or JSON. It utilizes a dual engine combining Vision Language Models (VLMs) and Optical Character Recognition (OCR) supporting 109 languages. MinerU carefully reconstructs document layouts, extracts mathematical formulas to LaTeX, and maps tables to HTML, ensuring the output maintains human reading order and is perfectly formatted for RAG or agentic workflows.

Data scientists, AI engineers, and developers building RAG systems or agentic workflows requiring high-quality document extraction.

  • VLM+OCR dual engine: Combines vision models and OCR for high-accuracy extraction across 109 languages.
  • Complex layout analysis: Accurately handles multi-column layouts, cross-page tables, and removes headers/footers.
  • Formula and table extraction: Converts mathematical formulas to LaTeX and complex tables to structured HTML.
  • Multi-format support: Natively parses PDF, DOCX, PPTX, XLSX, and images into Markdown/JSON.
  • Agent integration: Provides MCP Server and native integrations for LangChain, Dify, and FastGPT.

Where teams use it

RAG Data Preparation

Data engineers use it to parse vast archives of PDF reports into clean Markdown to feed into vector databases.

Scientific Paper Extraction

Researchers extract text and complex LaTeX formulas from academic papers for automated analysis.

Financial Document Processing

Financial analysts extract structured tables from corporate PPTX and XLSX files for LLM summarization.

Agentic Workflows

AI agents use the MCP server integration to dynamically read and understand complex documents during execution.

Getting started: Access the zero-install web version or deploy locally via provided guides.

README

master branch
MinerU — High-accuracy document parsing engine for LLM · RAG · Agent workflows Converts PDF · DOCX · PPTX · XLSX · Images · Web pages into structured Markdown / JSON · VLM+OCR dual engine · 109 languages
MCP Server · LangChain / Dify / FastGPT native integration · 10+ domestic AI chip support

🔍 Core Parsing Capabilities

  • Native support for DOCX, PPTX, and XLSX parsing
  • Formulas → LaTeX · Tables → HTML, accurate layout reconstruction
  • Supports scanned docs, handwriting, multi-column layouts, cross-page table merging
  • Output follows human reading order with automatic header/footer removal
  • VLM + OCR dual engine, 109-language OCR recognition

🔌 Integration

Use Case Solution
AI Coding Tools MCP Server — Cursor · Claude Desktop · Windsurf
RAG Frameworks LangChain · LlamaIndex · RAGFlow · RAG-Anything · Flowise · Dify · FastGPT
Development Python / Go / TypeScript SDK · CLI · REST API · Docker
No-Code mineru.net online · Gradio WebUI · Desktop client

🖥️ Deployment (Private · Fully Offline)

Inference Backend Best For
pipeline Fast & stable, no hallucination, runs on CPU or GPU
vlm-engine High accuracy, supports vLLM / LMDeploy / mlx ecosystem
hybrid-engine High accuracy, native text extraction, low hallucination

Domestic AI chips: Ascend · Cambricon · Enflame · MetaX · Moore Threads · Kunlunxin · Iluvatar · Hygon · Biren · T-Head

Changelog

  • 2026/06/18 3.4 Released

    This release focuses on OCR capability upgrades for the pipeline backend, OCR processing pipeline optimization, and model download experience improvements. The main updates include:

    • OCR model upgrade and processing acceleration

      • The OCR model for the pipeline backend has been upgraded to PP-OCRv6, improving OCR accuracy by about 11% on OmniDocBench v1.6.
      • Removed Japanese, Traditional Chinese, English, and Latin options from OCR language selection. These scenarios are now routed to the ch OCR model, simplifying model configuration and language selection.
      • Optimized the OCR inference and processing pipeline, increasing OCR processing speed by about 100% and significantly improving parsing efficiency for batch documents and OCR-intensive documents.
    • Model download logic optimization

      • Added automatic model source selection, allowing first-time installations to choose a better model source based on the current network environment.
      • Before downloading models, MinerU now prioritizes checking locally downloaded model cache files. Cache hits can be reused directly, reducing repeated downloads and unnecessary remote requests.
      • For more details about model source configuration, automatic source selection, and local model usage, see the Model Source Documentation.

    With the 3.4 release, MinerU further improves the parsing accuracy and processing efficiency of the pipeline backend in OCR scenarios. It also optimizes model downloads, cache reuse, and local configuration write-back, making first-time installation, model updates, and multi-environment deployment more stable and automated.

  • 2026/06/11 3.3 Released

    This release focuses on Hybrid parsing performance optimization and VLM model capability upgrades. The main updates include:

    • New effort parsing-strength parameter for the Hybrid backend

      • Added two parsing-strength levels, medium and high, allowing users to balance parsing speed, parsing accuracy, and feature requirements.
      • On OmniDocBench v1.6, medium reduces overall accuracy by only 0.13 points compared with high, while delivering 35% ~ 220% parsing speed improvements across different devices and scenarios:
        • Linux: about 80% faster for text PDF scenarios and about 35% faster for OCR scenarios
        • Windows: about 90% faster for text PDF scenarios and about 45% faster for OCR scenarios
        • macOS: about 220% faster for text PDF scenarios and about 50% faster for OCR scenarios
      • The default Hybrid backend now uses effort=medium, significantly improving overall parsing efficiency while maintaining high parsing accuracy.
      • The medium level does not support image analysis; for maximum parsing accuracy or image analysis support, switch to the high-strength parsing mode with effort=high, which may have an impact on parsing speed.
    • VLM model upgraded to MinerU2.5-Pro-2605-1.2B

      • Fixed multiple model issues found in the 2604 version, further improving parsing stability on complex documents.
      • Added native multilingual OCR support, reducing the need for extra language-parameter configuration and improving out-of-the-box usability for multilingual documents.

    With the 3.3 release, MinerU further improves Hybrid backend efficiency across platforms and scenarios while maintaining high-accuracy parsing. The default medium effort level is better suited for most day-to-day document processing tasks, while high is designed for scenarios that require maximum parsing accuracy or image analysis capabilities.

  • 2026/04/18 3.1.0 Released

    This release focuses on licensing openness, parsing accuracy, and full-format native support. The main updates include:

    • License upgrade
      • MinerU has officially moved from AGPLv3 to the MinerU Open Source License, a custom license based on Apache 2.0.
      • This change significantly reduces adoption friction for both community users and commercial deployments, making MinerU easier to integrate into real-world workflows.
    • VLM main model upgrade
      • The primary VLM model has been upgraded to MinerU2.5-Pro-2604-1.2B, bringing overall parsing accuracy to a state-of-the-art level.
      • The new model now supports image and chart parsing, truncated paragraph merging, cross-page table merging, and image recognition inside tables, further strengthening performance on complex document layouts.
    • Full-format native parsing support
      • Native parsing support has now been extended to PPTX and XLSX.
      • MinerU now fully supports parsing across images, PDF, DOCX, PPTX, and XLSX, providing a more complete multi-format document understanding workflow.

    With the 3.1.0 release, MinerU becomes more open, more accurate, and easier to adopt in production. The new license lowers the barrier for both community and commercial use, MinerU2.5-Pro-2604-1.2B improves parsing quality on complex content, and native PPTX / XLSX support completes end-to-end coverage of mainstream document formats.

  • 2026/03/29 3.0.0 Released

    This release delivers a systematic upgrade centered on parsing capability, system architecture, and engineering usability. The main updates include:

    • Native DOCX parsing
      • Official support for native DOCX parsing, delivering high-precision results without hallucinations.
      • Compared with the traditional workflow of first converting DOCX to PDF and then parsing it, end-to-end speed is improved by tens of times, making it better suited for scenarios with high requirements for both accuracy and throughput.
    • pipeline backend upgrade
      • The pipeline backend achieves a score of 86.2 on OmniDocBench (v1.5), surpassing the accuracy of the previous-generation mainstream VLM MinerU2.0-2505-0.9B.
      • Added support for parsing images/formulas inside tables, seal text recognition, vertical text support, and interline formula numbering recognition, continuously improving parsing quality for complex document scenarios.
      • While maintaining high accuracy, it keeps resource usage extremely low and continues to support inference in pure CPU environments.
    • API / CLI / Router orchestration upgrade
      • mineru now runs as an orchestration client based on mineru-api; when --api-url is not provided, it will automatically start a local temporary service.
      • mineru-api adds a new asynchronous task endpoint POST /tasks, supporting task submission, status querying, and result retrieval; meanwhile, it retains the synchronous parsing endpoint POST /file_parse for compatibility with legacy plugins.
      • Added mineru-router, designed for unified entry deployment and task routing across multiple services and multiple GPUs; its interfaces are fully compatible with mineru-api and support automatic task load balancing.
    • Deployment and usability improvements
      • Resolved compatibility issues with torch >= 2.8; the base image has been upgraded to vllm0.11.2 + torch2.9.0, unifying installation paths across different Compute Capabilities.
      • Optimized the parsing pipeline with a sliding-window mechanism, significantly reducing peak memory usage in long-document scenarios, so documents with tens of thousands of pages no longer need to be split manually.
      • Batch inference in pipeline now supports streaming writes to disk, allowing completed parsing results to be written out in time and further improving the experience for long-running tasks.
      • Completed thread-safety optimization and now fully supports multi-threaded concurrent inference; together with mineru-router, this enables one-click multi-GPU deployment and makes it easy to build high-concurrency, high-throughput parsing systems.
      • Completely removed the use of two AGPLv3 models (doclayoutyolo and mfd_yolov8) and one CC-BY-NC-SA 4.0 model (layoutreader).

    This update is not just a set of feature enhancements, but a key leap forward in MinerU's overall system capabilities. We specifically addressed the peak memory usage issue in long-document parsing. Through optimizations such as sliding windows and streaming writes to disk, ultra-long document parsing has moved from “requiring manual splitting and careful handling” to being “stable, scalable, and ready for production workloads.” At the same time, we completed thread-safety optimization and fully enabled multi-threaded concurrent inference, further improving single-machine resource utilization and runtime stability under high-concurrency workloads. On top of this, with mineru-router and the new API / CLI orchestration framework, MinerU now supports one-click multi-GPU deployment, unified access across multiple services, and automatic task load balancing, significantly reducing the difficulty of large-scale deployment. As a result, MinerU is evolving from a standalone data production tool into a large-scale document parsing foundation for high-concurrency and high-throughput scenarios, providing enterprise-grade document data processing with infrastructure that is more stable, more efficient, and easier to scale.

📝 View the complete Changelog for more historical version information

MinerU

Project Introduction

MinerU is a document parsing tool that converts PDF, image, DOCX, PPTX, and XLSX inputs into machine-readable formats such as Markdown and JSON for downstream retrieval, extraction, and processing. MinerU was born during the pre-training process of InternLM. We focus on solving symbol conversion issues in scientific literature and hope to contribute to technological development in the era of large models. Compared to well-known commercial products, MinerU is still young. If you encounter any issues or if the results are not as expected, please submit an issue on issue and attach the relevant document or sample file.

pdf_zh_cn.mp4

Key Features

  • Support PDF, image, DOCX, PPTX, and XLSX inputs.
  • Remove headers, footers, footnotes, page numbers, etc., to ensure semantic coherence.
  • Output text in human-readable order, suitable for single-column, multi-column, and complex layouts.
  • Preserve the structure of the original document, including headings, paragraphs, lists, etc.
  • Extract images, image descriptions, tables, table titles, and footnotes.
  • Automatically recognize and convert formulas in the document to LaTeX format.
  • Automatically recognize and convert tables in the document to HTML format.
  • Automatically detect scanned PDFs and garbled PDFs and enable OCR functionality.
  • OCR supports detection and recognition of 109 languages.
  • Supports multiple output formats, such as multimodal and NLP Markdown, JSON sorted by reading order, and rich intermediate formats.
  • Supports various visualization results, including layout visualization and span visualization, for efficient confirmation of output quality.
  • Built-in CLI, FastAPI, Gradio WebUI, for local orchestration and multi-service deployment.
  • Supports running in a pure CPU environment, and also supports GPU/MPS acceleration
  • Compatible with Windows, Linux, and Mac platforms.

Quick Start

Document parsing is a difficult and complex task. In scenarios such as complex layouts, scanned pages, and handwritten content, the parsing results may fall short of expectations. We recommend trying the online demo first to evaluate MinerU's parsing quality and suitability before choosing an appropriate deployment method based on your actual needs. If you have document samples with unsatisfactory parsing results, feel free to share them in an issue. We will continue improving the parsing capabilities. If you encounter any installation issues, please first consult the FAQ.

Online Experience

Official online web application

The official online version has the same functionality as the client, with a beautiful interface and rich features, requires login to use

  • OpenDataLab

Gradio-based online demo

A WebUI developed based on Gradio, with a simple interface and only core parsing functionality, no login required

  • ModelScope
  • HuggingFace

Local Deployment

Warning

Pre-installation Notice—Hardware and Software Environment Support

To ensure the stability and reliability of the project, we only optimize and test for specific hardware and software environments during development. This ensures that users deploying and running the project on recommended system configurations will get the best performance with the fewest compatibility issues.

By focusing resources on the mainline environment, our team can more efficiently resolve potential bugs and develop new features.

In non-mainline environments, due to the diversity of hardware and software configurations, as well as third-party dependency compatibility issues, we cannot guarantee 100% project availability. Therefore, for users who wish to use this project in non-recommended environments, we suggest carefully reading the documentation and FAQ first. Most issues already have corresponding solutions in the FAQ. We also encourage community feedback to help us gradually expand support.

Parsing Backend pipeline *-engine *-http-client
hybrid vlm hybrid vlm
Backend Features Good Compatibility High Hardware Requirements For OpenAI Compatible Servers2
Accuracy1 86.47 95.39 (high)
95.26 (medium)
95.30 95.39 (high)
95.26 (medium)
95.30
Operating System Linux3 / Windows4 / macOS5
Pure CPU Support
GPU Acceleration Volta and later architecture GPUs or Apple Silicon Not Required
Min VRAM 4GB 8GB 2GB
RAM Min 16GB, Recommended 32GB or more Min 16GB
Disk Space Min 20GB, SSD Recommended Min 2GB
Python Version 3.10-3.13

1 Accuracy metrics are the End-to-End Evaluation Overall scores from OmniDocBench (v1.6), based on the latest version of MinerU.
2 Servers compatible with OpenAI API, such as local model servers or remote model services deployed via inference frameworks like vLLM/SGLang/LMDeploy.
3 Linux only supports distributions from 2019 and later.
4 Since the key dependency ray does not support Python 3.13 on Windows, only versions 3.10~3.12 are supported.
5 macOS requires version 14.0 or later.

Install MinerU

Install MinerU using pip or uv

pip install --upgrade pip
pip install uv
uv pip install -U "mineru[all]"

Install MinerU from source code

git clone https://github.com/opendatalab/MinerU.git
cd MinerU
uv pip install -e .[all]

Tip

  • mineru[all] includes all core features, compatible with Windows / Linux / macOS systems, suitable for most users.
  • If CUDA acceleration is unavailable after installing on Windows, see the Windows CUDA acceleration FAQ.
  • If you need to specify the inference framework for the VLM model, or only intend to install a lightweight client on an edge device, please refer to the documentation Extension Modules Installation Guide.

Deploy MinerU using Docker

MinerU provides a convenient Docker deployment method, which helps quickly set up the environment and solve some tricky environment compatibility issues.

Tip

  • Docker deployment is only supported on Linux and Windows environments with WSL2 support;
  • macOS users should refer to the two installation methods above for installation instead of using Docker deployment.

You can get the Docker Deployment Instructions in the documentation.


Using MinerU

If your device meets the GPU acceleration requirements in the table above, you can use a simple command line for document parsing:

mineru -p <input_path> -o <output_path>

If your device does not meet the GPU acceleration requirements, you can specify the backend as pipeline to run in a pure CPU environment:

mineru -p <input_path> -o <output_path> -b pipeline

mineru currently supports local PDF, image, DOCX, PPTX, and XLSX file or directory inputs, and can be used for document parsing through the CLI, API, WebUI, and mineru-router. For detailed instructions, please refer to the Usage Guide.

FAQ

  • If you encounter any issues during usage, you can first check the FAQ for solutions.
  • If your issue remains unresolved, you may also use DeepWiki to interact with an AI assistant, which can address most common problems.
  • If you still cannot resolve the issue, you are welcome to join our community via Discord or WeChat to discuss with other users and developers.

All Thanks To Our Contributors

License Information

This repository is licensed under the MinerU Open Source License, based on Apache 2.0 with additional conditions.

Acknowledgments

Citation

@article{wang2026mineru2,
  title={MinerU2. 5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale},
  author={Wang, Bin and He, Tianyao and Ouyang, Linke and Wu, Fan and Zhao, Zhiyuan and Chu, Tao and Qu, Yuan and Jin, Zhenjiang and Zeng, Weijun and Miao, Ziyang and others},
  journal={arXiv preprint arXiv:2604.04771},
  year={2026}
}

@article{dong2026minerudiffusion,
  title={MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding},
  author={Dong, Hejun and Niu, Junbo and Wang, Bin and Zeng, Weijun and Zhang, Wentao and He, Conghui},
  journal={arXiv preprint arXiv:2603.22458},
  year={2026}
}

@article{niu2025mineru2,
  title={Mineru2. 5: A decoupled vision-language model for efficient high-resolution document parsing},
  author={Niu, Junbo and Liu, Zheng and Gu, Zhuangcheng and Wang, Bin and Ouyang, Linke and Zhao, Zhiyuan and Chu, Tao and He, Tianyao and Wu, Fan and Zhang, Qintong and others},
  journal={arXiv preprint arXiv:2509.22186},
  year={2025}
}

@article{wang2024mineru,
  title={Mineru: An open-source solution for precise document content extraction},
  author={Wang, Bin and Xu, Chao and Zhao, Xiaomeng and Ouyang, Linke and Wu, Fan and Zhao, Zhiyuan and Xu, Rui and Liu, Kaiwen and Qu, Yuan and Shang, Fukai and others},
  journal={arXiv preprint arXiv:2409.18839},
  year={2024}
}

@article{he2024opendatalab,
  title={Opendatalab: Empowering general artificial intelligence with open datasets},
  author={He, Conghui and Li, Wei and Jin, Zhenjiang and Xu, Chao and Wang, Bin and Lin, Dahua},
  journal={arXiv preprint arXiv:2407.13773},
  year={2024}
}

Star History

Star History Chart

Links

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

184 total
  1. MinerU 4.0.0a5v4.0.0a5Jul 30, 2026pre-release43 downloads

    **Full Changelog**: https://github.com/opendatalab/MinerU/compare/v4.0.0a4...v4.0.0a5

  2. MinerU 4.0.0a4v4.0.0a4Jul 27, 2026pre-release30 downloads

    **Full Changelog**: https://github.com/opendatalab/MinerU/compare/v4.0.0a3...v4.0.0a4

  3. MinerU 4.0.0a3v4.0.0a3Jul 22, 2026pre-release40 downloads

    ## What's Changed * Add model download completion markers by @johnking0099 in https://github.com/opendatalab/MinerU/pull/5285 * Clarify text file handling in doclib by @johnking0099 in https://github.com/opendatalab/MinerU/pull/5286 * Expose doclib visual blocks through locators by @johnking0099 in https://github.com/opendatalab/MinerU/pull/5306 * build: enable hf-xet for model downloads by @johnking0099 in https://github.com/opendatalab/MinerU/pull/5307 * fix: bind doclib discovery to endpoint instance by @johnking0099 in https://github.com/opendatalab/MinerU/pull/5308 * Handle legacy doclib server shutdown safely by @johnking0099 in https://github.com/opendatalab/MinerU/pull/5309 ## New Contributors * @johnking0099 made their first contribution in https://github.com/opendatalab/MinerU/pull/5285 **Full Changelog**: https://github.com/opendatalab/MinerU/compare/v4.0.0a2...v4.0.0a3

  4. MinerU 4.0.0a2v4.0.0a2Jul 15, 2026pre-release44 downloads

    **Full Changelog**: https://github.com/opendatalab/MinerU/compare/v4.0.0a1...v4.0.0a2

  5. MinerU 4.0.0a1v4.0.0a1Jul 14, 2026pre-release24 downloads

    ## What's Changed * master->dev by @myhloli in https://github.com/opendatalab/MinerU/pull/5042 * Enhance PDF processing and improve concurrency management by @myhloli in https://github.com/opendatalab/MinerU/pull/5062 * feat: add functionality to skip broken PDF pages during rewrite process by @myhloli in https://github.com/opendatalab/MinerU/pull/5064 * next with 3.2.2 by @myhloli in https://github.com/opendatalab/MinerU/pull/5174 * feat: integrate ImagePayloadCache for improved image handling and path management by @myhloli in https://github.com/opendatalab/MinerU/pull/5194 * Next with 3.3.1 by @myhloli in https://github.com/opendatalab/MinerU/pull/5196 * Next with 3.4.0 by @myhloli in https://github.com/opendatalab/MinerU/pull/5202 * feat: enhance formula content handling with normalization and tagging functions by @myhloli in https://github.com/opendatalab/MinerU/pull/5206 * feat: adjust block type handling for INDEX in hybrid magic model by @myhloli in https://github.com/opendatalab/MinerU/pull/5207 * feat: remove VLM backend by @myhloli in https://github.com/opendatalab/MinerU/pull/5214 * Next by @myhloli in https://github.com/opendatalab/MinerU/pull/5236 * Next by @myhloli i

Code frequency

additions and deletions
+197.9K-197.9KWeek of 2025-08-10: +2,484 linesWeek of 2025-08-10: -1,994 linesWeek of 2025-08-17: +3,073 linesWeek of 2025-08-17: -1,539 linesWeek of 2025-08-24: +1,403 linesWeek of 2025-08-24: -1,213 linesWeek of 2025-08-31: +2,119 linesWeek of 2025-08-31: -182 linesWeek of 2025-09-07: +560 linesWeek of 2025-09-07: -3,464 linesWeek of 2025-09-14: +1,869 linesWeek of 2025-09-14: -1,060 linesWeek of 2025-09-21: +60 linesWeek of 2025-09-21: -24 linesWeek of 2025-09-28: +22 linesWeek of 2025-09-28: -16 linesWeek of 2025-10-05: +193 linesWeek of 2025-10-05: -73 linesWeek of 2025-10-12: +197,924 linesWeek of 2025-10-12: -188,821 linesWeek of 2025-10-19: +4,790 linesWeek of 2025-10-19: -25,281 linesWeek of 2025-10-26: +925 linesWeek of 2025-10-26: -654 linesWeek of 2025-11-02: +608 linesWeek of 2025-11-02: -337 linesWeek of 2025-11-09: +610 linesWeek of 2025-11-09: -314 linesWeek of 2025-11-16: +6,887 linesWeek of 2025-11-16: -373 linesWeek of 2025-11-23: +713 linesWeek of 2025-11-23: -277 linesWeek of 2025-11-30: +375 linesWeek of 2025-11-30: -332 linesWeek of 2025-12-07: +1,567 linesWeek of 2025-12-07: -446 linesWeek of 2025-12-14: +2,304 linesWeek of 2025-12-14: -1,152 linesWeek of 2025-12-21: +2,707 linesWeek of 2025-12-21: -2,327 linesWeek of 2025-12-28: +502 linesWeek of 2025-12-28: -388 linesWeek of 2026-01-04: +6,125 linesWeek of 2026-01-04: -8,496 linesWeek of 2026-01-11: +4,438 linesWeek of 2026-01-11: -662 linesWeek of 2026-01-18: +941 linesWeek of 2026-01-18: -247 linesWeek of 2026-01-25: +2,506 linesWeek of 2026-01-25: -1,405 linesWeek of 2026-02-01: +633 linesWeek of 2026-02-01: -266 linesWeek of 2026-02-08: +405 linesWeek of 2026-02-08: -334 linesWeek of 2026-02-15: +0 linesWeek of 2026-02-15: -0 linesWeek of 2026-02-22: +2,143 linesWeek of 2026-02-22: -4,401 linesWeek of 2026-03-01: +1,373 linesWeek of 2026-03-01: -781 linesWeek of 2026-03-08: +3,294 linesWeek of 2026-03-08: -1,585 linesWeek of 2026-03-15: +14,105 linesWeek of 2026-03-15: -10,609 linesWeek of 2026-03-22: +8,570 linesWeek of 2026-03-22: -4,562 linesWeek of 2026-03-29: +2,801 linesWeek of 2026-03-29: -898 linesWeek of 2026-04-05: +3,362 linesWeek of 2026-04-05: -2,767 linesWeek of 2026-04-12: +4,441 linesWeek of 2026-04-12: -1,933 linesWeek of 2026-04-19: +2,312 linesWeek of 2026-04-19: -741 linesWeek of 2026-04-26: +1,039 linesWeek of 2026-04-26: -172 linesWeek of 2026-05-03: +1,599 linesWeek of 2026-05-03: -331 linesWeek of 2026-05-10: +6,329 linesWeek of 2026-05-10: -2,667 linesWeek of 2026-05-17: +4,145 linesWeek of 2026-05-17: -1,988 linesWeek of 2026-05-24: +1,603 linesWeek of 2026-05-24: -976 linesWeek of 2026-05-31: +2,595 linesWeek of 2026-05-31: -1,556 linesWeek of 2026-06-07: +1,372 linesWeek of 2026-06-07: -738 linesWeek of 2026-06-14: +21,047 linesWeek of 2026-06-14: -42,770 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +364 linesWeek of 2026-06-28: -50 linesWeek of 2026-07-05: +381 linesWeek of 2026-07-05: -26 linesWeek of 2026-07-12: +0 linesWeek of 2026-07-12: -0 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesAug 10, 2025Aug 2, 2026
+329.6K lines added, -321.2K removed over the last year.

Commits per week

last 52 weeks
870Week of 2025-08-10: 19 commitsWeek of 2025-08-17: 17 commitsWeek of 2025-08-24: 31 commitsWeek of 2025-08-31: 14 commitsWeek of 2025-09-07: 16 commitsWeek of 2025-09-14: 87 commitsWeek of 2025-09-21: 10 commitsWeek of 2025-09-28: 4 commitsWeek of 2025-10-05: 3 commitsWeek of 2025-10-12: 31 commitsWeek of 2025-10-19: 49 commitsWeek of 2025-10-26: 50 commitsWeek of 2025-11-02: 28 commitsWeek of 2025-11-09: 27 commitsWeek of 2025-11-16: 40 commitsWeek of 2025-11-23: 42 commitsWeek of 2025-11-30: 21 commitsWeek of 2025-12-07: 21 commitsWeek of 2025-12-14: 24 commitsWeek of 2025-12-21: 47 commitsWeek of 2025-12-28: 22 commitsWeek of 2026-01-04: 33 commitsWeek of 2026-01-11: 18 commitsWeek of 2026-01-18: 45 commitsWeek of 2026-01-25: 45 commitsWeek of 2026-02-01: 19 commitsWeek of 2026-02-08: 7 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 43 commitsWeek of 2026-03-01: 24 commitsWeek of 2026-03-08: 23 commitsWeek of 2026-03-15: 58 commitsWeek of 2026-03-22: 74 commitsWeek of 2026-03-29: 63 commitsWeek of 2026-04-05: 20 commitsWeek of 2026-04-12: 59 commitsWeek of 2026-04-19: 45 commitsWeek of 2026-04-26: 18 commitsWeek of 2026-05-03: 29 commitsWeek of 2026-05-10: 55 commitsWeek of 2026-05-17: 43 commitsWeek of 2026-05-24: 46 commitsWeek of 2026-05-31: 49 commitsWeek of 2026-06-07: 42 commitsWeek of 2026-06-14: 23 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 7 commitsWeek of 2026-07-05: 5 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 0 commitsAug 10, 2025Aug 2, 2026
1.5K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 8 commitsSun 1:00 — 14 commitsSun 2:00 — 16 commitsSun 3:00 — 15 commitsSun 4:00 — 6 commitsSun 5:00 — 1 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 9 commitsSun 11:00 — 11 commitsSun 12:00 — 3 commitsSun 13:00 — 6 commitsSun 14:00 — 3 commitsSun 15:00 — 5 commitsSun 16:00 — 5 commitsSun 17:00 — 8 commitsSun 18:00 — 8 commitsSun 19:00 — 5 commitsSun 20:00 — 11 commitsSun 21:00 — 4 commitsSun 22:00 — 9 commitsSun 23:00 — 9 commitsMon 0:00 — 8 commitsMon 1:00 — 11 commitsMon 2:00 — 4 commitsMon 3:00 — 6 commitsMon 4:00 — 4 commitsMon 5:00 — 2 commitsMon 6:00 — 3 commitsMon 7:00 — 3 commitsMon 8:00 — 0 commitsMon 9:00 — 4 commitsMon 10:00 — 32 commitsMon 11:00 — 31 commitsMon 12:00 — 24 commitsMon 13:00 — 10 commitsMon 14:00 — 41 commitsMon 15:00 — 65 commitsMon 16:00 — 64 commitsMon 17:00 — 72 commitsMon 18:00 — 68 commitsMon 19:00 — 72 commitsMon 20:00 — 41 commitsMon 21:00 — 19 commitsMon 22:00 — 24 commitsMon 23:00 — 9 commitsTue 0:00 — 14 commitsTue 1:00 — 14 commitsTue 2:00 — 18 commitsTue 3:00 — 12 commitsTue 4:00 — 1 commitsTue 5:00 — 1 commitsTue 6:00 — 5 commitsTue 7:00 — 2 commitsTue 8:00 — 1 commitsTue 9:00 — 10 commitsTue 10:00 — 49 commitsTue 11:00 — 78 commitsTue 12:00 — 20 commitsTue 13:00 — 9 commitsTue 14:00 — 55 commitsTue 15:00 — 70 commitsTue 16:00 — 82 commitsTue 17:00 — 90 commitsTue 18:00 — 68 commitsTue 19:00 — 80 commitsTue 20:00 — 55 commitsTue 21:00 — 28 commitsTue 22:00 — 19 commitsTue 23:00 — 10 commitsWed 0:00 — 23 commitsWed 1:00 — 19 commitsWed 2:00 — 11 commitsWed 3:00 — 6 commitsWed 4:00 — 2 commitsWed 5:00 — 0 commitsWed 6:00 — 4 commitsWed 7:00 — 4 commitsWed 8:00 — 3 commitsWed 9:00 — 7 commitsWed 10:00 — 36 commitsWed 11:00 — 53 commitsWed 12:00 — 15 commitsWed 13:00 — 6 commitsWed 14:00 — 65 commitsWed 15:00 — 51 commitsWed 16:00 — 76 commitsWed 17:00 — 69 commitsWed 18:00 — 55 commitsWed 19:00 — 61 commitsWed 20:00 — 44 commitsWed 21:00 — 11 commitsWed 22:00 — 24 commitsWed 23:00 — 11 commitsThu 0:00 — 15 commitsThu 1:00 — 14 commitsThu 2:00 — 13 commitsThu 3:00 — 6 commitsThu 4:00 — 3 commitsThu 5:00 — 0 commitsThu 6:00 — 1 commitsThu 7:00 — 3 commitsThu 8:00 — 3 commitsThu 9:00 — 3 commitsThu 10:00 — 34 commitsThu 11:00 — 56 commitsThu 12:00 — 19 commitsThu 13:00 — 10 commitsThu 14:00 — 56 commitsThu 15:00 — 73 commitsThu 16:00 — 80 commitsThu 17:00 — 77 commitsThu 18:00 — 82 commitsThu 19:00 — 70 commitsThu 20:00 — 25 commitsThu 21:00 — 37 commitsThu 22:00 — 15 commitsThu 23:00 — 21 commitsFri 0:00 — 20 commitsFri 1:00 — 18 commitsFri 2:00 — 17 commitsFri 3:00 — 8 commitsFri 4:00 — 2 commitsFri 5:00 — 5 commitsFri 6:00 — 1 commitsFri 7:00 — 4 commitsFri 8:00 — 6 commitsFri 9:00 — 10 commitsFri 10:00 — 45 commitsFri 11:00 — 52 commitsFri 12:00 — 21 commitsFri 13:00 — 18 commitsFri 14:00 — 51 commitsFri 15:00 — 98 commitsFri 16:00 — 73 commitsFri 17:00 — 87 commitsFri 18:00 — 101 commitsFri 19:00 — 52 commitsFri 20:00 — 25 commitsFri 21:00 — 23 commitsFri 22:00 — 8 commitsFri 23:00 — 16 commitsSat 0:00 — 17 commitsSat 1:00 — 19 commitsSat 2:00 — 22 commitsSat 3:00 — 9 commitsSat 4:00 — 20 commitsSat 5:00 — 2 commitsSat 6:00 — 0 commitsSat 7:00 — 2 commitsSat 8:00 — 1 commitsSat 9:00 — 1 commitsSat 10:00 — 4 commitsSat 11:00 — 7 commitsSat 12:00 — 6 commitsSat 13:00 — 7 commitsSat 14:00 — 14 commitsSat 15:00 — 16 commitsSat 16:00 — 7 commitsSat 17:00 — 20 commitsSat 18:00 — 26 commitsSat 19:00 — 11 commitsSat 20:00 — 3 commitsSat 21:00 — 1 commitsSat 22:00 — 5 commitsSat 23:00 — 7 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Jun 29, 2026daily#22+11
Jun 27, 2026daily#15+36
Jun 26, 2026daily#11+21
Jun 25, 2026daily#21+17