Trending repositories: data-extraction

13 tracked repositories tagged with data-extraction, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

13 of 13 repositories

  • microsoft/markitdown

    Python tool for converting files and office documents to Markdown.

    AI summary: A lightweight Python utility for converting various file formats into clean Markdown.

    172,163dataPythonMIT
  • firecrawl/firecrawl

    The context API to search, scrape, and interact with the web at scale. 🔥

    AI summary: An API service that crawls websites and turns them into clean, LLM-ready markdown data.

    162,732dataTypeScriptAGPL-3.0
  • D4Vinci/Scrapling

    🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

    AI summary: A fast, undetectable web scraping framework for Python.

    72,975dataPythonBSD-3-Clause
  • NanmiCoder/MediaCrawler

    小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫

    AI summary: A comprehensive crawler for extracting posts and comments from major Chinese social media platforms.

    60,171dataPythonOther
  • google/langextract

    A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

    AI summary: AI-powered tool for extracting structured data from unstructured text using large language models.

    37,990dataPythonApache-2.0
  • Yuan1z0825/nature-skills

    符合nature论文学术表达和科研绘图的Skill

    AI summary: A collection of AI skills for parsing, analyzing, and synthesizing scientific literature from Nature.

    33,823learningPythonApache-2.0
  • opendataloader-project/opendataloader-pdf

    PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

    AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.

    28,188dataJavaApache-2.0
  • baidu/Unlimited-OCR

    Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

    AI summary: A system designed for one-shot, long-horizon parsing in OCR tasks.

    22,391ai-mlPythonMIT
  • eigent-ai/eigent

    Eigent: The Open Source Cowork Desktop - Local and Free Alternative to Claude Cowork and Codex

    AI summary: An intelligent data extraction and transformation pipeline for unstructured enterprise documents.

    14,783dataTypeScriptApache-2.0
  • rahulnyk/knowledge_graph

    Convert any text to a graph of knowledge. This can be used for Graph Augmented Generation or Knowledge Graph based QnA

    AI summary: A tool to convert unstructured text into knowledge graphs for enhanced QnA and data generation.

    3,621dataJupyter NotebookMIT
  • jackwener/xiaohongshu-cli

    A CLI for Xiaohongshu (小红书) — search, read, interact via reverse-engineered API

    AI summary: Command-line interface tool for interacting with Xiaohongshu (Little Red Book).

    2,463developer-toolsPython
  • tnm/zclaw

    Your personal AI assistant at all-in 888KiB (~35KB in app code). Running on an ESP32. GPIO, cron, custom tools, memory, and more.

    AI summary: ZClaw is a highly concurrent, distributed web scraper built for massive data extraction.

    2,203dataCMIT
  • peteromallet/dataclaw

    Agent harness to publish your agent chat history as Huggingface datasets.

    AI summary: A lightweight Python library for declarative data extraction and transformation.

    2,109dataPythonMIT