Trending repositories: data-extraction
13 tracked repositories tagged with data-extraction, ordered by stars. Use the topic filters below to narrow further.
13 of 13 repositories
microsoft/markitdown
Python tool for converting files and office documents to Markdown.
AI summary: A lightweight Python utility for converting various file formats into clean Markdown.
172,163dataPythonMITfirecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
AI summary: An API service that crawls websites and turns them into clean, LLM-ready markdown data.
162,732dataTypeScriptAGPL-3.0D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
AI summary: A fast, undetectable web scraping framework for Python.
72,975dataPythonBSD-3-ClauseNanmiCoder/MediaCrawler
小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫
AI summary: A comprehensive crawler for extracting posts and comments from major Chinese social media platforms.
60,171dataPythonOthergoogle/langextract
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
AI summary: AI-powered tool for extracting structured data from unstructured text using large language models.
37,990dataPythonApache-2.0Yuan1z0825/nature-skills
符合nature论文学术表达和科研绘图的Skill
AI summary: A collection of AI skills for parsing, analyzing, and synthesizing scientific literature from Nature.
33,823learningPythonApache-2.0opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
AI summary: A high-accuracy PDF parser that extracts AI-ready Markdown, JSON, and HTML using deterministic and hybrid approaches.
28,188dataJavaApache-2.0baidu/Unlimited-OCR
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
AI summary: A system designed for one-shot, long-horizon parsing in OCR tasks.
22,391ai-mlPythonMITeigent-ai/eigent
Eigent: The Open Source Cowork Desktop - Local and Free Alternative to Claude Cowork and Codex
AI summary: An intelligent data extraction and transformation pipeline for unstructured enterprise documents.
14,783dataTypeScriptApache-2.0rahulnyk/knowledge_graph
Convert any text to a graph of knowledge. This can be used for Graph Augmented Generation or Knowledge Graph based QnA
AI summary: A tool to convert unstructured text into knowledge graphs for enhanced QnA and data generation.
3,621dataJupyter NotebookMITjackwener/xiaohongshu-cli
A CLI for Xiaohongshu (小红书) — search, read, interact via reverse-engineered API
AI summary: Command-line interface tool for interacting with Xiaohongshu (Little Red Book).
2,463developer-toolsPythontnm/zclaw
Your personal AI assistant at all-in 888KiB (~35KB in app code). Running on an ESP32. GPIO, cron, custom tools, memory, and more.
AI summary: ZClaw is a highly concurrent, distributed web scraper built for massive data extraction.
2,203dataCMITpeteromallet/dataclaw
Agent harness to publish your agent chat history as Huggingface datasets.
AI summary: A lightweight Python library for declarative data extraction and transformation.
2,109dataPythonMIT