google/langextractPublic

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

AI summary: AI-powered tool for extracting structured data from unstructured text using large language models.

Stars
38K
+13 today
Forks
2.6K
Watchers
168
Open issues
75
Open PRs
45
Contributors
~23
Commits
173
Branches
67

PythonApache-2.0Created Jul 8, 2025Last push 12d agoLatest release v1.6.0+53 stars this week+71 this month

Star history

since Jul 6, 2025
010K20K30KJul 2025Nov 2025Mar 2026Aug 2026
38K stars as of Aug 7, 2026, tracked back to Jul 6, 2025. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 3 commits2025-08-11: 1 commit2025-08-12: 0 commits2025-08-13: 9 commits2025-08-14: 10 commits2025-08-15: 3 commits2025-08-16: 0 commits2025-08-17: 3 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 1 commit2025-08-21: 2 commits2025-08-22: 1 commit2025-08-23: 2 commits2025-08-24: 3 commits2025-08-25: 4 commits2025-08-26: 0 commits2025-08-27: 2 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 3 commits2025-09-01: 3 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 1 commit2025-09-05: 1 commit2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 1 commit2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 1 commit2025-09-17: 0 commits2025-09-18: 1 commit2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 1 commit2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 1 commit2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 2 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 1 commit2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 3 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 1 commit2025-11-19: 0 commits2025-11-20: 1 commit2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 2 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 1 commit2025-12-28: 1 commit2025-12-29: 1 commit2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 1 commit2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 4 commits2026-03-22: 2 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 2 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 1 commit2026-04-06: 0 commits2026-04-07: 1 commit2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 1 commit2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 1 commit2026-04-15: 0 commits2026-04-16: 2 commits2026-04-17: 0 commits2026-04-18: 1 commit2026-04-19: 2 commits2026-04-20: 0 commits2026-04-21: 4 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 2 commits2026-04-29: 4 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 1 commit2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 2 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 3 commits2026-05-20: 1 commit2026-05-21: 1 commit2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 3 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 7 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
116 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    37,990 stars

  • Well documented

    High community health score

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    10 trending appearances

What langextract does

LangExtract is an advanced data extraction framework developed by Google that leverages large language models (LLMs) to pull structured entities and relationships from unstructured text. It simplifies the process of converting complex documents, such as full texts of literature or detailed radiology reports, into organized JSON formats. The tool abstracts away the complexities of prompt engineering and model interactions, allowing developers to define schemas and extract accurate, structured data seamlessly. It supports various providers, including Google Cloud, OpenAI, and local LLMs through Ollama, providing flexibility in model choice.

LangExtract is designed for data scientists, NLP engineers, and backend developers who need to convert unstructured text into structured formats reliably. It is highly beneficial for teams looking to integrate LLM-based extraction into their data pipelines without managing model intricacies.

  • LLM-powered extraction: Uses state-of-the-art language models to accurately identify and extract complex data structures.
  • Schema-driven parsing: Allows developers to define precise output formats (e.g., JSON) to ensure consistent data extraction.
  • Multi-provider compatibility: Works out-of-the-box with cloud APIs like OpenAI and Google, as well as local models via Ollama.
  • Complex document handling: Capable of processing lengthy texts, such as entire books or detailed medical reports.
  • Custom provider integration: Provides an extensible architecture for adding support for new or proprietary language models.

Where teams use it

Medical record structuring

Healthcare researchers use it to extract specific diagnoses and medications from unstructured radiology reports automatically.

Literature analysis

Data scientists process full texts of novels to extract character relationships and sentiment timelines into structured databases.

Automated document digitization

Enterprises convert thousands of unstructured contracts into queryable databases without writing complex regex patterns.

Private data extraction

Security-conscious organizations run the extraction pipeline locally using Ollama to ensure sensitive data never leaves their network.

Getting started: pip install langextract

README

main branch

LangExtract Logo

LangExtract

PyPI version GitHub stars Tests DOI Live demo

Table of Contents

Introduction

LangExtract is a Python library that uses LLMs to extract structured information from unstructured text documents based on user-defined instructions. It processes materials such as clinical notes or reports, identifying and organizing key details while ensuring the extracted data corresponds to the source text.

LangExtract end to end: unstructured text is chunked, extracted in parallel by an LLM, and every extracted value is grounded back to its exact character span in the source

Try the live demo →
Run grounded extraction on Romeo and Juliet in your browser, no install required.

Why LangExtract?

  1. Precise Source Grounding: Maps every extraction to its exact location in the source text, enabling visual highlighting for easy traceability and verification.
  2. Reliable Structured Outputs: Enforces a consistent output schema based on your few-shot examples, leveraging controlled generation in supported models like Gemini to guarantee robust, structured results.
  3. Optimized for Long Documents: Overcomes the "needle-in-a-haystack" challenge of large document extraction by using an optimized strategy of text chunking, parallel processing, and multiple passes for higher recall.
  4. Interactive Visualization: Instantly generates a self-contained, interactive HTML file to visualize and review thousands of extracted entities in their original context.
  5. Flexible LLM Support: Supports your preferred models, from cloud-based LLMs like the Google Gemini family to local open-source models via the built-in Ollama interface.
  6. Adaptable to Any Domain: Define extraction tasks for any domain using just a few examples. LangExtract adapts to your needs without requiring any model fine-tuning.
  7. Leverages LLM World Knowledge: Utilize precise prompt wording and few-shot examples to influence how the extraction task may utilize LLM knowledge. The accuracy of any inferred information and its adherence to the task specification are contingent upon the selected LLM, the complexity of the task, the clarity of the prompt instructions, and the nature of the prompt examples.

Quick Start

Note: Using cloud-hosted models like Gemini requires an API key. See the API Key Setup section for instructions on how to get and configure your key.

Extract structured information with just a few lines of code.

1. Define Your Extraction Task

First, create a prompt that clearly describes what you want to extract. Then, provide a high-quality example to guide the model.

import langextract as lx
import textwrap

# 1. Define the prompt and extraction rules
prompt = textwrap.dedent("""\
    Extract characters, emotions, and relationships in order of appearance.
    Use exact text for extractions. Do not paraphrase or overlap entities.
    Provide meaningful attributes for each entity to add context.""")

# 2. Provide a high-quality example to guide the model
examples = [
    lx.data.ExampleData(
        text="ROMEO. But soft! What light through yonder window breaks? It is the east, and Juliet is the sun.",
        extractions=[
            lx.data.Extraction(
                extraction_class="character",
                extraction_text="ROMEO",
                attributes={"emotional_state": "wonder"}
            ),
            lx.data.Extraction(
                extraction_class="emotion",
                extraction_text="But soft!",
                attributes={"feeling": "gentle awe"}
            ),
            lx.data.Extraction(
                extraction_class="relationship",
                extraction_text="Juliet is the sun",
                attributes={"type": "metaphor"}
            ),
        ]
    )
]

Note: Examples drive model behavior. Each extraction_text should ideally be verbatim from the example's text (no paraphrasing), listed in order of appearance. LangExtract raises Prompt alignment warnings by default if examples don't follow this pattern—resolve these for best results.

Grounding: LLMs may occasionally extract content from few-shot examples rather than the input text. LangExtract automatically detects this: extractions that cannot be located in the source text will have char_interval = None. Filter these out with [e for e in result.extractions if e.char_interval] to keep only grounded results.

2. Run the Extraction

Provide your input text and the prompt materials to the lx.extract function.

# The input text to be processed
input_text = "Lady Juliet gazed longingly at the stars, her heart aching for Romeo"

# Run the extraction
result = lx.extract(
    text_or_documents=input_text,
    prompt_description=prompt,
    examples=examples,
    model_id="gemini-3.5-flash",
)

For advanced constraints beyond examples, such as enum values on extraction attributes, Gemini and OpenAI support output_schema with or without few-shot examples. See Custom output schemas.

Model Selection: gemini-3.5-flash is the recommended default, offering strong extraction quality for LangExtract's schema-constrained workflows. For high-volume or cost-sensitive workloads, consider the current stable Flash-Lite model, gemini-3.1-flash-lite; for highly complex tasks requiring deeper reasoning, evaluate a current Gemini Pro model from the official model documentation. For large-scale or production use, a paid Gemini tier is suggested to increase throughput and avoid rate limits. See the rate-limit documentation for details.

Model Lifecycle: Note that Gemini models have a lifecycle with defined retirement dates. Users should consult the official model version documentation to stay informed about the latest stable and legacy versions.

3. Visualize the Results

The extractions can be saved to a .jsonl file, a popular format for working with language model data. LangExtract can then generate an interactive HTML visualization from this file to review the entities in context.

# Save the results to a JSONL file
lx.io.save_annotated_documents([result], output_name="extraction_results.jsonl", output_dir=".")

# Generate the visualization from the file
html_content = lx.visualize("extraction_results.jsonl")
with open("visualization.html", "w") as f:
    if hasattr(html_content, 'data'):
        f.write(html_content.data)  # For Jupyter/Colab
    else:
        f.write(html_content)

This creates an animated and interactive HTML file:

Romeo and Juliet Basic Visualization

Note on LLM Knowledge Utilization: This example demonstrates extractions that stay close to the text evidence - extracting "longing" for Lady Juliet's emotional state and identifying "yearning" from "gazed longingly at the stars." The task could be modified to generate attributes that draw more heavily from the LLM's world knowledge (e.g., adding "identity": "Capulet family daughter" or "literary_context": "tragic heroine"). The balance between text-evidence and knowledge-inference is controlled by your prompt instructions and example attributes.

Scaling to Longer Documents

For larger texts, you can process entire documents directly from URLs with parallel processing and enhanced sensitivity:

# Process Romeo & Juliet directly from Project Gutenberg
result = lx.extract(
    text_or_documents="https://www.gutenberg.org/files/1513/1513-0.txt",
    prompt_description=prompt,
    examples=examples,
    model_id="gemini-3.5-flash",
    extraction_passes=3,    # Improves recall through multiple passes
    max_workers=20,         # Parallel processing for speed
    max_char_buffer=1000    # Smaller contexts for better accuracy
)

This approach can extract hundreds of entities from full novels while maintaining high accuracy. The interactive visualization seamlessly handles large result sets, making it easy to explore hundreds of entities from the output JSONL file. See the full Romeo and Juliet extraction example → for detailed results and performance insights.

Vertex AI Batch Processing

Save costs on large-scale tasks by enabling Vertex AI Batch API with language_model_params that include vertexai=True, project, location, and a batch config.

See an example of the Vertex AI Batch API usage in this example.

Installation

From PyPI

pip install langextract

Recommended for most users. For isolated environments, consider using a virtual environment:

python -m venv langextract_env
source langextract_env/bin/activate  # On Windows: langextract_env\Scripts\activate
pip install langextract

From Source

LangExtract uses modern Python packaging with pyproject.toml for dependency management:

Installing with -e puts the package in development mode, allowing you to modify the code without reinstalling.

git clone https://github.com/google/langextract.git
cd langextract

# For basic installation:
pip install -e .

# For development (includes linting tools):
pip install -e ".[dev]"

# For testing (includes pytest):
pip install -e ".[test]"

Docker

docker build -t langextract .
docker run --rm -e LANGEXTRACT_API_KEY="your-api-key" langextract python your_script.py

API Key Setup for Cloud Models

When using LangExtract with cloud-hosted models (like Gemini or OpenAI), you'll need to set up an API key. On-device models don't require an API key. For developers using local LLMs, LangExtract offers built-in support for Ollama and can be extended to other third-party APIs by updating the inference endpoints.

API Key Sources

Get API keys from:

Setting up API key in your environment

Option 1: Environment Variable

export LANGEXTRACT_API_KEY="your-api-key-here"

Option 2: .env File (Recommended)

Add your API key to a .env file:

# Add API key to .env file
cat >> .env << 'EOF'
LANGEXTRACT_API_KEY=your-api-key-here
EOF

# Keep your API key secure
echo '.env' >> .gitignore

In your Python code:

import langextract as lx

result = lx.extract(
    text_or_documents=input_text,
    prompt_description="Extract information...",
    examples=[...],
    model_id="gemini-3.5-flash"
)

Option 3: Direct API Key (Not Recommended for Production)

You can also provide the API key directly in your code, though this is not recommended for production use:

result = lx.extract(
    text_or_documents=input_text,
    prompt_description="Extract information...",
    examples=[...],
    model_id="gemini-3.5-flash",
    api_key="your-api-key-here"  # Only use this for testing/development
)

Option 4: Vertex AI (Service Accounts)

Use Vertex AI for authentication with service accounts:

result = lx.extract(
    text_or_documents=input_text,
    prompt_description="Extract information...",
    examples=[...],
    model_id="gemini-3.5-flash",
    language_model_params={
        "vertexai": True,
        "project": "your-project-id",
        "location": "global"  # or regional endpoint
    }
)

Adding Custom Model Providers

LangExtract supports custom LLM providers via a lightweight plugin system. You can add support for new models without changing core code.

  • Add new model support independently of the core library
  • Distribute your provider as a separate Python package
  • Keep custom dependencies isolated
  • Override or extend built-in providers via priority-based resolution

See the detailed guide in Provider System Documentation to learn how to:

  • Register a provider with @router.register(...) from langextract.providers
  • Publish an entry point for discovery
  • Optionally provide a schema with get_schema_class() for structured output
  • Integrate with the factory via create_model(...)

Using OpenAI Models

LangExtract supports OpenAI models (requires optional dependency: pip install langextract[openai]):

import langextract as lx

# OPENAI_API_KEY in the environment is picked up automatically; pass
# api_key=... explicitly only if you need to override it.
result = lx.extract(
    text_or_documents=input_text,
    prompt_description=prompt,
    examples=examples,
    model_id="gpt-4o",  # Automatically selects OpenAI provider
)

The OpenAI provider uses structured outputs or JSON mode and auto-determines fence behavior — leave fence_output and use_schema_constraints unset. output_schema is also supported for OpenAI models that support structured outputs; provide a LangExtract output-envelope JSON schema, preferably with the lx.schema helpers.

For large, non-latency-sensitive OpenAI workloads, enable the OpenAI Batch API with language_model_params. Batch mode is opt-in and falls back to realtime calls when the prompt count is below the configured threshold.

result = lx.extract(
    text_or_documents=documents,
    prompt_description=prompt,
    examples=examples,
    model_id="gpt-4o-mini",
    language_model_params={
        "batch": {
            "enabled": True,
            "threshold": 50,
            "poll_interval": 10,
        }
    },
)

For OpenAI-compatible endpoints or non-GPT model IDs (which skip auto-routing), use ModelConfig with an explicit provider:

from langextract.factory import ModelConfig

result = lx.extract(
    text_or_documents=input_text,
    prompt_description=prompt,
    examples=examples,
    config=ModelConfig(
        model_id="my-openai-compatible-model",
        provider="openai",
        provider_kwargs={"api_key": "sk-...", "base_url": "https://..."},
    ),
)

Using Local LLMs with Ollama

LangExtract supports local inference using Ollama, allowing you to run models without API keys:

import langextract as lx

result = lx.extract(
    text_or_documents=input_text,
    prompt_description=prompt,
    examples=examples,
    model_id="gemma2:2b",  # Automatically selects Ollama provider
    model_url="http://localhost:11434",
)

The Ollama provider exposes FormatModeSchema for JSON mode. Leave fence_output and use_schema_constraints unset so the factory auto-configures from the provider's schema. Ollama does not currently support output_schema.

Quick setup: Install Ollama from ollama.com, run ollama pull gemma2:2b, then ollama serve.

For detailed installation, Docker setup, and examples, see examples/ollama/.

More Examples

Additional examples of LangExtract in action:

Romeo and Juliet Full Text Extraction

LangExtract can process complete documents directly from URLs. This example demonstrates extraction from the full text of Romeo and Juliet from Project Gutenberg (147,843 characters), showing parallel processing, sequential extraction passes, and performance optimization for long document processing.

View Romeo and Juliet Full Text Example →

Medication Extraction

Disclaimer: This demonstration is for illustrative purposes of LangExtract's baseline capability only. It does not represent a finished or approved product, is not intended to diagnose or suggest treatment of any disease or condition, and should not be used for medical advice.

LangExtract excels at extracting structured medical information from clinical text. These examples demonstrate both basic entity recognition (medication names, dosages, routes) and relationship extraction (connecting medications to their attributes), showing LangExtract's effectiveness for healthcare applications.

View Medication Examples →

Radiology Report Structuring: RadExtract

Explore RadExtract, a live interactive demo on HuggingFace Spaces that shows how LangExtract can automatically structure radiology reports. Try it directly in your browser with no setup required.

View RadExtract Demo →

Community Providers

Extend LangExtract with custom model providers! Check out our Community Provider Plugins registry to discover providers created by the community or add your own.

For detailed instructions on creating a provider plugin, see the Custom Provider Plugin Example.

Contributing

Contributions are welcome! See CONTRIBUTING.md to get started with development, testing, and pull requests. You must sign a Contributor License Agreement before submitting patches.

Thanks to everyone who has contributed.

Testing

To run tests locally from the source:

# Clone the repository
git clone https://github.com/google/langextract.git
cd langextract

# Install with test dependencies
pip install -e ".[test]"

# Run all tests
pytest tests

Or reproduce the full CI matrix locally with tox:

tox  # runs pylint + pytest on Python 3.10 and 3.11

Ollama Integration Testing

If you have Ollama installed locally, you can run integration tests:

# Test Ollama integration (requires Ollama running with gemma2:2b model)
tox -e ollama-integration

This test will automatically detect if Ollama is available and run real inference tests.

Development

Code Formatting

This project uses automated formatting tools to maintain consistent code style:

# Auto-format all code
./autoformat.sh

# Or run formatters separately
isort langextract tests --profile google --line-length 80
pyink langextract tests --config pyproject.toml

Pre-commit Hooks

For automatic formatting checks:

pre-commit install  # One-time setup
pre-commit run --all-files  # Manual run

Linting

Run linting before submitting PRs:

pylint --rcfile=.pylintrc langextract tests

See CONTRIBUTING.md for full development guidelines.

How to Cite

If you use LangExtract in your research, please cite it:

@software{goel_langextract,
  author  = {Goel, Akshay},
  title   = {{LangExtract}},
  year    = {2026},
  version = {1.6.0},
  doi     = {10.5281/zenodo.21126643},
  url     = {https://github.com/google/langextract}
}

Cite the version you used — each release has its own DOI on Zenodo. If your style rejects @software, use @misc.

Disclaimer

This is not an officially supported Google product. If you use LangExtract in production or publications, please cite accordingly and acknowledge usage. Use is subject to the Apache 2.0 License. For health-related applications, use of LangExtract is also subject to the Health AI Developer Foundations Terms of Use.


Happy Extracting!

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

18 total
  1. v1.6.0v1.6.0Jul 2, 2026

    ## Highlights - Add user-provided `output_schema` support for Gemini and OpenAI. - Add Ollama GPT-OSS JSON chat support. **Full Changelog**: https://github.com/google/langextract/compare/v1.5.0...v1.6.0

  2. v1.5.0v1.5.0May 20, 2026

    ## Highlights - Add OpenAI Batch API support. - Update the default Gemini Flash model to `gemini-3.5-flash`. **Full Changelog**: https://github.com/google/langextract/compare/v1.4.0...v1.5.0

  3. v1.4.0v1.4.0May 15, 2026

    ## Highlights - Add OpenAI structured output schema support. - Fix Ollama compatibility for newer thinking-capable models. - Apply `additional_context` consistently for `Document` inputs. **Full Changelog**: https://github.com/google/langextract/compare/v1.3.0...v1.4.0

  4. v1.3.0v1.3.0Apr 29, 2026

    ## What's New ### Features - Add automatic retry logic for transient Gemini API errors (503, 429) with jittered exponential backoff and keyword-only retry knobs (#385, fixes #240) ### Security - Make URL fetching opt-in by default (#449). URLs in input text are no longer auto-fetched; pass an explicit flag to enable. ### Performance - Replace difflib fuzzy aligner with an O(n·m²) LCS DP for significantly faster alignment on large extractions (#442) ### Bug Fixes - Close progress bars cleanly when save/download fails (#434) - Make `IssueKind` public and export it in `__all__` (#426) ### Documentation - Note that `output_name` is not sanitized in `save_annotated_documents` (#451) - Add `langextract-usage` Agent Skill (#448) **Full Changelog**: https://github.com/google/langextract/compare/v1.2.1...v1.3.0

  5. v1.2.1v1.2.1Apr 8, 2026

    ## What's New ### Bug Fixes - Pass `reasoning_effort` directly to OpenAI API as a top-level Chat Completions parameter (#429) - Fixes `unexpected keyword argument 'reasoning'` error when using reasoning models (o1, o3, o4-mini, gpt-5) - Suppress schema errors in `resolve()` when `suppress_parse_errors=True` (#435) - Extends suppression to `ValueError` from malformed-but-parseable LLM output, not just `FormatError` - Fix Ollama Docker example healthcheck (#360) - Replaces `curl` (not present in image) with `ollama list` **Full Changelog**: https://github.com/google/langextract/compare/v1.2.0...v1.2.1

Code frequency

additions and deletions
+6.5K-6.5KWeek of 2025-08-10: +6,481 linesWeek of 2025-08-10: -831 linesWeek of 2025-08-17: +4,269 linesWeek of 2025-08-17: -2,395 linesWeek of 2025-08-24: +359 linesWeek of 2025-08-24: -27 linesWeek of 2025-08-31: +1,273 linesWeek of 2025-08-31: -13 linesWeek of 2025-09-07: +2,470 linesWeek of 2025-09-07: -643 linesWeek of 2025-09-14: +38 linesWeek of 2025-09-14: -20 linesWeek of 2025-09-21: +1 linesWeek of 2025-09-21: -0 linesWeek of 2025-09-28: +48 linesWeek of 2025-09-28: -1 linesWeek of 2025-10-05: +0 linesWeek of 2025-10-05: -0 linesWeek of 2025-10-12: +0 linesWeek of 2025-10-12: -0 linesWeek of 2025-10-19: +0 linesWeek of 2025-10-19: -0 linesWeek of 2025-10-26: +1,455 linesWeek of 2025-10-26: -46 linesWeek of 2025-11-02: +229 linesWeek of 2025-11-02: -122 linesWeek of 2025-11-09: +2,147 linesWeek of 2025-11-09: -183 linesWeek of 2025-11-16: +1,497 linesWeek of 2025-11-16: -177 linesWeek of 2025-11-23: +134 linesWeek of 2025-11-23: -14 linesWeek of 2025-11-30: +0 linesWeek of 2025-11-30: -0 linesWeek of 2025-12-07: +0 linesWeek of 2025-12-07: -0 linesWeek of 2025-12-14: +0 linesWeek of 2025-12-14: -0 linesWeek of 2025-12-21: +135 linesWeek of 2025-12-21: -16 linesWeek of 2025-12-28: +608 linesWeek of 2025-12-28: -12 linesWeek of 2026-01-04: +0 linesWeek of 2026-01-04: -0 linesWeek of 2026-01-11: +0 linesWeek of 2026-01-11: -0 linesWeek of 2026-01-18: +0 linesWeek of 2026-01-18: -0 linesWeek of 2026-01-25: +0 linesWeek of 2026-01-25: -0 linesWeek of 2026-02-01: +0 linesWeek of 2026-02-01: -0 linesWeek of 2026-02-08: +0 linesWeek of 2026-02-08: -0 linesWeek of 2026-02-15: +0 linesWeek of 2026-02-15: -0 linesWeek of 2026-02-22: +6 linesWeek of 2026-02-22: -2 linesWeek of 2026-03-01: +0 linesWeek of 2026-03-01: -0 linesWeek of 2026-03-08: +0 linesWeek of 2026-03-08: -0 linesWeek of 2026-03-15: +147 linesWeek of 2026-03-15: -17 linesWeek of 2026-03-22: +169 linesWeek of 2026-03-22: -10 linesWeek of 2026-03-29: +76 linesWeek of 2026-03-29: -43 linesWeek of 2026-04-05: +754 linesWeek of 2026-04-05: -14 linesWeek of 2026-04-12: +1,987 linesWeek of 2026-04-12: -268 linesWeek of 2026-04-19: +922 linesWeek of 2026-04-19: -54 linesWeek of 2026-04-26: +866 linesWeek of 2026-04-26: -156 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +1,164 linesWeek of 2026-05-10: -94 linesWeek of 2026-05-17: +2,178 linesWeek of 2026-05-17: -575 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +0 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +2,315 linesWeek of 2026-06-28: -83 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +0 linesWeek of 2026-07-12: -0 linesWeek of 2026-07-19: +548 linesWeek of 2026-07-19: -29 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesAug 10, 2025Aug 2, 2026
+32.3K lines added, -5.8K removed over the last year.

Commits per week

last 52 weeks
260Week of 2025-08-10: 26 commitsWeek of 2025-08-17: 9 commitsWeek of 2025-08-24: 9 commitsWeek of 2025-08-31: 8 commitsWeek of 2025-09-07: 1 commitsWeek of 2025-09-14: 2 commitsWeek of 2025-09-21: 1 commitsWeek of 2025-09-28: 1 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 2 commitsWeek of 2025-11-02: 1 commitsWeek of 2025-11-09: 3 commitsWeek of 2025-11-16: 2 commitsWeek of 2025-11-23: 2 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 1 commitsWeek of 2025-12-28: 2 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 1 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 4 commitsWeek of 2026-03-22: 2 commitsWeek of 2026-03-29: 2 commitsWeek of 2026-04-05: 3 commitsWeek of 2026-04-12: 4 commitsWeek of 2026-04-19: 6 commitsWeek of 2026-04-26: 7 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 2 commitsWeek of 2026-05-17: 5 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 3 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 7 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 0 commitsAug 10, 2025Aug 2, 2026
116 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 1 commitsSun 2:00 — 0 commitsSun 3:00 — 1 commitsSun 4:00 — 0 commitsSun 5:00 — 4 commitsSun 6:00 — 1 commitsSun 7:00 — 0 commitsSun 8:00 — 1 commitsSun 9:00 — 1 commitsSun 10:00 — 1 commitsSun 11:00 — 0 commitsSun 12:00 — 1 commitsSun 13:00 — 1 commitsSun 14:00 — 0 commitsSun 15:00 — 4 commitsSun 16:00 — 0 commitsSun 17:00 — 1 commitsSun 18:00 — 2 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 2 commitsSun 22:00 — 5 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 1 commitsMon 11:00 — 2 commitsMon 12:00 — 0 commitsMon 13:00 — 1 commitsMon 14:00 — 1 commitsMon 15:00 — 3 commitsMon 16:00 — 2 commitsMon 17:00 — 2 commitsMon 18:00 — 1 commitsMon 19:00 — 0 commitsMon 20:00 — 0 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 0 commitsTue 0:00 — 1 commitsTue 1:00 — 2 commitsTue 2:00 — 1 commitsTue 3:00 — 0 commitsTue 4:00 — 1 commitsTue 5:00 — 0 commitsTue 6:00 — 1 commitsTue 7:00 — 0 commitsTue 8:00 — 2 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 1 commitsTue 12:00 — 1 commitsTue 13:00 — 0 commitsTue 14:00 — 1 commitsTue 15:00 — 0 commitsTue 16:00 — 1 commitsTue 17:00 — 2 commitsTue 18:00 — 6 commitsTue 19:00 — 2 commitsTue 20:00 — 2 commitsTue 21:00 — 1 commitsTue 22:00 — 0 commitsTue 23:00 — 2 commitsWed 0:00 — 3 commitsWed 1:00 — 0 commitsWed 2:00 — 1 commitsWed 3:00 — 1 commitsWed 4:00 — 1 commitsWed 5:00 — 3 commitsWed 6:00 — 2 commitsWed 7:00 — 0 commitsWed 8:00 — 1 commitsWed 9:00 — 2 commitsWed 10:00 — 1 commitsWed 11:00 — 1 commitsWed 12:00 — 0 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 1 commitsWed 16:00 — 0 commitsWed 17:00 — 0 commitsWed 18:00 — 1 commitsWed 19:00 — 2 commitsWed 20:00 — 1 commitsWed 21:00 — 1 commitsWed 22:00 — 2 commitsWed 23:00 — 4 commitsThu 0:00 — 5 commitsThu 1:00 — 3 commitsThu 2:00 — 5 commitsThu 3:00 — 3 commitsThu 4:00 — 1 commitsThu 5:00 — 1 commitsThu 6:00 — 0 commitsThu 7:00 — 5 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 0 commitsThu 12:00 — 0 commitsThu 13:00 — 0 commitsThu 14:00 — 2 commitsThu 15:00 — 0 commitsThu 16:00 — 2 commitsThu 17:00 — 0 commitsThu 18:00 — 1 commitsThu 19:00 — 1 commitsThu 20:00 — 1 commitsThu 21:00 — 1 commitsThu 22:00 — 1 commitsThu 23:00 — 0 commitsFri 0:00 — 1 commitsFri 1:00 — 2 commitsFri 2:00 — 0 commitsFri 3:00 — 4 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 1 commitsFri 7:00 — 2 commitsFri 8:00 — 1 commitsFri 9:00 — 0 commitsFri 10:00 — 0 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 1 commitsFri 14:00 — 0 commitsFri 15:00 — 0 commitsFri 16:00 — 0 commitsFri 17:00 — 2 commitsFri 18:00 — 0 commitsFri 19:00 — 2 commitsFri 20:00 — 0 commitsFri 21:00 — 1 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 1 commitsSat 1:00 — 4 commitsSat 2:00 — 2 commitsSat 3:00 — 1 commitsSat 4:00 — 5 commitsSat 5:00 — 1 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 1 commitsSat 14:00 — 2 commitsSat 15:00 — 1 commitsSat 16:00 — 0 commitsSat 17:00 — 1 commitsSat 18:00 — 0 commitsSat 19:00 — 2 commitsSat 20:00 — 1 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Feb 13, 2026daily#5+206
Feb 12, 2026daily#3+269
Feb 11, 2026daily#3+605
Feb 10, 2026daily#3+595
Feb 9, 2026daily#5+331
Feb 8, 2026daily#22+132
Jan 20, 2026daily#16+181
Jan 19, 2026daily#20+193
Jan 17, 2026daily#17+162
Jan 16, 2026daily#21+134
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    277.1K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    238.5K stars · JavaScript

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    234.7K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    227K stars · Python