VectifyAI/PageIndexPublic

📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG

AI summary: A document index designed specifically for vectorless, reasoning-based Retrieval-Augmented Generation.

Stars
38.6K
+234 today
Forks
3.3K
Watchers
158
Open issues
37
Open PRs
82
Contributors
~16
Commits
484
Branches
54

PythonMITCreated Apr 1, 2025Last push todayLatest release v0.2.21+2.8K stars this week+3.1K this month

Quick answers

What is PageIndex?
A document index designed specifically for vectorless, reasoning-based Retrieval-Augmented Generation.
What does PageIndex do?
PageIndex is a novel document indexing system tailored for vectorless, reasoning-based Retrieval-Augmented Generation (RAG) architectures. Unlike traditional RAG systems that rely on vector embeddings and similarity search, PageIndex focuses on preserving the structural and contextual integrity of documents. It prepares data in a way that allows large language models to reason over the content directly, rather than just matching semantic similarity. This approach aims to reduce hallucinations and improve the accuracy of answers by providing the LLM with a more coherent representation of the source material. It represents a shift towards utilizing the inherent reasoning capabilities of advanced LLMs over traditional retrieval techniques.
Who is PageIndex for?
PageIndex is designed for AI engineers and developers building next-generation RAG applications. It is for those looking to move beyond the limitations of traditional vector-based retrieval and leverage the reasoning power of modern LLMs.
How do I get started with PageIndex?
pip install pageindex
How popular is PageIndex on GitHub?
VectifyAI/PageIndex has 38,607 stars and 3,344 forks on GitHub, and gained 2,753 stars in the last 7 days.
What license does PageIndex use?
VectifyAI/PageIndex is released under the MIT license.

Star history

since Jul 28, 2026
010K20K30KJul 2026Aug 2026Sep 2026Oct 2026
38.6K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepOctMonWedFri2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 2 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 1 commit2025-10-28: 1 commit2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 2 commits2025-11-02: 3 commits2025-11-03: 2 commits2025-11-04: 0 commits2025-11-05: 15 commits2025-11-06: 1 commit2025-11-07: 2 commits2025-11-08: 0 commits2025-11-09: 1 commit2025-11-10: 0 commits2025-11-11: 1 commit2025-11-12: 0 commits2025-11-13: 2 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 1 commit2025-11-18: 1 commit2025-11-19: 3 commits2025-11-20: 1 commit2025-11-21: 2 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 1 commit2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 5 commits2025-12-20: 0 commits2025-12-21: 1 commit2025-12-22: 1 commit2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 1 commit2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 2 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 1 commit2026-01-25: 2 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 1 commit2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 2 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 7 commits2026-03-03: 0 commits2026-03-04: 1 commit2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 1 commit2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 1 commit2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 1 commit2026-03-27: 5 commits2026-03-28: 3 commits2026-03-29: 4 commits2026-03-30: 1 commit2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 1 commit2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 1 commit2026-04-24: 0 commits2026-04-25: 2 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 1 commit2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 1 commit2026-05-06: 2 commits2026-05-07: 1 commit2026-05-08: 2 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 1 commit2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 2 commits2026-05-31: 0 commits2026-06-01: 1 commit2026-06-02: 2 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 2 commits2026-06-06: 1 commit2026-06-07: 0 commits2026-06-08: 1 commit2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 1 commit2026-06-17: 3 commits2026-06-18: 1 commit2026-06-19: 1 commit2026-06-20: 1 commit2026-06-21: 0 commits2026-06-22: 1 commit2026-06-23: 2 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 1 commit2026-06-27: 0 commits2026-06-28: 1 commit2026-06-29: 0 commits2026-06-30: 1 commit2026-07-01: 0 commits2026-07-02: 1 commit2026-07-03: 4 commits2026-07-04: 0 commits2026-07-05: 1 commit2026-07-06: 1 commit2026-07-07: 0 commits2026-07-08: 1 commit2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 1 commit2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 1 commit2026-07-16: 3 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 3 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 10 commits2026-07-31: 1 commit2026-08-01: 1 commit2026-08-02: 5 commits2026-08-03: 3 commits2026-08-04: 4 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 1 commit2026-08-11: 1 commit2026-08-12: 0 commits2026-08-13: 3 commits2026-08-14: 0 commits2026-08-15: 2 commits2026-08-16: 0 commits2026-08-17: 4 commits2026-08-18: 0 commits2026-08-19: 4 commits2026-08-20: 2 commits2026-08-21: 1 commit2026-08-22: 1 commit2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 4 commits2026-08-27: 6 commits2026-08-28: 4 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 8 commits2026-09-01: 2 commits2026-09-02: 5 commits2026-09-03: 6 commits2026-09-04: 1 commit2026-09-05: 1 commit2026-09-06: 2 commits2026-09-07: 2 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 6 commits2026-09-11: 1 commit2026-09-12: 0 commits2026-09-13: 2 commits2026-09-14: 0 commits2026-09-15: 2 commits2026-09-16: 1 commit2026-09-17: 11 commits2026-09-18: 5 commits2026-09-19: 1 commit2026-09-20: 3 commits2026-09-21: 7 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 2 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 1 commit2026-09-29: 1 commit2026-09-30: 1 commit2026-10-01: 3 commits2026-10-02: 0 commits2026-10-03: 0 commits2026-10-04: 0 commits2026-10-05: 0 commits2026-10-06: 0 commits2026-10-07: 0 commits2026-10-08: 0 commits2026-10-09: 0 commits2026-10-10: 0 commits
263 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    38,607 stars

  • Rising fast

    +2,753 stars this week

  • Actively maintained

    Pushed within 48 hours

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    14 trending appearances

What PageIndex does

PageIndex is a novel document indexing system tailored for vectorless, reasoning-based Retrieval-Augmented Generation (RAG) architectures. Unlike traditional RAG systems that rely on vector embeddings and similarity search, PageIndex focuses on preserving the structural and contextual integrity of documents. It prepares data in a way that allows large language models to reason over the content directly, rather than just matching semantic similarity. This approach aims to reduce hallucinations and improve the accuracy of answers by providing the LLM with a more coherent representation of the source material. It represents a shift towards utilizing the inherent reasoning capabilities of advanced LLMs over traditional retrieval techniques.

PageIndex is designed for AI engineers and developers building next-generation RAG applications. It is for those looking to move beyond the limitations of traditional vector-based retrieval and leverage the reasoning power of modern LLMs.

  • Vectorless architecture: Eliminates the need for generating and storing expensive vector embeddings for document retrieval.
  • Reasoning-optimized indexing: Structures document data specifically to enhance the reasoning capabilities of large language models.
  • Context preservation: Maintains the original structural and semantic context of documents better than chunk-based vector methods.
  • Simplified infrastructure: Reduces the complexity of RAG pipelines by removing the dependency on dedicated vector databases.
  • Enhanced accuracy: Aims to provide more precise and contextually aware answers by allowing LLMs to analyze full document structures.

Where teams use it

Advanced reasoning RAG

Developers use it to build question-answering systems that require complex reasoning over documents rather than simple fact retrieval.

Vector database alternative

Teams use it to simplify their AI infrastructure stack by eliminating the need to manage and scale vector databases.

Context-heavy document analysis

Applications use it to process long-form documents like legal contracts where maintaining structural context is critical.

Hallucination reduction

Projects employ it to ground LLM responses more firmly in the source text, reducing the likelihood of generated hallucinations.

Getting started: pip install pageindex

README

main branch
pi_github_banner_low

VectifyAI%2FPageIndex | Trendshift

PageIndex: Vectorless, Reasoning-based RAG

Reasoning-based RAG  ◦  No Vector DB, No Chunking  ◦  Context-Aware Retrieval  ◦  Reads Like a Human

🌐 Website  •   ☁️ Cloud  •   📖 Docs  •   📝 Blog  •   ✉️ Contact 

Updates

  • [Aug '26] 🔥 PageIndex SDK: pip install -U pageindex now ships local mode: index, retrieve, and chat entirely on your machine with your own LLM key, or point the same client at PageIndex Cloud with an API key.
  • [Aug '26] ⚡ PageIndex Flash: fast tree index generation for text-based PDFs, now the default indexing method in PageIndex SDK local mode.
  • Scale PageIndex to Millions of Documents: PageIndex File System is a file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document.
  • PageIndex App: a human-like document analysis agent for long professional documents.

What is PageIndex?

Are you frustrated with vector database retrieval accuracy for long and complex documents? Vector-based RAG retrieves by semantic similarity. But similarity ≠ relevance — what retrieval actually needs is relevance, and relevance requires reasoning. On professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search misses what is relevant but not similar, and returns what is similar but not relevant.

Inspired by AlphaGo, PageIndex replaces the vector index with a hierarchical tree index and lets an LLM reason its way through it, the way a human expert turns to and reads the right section of a long report. Retrieval happens in two steps:

  1. Index: generate a tree-structure index for each document
  2. Retrieve: agentically search that tree with LLM reasoning

TL;DR

PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, with no vector DBs or chunking.

Compare with Vector RAG

Vector RAG PageIndex
Index vector index tree index
Retrieval semantic similarity search LLM reasoning over the tree
Result opaque, “vibe retrieval” traceable to explicit references
Context query embedding only full context: conversation history, domain knowledge, etc.

It is ideal for financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any other long, complex professional document.

Quickstart

pip install -U pageindex
import os
from pageindex import PageIndexClient

os.environ["OPENAI_API_KEY"] = "your-openai-key"

client = PageIndexClient(
    index="gpt-5.6-luna",               # model to build the tree index
    chat="gpt-5.6-sol",                 # model to search the tree
)
doc_id = client.submit_document("report.pdf")["doc_id"]

answer = client.chat("What was the 2023 operating margin?", doc_id=doc_id)
print(answer)

Model Recommendations

  • index=: a basic model is sufficient. The tree structure itself is extracted from the document layout without an LLM; the index model only summarizes and refines it, which a basic model does well.
  • chat=: use the best model you can afford. The chat model searches the tree to retrieve information. See Query cost and accuracy.

Configure other models, streaming, multi-document search, citations, and more.

Drop PageIndex tools into the OpenAI Agents SDK, the Claude Agent SDK, or any other framework.

Benchmarks

Local indexing cost and time

Building a tree locally runs about $0.001 per page with gpt-5.6-luna as the index model, so a 1,000-page textbook costs a little over a dollar and a few minutes, once, and every later question reuses it. PageIndex is designed not to rely heavily on the model used at index time, so in our experiments a basic model does not hurt quality.

Indexing cost against document length, log-log, for nine PDFs from 9 to 1,098 pages. Points track a $0.0011-per-page reference line; the spread around it is text density, not length.

Indexing time also scales predictably with document length. In the same local setup, the benchmark documents (9 to 1,098 pages) finished in roughly 13 seconds to 4.5 minutes.

Indexing time against document length, log-log, for nine PDFs from 9 to 1,098 pages. The measured indexing times range from about 13 seconds to 4.5 minutes and increase predictably with document length.

Query cost and accuracy

PageIndex-OSS-Benchmark measures exactly the setup in the quickstart above (PageIndexClient() in local mode, flash indexing, no OCR) on 62 lookup questions over 34 PDFs (1,945 pages) drawn from MMLongBench-Doc-V2. Every question's answer is a fact stated in running text, so a wrong answer is a retrieval or reading failure, not a reasoning one.

Accuracy against average cost per question. Each model forms a near-vertical reasoning-effort ladder; moving between models costs an order of magnitude a step.

Full results, data, and the runner are in the benchmark repo.

Cost per query vs. native PDF input

The alternative to retrieval is handing the model the whole PDF on every question. That cost grows with the document; PageIndex's does not, because it reads only the nodes its reasoning reaches. On documents where both routes return the same answer, native PDF input costs 2.1× more at 52 pages and 16.6× more at 420 (gpt-5.6-sol, prompt caching excluded) — and at 805 pages the document no longer fits in the context window at all.

Cost per query relative to PageIndex retrieval, for five PDFs from 52 to 805 pages. Passing the PDF natively costs 2.1x, 3.4x, 7.8x, and 16.6x more at 52, 85, 198, and 420 pages; at 805 pages it exceeds the model's context window.

Leading accuracy on FinanceBench

PageIndex reached a state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), vastly outperforming vector-based RAG.

Explore the full FinanceBench evaluation results and the blog post.

PageIndex Cloud

The open-source version is ideal for text-heavy PDFs and local workflows. With PageIndex Cloud, document indexing and storage run in the cloud: PageIndex handles parsing, OCR, image understanding, tree-index construction, and managed storage for you. The chat and retrieval layer remains compatible with your model, so you can search the cloud-hosted index using the model provider your application already uses.

Moving indexing and storage from Local to Cloud only requires a PageIndex API key:

import os
from pageindex import PageIndexClient

os.environ["PAGEINDEX_API_KEY"] = "your-pageindex-key"
os.environ["OPENAI_API_KEY"] = "your-openai-key"

client = PageIndexClient(
    index="cloud",                       # build and store the index in PageIndex Cloud
    chat="gpt-5.6-sol",                  # use your preferred compatible model for chat
)
doc_id = client.submit_document("report.pdf", wait=True)["doc_id"]
print(client.chat("What was the 2023 operating margin?", doc_id=doc_id))
Local Cloud
Handles Text-based PDFs Text-based, scanned, and image-rich documents
Indexing On your machine Managed by PageIndex
Storage Local directory Cloud storage
Citations Page-level Block-level
OCR & image understanding — ✓
Metadata — ✓
Folders — ✓
MCP server — ✓

More About PageIndex Cloud

Ready to Try It?

For dedicated deployment (VPC or on-premises), contact us or book a demo.


⭐ Support Us

Leave us a star 🌟 if you like our project. Thank you!

Please cite this work as:

Mingtian Zhang, Yu Tang and PageIndex Team,
"PageIndex: Next-Generation Vectorless, Reasoning-based RAG",
PageIndex Blog, Sep 2025.
Or use the BibTeX citation.
@article{zhang2025pageindex,
  author = {Mingtian Zhang and Yu Tang and PageIndex Team},
  title = {PageIndex: Next-Generation Vectorless, Reasoning-based RAG},
  journal = {PageIndex Blog},
  year = {2025},
  month = {September},
  note = {https://pageindex.ai/blog/pageindex-intro},
}

© 2026 PageIndex AI

pageindex-wordmark-animated-288-warm
View on GitHub

Recent activity

commits and pull requests

Releases and announcements

19 total
  1. v0.2.21v0.2.21Oct 1, 20268 downloads

    - The **PageIndex SDK**, **local** or **cloud** — vectorless, reasoning-based RAG, end to end. - **Much faster indexing** — the **PageIndex Flash** engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently. ```python client = PageIndexClient() client.submit_document("report.pdf") client.chat("What does the report conclude?") ``` Index, to chat, to agent integration, one client. Local mode needs no server, no vector DB, no PageIndex API key. ## Highlights - **Flash engine**: the local default (`mode="standard"` keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default — `optimize="merge"` for the deterministic LLM-free pass, `"full"` (default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time. - **One complete surface, local and cloud**: `PageIndexLocalClient(storage_path=...)` is the same client as cloud — submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged. - **Cloud documents, your ow

  2. v0.2.20v0.2.20Sep 28, 202632 downloads

    - The **PageIndex SDK**, **local** or **cloud** — vectorless, reasoning-based RAG, end to end. - **Much faster indexing** — the **PageIndex Flash** engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently. ```python client = PageIndexClient() client.submit_document("report.pdf") client.chat("What does the report conclude?") ``` Index, to chat, to agent integration, one client. Local mode needs no server, no vector DB, no PageIndex API key. ## Highlights - **Flash engine**: the local default (`mode="standard"` keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default — `optimize="merge"` for the deterministic LLM-free pass, `"full"` (default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time. - **One complete surface, local and cloud**: `PageIndexLocalClient(storage_path=...)` is the same client as cloud — submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged. - **Cloud documents, your ow

  3. v0.2.19v0.2.19Sep 21, 202622 downloads

    - The **PageIndex SDK**, **local** or **cloud** — vectorless, reasoning-based RAG, end to end. - **Much faster indexing** — the **PageIndex Flash** engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently. ```python client = PageIndexClient() client.submit_document("report.pdf") client.chat("What does the report conclude?") ``` Index, to chat, to agent integration, one client. Local mode needs no server, no vector DB, no PageIndex API key. ## Highlights - **Flash engine**: the local default (`mode="standard"` keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default — `optimize="merge"` for the deterministic LLM-free pass, `"full"` (default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time. - **One complete surface, local and cloud**: `PageIndexLocalClient(storage_path=...)` is the same client as cloud — submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged. - **Cloud documents, your ow

  4. v0.2.18v0.2.18Sep 16, 202628 downloads

    - The **PageIndex SDK**, **local** or **cloud** — vectorless, reasoning-based RAG, end to end. - **Much faster indexing** — the **PageIndex Flash** engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently. ```python client = PageIndexClient() client.submit_document("report.pdf") client.chat("What does the report conclude?") ``` Index, to chat, to agent integration, one client. Local mode needs no server, no vector DB, no PageIndex API key. ## Highlights - **Flash engine**: the local default (`mode="standard"` keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default — `optimize="merge"` for the deterministic LLM-free pass, `"full"` (default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time. - **One complete surface, local and cloud**: `PageIndexLocalClient(storage_path=...)` is the same client as cloud — submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged. - **Cloud documents, your ow

  5. v0.2.17v0.2.17Sep 13, 202616 downloads

    - The **PageIndex SDK**, **local** or **cloud** — vectorless, reasoning-based RAG, end to end. - **Much faster indexing** — the **PageIndex Flash** engine gets the tree from layout stats: no LLM involved for the structure generation itself, LLMs only write the node summaries, and tree expansion proposes a wave of nodes concurrently. ```python client = PageIndexClient() client.submit_document("report.pdf") client.chat("What does the report conclude?") ``` Index, to chat, to agent integration, one client. Local mode needs no server, no vector DB, no PageIndex API key. ## Highlights - **Flash engine**: the local default (`mode="standard"` keeps the classic LLM pipeline). Embedded bookmarks are consumed when trustworthy, and tree optimization is on by default — `optimize="merge"` for the deterministic LLM-free pass, `"full"` (default) adds LLM expand, which runs a wave of nodes concurrently instead of one round-trip at a time. - **One complete surface, local and cloud**: `PageIndexLocalClient(storage_path=...)` is the same client as cloud — submit, tree, page content, chat, and agent tools all present in both modes, so code moves between them unchanged. - **Cloud documents, your ow

Code frequency

additions and deletions
+17.7K-17.7KWeek of 2025-10-12: +6 linesWeek of 2025-10-12: -5 linesWeek of 2025-10-19: +0 linesWeek of 2025-10-19: -0 linesWeek of 2025-10-26: +669 linesWeek of 2025-10-26: -2 linesWeek of 2025-11-02: +122 linesWeek of 2025-11-02: -90 linesWeek of 2025-11-09: +17 linesWeek of 2025-11-09: -15 linesWeek of 2025-11-16: +1,356 linesWeek of 2025-11-16: -182 linesWeek of 2025-11-23: +0 linesWeek of 2025-11-23: -0 linesWeek of 2025-11-30: +3 linesWeek of 2025-11-30: -2 linesWeek of 2025-12-07: +0 linesWeek of 2025-12-07: -0 linesWeek of 2025-12-14: +66 linesWeek of 2025-12-14: -60 linesWeek of 2025-12-21: +18 linesWeek of 2025-12-21: -15 linesWeek of 2025-12-28: +0 linesWeek of 2025-12-28: -0 linesWeek of 2026-01-04: +1 linesWeek of 2026-01-04: -0 linesWeek of 2026-01-11: +0 linesWeek of 2026-01-11: -0 linesWeek of 2026-01-18: +12 linesWeek of 2026-01-18: -9 linesWeek of 2026-01-25: +4 linesWeek of 2026-01-25: -2 linesWeek of 2026-02-01: +0 linesWeek of 2026-02-01: -0 linesWeek of 2026-02-08: +19 linesWeek of 2026-02-08: -0 linesWeek of 2026-02-15: +0 linesWeek of 2026-02-15: -0 linesWeek of 2026-02-22: +5 linesWeek of 2026-02-22: -5 linesWeek of 2026-03-01: +1,494 linesWeek of 2026-03-01: -876 linesWeek of 2026-03-08: +0 linesWeek of 2026-03-08: -0 linesWeek of 2026-03-15: +81 linesWeek of 2026-03-15: -107 linesWeek of 2026-03-22: +5,732 linesWeek of 2026-03-22: -4,875 linesWeek of 2026-03-29: +78 linesWeek of 2026-03-29: -57 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +67 linesWeek of 2026-04-19: -2 linesWeek of 2026-04-26: +10 linesWeek of 2026-04-26: -9 linesWeek of 2026-05-03: +19 linesWeek of 2026-05-03: -15 linesWeek of 2026-05-10: +2 linesWeek of 2026-05-10: -2 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +10 linesWeek of 2026-05-24: -6 linesWeek of 2026-05-31: +23 linesWeek of 2026-05-31: -21 linesWeek of 2026-06-07: +73 linesWeek of 2026-06-07: -13 linesWeek of 2026-06-14: +82 linesWeek of 2026-06-14: -35 linesWeek of 2026-06-21: +68 linesWeek of 2026-06-21: -21 linesWeek of 2026-06-28: +203 linesWeek of 2026-06-28: -62 linesWeek of 2026-07-05: +42 linesWeek of 2026-07-05: -26 linesWeek of 2026-07-12: +123 linesWeek of 2026-07-12: -38 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +17,658 linesWeek of 2026-07-26: -73 linesWeek of 2026-08-02: +936 linesWeek of 2026-08-02: -115 linesWeek of 2026-08-09: +11,770 linesWeek of 2026-08-09: -2,413 linesWeek of 2026-08-16: +5,055 linesWeek of 2026-08-16: -1,634 linesWeek of 2026-08-23: +1,845 linesWeek of 2026-08-23: -367 linesWeek of 2026-08-30: +2,654 linesWeek of 2026-08-30: -1,919 linesWeek of 2026-09-06: +2,507 linesWeek of 2026-09-06: -813 linesWeek of 2026-09-13: +4,226 linesWeek of 2026-09-13: -2,683 linesWeek of 2026-09-20: +2,726 linesWeek of 2026-09-20: -689 linesWeek of 2026-09-27: +4,587 linesWeek of 2026-09-27: -4,059 linesWeek of 2026-10-04: +0 linesWeek of 2026-10-04: -0 linesOct 12, 2025Oct 4, 2026
+64.4K lines added, -21.3K removed over the last year.

Commits per week

last 52 weeks
230Week of 2025-10-12: 2 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 4 commitsWeek of 2025-11-02: 23 commitsWeek of 2025-11-09: 4 commitsWeek of 2025-11-16: 8 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 1 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 5 commitsWeek of 2025-12-21: 2 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 1 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 3 commitsWeek of 2026-01-25: 2 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 1 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 2 commitsWeek of 2026-03-01: 8 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 2 commitsWeek of 2026-03-22: 9 commitsWeek of 2026-03-29: 6 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 3 commitsWeek of 2026-04-26: 1 commitsWeek of 2026-05-03: 6 commitsWeek of 2026-05-10: 1 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 2 commitsWeek of 2026-05-31: 6 commitsWeek of 2026-06-07: 1 commitsWeek of 2026-06-14: 7 commitsWeek of 2026-06-21: 4 commitsWeek of 2026-06-28: 7 commitsWeek of 2026-07-05: 4 commitsWeek of 2026-07-12: 4 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 15 commitsWeek of 2026-08-02: 12 commitsWeek of 2026-08-09: 7 commitsWeek of 2026-08-16: 12 commitsWeek of 2026-08-23: 14 commitsWeek of 2026-08-30: 23 commitsWeek of 2026-09-06: 11 commitsWeek of 2026-09-13: 22 commitsWeek of 2026-09-20: 12 commitsWeek of 2026-09-27: 6 commitsWeek of 2026-10-04: 0 commitsOct 12, 2025Oct 4, 2026
263 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 1 commitsSun 1:00 — 1 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 2 commitsSun 5:00 — 2 commitsSun 6:00 — 0 commitsSun 7:00 — 3 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 2 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 1 commitsSun 15:00 — 0 commitsSun 16:00 — 1 commitsSun 17:00 — 10 commitsSun 18:00 — 4 commitsSun 19:00 — 5 commitsSun 20:00 — 2 commitsSun 21:00 — 3 commitsSun 22:00 — 1 commitsSun 23:00 — 1 commitsMon 0:00 — 3 commitsMon 1:00 — 3 commitsMon 2:00 — 1 commitsMon 3:00 — 3 commitsMon 4:00 — 1 commitsMon 5:00 — 4 commitsMon 6:00 — 2 commitsMon 7:00 — 5 commitsMon 8:00 — 1 commitsMon 9:00 — 4 commitsMon 10:00 — 0 commitsMon 11:00 — 1 commitsMon 12:00 — 3 commitsMon 13:00 — 1 commitsMon 14:00 — 3 commitsMon 15:00 — 2 commitsMon 16:00 — 4 commitsMon 17:00 — 7 commitsMon 18:00 — 5 commitsMon 19:00 — 1 commitsMon 20:00 — 2 commitsMon 21:00 — 2 commitsMon 22:00 — 4 commitsMon 23:00 — 1 commitsTue 0:00 — 3 commitsTue 1:00 — 2 commitsTue 2:00 — 1 commitsTue 3:00 — 3 commitsTue 4:00 — 0 commitsTue 5:00 — 1 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 3 commitsTue 11:00 — 0 commitsTue 12:00 — 6 commitsTue 13:00 — 0 commitsTue 14:00 — 1 commitsTue 15:00 — 4 commitsTue 16:00 — 3 commitsTue 17:00 — 1 commitsTue 18:00 — 5 commitsTue 19:00 — 4 commitsTue 20:00 — 2 commitsTue 21:00 — 4 commitsTue 22:00 — 5 commitsTue 23:00 — 5 commitsWed 0:00 — 3 commitsWed 1:00 — 8 commitsWed 2:00 — 2 commitsWed 3:00 — 0 commitsWed 4:00 — 6 commitsWed 5:00 — 4 commitsWed 6:00 — 1 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 3 commitsWed 11:00 — 1 commitsWed 12:00 — 1 commitsWed 13:00 — 0 commitsWed 14:00 — 4 commitsWed 15:00 — 6 commitsWed 16:00 — 3 commitsWed 17:00 — 6 commitsWed 18:00 — 6 commitsWed 19:00 — 3 commitsWed 20:00 — 6 commitsWed 21:00 — 5 commitsWed 22:00 — 3 commitsWed 23:00 — 7 commitsThu 0:00 — 4 commitsThu 1:00 — 3 commitsThu 2:00 — 2 commitsThu 3:00 — 6 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 1 commitsThu 9:00 — 7 commitsThu 10:00 — 1 commitsThu 11:00 — 4 commitsThu 12:00 — 1 commitsThu 13:00 — 7 commitsThu 14:00 — 5 commitsThu 15:00 — 3 commitsThu 16:00 — 5 commitsThu 17:00 — 4 commitsThu 18:00 — 7 commitsThu 19:00 — 8 commitsThu 20:00 — 10 commitsThu 21:00 — 5 commitsThu 22:00 — 3 commitsThu 23:00 — 6 commitsFri 0:00 — 4 commitsFri 1:00 — 8 commitsFri 2:00 — 3 commitsFri 3:00 — 7 commitsFri 4:00 — 2 commitsFri 5:00 — 1 commitsFri 6:00 — 0 commitsFri 7:00 — 3 commitsFri 8:00 — 1 commitsFri 9:00 — 1 commitsFri 10:00 — 5 commitsFri 11:00 — 1 commitsFri 12:00 — 1 commitsFri 13:00 — 1 commitsFri 14:00 — 4 commitsFri 15:00 — 4 commitsFri 16:00 — 1 commitsFri 17:00 — 4 commitsFri 18:00 — 3 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 1 commitsFri 22:00 — 11 commitsFri 23:00 — 0 commitsSat 0:00 — 4 commitsSat 1:00 — 6 commitsSat 2:00 — 2 commitsSat 3:00 — 1 commitsSat 4:00 — 2 commitsSat 5:00 — 4 commitsSat 6:00 — 1 commitsSat 7:00 — 0 commitsSat 8:00 — 1 commitsSat 9:00 — 1 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 2 commitsSat 14:00 — 2 commitsSat 15:00 — 0 commitsSat 16:00 — 2 commitsSat 17:00 — 2 commitsSat 18:00 — 2 commitsSat 19:00 — 1 commitsSat 20:00 — 2 commitsSat 21:00 — 0 commitsSat 22:00 — 2 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
May 8, 2026daily#25+47
May 7, 2026daily#18+56
May 6, 2026daily#11+83
May 5, 2026daily#18+62
Feb 26, 2026daily#25+168
Feb 23, 2026daily#16+145
Feb 3, 2026daily#20+128
Feb 2, 2026daily#13+207
Feb 1, 2026daily#8+256
Jan 31, 2026daily#20+142
Jan 26, 2026daily#14+204
Jan 25, 2026daily#11+274
Jan 24, 2026daily#4+408
Jan 23, 2026daily#13+268
  • public-apis/public-apis

    A collective list of free APIs

    486.1K stars · Python

  • openclaw/openclaw

    The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

    391.3K stars · TypeScript

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    373.2K stars · Python

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    295.2K stars · Shell

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    285.8K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript