← プロジェクト一覧に戻る

PageIndex

ベクトル検索を使わないRAG。文書の目次にあたる木構造を作り、モデルにそこをたどらせて該当箇所へ行き着かせる

MIT
スター
35.3k
フォーク
3.1k
オープンIssue
162
最終コミット
2026年8月26日

PageIndexとは

ベクトル検索が見つけるのは質問に似た文章ですが、似ていることと関係があることは別です。長い規制文書では、答えにあたる段落が質問とほとんど語彙を共有していないことがよくあります。PageIndexはベクトルの保存先を使いません。文書の実際の構造を節ごとに木として組み立て、人が目次を使うのと同じ手順でモデルにそこをたどらせます。範囲を絞り、答えがありそうな箇所を開き、読む、という流れです。文章を機械的に分割することはなく、答えは必ずどの節から来たかを示します。引き換えは1問あたりの費用です。検索が参照ではなく推論になるためで、想定されているのは短い文書を大量に持つ状況ではなく、長い個別の文書です。

PageIndexで何ができますか?

  • 類似ではなく推論で引き当てる — どの節に答えがあるはずかをモデルが判断して開くため、質問と違う言い回しで答えている箇所も見つかります。
  • 機械的な分割ではなく本来の節を単位にする — 検索の単位が文書の実際の節なので、定義とそれに依存する条文が別々の断片に切り離されることがありません。
  • 答えの出所を示す — 結果には必ず、木構造のどの部分から読んだかが付きます。確かめられる答えと、信じるしかない答えの違いはここにあります。
  • ベクトルデータベースを持たない — 索引を作る、埋め込む、調整する、同期を保つという作業がありません。文書を渡すと、保存されるのは木構造です。
  • 構造が明快な文書は数秒で取り込む — 目次のはっきりした資料向けに、モデルに推測させず文書そのものから構造を抽出する高速な方式が用意されています。
  • 手元でもホスティング型でも動かす — 同じクライアントで、自分のモデルの鍵を使って手元で索引と検索を行うことも、コードを変えずに開発元のクラウドに向けることもできます。

PageIndexを選ぶ前に

  • 検索そのものがモデルの呼び出しを伴うため、1問あたりの時間と費用はベクトル検索より大きくなります。安い問い合わせを大量に捌く用途では計算が変わります。
  • 想定されているのは長く構造の整った専門文書です。短く構造のないページの集合ではたどるべき木がほとんどなく、開発元が挙げる高い正答率も金融文書のベンチマークによるものです。

よくある質問

PageIndexは商用利用できますか?

PageIndexはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

PageIndexはどの形で使えますか?

PageIndexはローカル実行・セルフホスト・マネージドクラウドの形で利用できます。

ドキュメント

VectifyAI/PageIndex のREADMEより転載(MIT)。 原文を読む ↗

PageIndex: Vectorless, Reasoning-based RAG

  • [2026/08] 🔥 PageIndex SDK: pip install -U pageindex now ships local mode: index, retrieve, and chat entirely on your machine with your own LLM key, or point the same client at PageIndex Cloud with an API key.
  • [2026/08] ⚡ PageIndex Flash: tree structure generation from PDFs in seconds, with structure extracted heuristically instead of by an LLM.
  • PageIndex Chat: a human-like document analysis agent for long professional documents. Also available via MCP or API.

What is PageIndex?

Are you frustrated with vector database retrieval accuracy for long and complex documents? Vector-based RAG retrieves by semantic similarity. But similarity ≠ relevance — what retrieval actually needs is relevance, and relevance requires reasoning. On professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search misses what is relevant but not similar, and returns what is similar but not relevant.

Inspired by AlphaGo, PageIndex replaces the vector index with a hierarchical tree index and lets an LLM reason its way through it, the way a human expert turns to and reads the right section of a long report. Retrieval happens in two steps:

  1. Index: generate a tree-structure index for each document
  2. Retrieve: search that tree with LLM reasoning, agentically

Why it works

PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, without vector databases or chunking.

Compare with Vector RAG

Vector RAGPageIndex
Indexvector indextree index
Unitfixed-size chunksnatural sections
Retrievalsemantic similarity searchLLM reasoning over the tree
Resultopaque, “vibe retrieval”traceable to explicit references
Contextquery embedding onlyfull context: conversation history, domain knowledge, etc.

It is ideal for financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any other long, complex professional document.

PageIndex achieved state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), vastly outperforming vector-based RAG (see Benchmarks).

Quickstart

pip install -U pageindex
import os
from pageindex import PageIndexClient

os.environ["OPENAI_API_KEY"] = "your-openai-key"

client = PageIndexClient(                     
    index="gpt-5.6-luna",               # model to build the tree index
    chat="gpt-5.6-sol",                 # model to search the tree
)
doc_id = client.submit_document("report.pdf")["doc_id"]

answer = client.chat("What was the 2023 operating margin, and where is it stated?",
                     doc_id=doc_id)
print(answer)

Model Recommendations

  • index=: a basic model is sufficient. The index model generates the document’s tree index. A basic model is sufficient to produce a good tree structure.
  • chat=: use the best model you can afford. The chat model searches the tree to retrieve information. See Query cost and accuracy.

See the Detailed Usage Guide to configure other models, or integrate PageIndex with your own agent.

Get Answers with Citations

To request inline page-level citations, pass a system message together with the question:

messages = [
    {
        "role": "system",
        "content": (
            'Cite only statements supported by tool outputs using '
            '<cite doc="{docName}" page="{pageNumber}"/>'
        ),
    },
    {"role": "user", "content": "Summarize the document."},
]

answer = client.chat(messages, doc_id=doc_id)

The model fills in the document name and page number, for example:

Revenue increased during the reporting period. <cite doc="report.pdf" page="12"/>

Benchmarks

Indexing cost and time

Building a tree locally runs about $0.001 per page with index_model="gpt-5.6-luna", so a 1,000-page textbook costs a little over a dollar and a few minutes, once, and every later question reuses it. PageIndex is designed not to rely heavily on the model used at index time, so in our experiments a basic model does not hurt quality.

Indexing time also scales predictably with document length. In the same local setup, the benchmark documents (9 to 1,098 pages) finished in roughly 13 seconds to 4.5 minutes.

Query cost and accuracy

PageIndex-OSS-Benchmark measures exactly the setup in the quickstart above (PageIndexClient() in local mode, flash indexing, no OCR) on 62 lookup questions over 34 PDFs (1,945 pages) drawn from MMLongBench-Doc-V2. Every question’s answer is a fact stated in running text, so a wrong answer is a retrieval or reading failure, not a reasoning one.

Full results, data, and the runner are in the benchmark repo.

Detailed Usage Guide

⚙️ Step 1: Initialize the client

Create a local client and choose the models used for indexing and retrieval:

from pageindex import PageIndexClient
import os

client = PageIndexClient(
    index_model="gpt-5.6-luna",
    chat_model="gpt-5.6-sol",
    storage_path=".pageindex",
)
  • index_model builds the tree index. A basic model is sufficient.

  • chat_model searches the tree and answers questions. Use the best model you can afford.

  • storage_path specifies where indexed documents are stored locally.

index_model= / chat_model= are the flat spellings of the quickstart’s index= / chat=; either spelling works.

Model naming conventions

Model names follow LiteLLM’s naming convention. Choose the format that matches your provider:

OpenAI: use the model name directly and set OPENAI_API_KEY:

os.environ["OPENAI_API_KEY"] = "your-openai-api-key"
chat_model = "gpt-5.6-sol"

Anthropic: prefix the model name with anthropic/ and set ANTHROPIC_API_KEY:

os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-api-key"
chat_model = "anthropic/claude-sonnet-4-6"

OpenRouter: prefix the provider and model name with openrouter/ and set OPENROUTER_API_KEY:

os.environ["OPENROUTER_API_KEY"] = "your-openrouter-api-key"
chat_model = "openrouter/anthropic/claude-sonnet-4-6"

For model names and API key settings for other providers, see the LiteLLM provider documentation.

🌲 Step 2: Build the tree index

submit_document defaults to Flash indexing: the structure is extracted from the PDF’s own layout (no LLM), and a model is called only for node summaries and the tree-optimization expansion pass. It takes seconds.

doc_id = client.submit_document("report.pdf")["doc_id"]

Inspect what you got:

tree = client.get_document_structure(doc_id)    # titles, page ranges, summaries; no text
client.list_documents()                         # everything you have indexed

A PageIndex tree looks like a table of contents optimized for LLMs and agents:

{
  "title": "Financial Stability",
  "node_id": "0006",
  "start_index": 21,
  "end_index": 22,
  "summary": "The Federal Reserve ...",
  "nodes": [
    {
      "title": "Monitoring Financial Vulnerabilities",
      "node_id": "0007",
      "start_index": 22,
      "end_index": 28,
      "summary": "The Federal Reserve's monitoring ..."
    },
    {
      "title": "Domestic and International Cooperation and Coordination",
      "node_id": "0008",
      "start_index": 28,
      "end_index": 31,
      "summary": "In 2023, the Federal Reserve collaborated ..."
    }
  ]
}

See more example documents and generated tree structures.

💬 Step 3: Ask questions

chat() is the one-line surface. Underneath it is a document-QA agent, and you can talk to it over whichever protocol your stack already speaks:

Get a simple answer with chat():

client.chat("What changed in the risk factors?", doc_id=doc_id)

Pass a string or role/content history and get the answer back.

Stream the answer:

client.chat(question, doc_id=doc_id, stream=True)

Returns the answer as text chunks.

Use the OpenAI Chat Completions format:

client.chat_completions(messages, doc_id=doc_id)

Returns the full envelope, including token usage, streaming metadata, and finish_reason.

Use the OpenAI Responses format:

client.responses("...", doc_id=doc_id, reasoning={"effort": "high"})

Returns the agent’s process transcript in items. Append those items to the next call’s input to preserve memory and benefit from provider prompt caching. This requires a Responses-compatible backend in local mode.

Use the Anthropic Messages format:

client.messages("...", model="claude-sonnet-4-6", doc_id=doc_id)

Uses Anthropic’s native Messages API and tool runner. Install it with pip install 'pageindex[anthropic]'.

Pass a list of ids to doc_id to search several documents at once, and keep it identical across a conversation’s calls.

Integrate PageIndex with your own agent

Instead of calling PageIndex’s agent, hand PageIndex’s tools to yours. One call fills every slot:

OpenAI Agents SDK:

from agents import Agent, Runner

agent = Agent(**client.openai_agent_config(doc_id=doc_id))
result = Runner.run_sync(agent, "Summarize the auditor's concerns.")

openai_agent_config() provides the instructions and tools required by an OpenAI agent.

Anthropic SDK tool runner:

runner = anthropic_client.beta.messages.tool_runner(
    **client.anthropic_runner_config(model="claude-sonnet-4-6", doc_id=doc_id),
    messages=[{"role": "user", "content": "Summarize the auditor's concerns."}],
)

anthropic_runner_config() configures Anthropic’s native tool runner. Install the integration with pip install 'pageindex[anthropic]'.

Claude Agent SDK:

options = ClaudeAgentOptions(**client.claude_agent_config(doc_id=doc_id))

claude_agent_config() creates the options for the Claude Agent SDK. Install the integration with pip install 'pageindex[claude]'.

Other agent frameworks:

tools = client.agent_tools()

agent_tools() returns plain Python functions that work with LangChain, PydanticAI, and other agent frameworks.

Each *_config helper is sugar over the explicit pieces (client.agent_instructions() for the system prompt, client.as_openai_tools() / as_anthropic_tools() / as_claude_mcp() for the tools), so you can swap in your own prompt whenever you need to. Locally, doc_id is enforced at the tool layer, not just prompted: out-of-scope lookups return NOT_FOUND.

PageIndex Cloud

The open-source version is ideal for text-heavy PDFs and local workflows. With PageIndex Cloud, document indexing and storage run in the cloud: PageIndex handles parsing, OCR, image understanding, tree-index construction, and managed storage for you. The chat and retrieval layer remains compatible with your model, so you can search the cloud-hosted index using the model provider your application already uses.

Moving indexing and storage from Local to Cloud only requires a PageIndex API key:

import os
from pageindex import PageIndexClient

os.environ["PAGEINDEX_API_KEY"] = "your-pageindex-key"
os.environ["OPENAI_API_KEY"] = "your-openai-key"


client = PageIndexClient(
    index="cloud",                       # build and store the index in PageIndex Cloud
    chat="gpt-5.6-sol",                  # use your preferred compatible model for chat
)

# The rest of your code stays the same (wait=True: cloud indexing is asynchronous)
doc_id = client.submit_document("report.pdf", wait=True)["doc_id"]
print(client.chat("What was the 2023 operating margin?", doc_id=doc_id))
CapabilityLocal (this repo)Cloud (get an API key)
Best fortext-heavy PDFs and local workflowsscanned, image-heavy, and large document collections
Indexingruns locallyruns in PageIndex Cloud, with production OCR and image understanding
Storagelocalmanaged in PageIndex Cloud
Chat modelyour modelyour model, or the managed chat included with your key
Citationspage-levelline-level
Image understanding—✅
Multi-document scalemanualPageIndex File System
MCP server—✅

More About PageIndex Cloud

Ready to Try It?

For dedicated deployment (VPC or on-premises), contact us or book a demo.


⭐ Support Us

Leave us a star 🌟 if you like our project. Thank you!

Please cite this work as:

Mingtian Zhang, Yu Tang and PageIndex Team,
"PageIndex: Next-Generation Vectorless, Reasoning-based RAG",
PageIndex Blog, Sep 2025.
@article{zhang2025pageindex,
  author = {Mingtian Zhang and Yu Tang and PageIndex Team},
  title = {PageIndex: Next-Generation Vectorless, Reasoning-based RAG},
  journal = {PageIndex Blog},
  year = {2025},
  month = {September},
  note = {https://pageindex.ai/blog/pageindex-intro},
}

Connect with Us

         


© 2026 Vectify AI

PageIndex
AIに聞く
GitHub