PageIndex
View on GitHub📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
Python SDK for PageIndex, a vectorless RAG engine that builds a hierarchical tree index per document and lets an LLM reason over that tree to retrieve the right section. Supports local or cloud indexing, agent/MCP integration, and citations; no vector DB or chunking.
Use Cases
Question answering over long financial reportsLegal and regulatory document searchTechnical manual and spec lookupMedical literature reviewAcademic textbook QAMulti-document corpus reasoningCitation-backed retrieval for agentsAgent tool integration via OpenAI Agents SDK, Claude Agent SDK or MCPReplacing vector DB RAG pipelinesReducing per-query cost vs native PDF input
Built With
- Language
- Python
- Frameworks
- openai-agents · claude-agent-sdk · mcp · litellm · openai · anthropic · PyPDF2 · pypdfium2 · python-dotenv · pytest
Tags
rag · vectorless-rag · reasoning-based-retrieval · tree-index · document-retrieval · long-documents · pdf · agentic-retrieval · citations · context-engineering · information-retrieval · no-vector-db · no-chunking · llm · python · document-qa