oxidizePdf
View on GitHubPure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Pure Rust PDF toolkit whose headline feature is structure-aware RAG chunking: each chunk carries pages, bounding boxes, element types, heading context and token estimates, with no ML or C dependencies. Also parses, generates, encrypts and validates PDFs in one crate.
Use Cases
Build RAG ingestion pipelines from PDFsStructure-aware chunking with page/bbox/token metadataExtract text and tables from PDFs for embeddingsInvoice data extraction (ES/EN/DE/IT)Export LLM-ready Markdown/JSON chunks for vector storesGenerate and manipulate PDFs in pure RustValidate PDF/A conformance levelsVerify and prepare digital signaturesRead/write PDF encryption (AES, RC4)Corruption recovery and lenient PDF parsing
Built With
- Language
- Rust
- Frameworks
- LangChain · LlamaIndex · Rust · Cargo
Tags
pdf · rag · chunking · text-extraction · table-extraction · document-processing · embeddings · rust · pdf-parser · pdf-generation · invoice-extraction · digital-signatures · encryption · pdfa · no-c-deps · ai-ingestion