yfedoseev/pdf_oxide

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

/ 100

Established

This tool helps you quickly get information out of PDF documents, convert them to other formats, or even fill out forms. You can feed it individual PDF files or entire batches, and it will give you back the raw text, images, structured data like tables, or converted Markdown/HTML files. It's designed for anyone who needs to process many PDFs efficiently, such as data analysts, researchers, or operations managers.

421 stars and 6,692 monthly downloads. Available on PyPI.

Use this if you need to rapidly extract content from a large number of PDF files or integrate PDF processing into automated workflows.

Not ideal if you primarily need advanced PDF design capabilities or interactive editing for single documents with a graphical user interface.

document-processing data-extraction workflow-automation data-conversion information-retrieval

No Dependents

Maintenance 10 / 25

Adoption 19 / 25

Maturity 22 / 25

Community 16 / 25

How are scores calculated?

Stars

421

Forks

Language

Rust

License

Apache-2.0

Compare

pdf_oxide and kreuzberg

Related tools

kreuzberg-dev/kreuzberg

A polyglot document intelligence framework with a Rust core. Extract text, metadata, and...

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR...

opendataloader-project/opendataloader-pdf

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

AKSarav/pdfstract

PDFStract - The Extraction and Chunking Layer in Your RAG Pipeline - Available as CLI - WEBUI - API

NanoNets/docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking...

Explore RAG Tools

All categories Trending RAG directory Insights