yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
This tool helps you quickly get information out of PDF documents, convert them to other formats, or even fill out forms. You can feed it individual PDF files or entire batches, and it will give you back the raw text, images, structured data like tables, or converted Markdown/HTML files. It's designed for anyone who needs to process many PDFs efficiently, such as data analysts, researchers, or operations managers.
421 stars and 6,692 monthly downloads. Available on PyPI.
Use this if you need to rapidly extract content from a large number of PDF files or integrate PDF processing into automated workflows.
Not ideal if you primarily need advanced PDF design capabilities or interactive editing for single documents with a graphical user interface.
Stars
421
Forks
40
Language
Rust
License
Apache-2.0
Category
Last pushed
Mar 11, 2026
Monthly downloads
6,692
Commits (30d)
0
Get this data via API
curl "https://pt-edge.onrender.com/api/v1/quality/rag/yfedoseev/pdf_oxide"
Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.
Compare
Related tools
kreuzberg-dev/kreuzberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, and...
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR...
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
AKSarav/pdfstract
PDFStract - The Extraction and Chunking Layer in Your RAG Pipeline - Available as CLI - WEBUI - API
NanoNets/docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking...