yfedoseev/pdf_oxide

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

67
/ 100
Established

This tool helps you quickly get information out of PDF documents, convert them to other formats, or even fill out forms. You can feed it individual PDF files or entire batches, and it will give you back the raw text, images, structured data like tables, or converted Markdown/HTML files. It's designed for anyone who needs to process many PDFs efficiently, such as data analysts, researchers, or operations managers.

421 stars and 6,692 monthly downloads. Available on PyPI.

Use this if you need to rapidly extract content from a large number of PDF files or integrate PDF processing into automated workflows.

Not ideal if you primarily need advanced PDF design capabilities or interactive editing for single documents with a graphical user interface.

document-processing data-extraction workflow-automation data-conversion information-retrieval
No Dependents
Maintenance 10 / 25
Adoption 19 / 25
Maturity 22 / 25
Community 16 / 25

How are scores calculated?

Stars

421

Forks

40

Language

Rust

License

Apache-2.0

Last pushed

Mar 11, 2026

Monthly downloads

6,692

Commits (30d)

0

Get this data via API

curl "https://pt-edge.onrender.com/api/v1/quality/rag/yfedoseev/pdf_oxide"

Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.