Flagged: spawns subprocesses. Scanned against the SlowMist agent-security taxonomy, refreshed every 8 hours. Full audit →
by NameetP · MCP Server · ★ 79
Last updated: · Indexed by AgentSkillsHub · Auto-synced every 8h
🔒 Is pdfmux safe to install? View the security audit →
pdfmux Self-healing PDF extraction with per-page confidence scoring. Open-source LlamaParse alternative for RAG pipelines, MCP server for Claude Desktop, LangChain + LlamaIndex loaders. Ranked #2 on opendataloader-bench (0.900). The only PDF extractor that audits its own output. Catches blank pages, scrambled columns, broken tables — re-extracts them with a stronger backend. So your LLM gets clean data, not silent garbage. Routes each page to the best of 5 rule-based backends + BYOK LLM fallback (Gemini / Claude / GPT-4o / Ollama). One CLI. One API. Zero config. PDF ── pdfmux router ── best extractor per page ── audit ── re-extract failures ── Markdown / JSON / chunks | ├─ PyMuPDF (digital text, 0.01s/page) ├─ OpenDataLoader (complex layouts, 0.05s/page) ├─ RapidOCR (scanned pages, CPU-only)
| Stars | 79 |
| Forks | 12 |
| Language | Python |
| Category | MCP Server |
| License | MIT |
| Quality Score | 68.870137129299/100 |
| Open Issues | 4 |
| Last Updated | 2026-08-13 |
| Created | 2026-03-03 |
| Platforms | cli, mcp, python |
| Est. Tokens | ~19k |
Looking for a pdfmux alternative? If you're comparing pdfmux with other mcp server tools, these 6 projects are the closest alternatives on Agent Skills Hub — ranked by topic overlap, star count, and community traction.
MCP server that lets Claude Code and other AI agents work through large PDFs, and whole folders of them, witho
AI-Native document parser: PDF, Office & images → clean Markdown with LaTeX, tables & OCR. Zero-dependency CLI
Open-source protocol suite standardizing LLM, Vector, Graph, and Embedding infrastructure across LangChain, Ll
A Model Context Protocol (MCP) server implementation that integrates with the Nutrient Document Web Service (D
Python, LlamaIndex, LangChain, 15 Property Graph, 4 RDF , 10 Vector, OpenSearch, Elasticsearch, Alfresco, Nuxe
A self-healing scraper for hostile sites: broken selectors repair themselves, browser rendering kicks in when
Explore other popular mcp server tools:
pdfmux is PDF extraction that audits its own output — and certifies any other extractor's, catching pages they silently dropped. Verify signed manifests offline: free, MIT, no account. 0.903 on opendataloader-b. It is categorized as a MCP Server with 79 GitHub stars.
pdfmux is primarily written in Python. It covers topics such as ai-agent, docling, document-ai.
You can find installation instructions and usage details in the pdfmux GitHub repository at github.com/NameetP/pdfmux. The project has 79 stars and 12 forks, indicating an active community.
pdfmux is released under the MIT license, making it free to use and modify according to the license terms.
The top alternatives to pdfmux on Agent Skills Hub include pdf-mcp, MinerU-Skill, corpusos. Each offers a different approach to the same problem space — compare them side-by-side by stars, quality score, and community activity.