Lightweight PDF parser with layout, tables, formulas and bounding boxes
Summary
papero-pdf-text-extractor is an open-source tool for lightweight PDF structure extraction that preserves reading order, tables, formulas, and bounding boxes. It outputs Markdown, JSON, Word, and Excel, runs CPU-only in the browser, Python, or as an API, and uses a dual-engine approach (layout with PDFium and metadata/OCR with Apache Tika). The project emphasizes zero-ML pipelines and provides benchmarks, quick-start instructions, and Docker/API options.