PDF file │ Stage 1 – Ingest (pypdfium2 + pdfplumber) │ PDF → PageRaw list │ Extracts text spans (text, bbox, font name/size, bold, italic), │ vector drawings, and embedded images for every page. │ ...
These numbers are measured, not estimated. The ground truth and evaluation script are included so the methodology is verifiable. This is not 99% accuracy. Achieving 99% on heterogeneous OCR/PDF data ...