PDF to JSON
Extract every text run with its page, position and size as structured JSON. Runs entirely in your browser.
Drop a PDF here
or click to browse
0 bytes
- Processed on this device
- 0 bytes
- Files opened here
- 0
- Requests carrying a file
- 0
Something not working as it should? Report an issue with PDF to JSON
About PDF to JSON
Plain text extraction loses the one thing that often matters most: where each piece of text sat on the page. Without position, a table becomes a jumble and there is no way to tell a heading from a footnote.
This tool returns structured JSON in which every text run carries its page number, its coordinates, its font size and its font name. That is enough to rebuild tables, detect headings by size, associate labels with values in a form, or filter out headers and footers by position.
It is aimed at anyone parsing documents programmatically: extracting line items from invoices, pulling fields from standardised forms, or converting a stack of reports into something a script can read. The usual route to this is a paid document API that requires uploading every file.
Since parsing happens in your browser, that trade off disappears. Contracts, invoices and personal records can be turned into structured data without any of them being transmitted, which is often the difference between a workflow being allowed and not.
How to pdf to json
- Load the PDF
Drop the file in. It is parsed locally with PDF.js.
- Review the output
The JSON structure is previewed so you can check the shape before downloading.
- Download
Save the .json file and feed it into your own tooling.
Frequently asked questions
What is in the JSON?
Each text run with its page number, x and y coordinates, width, height, font size and font name. That is enough to reconstruct layout, detect headings or align table columns.
Does this work on scanned PDFs?
No, because a scan contains no text objects to describe. Run OCR PDF first to add a text layer, then extract to JSON.
Why use this instead of plain text extraction?
Because position carries meaning. Tables, multi-column layouts, form labels and headers are all impossible to reconstruct reliably from a flat string of text.
Is there a size limit?
None imposed. Very large documents produce large JSON, so a preview is shown rather than rendering the whole structure on screen.