PDF to text
Pull the text layer out of a PDF page by page - handy for quoting, translating or reusing content.
Private by design
Files are processed inside your browser and never uploaded anywhere.
No waiting
Results appear in seconds, with no queues or file size servers involved.
No sign-up
Use every tool as often as you like - no account, no watermarks.
When you need the words rather than the layout - to quote a clause, translate a report, feed a document into a spreadsheet or search a long contract in a plain editor - extracting the text layer is far quicker than copying page by page. PDF to text pulls the readable content out of your document in page order and hands you one plain text file.
How to extract text from a PDF
- 1Add the PDF you want to read the text out of.
- 2Run the extraction and wait while each page is read in turn.
- 3Download the plain text file and open it in any editor.
What extraction can and cannot recover
Text extraction is the right tool for reuse and analysis: quoting legislation, translating documentation, pasting content into a CMS, counting words, or searching a document that your reader handles awkwardly. It works on any PDF that carries a real text layer - anything produced digitally from Word, a browser, an accounting system or a design tool.
What comes out is the text itself, in reading order per page, without formatting: no fonts, colours, bold, headings, images or page furniture. Tables lose their grid and become sequences of cell values, and multi-column layouts can interleave if the source encodes them unusually, so a quick read-through before reuse is wise. The critical limitation is scans: a photographed or faxed page contains only a picture of words, with no text layer to extract, so the result will be empty or nearly so unless the file was already put through OCR. Extraction never modifies your PDF, and because everything runs in your browser, contracts and personal records are never uploaded - although a very long document takes time and memory, since each page is parsed individually on your device.
Frequently asked about PDF to text
Why is my output empty or almost empty?
The PDF is most likely a scan - an image of a page with no text layer. Extraction can only read text that is actually stored in the file, so such documents need OCR first.
Is the original formatting preserved?
No. You get plain text in reading order. Fonts, sizes, colours, headings and images are not part of the output.
How are tables handled?
Cell contents come through, but the grid does not. Expect a sequence of values that usually needs tidying before you paste it into a spreadsheet.
Can I tell which page a passage came from?
Yes. The text is written page by page in document order, so passages appear in the same sequence as in the PDF.
Does it work on password-protected files?
A document that requires a password to open must be unlocked in your reader and saved without encryption before extraction can read it.
Why does a long PDF take a while?
Each page is parsed separately in your browser, so time scales with page count. Keep the tab in the foreground and split very large files if needed.