FileFlipLocal
PDF tools/Ref 13/PDF

Extract text from a PDF

Pull the text out of a PDF into a plain .txt file. Works on documents that contain real text; a scan is a picture and has no text to pull.

No upload

This tool does not upload your files

The conversion runs entirely in your browser. Your files stay on your device from start to finish, there is no server copy, and nothing is stored once you leave the page. That is true for every tool on FileFlip.

About this tool

Sometimes you do not want the document, you want the words in it. Pasting from a PDF viewer tends to bring line breaks and column order along with it. This tool pulls the text out into a plain .txt file, working out where lines and paragraphs end from the position of the words on the page.

How to pDF to text

  1. 01

    Add the PDF

    One document at a time.

  2. 02

    Choose the pages

    Leave the field empty for the whole document, or list pages and ranges to take a section.

  3. 03

    Decide about page markers

    A short marker before each page is useful when you need to point back at where something came from, and noise when you just want the prose.

  4. 04

    Download the .txt

    Plain UTF-8 text, so accented characters and other scripts survive.

What to watch for

A scan has no text in it

A scanned page is a photograph. There are no words in the file to extract, only pixels that look like words. Reading those needs optical character recognition, which this tool does not do, and the tool will tell you rather than hand you an empty file.

Layout is not preserved

Tables come out as loose runs of text, and a two column page is read in whatever order the words were written into the file. Simple prose comes out cleanly; complex layouts do not.

Paragraphs are inferred

A PDF stores glyphs at coordinates and has no idea what a paragraph is. A large vertical gap is treated as a paragraph break and a small one as a line break, which is right most of the time.

Frequently asked

Are my documents uploaded anywhere?
No. The conversion runs inside your browser using code that was downloaded to your device. Your file is never sent to us or to anyone else, so there is no copy to store or leak.
Why is my file empty?
Almost always because the PDF is a scan. Open it in a viewer and try to select a word: if you cannot, there is no text in the file to extract.
Does it keep bold and headings?
No. A .txt file has no formatting. If you want an editable document that keeps its paragraphs, use the PDF to Word tool instead.
What about other alphabets?
The output is UTF-8, so Latvian, Greek, Cyrillic and the rest come through as long as the PDF stores real text rather than outlines.