Why PDFs fight back
Paste an article into a speed reader and it works. Open a PDF and it often produces nonsense. The reason is that a PDF does not store paragraphs. It stores instructions: put this glyph at these coordinates, in this font, at this size. Whether a run of glyphs is a heading, a footnote or a page number is something a human infers from position and typography — and something software has to reconstruct.
Five things routinely end up in the middle of your reading:
- Running headers and footers. The book title on every left page, the chapter on every right one. Extracted naively, they appear every few hundred words, mid-sentence.
- Page numbers. A lone "147" arriving between two clauses.
- Footnotes and references. Small type at the foot of the page, extracted in reading order as if it were the next sentence.
- Hyphenated line breaks. "infor-" and "mation" as two separate words, which in RSVP means two separate flashes of gibberish.
- Columns. Two-column academic layouts read left-to-right across the gutter, interleaving two unrelated sentences into one unreadable stream. This is the worst one, and the most common in papers.
What can be cleaned automatically
Most of that list is solvable without any intelligence at all, just arithmetic on the coordinates the PDF already provides:
- A line that repeats in the same vertical position on more than about a third of the pages is furniture, not text. That removes running headers and footers in one rule.
- A short line that is only digits, or digits with a slash, at the very top or bottom of the page, is a page number.
- A word ending in a hyphen at the end of a line, followed by a lower-case fragment at the start of the next, is one word split in two — rejoin them.
- A line set noticeably larger than the body text, on its own, is a heading. That is enough to rebuild a chapter list, which is what makes a long document navigable.
- Text set in a font distinctly smaller than the body, clustered at the foot of a page, is very probably a note.
Readly does all of this in your browser as the file loads. It is also why nothing is uploaded: the parsing happens on your device, so the file never reaches a server. On a four-page test PDF the rules removed three running headers, rejoined the hyphenated words and found three chapters, and the word count matched the source exactly.
What cannot, yet
Two cases defeat coordinate arithmetic.
Two-column layouts need the reader to work out that the page has two text blocks and that the left one is finished before the right one starts. It is doable — cluster the horizontal positions and you can see the gutter — but until it is done, an academic paper in two columns will read as alternating half-sentences. If that is what you are opening, check the first screen before you settle in.
Scans contain no text at all, only a picture of text. There is nothing to extract; it needs optical character recognition, which is a different problem with a different cost. If you can select a sentence with your mouse in a PDF viewer, the text is there. If you cannot, no reader can read it.
How to actually read a long PDF
- Look at the first screen before you start. Ten seconds of checking tells you whether the extraction worked. Mixed-up columns are obvious immediately.
- Read by chapter, not by document. A 300-page book is not a session. Turn on the pause at chapter starts, and treat each chapter as its own sitting.
- Slow down for the parts that matter. A methods section is not an introduction. Changing pace mid-document is one tap, and the readers who get the most out of this change it constantly.
- Skip the front matter. Title pages, copyright, dedications and tables of contents are all text, and all worthless in a stream. Drag the progress bar past them.
- Keep the original open if you will need to quote. RSVP is for taking something in, not for working with it. When you need the exact words, go back to the page.
Where AI fits, and where it does not
The obvious next step is to hand the extracted text to a language model and ask it to keep the body and drop everything else. It works, and it handles the cases arithmetic cannot — including columns and inconsistent footnotes. It also costs money per document, and it introduces a real risk: a model asked to clean text can quietly rewrite it. For a novel that is vandalism; for a contract it is dangerous.
Our position is that the rules run first and free, for everyone, and that anything involving a model has to be optional, visible and reversible — you should be able to see what was removed and go back to the unedited version. That feature is behind a flag while we find out from readers whether the problem is common enough to be worth the cost.
If your file is an ebook, not a PDF
Then you are in luck, and you should stop reading this page. An EPUB already knows what its chapters and paragraphs are, because it is HTML in a zip — none of the above applies. See reading an EPUB one word at a time. If you have a choice of formats for the same book, take the EPUB every time.