PDF text layer checker
Find out whether a PDF contains real selectable text or is a scanned image, and how many of its words are split across lines. The file is read in this browser tab and never uploaded.
Drop a PDF here, or choose a file
It is read in this browser tab. Nothing is uploaded, and there is no server to upload it to.
Why this matters
Almost everything that makes a PDF useful depends on its text layer. Search, selection, copy-paste, dictionary lookups, translation and every AI reading tool all read that layer rather than the picture on the page. When it is missing, none of them work, and most apps fail silently rather than telling you why.
A working text layer is not the end of it either. We measured 60 recent arXiv papers and found that 78% contained words the typesetter had cut in half, which is why selecting a word sometimes hands you a fragment. This tool counts the same thing in your own document.
How to read the result
- Choose a PDF. Drop a PDF onto the box above or pick one from your device. It is read in the browser tab and never uploaded.
- Read the verdict. The tool reports whether the document has real selectable text, very little text, or none at all — the last meaning it is a scanned image.
- Check the character count. A page of body text runs into the thousands of characters. A few hundred means the page is mostly figures, or the text extracted badly.
- Look at the split words. This counts hyphens at the end of a line. Around half are genuine compounds; the rest are words the typesetter cut in half, which is why selecting one can return a fragment.
Frequently asked questions
- Is my PDF uploaded anywhere?
- No. The file is read inside your browser tab using pdf.js and analysed there. There is no upload endpoint in this tool, so there is nothing for a server to receive. Closing the tab discards it.
- What is a PDF text layer?
- It is the machine-readable text stored alongside the visual page. Born-digital PDFs have one, so you can select, search and copy. Scanned documents are images of pages with no text layer, which is why selection does nothing until OCR adds one.
- My PDF says no text layer. What can I do?
- Run it through OCR. Uploading it to Google Drive and opening it with Google Docs is the quickest free route, though it means sending the document to Google. Dedicated tools such as Adobe Acrobat or ABBYY FineReader do it while preserving layout.
- Why does it only check the first 12 pages?
- Twelve pages is enough to characterise a document and keeps the check fast on a phone. It matches the sample size used in our study of 60 academic papers, so your result is comparable to those numbers.
- What does 'words split across lines' mean?
- When a long word will not fit, the typesetter hyphenates it and the two halves land on different lines. The PDF genuinely contains the fragments, so selecting the word can return half of it. We measured this in 78% of 60 recent arXiv papers.
Built by the people making FlowRead
A local-first PDF reader that explains what you are reading without uploading your documents. Coming to iOS and Android.
Join the waitlist