You've opened a PDF in your browser or Acrobat, pressed Ctrl+F to search for a word, and found nothing — even though you can clearly see that word on the page. This frustrating experience happens because the PDF doesn't actually contain text. It contains an image of text. Understanding why this happens, and how to fix it, saves a lot of searching time.

Why some PDFs aren't searchable

When a PDF is created by exporting from a word processor — Word, Google Docs, Pages, LibreOffice — the text is stored as actual text in the PDF's content stream. Your PDF viewer can read, select, copy, and search that text directly.

When a PDF is created by scanning a physical document — either on a flatbed scanner or by photographing a page with a phone — the scanner captures an image of the page. The resulting PDF contains that image, not text. To the PDF viewer, it's no different from a photograph. There are no text characters to search, select, or copy. The words you can see visually are just patterns of pixels in the image.

PDFs can also be partially searchable — if someone scanned a document and a previous tool has already run OCR on it, there may be a hidden text layer. But if that text layer is inaccurate (which early OCR systems often were) or covers only some pages, searches will miss content.

What is OCR and how does it work?

OCR stands for Optical Character Recognition. It's the technology that reads an image and identifies the text within it, character by character. Modern OCR — particularly Google's Tesseract engine and commercial systems like ABBYY — is remarkably accurate for clean, well-scanned documents in standard fonts. Accuracy degrades for handwriting, unusual fonts, poor scan quality, and non-Latin scripts.

When OCR processes a scanned PDF, it creates an invisible text layer that sits behind the visible image. The image itself doesn't change — the scan looks exactly the same. But now a text layer exists that search tools can read, and which allows you to select and copy text from the document.

Free methods to make a PDF searchable

Google Drive — the simplest free option. Upload your scanned PDF to Google Drive, right-click it, and choose "Open with Google Docs". Google will OCR the document and open it as a Google Doc with the recognised text. You can then copy the text, or export back to PDF with the text layer embedded. Quality is good for clean documents in major languages.

Adobe Acrobat free tier — Adobe's online Acrobat service offers a limited number of free OCR operations per month. Upload your scanned PDF, use the OCR feature to make it searchable, and download the result. The text layer is embedded in the PDF.

Microsoft Word — recent versions of Microsoft 365 Word can open PDF files and automatically apply OCR to scanned content. Open your PDF in Word (File → Open → select the PDF), Word converts it to a Word document with recognised text, then export back to PDF from Word to get a searchable PDF.

Small PDF or ILovePDF — both offer free OCR with limitations on their free tiers (typically a number of conversions per day).

What affects OCR accuracy?

Scan quality is the biggest factor. A clean, high-contrast scan at 300 DPI produces much better OCR results than a low-DPI photo taken at an angle under poor lighting. If you're going to be running OCR on a physical document, it's worth scanning at 300 DPI specifically for this purpose rather than using a quick phone photograph.

Font and layout matter too. Standard printed fonts in a clean layout produce near-perfect accuracy with modern OCR. Handwriting, unusual typefaces, two-column layouts, documents with watermarks or marks across text, and poor-quality photocopies of photocopies all reduce accuracy significantly.

Language support varies by OCR engine. English, French, German, Spanish, and other major European languages have excellent support. Less common languages, right-to-left scripts, and mixed-language documents may have lower accuracy depending on which OCR engine you use.

Checking if your PDF is already searchable

Before running OCR, check whether the PDF already has a text layer. Open it in any PDF viewer, press Ctrl+F (or Cmd+F on Mac), and search for a word you can see on the first page. If the search highlights that word, the document is already searchable. If no results appear, you need OCR.

An even quicker check: try to click and drag to select some text on the page. In a searchable PDF, dragging selects text and highlights it. In a scanned image PDF, dragging has no effect or shows a rubber-band selection that doesn't highlight any text.

Step-by-step: making a PDF searchable with Google Drive

Google Drive's OCR is the most accessible free option for most people. Here's the exact process.

  1. Go to drive.google.com and sign in to your Google account
  2. Click New → File upload and select your scanned PDF
  3. Once uploaded, right-click the PDF in Drive and choose Open with → Google Docs
  4. Google automatically runs OCR and opens the document in Google Docs with recognised text
  5. To get a searchable PDF back: in Google Docs, go to File → Download → PDF Document (.pdf)
  6. The downloaded PDF contains both the original scan image and a hidden text layer — it's now searchable

The quality of Google's OCR is excellent for clean, well-scanned text in major languages. The text layer is hidden behind the original scan image, so the document looks identical — you just gain the ability to search, select, and copy text.

Checking OCR accuracy before you rely on it

OCR is not perfect. Before depending on a searchable PDF for important searches — finding a specific clause in a contract, locating a name in a long document — it's worth checking the accuracy of the recognised text.

Open the OCR'd PDF in any viewer and try Ctrl+F to search for a word you can see on the first page. If it highlights correctly, the OCR captured that word accurately. Then try selecting text by clicking and dragging — see if the selected text matches what's visually on the page. For critical documents, copy a paragraph of text and paste it into a text editor to compare it with the original visually.

Common OCR errors include: confusing '1' (one) with 'l' (lowercase L) and 'I' (uppercase i), misreading digits in handwritten or stylised numbers, combining two short words into one, and breaking one long word into two. For standard typed documents at reasonable scan quality, accuracy is typically 98-99% — high enough for searching but worth verifying for critical data extraction.

Batch OCR for multiple documents

If you have many documents to make searchable — an archive of scanned files, a folder of historical records — processing them one at a time is tedious. Google Drive allows you to upload multiple PDFs and open each with Google Docs individually, but for large batches this is still time-consuming.

For bulk OCR, consider using Google Cloud Vision API (free for limited volumes, pay-per-use above that), Adobe Acrobat's batch OCR processing (requires a paid subscription), or the open-source Tesseract OCR engine which can be run from the command line on multiple files with a simple script. For non-technical users, ABBYY FineReader offers batch processing with a free trial.

What happens to the original scan?

Adding an OCR text layer does not change or overwrite the original scan. The page image remains exactly as scanned — every pixel is preserved. The text layer is invisible and sits beneath the image, used only by search functions and accessibility tools. This means the document looks identical before and after OCR. A human reading the document sees the original scan; a search engine or screen reader also sees the text layer.

File size typically increases slightly after OCR — the hidden text layer adds a small amount of data, usually a few percent of the original size. For a 10-page scanned PDF that's 5MB, the OCR'd version might be 5.2MB. This is negligible for most purposes.

Making future scans searchable from the start

The best solution for searchability is to use scanning software that applies OCR automatically during the scan. Google Drive's mobile scanner (available in the Google Drive app for Android and iOS) automatically creates searchable PDFs. Microsoft Lens also applies OCR during scanning. Apple's document scanner in iOS creates image-based PDFs by default, but uploading to Google Drive and opening with Google Docs immediately adds searchability.

For physical scanners and photocopiers in offices, many modern models have built-in OCR settings. Check your scanner's software settings for options like 'searchable PDF' or 'OCR output' — enabling these means every scan you produce is immediately searchable without a post-processing step.

PDF OCR — coming soon to PDF99

Browser-based OCR to make scanned PDFs searchable. Free, no sign-up.

View all PDF tools →