Anonimatum
Back to blog
AI technology 12 August 2026 15 min read

OCR and scanned PDFs: anonymize documents without a text layer

A scanned PDF does not contain text: it contains an image of the page. That is why string-only anonymizers fail silently, and manual black-out on the image often leaves recoverable text if layers are not destroyed properly or if later OCR «reads» what you thought was hidden.

In public bodies, healthcare, education and law firms, the scanner remains the entry door for case files. This guide shows how to design a serious flow: OCR → detection → review → redaction, when to generate searchable PDF, when to move to Word OCR, and how to do it with EU processing.

Three different problems (do not mix them)

1. Anonymize the scan

Goal: publish or share without PII. You need OCR to detect, then effective redaction in the output PDF.

2. Make the PDF searchable

Goal: keep the scan’s appearance but allow search/copy (searchable PDF).

3. Reuse the content

Goal: edit in Word. You need conversion OCR (PDF to Word OCR), not only an invisible layer.

Choosing the wrong goal is the #1 reason «OCR» projects satisfy neither business nor compliance.

Scan quality: what nobody mentions and OCR pays for

  • Resolution: aim for 300 dpi or more for IDs, stamps and small print.
  • Contrast and skew: crooked or grey pages increase confusions (0/O, 1/l).
  • Mixed pages: a file with native + scanned pages must detect which lack text.
  • Language: OCR models vary by language and typeface; validate with real samples from your archive.
  • Stamps and handwritten signatures: may need visual zones in addition to text OCR.

Why manual black-out on scans is fragile

Painting black on the image in a generic editor does not guarantee irreversibility: information may remain in other layers or thumbnails, or the operator misses a page. Without OCR you also will not know if a second ID sits in the margin. A flow with automatic detection + preview reduces that operational risk.

Flow A — Anonymize scanned PDFs (recommended for publication)

  • Upload the PDF (or images/HEIC depending on support).
  • The system detects missing text layer and runs OCR.
  • Entities are proposed (names, IDs, addresses, etc.).
  • You review inclusions/exclusions in preview.
  • You generate the redacted PDF and verify with search on the result (if a layer exists) plus visual review.

Note: product step detail may vary by plan; the critical point is not skipping human review on high-risk documents.

Anonymize scanned PDFs with OCR in Anonimatum’s flow.

Scanned PDF anonymizer

Flow B — Searchable PDF without changing appearance

When archive and search matter but the scan image must be preserved (stamps, original layout), generate a PDF with a text layer. Use it to index files *after* deciding whether they still contain PII; if they must be published, anonymize first or ensure the layer does not re-expose data you only hid visually.

Create searchable PDFs with OCR.

Searchable PDF OCR

Flow C — PDF to Word OCR (when you need to edit)

If the team must rewrite or reuse content, convert to DOCX via OCR. Then apply the same precautions as in the Word guide: comments, metadata and verification. Do not convert and share the Word file without anonymization if the original scan had PII.

Get editable Word from scans.

PDF to Word OCR

Compliance checklist for scans

  • Is the document image-only, hybrid or native?
  • Are legal basis / purpose clear for OCR and anonymization?
  • Does processing run in the EU without external LLMs on content?
  • Who validates preview on sensitive case files?
  • Are verification evidences kept (samples, date, reviewer)?
  • Was the output tested by searching known identifiers?

How Anonimatum approaches it

Anonimatum integrates OCR into the anonymization flow and offers dedicated searchable PDF and PDF to Word OCR tools, with local AI on European infrastructure. That avoids chaining three different vendors (scanner → cloud OCR → redactor) with inconsistent transfers and criteria.

Large remittances of scans in ZIP?

Batch processing

Validate your document typology with Politeia Soft.

Contact