Anonimatum

OCR + anonymization · Local AI in the EU

Scanned PDF anonymizer with OCR

Scanner or photo PDFs with no text layer: OCR inside the anonymization flow to detect and redact personal data before sharing.

A scanned PDF is not selectable text—it is a stack of images. Manual blacking-out or tools that only read the text layer miss names, dates and addresses on the scan. Anonimatum runs real OCR inside anonymization so those pixels become processable text and then automatic detection targets.

Processing stays on European infrastructure. Pages are not sent to external LLM APIs: detection combines patterns, word sets, whitelist/reverse lists and local AI. Review matches before redaction and export an anonymized PDF fit for archive, publication or third-party delivery.

If you only need a searchable text layer without redaction, use searchable PDF OCR. When the goal is minimization and GDPR on a scan, this flow is the right one: OCR + review + redaction in one environment.

How to anonymize a scanned PDF

  1. 1

    Upload the PDF or images

    Drop the scan into the app. The system detects a missing text layer and prepares server-side OCR in the EU.

  2. 2

    Review detection after OCR

    Patterns, word sets and local AI flag dates, addresses, amounts, QR codes and other sensitive data. Adjust matches before redacting.

  3. 3

    Download the anonymized PDF

    Export a file ready to archive, publish or share. Your original stays under account control.

What you get

  • Real OCR inside scanned PDF anonymization
  • Detection with patterns, word sets, whitelist and local EU AI
  • Active detectors: dates, addresses, amounts and QR
  • Preview and review before redaction
  • Processing on European infrastructure (GDPR)
  • Batch queue compatible when plan allows
  • Works alongside searchable PDF and PDF to Word OCR
  • No external LLM APIs in the detection pipeline

Who this flow is for

Public bodies and archives

Digitized case files that must be published or shared without identifying data.

Law and compliance teams

Scanned contracts or opinions where a text layer is missing or incomplete.

Healthcare and insurance

Photographed reports that still contain patient or policyholder data.

GDPR owners

Teams that need evidence the scan was OCR’d and reviewed before release.

Why anonymize scans here

OCR and redaction in the same EU environment—no third-party generative LLM APIs.

OCR inside anonymization

No external pipeline: the scan becomes text and sensitive data is flagged in one process.

Local AI in the EU

Contextual detection without shipping documents to third-party model APIs.

Patterns, word sets and whitelist

Deterministic rules plus human review to reduce false positives and negatives.

Batch queue

Process volumes of scans in queue when your plan allows, without skipping review.

Common use cases

Historical files

Old digitizations without prior OCR that must enter a transparency portal.

Mobile receipts

Photos of invoices as PDF that must be forwarded redacted.

Scanned minutes

Paper-signed minutes with names and addresses to hide.

Path to Word OCR

When you also need to edit, combine with PDF to Word OCR around anonymization as needed.

FAQ — scanned PDFs

Can I anonymize a PDF without selectable text?

Yes. The flow enables OCR to extract text from image pages, then runs detection and redaction.

Does OCR run outside the EU?

No. OCR and anonymization run on Anonimatum’s European infrastructure.

What if I only need a searchable PDF?

Use the searchable PDF OCR tool when you do not need redaction.

Which AI detectors work on scans?

Active options include dates, addresses, amounts/quantities and QR codes, plus patterns and word sets.

Is manual review required?

Review is recommended. Whitelist and reverse lists help tune what is redacted vs kept.

Where do I open the flow?

At https://app.anonimatum.com in the anonymization area.

Anonymize your scanned PDFs

Try the OCR-enabled flow in the Anonimatum app.

Open anonymization