OCR + anonymization · Local AI in the EU
Scanner or photo PDFs with no text layer: OCR inside the anonymization flow to detect and redact personal data before sharing.
A scanned PDF is not selectable text—it is a stack of images. Manual blacking-out or tools that only read the text layer miss names, dates and addresses on the scan. Anonimatum runs real OCR inside anonymization so those pixels become processable text and then automatic detection targets.
Processing stays on European infrastructure. Pages are not sent to external LLM APIs: detection combines patterns, word sets, whitelist/reverse lists and local AI. Review matches before redaction and export an anonymized PDF fit for archive, publication or third-party delivery.
If you only need a searchable text layer without redaction, use searchable PDF OCR. When the goal is minimization and GDPR on a scan, this flow is the right one: OCR + review + redaction in one environment.
Drop the scan into the app. The system detects a missing text layer and prepares server-side OCR in the EU.
Patterns, word sets and local AI flag dates, addresses, amounts, QR codes and other sensitive data. Adjust matches before redacting.
Export a file ready to archive, publish or share. Your original stays under account control.
Digitized case files that must be published or shared without identifying data.
Scanned contracts or opinions where a text layer is missing or incomplete.
Photographed reports that still contain patient or policyholder data.
Teams that need evidence the scan was OCR’d and reviewed before release.
OCR and redaction in the same EU environment—no third-party generative LLM APIs.
No external pipeline: the scan becomes text and sensitive data is flagged in one process.
Contextual detection without shipping documents to third-party model APIs.
Deterministic rules plus human review to reduce false positives and negatives.
Process volumes of scans in queue when your plan allows, without skipping review.
Old digitizations without prior OCR that must enter a transparency portal.
Photos of invoices as PDF that must be forwarded redacted.
Paper-signed minutes with names and addresses to hide.
When you also need to edit, combine with PDF to Word OCR around anonymization as needed.
Yes. The flow enables OCR to extract text from image pages, then runs detection and redaction.
No. OCR and anonymization run on Anonimatum’s European infrastructure.
Use the searchable PDF OCR tool when you do not need redaction.
Active options include dates, addresses, amounts/quantities and QR codes, plus patterns and word sets.
Review is recommended. Whitelist and reverse lists help tune what is redacted vs kept.
At https://app.anonimatum.com in the anonymization area.