OCR for Arabic scanned documents — accuracy & steps | Itqan

Blurry scans: repair, recognize, then compress.

Arabic offices still live on scanners: ministry letters, court annexes, bank statements, university transcripts, and stapled invoices arrive as photo-like PDFs. You can read them with your eyes, but you cannot search, copy a clause, or feed them cleanly into Word. This problem/solution guide focuses on Arabic OCR for scanned documents on Itqan — how to prepare pages, run OCR PDF, judge quality, repair broken files, compress oversized packs, and leave with a searchable PDF you can actually use.

The problem: why Arabic scans fight OCR

Optical Character Recognition turns page images into digital text. Arabic adds challenges that Latin-script scans often dodge:

  • Connected letters and shaping: forms change with position; poor resolution breaks joins.
  • RTL layout mixed with LTR numbers and English names on the same line (common in GCC letterheads).
  • Diacritics (tashkeel) that may be faint on photocopies.
  • Stamps, signatures, and official seals overlapping text.
  • Skewed phone photos of papers on desks instead of flatbed scans.
  • Low-contrast blue or red ink on forms.
  • Multi-generation photocopies with moiré and speckles.

Teams then waste hours retyping amounts into Excel, miss a clause in a 200-page annex, or fail accessibility checks because screen readers hear nothing. Compressing a terrible scan without OCR only makes a smaller terrible scan. Converting straight to Word without OCR often embeds a giant image per page — useless for editing.

The solution path on Itqan

Treat OCR as a pipeline, not a single miracle button:

  1. Stabilize the file. If the PDF will not open or pages are blank after chat-app forwarding, run PDF repair first.
  2. Prep the scan. Prefer 300 DPI greyscale or colour flatbed scans; avoid extreme JPEG compression; straighten pages; crop black borders; reshoot phone captures under even light.
  3. Run OCR. Open OCR PDF, upload, choose searchable PDF when you need to keep the visual page, or TXT when you only need the words.
  4. Quality check. Search three known unique tokens (a ministry name, a date, a contract number). Spot-check tables and stamped corners.
  5. Size for delivery. If email limits block you, use Compress PDF after verifying OCR — or compress carefully before OCR if upload limits force it, then re-check recognition.
  6. Next workflow. Edit in PDF to Word only after a text layer exists; archive with PDF/A when policy requires; protect or redact before external share.

Detailed button labels live in How to OCR a PDF. This article stays on Arabic scan reality: prep, quality, and failure modes.

Before / after — what success looks like

BeforeAfter a good Arabic OCR pass
Ctrl+F finds nothingSearch finds أسماء، أرقام قضايا، تواريخ
Copy-paste yields empty or mojibakeSelectable Arabic in correct reading order
Word conversion is one image per pageEditable paragraphs after OCR → Word
Screen readers silentAssistive tech can read the text layer
50 MB photo PDF for ten pagesSearchable PDF, optionally compressed for mail
Staff retype invoice totalsTotals copied once and spot-checked

Preparing Arabic scans for better accuracy

Capture settings

Aim for about 300 DPI for text documents. 200 DPI may limp for large headings; 150 DPI from a phone often fails on small footnotes. Use greyscale for black text on white; keep colour when seals and highlighter colours matter for humans even if OCR focuses on text. Disable “strong compression” presets on scanners when the destination is OCR rather than quick email.

Geometry and cleanliness

Flatten curled pages. Deskew so lines are horizontal. Remove dark binder edges that confuse layout analysis. Avoid fingers in the frame. If a stamp covers a critical amount, rescan with the stamp outside the number when legally acceptable, or accept manual correction for that cell.

Language expectations

Itqan’s OCR path is built with Arabic RTL handling so copied text keeps sensible order — a frequent failure mode in generic engines that emit reversed or disconnected letters. Mixed Arabic/English letterheads usually work when both scripts are sharp. Extremely stylized calligraphy or decorative basmalah lines may not be machine-readable; treat those as graphics.

When to repair or split first

Huge binders: split into chapters, OCR each, then merge — easier to reprocess a bad chapter. Corrupt files: repair before OCR. Password-locked scans: unlock if authorized. Education and records teams can also review the wider education solutions context for thesis and transcript packs.

Judging OCR quality like a reviewer

Never trust a green “success” toast alone. Run a human checklist:

  • Search a rare proper noun from page 1 and from a middle page.
  • Copy a full Arabic paragraph into a notepad; confirm letters are connected properly and not mirrored.
  • Check digit-heavy lines (IBAN fragments you are allowed to test, invoice totals, hijri/gregorian dates).
  • Inspect tables: are columns jumbled into one stream? If yes, keep the searchable PDF for archive and retype critical tables, or rescan clearer sources.
  • Compare page count before and after.
  • If you will redact later, remember OCR text layers can expose secrets — redact thoughtfully (see the redact vs password guide).

Accept that 100% accuracy on noisy Arabic photocopies is unrealistic. Target “searchable and mostly correct for operations,” then manually fix legal operative clauses.

Scenarios — problem to solution

Accounts payable: month-end Arabic invoices

Problem: fifty supplier scans, staff retyping VAT and totals. Solution: batch-prep at 300 DPI, OCR to searchable PDF, copy figures into the sheet, spot-check stamped pages. Compress only the email zip after QA. Hub entry points sit on PDF tools.

Legal: finding a clause in a scanned contract annex

Problem: 300 image pages, hearing tomorrow. Solution: repair if needed, OCR overnight in chunks, search the key phrase, export the hit pages, protect the working bundle. Do not watermark over text so heavily that OCR of future rescans fails.

University: thesis scan before editing

Problem: PDF→Word produces pictures. Solution: OCR first, then Word conversion, then human language edit. For final deposit requiring archival format, convert to PDF/A after the text is stable.

HR: bilingual policy handbook

Problem: accessibility audit fails. Solution: OCR, test with a screen reader on Arabic RTL sections, fix misread headings manually in an exported copy if required.

Common mistakes

  • Photographing pages at an angle in yellow lamp light, then blaming the OCR engine.
  • Heavy compress → OCR → surprise that digits died.
  • Skipping repair on a truncated WhatsApp PDF.
  • Assuming English-only OCR settings on a predominantly Arabic letter (use the Arabic-capable path on Itqan).
  • Deleting the original scan before verifying the searchable output.
  • Sending OCR output externally without redacting personal data.

A fuller pipeline for messy Arabic binders

Real binders rarely arrive as one clean PDF. You may receive a mix of native digital letters, phone photos, and flatbed scans stapled into one file by a previous clerk. Start by opening the document and classifying pages: already selectable text versus image-only. You can OCR the whole file, but when time is short, extract image-only ranges, OCR those, and merge back with the native-text sections using merge tools from the hub. That hybrid approach avoids re-OCRing crisp digital pages (which can occasionally degrade layout analysis) while still unlocking the photographic chapters.

For phone photos, preprocess mentally before upload: one page per frame, no perspective trapezoids, no glossy reflections from plastic sleeves. If the only copy is a warped photo, OCR will still run — expect more digit errors near edges. Correct financial figures by hand; use search mainly to navigate.

After OCR, decide the output role. Searchable PDF is the default archive and share format: humans see the familiar scan, machines use the hidden text layer. TXT is for dumping into translation tools or spreadsheets when layout does not matter. If you need both, run once to searchable PDF, then copy text from high-value pages rather than trusting a full TXT dump of a 400-page binder without review.

Compression strategy matters. If the upload fails due to size, compress lightly, OCR, then evaluate. If OCR is already done and only email rejects the attachment, compress the searchable PDF and re-test Ctrl+F on a sample of pages — aggressive downsampling can blur glyphs so future re-OCR (if someone tries) becomes worse; keep a master uncompressed searchable copy in records when possible.

Finally, integrate with protection: once the dossier is searchable, secrets are easier to discover by friend and foe. Apply redaction on release copies and password protection for narrow distribution lists, following the compare guide on redact versus watermark versus password. Business admin teams coordinating invoice and contract scans can also skim business PDF workflows for neighbouring steps such as merge and e-sign prep.

Pre-flight checklist (print or pin)

  • File opens? If not → repair.
  • DPI roughly 300? If not → rescan critical pages.
  • Pages upright and deskewed?
  • Arabic letterforms sharp, not ink-bled blobs?
  • OCR to searchable PDF.
  • Three successful searches across the document.
  • Copy-paste sample paragraph looks correct RTL.
  • Compress only if delivery requires it; keep a master.
  • Redact/protect before external send.
  • Store under the correct retention folder — OCR is not a backup policy.

Privacy note for scanned personal data

Scans often contain national IDs and signatures. Prefer temporary processing awareness from the security page and the privacy guide; download to controlled storage; redact before wide distribution. OCR makes secrets easier to find — which is good for you and risky if the file leaks.

Where to go next

Start with OCR PDF. If the file is sick, repair first. If it is huge, compress with eyes open. Learn click steps in the OCR how-to, and keep the PDF hub open for merge, Word conversion, and archival PDF/A when your Arabic dossier must both search and last.

FAQ

Should I OCR before or after organizing pages?

Organize and rotate first when page order is wrong. OCR on shuffled pages creates searchable but misleading archives. Fix orientation in your scanner software or with PDF organize tools, then run recognition.

Does OCR add tashkeel automatically?

Usually not for modern Arabic business text. Expect bare letters; add diacritics manually if your audience requires them for liturgical or pedagogical publishing.

Is one OCR pass enough for legal archives?

OCR plus human review on operative clauses is the practical standard. Certified copies may still require original scans per local law. Keep the image-bearing searchable PDF as the working archive, not a TXT dump alone.

Where is Arabic-specific troubleshooting?

This guide covers RTL, stamps, and prep. Click-by-click tool buttons are in the OCR how-to, and neighbouring utilities live on the PDF hub.

A five-minute QA desk after every OCR batch

Do not close the batch when the progress bar finishes. Spend five minutes on quality: search for a party name, an article number, and a date. If one fails, open the page visually and decide whether to rescan or make a limited manual fix. Log an approximate success rate in the team note so you know which scanner or staffer needs coaching. Before external sharing, apply redaction or password protection and follow the habits on the security page.

For documents you will later translate or summarize with AI-assisted tools, wait until text is searchable — models on image-only pages produce weak results. Run OCR first, spot-check, then move to language tools.

If Arabic and Latin share a page, run QA in both languages. A perfect Arabic hit with broken English numbers still fails finance handoff. Compress only after QA, not before, unless upload limits force an earlier pass — then re-OCR after a clearer rescan when possible. Education teams packaging theses can also start from the education solutions page once the text layer is trustworthy.

Back to blog