DetermindsStart a project
Services
Industries
Case studies
Products
Company
Start a project
Engineering

Why Arabic OCR is still hard, and how to build reliable pipelines

In Arabic document pipelines, the condition of the page affects accuracy before the script does, because stamps, skew, shadow, and handwritten corrections arrive in the same batch as clean printed fields.

Why Arabic OCR is still hard
Contents

A document intelligence requirement usually arrives as a single line asking for structured data, while the archive behind it holds Arabic text, stamps, handwriting, and photographs taken at an angle.

Closing the gap between a client’s first sample and the full archive behind it is where most of the engineering effort goes.

The Arabic script is the first challenge

Arabic is cursive. Letters join, and a letter’s shape changes depending on whether it sits at the start, middle, or end of a word. Diacritics may or may not be present. Text runs right to left while embedded numbers run left to right, so a single line can change direction twice. A recognition model tuned on Latin script does not degrade gracefully on Arabic, and its errors look like plausible words.

The common response is a bigger model. That helps, and the larger lever lies elsewhere: on real documents, recognition quality is dominated by what happened to the page before it reached the pipeline.

Page condition drives recognition quality

A document that has been printed, signed, stamped, filed in a folder, retrieved, photocopied, and then photographed on a phone under fluorescent light carries every one of those events in its pixels. The recurring issues are:

  • A stamp printed over the field you need, usually in a color chosen for contrast against paper rather than against text
  • Skew, because pages are rarely photographed square, and perspective distortion from a phone held at an angle over a desk
  • Shadow across a third of the page, cast by the person taking the photograph
  • A handwritten correction next to a printed total, where the handwriting is the operative value
  • Two documents in one image, captured together to save time

Preprocessing for these issues moves accuracy more than a model change does. Deskew, dewarp, shadow normalization, and stamp channel separation are the steps that make the difference.

Recognition is the more straightforward half. The larger task is knowing that a particular number is the total rather than a line item, on a layout the pipeline has never seen.

Extraction depends on business rules as much as machine learning

Once the characters are recognized, the pipeline still has to decide what they mean. On an invoice, the total, the tax, the date, the supplier identifier, and the reference number all need to land in specific fields, and the layout carrying them differs for every supplier and changes without notice.

The approach that works is specific. It starts with a real sample, grouped into document families to show which ones carry the volume. A small number of layouts nearly always produces most of the pages, and the long tail is best left out of the first phase of automation. Extraction is then written against those families with the business rules attached, so the pipeline can tell when it is wrong:

  • Totals must reconcile against the sum of line items
  • Dates must fall inside a plausible range for the batch
  • Identifiers must match a known format, and where possible resolve against a master record
  • Currency and tax rate must be consistent with the supplier’s history

A field that fails validation shows the system working as designed. What matters is where that field goes next.

Confidence routing is the core of the design

Human review is part of what makes the economics work. A pipeline that passes most documents automatically with measured per-field accuracy, and routes the rest to a review screen built for speed, delivers more value than full automation with a higher headline figure and unchecked errors on individual fields.

Confidence thresholds are therefore set per field rather than per document, because a slightly wrong supplier name is recoverable and a wrong total is not. Corrections made in review feed back as training signal, so the review queue shrinks over the first few months.

Accuracy measured per field, on the client’s own documents

An overall accuracy figure for a document pipeline says little, because it averages the easy fields, which are printed and machine-set, with the hard ones, which are handwritten or overprinted. Determinds reports accuracy per field on the client’s own documents and runs the measurement on a held-out sample the client chooses.

For the same reason, a short paid proof comes before a full proposal. A proof run on your most difficult documents produces a figure you can trust and a clear view of which families are worth automating, and which are best kept manual.

Where to start

For an organization with a stack of paper and a team retyping it, the first step is a representative sample of real documents, deliberately including the most difficult ones. Counting what the sample contains settles most of the decisions that follow, because they depend on the distribution.

Determinds delivers that first step as a short paid proof, and its findings hold whoever builds the pipeline afterwards.

01  Frequently asked

Questions about Arabic document pipelines

Common questions before a proof begins.

What accuracy should we expect on our documents?

Determinds measures accuracy per field, on a held-out sample you choose from your own documents, and quotes a figure once it has seen that sample. An overall accuracy figure would average easy printed fields with hard handwritten ones and hide the fields that matter most to you.

Does this work on photographs rather than scans?

Yes, and in this region most of the input is photographs. Photographs are also where most of the accuracy is won or lost, through deskew, dewarp, shadow normalization, and separating a stamp from the text underneath it. That preprocessing moves results more than a model change does.

Can it read handwriting?

Yes, for handwritten numerals and short corrections beside printed fields, with a review step. Free-form handwritten paragraphs go to a person, and wherever handwriting carries the operative value, the pipeline routes it to review by design.

Can the pipeline run inside our own environment?

Yes, on-premise or in your private cloud tenancy, with self-hosted models where documents must stay inside your boundary. That requirement affects model selection and cost, so it is best settled early.

How many documents do you need to begin?

A representative sample that reflects the real spread of your documents, including the difficult ones. That is enough to identify the document families and show which are worth automating.

How long from proof to production?

A first production release typically follows 8 weeks after the short proof, with the timeline usually set by integration and access more than by the extraction work itself.

Paid proof

A short paid proof on your own documents

You will get measured per-field accuracy and a clear view of what to automate first.