Use case · Arabic documents

Arabic OCR API

Arabic OCR with Mahad OCR reads Arabic and bilingual Arabic/English documents — GCC identity cards, CVs, invoices and letters — and returns the text and fields as JSON. Text is kept in the script it was printed in; on identity cards the Arabic name is returned beside the English one, and Arabic CVs get Arabic, right-to-left headings in the finished CV.

Fields returned

KeyWhat it holds
name_arabicOn ID cards, beside full_name (shown side by side in field_review)
valuesFields in the script printed; a bilingual non-name value such as “WORK / عمل” publishes its Latin part
textThe document’s text (general documents), up to 20,000 characters
languagesLanguages the reader saw on the document
cvArabic CVs: the full structure, in Arabic

How ✓ and ⚠ work here

  • Arabic identity cards follow the same ✓ / ⚠ rules as any card: MRZ check digits, two readers agreeing, and the country number rules.
  • Saudi documents: a date printed in both Hijri and Gregorian is cross-checked when both are returned.
  • ⚠ Dates in a non-Gregorian calendar are not converted silently; they go to review.
  • Names are always kept in the script printed — never transliterated by the engine.

Example request and response

Fictional data, real response shape (trimmed to the keys that matter here). Full field list in the API reference.

curl -X POST https://api.mahadocr.com/v1/documents/upload \
  -H "X-MahadDoc-API-Key: mk_live_YOUR_KEY" \
  -F "[email protected]"

# 202 Accepted
# { "id": "DOC_3f9a1c7e2b40", "status": "processing", ... }

curl https://api.mahadocr.com/v1/documents/DOC_3f9a1c7e2b40 \
  -H "X-MahadDoc-API-Key: mk_live_YOUR_KEY"
{
  "id": "DOC_3f9a1c7e2b40",
  "status": "verified",
  "result": {
    "document_type": "letter",
    "values": { "issuer": "شركة الاختبار للتجارة", "date": "2026-09-14", "subject": "خطاب تعريف" },
    "field_evidence": { "issuer": "model", "date": "model", "subject": "model" }
  }
}

Limits

  • Arabic and English are the first-class languages. Other scripts depend on the reading route; a script that route cannot read is reported as UNSUPPORTED_LANGUAGE_SCRIPT rather than returned as noise.
  • Handwritten Arabic is not a supported use case — test your own samples first.
  • General documents are read by one reader, so their fields are ⚠ for a person to check.
  • Files: PNG, JPG, PDF or Word, up to 10 MB; a PDF is read up to its first 5 pages.

Frequently asked questions

Does Mahad OCR translate Arabic to English?

No. Text is returned in the script printed. On bilingual documents both are kept where the fields allow (for example name_arabic beside full_name).

Can I try Arabic OCR without an account?

Yes — the free Arabic OCR tool reads the text of an image or PDF on our servers, without an account, and does not store the file. The API adds structured fields, evidence and review.

Which Arabic documents work best?

Printed documents: GCC ID cards, residency permits, CVs, invoices and letters. Photos should be sharp, flat and without glare.

How are Hijri dates handled?

Where a date appears in both Hijri and Gregorian form, both are compared. A date only in Hijri is sent to review instead of being converted without a check.

Are Arabic names transliterated?

No. Names are kept as printed. If a card prints the name in both scripts, both are returned.

What other languages are supported?

English and Arabic first. Hindi is supported for CVs. For other scripts, test your documents in the live demo before you rely on them.

Related