Most OCR tools give you text and a number between 0 and 1. The number looks precise, but it does not tell you why a value should be trusted, or what to do when it should not. Mahad OCR takes a different approach: every field it returns carries a plain mark and the evidence behind it.
✓ Confirmed — and why
A field is marked ✓ only when there is evidence for it. The evidence is named in the result, so you can see it:
- Two independent readers agree. Two separate AI readers read the document on their own. When both return the same value, that value is confirmed.
- The MRZ check digits prove it. On passports and many ID cards, the machine-readable zone carries check digits. When a value passes its check digit, it is proven by arithmetic, not by opinion.
- Taken from the file's own text. A PDF with a real text layer already contains the exact characters; they are read, not guessed from pixels.
⚠ Check it — and the reason
When the evidence is not there, the field is marked ⚠ with the reason, so a person knows exactly what to look at:
- Disputed: the two readers read different values. Both readings are shown side by side.
- Unreadable: the field is printed, but glare, blur or a cut-off edge hid it.
- Missing: a field the document type normally has was not found.
- Read once: only one reader returned a value.
It never invents a value
If a field cannot be read, it is left empty and marked ⚠ — it is never filled with a guess. For identity documents this matters more than anything: a wrong passport number is worse than an empty one.
What you do with it
The result has a single decision for the whole document: verified when every required field is
confirmed, review when at least one field needs a person, or rescan when the photo itself is
the problem. Your system can accept verified documents automatically and send the rest to the built-in review screen,
where a person corrects the ⚠ fields and approves.
{
"status": "review",
"result": {
"decision": "review",
"values": { "surname": "ERIKSSON", "document_number": "L898902C3" },
"field_evidence": { "surname": "corroborated", "document_number": "deterministic" },
"review_reasons": ["Expiry date: glare hid it — please check"]
}
}
That is the whole idea: fewer numbers to interpret, and a clear next step for every field.