India · Mumbai
Document OCR API for Indian businesses
Mahad OCR is a document OCR API built and operated by Mahad Information Technology Private Limited, a company registered in Mulund, Mumbai. It reads Aadhaar and PAN cards (format checks as evidence only), Indian passports (MRZ check digits verified on the server), CVs in English and Hindi, and invoices, and returns every field as JSON marked ✓ confirmed or ⚠ check it.
Fields returned
| Key | What it holds |
|---|---|
| Aadhaar | id_number, full_name, date_of_birth, sex — Verhoeff check digit on a full 12-digit number (evidence only) |
| PAN | id_number, full_name, date_of_birth — 10-character format check (evidence only) |
| Indian passport | All data-page fields including father_name where printed — MRZ check digits on number, birth date and expiry |
| CV / resume | Contact details, every job, education, skills, languages — English, Hindi or mixed |
| Invoice | Supplier, invoice number, dates, subtotal, tax, total, line items — totals checked by arithmetic |
How ✓ and ⚠ work here
- ✓ Passport number, birth date and expiry are proven by ICAO 9303 check digits where the MRZ is readable.
- ✓ Identity fields that two independent readers read the same way.
- Aadhaar and PAN number checks are structure checks only — there is no UIDAI or Income Tax Department lookup.
- ⚠ Anything read by one reader only, disputed, or unreadable is marked for a person to check.
GST invoices
Invoices are read with the standard invoice fields (supplier, number, dates, subtotal, tax, total, line items). The engine has no dedicated GSTIN field and does not validate a GSTIN against the GST portal; a GSTIN appears only if the reader returns it as an extra printed field.
Hindi and other Indian languages
English and Hindi (Devanagari) are read, both by the AI readers and by the on-server reader, which has Hindi, Urdu, Arabic and English language packs. Other Indian scripts such as Tamil, Telugu, Bengali or Gujarati have no on-server language pack; a script the active route cannot read is reported as UNSUPPORTED_LANGUAGE_SCRIPT instead of returning noise. Test your own documents first.
Where data is processed
The servers that run the service are located in India. Reading a document uses AI sub-processors named in the privacy policy (Google Gemini and OpenAI for reading; DeepSeek for structuring CV text and optional CV wording), and Cloudflare sits in front of the service for network protection. The MRZ and on-server text reader involve no sub-processor.
How long documents are kept
Mahad OCR reads documents and returns data; it is not an archive. Original files are kept 30 days after reading, extracted results 90 days, failed reads 7 days, items waiting for review until reviewed plus 30 days, and documents of test accounts 24 hours. You can delete any document earlier with DELETE /v1/documents/{id}.
Mumbai team and contact
The company is registered in Mulund, Mumbai (CIN U62099MH2025PTC437790). For sales questions write to [email protected], for technical support to [email protected], or message us on WhatsApp at +974 5103 8000.
Example request and response
Fictional data, real response shape (trimmed to the keys that matter here). Full field list in the API reference.
curl -X POST https://api.mahadocr.com/v1/documents/upload \
-H "X-MahadDoc-API-Key: mk_live_YOUR_KEY" \
-F "[email protected]"
# 202 Accepted
# { "id": "DOC_3f9a1c7e2b40", "status": "processing", ... }
curl https://api.mahadocr.com/v1/documents/DOC_3f9a1c7e2b40 \
-H "X-MahadDoc-API-Key: mk_live_YOUR_KEY"{
"id": "DOC_3f9a1c7e2b40",
"status": "verified",
"result": {
"status": "verified",
"decision": "verified",
"document_type": "national_id",
"values": { "id_number": "ABCPT1234K", "full_name": "RAVI SPECIMEN TESTKUMAR",
"date_of_birth": "1990-04-15" },
"field_evidence": { "full_name": "corroborated", "date_of_birth": "corroborated" },
"verification": {
"evidence": [
{ "level": "L1", "rule": "pan_structure", "result": "passed", "fields": ["id_number"] }
]
}
}
}Limits
- No government lookups: Aadhaar, PAN, passport and GSTIN are never checked with UIDAI, the Income Tax Department, Passport Seva or the GST portal.
- Mahad OCR is not a KYC provider or an Aadhaar authentication user agency; it extracts data for your own process.
- A full Aadhaar number is returned as read — the API does not mask it.
- Files: PNG, JPG, PDF or Word, up to 10 MB; a PDF is read up to its first 5 pages.
How to do it
- 1Create a free Sandbox account at app.mahadocr.com and verify your email and phone.
- 2Create an API key and POST a document to /v1/documents/upload.
- 3Poll GET /v1/documents/{id} and use result.values; show ⚠ fields to a person.
- 4Move to a paid plan on the pricing page when you need more than 50 documents a month.
Frequently asked questions
Is there a free plan?
Yes. The Sandbox plan is free and includes 50 documents a month. Paid plans with higher quotas are on the pricing page.
Do you store documents?
Only for fixed periods: original files 30 days after reading, extracted results 90 days, failed reads 7 days, items waiting for review until reviewed plus 30 days, and test-account documents 24 hours. You can delete a document earlier at any time.
Where is my data processed?
The service runs on servers located in India. Reading uses the AI sub-processors listed in the privacy policy (Google Gemini, OpenAI, DeepSeek), and Cloudflare protects the network in front of the service.
Can you read GST invoices?
Yes, as invoices: supplier, number, dates, amounts and line items, with totals checked by arithmetic. There is no dedicated GSTIN field and no GST portal check.
Do you verify Aadhaar or PAN with the government?
No. The Aadhaar check digit and PAN format are checked as evidence that the number was read correctly. There is no UIDAI or Income Tax Department lookup.
Can we meet the team in Mumbai?
Write to [email protected] or message us on WhatsApp and we will arrange a call.
Which Indian languages are supported?
English and Hindi. Other Indian scripts are not covered by the on-server reader; test them before relying on them.