Use case · PDF documents
PDF to structured data API
PDF to data with Mahad OCR reads a PDF — born-digital or scanned — and returns its fields as JSON: invoices, receipts, bank statements, certificates, CVs, ID documents, or any other document as general fields and text. When every page of the PDF has its own text layer, that exact text is used, so characters are not re-guessed; scanned pages are read by OCR. You can also ask for the result as Excel, Word or PDF files.
Fields returned
| Key | What it holds |
|---|---|
| result.document_type | Detected type (invoice, receipt, bank_statement, certificate, cv, passport, id, visa, or a general document) |
| result.values | The fields for that type (for example vendor, invoice_number, total_amount on an invoice) |
| result.field_evidence | Why each field is trusted: e.g. arithmetic, mrz, corroborated — or a single reader (⚠) |
| result.missing_required | Required fields of that type that were not found |
| text | General documents: the document’s text, up to 20,000 characters |
| result.outputs | Links to the files asked for with outputs=xlsx, docx or pdf |
How ✓ and ⚠ work here
- ✓ Born-digital PDFs: when every page has a text layer, the exact text is used instead of OCR.
- ✓ Invoices and receipts: totals are checked by arithmetic (subtotal + tax = total).
- ✓ Identity documents inside a PDF follow the identity rules: two readers, MRZ check digits, country number rules.
- ⚠ Other fields are read by one reader and marked for a person to check.
Outputs
JSON is always returned. Add outputs=xlsx, outputs=docx or outputs=pdf (comma-separated for several) to the upload to also get a spreadsheet, a Word file or a clean PDF of the result; each is downloaded from its own link once reading has finished.
Want to try it without code?
The free PDF to text tool reads a PDF in your own browser — nothing is uploaded — and the free PDF to Excel tool turns a table into a spreadsheet. The API adds structured fields, evidence, review and bulk processing.
Example request and response
Fictional data, real response shape (trimmed to the keys that matter here). Full field list in the API reference.
curl -X POST https://api.mahadocr.com/v1/documents/upload \
-H "X-MahadDoc-API-Key: mk_live_YOUR_KEY" \
-F "[email protected]" \
-F "outputs=xlsx"{
"id": "DOC_3f9a1c7e2b40",
"status": "review",
"result": {
"document_type": "bank_statement",
"values": { "bank_name": "Sample Test Bank", "account_holder": "TEST TRADING LLC",
"account_number": "0000123456", "currency": "QAR",
"period_start": "2026-08-01", "period_end": "2026-08-31" },
"field_evidence": { "bank_name": "model", "account_number": "model" },
"outputs": { "json": "/v1/documents/DOC_3f9a1c7e2b40",
"xlsx": "/v1/documents/DOC_3f9a1c7e2b40/outputs/xlsx" }
}
}Limits
- A PDF is read up to its first 5 pages; longer files go to review with “pages not processed”. Split long files before uploading.
- Files up to 10 MB.
- Handwriting is not a supported use case — test your own samples first.
- There is no direct connector to accounting or ERP systems; use the JSON or the Excel output.
How to do it
- 1POST the PDF to /v1/documents/upload; add doc_type if you know it, and outputs=xlsx (or docx, pdf) if you want files.
- 2Poll GET /v1/documents/{id} until result.status is not processing.
- 3Map result.values into your system; check the fields that are not confirmed in result.field_evidence.
- 4Download any output from result.outputs.
Frequently asked questions
Does it work with scanned PDFs?
Yes. Pages without a text layer are read by OCR. Clear, straight scans give the best results.
Is a digital PDF read differently?
Yes. When every page has its own text layer, that exact text is used, so characters are not re-read from an image.
How many pages can a PDF have?
Up to 5 pages are read per document. Longer files go to review with “pages not processed”.
Can I get Excel instead of JSON?
JSON is always returned. Add outputs=xlsx to also get a spreadsheet, or docx / pdf for Word and PDF files.
Which document types are recognised?
Invoices, receipts, bank statements, certificates, CVs, passports, ID cards and visas. Anything else is read as a general document: its fields and text, with nothing required.
Is there a free way to try?
Yes. The free browser tools read a PDF without uploading it, and the free Sandbox plan includes 50 documents a month through the API.