Feature
Duplicate Detection
Identify when the same file has already been submitted — and reuse the result for free.
What is Duplicate Detection in Mahad OCR?
Available now. Every upload is hashed with SHA-256. Input: Any uploaded document. Limitation: Detection is byte-exact.
What it does
Every upload is hashed with SHA-256. An identical file already processed for your tenant returns the stored extraction instead of re-running the engines, so you are not charged twice. Separately, any image identical to a previously blocked document is sent to review.
How it works
- 1File bytes hashed on upload
- 2Existing identical document found → cached result returned at zero engine cost
- 3Hash matched against blocked documents → sent to review
Sample result
Illustrative response shape using synthetic data.
{ "cache": { "hit": true, "source_document_id": "DOC_…" }, "engine": { "cost_usd": 0.0 } }Supported inputs
Any uploaded document.
Fields returned
- cache.hit
- cache.source_document_id
- pixel_hash duplicate flag
API
Automatic on every upload.
Security & privacy
Hash comparison never crosses tenant boundaries for cached results.
Current limitations
Detection is byte-exact. A re-photographed or re-compressed version of the same document is a different file and will not match.
Related document types
Try it on your own document
Run the live demo, or create an account and get an API key.