Feature

Duplicate Detection

Identify when the same file has already been submitted — and reuse the result for free.

What is Duplicate Detection in Mahad OCR?

Available now. Every upload is hashed with SHA-256. Input: Any uploaded document. Limitation: Detection is byte-exact.

Available now

What it does

Every upload is hashed with SHA-256. An identical file already processed for your tenant returns the stored extraction instead of re-running the engines, so you are not charged twice. Separately, any image identical to a previously blocked document is sent to review.

How it works

  1. 1File bytes hashed on upload
  2. 2Existing identical document found → cached result returned at zero engine cost
  3. 3Hash matched against blocked documents → sent to review

Sample result

Illustrative response shape using synthetic data.

{ "cache": { "hit": true, "source_document_id": "DOC_…" }, "engine": { "cost_usd": 0.0 } }

Supported inputs

Any uploaded document.

Fields returned

  • cache.hit
  • cache.source_document_id
  • pixel_hash duplicate flag

API

Automatic on every upload.

Security & privacy

Hash comparison never crosses tenant boundaries for cached results.

Current limitations

Detection is byte-exact. A re-photographed or re-compressed version of the same document is a different file and will not match.

Try it on your own document

Run the live demo, or create an account and get an API key.