TIL / Checksum the file before indexing it, not after
Checksum the file before indexing it, not after
The problem
In a document Q&A system, users re-upload the same attachment more than you’d expect: the same PDF shared in two threads, or the same file re-attached after an edit. Indexing it twice wastes embedding calls and, worse, doubles up retrieval results for the same content.
The fix
Hash the file’s bytes (SHA-256 is plenty) as soon as it’s uploaded, before any parsing or embedding work starts, and check that hash against what’s already indexed for this tenant.
import hashlib
def file_checksum(data: bytes) -> str:
return hashlib.sha256(data).hexdigest()
checksum = file_checksum(upload_bytes)
existing = await db.get_document_by_checksum(tenant_id, checksum)
if existing:
return existing # skip re-parsing and re-embedding entirely
Gotcha
Scope the uniqueness check to the tenant (or whatever isolation boundary the system has), not globally - two different tenants uploading the same public PDF should not end up sharing one indexed copy across a multi-tenant boundary, even though the bytes are identical.