TIL / A headless-LibreOffice fallback for documents an AI service can't read
A headless-LibreOffice fallback for documents an AI service can't read
The problem
A document-intelligence service (in this case, Azure AI Document Intelligence) is tuned for PDFs. Some DOCX files it can’t read directly, even though the file itself opens fine in a word processor. Failing the whole upload on that is a bad experience for a file that’s perfectly readable.
The fix
On a DOCX parse failure, convert the file to PDF with headless LibreOffice and retry the same document-intelligence call against the converted file instead of the original.
libreoffice --headless --convert-to pdf --outdir /tmp/converted input.docx
try:
result = await doc_intelligence.analyze(docx_bytes)
except UnsupportedFormatError:
pdf_bytes = convert_to_pdf(docx_bytes) # shells out to the command above
result = await doc_intelligence.analyze(pdf_bytes)
Gotcha
Headless LibreOffice is not thread-safe against a shared user profile - running two
conversions concurrently against the same profile directory corrupts both. Give each
conversion its own isolated profile directory (-env:UserInstallation=file:///tmp/<job-id>)
rather than serializing every conversion through a single lock.