The problem
Invoices arrive in different formats and languages. Copying them into a sheet takes time, but an extractor that invents missing fields or loses an attachment creates a different problem.
What I built
The 32-node n8n workflow runs every 15 minutes. It reads the existing sheet once per batch, collects untriaged email, and splits out the attachments. It checks the text layer first and sends scans to Claude Vision OCR.
Checking the records
The workflow standardizes and validates the fields, then checks four duplicate keys. A row key prevents duplicate writes. Unreadable items leave a diagnostic record, and Gmail labels keep already-triaged messages from causing another model call.
What I tested
The records cover different invoice layouts, invoices written in an email body, missing invoice numbers, repeats, corrupt files, image-only PDFs, and European number formats. These are documented workflow results from the build.
What testing caught
One email can contain several invoices. I had to split the attachments into separate items so a successful first extraction didn’t hide the second one.
Recorded workflow evidence · August 2026. The figures here come from those test records.