AI Integration

Invoice Data Extraction: What the Cloud Vendors Actually Document

Every invoice automation pitch leads with an accuracy percentage. None of the three major document APIs publishes one. Here is what Google, AWS and Microsoft do document, the page limits that shape your architecture, and the work that is still yours after the API returns.

Every pitch for invoice automation opens with an accuracy number. Ninety-nine percent, ninety-nine point five, straight through on the first pass. None of the three major document extraction APIs publishes such a number for invoices, and that absence is the single most useful fact when you are scoping a build. What they do publish is precise, checkable and much more useful: the shape of the response, the page and size limits, and the languages covered. Everything below is read from the vendors’ own documentation on 30 September 2026, and it is the material we use when a finance team asks us what an invoice pipeline will actually cost to run.

Accuracy on your invoices is a property of your suppliers, not of the model. The vendors know this, which is why they document response structure and limits and leave the percentage to the people selling implementations.

The three services, and the shapes they return

The services are more different than the marketing suggests, and the difference that matters is what comes back. Amazon returns a labelled expense structure. Microsoft returns key-value pairs plus line items. Google returns extracted entities from a specialised invoice processor. You will write different post-processing for each, and swapping one for another later is not a configuration change.

Amazon is the most explicit about structure. The AnalyzeExpense documentation describes a response of ExpenseDocuments containing SummaryFields for document level values and LineItemGroups holding LineItems and their LineItemExpenseFields. Every field carries a LabelDetection, a ValueDetection, a normalised Type and a Geometry, and both the label and the value carry their own confidence.

Microsoft’s prebuilt invoice model states that the key-value pairs and line items are returned in the documentResults section of the JSON output, and that it is “where to find all the fields from the invoice such as invoice ID, ship to, bill to, customer, total, line items and lots more”. The documentation also states plainly that “The model currently supports invoices in 27 languages”, and points at a separate schema page for the full field list rather than printing it.

Google ships a dedicated Invoice Parser within Document AI, and documents it mainly through its limits rather than its field list. That is the section worth reading before you design anything.

ServiceWhat comes backDocumented in the vendor docs
AWS Textract AnalyzeExpenseExpenseDocuments with SummaryFields and LineItemGroupsNormalised field Type, plus separate confidence on label and value detection
Azure AI Document IntelligenceKey-value pairs and line items in documentResults27 invoice languages; the full field list sits in a separate schema document
Google Document AI Invoice ParserExtracted entities from a specialised processorExplicit page and file size limits per request mode
Read from each vendor’s own documentation on 30 September 2026. The row that decides most integrations is the middle column, because it is the one your code has to be written against.

Page and size limits shape the architecture before accuracy does

Google is the most specific of the three about what one request may contain, and the numbers are smaller than most people assume. The Document AI limits page gives a maximum file size of 40 MB for online processing requests and 1 GB for batch, a cap of 5,000 files per batch processing request, and a 40 megapixel image resolution limit that does not apply to PDF files.

The invoice specific numbers are tighter still: 15 pages maximum for online synchronous requests, 200 pages maximum for batch asynchronous requests, and 30 pages in imageless mode online. A consolidated statement from a large supplier goes past fifteen pages regularly, which means the synchronous path you prototyped with is not the path you can run in production.

  • Decide the split before you build. Synchronous for the common case, batch for anything over the page limit, and a deterministic rule that chooses between them rather than a try-and-fall-back.
  • Design for the multi-invoice PDF. Suppliers email one file containing eleven invoices. Splitting that reliably is its own piece of work and it belongs before extraction, not after.
  • Hold the original. Whatever you extract, the PDF is the record. Auditors ask for the document, not for your parsed JSON.
  • Expect the limits to be per request, not per day. Throughput planning is a queue problem, and the queue is where an invoice pipeline usually falls over first.

None of this is exotic engineering. It is the difference between a demo that reads one clean PDF and a service that receives whatever a thousand suppliers send on the last working day of the month. The same distinction runs through our notes on AI workflow automation, where the interesting work is almost always the queue rather than the model.

Nobody publishes an accuracy number, and that is honest

Search the three sets of documentation for a headline accuracy figure on invoices and you will not find one. This is not an oversight. Extraction quality on a given estate depends on how many of your suppliers send native PDFs against scans, how many send them in a language the model covers, how consistent their layouts are, and how many of your fields are standard rather than a purchase order reference printed in the footer.

A vendor cannot know any of that, so a published number would either be measured on a benchmark that is not your mail or be marketing. The practical consequence is that the only accuracy figure worth anything is one you measure yourself, on a sample of your own invoices, before you commit to a design.

  • Take 200 real invoices, not clean ones. Sample across your top suppliers by volume, and include the scanned ones and the awkward ones on purpose.
  • Score field by field, not document by document. An invoice with the total right and the purchase order number wrong is not eighty percent correct, it is an exception.
  • Separate extraction failure from validation failure. A correctly read total that does not match the purchase order is a business problem, not a model problem, and the two get fixed by different people.
  • Record the straight through rate. The share of invoices that reach the ledger with no human touch is the only number a finance director should be asked to care about.

That exercise takes a few days and settles arguments that otherwise run for months. It also tells you the shape of your exception queue, which is the part of the system that actually needs designing.

Confidence is a routing signal, not a quality score

Both label and value detections in the Amazon response carry a confidence, and the temptation is to treat a high number as a guarantee. It is not. Confidence tells you how sure the model is that it read the characters it thinks it read. It says nothing about whether the field it attached them to is the right one, which is the error that actually costs money on an invoice.

Use it as a router. Above your threshold, post automatically. Below it, queue for review. Choose the threshold from the sample above rather than from a blog post, and then watch what it does to the size of the queue, because a threshold that produces a queue nobody can clear is the same as having no automation at all.

One detail is worth knowing before you plan around a managed review product: Google’s Human in the Loop feature for Document AI was deprecated on 16 January 2024, as the documentation for it now states. The review queue, the reviewer interface, the audit trail of who corrected what and the feedback loop back into your rules are yours to build and yours to maintain. Budget for them as a product, because that is what they are.

“Human in the Loop (HITL) was deprecated on January 16, 2024.”

Google Cloud, Document AI documentation, read 30 September 2026

What is still yours after the API returns

The extraction call is the smallest part of an accounts payable pipeline and the only part a vendor sells you. Everything below is work you own regardless of which of the three you choose, and it is where implementation budgets actually go.

StageWhat it has to doWhere it usually goes wrong
IntakeReceive from mail, portal and EDI, deduplicate at the file level, split multi-invoice PDFsThe same invoice arriving twice by two routes a week apart
ExtractionCall the API on the right path for the page count, store the raw responseThrowing away the raw response and being unable to explain a value later
NormalisationDates, currencies, tax codes, supplier identity resolutionSupplier matching on name rather than on a resolved vendor record
ValidationPurchase order matching, tolerance rules, duplicate payment checksTolerances set once and never revisited as prices move
Exception handlingA queue, a reviewer interface, an audit trail, a route back into the rulesTreating it as a spreadsheet rather than as part of the system
PostingWrite to the ledger idempotently, with retries that cannot double postRetry logic that creates the duplicate payment the validation stage was built to prevent
The six stages of an invoice pipeline. Only the second one is bought; the other five are built, and the fifth is the one most often discovered late.

Two of those stages deserve naming as risks rather than tasks. Idempotent posting is the one that causes real financial damage when it is wrong, because a retry that posts twice creates a duplicate payment that your duplicate checks have already waved through. And the audit trail on the exception queue is the one auditors ask about, which makes it a design requirement rather than a nice to have: who saw this invoice, what did they change, and when. The same discipline we apply to authorization and audit on customer portals applies here for the same reason.

If you want the wider implementation frame, with workflow architecture, governance and rollout laid out end to end, our guide to AI invoice processing automation covers the programme around this. This post is the narrower question underneath it: what the extraction layer actually promises, in the vendors’ own words.

If you are scoping an invoice pipeline and want the 200 invoice measurement run properly before anyone commits to a vendor, or you have one in production and cannot explain its straight through rate, talk to us. Examples of the systems we have built are on the work page.

How accurate is AI invoice data extraction?

No major vendor publishes an accuracy figure for invoices, and the omission is reasonable: results depend on your suppliers, their layouts, their languages and whether they send native PDFs or scans. Measure it on 200 of your own invoices, score field by field rather than document by document, and report the straight through rate.

What is the page limit for Google Document AI invoice processing?

The Document AI limits page gives 15 pages maximum for online synchronous requests to the Invoice Parser, 200 pages for batch asynchronous requests and 30 pages in imageless mode online, with a 40 MB file size limit online and 1 GB for batch. Consolidated supplier statements exceed the synchronous limit routinely, so the batch path is not optional.

How many languages does Azure Document Intelligence support for invoices?

The prebuilt invoice model documentation states that the model currently supports invoices in 27 languages, with the full list on a separate language support page. If you receive invoices outside that set, plan for a manual path rather than expecting degraded extraction.

What does AWS Textract AnalyzeExpense return?

A set of ExpenseDocuments, each with SummaryFields for document level values and LineItemGroups containing LineItems and their LineItemExpenseFields. Every field carries a label detection, a value detection, a normalised type and geometry, and both the label and the value detection carry their own confidence score.

Should low confidence invoices go to a human?

Yes, and that queue is yours to build. Google deprecated its Human in the Loop feature for Document AI on 16 January 2024, so the reviewer interface, the audit trail and the feedback loop into your rules are part of your system. Set the threshold from your own sample and watch the queue size, because a queue nobody can clear is the same as no automation.

Which invoice extraction service should we choose?

Choose on response shape and on limits rather than on claimed accuracy, because the response shape is what your code is written against and the limits decide your architecture. Then run the same 200 invoice sample through the shortlist and compare straight through rates on your own mail, which is the only comparison that transfers.

Need help building this?

Let our team build it for you.

Dude Lemon builds production-grade web apps, APIs, and cloud infrastructure. Get a free consultation and project proposal within 48 hours.

Start a project