OCR
Long documents: the lines do not vanish on page thirty
Multi-page invoice extraction rarely fails loudly: the lines vanish in silence, with no error shown. Where OCR (optical character recognition) reads page by page, the Capture agent follows the continuity of the table, strips repeated headers and compares the sum of the lines with the stated total.
The problem
Lines disappear quietly
A maintenance contract, a works rate schedule, a consumption statement: these documents routinely run past thirty pages and carry tables that continue from one end to the other. This is the breaking point of most tools, and it is invisible: no error is shown, there are simply lines missing at the end. To place this capability in the chain, see the steps that follow OCR.
The fault cannot be spotted by eye. Nobody recounts four hundred lines to check that twelve are not missing. The exported file looks clean; it is merely incomplete. And incompleteness spreads into everything built on top of it: matching, spend analysis, claims.
Volume is not an exception reserved for construction contracts. A multi-site maintenance contract, a telecoms statement, an energy invoice with several metering points, a site batch with its measurement sheets: in indirect procurement as much as in production procurement, documents of several dozen pages are the rule as soon as the service is recurring. They also carry the highest amounts, and are therefore the ones it would be most useful to check.
A truncated document does not produce a discrepancy: it produces an absence of discrepancy. That costs more, because nothing signals it.
How continuity is held
The difficulty is not reading a page. It is knowing what connects page twenty-nine to page thirty.
- Table continuity across pages
- A line cut by a page break is put back together. A table resuming without a header is recognised as the continuation of the previous one, not as a new table.
- Repeated headers and footers
- Column labels reprinted on every page, legal mentions, pagination: identified as furniture, they never enter the lines.
- Intermediate subtotals
- Section, lot and page totals are marked as subtotals. Added to the lines, they would double part of the document.
- Carry-forwards from page to page
- Carry-forward mentions connect two pages without creating an extra amount. They behave as a pivot, not as a billed line.
- Sections, lots and levels
- A lump-sum price breakdown or a rate schedule is organised in nested sections. The position of each line in that tree is preserved.
- End-of-document consistency check
- The sum of the extracted lines is compared with the total shown by the document. Any gap is flagged as a reading anomaly before anything is used.
Two traps deserve naming, because they produce silent errors. The first is the column that changes meaning mid-document: it can carry a unit price in one section and a lump sum in the next. The second is the subtotal presented as an ordinary line, with no distinctive label. Add a third, rarer but costly: the appendix that restates a table already invoiced earlier in the document. In all three cases the structure of the document is read before the values, and it is the structure that decides how to interpret them.
The result is not declared correct: it is verifiable. Line count, recalculated total and printed total are displayed side by side.
What it changes
A long document stops being a gamble.
- The recalculated total and the printed total are shown together: completeness is checked at a glance, without recounting lines.
- Nested tables stay usable: a line keeps its section, its lot and its level, and therefore its reference price.
- Documents nobody reread come into the controlled scope: multi-year contracts, statements, pricing appendices.
- Comparing two versions of the same long document becomes possible, line by line, without manual cross-reading.
The benefit is measured less in time saved than in scope recovered. Long documents are precisely the ones people give up checking for lack of time, and therefore the ones where discrepancies pile up unseen. Bringing them into the controlled flow requires no change of tool: your ERP (enterprise resource planning system) remains the source of truth, and the extracted data is added to it. The exports come out in the spreadsheet and structured formats your teams already use, so the first analysis can be run in the same week, on real documents.
Capture
Gather and consolidate all your procurement data.
Collection by dedicated mailbox, upload, SFTP, scan, API or ERP export. Reads every format (native PDF, scan, photo, spreadsheet, structured feed), then extracts line by line: supplier, references, quantities, unit prices, discounts, taxes, terms.
The end of re-entry. A procurement history you can finally query.
What Capture detects
- Long documents and multi-page tables
- Unknown layouts
- Several languages and currencies
- Poor scans and handwriting
The subscription covers extraction and is counted in pages processed, with no length limit per document. Checking your negotiated terms is the next step and is sized with you: the price of extraction, tier by tier is set out there.
Nothing is lost. Everything can be checked, everything can be proven.
- The contract clause and the invoice line, highlighted side by side.
- Every extracted value stays linked to the exact place in the document where it was read.
- The same case produces the same decision, today as in six months: the rules are applied deterministically.
- No discrepancy is set aside in silence. Anything that matches no rule is raised, with its reason.
- The agent records what it did, in the order it did it: who, what, how much, when.
Frequently asked questions
Is there a limit to the number of pages in a document?
No. A document of several hundred pages is handled as a whole: tables keep their continuity and the line count is verified at the end. The subscription is counted in pages processed, not in documents.
How do I know no line has been lost?
The sum of the extracted lines is compared with the total printed on the document, and the number of lines read is reported. Any gap is flagged as a reading anomaly, with the page concerned.
Is a table that changes structure mid-document handled?
Yes. The structure is re-read at every section: a column carrying a unit price in the first part and a lump sum in the second is interpreted according to the section it sits in. Where the structure becomes uncertain, the area is flagged rather than guessed.
How do I check a Zylio conclusion?
Every discrepancy opens onto its evidence: the contract clause and the invoice line, highlighted side by side, with the calculation shown.
Measurable impact in every environment
More than 5 million procurement documents analysed
Between 1 and 7% of margin recovered
on the scope analysed
From 15 to 45% of time given back to teams, per FTE
depending on the scope and on data maturity
Zylio fits into your existing ecosystem.
The ERP runs the process. Zylio handles the exception and recovers the value that escapes it: invoices without a purchase order, line-by-line price discrepancies, duplicates and overbilling, off-contract spend.
Your data under high security.
Zylio meets the most demanding standards, and nothing is committed without your approval.
- Certifications
- Hosting
- Encryption
- Access
Continue on OCR
- Supplier contract OCRTurning a signed PDF into a price schedule you can genuinely check.
- Attachments and batchesFour invoices stapled into one PDF, four separate processes.
- What every invoice line carriesReference, quantity, unit, unit price, discount, VAT rate.
- Supplier statement OCRSpotting what is missing and what was paid twice: period, balances, open invoices, credit notes, payments on account.
See what this looks like on your own data
Twenty minutes, on a spend category of your choosing. We show you what the agents detect, with the evidence behind it.
- No commitment, on your own data
- Result in 3 weeks
- 20 minutes, no sales pitch
- Your data stays hosted in France
- No change of tool or process

