Data extraction
PDF extraction of purchase data for analysis: consistent, sourced lines instead of files
PDF extraction of purchase data means reading quotes, orders, invoices and contracts received as PDFs, pulling out every line (reference, quantity, unit price, discount, tax) and placing them in a single data model. Zylio does this with the Capture agent, keeps the link to the original page and delivers data ready for its dashboards or your BI tool.
The problem
A PDF is not data
Most purchase data arrives as PDF: the supplier’s quote, its invoice, the signed framework contract, the rate card in the appendix. A PDF can be read on screen, printed and filed. It cannot be filtered, summed or compared. For analysis, it is barely better than a sheet of paper.
The usual answer is retyping: someone copies the totals into a spreadsheet, sometimes the main lines, rarely everything. Time and accuracy are lost, and above all substance: percentage discounts, delivery charges and packaging units do not survive the copy. The analysis that follows covers a simplified spend, not the real one.
Even when the ERP holds the order, it does not hold what the supplier wrote. The gap between the two, invoiced price against negotiated price, delivered quantity against ordered quantity, is exactly what you would want to analyse. It exists nowhere until the PDF has been extracted line by line.
Scope
What Zylio extracts from each document
Not just the header and the total: every line, with the fields that make analysis possible afterwards.
Entities
Supplier, invoiced entity, delivery site, governing contract. The supplier is recognised even under a different legal name or address.
References and descriptions
Supplier item code, internal code where present, full description. Two different descriptions for the same item are matched, not duplicated.
Quantities and units
Ordered, delivered and invoiced quantity; sales unit and packaging unit. A box of twelve and a single piece cannot be compared without conversion.
Prices and discounts
Gross unit price, discount as a percentage or an amount, net price, volume tier applied. The net price is calculated and checked, not merely copied.
Taxes and charges
Tax rate per line, delivery charges, free-freight threshold, eco-contributions, handling fees. What is usually buried in the total at the bottom of the page.
Dates and terms
Issue, delivery and due dates; payment terms, early-payment discount, currency. Everything that weighs on the real cost without appearing in the price.
Mechanism
From PDF to table: the steps
- 01
Intake
PDFs arrive through a dedicated mailbox, upload, SFTP or API. A scan, a photo or a spreadsheet is accepted just like a native PDF.
- 02
Reading
The Capture agent identifies the document type and its structure, with no template to configure. A table running over several pages is read as one.
- 03
Line-by-line extraction
Each line becomes a record with its fields. A confidence score accompanies every value; whatever could not be read is declared, never estimated.
- 04
Normalisation
Fields are brought into a single model: same units, same currencies, same categories, whatever the supplier or the original format.
- 05
Validation
Values below the confidence threshold are shown to a person, with the area of the PDF alongside. They confirm or correct in one move.
- 06
Delivery
The data feeds the Zylio repository and its dashboards, or flows out to your data warehouse and your BI tool.
Multi-page tables, irregular columns, descriptions spanning two lines.
Analysis
What you can finally analyse
Once the lines are extracted and consistent, the questions that went unanswered get an answer. How much did we spend per supplier, per site, per category, per period? What share of that spend runs under contract? Has the price paid for a reference moved, and since when? Do the accepted quote and the received invoice say the same thing?
These answers no longer require a data project. They are asked in plain language to Insight, the procurement intelligence layer, which builds the figure, the table or the curve, and shows for every value the lines that make it up, and for every line, the PDF it came from. A figure challenged in a meeting can be reopened down to its original page.
If your analyses already live in a BI tool, the extracted data goes there in a stable model: same columns, same supplier and item keys, same units. What your reports were missing was not the tool; it was line-level data.
Data model, refresh, join keys.
Outcome
What changes for your analyses
The spend you analyse is the real spend. An invoice of more than a hundred lines comes in whole, with its discounts and charges, not as a copied total. A category analysis no longer rests on the rough allocation typed at order time, but on what the supplier actually delivered and invoiced.
Time changes sides. The hours spent copying PDFs into a spreadsheet before each review go back to the teams: between 15% and 45% of time handed back per FTE, depending on scope and data maturity. That time goes to preparing a negotiation with the real history of invoiced prices, or to documenting a gap found in the analysis.
The scope widens. When extraction costs nothing, there is no reason to analyse only the large suppliers. Small accounts, site-level purchases and recurring services enter the analysis, and that is often where gaps pile up while nobody is looking.
Your purchasing data already exists. It is just locked inside PDFs.
Multi-site
A particular lever for groups and networks
When several sites or subsidiaries buy separately, PDFs are even more scattered and consolidated analysis even rarer. Extraction into a single model reveals what nobody could see: the same supplier serving three sites under three legal names, the same reference invoiced at different prices depending on the entity, a framework contract negotiated at headquarters and ignored in the regions.
These are facts, with the invoice and the page as proof. They change a negotiation: you no longer arrive with a hunch about volume, but with consolidated spend and observed gaps. They also change the organisation: purchasing processes are easier to streamline when every site reports in the same format.
Same reference, same supplier, different prices depending on the site.
Capture
Gather and consolidate all your procurement data.
Collection by dedicated mailbox, upload, SFTP, scan, API or ERP export. Reads every format (native PDF, scan, photo, spreadsheet, structured feed), then extracts line by line: supplier, references, quantities, unit prices, discounts, taxes, terms.
The end of re-entry. A procurement history you can finally query.
What Capture detects
- Long documents and multi-page tables
- Unknown layouts
- Several languages and currencies
- Poor scans and handwriting
Nothing is lost. Everything can be checked, everything can be proven.
- The contract clause and the invoice line, highlighted side by side.
- Every extracted value stays linked to the exact place in the document where it was read.
- The same case produces the same decision, today as in six months: the rules are applied deterministically.
- No discrepancy is set aside in silence. Anything that matches no rule is raised, with its reason.
- The agent records what it did, in the order it did it: who, what, how much, when.
Frequently asked questions
Do I need to set up a template per supplier or per PDF type?
No. The Capture agent recognises the structure of each document as it reads it, including a layout it has never seen. A new supplier requires no configuration. Whatever cannot be read with confidence is flagged for validation, with the area of the PDF alongside.
What happens to the extracted data?
It feeds the Zylio procurement repository and its dashboards, and can flow out to your data warehouse or BI tool in a stable model. Every value keeps its link to the original document, page and line. No data is used to train a model.
Does extraction change anything in the ERP?
No. Zylio reads the ERP’s exports and APIs to match extracted lines with your orders; it writes nothing into the system and adds no development to its core. Extracted data lives alongside the ERP, not inside it.
How many documents can be extracted?
The platform absorbs over 10,000 supplier documents per week, across all formats, languages and structures, including documents of more than a hundred lines. Volume is not the constraint; the scope you choose to open is, and it widens at your own pace.
Measurable impact in every environment
More than 5 million procurement documents analysed
Between 1 and 7% of margin recovered
on the scope analysed
From 15 to 45% of time given back to teams, per FTE
depending on the scope and on data maturity
Zylio fits into your existing ecosystem.
The ERP runs the process. Zylio handles the exception and recovers the value that escapes it: invoices without a purchase order, line-by-line price discrepancies, duplicates and overbilling, off-contract spend.
Your data under high security.
Zylio meets the most demanding standards, and nothing is committed without your approval.
- Certifications
- Hosting
- Encryption
- Access
Read next
- Automatic integration of supplier quotationsThe quote received by email or portal, read, checked and returned to your ERP as structured lines, without re-keying.
- AI standardizationQuotes, price schedules, catalogues and rate cards brought to one usable structure, with proof on each line.
- Automatic standardization of quotationsEvery quote received, whatever its format, brought to a single, verifiable structure.
- Standardizing contractsBring every supplier contract into one common structure, without rewriting a single word, so you can compare and query them at will.
- The 3-week diagnosticYour documents extracted and analysed, with observed gaps as proof.
See what this looks like on your own data
Twenty minutes, on a spend category of your choosing. We show you what the agents detect, with the evidence behind it.
- No commitment, on your own data
- Result in 3 weeks
- 20 minutes, no sales pitch
- Your data stays hosted in France
- No change of tool or process

