Quotes, unit price schedules, priced bills of quantities, invoices, framework contracts, price lists. Why has the mass of supplier documents become the first blind spot of procurement performance?

Ten years ago, a buyer managed a few dozen suppliers, an Excel workbook and contracts signed in duplicate. Today, that same buyer juggles hundreds of references, dozens of document formats, pricing terms with moving parts and regulatory obligations that keep piling up. Supplier data has exploded. Its complexity, however, has remained largely invisible, and that is precisely what makes it expensive.

When supplier data became unmanageable

The growth of supplier data is not new, but its acceleration is brutal. Three converging factors explain why most procurement functions have reached a breaking point, and why adding a person to the team no longer closes the gap.

The explosion in the number of suppliers and references

Globalised supply chains, the growing specialisation of service providers and the fragmentation of markets have mechanically multiplied the number of active suppliers. A mid-sized company now deals with several hundred distinct suppliers; a large group, with several thousand. Each produces its own documents (quotes, purchase orders, delivery notes, invoices, amendments) in its own format, with its own conventions, its own item codes, its own line descriptions.

The result is mechanical: where a procurement department once handled a few dozen documents a week, it now handles hundreds. And each one must be read, interpreted, compared with one or more references, then approved or disputed, by someone who also has a negotiation to prepare and a supplier review to run. The volume of data to process has grown far faster than the headcount assigned to it, and no recruitment plan closes that gap.

The multiplication of formats and document conventions

What makes this growth hard to absorb is not just the volume. It is the heterogeneity. Supplier data arrives in dozens of formats, often unstructured, rarely standardised.

A quote may come as an Excel table, a free-form PDF, a hand-formatted Word document or a simple summary email. A unit price schedule (the BPU of French public procurement) follows public-sector conventions, but every authority produces its own variant. A priced bill of quantities (the DPGF) can run to hundreds of lines with codes, units and terms specific to each contract. Price lists carry prices subject to variation rules indexed on official rates, which move independently of contractual commitments.

Faced with this heterogeneity, every manual or semi-automated process hits the same wall: you cannot directly compare what has not been put into the same format. Before you can even detect a discrepancy, you have to normalise data that will not normalise itself (same units, same references, same scope for each line), and that preliminary step alone consumes most of the time available.

Regulatory pressure as an accelerator of complexity

The third force comes from outside: the intensification of regulatory obligations. The CSRD (Corporate Sustainability Reporting Directive), the CS3D (Corporate Sustainability Due Diligence Directive) and the general tightening of traceability requirements oblige companies to document not only what they buy and at what price, but also from whom, under what conditions and with what consequences for their value chain. Every supplier becomes a source of data to qualify, monitor and document over time.

The invisible cost of unmanaged complexity

The complexity of supplier data has a cost. But that cost is rarely measured, because it is diffuse, spread across several functions and, above all, absent from the balance sheet. No budget line reads “time spent re-keying quotes” or “rebates that were never applied”. It nevertheless shows up in three concrete forms, and each of them can be observed on a real sample of documents.

The time cost: expert hours spent on mechanical tasks

The first consequence is the most obvious, but the least well measured. In most organisations, a large share of a buyer’s time is absorbed by document-processing tasks: compiling quotes, re-keying invoice data, trying to reconcile by hand documents that do not speak the same format. That is not time spent on negotiation, supplier intelligence or building partnerships. It is time spent on data entry and mechanical checking, paid at an expert’s rate.

On complex files it is even starker: an invoice tied to a construction contract, with a unit price schedule of several hundred lines, can occupy a buyer for half a day before approval or dispute. Multiply by the number of invoices per week, then by the number of buyers.

The error cost: overbilling that slips under the radar

The second consequence is financial, and often startling the first time it is measured. In any significant invoicing volume, a share of the invoices received contains anomalies against the negotiated terms. These anomalies are not necessarily fraudulent: they usually result from configuration errors, price updates not passed through, ancillary charges added by habit rather than by agreement, or diverging interpretations of a clause.

A rebate obtained after three months of negotiation disappears from the next invoice. A price per kilo slightly above the contractual rate has been applied for several deliveries without anyone noticing. On a single invoice, each discrepancy is marginal. Over a year, across every supplier, across a network of several dozen sites, these discrepancies represent points of margin lost in silence.

Across more than 5 million purchasing documents analysed, our first analyses systematically reveal invoicing errors on a significant share of the documents processed: discrepancies that, without a suitable tool, would never have been detected. This finding, observed in organisations of every size, is very probably the norm wherever controls are limited, precisely because undetected errors are not counted.

The decision cost: strategic choices made without reliable data

The third consequence is perhaps the most serious in the long run, because it is the least visible. When supplier data is scattered, unnormalised and uncompared, strategic purchasing decisions (renewing a contract, changing supplier, opening a renegotiation, consolidating volumes) are made on the basis of intuition and partial history.

A procurement director who wants to know whether the main supplier honoured its commitments over the last twelve months has to painstakingly reconstruct contracts, invoices and initial terms. With several hundred suppliers across several dozen sites, an exhaustive exercise is impossible. Decisions are made with incomplete data, or not made at all.

By processing its existing supplier data properly, an organisation can nevertheless obtain between 1 and 7% of margin recovered on the scope analysed, without changing supplier or renegotiating.

Anatomy of supplier data: what your documents really contain

To understand why managing supplier data is so complex, you have to look closely at what the documents circulating through a purchasing process actually contain. Each type of document has its own logic, its own critical data and its own traps, and each one is compared with a different reference, which is why no single check covers them all.

Quotes: the pricing promise in free format

The quote is the first contractual act of the supplier relationship. It contains the information that matters most for what follows: identification number, validity period, details of both parties, description of the services, quantities, unit prices, rebates, volume tiers, payment terms.

The problem is structural: no legal standard imposes a quote format. Every supplier uses its own template: a PDF generated from its ERP with proprietary item codes, a hand-formatted Excel table, a Word document, sometimes a photograph of a printed page. Comparing two quotes from different suppliers for the same service is like comparing two languages without a dictionary: the words differ, the units differ, and what one supplier includes in the price the other invoices separately.

Unit price schedules and priced bills of quantities: the contractual complexity of public procurement and construction

Unit price schedules and priced bills of quantities are specific to public contracts and to large construction or services operations. They are by nature extremely dense: a schedule often has several hundred lines of unit services, each with its code, description, unit of measure and price.

These documents are indispensable for invoice control on long-term projects. In practice, the manual comparison between a schedule of several hundred lines and a site invoice is time-consuming, error-prone, and often skipped for lack of time. The result: undetected discrepancies that accumulate, progress statement after progress statement.

Framework contracts and price lists: the underused reference data

Framework contracts and price lists are the most strategic documents of the supplier relationship, and yet the most underused day to day. A price list sets out the references whose price has been negotiated for a given period. It is the reference basis for assessing whether each invoice received honours the supplier’s commitments.

The problem is twofold. First, these documents are mostly stored in static formats (Excel or PDF) that were never designed to be automatically confronted with invoice flows. Second, their terms are complex: rebates tied to ordered volumes, variations indexed on official rates, special terms for certain sites or periods, amendments that override a clause without restating the whole schedule. A framework contract can carry more than 12,000 price lines, pricing rules included. This complexity makes manual checking laborious, and systematic checking practically impossible.

Invoices: the document of truth, hard to verify

The invoice crystallises every tension in supplier data management. It has to be compared with the quote, the purchase order, the framework contract, the price list, the technical specification where applicable. It must reflect the quantities actually delivered, the contractual prices, the negotiated rebates, the transport terms, the right VAT rate.

Yet the invoice arrives in the supplier’s format, not the format the buyer expects. It may use different descriptions from the quote for the same service, aggregate several deliveries, include unplanned charge lines, or apply prices slightly different from the contractual ones without the gap being visible to the naked eye on a multi-page document. The person approving it usually has the invoice on screen and the contract somewhere else, if at all.

ERP, Excel, OCR: why the tools in place are no longer enough

The ERP: the reference tool, but not a control tool

For decades the ERP has been the backbone of procurement, finance and accounting departments. It centralises orders, tracks budget commitments, manages supplier payments. On those missions it is irreplaceable. But faced with today’s supplier data, heterogeneous, unstructured and multi-format, it shows three structural limits.

First: an ERP only works on data it has been taught beforehand. Every new supplier requires several days of configuration: coding of references, pricing terms, validation rules. And the most underestimated cost is maintenance: as soon as a supplier changes its format or its terms, the configuration is obsolete. The ERP then keeps approving invoices against outdated references, and the discrepancies pile up silently until the audit.

Second: the ERP is built for data structured according to its own reference model. A document that does not follow that model (that is, almost every quote, price schedule and invoice received in free format) must first be re-keyed, reformatted and mapped to internal codes. That work falls on the teams, and every re-keying is an opportunity to introduce an inaccuracy: a miscoded reference, a rounded rebate, a simplified condition. Hence a well-known paradox: the ERP gives an impression of control. The data is entered, the flows are traced, but it does not detect that a rebate has vanished, that a price has moved or that an invoiced service does not match the one ordered. The ERP records; it does not check.

The third limit is organisational as much as technical. Every adaptation goes through IT or the integrator, and the delay between a need expressed by the business and its translation into the system is counted in weeks, sometimes months. Meanwhile, procurement, finance and accounting teams see the blind spots but have no hand on the tools, and invoicing discrepancies keep accumulating.

That friction is what AI agents are designed to remove. An agent specialised in purchasing data requires no specific development and changes nothing in the ERP: it reads what comes out of it (orders, receipts, supplier master data) and what arrives from suppliers, and returns its findings where the teams already work. Where the ERP records and traces, the agent checks, compares and alerts.

Excel: the universal fallback that quickly reaches its structural limits

Faced with the ERP’s gaps on unstructured documents, almost every procurement team has the same reflex: the spreadsheet. Excel is flexible, accessible, mastered by everyone, and it does the job on limited volumes. But as soon as document volume grows and pricing terms get more complex, its limits appear.

Excel compares numbers in cells, not contexts. It does not understand that a line reading “handling included” corresponds to “free delivery” on the quote, does not verify that a rebate was applied on every relevant line, and does not cross an invoice with the price list in force. Every check remains manual, which means most of them are not done.

Its second limit is organisational: every employee builds their own file with their own logic. When they leave, the file becomes unreadable. No collective memory, no traceability of decisions, no way for an auditor to see why a discrepancy was accepted. Excel is an individual tool in a process that involves several functions and several sites. And a formula error raises no alert: it propagates silently through every calculation that depends on it.

OCR: a useful first step, but an operational dead end on its own

OCR solutions digitise documents and extract text. That is useful, but insufficient. They stop at the surface: they do not know that the price per kilo just read must be compared with the price list negotiated six months earlier, do not detect that a different description refers to the same service, and do not understand that a VAT rate is wrong. The extracted data lands in a spreadsheet that an employee still has to analyse, with the same time constraints as before. OCR has moved the problem; it has not solved it.

Agentic AI: turning document complexity into prepared decisions

Agentic artificial intelligence is a break from previous approaches, not because it is more powerful, but because it addresses the right problem. Contrary to a common assumption, its value for procurement is not creativity. It is the ability to read heterogeneous documents at scale, without a predefined template, and to spot quickly what does not add up. Three capabilities, in sequence, make the difference.

Reading every format, extracting the data that matters

An agent built for purchasing data combines natural language processing and vision models to read a supplier document whatever its format: a price schedule of several hundred lines in PDF, a hand-formatted Excel quote, an invoice generated from a supplier’s ERP with proprietary codes. That is Capture’s role at Zylio: extracting, line by line, the nature of the service, quantities, unit prices, rebates, payment terms and due dates (100+ lines per document, with no manual entry) and keeping every value linked to the page and line where it was read.

Crossing several references at once

Once the data is extracted, the agent confronts it with every relevant reference: the initial quote, the framework contract, the price list in force at the invoice date, the technical specification for qualitative requirements, the history of previous invoices to detect gradual drift. That is Compliance’s job: every invoiced line is compared with the clause that governs it. The cross-check extends to the content of the services: does the quantity delivered match the quantity ordered? Does the lead time respect the specification? The agent detects discrepancies the human eye cannot see in the time it has.

Alerting, documenting, leaving the decision to you

An effective agent does not stop at detection. It documents and prepares the resolution. As soon as a discrepancy is identified, it builds the file: the precise contractual reference, the problem line, the amount at stake, the proof side by side. It informs the profit centres concerned and quantifies the impact of the discrepancy on margins. What it does not do is act in your place: Zylio contacts no supplier and raises no claim. The agent prepares the case; your team decides, with the file in hand. It is this complete chain (reading, extraction, cross-checking, detection, file) that turns document complexity into prepared decisions.

Conclusion: supplier data, the deposit you have not yet mined

Most organisations look for their next points of performance in new negotiations, new suppliers, new markets. They overlook a deposit that already exists, in plain sight, in the documents that circulate every week between their teams and their suppliers: rebates not applied, overbilling not detected, contractual terms ignored for lack of time to check them, strategic decisions made without compared data.

This deposit is all the more accessible in that it requires no renegotiation, no change of provider, no months-long transformation project. It requires one change: giving your teams a tool able to read what your documents contain, cross what they reveal, and flag what is wrong before payment rather than at audit, with the proof attached, so that the decision that follows is a short one.

That is the promise of agentic AI applied to supplier data: not replacing the expertise of your buyers, controllers and accountants, but giving them back the time they spend on tasks the machine does better and faster: 15 to 45% of time given back to teams, per FTE, depending on the scope and the maturity of the data. So that they can focus on what creates value: supplier strategy, substantive negotiation, building lasting partnerships.

Every unchecked document is a potential discrepancy that will never be detected, or a margin that will never be recovered. To find out what yours contain, the most direct route is a three-week diagnostic, on your data.