OCR
Document extraction API: integrate without a project
A document extraction API (an application programming interface) should slip into an existing flow, not open a project. API, SDK, webhook and a dedicated inbox are included in every plan, at no extra cost. Exports come out as XLSX, CSV or JSON. Your ERP stays the source of truth.
The problem
Integration often costs more than extraction
A technical team assessing a document extraction API is not looking for a demonstration. It wants to know how many weeks the plumbing will cost: where documents come in, what the processing returns, where the result lands. What is actually read is set out on the pillar page, with what extraction covers, document by document. This page deals with the rest: how the documents arrive, and in what form the data leaves.
The usual trap closes after signature, when the dedicated inbox, the webhook or the export turn out to be chargeable options, each opening its own piece of work. That is where the real cost of an extraction tool sits: in the number of people you have to mobilise before the first document travels end to end.
The rule applied here is explicit. Whatever lets you plug extraction into an existing flow is included in every plan, at no extra cost. Whatever requires work inside an ERP (enterprise resource planning system) or S2P (source-to-pay) environment belongs to the enterprise offer, and that is said up front, not at connection time.
Five ways for a document to come in
None excludes the others: the same organisation often receives PDFs by e-mail, scans dropped on a shared space and a regular feed from head office. All five channels feed the same processing.
- Dedicated inbox
- An address to give to suppliers, or to feed with a forwarding rule from the accounts mailbox. Nothing to develop on the mail side.
- Upload
- Manual submission, for scanned paper, documents found late and the one-off cases that fit no flow.
- SFTP
- The channel for regular feeds, when a subsidiary or a platform pushes batches of documents at fixed times.
- Storage space synchronisation
- The shared folder where documents already land: it is read where it is, without asking teams to change habits.
- API
- The programmatic call, for teams that want to control the exact moment a document goes for processing.
What comes back, and in what form
The output reads as well by machine as by spreadsheet. That is deliberate: the first integration is often an export, the second a webhook.
- XLSX, CSV and JSON exports, depending on whether the data goes to a spreadsheet, a warehouse or a service.
- Webhook: your system is told when a document has been processed, rather than polling in a loop.
- SDK and API, included in every plan on the same footing as the intake channels.
What the returned data holds
This is the question a technical team asks, and rarely the one it gets answered. Three elements decide what you will be able to automate.
- The requested fields, line by line
- Reference, description, quantity, unit, unit price, discount and tax rate for each line, on top of the header and footer data.
- A confidence score per field
- Not one global percentage for the document: a certainty level on each value, because the field the payment depends on is the one that matters.
- The position of the value in the document
- Every extracted value stays tied to the exact place where it was read: page, zone, line. The data stays verifiable at source.
Those last two decide your automation threshold. A field below the threshold you set goes to human review; the rest pass through untouched. You move the slider, not us. Every extracted value stays linked to the exact place in the document where it was read.
Capture
Gather and consolidate all your procurement data.
Collection by dedicated mailbox, upload, SFTP, scan, API or ERP export. Reads every format (native PDF, scan, photo, spreadsheet, structured feed), then extracts line by line: supplier, references, quantities, unit prices, discounts, taxes, terms.
The end of re-entry. A procurement history you can finally query.
What Capture detects
- Long documents and multi-page tables
- Unknown layouts
- Several languages and currencies
- Poor scans and handwriting
What belongs to the enterprise offer
Connectors into an ERP or S2P environment, and the support that goes with them, are not included in the self-serve plan, taken out on the pricing page. They are scoped with the sales team, because they depend on your version, your reference data and your coding rules. Saying so beforehand avoids the unpleasant surprise at connection time.
Your ERP stays the source of truth
Zylio works on top of it and adds no development inside it. The system keeps carrying the process, the approvals and the postings; extraction sits upstream and returns its results in the formats your tools already read.
We fit around existing processes, we do not replace them. No development added inside the core of the ERP.
The systems in place and the connection modes are set out on the integrations page.
The plan covers extraction. Checking your negotiated terms is the next step and is scoped with you.
Check the plans and the included volumeNothing is lost. Everything can be checked, everything can be proven.
- The contract clause and the invoice line, highlighted side by side.
- Every extracted value stays linked to the exact place in the document where it was read.
- The same case produces the same decision, today as in six months: the rules are applied deterministically.
- No discrepancy is set aside in silence. Anything that matches no rule is raised, with its reason.
- The agent records what it did, in the order it did it: who, what, how much, when.
Frequently asked questions
Is the API charged as an extra?
No. The API, the SDK, the webhook and the dedicated inbox are included in every plan. Billing is based on the volume of pages analysed, not on the channel the documents come through. Only connectors into an ERP or S2P environment, and the support around them, belong to the enterprise offer.
What does the data returned by the API contain?
The fields you asked for, extracted line by line, each with its confidence level and its position in the document. Exports come out as XLSX, CSV or JSON, and a webhook tells your system as soon as a document has been processed.
Does Zylio replace my ERP?
No. Your ERP runs the process and remains the source of truth. Zylio handles the exception, on top of it, and feeds the results back. No additional development inside your system.
Which tools does Zylio connect to?
To the ERPs and management tools already in place, including SAP, Sage, Oracle and Pennylane, as well as to existing document repositories, mailboxes and feeds.
Measurable impact in every environment
More than 5 million procurement documents analysed
Between 1 and 7% of margin recovered
on the scope analysed
From 15 to 45% of time given back to teams, per FTE
depending on the scope and on data maturity
Zylio fits into your existing ecosystem.
The ERP runs the process. Zylio handles the exception and recovers the value that escapes it: invoices without a purchase order, line-by-line price discrepancies, duplicates and overbilling, off-contract spend.
Your data under high security.
Zylio meets the most demanding standards, and nothing is committed without your approval.
- Certifications
- Hosting
- Encryption
- Access
Read next
- The OCR pillar: what gets extractedWhat Zylio reads on each type of document, and what comes after that.
- Automated invoice intakeE-mail, paper, portal, EDI, platform: five channels, one flow.
- Invoice OCR accuracyThe score per field and the review threshold, seen from the returned data.
- Line-by-line extractionEvery item, with its reference, quantity, unit price and discount, made comparable with the contract.
- Multilingual and multi-currencyDate, thousands separator, currency symbol, tax mention: everything is normalised on output, the original value kept beside it.
See what this looks like on your own data
Twenty minutes, on a spend category of your choosing. We show you what the agents detect, with the evidence behind it.
- No commitment, on your own data
- Result in 3 weeks
- 20 minutes, no sales pitch
- Your data stays hosted in France
- No change of tool or process

