Skip to content

OCR

Document extraction API: integrate without a project

A document extraction API (an application programming interface) should slip into an existing flow, not open a project. API, SDK, webhook and a dedicated inbox are included in every plan, at no extra cost. Exports come out as XLSX, CSV or JSON. Your ERP stays the source of truth.

The problem

Integration often costs more than extraction

A technical team assessing a document extraction API is not looking for a demonstration. It wants to know how many weeks the plumbing will cost: where documents come in, what the processing returns, where the result lands. What is actually read is set out on the pillar page, with what extraction covers, document by document. This page deals with the rest: how the documents arrive, and in what form the data leaves.

The usual trap closes after signature, when the dedicated inbox, the webhook or the export turn out to be chargeable options, each opening its own piece of work. That is where the real cost of an extraction tool sits: in the number of people you have to mobilise before the first document travels end to end.

The rule applied here is explicit. Whatever lets you plug extraction into an existing flow is included in every plan, at no extra cost. Whatever requires work inside an ERP (enterprise resource planning system) or S2P (source-to-pay) environment belongs to the enterprise offer, and that is said up front, not at connection time.

Five ways for a document to come in

None excludes the others: the same organisation often receives PDFs by e-mail, scans dropped on a shared space and a regular feed from head office. All five channels feed the same processing.

Dedicated inbox
An address to give to suppliers, or to feed with a forwarding rule from the accounts mailbox. Nothing to develop on the mail side.
Upload
Manual submission, for scanned paper, documents found late and the one-off cases that fit no flow.
SFTP
The channel for regular feeds, when a subsidiary or a platform pushes batches of documents at fixed times.
Storage space synchronisation
The shared folder where documents already land: it is read where it is, without asking teams to change habits.
API
The programmatic call, for teams that want to control the exact moment a document goes for processing.

What comes back, and in what form

The output reads as well by machine as by spreadsheet. That is deliberate: the first integration is often an export, the second a webhook.

  • XLSX, CSV and JSON exports, depending on whether the data goes to a spreadsheet, a warehouse or a service.
  • Webhook: your system is told when a document has been processed, rather than polling in a loop.
  • SDK and API, included in every plan on the same footing as the intake channels.

What the returned data holds

This is the question a technical team asks, and rarely the one it gets answered. Three elements decide what you will be able to automate.

The requested fields, line by line
Reference, description, quantity, unit, unit price, discount and tax rate for each line, on top of the header and footer data.
A confidence score per field
Not one global percentage for the document: a certainty level on each value, because the field the payment depends on is the one that matters.
The position of the value in the document
Every extracted value stays tied to the exact place where it was read: page, zone, line. The data stays verifiable at source.

Those last two decide your automation threshold. A field below the threshold you set goes to human review; the rest pass through untouched. You move the slider, not us. Every extracted value stays linked to the exact place in the document where it was read.

Capture

Gather and consolidate all your procurement data.

Collection by dedicated mailbox, upload, SFTP, scan, API or ERP export. Reads every format (native PDF, scan, photo, spreadsheet, structured feed), then extracts line by line: supplier, references, quantities, unit prices, discounts, taxes, terms.

The end of re-entry. A procurement history you can finally query.

What Capture detects

  • Long documents and multi-page tables
  • Unknown layouts
  • Several languages and currencies
  • Poor scans and handwriting

What belongs to the enterprise offer

Connectors into an ERP or S2P environment, and the support that goes with them, are not included in the self-serve plan, taken out on the pricing page. They are scoped with the sales team, because they depend on your version, your reference data and your coding rules. Saying so beforehand avoids the unpleasant surprise at connection time.

Your ERP stays the source of truth

Zylio works on top of it and adds no development inside it. The system keeps carrying the process, the approvals and the postings; extraction sits upstream and returns its results in the formats your tools already read.

We fit around existing processes, we do not replace them. No development added inside the core of the ERP.

The systems in place and the connection modes are set out on the integrations page.

The plan covers extraction. Checking your negotiated terms is the next step and is scoped with you.

Check the plans and the included volume

Nothing is lost. Everything can be checked, everything can be proven.

  • The contract clause and the invoice line, highlighted side by side.
  • Every extracted value stays linked to the exact place in the document where it was read.
  • The same case produces the same decision, today as in six months: the rules are applied deterministically.
  • No discrepancy is set aside in silence. Anything that matches no rule is raised, with its reason.
  • The agent records what it did, in the order it did it: who, what, how much, when.

Frequently asked questions

Is the API charged as an extra?

No. The API, the SDK, the webhook and the dedicated inbox are included in every plan. Billing is based on the volume of pages analysed, not on the channel the documents come through. Only connectors into an ERP or S2P environment, and the support around them, belong to the enterprise offer.

What does the data returned by the API contain?

The fields you asked for, extracted line by line, each with its confidence level and its position in the document. Exports come out as XLSX, CSV or JSON, and a webhook tells your system as soon as a document has been processed.

Does Zylio replace my ERP?

No. Your ERP runs the process and remains the source of truth. Zylio handles the exception, on top of it, and feeds the results back. No additional development inside your system.

Which tools does Zylio connect to?

To the ERPs and management tools already in place, including SAP, Sage, Oracle and Pennylane, as well as to existing document repositories, mailboxes and feeds.

Measurable impact in every environment

More than 5 million procurement documents analysed

Between 1 and 7% of margin recovered

on the scope analysed

From 15 to 45% of time given back to teams, per FTE

depending on the scope and on data maturity

Zylio fits into your existing ecosystem.

The ERP runs the process. Zylio handles the exception and recovers the value that escapes it: invoices without a purchase order, line-by-line price discrepancies, duplicates and overbilling, off-contract spend.

  • SAP
  • Sage
  • Oracle
  • NetSuite
  • Microsoft Dynamics 365
  • Pennylane
All integrations

Your data under high security.

Zylio meets the most demanding standards, and nothing is committed without your approval.

Certifications
SOC 2 Type II · ISO 27001
Hosting
Hosted in France
Encryption
End-to-end AES-256 encryption
Access
Enterprise SSO · multi-factor authentication · Zero Trust approach
Security and compliance

See what this looks like on your own data

Twenty minutes, on a spend category of your choosing. We show you what the agents detect, with the evidence behind it.

  • No commitment, on your own data
  • Result in 3 weeks
  • 20 minutes, no sales pitch
  • Your data stays hosted in France
  • No change of tool or process