Data entry holds up month end
Supplier invoices, receipts and delivery notes are downloaded from email, opened one by one and keyed into the accounting system. Closing week goes on this, leaving little time for review and analysis.
AI solution
AI document processing is a workflow that reads the information in invoices, delivery notes, contracts, application forms and email attachments, splits it into fields and writes it to the right system. We do not set it up as a standalone OCR tool but as a process with reading, validation, human review and an audit trail.
In short
AI document processing turns scanned, photographed or PDF documents into records in your accounting, ERP or CRM system by extracting dates, amounts, suppliers, line items and other fields. A well built workflow scores its confidence for each field, cross checks values with rules, sends documents it is unsure about to a person and keeps the original file untouched. Documents that already arrive as structured e-invoices are read directly, without AI.
Talha Aslan and teamLast updated:
When you need it
For a team that handles a few documents a month, an AI workflow costs more effort than it saves. If most of your documents already arrive as structured e-invoices, reading that data directly comes first. If the situations below sound familiar, document processing automation is worth a look.
Supplier invoices, receipts and delivery notes are downloaded from email, opened one by one and keyed into the accounting system. Closing week goes on this, leaving little time for review and analysis.
The amount on a quote is copied into the order, the address on the order into the shipping papers. When a figure slips, the first to notice is often the customer or the supplier.
Tools that read fixed coordinates only work on layouts they already know. A new supplier, a different invoice layout or a skewed phone photo mixes up the fields, and corrections are made by hand again.
Contracts, applications and correspondence are spread across shared folders and inboxes. When someone asks about a renewal date or a specific clause, every file is read again from the start.
Our approach
We start with your document inventory, not with software. From real samples of recent months we list which document types arrive through which channel, which fields each one needs and which system those fields belong in. The same samples become the test set we use to measure accuracy.
The workflow first identifies the document type, then extracts fields with a model that reads image and text together, or with a reading service such as Google Document AI or Azure AI Document Intelligence. The values are then tested against rules: do the totals add up, is the VAT number in a valid format, has this invoice been processed before. Documents with low confidence or a failed rule land on a review screen; clean ones go straight into your accounting, ERP or CRM. Setup and upkeep run as part of our AI automation services.
Documents often arrive through a form or portal where customers upload files. When your customers need a place to upload documents, we build on a membership website; when your team needs its own internal dashboard, that work runs through custom software development.
Every step leaves a record: which model version read the document, who approved it and which entry it became can all be traced later.
Which workflow?
Reading invoices and reviewing contracts are different jobs, so we first pick the document type that costs your team the most time.
Accounting and purchasing
Extracts fields from incoming invoices and expense documents, matches them to orders and prepares the accounting entry.
Operations and onboarding
Sorts the applications, declarations and attachments customers upload, spots what is missing and gets the file completed.
Legal and management
Pulls parties, term, renewal and notice dates from contracts and adds them to your calendar and contract register.
Essentials
These points decide where the workflow catches mistakes and how it protects the personal data inside your documents.
Many invoices already arrive as structured data in formats such as XRechnung, ZUGFeRD or other EN 16931 e-invoices. Germany's Federal Ministry of Finance, for example, states that domestic businesses must be able to receive e-invoices since 1 January 2025 and that a plain PDF does not count as one. We read such files directly and keep AI for paper, scans and PDFs.
The model returns a value for every field but cannot prove that it is right. Rules such as matching totals, date and tax ID formats and duplicate invoice checks run on every document; anything that misses the threshold is not posted without approval.
Article 9 of the GDPR treats health, genetic and biometric data as special categories, and Article 10 covers criminal records. Documents that contain them run in a separate workflow with narrower access, and where needed the model runs on your own server.
Article 22 of the GDPR gives people the right not to be subject to decisions based solely on automated processing that significantly affect them, and the EU AI Act lists AI used for recruitment or creditworthiness as high risk in Annex III. Rejecting an application or stopping a payment therefore always stays with a person.
If the reading service sits outside the EU or EEA, document content leaves it. Chapter V of the GDPR then requires an adequacy decision or safeguards such as standard contractual clauses, and the provider signs a processing agreement under Article 28. Your legal adviser makes the final call.
The source document is never overwritten; extracted data is stored separately with the model version and the person who approved it. If an entry is questioned, you can show which document it came from and how, and roll it back.
Sources: Federal Ministry of Finance (Germany): FAQ on e-invoicing · General Data Protection Regulation (2016/679), Articles 9, 10, 22 and 28, EUR-Lex · EU AI Act (Regulation 2024/1689), Annex III, EUR-Lex
Comparison
| Topic | Template OCR | AI assisted workflow |
|---|---|---|
| New layouts | A new template for each layout | Introduced with samples, no template |
| What is read | The text on the page | Text, tables and what each field means |
| Catching errors | Misreads go into the record | Confidence scores and rule checks |
| Unclear documents | Fields silently left blank | Sent to the review queue |
| Cost per document | Low and fixed | Varies with model usage |
| Setup | Quick, for one layout | Inventory, test set and pilot |
Quick check
Must haves: is your process ready?
0 of 6 in place Tick the boxes to see how ready you are for automation.
Added as needed
We choose which of these you need together during the first call.
Share ten to twenty anonymised samples of one document type and the system the data should go to; we will define the fields, the review rules and a written quote.
Process
We listen to your processes in a free 15-minute call. Then discovery maps your tools and tasks, scores the opportunities and ends with a written scope and fee for your approval.
We build the first workflow in your accounts and test it with real but masked examples. Approval steps, error scenarios and alerts go in before anything reaches a customer.
We switch the workflow on step by step, watch the logs and adjust thresholds with your team. You get documentation and a short training session.
On the monthly plan, we monitor running workflows, adapt them to model and API changes and add new workflows from the priority list, with a monthly report.
Data, security and measurement
Documents in the test set are compared with values entered by hand. We measure per field rather than per document, because a single wrong amount spoils the whole entry.
We track the share of documents posted without review. When that share rises, the threshold is not loosened until sample entries have been checked by a person.
The time from a document's arrival to its entry in your system is compared with manual logs kept in the weeks before automation.
Access to documents and extracted data is limited by role. Retention periods are set in writing with your accountant and legal adviser according to your legal obligations.
Free tools
Check VAT on invoice amounts, convert foreign currency invoices, work out due dates, create file hashes to catch duplicates, measure the hours spent on manual entry and generate strong passwords for new accounts.
Finance
Add VAT and extract it with the correct formula; preset + custom rates.
Currency
Convert euros and 30 currencies with official ECB reference rates, look up any past date since 1999 and see monthly averages.
Calculator
Days, weeks and working days between dates; add/subtract from a date.
Security
Create MD5, SHA-1, SHA-256 and SHA-512 hashes of text and files in your browser, verify a download against its checksum and compare hashes. Nothing is uploaded.
Work
Calculate daily and weekly working hours after breaks, in hours and decimals, and check legal breaks and rest periods for the UK and EU.
Security
Cryptographically random strong passwords + strength meter + crack time.
How we work
We do not yet have a live client project in AI document processing that we can show as a reference, so instead of promising results we describe our method. You can see our automation, software and web projects on the references page.
Before any build, we put together an anonymised sample set from your recent documents; accuracy is measured on that set.
In the first weeks the workflow runs alongside your team without posting anything and only shows its results. It goes live once the list of differences is clear.
If the model or reading service changes, rules, review screen and integrations stay in place, and the new layer is measured again on the same test set.
Model, reading service and automation accounts are opened in your company's name; field definitions, rules and documentation are handed over to you.
FAQ
If your question is not here, write to us; we will send you an answer and a written quote.
Next step
Tell us which documents you handle, the rough monthly volume and the system the data should go to; after a free 15 minute call we will send the scope and a written quote.
In-depth guide
Most document automation projects succeed or fail on small decisions made long before anyone picks a model: which document type goes first, which fields are mandatory, which rule stops an entry and who looks at the review screen. This guide walks through those decisions in the order a business owner faces them when planning AI document processing.
It is a working checklist, not a brochure: criteria to put to any vendor, preparation you can do with your own files, and the cases where automation is not worth it. Technical terms are explained the first time they appear.
Whether automation suits you becomes clear only when you count the documents that actually arrived over the last three months; gut feeling tends to produce needless or half finished projects. Go through shared folders, inboxes and scanner output for a week and fill in a simple sheet.
The sheet argues against automation in three cases: volume is low, most documents already arrive as structured electronic invoices, or each document type shows up only a few times a year. Then the import feature of your accounting software plus a tidy folder structure gives the same relief for less effort. If one document type with shifting layouts holds up every close, you have found your first workflow.
Start with the document type that eats the most staff time and whose mistakes cost the most; automating everything at once makes measurement impossible. Typical candidates:
An accounting firm that collects paperwork from clients should get documents arriving through one channel before adding AI document processing; a secure upload area on a website for accounting firms is a sensible place to start. The first type you choose becomes the template: the same review screen, audit trail and measurement sheet are reused for the second.
A robust document workflow consists of separate layers that can each be tested on their own. Ask any vendor to walk you through them; "the AI reads it and posts it" describes a system where errors cannot be traced.
The payoff is replaceability: when the reading service changes, only extraction is measured again.
Sending structured documents through an AI model adds cost and risk for no gain. Formats such as Peppol UBL, XRechnung, ZUGFeRD and other invoices following the European standard EN 16931 already carry amounts, tax and line items as separate fields that your accounting software can import directly.
Confusion starts when a supplier emails a PDF rendering while the structured file sits elsewhere. Germany's Federal Ministry of Finance, for instance, states that domestic businesses must be able to receive electronic invoices since 1 January 2025 and that a plain PDF does not count as one. A workflow should handle every incoming PDF in this order:
This split does two things at once. The model reads fewer documents, so usage fees drop, and the same invoice can no longer be posted twice, once from XML and once from PDF.
A field schema is the written list of what gets extracted from each document type and in which format; it is the contract of the workflow. Without one, output drifts and the target system quietly accepts it. For every field, write down the name, data type, whether it is mandatory and its standard format: dates in one format, one decimal separator, currency as the three letter international code.
Next, list the cross check rules for each type. A typical starting set for invoices:
To catch the same file uploaded twice under another name, compare content hashes, fixed length fingerprints computed from each file; try it with our hash generator.
Entry and exit points matter as much as the reading layer; if the target system cannot accept entries, even an accurate model is of little use. On the intake side, write down for each channel who may send documents, which file types are accepted and the size limit.
On the output side, the first question is how the target system takes in records. Accounting packages, ERP and CRM software are usually fed in one of three ways:
Whichever route you choose, two rules hold. First, resending an entry must never create a second one; a retry after a network outage should not produce a duplicate bill. Second, the workflow may suggest a general ledger account but should not finalize it in the first months; the suggestion appears on the review screen and the accountant decides. If the target system has no connection point, a small middleware service is planned separately as custom software development.
There are three main options for the reading layer, and your own test set should decide between them, not a vendor's landing page. Prebuilt document services offer a fast start on common types such as invoices and receipts; general purpose models that read image and text together cope better with changing layouts and free text; open models running on your own server keep data in house.
Compare them on the same test set using these criteria:
For companies handling medical reports or identity documents, a local LLM setup is a serious option, though hardware, updates and security become your responsibility. Many businesses settle on a hybrid: sensitive types local, everything else on a reviewed cloud service.
The review screen is where people spend most of their time in AI document processing, and a poorly designed one gives back the hours that automation saved. A good screen shows the document image and the extracted fields side by side, highlights the region where each value was read and colors only the doubtful fields. The reviewer checks the flagged fields, not the whole page.
Set thresholds per field rather than as one number. A small misreading of a supplier name usually corrects itself when matched against master data, while one wrong digit in the gross total is money lost. So amounts and tax IDs get strict thresholds and description fields looser ones. Above a certain invoice amount you can require a second approver regardless of the score.
The audit trail for each document should hold at least:
Review corrections weekly and add recurring ones to the test set.
Documents often carry personal data: a sole trader's name on an invoice, an ID number on an application, a diagnosis on a medical certificate. Before building, draw a data flow map: which personal data each type contains, where it travels and how long it is kept.
The EU AI Act lists AI used in recruitment and creditworthiness assessment as high risk in Annex III, so a workflow that screens CVs or scores loan files carries a very different set of obligations from simple invoice capture. Your legal adviser makes the assessment; the technical team documents the data flow and settings.
A workflow goes live in stages, each based on the measurements of the one before. This skeleton suits a single document type in most businesses.
Shadow mode is the stage most often cut short, yet it is the only place to see behavior on real documents. Keep it running through at least one month end close; hastily scanned paper in closing week reveals errors quiet days never show.
Value shows up in several measures that balance each other, not in a single rate; if the straight through share rises while corrections rise too, the workflow got faster and worse. Record the starting point before anything changes: tracking time per document for a few weeks with our working hours calculator is the only solid basis for a later comparison.
Read the numbers with seasonality in mind; year end or holiday peaks can make AI document processing look better or worse than it is.
AI document processing reduces reading errors but does not remove them, and some risks did not exist with classic OCR at all. Write down a countermeasure for each limit before the build.
These measures make risk manageable, not zero, so the review queue is never switched off; it only gets narrower.
Most problems come from preparation and rollout order, not from the model.
One more turns up often: involving the accounting team late. People who will use the review screen should see the field schema and rules from day one.
A good partner asks about your samples, target system and closing calendar before showing a demo. When requesting quotes for AI document processing, ask for written answers on these points:
We deliver this as part of our AI automation services. We have no live client project in AI document processing to show as a reference yet, so we describe our method rather than promise results; our automation, software and web work is on the references page.
Cost depends on the number of document types, volume, connected systems and the scope of upkeep; fixed price options for discovery and the first workflow are on our pricing page. Send the most time consuming document type, rough monthly volume and target system via the contact form, and we will define the field schema, review rules and a written quote with you.
Start a Project
Thanks {name}, we've received your brief. We usually reply within the same day.
What happens next?