AI solution

AI Document Processing

AI document processing is a workflow that reads the information in invoices, delivery notes, contracts, application forms and email attachments, splits it into fields and writes it to the right system. We do not set it up as a standalone OCR tool but as a process with reading, validation, human review and an audit trail.

Invoice and delivery note captureDocument classificationField level confidence scoresHuman review queueERP and accounting integration
  • Google Partner
  • Talha Aslan and team
  • English, German, Turkish

In short

AI document processing turns scanned, photographed or PDF documents into records in your accounting, ERP or CRM system by extracting dates, amounts, suppliers, line items and other fields. A well built workflow scores its confidence for each field, cross checks values with rules, sends documents it is unsure about to a person and keeps the original file untouched. Documents that already arrive as structured e-invoices are read directly, without AI.

Talha Aslan and teamLast updated:

When you need it

When does paperwork call for automation?

For a team that handles a few documents a month, an AI workflow costs more effort than it saves. If most of your documents already arrive as structured e-invoices, reading that data directly comes first. If the situations below sound familiar, document processing automation is worth a look.

Data entry holds up month end

Supplier invoices, receipts and delivery notes are downloaded from email, opened one by one and keyed into the accounting system. Closing week goes on this, leaving little time for review and analysis.

The same data is typed twice and goes wrong twice

The amount on a quote is copied into the order, the address on the order into the shipping papers. When a figure slips, the first to notice is often the customer or the supplier.

Template OCR breaks when the layout changes

Tools that read fixed coordinates only work on layouts they already know. A new supplier, a different invoice layout or a skewed phone photo mixes up the fields, and corrections are made by hand again.

Documents cannot be found when needed

Contracts, applications and correspondence are spread across shared folders and inboxes. When someone asks about a renewal date or a specific clause, every file is read again from the start.

Our approach

A document workflow that reads, double checks and asks when unsure

We start with your document inventory, not with software. From real samples of recent months we list which document types arrive through which channel, which fields each one needs and which system those fields belong in. The same samples become the test set we use to measure accuracy.

The workflow first identifies the document type, then extracts fields with a model that reads image and text together, or with a reading service such as Google Document AI or Azure AI Document Intelligence. The values are then tested against rules: do the totals add up, is the VAT number in a valid format, has this invoice been processed before. Documents with low confidence or a failed rule land on a review screen; clean ones go straight into your accounting, ERP or CRM. Setup and upkeep run as part of our AI automation services.

Documents often arrive through a form or portal where customers upload files. When your customers need a place to upload documents, we build on a membership website; when your team needs its own internal dashboard, that work runs through custom software development.

  • Document type is classified before it is read
  • Confidence score and rule checks for every field
  • Documents below the threshold go to a review queue
  • Structured data such as e-invoice XML is taken as is
  • Originals kept unchanged, every step logged
Anatomy of a document processing workflow
  1. Document intakeEmail attachment, scan, photo or upload form
  2. ClassificationInvoice, delivery note, contract, application
  3. Field extractionDate, amount, supplier, line items, with confidence
  4. Rule checksTotals, tax IDs, duplicate detection
  5. Review queueDocuments below the threshold go to a person
  6. Posting to your systemsAccounting, ERP or CRM entry with audit trail

Every step leaves a record: which model version read the document, who approved it and which entry it became can all be traced later.

Which workflow?

Document type and destination shape the setup

Reading invoices and reviewing contracts are different jobs, so we first pick the document type that costs your team the most time.

Accounting and purchasing

Invoice, receipt and delivery note capture

Extracts fields from incoming invoices and expense documents, matches them to orders and prepares the accounting entry.

  • E-invoices from XML, PDFs and photos via the model
  • Three way match with order and delivery note
  • Review queue whenever amounts differ

Operations and onboarding

Applications and supporting documents

Sorts the applications, declarations and attachments customers upload, spots what is missing and gets the file completed.

  • Checklist of missing items per document type
  • Draft request for missing documents
  • Tighter access for special category data

Legal and management

Contract and correspondence summaries

Pulls parties, term, renewal and notice dates from contracts and adds them to your calendar and contract register.

  • Clause level summary with its source
  • Reminder ahead of renewal dates
  • Extraction only; assessment stays with your lawyer

Essentials

The building blocks of a reliable document workflow

These points decide where the workflow catches mistakes and how it protects the personal data inside your documents.

Structured data first

Many invoices already arrive as structured data in formats such as XRechnung, ZUGFeRD or other EN 16931 e-invoices. Germany's Federal Ministry of Finance, for example, states that domestic businesses must be able to receive e-invoices since 1 January 2025 and that a plain PDF does not count as one. We read such files directly and keep AI for paper, scans and PDFs.

Confidence thresholds and rule checks

The model returns a value for every field but cannot prove that it is right. Rules such as matching totals, date and tax ID formats and duplicate invoice checks run on every document; anything that misses the threshold is not posted without approval.

A separate path for special category data

Article 9 of the GDPR treats health, genetic and biometric data as special categories, and Article 10 covers criminal records. Documents that contain them run in a separate workflow with narrower access, and where needed the model runs on your own server.

No solely automated decisions

Article 22 of the GDPR gives people the right not to be subject to decisions based solely on automated processing that significantly affect them, and the EU AI Act lists AI used for recruitment or creditworthiness as high risk in Annex III. Rejecting an application or stopping a payment therefore always stays with a person.

Transfers outside the EU

If the reading service sits outside the EU or EEA, document content leaves it. Chapter V of the GDPR then requires an adequacy decision or safeguards such as standard contractual clauses, and the provider signs a processing agreement under Article 28. Your legal adviser makes the final call.

Original files and audit trail

The source document is never overwritten; extracted data is stored separately with the model version and the person who approved it. If an entry is questioned, you can show which document it came from and how, and roll it back.

Sources: Federal Ministry of Finance (Germany): FAQ on e-invoicing · General Data Protection Regulation (2016/679), Articles 9, 10, 22 and 28, EUR-Lex · EU AI Act (Regulation 2024/1689), Annex III, EUR-Lex

Comparison

Template OCR or AI document processing?

TopicTemplate OCRAI assisted workflow
New layoutsA new template for each layoutIntroduced with samples, no template
What is readThe text on the pageText, tables and what each field means
Catching errorsMisreads go into the recordConfidence scores and rule checks
Unclear documentsFields silently left blankSent to the review queue
Cost per documentLow and fixedVaries with model usage
SetupQuick, for one layoutInventory, test set and pilot

Quick check

Document processing feature list

Must haves: is your process ready?

0 of 6 in place Tick the boxes to see how ready you are for automation.

Added as needed

  • E-invoice XML integration
  • Matching with orders and delivery notes
  • Handwriting and poor scan support
  • Documents in several languages
  • Model hosted on your own server
  • Weekly accuracy report

We choose which of these you need together during the first call.

Let us start with the document type that eats most of your time

Share ten to twenty anonymised samples of one document type and the system the data should go to; we will define the fields, the review rules and a written quote.

Process

From discovery to launch in four steps

  1. First call and discovery

    We listen to your processes in a free 15-minute call. Then discovery maps your tools and tasks, scores the opportunities and ends with a written scope and fee for your approval.

  2. Build and test

    We build the first workflow in your accounts and test it with real but masked examples. Approval steps, error scenarios and alerts go in before anything reaches a customer.

  3. Go live and tune

    We switch the workflow on step by step, watch the logs and adjust thresholds with your team. You get documentation and a short training session.

  4. Monitor and expand

    On the monthly plan, we monitor running workflows, adapt them to model and API changes and add new workflows from the priority list, with a monthly report.

Free tools

Prepare your document workflow with free tools

Check VAT on invoice amounts, convert foreign currency invoices, work out due dates, create file hashes to catch duplicates, measure the hours spent on manual entry and generate strong passwords for new accounts.

Finance

VAT Calculator

Add VAT and extract it with the correct formula; preset + custom rates.

Currency

Euro Converter (ECB Rates)

Convert euros and 30 currencies with official ECB reference rates, look up any past date since 1999 and see monthly averages.

Calculator

Date Calculator

Days, weeks and working days between dates; add/subtract from a date.

Security

Hash Generator

Create MD5, SHA-1, SHA-256 and SHA-512 hashes of text and files in your browser, verify a download against its checksum and compare hashes. Nothing is uploaded.

Work

Working Hours Calculator

Calculate daily and weekly working hours after breaks, in hours and decimals, and check legal breaks and rest periods for the UK and EU.

Security

Password Generator

Cryptographically random strong passwords + strength meter + crack time.

All free tools

How we work

We start with one document type and expand on evidence

We do not yet have a live client project in AI document processing that we can show as a reference, so instead of promising results we describe our method. You can see our automation, software and web projects on the references page.

Real samples first

Before any build, we put together an anonymised sample set from your recent documents; accuracy is measured on that set.

Shadow run

In the first weeks the workflow runs alongside your team without posting anything and only shows its results. It goes live once the list of differences is clear.

Replaceable reading layer

If the model or reading service changes, rules, review screen and integrations stay in place, and the new layer is measured again on the same test set.

Accounts and rules stay with you

Model, reading service and automation accounts are opened in your company's name; field definitions, rules and documentation are handed over to you.

All references

FAQ

Questions about AI document processing

If your question is not here, write to us; we will send you an answer and a written quote.

Next step

Let us plan your first document workflow

Tell us which documents you handle, the rough monthly volume and the system the data should go to; after a free 15 minute call we will send the scope and a written quote.

In-depth guide

AI Document Processing: From Inventory to Approved Entries

Talha Aslan and teamLast updated: 15 min read

Most document automation projects succeed or fail on small decisions made long before anyone picks a model: which document type goes first, which fields are mandatory, which rule stops an entry and who looks at the review screen. This guide walks through those decisions in the order a business owner faces them when planning AI document processing.

It is a working checklist, not a brochure: criteria to put to any vendor, preparation you can do with your own files, and the cases where automation is not worth it. Technical terms are explained the first time they appear.

Measure fit with a document inventory

Whether automation suits you becomes clear only when you count the documents that actually arrived over the last three months; gut feeling tends to produce needless or half finished projects. Go through shared folders, inboxes and scanner output for a week and fill in a simple sheet.

  • Type and source: Supplier invoice, expense receipt, delivery note, contract or application form, and which channel and sender each one comes from.
  • Monthly volume and timing: Whether documents arrive evenly through the month or pile up in closing week.
  • Layout variety: How many different suppliers or templates exist for the same type; one large supplier or hundreds of small senders.
  • Manual time: How many minutes it takes on average to open, check and key in one document.
  • Cost of an error: When a wrong amount gets noticed, by whom, and what fixing it takes.

The sheet argues against automation in three cases: volume is low, most documents already arrive as structured electronic invoices, or each document type shows up only a few times a year. Then the import feature of your accounting software plus a tidy folder structure gives the same relief for less effort. If one document type with shifting layouts holds up every close, you have found your first workflow.

The first workflow by type of company

Start with the document type that eats the most staff time and whose mistakes cost the most; automating everything at once makes measurement impossible. Typical candidates:

  • Wholesaler or distributor: Purchase invoices and delivery notes from many suppliers; the first goal is matching them to purchase order lines.
  • Logistics or freight company: Proof of delivery forms, photos of signed delivery notes and customs paperwork; the first goal is linking each document to the right shipment and customer.
  • Accounting or bookkeeping firm: Receipts, bills and bank statements that clients send in a jumble; the first goal is sorting them by client and period.
  • Insurance broker or claims handler: Adjuster reports, registration papers and repair invoices attached to a claim; the first goal is a list of what is still missing.
  • Project or services business: Customer and vendor contracts; the first goal is a register of parties, term, renewal and notice dates.

An accounting firm that collects paperwork from clients should get documents arriving through one channel before adding AI document processing; a secure upload area on a website for accounting firms is a sensible place to start. The first type you choose becomes the template: the same review screen, audit trail and measurement sheet are reused for the second.

Plain language architecture in seven layers

A robust document workflow consists of separate layers that can each be tested on their own. Ask any vendor to walk you through them; "the AI reads it and posts it" describes a system where errors cannot be traced.

  • Intake: An inbox, shared folder, scanner or upload form is watched; each file gets a unique ID and its original is copied to storage nobody can edit.
  • Preprocessing: Multipage scans are split into documents, skewed images straightened and blank pages removed.
  • Classification: The document type is identified, because invoices, delivery notes and contracts follow different field schemas.
  • Extraction: OCR (optical character recognition, turning the letters in an image into text) and the model on top name the fields and give each a confidence score.
  • Validation: Rules run on totals, formats, duplicates and order matching.
  • Review: A document that fails a rule or scores below threshold lands on a person's screen.
  • Posting: The approved entry is written to the target system and stored with a link to its source file.

The payoff is replaceability: when the reading service changes, only extraction is measured again.

Do not read structured invoices as images

Sending structured documents through an AI model adds cost and risk for no gain. Formats such as Peppol UBL, XRechnung, ZUGFeRD and other invoices following the European standard EN 16931 already carry amounts, tax and line items as separate fields that your accounting software can import directly.

Confusion starts when a supplier emails a PDF rendering while the structured file sits elsewhere. Germany's Federal Ministry of Finance, for instance, states that domestic businesses must be able to receive electronic invoices since 1 January 2025 and that a plain PDF does not count as one. A workflow should handle every incoming PDF in this order:

  • Look for the structured file first: If the same invoice number exists as XML in your electronic invoicing inbox, the PDF is treated as a visual copy and not read.
  • Then link the two: When PDF and XML belong to the same invoice, the entry is built from the XML and the PDF is attached to the archive.
  • Only then send it to the model: Paper, scans, photos and plain PDFs with no structured counterpart go to the reading layer.

This split does two things at once. The model reads fewer documents, so usage fees drop, and the same invoice can no longer be posted twice, once from XML and once from PDF.

Field schema and cross check rules

A field schema is the written list of what gets extracted from each document type and in which format; it is the contract of the workflow. Without one, output drifts and the target system quietly accepts it. For every field, write down the name, data type, whether it is mandatory and its standard format: dates in one format, one decimal separator, currency as the three letter international code.

Next, list the cross check rules for each type. A typical starting set for invoices:

  • Arithmetic: Line amounts add up to the net total, net total times rate gives the tax, and both together give the gross total.
  • Tax ID format: A US employer identification number has nine digits, while EU VAT numbers carry a country prefix and can be checked against the European Commission's VIES service; an ID that fails is not considered read.
  • Date logic: An invoice date cannot lie in the future, and a due date cannot precede the invoice date.
  • Duplicates: If the same sender, invoice number and amount were already posted, the document stops.
  • Three way match: Quantity and price are compared with the purchase order and delivery note; any gap beyond your tolerance goes to review.

To catch the same file uploaded twice under another name, compare content hashes, fixed length fingerprints computed from each file; try it with our hash generator.

Intake channels and target system connections

Entry and exit points matter as much as the reading layer; if the target system cannot accept entries, even an accurate model is of little use. On the intake side, write down for each channel who may send documents, which file types are accepted and the size limit.

On the output side, the first question is how the target system takes in records. Accounting packages, ERP and CRM software are usually fed in one of three ways:

  • Documented API: The entry is created at once, and errors flow back into the queue.
  • Import file: Approved entries are exported as a file at set intervals; slower, but often the safest route with older accounting software.
  • Staging table: The workflow writes to its own database and the target system reads it; a last resort for closed legacy systems.

Whichever route you choose, two rules hold. First, resending an entry must never create a second one; a retry after a network outage should not produce a duplicate bill. Second, the workflow may suggest a general ledger account but should not finalize it in the first months; the suggestion appears on the review screen and the accountant decides. If the target system has no connection point, a small middleware service is planned separately as custom software development.

Choosing and hosting the reading model

There are three main options for the reading layer, and your own test set should decide between them, not a vendor's landing page. Prebuilt document services offer a fast start on common types such as invoices and receipts; general purpose models that read image and text together cope better with changing layouts and free text; open models running on your own server keep data in house.

Compare them on the same test set using these criteria:

  • Character accuracy: Whether accented names, umlauts and special characters in supplier names and addresses come out right.
  • Tables across pages: Whether a line item table split by a page break is merged back into one table.
  • Meaningful confidence: Whether high scoring fields really contain fewer errors; if score and error rate are unrelated, a threshold is pointless.
  • Version pinning: Whether you can lock a specific model version and how much notice you get before it is retired.
  • Data terms: Whether submitted documents are used for training, how long they are kept and in which region they are processed.

For companies handling medical reports or identity documents, a local LLM setup is a serious option, though hardware, updates and security become your responsibility. Many businesses settle on a hybrid: sensitive types local, everything else on a reviewed cloud service.

Review screen, thresholds and audit trail

The review screen is where people spend most of their time in AI document processing, and a poorly designed one gives back the hours that automation saved. A good screen shows the document image and the extracted fields side by side, highlights the region where each value was read and colors only the doubtful fields. The reviewer checks the flagged fields, not the whole page.

Set thresholds per field rather than as one number. A small misreading of a supplier name usually corrects itself when matched against master data, while one wrong digit in the gross total is money lost. So amounts and tax IDs get strict thresholds and description fields looser ones. Above a certain invoice amount you can require a second approver regardless of the score.

The audit trail for each document should hold at least:

  • Arrival time, channel and content hash of the file.
  • The service or model used for reading, with its version.
  • Each field's first value, confidence score and corrected value, if any.
  • Who approved or rejected it, when, and for what reason.
  • The entry number created in the target system and any reversal.

Review corrections weekly and add recurring ones to the test set.

GDPR, special categories and transfers

Documents often carry personal data: a sole trader's name on an invoice, an ID number on an application, a diagnosis on a medical certificate. Before building, draw a data flow map: which personal data each type contains, where it travels and how long it is kept.

  • Special categories: Article 9 of the GDPR covers health, genetic and biometric data, and Article 10 covers criminal convictions. Documents containing them run in a separate workflow with narrower access and, where possible, masking.
  • Processors and transfers: The reading service signs a processing agreement under Article 28. If it sits outside the EU or EEA, Chapter V requires an adequacy decision or safeguards such as standard contractual clauses.
  • Automated decisions: Article 22 gives people the right not to be subject to decisions based solely on automated processing that significantly affect them. The workflow reads a file and lists what is missing; it does not reject the application.
  • Health records in the US: A HIPAA covered entity needs a business associate agreement with any vendor handling protected health information for it.

The EU AI Act lists AI used in recruitment and creditworthiness assessment as high risk in Annex III, so a workflow that screens CVs or scores loan files carries a very different set of obligations from simple invoice capture. Your legal adviser makes the assessment; the technical team documents the data flow and settings.

Rolling out from pilot to full use

A workflow goes live in stages, each based on the measurements of the one before. This skeleton suits a single document type in most businesses.

  1. Collect the sample set: Set aside anonymized documents of the chosen type from recent months, including different senders and bad scans.
  2. Write the correct answers: Key in the true value of every field for each sample; this sheet is the reference for every later measurement.
  3. Measure the prototype: Run reading and rules on this set only and produce an error list per field.
  4. Run in shadow mode: The workflow reads live documents but posts nothing; staff work as usual and both results are compared daily.
  5. Go live narrowly: First only one group of senders, or documents below a set amount, may pass without review; the rest stay on the review screen.
  6. Widen the scope: As long as the error list stays acceptable, add more senders, then a second document type.

Shadow mode is the stage most often cut short, yet it is the only place to see behavior on real documents. Keep it running through at least one month end close; hastily scanned paper in closing week reveals errors quiet days never show.

Measuring time saved and accuracy

Value shows up in several measures that balance each other, not in a single rate; if the straight through share rises while corrections rise too, the workflow got faster and worse. Record the starting point before anything changes: tracking time per document for a few weeks with our working hours calculator is the only solid basis for a later comparison.

  • Error per field: In a weekly random sample, the share of each field that was correct, corrected or left empty.
  • Time in the review queue: From the moment a document lands in the queue to approval; if the queue piles up, change the staffing plan rather than the threshold.
  • Close calendar: On which working day the month end close is finished, compared with the months before automation.
  • Errors found later: Mistakes caught after posting by a supplier, a customer or an auditor.
  • Usage fee per document: The model and reading service bill divided by documents processed, tracked monthly.

Read the numbers with seasonality in mind; year end or holiday peaks can make AI document processing look better or worse than it is.

Limits and overlooked risks

AI document processing reduces reading errors but does not remove them, and some risks did not exist with classic OCR at all. Write down a countermeasure for each limit before the build.

  • Invented values: Language models may fill an unreadable field with a plausible value instead of leaving it blank. Instructions must require empty output for unreadable fields, and rules must catch those gaps.
  • Instructions hidden in documents: Someone can embed text invisible to the human eye in a PDF to steer the model. The reading layer should only have permission to extract fields; no document content can issue commands to the workflow.
  • Silent version changes: When a provider updates its model, the same document may be read differently. Pin the version and rerun the test set on every change.
  • Stamps, signatures and handwritten notes: A stamp over the total or a correction in the margin leaves the model unsure which value to pick; such documents go to review under a strict threshold.
  • Long documents: In contracts running to dozens of pages the model can miss the relevant clause; require every extracted item to cite its page and clause number.

These measures make risk manageable, not zero, so the review queue is never switched off; it only gets narrower.

Common mistakes in document automation

Most problems come from preparation and rollout order, not from the model.

  • Deciding on a demo invoice: Extraction that looks flawless on a vendor's clean sample behaves differently on your blurry receipt photos; decide with your own sample set.
  • Sending structured invoices to the model: Reading the PDF rendering when the XML exists creates fees and duplicate entries; separate structured data first.
  • One global threshold: The same confidence threshold for all fields produces either needless review work or dangerous pass throughs; set it by field and amount.
  • Leaving the review screen for later: Treating it as a plain table while polishing extraction burns the saved time in the queue; design it alongside the first prototype.
  • Skipping shadow mode: Connecting a workflow straight to the books because it did well on the test set carries real world errors into accounting; run in parallel for at least one close.
  • Losing corrections: If fixes on the review screen are not recorded, the workflow repeats the same mistakes; tie every correction to the audit trail and the test set.

One more turns up often: involving the accounting team late. People who will use the review screen should see the field schema and rules from day one.

Choosing a partner and next step

A good partner asks about your samples, target system and closing calendar before showing a demo. When requesting quotes for AI document processing, ask for written answers on these points:

  • On which sample set will accuracy be measured, and per field or per document?
  • How will structured electronic invoices be kept out of the reading layer?
  • If the reading service is replaced, do the rules, review screen and integration have to be rebuilt?
  • In whose name are model, reading service and automation accounts opened, and who receives the field schema and rules at handover?
  • Are the data flow map and a vendor contract checklist part of the scope?

We deliver this as part of our AI automation services. We have no live client project in AI document processing to show as a reference yet, so we describe our method rather than promise results; our automation, software and web work is on the references page.

Cost depends on the number of document types, volume, connected systems and the scope of upkeep; fixed price options for discovery and the first workflow are on our pricing page. Send the most time consuming document type, rough monthly volume and target system via the contact form, and we will define the field schema, review rules and a written quote with you.