Short answer

Business process automation with AI pays off most where the input is unstructured text: invoices as PDFs, emails, long documents, messy data. Where a rule can be written down in advance, ordinary code is cheaper, faster and predictable. In both cases the most reliable model is the same: AI prepares a proposal, the system checks it against rules, and a person approves decisions where a mistake would be costly.

When AI automation helps a business, and when rules are enough

In this article, AI automation for business means large language models (LLMs) that read text or images and return a structured result: fields, a category, a summary or a draft reply. Rule-based automation is ordinary code: "if the sender is X, route to Y", "if the amount exceeds Z, require approval".

The difference is practical. A rule gives the same result every time, and its error can be found in the code. A language model copes with wording nobody anticipated, but sometimes makes mistakes that look convincing. So the first question before every project is: can we write down a rule?

Task Rules are enough AI makes sense
Supplier invoices Data arrives in a structured format (XML, API) PDFs and scanned documents arrive in different layouts
Email sorting The category is determined by sender, subject or form The category is determined by free-form text
Customer replies Standard status notifications Free-form questions where the answer depends on context
Data validation Format, required fields, totals Matching names, finding duplicates by meaning
Reports Figures and charts from the database Summaries of long texts or conversations

The best solution is usually a mix. Rules handle most of the flow, and AI gets only the cases the rules do not recognise.

The market is moving in this direction. According to Eurostat, in 2025 21.3% of Lithuanian enterprises with 10 or more employees used at least one AI technology, against an EU average of 19.95%. Lithuania's figure rose by 12.5 percentage points in a year (it was 8.76% in 2024). Across the EU, the most commonly used technology was text analysis (11.75% of enterprises). Enterprises that had considered AI but not adopted it cited as the main barriers a lack of expertise (70.89%), unclear legal consequences (52.52%) and data protection concerns (48.83%). We address all three below.

1. Reading documents and invoices

A typical case: accounts or purchasing receives invoices, delivery notes and contracts as PDFs or photos. A language model reads the document and fills in predefined fields: supplier, company registration number, date, line items, amounts, VAT.

The capability here is real but limited. For its Structured Outputs feature, OpenAI states that the model's response will conform to the supplied JSON schema, meaning every field will be in its place. But the same vendor openly writes that the feature does not prevent all mistakes: the model can still get the values themselves wrong. In other words, the structure is guaranteed, but the correctness of the values has to be checked separately.

So after AI extraction we recommend automated checks:

  • line totals must match the grand total, and VAT must match the rate;
  • the supplier is matched against the existing supplier list by company registration number, because names are often written differently;
  • the order number must exist in the system;
  • an unusual amount or a new supplier is sent to a person for review.

A document that passes all checks is processed further automatically. Mismatches and unclear cases go into a review queue, where an employee sees the original next to the extracted fields. Passing the data to the accounting system is a standard integration, which we describe on our integrations and automation page.

2. Classifying and routing emails

A shared inbox receives enquiries, complaints, invoices, job applications and advertising. A language model assigns a category, determines urgency, extracts the order number and routes the email to the responsible person or department.

It is worth exhausting rules first. Invoices from known suppliers, automated notifications and form submissions can be recognised without AI. The model is left with emails written in free form.

The most important design decision is an "unclear" category. A model forced to choose one of five categories will choose, even if none fits. If it is allowed to answer "unclear", such emails go to a shared queue instead of the wrong department. Accuracy should be measured before launch: take a few hundred emails that have already been sorted, compare the model's decisions with people's decisions, and only then decide which categories can be routed automatically.

3. Draft replies in customer service

Here a language model prepares a draft reply based on the customer's question, the order data and the company's knowledge base: policies, price lists, previous replies. This architecture is called RAG (retrieval-augmented generation): the system first finds the relevant documents, and the model writes the reply based on them, not only on its general knowledge.

When an employee reviews and sends the draft, it is a work tool, and responsibility for the reply stays with the person. When a customer talks directly to a chatbot, different rules apply. Under the EU Artificial Intelligence Act (Regulation (EU) 2024/1689), systems that interact directly with people must be designed so that the person is informed they are interacting with an AI system, unless this is obvious from the context. Most of the Act applies from 2 August 2026.

In practice we recommend starting with drafts. They save time straight away, and employees' corrections show which topics the model gets wrong. Only then can you decide which types of question are worth handing over to autonomous replies. If you need a system that not only writes but also takes actions, for example checking an order status or creating a return request, that is already an AI agent. We write about the limits and permissions of such agents on our AI agent development page.

4. Summaries of long texts and conversations

Summaries are one of the safest uses of AI, because a person reads the result anyway. Typical cases: reviewing procurement documents or contracts, summarising meeting recordings, condensing a customer's history in the CRM into a few lines before a call.

There are two limits. First, a summary can miss an important detail, such as a penalty clause in a contract. So a summary should indicate which part of the document each statement comes from, so that key points can be checked against the original. Second, conversation recordings and customer history are personal data. Processing them falls under the General Data Protection Regulation (GDPR): you need a legal basis, and the model provider becomes a data processor. Article 28 of the GDPR requires using only processors that provide appropriate technical and organisational measures, and concluding a contract with them.

Data location can also be controlled. For example, since February 2025 OpenAI has offered API projects with European data residency: in such projects requests are processed in the region without storing requests and responses. The vendor also states that by default API data is not used to train models. Other providers offer similar options, but the terms should be checked in the specific contract and data processing addendum.

5. Cleaning data before migration

When replacing a CRM or business management system, the old data is rarely tidy: the same customer entered three times in different ways, addresses in a single field, product categories changed several times. Cleaning it by hand takes weeks.

A split approach works well here. Format issues (phone numbers, dates, company registration numbers, email addresses) are handled by ordinary code: it is cheaper and more accurate. The language model is used where meaning has to be understood: suggesting which records are duplicates, splitting a free-text address into fields, mapping an old category to the new structure. All suggestions are presented for review in batches, and approved changes are saved with a reference to the original record, so that any decision can be reversed.

This work is usually part of a wider project, such as custom CRM development, when old data is moved into a new structure.

Why a human approval step is essential

The first reason is legal. Article 22 of the GDPR gives a person the right not to be subject to a decision based solely on automated processing that produces legal effects concerning them or similarly significantly affects them. Where such a decision is based on a contract or explicit consent, the data controller must at least ensure the right to obtain human intervention, to express their point of view and to contest the decision.

The AI Act places some areas in the high-risk category. These include AI systems used in recruitment, for example to filter job applications and evaluate candidates. Such systems are subject to additional requirements, including human oversight. Regulation (EU) 2026/1744, which entered into force on 27 July 2026, postponed the application of these requirements: until 2 December 2027 for stand-alone high-risk systems, and until 2 August 2028 for those embedded in regulated products. The same regulation amended the AI literacy provision: providers and deployers must take measures that help develop their staff's AI literacy.

The second reason is practical. A language model's mistake often looks like a correct answer: a tidy amount, a convincing sentence, a plausible category. Rule errors are usually visible straight away, for example as an empty field or an error message, whereas a model's error has to be looked for. Human approval at critical points is the simplest way to stop such errors before they reach the customer or the accounts.

Approval does not necessarily mean a person reviews everything. Usually it is enough to show them what failed the checks, plus periodically sampled automatically processed records to monitor quality.

What a mistake costs and how to limit it

We recommend choosing the level of automation according to what a mistake would cost. A misrouted email costs a few minutes. A wrong amount in the accounts costs a correction and possible tax questions. A wrong answer to a customer about contract terms can cost a dispute.

Practical measures to limit mistakes:

  • Automatic execution only after checks. Only actions whose result has passed rule-based checks run automatically. Everything else goes to the review queue.
  • Least privilege. The AI component gets only the data and actions the task requires. A system that writes summaries must not be able to change orders.
  • Action log. Every AI suggestion, check result and human decision is recorded, so that a mistake can be traced and corrected.
  • Accuracy measurement. Before launch and periodically afterwards, AI results are compared with people's decisions. When the model version changes, the measurement is repeated.
  • Fallback path. If the model provider's service is down, the process falls back to a manual queue and does not stop.

It is also worth estimating running costs from the start: every request to the model is charged by the amount of text processed, so for large volumes it is cheaper to filter out first what rules can handle.

What to do next

If you want to assess which of your processes are worth automating with rules and which suit AI, start by describing a single process: what goes in, what comes out and what a mistake costs. How we structure such projects and where we place the checks is described on our integrations and automation page.

Sources