No accuracy promises upfront
We do not quote an accuracy figure before seeing your data. Anyone who does is quoting someone else’s benchmark.
Four kinds of applied AI work and the four questions we ask before agreeing to any of them.
All of it applied to a specific business process with a measurable before and after. None of it sold on the promise that the technology is impressive.
Extracting structured data from invoices, contracts, specifications and the assorted PDFs a business drowns in. Measured against a labelled set from your own documents, not a public benchmark.
Sorting incoming work — tickets, applications, messages — to the right queue. Usually the highest-value AI application in a company and the least glamorous.
Internal assistants that answer from your own material, cite where the answer came from, and say "I do not know" instead of inventing. That last property is engineered, not hoped for.
The part almost everyone skips: a repeatable test set that tells you whether a prompt or model change made things better or quietly worse.
Four limits, stated before you spend a call finding them.
We do not quote an accuracy figure before seeing your data. Anyone who does is quoting someone else’s benchmark.
Client data is not used to train shared models. The processing arrangement is written down before anything moves.
We flag where a use case looks higher-risk and recommend a lawyer. We are engineers.
Customer-facing decisions get a human in the loop. That is a design rule, not a phase we remove later.
Evaluation first. It is the cheapest stage and the one that most often ends the project early and correctly.
How is the task done today, how long does it take and how often is it wrong? Without this, no later claim of improvement means anything.
A few hundred labelled examples from your real data, including the awkward ones. This is the honest part of the project.
The smallest system that beats the baseline, with a human review step wherever an error would reach a person.
Monitoring in production, because model behaviour drifts and a system nobody watches degrades quietly.
If the answers are unsatisfying, the correct outcome is a short invoice and no project.
An error in an internal draft is cheap. An error in a customer-facing decision is not. The answer determines how much human review the design needs.
Time per case and current error rate. Nobody measures this before an AI project, which is exactly why so many are declared successful without evidence.
A few hundred examples with correct answers. If nobody in the business can produce them, nobody in the business can tell whether the system works.
A model endpoint will be unavailable at some point. If there is no fallback, you have introduced a dependency rather than a capability.
Regularly. A rule engine or a better form is often cheaper, more predictable and easier to defend. That answer costs us a project and saves you one.
Yours stays yours. We do not use client data to train shared models, and any processing arrangement is written down before data moves.
Most internal document and classification work is low-risk, but the obligations depend on the use case. We flag where a use case looks like it could fall into a higher-risk category and recommend legal review. We are engineers, not lawyers.
Yes, and the first step is building the evaluation set that should have existed. Without it, nobody can tell whether a change helped.
We integrate whatever fits — hosted APIs or self-hosted models. The choice depends on the data sensitivity and the cost profile, not on a preference.
Registered in Cyprus, delivering across the EU, in English and German.
Tell us the task and how it is done today. If the answer is that you do not need a model for it, that is what you will hear.