Most AI demos work. Very few systems do.
ZORATHEN LTD builds AI-assisted software and, just as importantly, the evaluation harnesses that show whether it holds up on your data rather than on a curated example. Registered in Cyprus, delivering across the EU.
- A labelled evaluation set from your own data
- Measured accuracy, and where it fails
- A human review step wherever errors reach people
- A fallback for when the model is unavailable
No AI feature goes live without a measured baseline
The gap between a demo and a system is entirely made of the cases nobody put in the demo. Our job is to find those first, measure them, and decide honestly whether the thing is ready.
Four kinds of AI work
All of it applied to a specific business process with a measurable before and after. None of it sold on the promise that the technology is impressive.
Document processing
Extracting structured data from invoices, contracts, specifications and the assorted PDFs a business drowns in. Measured against a labelled set from your own documents, not a public benchmark.
Classification & routing
Sorting incoming work — tickets, applications, messages — to the right queue. Usually the highest-value AI application in a company and the least glamorous.
Assistants with guardrails
Internal assistants that answer from your own material, cite where the answer came from, and say "I do not know" instead of inventing. That last property is engineered, not hoped for.
Evaluation harnesses
The part almost everyone skips: a repeatable test set that tells you whether a prompt or model change made things better or quietly worse.
How an AI project runs
Evaluation first. It is the cheapest stage and the one that most often ends the project early and correctly.
Baseline
How is the task done today, how long does it take and how often is it wrong? Without this, no later claim of improvement means anything.
Eval set
A few hundred labelled examples from your real data, including the awkward ones. This is the honest part of the project.
Build
The smallest system that beats the baseline, with a human review step wherever an error would reach a person.
Watch
Monitoring in production, because model behaviour drifts and a system nobody watches degrades quietly.
Four questions we ask before building anything
If the answers are unsatisfying, the correct outcome is a short invoice and no project.
What does wrong cost?
An error in an internal draft is cheap. An error in a customer-facing decision is not. The answer determines how much human review the design needs.
What is the baseline?
Time per case and current error rate. Nobody measures this before an AI project, which is exactly why so many are declared successful without evidence.
Can we label the data?
A few hundred examples with correct answers. If nobody in the business can produce them, nobody in the business can tell whether the system works.
What happens when it is down?
A model endpoint will be unavailable at some point. If there is no fallback, you have introduced a dependency rather than a capability.
Substance you can look up
ZORATHEN LTD is a private limited company registered in the Republic of Cyprus, directed by Diana Oprea from Larnaca.
The company builds AI-assisted software for specific business processes, together with the evaluation harnesses that make improvement claims checkable.
Everything on this page can be checked against the public register. The registration number, the registered office and the director are printed below and repeated in the imprint.
- Registered EU company — Republic of Cyprus
- Every AI feature ships with a measured baseline
- Source code and evaluation sets transfer to the client
Clearly answered
Will you tell us AI is the wrong tool?
Regularly. A rule engine or a better form is often cheaper, more predictable and easier to defend. That answer costs us a project and saves you one.
Whose data trains what?
Yours stays yours. We do not use client data to train shared models, and any processing arrangement is written down before data moves.
What about the EU AI Act?
Most internal document and classification work is low-risk, but the obligations depend on the use case. We flag where a use case looks like it could fall into a higher-risk category and recommend legal review. We are engineers, not lawyers.
Can you improve an AI feature we already have?
Yes, and the first step is building the evaluation set that should have existed. Without it, nobody can tell whether a change helped.
Do you host models?
We integrate whatever fits — hosted APIs or self-hosted models. The choice depends on the data sensitivity and the cost profile, not on a preference.
Which markets do you serve?
Registered in Cyprus, delivering across the EU, in English and German.
An AI project that needs an honest second opinion?
Tell us the task and how it is done today. If the answer is that you do not need a model for it, that is what you will hear.