Measured, not demonstrated
Every AI feature ships with a labelled evaluation set and a number. Without one, "it works well" is a feeling.
A Cyprus product studio that treats evaluation as the deliverable and the model as an implementation detail.
Building something impressive with a language model takes an afternoon. Knowing whether it is good enough to put in front of anybody takes considerably longer.
The demo is easy because the demo is curated. The examples were chosen, the prompt was tuned to them, and the failures were not shown. Everyone doing this work knows it, and the incentive to keep quiet about it is substantial.
ZORATHEN inverts the order. We build the evaluation set before the feature: a few hundred labelled cases from real data, including the awkward ones. Then a claim of improvement is a measurement rather than an impression.
It makes for slower-looking projects and considerably fewer systems quietly withdrawn six months later.
Values are cheap to write down. These are the ones that cost us something when we follow them.
Every AI feature ships with a labelled evaluation set and a number. Without one, "it works well" is a feeling.
Automation is not the goal; the correct outcome is. Where a mistake would reach a customer, a person reviews it.
Rules, forms and better data beat a model more often than the market admits. We say so, at our own cost.
Diana Oprea is the director of ZORATHEN LTD and decides what the company takes on.
The evaluation-first approach is a commercial decision as much as a technical one: it makes it easy to end a project honestly before it becomes expensive.
No AI feature goes live without a number you could show to a sceptic.
Tell us the task and how it is done today. If the answer is that you do not need a model for it, that is what you will hear.