People have no idea how to select the right AI for their tasks.


Elena, founder of Plain AI
Cyprus

A 99helpers Maker Interview
Conversations with independent makers about how they make, host and market the things they build — and the lessons they learned along the way. See all interviews
Make
Host
Market
The product

Plain AI
plainai.tech/
We test AI models on real business processes, like invoice processing, contract review and resume screening, and publish the results in plain language.
01
Who are you, and what did you build?
I'm Elena, I'm 38 and I live in Cyprus. My background is corporate finance and business automation: the part of a company where invoices, contracts, expenses and hiring turn into processes, and where a wrong decision costs real money.
I've been a founder several times. Most of my startups failed, and one was sold. I've also moved around a lot: I lived in Russia, then in Dubai, then in Latvia, and now in Cyprus. In Dubai I joined the in5 tech accelerator, and that's where I met Olga Nayda, my friend and now my co-founder.
Together we're building Plain AI. We run custom benchmarks of AI models on real business cases: we take a process a finance or operations team does every day, test the models on it, and publish the results in plain language, so anyone can compare them.
02
Why did you build it?
I noticed that people have no idea how to select the right AI tool for their tasks. Which model is the most efficient for this job, given what it can do, what it costs, and where it's available and allowed to be used? There are dozens of models, their prices differ enormously, and the well-known benchmarks are made for AI researchers. None of them tells a CFO which model should read their supplier invoices.
So we test the models on the work itself. For each benchmark we build a realistic set of documents and run every model through the same task. For contract review, for example, that was eleven contracts (NDAs, a lease, a software licence, a supply agreement, a bilingual German and English SaaS agreement) plus companion documents like a budget memo and a purchase order, and one vendor agreement engineered to look fraudulent.
Every model is scored against a weighted rubric: reading the data, the legal review, the accounting and tax review, fraud detection, whether it makes things up. We don't only count how often a model is wrong, but how expensive its mistakes would be. A model that misses the fraud has its score capped, however well it did on everything else.

03
Why should people use your product?
The product is for business people: CEOs, CFOs and other C-level managers. When they're faced with a choice about something they don't fully understand, they need a starting point, and they need reasoning they can give to other stakeholders: the board, the finance team, IT.
“They need a starting point, and they need reasoning they can give to other stakeholders.”
That's how we write the reports. Each one begins with how to read it and an executive summary a CFO can get through in a few minutes; the technical detail and the full results come later, for whoever has to check them. We also say when AI is not the answer. Sometimes redesigning the process, or ordinary rule-based automation, solves the problem better and cheaper.
And the results are often not what people expect, because price and quality don't move together. In our contract review test, five models got a perfect score. The cheapest of them cost nine cents for the whole batch, the most expensive $5.63: sixty-four times more for the same result. In invoice processing, the top model cost less than a cent per invoice, while the most expensive model tested finished second from the bottom.

Right now we're doing custom benchmarks for free. If there's a process in your business you're thinking of handing to AI, your readers can request a benchmark through our website, and we'll test the models on it.
04
How is the business doing today?
We're validating the idea. Three benchmarks are published so far: contract review, invoice processing and resume screening, and one on AI SEO translation is coming next. Now we're actively looking for people who could benefit from a custom benchmark and who are willing to have a product interview with us. We want to understand how decision-makers choose AI today, and what they need to see before they say yes.
- Benchmarks published
- 3
- Models per benchmark
- Up to 19
- Custom benchmarks
- Free
05
What has worked well so far?
As soon as we start talking about our project, our friends and family ask which AI is better for their specific task, and whether we can run a custom benchmark for them. People around us have exactly the problem we're solving, and they tell us without being asked.
So it's manual for now: we take those requests one by one. That's fine at this stage. Every custom benchmark teaches us what people actually want to compare, and how they read the results.
06
What has been challenging?
Actually doing the benchmarks, because it takes time, tokens and expertise. You have to design a test set that looks like the real thing, with traps in it, like a fake vendor or a candidate who isn't allowed to work in the country. You have to decide what the correct answer is and write the scoring rules. Then you run close to twenty models through it and check the results. The frontier models are not cheap to run, and the checking is where the expertise goes.
“Actually doing the benchmarks, because it takes time, tokens and expertise.”
And marketing!!! Building is the part we know how to do. Getting in front of CEOs and CFOs is much harder, because they're not in the places where AI people talk about models.
07
What are your goals, and what does success look like for you?
Money in the bank is the ultimate success. After several startups, that's what I measure in the end.
“Money in the bank is the ultimate success.”
While we're getting there, success is making our website a popular place for people looking for AI: the site you open before you choose a model or a tool, and the report you send to your board to explain the choice.
08
What's your tech stack?
Claude Code is our CTO. We don't have a developer on the team: Claude Code writes and maintains the code behind plainai.tech, and we decide what to build and check what comes out.
“Claude Code is our CTO.”
09
How do you use AI?
We use it to code, write, research, review and publish. Basically, what we do is start and check. Everything else is done with AI.
For the benchmarks, AI is also the subject. We test the models, and we use them to help with the work around the tests. But the judgment stays with us: what counts as a correct answer, and which mistake would be expensive for a real business. That's where the finance background matters.
10
What advice would you give someone who wants to build something of their own?
I have too many failed projects to give advice on how to be successful.
“I have too many failed projects to give advice on how to be successful.”
What I do differently this time is validate first. We talk to the people who might use the product, do the work by hand before we automate it, and only build what people actually ask for.
11
Any resources you'd recommend?
I personally like Y Combinator's Startup School. It's free, and it covers the basics that are easy to skip when you're excited about an idea: talking to users, finding out whether anyone really needs what you build, and launching.
12
Where can people find you?
You can find me on X at @ehmlvs, and Plain AI at plainai.tech. If you'd like a free custom benchmark for a process in your business, request one through the website.
At a glance
- Founders
- Elena and Olga Nayda
- Based in
- Cyprus
- Met at
- in5 tech accelerator, Dubai
- What it does
- Benchmarks AI models on business processes
- Benchmarks published
- 3
- For
- CEOs, CFOs and other C-level
- Stage
- Validating the idea
- Custom benchmarks
- Free for now
Are you a maker?
Tell us how you make, host and market what you build.
We feature independent makers and small teams — the wins, the mistakes and the lessons in between. If you've shipped something, we'd like to hear about it.