How to Test a New AI Model on Your Own Business Data Before You Trust the Hype

Published 2026-10-07 by Dale Taylor

In short

To test a new AI model on your own business data, check two things before you run it: whether almost every answer is already "yes", and whether the inputs the model needs are actually filled in on every row. If either check fails, no model will help, however good it is. I ran a new decision model called Jev on three jobs in my own business. Two failed before or during the run for exactly these reasons, and the third worked because both checks passed.

Why a list of use cases tells you very little

A new model turns up and your feed fills with lists. "19 use cases that will blow your mind." "8 things that feel like cheating." The numbers in those videos are usually reported, not measured. Six times faster, two hundred times cheaper. I have not repeated any of them here, because I have no way to check them.

What a list cannot tell you is whether the model suits your data. That depends on what you have already got sitting in your CRM or your spreadsheets, and nobody else can see that.

So I tried three of the use cases on my own business. I run Aigentic, a one-person AI automation business in Skipton, and I do most of the work through Claude Code. That means I can usually put a new model against real data the same day it appears. This one was a decision model called Jev. Two of the three jobs did not work. The reasons were visible in my own data before I spent a penny on the model.

What is a decision model?

Most language models give you text back, and you then parse it. A decision model gives you a typed answer. Ask it "is this the same product, yes or no" and you get a true or false value with a confidence number, not a paragraph to read.

That makes it handy for any job that is really a yes or no question asked many times. It is also cheap and quick. My lead scoring run, which I will come to, was 92 calls for $0.0022 in total, at about 1.4 seconds each.

One practical note, because it cost me some time. Decision models are rejected by the normal chat completions endpoint on OpenRouter. You need the separate decisions endpoint, and you do not need a key from the model's vendor to use it.

Test one: scoring the videos I watch

I have a pipeline that reads the AI videos I save and scores each one out of ten for how relevant it is to my business. It had been doing that across 63 daily digests. The obvious job for a decision model was to score them better.

Before wiring anything up, I took 73 scores the pipeline had already produced. I labelled 34 of them myself, blind. I could see the title, the channel and the summary, but not the score the pipeline had given.

The correlation between my scores and the pipeline's was minus 0.12. That is not a weak relationship. It is no relationship. The score had been shaping what I read for two months and it was close to noise.

Then I found the thing that actually ended the idea. Of the 49 items I labelled, 42 were things I would act on. That is 86 percent.

When 86 percent of the answers are yes, the best classifier available is one that says yes every time. There is no accuracy left for a better model to win. So I never ran Jev on this job. I removed the score from the pipeline instead.

This is the first check. **Work out the base rate before you reach for a model.** If almost everything passes, you do not have a classification problem. You have a filter that filters nothing.

Test two: predicting who replies to a cold email

The second job looked much better suited. I had 92 companies I had cold emailed, and four of them sent back a genuine human reply. Could the model pick which four?

This is the right shape of problem. The base rate is low, the real outcomes were already in my CRM, and the answer would be useful if it worked.

I scored all 92 blind, using only the enrichment data that existed before each email went out. The result was an area under the curve of 0.476. A coin toss scores 0.5, so this was slightly worse than guessing. The four companies that replied ranked 30th, 32nd, 40th and 89th out of 92. None were in the top ten.

Four replies is a very small sample, and I would not lean on the ranking. The second finding is not underpowered at all. The model put the average company's chance of replying at 30 percent. The real rate was 4.3 percent. It was not occasionally wrong. It was confidently wrong across the whole list.

The reason is the useful part. Five of the fields it was reading were empty: tech stack, current tools, automation maturity, pain points and decision maker type. Not patchy. Zero populated out of 92. What the model actually received was a name, an industry, a one-line description and a headcount band, for 92 near-identical UK ecommerce agencies. Nothing in that tells you who replies.

I had built a lead scoring system with no signal in it, and then blamed the scorer.

This is the second check. **Count how many rows have each input filled in before you test.** A model cannot find a pattern in a column that is empty. An empty column is also worth a second look at the code that writes it, because it may be deferred or optional rather than broken. In my case it was simply never populated at scoring time.

Test three: the one that worked

The third job is why this is not just another post telling you a new tool is rubbish.

A client's product catalogue listed the same stock on two marketplaces under two different titles. The job is deciding whether two listings are the same product in the same pack size. I am not naming the client.

I had 134 pairs with a verdict. Jev got 96.3 percent of them right. The rule I had written by hand to do the same job got 86.6 percent.

Ten points sounds modest until you look at the mistakes. My rule let through 17 pairs that were not the same product. Jev let through one. Each of those 17 is a wrong price going into a valuation, so that is the difference that matters.

There is an honest caveat, and it is the reason I am not printing a bigger claim. I also ran the same set past Claude Opus 5.5, which scored 91.8 percent. Jev came out about four and a half points ahead. I do not believe that gap.

My ground truth is not all mine. I labelled 69 pairs by hand and let the other 84 stand where three scorers had already agreed. I spot-checked 15 of those and disagreed with the agreed answer on one. Scaled up, that suggests around six wrong answers hiding in the set, and six wrong answers is roughly the size of the gap. So "it beat the frontier model" is a claim I cannot make. "It beat my hand-written rule by about ten points and cut the bad matches from 17 to one" is one I can, because that gap is far too big to be explained by a handful of bad labels.

It worked because it is the opposite of the first two. Every pair has two titles. Nothing is empty, and the answers are not almost all the same.

What the three tests have in common

None of this was about the model. It was about whether the inputs existed.

Both failures were visible in my own data before I spent a single call on a model. Two database queries would have saved me the afternoon. One counts how often the answer is yes. The other counts how many rows have each input filled in.

A short checklist before you test any model

1. **Pick a job with a known right answer.** You need outcomes you already trust, such as replies, matches or refunds, not outcomes you will have to guess. 2. **Label a sample yourself, blind.** Do it before you look at what the existing system or the new model says. 3. **Check the base rate.** If one answer makes up most of the set, stop. 4. **Count the filled-in fields.** If the columns the model would use are mostly empty, fix the data first. 5. **Say how reliable your answer key is.** If part of it came from machines agreeing with each other, spot-check it and work out how many wrong answers could be hiding in it. If that is as big as the gap you are measuring, you have not measured it. 6. **Compare against what you already have.** A new model has to beat the current rule, not beat nothing.

Key takeaways

  • Do not judge a new AI model by a list of use cases. Test it on your own data.
  • Before running anything, check the base rate. If 86 percent of answers are already yes, no classifier has room to improve on it.
  • Before running anything, count how many rows have each input filled in. Five empty columns out of five meant my lead scoring had no signal to find.
  • Of three jobs I tried with a decision model, one worked: matching duplicate product listings, at 96.3 percent against 86.6 percent for my hand-written rule, with bad matches cut from 17 to one.
  • Be honest about your answer key. If its likely error is the same size as the gap you are claiming, you cannot claim the gap.

Start with the Sample Five, five leads in your niche with a named decision-maker and a verified work email.