Training AI on your own data: what works and what doesn't
Almost every business owner who wants AI running on their own company data asks for something they don't need. The cheaper option works better and is easier to check.
TerenceTraining AI on your own data. That's how most business owners put it when they call. They want an assistant that knows their price lists, their contracts, their manuals and the agreements once made with a customer. Not an AI that knows the internet, but one that knows their company. Perfectly logical. Except that nine times out of ten, training isn't what needs to happen, and that saves a lot of money.
This article covers three things: the three meanings hiding behind that one sentence, which one you probably mean, and where it goes wrong in practice. With the numbers included, and without the tech talk.
Three things training AI on your own data can mean
When three business owners say the same sentence, they often mean three different things. It pays to pull them apart, because the price difference between them is a factor of a thousand.
- Attaching documents to a conversation. You drag a pdf into a chat window and ask questions about it. Free, instant, and gone once that conversation ends.
- Letting AI look things up in your own archive. Your documents are processed so that every question first pulls up the right passages, and the answer is based on those, with the source document named.
- Adjusting the model itself. You have an existing AI model learn from your material so that its behaviour changes.
Almost everyone already uses number one. Number three is what most people mean when they say training. And number two is what they actually want.
We rarely build number three. Not because it can't be done, but because in nearly every case you get nothing out of it and still pay for it. If someone is selling you a model trained on your data, first ask what goes wrong if they don't do it.
Why fine-tuning an AI model is almost never the answer
Fine-tuning changes how a model behaves: the tone it strikes, the format it keeps, the trade terms it uses. What it does not do reliably is remember facts. And facts are exactly what you're after when you want to know what warranty period is in that 2023 contract.
- Facts blur. A fine-tuned model gives an answer that sounds like your documents sound, not necessarily what they say.
- You can't check where the answer came from. No document, no page number, no verification.
- Every change costs another round. New price list in January? Start again.
- Removing things is hard. A customer asking whether their data is really gone won't accept mostly as an answer.
- It is considerably more expensive than the alternatives, in money and in time.
There is one exception, and honestly it's rare in a small business. If you have to produce the exact same kind of text many times a day in your own fixed format, think thousands of short standardised assessments, fine-tuning can make sense. Then you're buying form, not knowledge. If that isn't you, this isn't for you.
What does work: look it up first, then answer
What you almost always need is easier to grasp. Your documents, so manuals, quotes, contracts, minutes and price agreements, are processed so they can be searched by meaning and not only by exact wording. Someone asks a question in plain language. The system pulls up the handful of passages that actually matter. Only then is the answer written, based on those passages.
The difference with the search function you have now: that one finds files, this one finds answers. Search for solar panel warranty today and you get twelve file names back and you get to start reading. In the other case you get: five years on the installation, twenty-five years on the panel itself, with the source document listed underneath.
That source reference isn't decoration. It is the only thing that lets someone check whether the answer is right before acting on it. A system that can't name a source isn't finished for business use.
one-off processing of a thousand pages, at published rates
I start with a free introductory conversation. We discuss what takes time, which systems you use and where automation could help. You then receive a proposal with automation opportunities, integrations, a schedule and costs.
Where it goes wrong: not in the technology, but in your documents
In practically every project that stalls on this, the cause isn't the AI. Five classics.
- Conflicting versions. If there are three price lists in the folder and none of them is the leading one, you'll sometimes get one and sometimes the other. That's not an AI fault, that's your archive.
- Scanned documents without text. A photographed attachment is just an image to a computer. Solvable, but it is work.
- Access rights. If everything becomes searchable, the personnel file becomes searchable too. Who may see what is arranged beforehand, not after someone has found it.
- Knowledge that only exists in people's heads. What isn't written down anywhere can't be found either. Sometimes the first step is: write down those ten frequently asked things.
- Too few questions. If something gets looked up five times a week, you won't earn a system back. Then a tidy folder is the answer.
An assistant searching through messy documents makes the mess visible rather than solving it. If we see in a first conversation that the real problem is the archive, we say so. Then the first step is tidying up, not building.
We want a Dutch AI model, and why that's a misunderstanding
Logical thought: your documents are in Dutch, so you want a Dutch model. In September 2025 the first serious study testing that appeared. Language technologists at the University of Antwerp built MTEB-NL, a yardstick with forty Dutch test sets, including legal texts and public tenders.
The result is counter-intuitive. Models built specifically on Dutch text, BERTje and RobBERT, scored 17.6 and 18.3 points on retrieving the right passage. Multilingual, international models reached 59.1. On recognising and classifying texts those same Dutch models did reasonably well, 50.9 against 60.2. They understand the language fine. They just don't find anything.
search score of a Dutch model against a multilingual model (MTEB-NL, University of Antwerp, September 2025)
The explanation is in the same study: searching is a separate skill a model has to be tuned for separately. Take the same Dutch model and do tune it for search, and the score jumps from 18.3 to 51.6. So it isn't about the language of the model, but about whether it has been tested on searching in your language. Ask about that. It's a Dutch model is not an answer to that question.
Do your documents stay in the country?
This is the question asked most often and answered properly least often. For the search part there's good news: the best performing models in that study are freely available and can run on your own server. One of the good scorers is small enough that no heavy machine is needed. So your archive doesn't have to leave the building.
For writing the answer it's different. That usually involves a large model from a large provider, and then the question is where that runs. Some providers offer processing inside Europe, others don't, and it differs per part of their range. Ask about it, have it written into the contract, and don't settle for that'll be fine.
How big does your archive need to be before this pays off?
There's a rule of thumb. With a handful of documents you can simply attach them every time. That works fine and is cheaper. Above roughly forty pages it tips over: sending everything becomes more expensive than looking things up, however often you use it. And the bigger your archive gets after that, the less it matters. With targeted lookup, a question about a thousand pages costs practically the same as a question about forty.
In practice: a product catalogue, a folder of contracts or a collection of installation instructions is always well above that. Five procedures and a staff handbook are not.
Where small business stands today
This isn't a rearguard action, but it isn't a race you've already lost either. According to Statistics Netherlands, in 2025 33 percent of companies with ten or more employees used AI, against 23 percent in 2024 and 14 percent in 2023. In the ten to fifty employee band it sits at 28 percent.
AI use at companies with 10 or more employees in 2025; at 10 to 50 employees 28 percent (Statistics Netherlands)
And that use is shallow. Research into AI use in small business published by the Dutch government in September 2025 shows that it is mostly aimed at writing text, communication and summarising documents. More specialised applications are rare. Which means exactly there, in the work tied to your own documents, it is still wide open.
More interesting than the percentage is the reason companies don't start. In those same figures, non-users name a lack of knowledge and experience about four times as often as excessive cost. So it isn't a money problem. It is an I-don't-know-where-to-start problem.
How to go about it
- For two weeks, note which questions actually get asked and where the answer was. Not from memory, but written down, with the time the search took.
- Count how many pages it genuinely involves. Usually that's a fraction of what's in the folder.
- Designate one leading document per subject and take the rest out of the source. This is the step the project stands or falls on.
- Start with one kind of question from one team. Not everything searchable, but our engineers can pull up the installation instructions.
- Demand a source reference with every answer and have two people spot-check the answers for a week.
- After a month, measure the same thing as in step one. If search time hasn't dropped, you stop.
What this delivers is rarely spectacular to look at. No clever robot appears. What appears is someone who no longer has to phone three colleagues to find out which warranty period applies. The comment we hear most often after delivery isn't about the AI, by the way, but about the documents: we had no idea it was such a mess.
Don't do it if your documents still have to be written rather than found, if there are a handful of questions a week, or if the real problem is that nobody knows which document applies. That last one you solve with an agreement, not with software.
Ready to get started?
Request a free consultation. We look together at where you are losing time.
Schedule free call