By now, almost every Mittelstand company has done something with artificial intelligence. A workshop, a test with a chatbot, a department that got a license. What’s striking is what happens next: often nothing. The topic disappears from leadership meetings, the license keeps running, and three people actually use it. The reason is almost never the technology. It lies in how the project was set up.
The five typical failure points
1. The pilot project had no problem to solve
The most common mistake is starting without a concrete trigger. AI gets introduced “because something has to be done about it.” The result is a tool that’s available but has no task. Employees try it out for two weeks, find it impressive, and then go back to their usual way of working — because the usual way works.
Successful rollouts start the other way round: with a task that measurably costs time and that nobody enjoys doing. The proposal that takes three hours every time because text blocks have to be pieced together from five old documents. The invoice check where someone matches line items against the purchase order. The report that’s built every month from the same sources. Anyone who starts with a task like this can say after four weeks whether anything has actually changed.
2. The model didn’t know the company
A general-purpose language model knows a lot about the world and nothing about your company. It doesn’t know your price list, your delivery terms, or the history with a particular customer. Ask it to write a quote and you get a well-worded text full of invented details.
Users notice this immediately and usually draw the wrong conclusion: “AI makes things up, we can’t use it.” The right conclusion would be: this model had no access to the information it needed for the task. The technical answer is connecting it to your own documents and data. Only then does the system work with your knowledge instead of a plausible guess.
3. There was no owner with time for it
AI projects are often looked after on the side — by IT, which is already stretched thin, or by an interested colleague from the department. Neither has the mandate to change processes. But that’s exactly the job: a tool that cuts a three-hour task down to twenty minutes changes the order of work steps, who’s responsible for what, and how approvals flow.
That’s why you need one person who is responsible for the rollout, has time set aside for it, and is allowed to make decisions. In companies with 30 to 250 employees, that’s often a team lead from the department where the first use case lives — not IT.
4. The legal questions came too late
A pilot is running, the results are convincing, and then someone asks: are we actually allowed to feed applicant data into this? What about the works council? Where is the data stored? This is where many projects stall for months — not because the answers are hard, but because nobody prepared them.
These questions can be answered, and they’re answered faster when you ask them at the start. Clarify the hosting location, the data processing agreement, how personal data is handled, and the involvement of employee representatives before the first test, and you won’t lose time later.
5. Nobody learned how to use it
Working with a language model is a skill. Anyone who just types “create a quote for me” gets a mediocre result and concludes the tool is useless. Anyone who has learned to provide context, check intermediate results, and work in steps gets usable results.
This skill doesn’t develop on its own. Two hours of hands-on training with the team’s real tasks achieve more than any generic course — and since February 2025 it’s also a legal requirement: Article 4 of the EU AI Act obliges companies to ensure adequate AI literacy among the people who operate such systems on their behalf.
What the successful rollouts have in common
Comparing projects that are still running after a year reveals four patterns.
They started small, but deep. Not five departments at once and superficially, but one department properly. A use case that’s fully thought through — including what happens to the output, who reviews it, and where it’s stored.
They connected the system to their existing data. Documents in SharePoint, folders on the file server, emails, the ERP system. The difference between “AI that answers in general terms” and “AI that knows our documents” is the difference between a toy and a tool.
They took permissions seriously. An assistant that accesses company data may only show an employee what that employee is already allowed to see. That sounds obvious, but it’s the point where many in-house builds fail.
They made the benefit visible. Not as a percentage in a slide deck, but concretely: creating a quote now takes one hour instead of three. Responding to a tender is done in a day instead of a week. Sentences like that convince the next department — slide decks don’t.
An approach that works
These observations point to a simple process that has proven itself in Mittelstand structures.
Weeks 1–2 — Selection. With three to five managers, gather tasks that cost a lot of time, occur frequently, and rely heavily on text and documents. Rate them by effort and impact. Pick exactly one.
Weeks 3–4 — Define the scope. Which data is needed? Where is it stored? Does it include personal data? Who needs to be involved? Which permissions apply? The result is one page, not a concept paper.
Weeks 5–8 — Build and test with real cases. Not with sample data, but with the actual cases from recent weeks. Only this shows whether the system can handle the edge cases that make up the majority in every company.
Weeks 9–12 — Team rollout and measurement. Everyone involved works with it, there’s a fixed point of contact for questions, and after four weeks you evaluate whether processing time has changed.
After these twelve weeks, you know whether the case holds up. If it does, the next one is easier, because the connections, permissions, and operating model are already in place. If it doesn’t, you’ve invested a quarter, not a year.
The question that comes first
Before you think about models, providers, and licenses, answer one question: which task in your company regularly costs time, relies on existing information, and frustrates the people involved? That task is your first use case. Everything else follows from it.
If you’d like an outside perspective: in an initial consultation, we’ll look at processes like this together and tell you honestly whether using AI is worthwhile here — or whether a simpler solution gets you there faster.
Three patterns that reveal a project at risk early
There are warning signs visible long before a project actually stalls. Anyone who knows them can course-correct while it’s still cheap to do so.
The project has no date on which a decision gets made. As long as nobody says when it will be evaluated and against what, it runs with neither brake nor gas pedal. It doesn’t end with a decision, but with waning interest. So set a fixed date from the start at which you decide to scale up, improve, or stop.
The enthusiasm sits with the people who don’t actually use it. If management and IT are convinced but the department stays politely silent, the foundation is missing. The benefit has to be felt where the work happens, or usage never takes off.
People talk about tools, not about processes. As soon as discussions revolve around vendors, models, and features without naming a concrete process, the order has slipped. Returning to the question “which task should get better?” is then the most effective course correction.
What you should measure — and what not to
The most common measurement mistake is the usage rate. It’s easy to track but says little and tempts you to force usage. Three other metrics are more meaningful.
- Process turnaround time. From intake to completion, measured before the start and three months after.
- Number of follow-up questions. A good indicator of whether results are actually usable.
- Voluntary continued use. Would the people involved miss it if you took it away again? This question takes two minutes to ask and is more honest than any statistic.
Frequently asked questions
We already have a failed project behind us. How do we start again?
Name openly what went wrong — it’s usually one of the five points above. Then deliberately choose a smaller, clearly defined case and set a decision date. A second attempt with a tight scope and a visible result rebuilds trust faster than a big relaunch.
Should IT or the department take the lead?
The department leads, IT supports. The reason is simple: the decisions to be made are functional ones — which process, which documents count, who reviews what. IT provides the connections, permissions, and operations.
How big can the first use case be?
Small enough to run through completely in twelve weeks, and big enough that its absence would be noticed. A process that occurs once or several times a day and noticeably costs time today usually fits that description.
What if the state of our documents is poor?
That’s the normal case, not a disqualifier. Don’t clean up everything — just the area the first case needs. The experience from that cleanup is also the best basis for deciding where to go next.
A worked example
What this process looks like in practice can be shown with a typical case. Take a technical services provider with around eighty employees that largely assembles its quotes from previous ones. The figures are placeholders — the pattern is transferable.
Weeks 1–2. In the session with three managers, five candidates come up: quote creation, invoice checking, service reports, applicant management, internal inquiries. They’re scored by frequency and how mechanical the work is. Applicant management is out immediately — a high-risk area, not a starting case. Quote creation wins: forty cases a month, a high text share, clear frustration in the team.
Weeks 3–4. The data sources are quickly identified: the quote archive on SharePoint, the price list in the ERP system, reference texts in a marketing folder. Personal data exposure is low. The works council is informed, since activity is logged. In parallel, the test list takes shape: twenty-five real inquiries from recent months, along with the quotes that came out of them.
Weeks 5–6. Reviewing the quote archive reveals the usual picture: three price lists with no clear validity date, folders with direct permissions for colleagues who left long ago, drafts sitting next to final versions. The cleanup takes two weeks — and everyone agrees it was overdue.
Weeks 7–9. The test with the twenty-five cases is disappointing at first: eight responses are unusable. The analysis shows that six of those trace back to missing or outdated documents, and only two to genuinely weak answers. After fixing the gaps, the score is twenty-two out of twenty-five.
Weeks 10–11. Four people in inside sales start using it. Two hours of hands-on training with real inquiries, eight templates for the most common cases, one named point of contact.
Weeks 12–13. Processing time per quote has dropped from around three hours to just under two — less than hoped for, but reliably measured. Something else stands out more: two tenders that would have been declined the previous year for lack of time were actually completed.
The decision is to scale up. The second case — service reports — starts in week 17 and goes live in eight weeks instead of thirteen, because the connections, permission logic, and operating model are already in place.
What makes the difference in this example
Three details in this process are what determine success, and all three seem unspectacular.
The case was chosen by frequency, not by enthusiasm. Forty cases a month build routine; twelve a year don’t.
The disappointing first test round didn’t lead to cancellation, but to root-cause analysis. Had they given up after eight failures, the verdict would have been “it doesn’t work” — even though the problem was in their own filing.
The baseline was measured beforehand. Without the measured three hours, everyone would have ended up with a different gut feeling, and the discussion would have been impossible to settle.
Want to know whether this pays off in your company? We’ll take a look at one concrete process with you and tell you honestly even if using AI isn’t worth it here.
Your secure AI platform for the Mittelstand. Secure. Intelligent. Integrated. Custom database integration, personally supported.
novendix GmbH · Industriestraße 6 · 91126 Schwabach
Locations: Schwabach · Weißenburg · Nuremberg
A company of the L&S Lange & Schermer Group
