When a Mittelstand company is asked where its knowledge actually lives, the answer is astonishingly often: in Microsoft 365. Proposals and project documents in SharePoint, personal work in progress in OneDrive, the coordination around it in Teams, and the actual history of decisions in Outlook. That’s exactly why connecting these systems is usually the first sensible step when AI is meant to be more than a drafting aid.

The step is doable, but it comes with prerequisites that rarely appear in a proposal. This article describes them.

What “connecting” actually means

Connecting doesn’t mean your files get copied somewhere. It means an assistant can search your holdings with proper authorization when asked a question, read the relevant passages, and formulate an answer with a source citation from them.

Technically, this happens through the interfaces Microsoft provides for this purpose. A service authenticates with its own identity, receives precisely defined permissions, and operates — this is the decisive part — on behalf of whichever user is asking. The result: whatever an employee isn’t allowed to open in the browser, they also can’t find through the assistant.

Two approaches are worth distinguishing. With live querying, the system searches directly on every question. This is always current and permission-compliant, but slower and less precise with large corpora. With a dedicated index, content is prepared and made searchable in advance. This delivers better results, but requires that permissions be carried along and changes be continuously synced. In practice, hybrid approaches are common — and how permissions are represented in the index is the single most important question to ask any vendor.

The groundwork that makes the difference

Cleaning up permissions

This is the uncomfortable part, and also the most important one. In environments that have grown organically, the same patterns show up almost every time: libraries accidentally shared with “Everyone.” Direct assignments to individual people instead of groups. Former employees still sitting in distribution lists. Teams areas for closed-out projects that nobody tidies up anymore.

As long as no assistant exists, this goes unnoticed, because nobody specifically searches for it. An AI assistant makes these mistakes visible — not because it’s insecure, but because it reliably finds what a human would never have found. This, incidentally, is a frequently underestimated side benefit: the preparation forces a permissions clean-up that was overdue anyway.

Sorting file storage — but only where it counts

Nobody has to clean up twenty years of file storage before the first use case can start. The opposite is the sensible approach: define a clearly bounded, well-maintained area and connect only that. For sales, that’s typically the approved product documents, the current price list, and a collection of good reference material.

What matters is that there’s a single valid version. If three price lists exist side by side and none is marked as current, even the best system will give the wrong answer — and no one notices, because it sounds plausible.

Deliberately excluding sensitive areas

Personnel files, salary data, documents related to ongoing legal disputes, company acquisitions: such areas don’t belong in scope at the start, even if the permissions are technically in order. They can be added later, once the permissions model has proven itself.

Clarifying the contractual basis

When content from Microsoft 365 is passed to a language model, that content leaves your Microsoft environment during this process. What needs to be clarified, therefore: where does the model run? What contract underlies it? Is content stored, and if so, for how long? Is it used for training? For German companies, hosting in Germany or at least in the EU with a solid data processing agreement is the standard that data protection officers expect.

Outlook: the most sensitive and most valuable holding

Email mailboxes are the most interesting material substantively, because they contain agreements, commitments, and decisions that don’t show up in any document. At the same time, they’re the most sensitive area of all.

Three points need to be clarified beforehand. First, co-determination: as soon as systems are introduced that are capable of monitoring behavior or performance, the works council has to be involved. Even if no monitoring is intended, the mere capability can be enough. Second, private use: if it’s permitted or tolerated, analysis becomes considerably harder. Third, scope: access to one’s own mailbox is something entirely different from a search across all mailboxes. The first case is usually unproblematic and already very useful; the second requires careful justification and involvement.

A pragmatic starting point is to initially include only the requester’s own mailbox. This lets you demonstrate the value without having to answer the hardest questions right away.

Where projects fail in practice

Wrong expectations about completeness. “The assistant should know everything” leads to a system that’s mediocre at everything. An assistant for product questions that works reliably gets used daily; a jack-of-all-trades with inconsistent quality gets ignored again after three weeks.

Permissions as an afterthought. If everyone involved has administrator rights during testing, everything works beautifully — until the first regular user discovers they can’t find anything, or worse: finds too much.

Underestimated variety of formats. Scanned delivery notes, drawings, multi-column data sheets, presentations with text embedded in graphics. Test with exactly these documents, not with the clean Word file.

No business owner. Someone has to be responsible for keeping the connected content current. Without this role, the corpus ages, answers get worse, and trust erodes faster than it was built.

A realistic timeline

Phase 1 — Stocktaking. Which areas exist, how are they permissioned, which are well-maintained, which are confidential? The result is a list with a traffic-light rating.

Phase 2 — Selection and cleanup. An area is selected, permissions are corrected, outdated documents are archived or flagged.

Phase 3 — Connection and testing. The area is connected and tested with twenty to thirty real questions whose correct answer is known. Testing is done with the rights of a regular user, not as an administrator.

Phase 4 — Team rollout. One department works with it, a named person collects feedback, obvious gaps are addressed.

Phase 5 — Expansion. The next area is added. This step is noticeably faster, because the connection, permissions logic, and operating model are already in place.

The payoff, once it clicks

The most noticeable effect isn’t the minute saved on an individual search. It lies in the fact that knowledge becomes accessible again that was effectively lost: the reasoning behind a special discount from last year, the solution to a problem a colleague already ran into two years ago, the wording that won a difficult tender.

This knowledge was there the whole time. It just wasn’t findable. This is exactly the point where a well-connected AI shows its value — and the reason we almost always start projects with a sober stocktaking of your Microsoft 365 environment before talking about models.

What the connection technically requires

To make planning realistic, here are the points your IT should clarify in advance. They’re rarely a lot of work, but they hold up projects when they only surface once things are already running.

  • Sign-in: Should users sign in through your existing directory? That’s the clean approach, because access is granted and revoked centrally.
  • Administrator consent: Access to company data has to be approved by an administrator — with precisely defined permissions, not blanket approval.
  • Scope: Which libraries, Teams areas, and mailboxes are included? An explicit allow-list is better than excluding individual areas.
  • Propagating changes: How quickly does a revoked permission or a deleted document take effect?
  • Logging: Where do access logs get recorded, and who’s allowed to view them?

A realistic timeframe

For a first, clearly bounded area, the technical connection is usually done within a few days. What actually takes time is the groundwork: reviewing and correcting permissions, flagging outdated documents, establishing a single valid version. Plan for two to six weeks here, depending on how organically your environment has grown.

This split surprises many people and is the most important expectation to set in the project: the technology is rarely the bottleneck.

Frequently asked questions

We still have data on a classic file server. Does this work too?

Yes, that’s common. What needs to be clarified is how permissions are carried over from the file system and how changes are propagated. Often the file server is even the better starting point, because the holdings there are more clearly bounded than in Teams structures that have grown organically.

What about Teams chats?

Technically possible, often valuable in substance, sensitive from a co-determination standpoint — chat history is particularly well suited to reflecting behavior. We recommend not starting there.

Do employees see documents through the assistant that they otherwise wouldn’t find?

Only ones they already have access to — provided the permissions check works correctly. This is exactly why this question is the most important one to ask any vendor, and exactly why cleaning up permissions isn’t busywork, but the actual safeguard.

What happens with very large corpora?

Result quality drops when too much disorganized material is included. A narrow, well-maintained scope delivers better answers than a comprehensive one. Expand gradually and test with the same test list after each expansion.

Permissions in Microsoft 365, concretely

When talking with vendors, it’s worth getting precise here, because this is where the biggest differences lie. Two models are worth distinguishing.

Access on behalf of the signed-in user. The service acts with the requester’s identity. It can find exactly what that person could open themselves — nothing more. Permission changes take effect immediately, because every request is checked freshly. This is the clean approach, because no second version of the truth about access rights exists.

Access with a service identity plus its own index. The service reads with broad permissions and builds its own search index, in which it stores, per section, who’s allowed to see it. This is faster and delivers better results, but requires that changes be propagated. The decisive questions here are: how quickly does a revoked permission take effect, how quickly does a deleted document disappear, and what happens to the index when an employee leaves the company?

Both models are defensible. Only the third is unacceptable: everything in one shared index, access control handled through instructions to the model.

Also ask to be shown exactly which permissions the service requests before an administrator approves it. A blanket approval covering all files and mailboxes should have to be justified — and for a first, bounded use case, it usually isn’t.

What’s different with Teams, Lists, and Planner

SharePoint libraries and OneDrive are the easy part. For the other building blocks of Microsoft 365, there are particularities worth knowing beforehand.

Teams files live technically in SharePoint, so they’re unproblematic. Teams chats and channel conversations are something else: often valuable in substance, sensitive from a co-determination standpoint because they directly reflect behavior. For a first step, they don’t belong in scope.

SharePoint lists are structured data, not documents. They’re poorly suited to plain text search and well suited to targeted queries — provided the system can treat them as a data source rather than as running text.

Planner and tasks rarely contain knowledge worth searching for, but often contain personal assignments. Usually better left out.

OneNote is a special case: often a goldmine content-wise, technically difficult because the structure is unstructured and content can be tucked away in images and handwritten notes. Test this specifically if it plays a role for you.

Handling retention policies

A point that surfaces only late: many companies have configured retention and deletion policies in Microsoft 365. If an AI system builds its own index, that index has to follow those policies — a document that gets deleted once a retention period expires must not keep living on through the index.

This isn’t a theoretical nicety. It’s exactly the case where a deletion concept fails silently and comes up during an audit. Ask your vendor how deletions and retention policies are represented in the index, and get the answer in writing.

Want to know whether this pays off in your company? We’ll take a look at one concrete process with you and tell you honestly even if using AI isn’t worth it here.

Your secure AI platform for the Mittelstand. Secure. Intelligent. Integrated. Custom database integration, personally supported.

novendix GmbH · Industriestraße 6 · 91126 Schwabach
Locations: Schwabach · Weißenburg · Nuremberg
A company of the L&S Lange & Schermer Group

© 2026 novendix GmbH — All rights reserved.A New Era of Thinking · Built for the German Mittelstand 🇩🇪