AI search over company documents, and where it still invents
The best measured result on retrieval over a document set is that the leading professional tools still got it wrong in between one in six and one in three answers. That is the number to design around, not the demo.
Every company we work with has the same buried asset: a decade of policies, contracts, SOPs, quotes and email threads that nobody can find anything in. The pitch for AI search over that pile is irresistible, and it is usually demonstrated with a question somebody already knew the answer to.
It does work. It works considerably better than keyword search over the same folder. But the way it fails is specific and quiet, and if you scope a project without knowing the failure mode you will ship something that is confidently wrong at a rate nobody is measuring.
The one study that measured it properly
Vendors do not publish accuracy rates for retrieval over your documents, so the closest usable evidence comes from a domain where somebody checked: legal research. Stanford's RegLab and Institute for Human-Centered AI benchmarked the leading commercial legal research tools, all of them retrieval systems over a curated document set, and published the result as AI on Trial: Legal Models Hallucinate in 1 out of 6 or More Benchmarking Queries in May 2024.
They ran over 200 open-ended queries from a pre-registered dataset. Two of the three tools produced incorrect information more than 17 percent of the time, and the third hallucinated more than 34 percent of the time.
| System type | Hallucination rate found | What it was retrieving over |
|---|---|---|
| General purpose chatbot, no retrieval | 58% to 82% | Model weights only |
| Two of three commercial legal tools | More than 17% | A curated legal corpus |
| The third commercial legal tool | More than 34% | A curated legal corpus |
The direction of that table is the good news. Retrieval cuts the error rate by a large multiple against a model answering from memory, which is exactly why you ground a system on your own documents. The level is the bad news. These are well-funded products built on a clean, curated, professionally maintained corpus, and they were still wrong often enough that a human has to check. Your SharePoint is not a curated corpus.
The authors are explicit that retrieval does not solve this. Their conclusion is that even retrieval-augmented systems are not hallucination-free, written against vendor claims that retrieval reduces hallucination to nearly zero. We have kept the same posture in what to trust a support chatbot with and in where AI document processing still needs a person, for the same reason.
Two different failures, and only one of them is visible
The study separates the error into two kinds, and that distinction is the single most useful thing to take from it when scoping a build.
- Incorrect. The answer misstates the thing or contains a plain factual error. Somebody who knows the subject catches this on sight.
- Misgrounded. The answer is correct, and the citation under it does not support the claim. The sentence is right and the reference is wrong.
Misgrounded is the dangerous category in a company deployment, because every mitigation people reach for makes it worse rather than better. The standard answer to hallucination is show the source, and showing a source is precisely what makes a misgrounded answer look verified. The citation is the trust signal, so an answer with a confident wrong citation is more likely to be acted on than one with none.
In practice this looks like an answer about leave entitlement that is right for one staff category and cites the policy for a different one, or a quoted payment term that is correct for one client and sourced from another client's contract. Both read as impeccable. Both are the kind of thing somebody forwards.
The problem underneath is permissions, not the model
The second failure mode has nothing to do with accuracy. Retrieval makes everything in the index reachable by question, and most companies have no idea what is in their index because sharing accumulated over a decade of people being helpful.
Microsoft's own documentation is unusually candid about this for Copilot. It states that Copilot operates within existing permissions and access controls, and that overshared or poorly governed content can affect results and increase risk. Note what that means: the system is behaving correctly and the outcome is still a leak, because the permission was wrong before anyone asked a question.
Three details from Microsoft's security guidance for Copilot and the SharePoint governance pages are worth knowing before you index anything:
- SharePoint's sharing settings default to the most permissive option, so the baseline is open unless somebody tightened it.
- The Everyone Except External Users report exists to identify the top 100 sites where content was shared with the entire organisation in the past 28 days, which tells you how routine that is.
- Restricted Content Discovery keeps a site out of assistant results without changing who can open it, and Microsoft describes it as a temporary governance control rather than a fix. Restricted SharePoint Search is labelled explicitly as not a security boundary and not intended as a long-term solution.
A vendor telling you its own search feature is not a security boundary is the most useful sentence in that documentation. It means the permissions work is yours and it has to happen before the index exists, not after somebody asks the wrong question successfully.
What it is genuinely good at
None of the above is an argument against building it. It is an argument against scoping it as an oracle. The tasks where retrieval over company documents pays for itself quickly all share one property: the answer is a pointer rather than a conclusion.
- Finding which document covers a thing, when nobody remembers what it was called or who wrote it. This is the highest-value and least risky use and it is usually enough.
- Summarising a long document somebody is about to read anyway, so the summary is checked by construction.
- Drafting a first answer for a person who then edits it, which is the same shape as what AI meeting notes are actually good at.
- Answering a question where being wrong is cheap and obvious, like where is the template, not like what is our notice period.
The pattern is that a human stays in the loop where the cost of error is real, and the system saves the twenty minutes of looking rather than the two minutes of deciding.
How to scope it so the failure mode is survivable
Five decisions, in this order. Each one is cheaper before the build than after.
- Fix the permissions first, on a smaller set. Index one well-governed collection rather than everything. Scope is the only permissions control that always works.
- Pick the corpus deliberately and date it. Three superseded versions of a policy in the index produce a confidently wrong answer from a genuinely real document. Retiring old versions does more for accuracy than any model change.
- Make the citation open the document at the passage. If checking takes one click, misgrounded answers get caught. If it takes three, they do not.
- Decide which questions it refuses. A deliberate refusal list for anything legal, contractual, HR-sensitive or financial, routed to a person, is worth more than any accuracy improvement.
- Log the questions and read them weekly. The questions people actually ask are never the ones in the requirements document, and they tell you what to fix next.
The second item is where most of the real gain sits, and it is a records problem rather than an AI problem. The same discipline applied to sales records is set out well in BDG Labs' piece on keeping CRM records clean enough to trust, and the argument transfers exactly: retrieval quality is a function of what you kept, not of what you asked. The structural version of that cleanup is migrating off spreadsheets without losing the history, and the broader case for when to build at all is in internal tools for a small business.
Honest limits
The numbers above are from legal research tools on a legal corpus, not from an internal document search on your drive. We are using them because they are the only pre-registered, independently run measurement we know of on this exact shape of system, and because the direction is informative. The specific percentages should not be quoted as applying to your deployment.
The study was also contested on methodology when it came out, and the authors added analysis in response. We still treat it as the best available evidence rather than the final word, and both tooling and models have moved since May 2024. The honest reading is that the error rate is probably lower now and is certainly not zero.
And we have no figure for how much time this saves. We have seen it save a lot of looking, and we have also seen a deployment abandoned because nobody trusted it after two confidently wrong answers in the first week. Which of those you get is decided by the five scoping decisions above rather than by the model you pick.
Describe it. We build it.
Seven or twelve days, pay on delivery, a year of maintenance included. Bring the problem, not a spec.
Book a meetingRead next
AI assisted software development speed, measured properly
AI assisted software development speed, from the only randomised trials on it: what they found, why the numbers moved, and what it means for a real build.
ReadA restaurant order management system, scoped around the channels
A restaurant order management system earns its keep by consolidating channels, not by taking orders. What it must own, and what the tax receipt forces.
Read