Why an Internal AI Assistant Surfaces Documents You Were Never Meant to Find
An internal AI assistant surfaces documents you were never meant to find because it enforces the permissions you have, not the permissions you were meant to have. Retrieval removes the obscurity that used to protect a badly permissioned folder, and leaves the permission itself untouched. Nothing has been granted. Something has been found.
That distinction is the whole failure mode, and it is worth being precise about who is at fault, because the usual write-up gets it wrong. This is not one vendor being careless. It is the architecture of every assistant that mirrors your existing access controls, which is every assistant a serious buyer would consider. The evidence below is deliberately weighted towards vendors documenting the problem against their own products, because that is the only kind of source that cannot be dismissed as a competitor talking.
Every quotation on this page was read from the vendor page or the paper it is attributed to, on 6 September 2026. Nothing here comes from a competitor summary or a secondary article. Where something could not be verified, it is named in the last section rather than quietly used.
Why Does an Internal AI Assistant Surface Documents You Were Never Meant to Find?
Because two different things have always been conflated in a file system, and an assistant separates them. The first is permission: whether your account is technically allowed to open a document. The second is discoverability: whether you would ever have known the document was there. For twenty years, the second quietly compensated for the first. A budgeting spreadsheet on a site nobody linked to, with an access list nobody reviewed, was safe because it was invisible.
An assistant indexes it, reads it, and answers a question with it in about a second. The permission was already wrong. The assistant simply removed the thing that was hiding the mistake. Which is why every control that vendors ship for this problem is a discoverability control rather than a permission control, and why the vendors keep saying so out loud.
What Do Vendors Actually Guarantee About Permissions?
Read the promises closely and they are all the same promise: the assistant will not widen your access. None of them is a promise that your access was correct. That gap is where the incidents live.
| Vendor | What the documentation promises | What it does not promise |
|---|---|---|
| GleanSource: Glean docs, how Glean accesses information | “Glean respects the permissions set in your company’s connectors. If you have permission to view a document in Google Drive or a thread in a public Slack channel, it can appear in your results. If your coworker has different permissions, they might see different results.” | That the permission was the right one. Glean mirrors the connector’s answer at speed; it does not review it. Changes propagate: “If any permissions change in a connector, Glean reflects those changes quickly.” |
| Microsoft 365 CopilotSource: Microsoft Learn, Restricted SharePoint Search | Content is surfaced according to the permissions that already exist in SharePoint and Microsoft Graph, and Microsoft ships two separate controls to narrow what Copilot may discover while an organisation cleans those permissions up. | That either control is a security boundary. Microsoft: Restricted SharePoint Search “isn’t a security boundary and doesn’t change any permissions on SharePoint sites.” |
| Atlassian RovoSource: Atlassian Support, Rovo data privacy and usage guidelines | “We index the entire workspace of the third-party app you connect (for example, the entire Google Drive or the SharePoint workspace). You can refine the scope of the data indexed with some specific connectors (such as Google Drive and SharePoint) using a blocklist.” | Least privilege. Index everything, then subtract by blocklist, is the inverse of it, and the subtraction is available only on some connectors. |
| Guru, default access modelSource: Guru Help, connecting sources for AI answers | The model Guru labels Guru Groups: “You manually assign which Guru Groups can access the Source’s content. Guru connects as a single user and indexes only what that user can access. Access is controlled entirely within Guru, independent of Source app permissions.” | That the source system’s access list is consulted when someone asks a question. In this model it is not. Whatever the connecting account could see becomes available to whichever Guru Groups an admin assigns. |
| Guru, inherited access modelSource: Guru Help, connecting sources for AI answers | “Inherited permissions (this is supported for some sources, for more information see here). Guru automatically respects the native access controls from your Source app.” It “requires an admin from the Source system (like a Slack Admin or Google Admin) to connect the Source.” | That this is how Guru works everywhere. Guru’s own pricing page states it without any of those conditions: “Permissions are inherited from your existing systems and enforced in real time.” |
| DustSource: Dust docs, workspace governance, roles, groups and permissions | Access is re-declared inside Dust after ingestion, through spaces and groups. “Permissions are granted to groups rather than directly to individual people.” | Any mechanism that narrows access. “Permissions are additive. If a person belongs to several groups, they receive the combined permissions granted by those groups. A group grant adds access and does not remove access granted elsewhere.” |
The Guru pair is worth pausing on, because the two rows come from the same company on the same day. The help documentation describes inheritance as one of two models, supported for some sources, and dependent on a source-system administrator connecting it. The pricing page describes it as how the product works: “Permissions are inherited from your existing systems and enforced in real time.” Both sentences are quoted above and both are live. A buyer who reads only the second one will design a rollout around a guarantee the first one does not give.
Microsoft Writes the Failure Down Itself: Alex Wilber and the Budgeting Site
The single best primary source in this field is not an analyst report or a security vendor's white paper. It is a page of Microsoft administrator documentation, in which Microsoft walks a named fictional employee through exactly the thing buyers are afraid of.
Microsoft’s own worked example, quoted in full
“The site might be open to some users who aren’t allowed to see it, such as Alex.” That one sentence is the entire argument of this page, written by the vendor with the largest deployed base of internal AI assistants, in the document it hands to the administrators rolling one out. Note what it does and does not say. It does not say Copilot broke a permission. It says the permission was already open, to a person who was not allowed to see the site, and that the assistant is what turned that latent fact into an answer on a screen.
Notice also which detail carries the weight: “Most people don’t know about this site.” The control that was working was ignorance. That is a real control, it works for years, and it evaporates on the day an index finishes building.
Why Has Microsoft Shipped Two Controls for This and Called Both Temporary?
Since Copilot shipped, Microsoft has released two successive features whose stated purpose is to stop Copilot surfacing content people can technically reach but were not meant to find. It describes both of them as temporary. The first is being retired. The second requires an entitlement above the Copilot licence, and Microsoft warns that using it too much makes Copilot worse at its job.
| Control | What Microsoft says it is for | The limit Microsoft states | Status |
|---|---|---|---|
| Restricted SharePoint SearchSource: Microsoft Learn, Restricted SharePoint Search | “Use Restricted SharePoint Search as a temporary measure to prevent certain content on SharePoint sites from being shared too widely.” It is “a short-term solution that gives your organization’s administrators time to review and audit site and file permissions.” | “It’s not intended or scalable for long-term use.” It “isn’t a security boundary and doesn’t change any permissions on SharePoint sites.” It “limits to 100 sites, which isn’t sustainable as your organization scales Copilot and agentic operations.” | Retiring. “Starting July 31, 2026, new enablement is blocked.” |
| Restricted SharePoint Search, the leak it does not closeSource: Microsoft Learn, Restricted SharePoint Search | Nothing further. This row is what the same page says the control cannot do. | “Restricted SharePoint Search doesn’t guarantee that only sites on the allow list show up in search or in Copilot chat and agentic experiences. If a user recently accessed a site, or the site was shared with that user in Teams or Outlook, the site appears in the user’s results and responses, even if it’s not on the allow list.” | Also: “Neither Copilot nor Restricted SharePoint Search prevents users from accessing content they own or previously accessed.” |
| Restricted Content DiscoverySource: Microsoft Learn, Restricted Content Discovery | It “helps you limit discovery of content from specific SharePoint sites, including recently interacted files, in organization-wide search results and Microsoft Copilot responses while those reviews are taking place.” | “Restricted Content Discovery is designed as a temporary governance control that gives organizations time to review and right-size access while continuing their Copilot deployment.” And: “Excessive use can reduce the amount of content available to organization-wide search and Microsoft Copilot experiences, which can affect the completeness and relevance of search results and AI-generated responses.” | Licence-gated on top of Copilot. “Customers who are licensed for Copilot and have SharePoint Advanced Management available to them can configure Restricted Content Discovery.” |
Two clauses in that table deserve to be read together, because they describe a genuine trade rather than a bug. Restricted Content Discovery works by removing content from what Copilot may discover. So Microsoft has to warn you that “excessive use can reduce the amount of content available to organization-wide search and Microsoft Copilot experiences, which can affect the completeness and relevance of search results and AI-generated responses.” Hiding content from the assistant makes the assistant less useful. There is no setting that gives you both. The only thing that gives you both is correct permissions, which is a governance project, not a product feature.
And the older control does not even fully hide things. Microsoft states that Restricted SharePoint Search “doesn’t guarantee that only sites on the allow list show up in search or in Copilot chat and agentic experiences,” because a site the user recently accessed, or one shared with them in Teams or Outlook, still appears. If your plan was to switch a feature on and stop thinking about SharePoint, the vendor has already told you that plan does not work.
Where Does the Same Failure Show Up in Other Vendors’ Defaults?
Microsoft is the most quotable because it documents the failure most candidly. It is not the only one whose defaults create it. Three other vendors publish defaults that produce the same outcome by different routes, and all three sentences are from their own documentation.
Atlassian Rovo indexes everything first and narrows afterwards. Atlassian writes: “We index the entire workspace of the third-party app you connect (for example, the entire Google Drive or the SharePoint workspace). You can refine the scope of the data indexed with some specific connectors (such as Google Drive and SharePoint) using a blocklist.” Index-everything-then-blocklist is the inverse of least privilege. It means the default state of a new connector is maximum exposure, and every reduction is something a human has to remember to write down. It also means the blocklist is only as good as somebody’s knowledge of what is in the drive, which is precisely the knowledge that was missing in Microsoft’s budgeting-site example.
Guru’s default model does not consult the source system at all. In the Guru Groups model, “Guru connects as a single user and indexes only what that user can access. Access is controlled entirely within Guru, independent of Source app permissions.” Read that twice. The access decision at question time is made by Guru group membership, not by the access list on the original document. Whatever the connecting account could reach is now available to whichever Guru Groups an administrator ticked. Guru does offer an inherited model that respects native controls, but its own documentation qualifies it as “supported for some sources” and requiring a source-system administrator to set up.
Dust permissions only ever widen. Dust re-declares access inside its own product after ingestion, and its governance documentation states the arithmetic: “Permissions are additive. If a person belongs to several groups, they receive the combined permissions granted by those groups. A group grant adds access and does not remove access granted elsewhere.” A mis-scoped group therefore over-grants silently, and nothing anywhere in the model narrows it back. That is a normal and defensible design. It is also one that punishes a messy group structure, and most companies have a messy group structure.
Do the Models Themselves Know What Is Confidential?
A reasonable hope is that the model catches what the permissions missed: that even if a salary spreadsheet is technically reachable, the assistant recognises what it is holding and declines. The best-sampled published test of that hope says it does not.
CRMArena-Pro is a benchmark of nineteen expert-validated tasks across sales, service and configure-price-quote processes, run in a synthetic customer-relationship-management organisation, covering both business-to-business and business-to-consumer scenarios, with multi-turn interactions driven by different personas. Alongside the task work, the authors created eighty queries per organisation designed specifically to probe whether an agent recognises a request for information it should not disclose. The result, verbatim from the abstract:
CRMArena-Pro, on confidentiality
Disclosure: this is vendor-affiliated research. The paper’s authors are at Salesforce AI Research and the sandbox environment is a Salesforce org. The findings are unflattering to the category the vendor sells into, which cuts against the usual direction of vendor bias, but the affiliation belongs on the page rather than in a footnote.
The second half of that sentence is the part that matters operationally, and it is the part everyone drops. Prompting for confidentiality helps, and it costs task performance. So the choice an administrator actually faces is not “careful or careless”. It is a dial between an assistant that is more useful and one that is more discreet, and the vendor documentation quoted earlier describes the same dial from the other end when it warns that hiding content degrades answers. Two independent sources, one benchmark and one product manual, describing the same trade.
Which is why the durable fix is upstream of both. If the assistant never had a path to the document, neither dial matters.
Has This Produced an Actual Vulnerability?
One is on the public record, and it needs to be described precisely, because it is routinely overstated.
Microsoft assigned itself a CVSS 9.3 information-disclosure CVE against Microsoft 365 Copilot in June 2025 for an AI command-injection flaw requiring no user interaction, and states it was fully mitigated and not exploited.
The details, from the two primary records. The National Vulnerability Database entry for CVE-2025-32711, published 11 June 2025, describes it as “Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network.” Microsoft scored it 9.3 with the vector CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N; NIST scored the same issue 7.5. The two components worth reading are UI:N, meaning no user interaction was required, and PR:N, meaning no privileges were required.
What the CVE does and does not establish
Publicly Disclosed:No;Exploited:No;Latest Software Release:Exploitation Less Likely.So this is a disclosed and patched vulnerability class, not a confirmed real-world breach. No customer data is known to have been disclosed, and nothing on this page should be read as saying otherwise. What it establishes is narrower and still useful: that a hosted internal AI assistant is a network-reachable information-disclosure surface in its own right, serious enough for its vendor to publish a 9.3 against itself, and that the vendor chose to disclose it when it did not strictly have to.
The reason to include it at all is that it changes the shape of the threat model. Permission leakage is usually imagined as an employee typing a curious question. A command-injection path means content that arrives in the system can steer the assistant, so the attacker does not need an account and the employee does not need to be curious. That is a different class of problem from over-shared SharePoint, and it lands on the same infrastructure.
Why This Is Structural, Not One Vendor Being Careless
Everything above comes from five vendors and one benchmark, and it converges on a single design fact rather than on a list of mistakes.
Any assistant that answers questions about internal knowledge has to decide what a given person may see. There are exactly two ways to make that decision. It can mirror the access controls that already exist in the source systems, in which case it inherits every error in them. Or it can maintain its own access model, in which case somebody has to re-declare, by hand, who may see what, for every system, forever. Glean and Guru’s inherited mode take the first route. Guru’s default mode and Dust take the second. Microsoft takes the first and has now shipped two temporary discoverability controls on top of it. There is no third option, and neither of the two can be correct if the underlying permissions are not.
The sentence to take into a vendor call
This is also why framing it as vendor carelessness leads to a bad purchase. The buyer who believes vendor A has a permissions bug goes looking for vendor B without one, finds a marketing page that promises real-time inherited permissions, and deploys onto the same broken access control lists with less documentation about what is happening. The buyer who understands it as structural does the access review first and then treats the vendor choice as what it is: a question about which of the two models fits how their systems are actually administered.
What to Fix Before You Switch an Assistant On
Five things, in this order. None of them is an AI project.
1. List the repositories that hold material an internal leak would actually hurt. Payroll, board papers, legal, unreleased pricing, personal data. In most companies this is ten to thirty locations, not the whole estate, and the review is a week of work rather than a programme.
2. Review the access lists on those, and only those, before the first connector is authorised. Microsoft’s worked example is a site whose owner “hasn’t set up proper permissions and hasn’t followed correct data governance process”. You are looking for the sites where the honest answer to “who can open this?” is “nobody has checked”.
3. Scope the first connector to systems, not to the workspace. Where a vendor indexes an entire connected workspace by default, that default is the decision unless somebody overrides it. Atlassian’s documentation tells you this outright, and tells you the blocklist is available on some connectors rather than all of them.
4. Find out which identity the connector uses. If the product connects “as a single user”, that account’s reach is the ceiling of everything the assistant can ever surface, and it is usually an administrator’s account because that was the easiest way to make the integration work. Connect with the narrowest account that makes the pilot useful.
5. Keep the pilot group small enough that a wrong answer is a conversation. The failure you want to discover is somebody saying “why can I see this?”, and you want to hear it from twelve people who expected to be testing, not from four hundred who did not.
The controls the vendors ship are worth using, on the terms the vendors state. Restricted Content Discovery buys time during a review, which is what Microsoft says it is for. It is not a permissions model, and it costs answer quality when it is overused. Treat any such feature as a tourniquet with a stated shelf life, because in both Microsoft cases that shelf life is written on the box.
One structural decision makes the review smaller rather than larger, and it is worth taking before the tooling conversation. An assistant that is connected to a named set of systems, with a named account per connection, is a scope you can write down on one page and re-check every quarter. An assistant that is pointed at everything is not. That is how we build an AI teammate for a European company: per-integration, so that the answer to “what can it reach?” is a list rather than a shrug. It is a smaller starting surface, not a certification, and this page has already been clear about what we do and do not hold.
The same instinct applies to the assistant you buy. What data leaves your systems, where it is processed, and whether the vendor trains on it are separate questions from permissions, and they are answered in different documents: our tier-by-tier reference on whether your AI vendor trains on your data covers the first two, and whether an internal assistant’s EU data residency actually covers the AI covers where the processing happens. If the assistant can also change records rather than only read them, the permission question gets larger rather than smaller, which is the subject of the difference between a RAG assistant and an agentic one. And if the reason you are buying one is the amount of time people spend hunting for documents, the numbers usually quoted for that are worth checking before they go in a business case: we traced them in the hours-wasted-searching statistic and where it actually comes from.
What This Page Does Not Claim
Last verified 6 September 2026
- No claim that customer data leaked. CVE-2025-32711 is recorded by Microsoft as fully mitigated, not publicly disclosed and not exploited. A page that turns it into a breach story is describing something the sources do not support.
- No use of the security researchers’ own write-up of that CVE. It is widely quoted elsewhere. The page hosting it returned HTTP 403 to us, so we could not read it, so nothing from it appears here. The CVE section rests entirely on Microsoft’s advisory and the NVD record.
- No claim that any of these vendors is insecure. Glean publishes SOC 2 Type II, ISO 27001, ISO 42001, HIPAA, TX-RAMP Level 2 and GDPR on its security page. Guru publishes an annual SOC 2 Type II audit. Ainora holds no SOC 2 and no ISO certification of any kind, so on that specific measure these vendors document more than we do, and a comparison that pretended otherwise would be worth nothing to you.
- No population claims. There is no figure on this page for how often oversharing happens in the field, because no study we could find measures it. What exists is vendor documentation of the mechanism and one benchmark measuring model behaviour, and both are labelled as what they are.
- No numbers from the Dust founder’s widely circulated remark about connecting multiple systems. It is often cited as an admission that cross-system permissions are unsolved. Read in full, in its own thread, it is a comment about query languages. It does not belong in this argument and is not used here.
Frequently Asked Questions
It can show you anything your account can technically open, which is not the same set as the documents you were meant to have. Glean states the rule plainly: “Glean respects the permissions set in your company’s connectors. If you have permission to view a document in Google Drive or a thread in a public Slack channel, it can appear in your results.” That is a promise about enforcement, not about correctness. If a site was left open because nobody expected anyone to find it, the assistant finds it, and the permission that made that possible was already wrong before the assistant arrived.Source: Glean docs, how Glean accesses information
Yes, in its own administrator documentation, with a named worked example. Microsoft describes a marketing specialist at a fictional company who “can see not only his own personal contents, like his OneDrive files, chats, emails, contents that he owns or visited, but also content from some sites that haven’t undergone access permission review or Access Control Lists (ACL) hygiene.” It then gives the case: a budgeting site whose owner “hasn’t set up proper permissions,” where “the site might be open to some users who aren’t allowed to see it, such as Alex. When Alex asks Copilot for some budgeting information, Copilot gets information from the budgeting site.” This is a vendor writing down the exact failure buyers are afraid of, in a page aimed at the administrators who have to prevent it.Source: Microsoft Learn, Restricted SharePoint Search
No, and Microsoft says so in the same sentence it introduces the feature: Restricted SharePoint Search “isn’t a security boundary and doesn’t change any permissions on SharePoint sites.” It also does not fully contain discovery. Microsoft notes that the feature “doesn’t guarantee that only sites on the allow list show up in search or in Copilot chat and agentic experiences,” because recently accessed sites and sites shared with the user in Teams or Outlook still appear. It caps at 100 sites, which Microsoft itself calls unsustainable as an organisation scales Copilot, and new enablement is blocked from 31 July 2026.Source: Microsoft Learn, Restricted SharePoint Search
Restricted Content Discovery, and Microsoft describes it in the same temporary terms: “a temporary governance control that gives organizations time to review and right-size access while continuing their Copilot deployment.” Two things are worth knowing before you plan around it. It needs an entitlement above Copilot itself, because only “customers who are licensed for Copilot and have SharePoint Advanced Management available to them can configure Restricted Content Discovery.” And Microsoft warns that leaning on it degrades the product you bought: “Excessive use can reduce the amount of content available to organization-wide search and Microsoft Copilot experiences, which can affect the completeness and relevance of search results and AI-generated responses.” Neither control is a substitute for fixing the permissions.Source: Microsoft Learn, Restricted Content Discovery
The best-sampled published evidence says no. CRMArena-Pro, a benchmark of nineteen expert-validated tasks across sales, service and configure-price-quote processes inside a synthetic CRM organisation, reports that “agents exhibit near-zero inherent confidentiality awareness; though targeted prompting can improve this, it often compromises task performance.” The same paper reports leading agents achieving “only around 58% single-turn success” and roughly 35% in multi-turn settings. One disclosure belongs with the finding: the paper comes from Salesforce AI Research and the environment is a Salesforce org, so it is vendor-affiliated work, although the results are unflattering to the agents the vendor sells.Source: arXiv:2505.18878, CRMArena-Pro
One is documented, patched and vendor-acknowledged. Microsoft assigned itself a CVSS 9.3 information-disclosure CVE against Microsoft 365 Copilot in June 2025 for an AI command-injection flaw requiring no user interaction, and states it was fully mitigated and not exploited. The NVD record for CVE-2025-32711 describes it as “Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network,” with Microsoft scoring it 9.3 and NIST scoring it 7.5. Microsoft’s own advisory says “this vulnerability has already been fully mitigated by Microsoft. There is no action for users of this service to take. The purpose of this CVE is to provide further transparency,” and its threat metadata records “Publicly Disclosed:No;Exploited:No.” Treat it as a disclosed and patched vulnerability class, not as evidence that customer data leaked. It did not, on the published record.Source: NVD API record, CVE-2025-32711
That is the wrong trade, because permissions are never finished and the discovery problem exists with or without the assistant. The order that works is: run an access review on the ten or twenty repositories that actually hold sensitive material, scope the first connector to those systems rather than to the whole workspace, and make the pilot group small enough that a wrong answer is a conversation rather than an incident. Every vendor control described on this page is designed to buy time for exactly that review. None of them is designed to replace it, and two of them say so in their own documentation.Source: Microsoft Learn, Restricted Content Discovery
Founder & CEO, AInora
Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.
View all articlesReady to try AI for your business?
Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.
Related Articles
RAG or Agentic: What Actually Separates the Two
A three-part test, from vendor admin documentation, for what a product is really selling you.
Does Your AI Vendor Train on Your Data?
Training defaults, EU residency, retention and admin controls, tier by tier, sourced to the vendor docs.
Does an Internal AI Assistant’s EU Residency Cover the AI?
Nine vendors, one question, and what each of them actually commits to in writing.
The Hours-Wasted-Searching Statistic, Traced to Origin
Where the number in every internal-search business case actually comes from.