Enterprise RAG: Why Permission-Aware Retrieval Is the Hard Part
Enterprise RAG can only be rolled out to a whole company if every answer is assembled from documents the person asking was already allowed to read. That means enforcing permissions inside the retrieval query, not on the model's output; carrying each source's access list onto every chunk; and failing closed whenever a permission is missing or unclear.
Key takeaways
- Permission-aware retrieval enforces access inside the retrieval query, not on the model's output.
- Every indexed chunk keeps the access-control list of the source it came from.
- A revoked grant stops answering on the next query; deletions reconcile on sync with a full-sweep backstop.
- A source with no permission rows is closed, not public — so mistakes fail toward safety.
- Good retrieval is still hybrid: vector search plus BM25, fused and optionally reranked.
An enterprise RAG system has one job the demo never tests. When a person asks a question, the answer must be assembled only from documents that person was already allowed to read — and from nothing else.
Get the retrieval right and the permissions wrong, and you have built a very fast way to leak. The model will happily summarize a document the reader was never cleared to open, cite it, and sound authoritative doing it. Retrieval-augmented generation makes that failure quieter, not louder: the source is folded into an answer, so nobody sees the document that should not have been reachable.
This is the part of enterprise AI search that separates a tool you can put in front of a whole company from a tool you can put in front of one team.
Why isn't filtering after the model enough?
There are two places to enforce a permission, and only one of them is safe.
The unsafe pattern retrieves the most relevant passages first and filters afterwards — sometimes after the model has already read them. Anything the retriever surfaced has already left the boundary; a filter on the way out is a promise, not a wall. The leak modes are well known: a passage that shapes the answer without appearing in the citations, a summary that survives after its source is filtered, a cache that outlives the grant.
The safe pattern makes the permission part of the query. In SphereIQ, the access check is part of the retrieval query itself, so a passage the person asking cannot read is never a candidate — not reranked, not read by the model, not in the context window. The same rule runs inside every retrieval path: chat, agent flows and the MCP server. There is no second path where it quietly does not hold.
A permission check on the model's output is a promise. A permission check inside the retrieval query is a wall. Only the second one survives a determined question.
Why must the permission travel with the document?
The signature mistake in this field is to treat permissions as a setting on the knowledge base. They are a property of the source.
Every chunk carries the access-control list it inherited from where it came from. A permission names a principal — a user, a team, a group synced from your identity provider, or explicit public access — and, optionally, a folder path. A grant with no path opens the whole source; a grant with a path opens only that part of it. Paths match on whole segments, so a grant on /Eng never accidentally opens /Engineering.
This is where the unglamorous truth of the whole category lives. A connector's real job is not to copy documents into an index. It is to carry each document's permissions across intact, and to keep the identity graph — which person belongs to which group — current enough that the query-time filter means something. The retrieval is a component. The permission ingestion is the project. It is why the Company Brain treats permissions as data that moves with the content, and why Connect is built around them.
A source with no permission rows is closed — readable by nobody but an administrator, not by everybody. It reads as a small default, and it is the whole game: the moment "unconfigured" means "public", every mistake fails in the direction of exposure.
How do you keep an index's permissions from going stale?
An index is a copy. If a knowledge base holds its own snapshot of your documents and their permissions, that snapshot can drift from the source between syncs. The useful question is not "is it ever stale" but "how far can it drift, and which way does it fail". SphereIQ answers it in three parts:
- Permission changes apply at query time. Removing a grant, a team membership or a whole user — including through SCIM deprovisioning — takes effect on the next question rather than the next crawl.
- Content changes reconcile incrementally. A document is fingerprinted; an unchanged one is skipped, a changed one has its old chunks replaced, and a deleted one is pruned. For sources that cannot report deletions, a periodic full sweep makes sure a removed document stops answering.
- Permissions are rewritten atomically. A source being re-synced is never briefly visible to everyone while its permissions are rebuilt.
Together these bound the window, state which way it fails, and refuse the defaults that turn a small window into an open door.
What does good hybrid retrieval look like?
Permission-correct retrieval still has to be good retrieval, and good retrieval in 2026 is hybrid. SphereIQ runs semantic search over pgvector embeddings and lexical BM25 search in PostgreSQL, with the title weighted above the body, fuses the two with reciprocal rank fusion, and can rerank the top candidates. Vector search finds the passage that means the same thing in different words; BM25 finds the exact identifier, error code or clause that a semantic model rounds off. You want both.
One detail matters more than it looks: embeddings from different models are not comparable. A query embedded with one model and a chunk embedded with another produce a similarity score that looks plausible and means nothing. So each search runs against one embedding model at a time, and every chunk records the model that embedded it — the same discipline as permissions, applied to relevance.
How to evaluate enterprise RAG
The demo will show you retrieval quality, because retrieval quality is what demos are good at. The two questions that decide whether you can deploy company-wide are quieter: where is the permission enforced — in the query, or on the output — and how long can a revoked permission keep answering. Ask both, and ask for a number on the second.
| Question | What good looks like |
|---|---|
| Where are permissions enforced? | Inside the retrieval query |
| Can an inaccessible chunk reach the model? | No — it is never a candidate |
| How do group and role changes propagate? | At query time, against current identity |
| What happens if a permission lookup fails? | Fail closed |
| How long can a revoked grant keep answering? | Bounded, and stated |
The same boundary applies when the platform runs in your own cloud — see self-hosted LLMs in your VPC. For the bigger picture of what a governed knowledge layer is, start with what a Company Brain is.
Frequently asked questions
What is permission-aware retrieval?
How is it different from role-based access control?
How quickly does a revoked permission take effect?
Does permission-aware RAG work self-hosted?
Ask a question as someone without access.
In a walkthrough we connect a restricted source and query it as a user who was never granted access — the right answer is nothing.