What Is Retrieval-Augmented Generation (RAG)?
Ask a general-purpose AI model a question about your company’s internal policies, and it will either guess or admit it doesn’t know — its training data stopped at some point in the past and never included your files in the first place. Retrieval-augmented generation, or RAG, is the fix most enterprises have landed on.
RAG is a technique that connects a large language model to an external knowledge source — a document repository, a database, or more commonly a vector database — and retrieves relevant information at the moment a question is asked, feeding it into the model’s context before it generates a response. Instead of relying only on what the model learned during training, RAG lets the model reach for current, organization-specific information: internal wikis, support tickets, contracts, code repositories, or compliance documentation.
How Does RAG Work?
- Ingestion and embedding. Source documents — PDFs, wikis, tickets, emails, code — are broken into chunks and converted into numerical representations called embeddings, which capture the semantic meaning of the text rather than just its keywords. These embeddings are stored in a vector database, purpose-built for fast similarity search.
- Query and retrieval. When a user asks a question, that query is also converted into an embedding, and the system searches the vector database for the chunks most semantically similar to it — not necessarily the ones containing the exact keywords, but the ones that mean something close to what was asked.
- Context injection. The retrieved chunks are inserted into the model’s context window alongside the user’s original question, effectively saying “here’s some relevant background — now answer using this.”
- Generation. The language model produces its response, grounded in the retrieved material rather than relying purely on its training data. Well-implemented RAG systems can also cite which source documents informed the answer, which is a meaningful improvement over an ungrounded model that may hallucinate a plausible-sounding but fabricated answer.
What Are the Security Risks of RAG?
Most organizations secure the model and secure the application layer around it, but the retrieval pipeline in between often falls through the cracks — it’s neither classic data storage nor classic application logic, and few existing tools were built to monitor it.
- Sensitive data exposure through the vector database. If source documents containing personal information, financial data, or trade secrets are ingested into the knowledge base, that data effectively becomes retrievable by any query that’s semantically close enough — even if the original document was never meant to be broadly accessible.
- Broken access control between systems and the model. This is arguably the sharpest risk. Source systems like SharePoint, Confluence, or a ticketing platform typically enforce their own permissions — not everyone can see every document. But if the RAG pipeline ingests content without preserving those original permissions, the model can retrieve and surface information to users who were never authorized to see it in the source system. The access control effectively evaporates at the point of ingestion unless it’s deliberately rebuilt into the retrieval layer
- Indirect prompt injection via retrieved content. Attackers don’t need to interact with the model directly — they only need to get malicious instructions into a document that’s likely to be retrieved. A poisoned support ticket, a manipulated wiki page, or a booby-trapped PDF can carry hidden instructions that the model treats as trusted context once it’s pulled into a response, potentially causing it to leak data or take unintended action.
- Knowledge base poisoning. Because RAG knowledge bases update continuously as source documents change, they’re an ongoing target — an attacker who can write to any ingested source (a shared drive, a wiki, a ticketing system) can plant content designed to corrupt future answers, not just a one-time exploit.
- Third-party and infrastructure exposure. RAG deployments typically involve several moving pieces — vector database software, orchestration frameworks, embedding models — each with its own vulnerability surface. Unpatched or misconfigured components in that stack have already been exploited in the wild to gain unauthorized access to connected systems.
How Can Enterprises Secure RAG Deployments?
- Preserve source-system permissions in the retrieval layer. The knowledge base should enforce the same access boundaries as the original documents, so a user querying the assistant only retrieves what they’d be authorized to see directly — a form of the same least-privilege principle behind zero trust network security.
- Classify data before ingestion, so sensitive categories are excluded from the knowledge base entirely or routed through stricter access rules.
- Treat retrieved content as untrusted input, applying input validation and monitoring similar to how you’d treat any external data reaching an application — a discipline borrowed from application security more broadly.
- Scope what connected agents can do, not just what they can see, particularly once a RAG system can take action rather than only answer questions — see AI agent access control for how to define those boundaries.
- Keep the surrounding infrastructure patched and monitored, including vector databases and orchestration frameworks, the same as any other production system.