Eight RAG architectures cover most real projects: naive, hybrid, HyDE, corrective, multimodal, graph, adaptive and agentic RAG. They differ in how they prepare content, how they search and whether they double-check what they find. The right one depends on why your current answers go wrong. NeoTek Solutions in Nashville designs and builds RAG systems for business and IT teams.
Retrieval-augmented generation (RAG) lets a large language model (LLM), the kind of AI behind chat assistants, answer from your own content. A basic version is easy to demo. Getting it to answer correctly for real users, on real documents, is harder. Most of that difficulty sits in one step: getting the right passages in front of the model.
That is why so many RAG “architectures” exist. They are not rival products. Each is a response to a specific way retrieval breaks down. This guide explains all eight in plain terms, from the simplest to the most involved. For each one, it covers how it works, where it breaks and use cases from the industries we serve.
What Do All RAG Architectures Share?
Every pattern here runs the same basic cycle. A person asks a question. The system gathers supporting material, such as passages, images or database rows. The LLM reads the question and that material together and writes an answer, ideally with citations to its sources.
The architectures differ in three places:
- Content preparationPlain text chunks, transcripts and image captions, or a graph of entities and relationships.
- Search methodBy meaning, by exact keywords, with a drafted answer or through several tools.
- Quality checksSome patterns trust the first search, others grade results, retry or loop.
Each added step can improve answers. Each one also adds delay, cost and parts to maintain. The goal is not the most sophisticated design. It is the simplest design that meets your quality bar.
New to the building blocks, such as chunking, embeddings and vector search? Start with our RAG architecture reference. Still deciding between RAG and training a model? Read RAG vs. fine-tuning first.
How Do the Eight Architectures Compare?
| Architecture | Problem it addresses | Typical use | What it adds |
|---|---|---|---|
| Naive RAG | The model lacks your information | Direct questions over one document set | Nothing; this is the baseline |
| Hybrid RAG | Vector search misses exact terms | SKUs, policy numbers, names, error codes | A keyword index and result merging |
| HyDE | Users word questions unlike the documents | Informal questions over technical content | One extra LLM call per question |
| Corrective RAG | Search returns weak or off-topic passages | Collections with gaps or uneven quality | A grading step and a fallback search |
| Multimodal RAG | Answers sit in diagrams, scans or recordings | Manuals, slide decks, calls, training video | Media processing and a vision-capable model |
| Graph RAG | Answers depend on links between facts | Relationship and “big picture” questions | A knowledge graph to build and maintain |
| Adaptive RAG | Every question gets the same effort | Mixed simple and complex traffic | A router that must be tested |
| Agentic RAG | Questions need several steps or systems | Multi-part questions across data sources | An agent loop, tools and strict limits |
1. Naive RAG: The Baseline to Measure First
Naive RAG, also called standard RAG, searches once and passes the closest passages to the model. The name sounds dismissive, but plenty of useful assistants run on it. With well-prepared content, it is often good enough, and it is always the benchmark for anything fancier.
How it works:
- Documents are split into chunks, typically a few hundred words each, following headings and sections where possible.
- An embedding model turns each chunk into a vector, a list of numbers that captures meaning, and the vector index stores it with its source and metadata.
- A user’s question is converted with the same embedding model.
- The index returns the few chunks whose vectors sit closest to the question.
- The LLM receives instructions, the question and those chunks, and answers only from them.
The weak point is step 4. If the right passage is not among the top results, the model answers from the wrong text or fills the gap itself. Stuffing in more chunks does not reliably help, because models can overlook facts buried in the middle of long prompts.
- Where it fitsDirect questions, well-organized content and answers that usually live in one place.
- Where it breaksChunks that cut tables in half, exact codes that meaning-based search blurs, and quality judged from a few demo questions.
- Practical tipBefore tuning anything, collect real questions with answers agreed by subject-matter experts. That evaluation set is how you will judge every pattern below.
Industry use cases:
- Retailstore associates ask about opening, closing and safety procedures from one approved operations manual.
- Governmentstaff search grant rules and internal policy manuals, with answers that cite the section.
- Manufacturingnew operators look up standard operating procedures (SOPs) written for one line or work cell.
2. Hybrid RAG: When Exact Codes and Plain Language Mix
Vector search is good at meaning but poor at exact strings. To an embedding model, part numbers “AX-4471” and “AX-4417” look almost the same. Keyword search has the opposite strength. Hybrid RAG runs both and merges the results, and for business content it has become a sensible default.
How it works:
- The question goes to a keyword index, usually scored with BM25, a long-established ranking formula based on matching terms.
- In parallel, it goes to a vector index that matches by meaning.
- The two ranked lists are merged. A popular method is reciprocal rank fusion (RRF). Each passage scores 1 ÷ (k + its rank) in each list, and the scores are added up; k is a constant, often 60. Passages ranked well by both methods rise to the top.
- Optionally, a re-ranker, a model that reads the question and each candidate passage side by side, reorders the leading results more precisely.
- The LLM answers from the strongest few passages.
Many search engines handle both methods in one index, including Azure AI Search, OpenSearch, Elasticsearch and PostgreSQL with pgvector plus its built-in full-text search. That keeps the added operational work small.
- Where it fitsQuestions that combine identifiers, such as SKUs, claim numbers and error codes, with everyday wording.
- Where it breaksMerge settings tuned on a handful of examples, and filters such as date or department applied to one search but not the other.
- Practical tipMeasure hybrid search alone before adding a re-ranker. Re-rankers often help, but they add delay and cost to every question.
Industry use cases:
- Manufacturingtechnicians search manuals and past work orders by fault code, part number or machine model.
- Financial Servicescompliance staff ask policy questions that name regulations and internal policy numbers.
- Entertainmentlicensing teams find tracks and contracts by exact title, writer name or identifier plus a plain description.
3. HyDE: Closing the Gap Between Questions and Documents
HyDE stands for hypothetical document embeddings. It addresses a quiet problem: people ask questions in different words than documents use. An employee types “can I carry over vacation days?” The policy reads “unused paid time off may be rolled into the following plan year.” Same meaning, very different wording.
How it works:
- The LLM writes a short, invented answer to the question, phrased like a passage from a policy or manual.
- That draft is converted into a vector, instead of or alongside the question itself. Some implementations write several drafts and average their vectors.
- The system searches the real documents with that vector. Text shaped like an answer tends to sit near real answer passages.
- The LLM answers using the original question and the real passages it found. The draft is discarded.
The draft can contain wrong facts. That is acceptable, because it is only a search aid. It must never be shown to users or treated as a source. HyDE also struggles when the model knows little about your field, because a draft full of the wrong vocabulary sends search in the wrong direction.
- Where it fitsShort or casual questions against formal, technical or jargon-heavy documents.
- Where it breaksUnfamiliar domains, and high-volume systems where an extra LLM call per question adds up.
- Practical tipTest HyDE against simpler options, such as rewriting the question or searching with several rephrased versions. Keep whichever finds the right passages most often at the lowest cost.
Industry use cases:
- Governmentresidents ask everyday questions, while the answers sit in formal ordinances and department pages.
- Healthcarecontact center agents type a member’s question, while benefit and payer documents use formal coverage terms.
- Retailshoppers ask about sizing or returns in casual words that rarely match the policy text.
4. Corrective RAG: Grading Results Before the Model Sees Them
Naive RAG uses whatever search returns. Corrective RAG, often shortened to CRAG, adds a quality gate: it scores the retrieved passages before any answer is written, and changes course when they fall short.
How it works:
- The system retrieves passages as usual.
- A grader, either a small trained model or an LLM with a scoring prompt, rates how well the passages address the question.
- Relevant passages are kept, and sentences that do not bear on the question are trimmed away.
- Off-topic results are dropped, and the system runs a fallback search, often with a reworded query.
- When the grade is unclear, the trimmed passages and the fallback results are used together.
- The LLM answers from this cleaned-up context.
The original research used web search as the fallback. For internal business questions, the open web is often the wrong place to look. Practical alternatives are a wider internal index, another department’s collection, or an honest “our sources don’t cover this” with a handoff to a person.
- Where it fitsContent with gaps or uneven quality, where a wrong answer costs more than a slower one.
- Where it breaksEvery question gets slower, graders need tuning, and a web fallback can send sensitive questions outside your environment or pull untrusted pages into the prompt.
- Practical tipLog every grade. Topics that keep triggering the fallback show your content team exactly what to write or fix next.
Industry use cases:
- Governmenta resident assistant sets aside outdated or off-topic pages and points to the right department instead of guessing.
- Financial Servicesa customer service assistant confirms product documents really answer the question, and escalates when they do not.
- Healthcarepayer rule lookups grade the retrieved policy passages and hand unclear cases to staff.
5. Multimodal RAG: Answers Hidden in Diagrams, Scans and Recordings
Much business knowledge is not plain text. It sits in wiring diagrams, scanned forms, slide decks, recorded trainings and call recordings. Multimodal RAG makes that material searchable and gives the model the evidence in a form it can use.
Teams usually build it in one of two ways, or mix both:
- Turn media into text. Optical character recognition (OCR) reads scans, speech recognition produces timestamped transcripts, and a vision model describes images and charts. All of it is indexed as text linked back to the original file.
- Search the media itself. Multimodal embedding models place images and text in a shared vector space. Newer retrievers embed entire page images, which preserves charts and layout that text extraction flattens.
At answer time, the system retrieves a mix of passages, images and clips, and a vision-capable model reads the text and examines the images. Citations should point to a page, a figure or a timestamp so people can verify them.
- Where it fitsAnswers that depend on diagrams, tables inside scanned PDFs, photos or recorded audio and video.
- Where it breaksVague image descriptions that hide what matters, heavy processing costs for long video, and models that cannot accept the formats you retrieve.
- Practical tipStart with the single format behind the most missed answers, often scanned PDFs or diagrams. For forms and field extraction, intelligent document processing is usually the better tool.
Industry use cases:
- Manufacturingmaintenance assistants read wiring diagrams, exploded part drawings and photos inside equipment manuals.
- Logisticsteams search scanned and photographed bills of lading and proof-of-delivery documents, such as for damage notes.
- Entertainmentcatalog search covers audio, video and transcripts, with results that link to a timestamp.
6. Graph RAG: Questions About Connections and Themes
Some questions cannot be answered from any single passage. “Which suppliers are tied to the parts in last quarter’s recalls?” or “What issues keep coming up across all our customer complaints?” require facts linked across many documents. Graph RAG maps those links in advance.
How it works:
- An LLM reads the source documents and pulls out entities, such as people, products, suppliers and policies, along with how they relate.
- These become a knowledge graph. In Microsoft’s GraphRAG approach, closely related entities are grouped into communities, and the LLM writes a summary of each community.
- Local search handles specific questions: it starts at the entities mentioned, follows their relationships and collects the linked source text.
- Global search handles broad questions: it draws on the community summaries and combines what they say about the topic.
- The LLM answers and cites the original source text, not just the graph.
Graph RAG can also run on an existing graph database built from structured records, such as assets, parts or reporting lines. That skips LLM extraction for data that is already structured.
- Where it fitsQuestions about how things connect, and requests for themes or summaries across large collections.
- Where it breaksIndexing is expensive because it takes many LLM calls, extraction errors turn into false “facts,” and the graph falls behind as documents change.
- Practical tipDesign for permissions from day one. A community summary can mix content from documents some users cannot open. Build separate graphs per access boundary, or restrict broad search on sensitive collections.
Industry use cases:
- Entertainmentrights teams trace how works, writers, publishers, territories and contracts connect when royalties go unmatched.
- Financial Servicesinvestigators see how applicants, businesses, accounts and addresses relate, while people make every decision.
- Manufacturingsafety teams find recurring themes across near-miss reports and link incidents to equipment and suppliers.
7. Adaptive RAG: Spending Effort Only Where It Is Needed
Questions vary widely in difficulty. “What does PTO stand for?” needs no search. “What is our international travel policy?” needs one. “How did last year’s policy change affect approval times for claims over $10,000?” needs several. Adaptive RAG adds a router that decides how much work each question gets.
How it works:
- A router assesses each incoming question. It can be a small classifier trained on labeled examples, or an LLM following clear routing rules.
- Simple questions skip retrieval, and the model answers directly.
- Standard questions get one retrieval pass, as in naive or hybrid RAG.
- Complex questions get iterative retrieval: search, read, then search again until the answer is supported.
A related research idea, Self-RAG, trains the model itself to decide when to retrieve and to critique its own drafts. Most business teams get similar gains from an explicit router, which is easier to test and to explain to auditors.
- Where it fitsHigh volumes of mixed simple and complex questions, where speed and cost matter.
- Where it breaksA router that skips retrieval for a question that needed company facts. The model then answers confidently from general knowledge, the very problem RAG is meant to solve.
- Practical tipWhen the router is unsure, have it retrieve. Test the router on its own labeled questions, separately from answer quality.
Industry use cases:
- Retaila service assistant answers “what are your hours?” directly and sends a gift-return question through full retrieval.
- Healthcarecontact center traffic mixes quick directory questions with detailed benefit and billing questions.
- Government311-style assistants handle simple service questions quickly and give questions that span departments more effort.
8. Agentic RAG: Multi-Step Questions Across Systems
Agentic RAG hands retrieval to an AI agent, a system that works toward a goal and chooses its own tools. Rather than searching once, the agent decides what information it needs, collects it from different sources and stops when it has enough to answer.
How it works:
- The agent reads the question and breaks it into steps or smaller questions.
- For each step, it picks a tool: document search, a read-only database query, a business API or, where approved, web search.
- It reviews what came back and keeps what is useful.
- It decides whether the evidence is now sufficient. If not, it plans another step.
- Once the evidence is sufficient, or a limit is hit, it writes the final answer with citations to every source used.
Example scenario: A claims operations manager asks how a policy change affected denial rates. The agent finds the policy memo in the document index, queries the claims database for denial rates before and after the effective date, and answers with both sources cited.
This is the most capable pattern and the hardest to predict. Steps, cost and response time change from one question to the next. An agent with tools also has more ways to do damage if manipulated, so guardrails are essential. Our agentic AI architecture guide covers the controls in more depth.
- Where it fitsQuestions that span several systems, involve calculations or comparisons, or cannot be settled with one search.
- Where it breaksLoops that never end, token costs that climb, tools with broader access than the task needs, and hidden instructions in retrieved text that try to trigger actions.
- Practical tipSet step limits, token budgets and timeouts. Run tools under the signed-in user’s permissions, keep database access read-only and log every tool call.
Industry use cases:
- Logisticsan exception assistant checks shipment status in the TMS, reads carrier emails and finds contract terms before drafting a reply.
- Healthcarea prior authorization assistant gathers payer criteria and supporting notes into a summary for staff to review.
- Financial Servicesan underwriting assistant pulls application data, guidelines and third-party reports into one summary for the underwriter.
Which RAG Architecture Fits Each Industry?
The table maps a common use case in each industry we serve to a sensible starting architecture and the patterns most often added later. Each use case links to a detailed guide covering sample questions, content sources, controls and rollout. Most projects start simple and add a pattern only when evaluation results show the gap it fills.
| Industry | Common use case | Start with | Add when needed |
|---|---|---|---|
| Healthcare | Contact center benefit and billing questions | Hybrid RAG | HyDE for member wording, adaptive routing, agentic prior authorization summaries |
| Financial services and insurance | Compliance and product policy questions | Hybrid RAG | Corrective checks, graph RAG for investigations, agentic underwriting summaries |
| Government | Resident-service assistant on official content | Naive RAG | HyDE for everyday wording, corrective checks for outdated pages |
| Manufacturing and automotive | Maintenance and SOP knowledge assistant | Hybrid RAG | Multimodal for diagrams and photos, graph RAG for incident themes |
| Retail and consumer brands | Customer service and store operations answers | Naive or hybrid RAG | Adaptive routing, agentic returns and order status |
| Logistics and supply chain | Exception management | Hybrid RAG | Multimodal for scanned delivery documents, agentic lookups across TMS and email |
| Music, media and entertainment | Contract, rights and catalog questions | Hybrid RAG | Graph RAG for rights relationships, multimodal for audio and video |
Regulated industries such as healthcare, financial services and government need permission-aware retrieval, citations and audit logs whichever pattern they choose. The security section below covers these controls.
Which RAG Architecture Should You Use?
We don’t pick an architecture up front. We launch naive or hybrid RAG with an evaluation set and logging, then study the questions that fail and let the symptom point to the fix:
- Exact codes, IDs or names get missedAdd hybrid search.
- The right document exists, but users phrase things differentlyTry HyDE or question rewriting.
- Search keeps returning off-topic passagesAdd a re-ranker, then corrective grading.
- The answer is inside a diagram, scan or recordingAdd multimodal processing for that format.
- Questions ask how things connect, or what patterns run across many documentsEvaluate graph RAG.
- Easy questions are slow or expensivePut an adaptive router in front.
- Questions need several systems or stepsConsider agentic RAG with strict guardrails.
- The content itself is wrong or missingFix the content. No architecture can retrieve what was never written down.
Can You Combine RAG Architectures?
Yes. Most production systems use two or three patterns together:
- Hybrid search with a re-rankeris a dependable foundation for most document assistants.
- An adaptive router over hybrid RAG and an agentic pathkeeps easy questions fast and gives hard ones full treatment.
- Corrective grading inside an agent looplets the agent judge whether each search helped before choosing its next move.
- Graph search next to hybrid searchcovers both relationship questions and exact lookups.
Add one pattern at a time and measure after each change. If a pattern does not improve results on your evaluation set, take it out. Every extra step is something your team has to run, monitor and pay for.
How Do You Know a Pattern Is Helping?
A new architecture should prove itself with measurements, not demos. Score retrieval and answers separately so you can tell which half went wrong:
- Context recallDid retrieval find the passages needed to answer?
- Context precisionHow much of what came back was actually relevant?
- FaithfulnessDoes the answer stay within the retrieved material, with no added claims?
- Answer relevanceDoes the answer address the question that was asked?
- Citation accuracyDo the cited sources really support each statement?
- LatencyTypical and slowest response times, since users remember the slow ones.
- Cost per answerModel, search and tool calls combined.
Build the evaluation set with subject-matter experts from real questions, including difficult ones and questions your sources cannot answer. Open-source frameworks such as Ragas can automate some scoring, but review a sample by hand as well, because automated judges make mistakes too.
What Security Controls Does Every RAG System Need?
Better retrieval also creates more paths to sensitive data. Whatever the architecture, these controls matter:
- Permission-aware searchFilter results by the signed-in user’s access before the model sees any text.
- Untrusted inputsTreat retrieved documents, web pages and tool results as data, never as instructions, to guard against prompt injection.
- Least-privilege toolsGive agents only the tools and access each task requires.
- Data boundariesDecide which questions and documents may leave your environment, especially with web fallback or hosted models.
- Audit trailsRecord questions, sources, tool calls and answers for later review.
An LLM gateway with guardrails puts many of these controls in one place. For regulated data, our AI governance, security and compliance team can help set policies before launch.
How Can NeoTek Solutions Help?
NeoTek Solutions, an AI company headquartered in Nashville, Tennessee, designs and builds RAG systems for organizations in Middle Tennessee and across the United States. We start with your questions and your content, not with a favorite architecture.
- Start from the failuresWe review your content sources, users, permissions and the answers that go wrong today, then recommend a starting pattern.
- Prove one assistantWe build a single assistant against an evaluation set your subject-matter experts agree on, so quality is a number rather than an impression.
- Harden it for productionWe add permission-aware search, monitoring and guardrails, and we stay vendor-neutral across Azure, AWS, Google Cloud and open-source options.
Your data is never used to train public models. Learn more about our generative AI solutions.
Frequently Asked Questions
Which RAG architecture do you usually recommend first?
No pattern is best in the abstract. For most business document assistants we start with hybrid search and a re-ranker, then add another pattern only when your evaluation results show the gap it fills.
Do you build agentic RAG?
Yes, where the questions need it. Agentic RAG answers multi-step questions one search cannot, and it is slower, costlier and harder to control. We usually put it behind an adaptive router, with step limits, read-only data access and logging.
When do you choose graph RAG over hybrid RAG?
We choose hybrid when the answer sits in a passage, and graph when it depends on how entities relate. Hybrid improves how passages are found, while graph changes what gets indexed, and it costs more to build and keep current.
Do you combine these architectures in one system?
Yes. Most of the systems we put into production use two or three, often an adaptive router in front of hybrid RAG with an agentic path for complex questions. We add one pattern at a time and measure each change.
What do you recommend for regulated teams in healthcare, financial services or government?
We usually start those teams on hybrid RAG, because their questions mix exact identifiers, such as policy, claim or ordinance numbers, with plain language. We treat permission-aware search, citations and audit logs as more important than the pattern, and we add corrective checks and human review for high-stakes answers.
Can you stop the system from making things up?
No design removes hallucinations, and we do not claim otherwise. We reduce them with stronger retrieval, evidence checks, citations and instructions to admit uncertainty, and we keep a person in the loop for high-stakes answers.
Can NeoTek Solutions build a RAG system for our organization?
Yes. We design, build and support RAG systems, from a first pilot through production, for teams in Middle Tennessee and across the United States. We work by video, by phone or onsite at your office.
Further Reading
These research papers introduced or shaped several of the patterns above:
- Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE), Gao and others, 2022.
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Asai and others, 2023.
- Corrective Retrieval Augmented Generation (CRAG), Yan and others, 2024.
- Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity, Jeong and others, 2024.
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization (Microsoft GraphRAG), Edge and others, 2024.
Find the Right RAG Architecture for Your Data
Tell us what your teams ask, where the answers live and where your current assistant falls short. We will recommend a practical architecture and help you build, test and run it. Learn more about our generative AI solutions or book a free AI consultation.