Back to BlogAI & Machine Learning

A Guide to RAG (Retrieval-Augmented Generation)

RAG (Retrieval-Augmented Generation)

RAG is a technique that combines two components:

  • Retrieval: Retrieve relevant information from a data store (vector database, search engine, etc.).
  • Generation: Generate an answer based on the retrieved information using a large language model (LLM).

Process:

  1. The user asks a question.
  2. The system searches for relevant text passages in the data store (retriever).
  3. These passages are placed into the LLM prompt to generate the final answer (generator).

COMMON RAG TECHNIQUES

Classic RAG

  • Description: Retrieve relevant text passages from a data store, then place them into the LLM prompt to generate an answer.
  • Applications:
    • Internal document Q&A (FAQ, user guides, technical documents).
    • Enterprise virtual assistants.
    • Customer-support chatbots.
  • Techniques:
    • Split documents into smaller passages (chunks).
    • Create an embedding for each chunk and store it in a vector database.
    • When a question arrives, convert the question into an embedding and find the most relevant passages (top-k).
    • Place these passages into the LLM prompt to generate an answer.
  • Tools: Chroma, FAISS, Pinecone, LangChain, HuggingFace.

Multi-Document RAG

  • Description: Retrieve and synthesize information from many different documents.
  • Applications:
    • Legal search, compiling medical records.
    • Synthesizing reports from multiple data sources.
  • Techniques:
    • Store embeddings of many different documents (optionally tagged with source labels).
    • Retrieve relevant passages from multiple documents at the same time.
    • Optionally synthesize or group results by document source.
  • Tools: Vector DBs with metadata support (Chroma, Weaviate, Qdrant), LangChain.

Multi-modal RAG

  • Description: Retrieve and combine information from multiple data types: text, images, audio, and more.
  • Applications:
    • Medical assistants (combining X-ray images and medical records).
    • Product search via images and descriptions.
  • Techniques:
    • Create embeddings for multiple data types (text, images, audio) using different models.
    • Store multimodal embeddings in a vector DB.
    • At query time, convert the question or input data into a suitable embedding to retrieve different data types.
    • Combine results from multiple data types and feed them into the LLM.
  • Tools: CLIP, BLIP, Weaviate, Milvus, LangChain.

Conversational RAG

  • Description: Keep conversation history and retrieve information based on the full conversation context.
  • Applications:
    • Intelligent chatbots, personal assistants.
    • Multi-turn customer support.
  • Techniques:
    • Keep conversation history (context window).
    • At query time, combine the new question with conversation history to retrieve suitable information.
    • Place context and conversation history into the LLM prompt.
  • Tools: LangChain Memory, ConversationBuffer, vector DB.

Hybrid RAG (Combining multiple retrievers)

  • Description: Combine multiple retrieval methods (semantic search, keyword search, rule-based, and more) to increase accuracy.
  • Applications:
    • Enterprise digital search at large scale.
    • In-depth Q&A systems (medical, legal).
  • Techniques:
    • Use multiple retrieval methods in parallel: semantic search (embedding), keyword search (BM25), rule-based, and more.
    • Merge or filter results from different retrievers.
    • Place the most relevant passages into the LLM.
  • Tools: LangChain MultiRetriever, ElasticSearch, FAISS, BM25.

RAG with dynamic chunking (Dynamic Chunking RAG)

  • Description: Split documents by meaning or dynamic structure, optimizing the context sent to the LLM.
  • Applications:
    • Processing long documents, technical reports, books.
    • Domain-specific Q&A systems.
  • Techniques:
    • Split documents based on meaning, structure, or dynamic length (not a fixed size).
    • May use semantic chunking, sliding window, or adaptive chunking.
    • Optimize the number and length of chunks to fit the LLM context window.
  • Tools: LangChain SemanticChunker, custom splitter, HuggingFace tokenizer.

COMPARING RAG WITH OTHER CHATBOT TECHNIQUES

Rule-based Chatbot (Hard rules)

  • Principle: Based on a set of IF-THEN rules, predefined scripts, or a decision tree.
  • Advantages:
    • Easy to control; responses are predictable.
    • Does not require large datasets; easy to deploy for simple tasks.
  • Disadvantages:
    • Inflexible; does not understand complex context.
    • Does not scale well across many topics.
  • Applications:
    • Simple FAQ, choice menus, automated call centers.

Retrieval-based Chatbot (Pure retrieval)

  • Principle: Search for the best answer from an existing data store (FAQ, documents) using keyword or semantic search.
  • Advantages:
    • Answers accurately if the information already exists in the data store.
    • Easy to control content; does not generate incorrect information.
  • Disadvantages:
    • Does not synthesize or rephrase information.
    • Cannot answer novel or complex questions.
  • Applications:
    • Internal search systems, document assistants.

Generative Chatbot (Pure text generation)

  • Principle: Use an LLM (GPT, Llama, etc.) to generate answers from a prompt, without retrieving external data.
  • Advantages:
    • Flexible, creative answers, strong context understanding.
    • Can answer many topics, including ones not present in the data.
  • Disadvantages:
    • Prone to fabricating information (hallucination).
    • Hard to control content; may generate incorrect information.
  • Applications:
    • Natural-conversation chatbots, personal assistants.

RAG (Retrieval-Augmented Generation)

  • Principle: Combine retrieving relevant information from a data store (retriever) with generating an answer based on that information using an LLM (generator).
  • Advantages:
    • Flexible, creative answers that are still grounded in real data.
    • Reduces the risk of fabricating information and increases reliability.
    • Easy to extend; new data can be updated without retraining the LLM.
  • Disadvantages:
    • More technically complex (requires a vector DB, embeddings, pipeline).
    • Performance depends on retrieval quality and chunking.
  • Applications:
    • Document Q&A chatbots, enterprise assistants, knowledge search, report synthesis.

Comparison summary table

TechniqueFlexibilityReliabilityContext understandingEasy to controlEasy to scaleMain applications
Rule-basedLowHighLowHighLowFAQ, menus, call centers
RetrievalMediumVery highLowHighMediumSearch, document assistants
GenerativeVery highLowVery highLowHighNatural chat, creative use
RAGHighHighHighMediumHighDocument QA, enterprise assistants

TECHNICAL ANALYSIS & PROGRAMMING FLOW OF THE BASE CODE

Pipeline

  • Step 1. Upload a PDF document: Use file_uploader

  • Step 2: Split the PDF document: Use PyPDFLoader to read the PDF content.

  • Step 3: Semantic Chunking: Use SemanticChunker to split the document into meaningful passages (chunks), which improves retrieval effectiveness.

    Semantic Chunking

    This is a technique for splitting (chunking) text into passages based on meaning or semantics, rather than only character length, sentence count, or word count. The goal is for each chunk to contain a complete idea, helping the model retrieve and understand context better when performing tasks such as RAG.

    Common types of chunking:

    1. Fixed-size Chunking

      • Split text into passages of a fixed length (by number of words, characters, or sentences).
      • Simple and fast, but easy to cut ideas apart.
    2. Sliding Window Chunking

      • Split into overlapping passages (overlap), which helps preserve context between chunks.
      • More effective than fixed-size, but still not optimal semantically.
    3. Semantic Chunking

      • Split based on meaning, grammatical structure, or natural breakpoints (for example: paragraphs, headings, topic sentences).
      • May use a language model, grammatical analysis, or embeddings to determine split points.
      • Preserves the complete meaning of each passage; suitable for RAG and advanced NLP applications.

  • Step 4: Embeddings: Use an embedding model to convert text passages into numeric vectors (embeddings).

  • Step 5: Store in a Vector Database: Store embeddings with Chroma (Vector database) to retrieve relevant passages quickly.

  • Step 6: Retriever: Retrieve the text passages most relevant to the question.

  • Step 7: Prompt Template: Use a sample prompt from LangChain Hub to combine context and the question.

  • Step 8: Use an LLM to generate the answer: Use an LLM model to generate an answer based on the retrieved context.

  • Step 9: Use Streamlit to build the interface and display the LLM answer: Build a simple web interface so users can upload a file, ask a question, and receive an answer.

Common types of vector databases

  1. Chroma
    • Open source, easy to integrate with Python, supports local and cloud.
    • Suitable for small projects, demos, and research.
  2. Pinecone
    • Cloud service, high performance, scales well, easy to manage.
    • Suitable for real products, large data, and many users.
  3. Weaviate
    • Open source, cloud support, integrates many data types (text, images, graphs).
    • Scalable, with many AI features.
  4. Qdrant
    • Open source, high performance, cloud support, strong REST API.
    • Optimized for semantic search and AI applications.
  5. FAISS
    • Facebook's C++/Python library, very high performance, mainly used locally.
    • Does not have data-management features like other vector DBs.
  6. Milvus
    • Open source, high performance, cloud support, manages large data.
    • Suitable for large-scale AI systems.
VDB nameOpen sourceCloudPerformanceData managementEase of useSuitable for
ChromaYesYesMediumBasicVery easyDemos, small projects
PineconeNoYesHighCompleteEasyProducts
WeaviateYesYesHighCompleteEasyProducts
QdrantYesYesHighCompleteEasyProducts
FAISSYesNoVery highNoMediumResearch, local
MilvusYesYesVery highCompleteMediumLarge scale

Sample prompts

A sample prompt (prompt template) is a structured piece of text that contains variables (placeholders) for inserting dynamic content (for example: context, question, data...). A sample prompt makes it easy to create consistent input instructions for a large language model (LLM), which improves the quality and stability of generated results.

Sample prompt example:

Answer the question based on the following information:
{context}

Question: {question}
Answer:

Prompts from LangChain

LangChain Hub is a repository of sample prompts, chains, and agent templates shared by the community and optimized for many purposes (QA, summarization, classification, etc.). You can download and use them directly in code with LangChain.

LangChain Hub provides many types of sample prompts, commonly including:

  • Question Answering (QA) Prompt: Answer a question based on context or documents.
  • Summarization Prompt: Summarize text.
  • Classification Prompt: Classify text, intent, sentiment, and more.
  • Translation Prompt: Translate language.
  • Chat/Conversation Prompt: Create dialogue, chatbots.
  • Extraction Prompt: Extract information (entities, key-value pairs, and more).
  • Rewriting/Paraphrasing Prompt: Rewrite or rephrase text.
  • Custom Prompt: Prompts created by the community for special purposes.

Note: The number and types of prompts on LangChain Hub are continuously updated and expanded by the community.

Common formats used for sample prompts

1. Template-style prompt with variables (placeholders)

Answer the question based on the following information:
{context}

Question: {question}
Answer:

• {context} and {question} are variables that will be replaced with actual data at runtime.

2. JSON or YAML-style prompt (for more complex systems)

system: |
  You are an intelligent AI assistant.
user: |
  Based on the following information, please answer the question:
  {context}

  Question: {question}
assistant: |
  Answer:

3. Python f-string prompt

prompt = f"""
Please answer based on the following information:
{context}

Question: {question}
Answer:
"""

4. LangChain PromptTemplate-style prompt

from langchain.prompts import PromptTemplate

template = """
Based on the following information, please answer the question:
{context}

Question: {question}
Answer:
"""
prompt = PromptTemplate(template=template, input_variables=["context", "question"])

5. Chat-style prompt (multi-turn)

messages = [
    {"role": "system", "content": "You are an AI assistant."},
    {"role": "user", "content": "Based on the following information, please answer: {context}"},
    {"role": "user", "content": "Question: {question}"}
]

Common LLM models with good Vietnamese support

  1. BKAI-ViLM
  • Author: BKAI (BKAV Institute of Artificial Intelligence)
  • Characteristics:
    • Trained in depth for Vietnamese.
    • Has encoder versions (bi-encoder, cross-encoder) and a decoder.
    • Highly effective for Vietnamese question answering, classification, and semantic search.

  1. PhoGPT
  • Author: VinAIResearch
  • Characteristics:
    • A general-purpose LLM trained on large Vietnamese data.
    • Strong support for conversation, summarization, translation, and question answering.
  1. VietAI/vietnamese-llama
  • Author: VietAI
  • Characteristics:
    • Fine-tuned from Llama for Vietnamese.
    • Strong support for general tasks, conversation, and question answering.
  1. VinAI PhoBERT / ViT5
  • Author: VinAIResearch
  • Characteristics:
    • PhoBERT: A strong encoder model for Vietnamese (suitable for embeddings and classification).
    • ViT5: An encoder-decoder model for text generation, summarization, and translation.
  1. Mistral, Llama-2, GPT-3.5/4 (multilingual, with a certain level of Vietnamese support)
  • Characteristics:
    • These models support many languages, including Vietnamese, but quality is not equal to models specialized for Vietnamese.
    • Suitable when you need multilingual support or fast integration.

ADVANTAGES/DISADVANTAGES OF THE BASE CODE

Advantages

  • Easy to extend: You can change the embedding model, LLM, or vector DB easily.
  • Vietnamese support: Uses Vietnamese embeddings and a strong LLM.
  • Smart splitting: Semantic chunking improves retrieval quality.
  • Friendly interface: Streamlit makes demos fast.

Disadvantages

  • Performance: Processing large documents or many users will be slow because it runs on a personal CPU/GPU.
  • Version management: Easy to hit dependency errors between packages (as you have already encountered).
  • Security: There is no access control; anyone can upload documents.
  • Scalability: Not suitable for production or large data volumes.
  • Error handling: Limited error control or detailed logging.

SUGGESTED UPGRADES FOR THE BASE CODE

  • Use cloud services: Deploy the vector DB (Pinecone, Weaviate, Qdrant) and LLM (OpenAI, Azure, etc.) on the cloud to increase performance and scalability.
  • Dynamic chunking: Apply adaptive chunking or sliding window to optimize context.
  • Caching & batching: Cache embedding, retrieval, and answer-generation results to reduce latency.
  • Package version management: Use a requirements.txt file or conda env to pin versions and avoid dependency errors.
  • Security: Add user authentication and control over upload/download.
  • Multi-document processing: Support uploading multiple files, classification, and multi-document search.
  • Prompt optimization: Customize prompts for each type of question or document.
  • Logging & monitoring: Add logging and monitor performance and errors.

SUMMARY

The current base code is suitable for demos, research, or small applications. For real-world deployment, you should move to a microservice architecture, use the cloud, optimize chunking, and tighten security and version management.

Huỳnh Phước Nguyên

Written by Huỳnh Phước Nguyên

AI Engineer, BK Hightech

Ready to build something great?

Tell us about your project and we'll get back to you within a day.

Get in Touch