{"id":28709,"date":"2025-10-20T09:27:00","date_gmt":"2025-10-20T08:27:00","guid":{"rendered":"https:\/\/www.web-systems.pl\/rag-retrieval-augmented-generation-large-language-models\/"},"modified":"2025-10-20T09:27:00","modified_gmt":"2025-10-20T08:27:00","slug":"rag-retrieval-augmented-generation-large-language-models","status":"publish","type":"post","link":"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/","title":{"rendered":"RAG (Retrieval-Augmented Generation) &#8211; how a new architecture is changing work with large language models"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">RAG, or Retrieval-Augmented Generation, is a technique that has become one of the most important trends in the development of artificial intelligence (AI). It combines the capabilities of large language models (LLM) with dynamic access to external knowledge sources: documents, databases, company archives or expert repositories.<br>Why is RAG gaining so much popularity? Because it solves the fundamental problems that today&#8217;s language models face &#8211; a lack of up-to-date information, a lack of specialist knowledge and so-called hallucinations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In this article we explain exactly how Retrieval-Augmented Generation works, where it performs best and why it is becoming the standard in modern AI implementations &#8211; in startups and large corporations alike.<\/p>\n\n\n\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_86 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Spis tre\u015bci<\/p>\n<span class=\"ez-toc-title-toggle\"><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><ul class='ez-toc-list-level-2' ><li class='ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#What_does_Retrieval-Augmented_Generation_RAG_mean\" >What does Retrieval-Augmented Generation (RAG) mean?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#1_Retrieval_%E2%80%93_intelligent_data_search\" >1. Retrieval &#8211; intelligent data search<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#2_Augmented_%E2%80%93_enriching_the_query_with_context\" >2. Augmented &#8211; enriching the query with context<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#3_Generation_%E2%80%93_producing_an_answer_based_on_context\" >3. Generation &#8211; producing an answer based on context<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#How_does_RAG_solve_the_problems_of_large_language_models\" >How does RAG solve the problems of large language models?<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#1_No_up-to-date_knowledge\" >1. No up-to-date knowledge<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#2_No_specialist_knowledge\" >2. No specialist knowledge<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#3_Hallucinations\" >3. Hallucinations<\/a><\/li><\/ul><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#RAG_AI_%E2%80%93_a_breakthrough_in_information_processing\" >RAG AI &#8211; a breakthrough in information processing<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#The_benefits_of_using_RAG_accuracy_context_savings\" >The benefits of using RAG: accuracy, context, savings<\/a><ul class='ez-toc-list-level-2' ><li class='ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#Accuracy\" >Accuracy<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#Context\" >Context<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#Savings\" >Savings<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#RAG_LLM_vs_fine-tuning_%E2%80%93_which_one_to_choose\" >RAG LLM vs. fine-tuning &#8211; which one to choose?<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#When_to_use_RAG\" >When to use RAG?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#When_to_use_fine-tuning\" >When to use fine-tuning?<\/a><\/li><\/ul><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#Applications_of_RAG_in_companies_and_projects\" >Applications of RAG in companies and projects<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-18\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#1_Company_chatbots_and_AI_assistants\" >1. Company chatbots and AI assistants<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-19\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#2_HR_and_onboarding\" >2. HR and onboarding<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-20\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#3_Sales_and_marketing\" >3. Sales and marketing<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-21\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#4_Law_and_medicine\" >4. Law and medicine<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-22\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#5_Education_and_training\" >5. Education and training<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-23\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#6_IT_and_DevOps\" >6. IT and DevOps<\/a><\/li><\/ul><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-24\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#How_to_implement_RAG_in_your_application_or_company\" >How to implement RAG in your application or company?<\/a><ul class='ez-toc-list-level-2' ><li class='ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-25\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#1_Preparing_the_data\" >1. Preparing the data<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-26\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#2_Creating_embeddings\" >2. Creating embeddings<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-27\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#3_Indexing_in_a_vector_database\" >3. Indexing in a vector database<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-28\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#4_Searching_for_context\" >4. Searching for context<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-29\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#5_Generating_the_answer\" >5. Generating the answer<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-30\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#6_Continuous_optimization\" >6. Continuous optimization<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-31\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#The_future_of_RAG_%E2%80%93_where_is_the_technology_heading\" >The future of RAG &#8211; where is the technology heading?<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-32\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#1_Intelligent_data_discovery\" >1. Intelligent data discovery<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-33\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#2_Multimodality\" >2. Multimodality<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-34\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#3_Personalization\" >3. Personalization<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-35\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#4_Transparency\" >4. Transparency<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-36\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#5_RAG_agents\" >5. RAG agents<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-37\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#6_Working_on_local_and_private_data\" >6. Working on local and private data<\/a><\/li><\/ul><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-1'><a class=\"ez-toc-link ez-toc-heading-38\" href=\"https:\/\/www.web-systems.pl\/en\/rag-retrieval-augmented-generation-large-language-models\/#Summary_%E2%80%93_why_is_it_worth_implementing_RAG\" >Summary &#8211; why is it worth implementing RAG?<\/a><\/li><\/ul><\/nav><\/div>\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"What_does_Retrieval-Augmented_Generation_RAG_mean\"><\/span>What does Retrieval-Augmented Generation (RAG) mean?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The name RAG comes from the three stages that make up the complete process of generating an answer:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Retrieval<\/strong> &#8211; searching for information in external sources<\/li>\n\n\n\n<li><strong>Augmented<\/strong> &#8211; enriching the query with context<\/li>\n\n\n\n<li><strong>Generation<\/strong> &#8211; generating the answer with a large language model (LLM)<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">These three processes work as a single organism. RAG simultaneously <strong>searches for data<\/strong>, <strong>adds it to the query<\/strong> and <strong>generates an answer based on it<\/strong>, which significantly increases the precision and reliability of the results. In practice this means the model does not rely solely on its built-in knowledge, but uses data supplied by the organization &#8211; current, domain-specific and verified.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Retrieval_%E2%80%93_intelligent_data_search\"><\/span>1. Retrieval &#8211; intelligent data search<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The first step consists of analyzing the user&#8217;s question and searching the available resources:<br>&#8211; PDF documents,<br>&#8211; manuals and regulations,<br>&#8211; SQL databases,<br>&#8211; company wikis,<br>&#8211; code repositories,<br>&#8211; articles and training materials.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">RAG does not work like a traditional keyword search engine. It uses <strong>semantic search<\/strong>, which analyzes the meaning and context of the question. Thanks to embeddings and vector databases (FAISS, Qdrant, Pinecone, Weaviate) it can find even those passages that do not literally contain the words used, but semantically match the intent of the question.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The effect? The search is far more precise and takes into account the industry context, the company&#8217;s internal vocabulary and the complex relationships between pieces of information.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_Augmented_%E2%80%93_enriching_the_query_with_context\"><\/span>2. Augmented &#8211; enriching the query with context<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">After the relevant passages have been found, the system <strong>does not return them to the user in raw form<\/strong>. Instead it attaches them as context to the query sent to the model. The LLM &#8220;reads&#8221; this data first, so the generated answer is:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>current<\/strong>, because it is based on fresh documents,<\/li>\n\n\n\n<li><strong>accurate<\/strong>, because it comes from the organization&#8217;s own sources,<\/li>\n\n\n\n<li><strong>safe<\/strong>, because it limits the risk of hallucinations.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The augmentation stage is exactly what the whole RAG concept is built on. The model stops being a &#8220;general&#8221; tool and starts to act like an expert working with the company knowledge base.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"3_Generation_%E2%80%93_producing_an_answer_based_on_context\"><\/span>3. Generation &#8211; producing an answer based on context<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In the final step the LLM generates an answer based on:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>the user&#8217;s question,<\/li>\n\n\n\n<li>the enriched context,<\/li>\n\n\n\n<li>data retrieved in real time.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">As a result the model does not guess, it <strong>draws conclusions from real information<\/strong>. RAG also makes it possible to cite sources, summarize large documents, combine information from many places and eliminate errors caused by gaps in the model&#8217;s training.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The user therefore receives an answer that is reliable, verified and tailored to the specifics of their organization.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"How_does_RAG_solve_the_problems_of_large_language_models\"><\/span>How does RAG solve the problems of large language models?<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_No_up-to-date_knowledge\"><\/span>1. No up-to-date knowledge<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">LLMs have no access to the internet or to new information after their training ends.<br>RAG pulls data on an ongoing basis from current documents and knowledge bases, so the answers match the actual state of affairs.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_No_specialist_knowledge\"><\/span>2. No specialist knowledge<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Language models handle natural language brilliantly, but they do not know internal company procedures.<br>RAG connects them with your own knowledge base &#8211; the model automatically becomes an expert in the field the organization works in.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"3_Hallucinations\"><\/span>3. Hallucinations<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">LLMs sometimes generate answers that sound convincing but are untrue.<br>RAG significantly reduces this problem, because the answers are built on verified sources. If the information is missing, the model can safely reply that it has no knowledge on the subject.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"RAG_AI_%E2%80%93_a_breakthrough_in_information_processing\"><\/span>RAG AI &#8211; a breakthrough in information processing<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<p class=\"wp-block-paragraph\">Classic language models such as GPT-4 or LLaMA are powerful tools, but their knowledge is static.<br>RAG introduces a dynamic component: the ability to use data in real time.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As a result:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>answers are current,<\/li>\n\n\n\n<li>they can be verified,<\/li>\n\n\n\n<li>sources can be indicated,<\/li>\n\n\n\n<li>and the model does not need to be retrained.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This makes RAG a key element of modern AI infrastructure &#8211; especially in companies that work with a large number of documents and procedures.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_benefits_of_using_RAG_accuracy_context_savings\"><\/span>The benefits of using RAG: accuracy, context, savings<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Accuracy\"><\/span><strong>Accuracy<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The model generates answers based on specific passages of documents.<br>The result: fewer errors and greater credibility.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Context\"><\/span><strong>Context<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">RAG takes into account:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>the conversation history,<\/li>\n\n\n\n<li>the specifics of the user,<\/li>\n\n\n\n<li>industry vocabulary,<\/li>\n\n\n\n<li>the structure of company documents.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Thanks to this, answers are precisely matched to the actual case at hand.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Savings\"><\/span><strong>Savings<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Instead of expensive retraining of the model, it is enough to update the documents.<br>RAG automatically starts using them. That is a significant reduction in the maintenance cost of AI systems.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"RAG_LLM_vs_fine-tuning_%E2%80%93_which_one_to_choose\"><\/span>RAG LLM vs. fine-tuning &#8211; which one to choose?<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"When_to_use_RAG\"><\/span>When to use RAG?<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<ul class=\"wp-block-list\">\n<li>when the knowledge changes frequently,<\/li>\n\n\n\n<li>when you need data from many sources,<\/li>\n\n\n\n<li>when the answers have to be grounded in facts,<\/li>\n\n\n\n<li>when you want the model to work on company documentation,<\/li>\n\n\n\n<li>when the volume of data is too large for fine-tuning.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"When_to_use_fine-tuning\"><\/span>When to use fine-tuning?<span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<ul class=\"wp-block-list\">\n<li>when you want to teach the model a specific style of expression,<\/li>\n\n\n\n<li>when the data is stable and does not change too often,<\/li>\n\n\n\n<li>when a command of specific domain language matters.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">In practice <strong>the best results come from combining both approaches<\/strong> &#8211; a fine-tuned model can be fed with knowledge through RAG at the same time.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Applications_of_RAG_in_companies_and_projects\"><\/span>Applications of RAG in companies and projects<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<p class=\"wp-block-paragraph\">RAG is already used in hundreds of commercial applications. Most often in:<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Company_chatbots_and_AI_assistants\"><\/span><strong>1. Company chatbots and AI assistants<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Customer service, helpdesk, answers to employee questions, analysis of regulations, policies and procedures.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_HR_and_onboarding\"><\/span><strong>2. HR and onboarding<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Automatic answers to new employees&#8217; questions, personalized onboarding processes.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"3_Sales_and_marketing\"><\/span><strong>3. Sales and marketing<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Recommendation systems, analysis of customer enquiries, generating offers based on current product data.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"4_Law_and_medicine\"><\/span><strong>4. Law and medicine<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Searching legal provisions, court rulings, case descriptions and test results &#8211; speeding up the work of experts.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"5_Education_and_training\"><\/span><strong>5. Education and training<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Assistants that learn from course materials and current publications.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"6_IT_and_DevOps\"><\/span><strong>6. IT and DevOps<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">RAG can analyze documentation, code and repositories and support developers in real time.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The technology is remarkably flexible &#8211; it can be implemented in practically any organization that works with knowledge.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"How_to_implement_RAG_in_your_application_or_company\"><\/span>How to implement RAG in your application or company?<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<p class=\"wp-block-paragraph\">The RAG implementation process consists of several stages:<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Preparing_the_data\"><\/span><strong>1. Preparing the data<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Documents are split into smaller fragments (chunks), which makes precise searching possible.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_Creating_embeddings\"><\/span><strong>2. Creating embeddings<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Every text fragment is converted into a numerical vector that reflects its meaning.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"3_Indexing_in_a_vector_database\"><\/span><strong>3. Indexing in a vector database<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The most popular solutions:<br>FAISS, Qdrant, Weaviate, Milvus, Pinecone.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"4_Searching_for_context\"><\/span><strong>4. Searching for context<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Once a question is asked, the system finds the best matching fragments.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"5_Generating_the_answer\"><\/span><strong>5. Generating the answer<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The LLM creates an answer based on the question and the context provided.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"6_Continuous_optimization\"><\/span><strong>6. Continuous optimization<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This includes:<br>&#8211; improving the chunking,<br>&#8211; filtering documents,<br>&#8211; automatic updates of the knowledge base,<br>&#8211; advanced prompt engineering.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">RAG can be integrated with existing systems &#8211; CRM, ERP, DMS, the intranet, a helpdesk, SQL and NoSQL databases.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_future_of_RAG_%E2%80%93_where_is_the_technology_heading\"><\/span>The future of RAG &#8211; where is the technology heading?<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<p class=\"wp-block-paragraph\">RAG is developing at a fast pace. The most important directions of development are:<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"1_Intelligent_data_discovery\"><\/span><strong>1. Intelligent data discovery<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Models will automatically assess the relevance of results, combine multiple sources and pick the best fragments.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"2_Multimodality\"><\/span><strong>2. Multimodality<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">RAG will cover not only text, but also:<br>&#8211; images,<br>&#8211; video,<br>&#8211; audio,<br>&#8211; charts,<br>&#8211; tabular data.<br>AI will retrieve a fragment of a film, a technical diagram or a statistical chart as part of the answer.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"3_Personalization\"><\/span><strong>3. Personalization<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Systems will learn the user&#8217;s preferences and predict what information they need.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"4_Transparency\"><\/span><strong>4. Transparency<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">More and more emphasis is placed on explaining where an answer comes from and why it was generated.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"5_RAG_agents\"><\/span><strong>5. RAG agents<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">LLMs will decide on their own when to run a search, how to combine the results and what actions to take later in the process &#8211; for example send an email, prepare an analysis or put together a report.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"6_Working_on_local_and_private_data\"><\/span><strong>6. Working on local and private data<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">RAG integrates with local LLMs, which makes it possible to process sensitive data without sending it to the cloud.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Everything indicates that in the coming years RAG will become the standard in AI systems &#8211; just as SQL databases became the standard in business applications.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Summary_%E2%80%93_why_is_it_worth_implementing_RAG\"><\/span>Summary &#8211; why is it worth implementing RAG?<span class=\"ez-toc-section-end\"><\/span><\/h1>\n\n\n\n<p class=\"wp-block-paragraph\">RAG is a key technique that lets companies move from the stage of AI experiments to real, production-grade business value.<br>By combining knowledge and language in a single process it:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>increases the precision and credibility of answers,<\/li>\n\n\n\n<li>makes it possible to update knowledge in real time,<\/li>\n\n\n\n<li>improves customer service and internal work processes,<\/li>\n\n\n\n<li>speeds up document analysis,<\/li>\n\n\n\n<li>lowers the cost of AI development in the organization.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">If you are thinking about implementing your own AI chatbot or a RAG-based system that genuinely supports users and employees &#8211; <a href=\"https:\/\/www.web-systems.pl\/en\/contact\/\" data-type=\"page\" data-id=\"21033\">get in touch with us<\/a>.  We will show you how to build a solution matched to the needs of your company and how to make effective use of the potential of AI in practice. We have already delivered several RAG systems and we know how to do it; right now we are working on a RAG for the City of Warsaw municipal office.<\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>RAG, or Retrieval-Augmented Generation, is a technique that has become one of the most important trends in the development of artificial intelligence (AI). It combines the capabilities of large language models (LLM) with dynamic access to external knowledge sources: documents, databases, company archives or expert repositories.Why is RAG gaining so much popularity? Because it solves [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":27685,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[816,820,808,805,834],"tags":[856,899,1008,1101,1103,1188],"class_list":["post-28709","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-applications","category-business","category-it-en","category-tools","category-news-trends","tag-app","tag-databases","tag-llm-en","tag-rag-en","tag-retrieval-augmented-generation-en","tag-knowledge"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/posts\/28709","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/comments?post=28709"}],"version-history":[{"count":0,"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/posts\/28709\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/media\/27685"}],"wp:attachment":[{"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/media?parent=28709"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/categories?post=28709"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.web-systems.pl\/en\/wp-json\/wp\/v2\/tags?post=28709"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}