A RAG chatbot is a chatbot that searches your own approved documents for the passages relevant to a question, then has a language model write the answer from those passages. RAG stands for retrieval augmented generation: retrieve first, then generate. Benian Technologies, an AI implementation partner, builds them for businesses whose staff keep answering the same questions by hand.
The money problem is simple. When a customer asks about your return window at 9pm and gets no answer, or gets a confident wrong one, you lose the sale or you pay for it later in a refund, a chargeback or a support thread. A plain AI chatbot answers from whatever it learned in training, which knows nothing about your policies. A RAG chatbot answers from the policy itself and can show which document it used.
This page explains how retrieval works, how a RAG chatbot differs from an LLM chatbot and an older NLP intent bot, what preparing documents involves, how answers stay tied to sources, and what can still go wrong. It also covers when a simple FAQ page or a scripted menu is the better choice.
How a RAG chatbot finds the answer: sources, chunks and search
It starts with sources: the documents you approve. Typical ones are shipping and return policies, product spec sheets, help articles, warranty terms, service area rules and internal procedures. Each document is split into chunks, short passages of a few paragraphs, because a model answers better from the three passages that matter than from a whole manual.
Each chunk is converted into an embedding, a list of numbers that represents its meaning, and stored in a search index. When a visitor types a question, the question is converted the same way and the index returns the chunks closest in meaning. Many builds also run a keyword search alongside it, because exact terms such as a part number or a state name are easy for meaning search to miss.
The retrieved chunks go to the language model with instructions: answer only from these passages, cite which ones you used, and say you do not know if they do not cover the question. The model writes a normal sentence in the visitor's own terms. That last step is the generation part, and it is the only part the visitor sees.
RAG chatbot versus LLM chatbot versus NLP intent bot
An NLP intent bot is the older design. Natural language processing in a chatbot of this kind means the software classifies each message into a predefined intent, such as track order or reset password, and returns a scripted reply. It is predictable and never invents a policy, but every question needs an intent someone wrote, and anything outside the list gets a fallback message.
An LLM chatbot is a large language model answering on its own, which is what people mean when they compare a bot to ChatGPT. It handles any phrasing and writes fluently, but its knowledge is general and frozen at training time. Ask it your restocking fee and it will either decline or guess. A guess that sounds right is worse than no answer.
A RAG chatbot uses the same kind of model but changes where the facts come from. The language understanding is the model's; the facts are yours, pulled at the moment of the question. Update the return policy document and the next answer reflects it, with no retraining. For most business question answering this is the design that fits, and an intent menu still beats it for short fixed processes like collecting a return reason.
Preparing approved documents is most of the work
Answer quality depends more on the documents than on the model. Before developing a chatbot, the useful exercise is to collect the twenty or thirty questions customers actually ask, from your inbox, chat logs and phone notes, and check whether a current document answers each one. Usually several do not, and someone has to write the missing answer once.
Common problems: two versions of the same policy that disagree, a PDF that is a scanned image with no readable text, pricing tables that change weekly, and documents written for staff that should not be shown to customers. Each source needs an owner and a decision: public, internal only, or excluded. Expired promotions and old terms should be removed, not left for the chatbot to find.
Some questions should never be answered from documents at all, such as a specific order status or a refund decision. Those need a live lookup in your store or system with its own permissions, or a handoff to a person.
Source references and refusing to guess
A good RAG chatbot shows where an answer came from, for example a link to the shipping policy section it used. That lets a customer check, and it lets your team audit a bad answer in minutes by seeing which passage misled it.
Refusal is a feature. When retrieval returns nothing relevant, the right reply is that the bot does not have that information, followed by a handoff with the question attached. Benian's Chat AI answer from approved material with source references, and Benian configures and tests handoffs to your team. RAG reduces made up answers but does not end them: the model can misread a passage, merge two passages, or retrieve an outdated one. AI can still make mistakes, so anything with legal, medical or financial weight needs human review.
Developing and testing a RAG chatbot, then keeping it current
Creating an AI chatbot on company documents follows a short sequence: confirm source access, channels such as your website, WhatsApp, Slack or Teams, and languages; agree acceptance checks; build the index and instructions; then test. Testing means a written set of real questions with the expected answer and source for each, including questions the bot should refuse and questions that must go to a person. Run the set before launch and again after every document change.
After launch, measure what the bot could not answer, how often visitors asked for a human, which sources get cited most, and a weekly sample of transcripts read by someone who knows the business. Unanswered questions are a list of documents to write. Keeping the knowledge current is a routine, not a project: when a policy changes, the document changes, and the index refreshes.
What drives cost: the number and condition of sources, how many channels and languages, whether answers need live lookups in other systems, how strict the handoff rules are, and how much testing the risk of a wrong answer justifies. Benian publishes no price; each build is scoped after we see the documents.
An example, and when not to build one
VOT Distribution, a multi-brand e-commerce distributor, has two AI storefront assistants in production built by Benian. One is the assistant at shopfreezo.com, which answers product, compliance and shipping questions around the clock. Those are the questions a RAG chatbot suits: high volume, answerable from written material, and costly when answered wrong.
Do not build one if customers ask fewer than a handful of questions a week, if your answers live in someone's head rather than in documents, or if most questions need a judgment call. Start smaller: write a clear FAQ page from the questions you already get. If that page answers most of them, you may not need a chatbot yet.