Generation

The last, deliberately simple stage: retrieved chunks and a question go into a prompt template, the template goes to the LLM, and a parsed answer comes out the other side.

01

Filling the prompt template

retrieved chunks
question

Generation starts by slotting two things into a prompt template: the chunks retrieval just returned, and the user's original question. Nothing new happens here computationally, it's just assembling the actual text that's about to be sent to the model.

A typical template reads something like "Answer the question based on this context: {context} Question: {question}," with retrieval's output and the question dropped straight into those placeholders. The instruction to answer only from the given context is what keeps the model grounded instead of free-associating from its training data.

02

The chain: prompt to grounded answer

prompt -> LLM -> output parser -> grounded answer

The filled prompt goes to the LLM, and its response passes through an output parser, typically just extracting plain text, before it's returned as the final answer. Three stages, prompt, model, parser, chained together in a fixed sequence.

This chain is deliberately the simplest possible link in the whole pipeline. Every more advanced RAG technique, better chunking, smarter retrieval, a rewritten query, changes one of the stages upstream of this one; the generation step itself rarely needs to get more complicated to benefit from all of that upstream work.