Prompt Techniques

Rewriting the question itself before retrieval runs, to catch what a single phrasing would miss: multi-query, RAG-Fusion, decomposition, and step-back prompting.

01

Multi-query and RAG-Fusion

one question
rewrite 1
rewrite 2
rewrite 3
union of all retrieved chunks

A single phrasing of a question might just miss the chunk that actually answers it, not because retrieval failed, but because the question and the passage happen to use different words for the same idea. Multi-query prompting has an LLM reword the question three or four different ways, then retrieves separately for each rewrite.

The simplest way to combine the results is a plain union of everything retrieved across all the rewrites. RAG-Fusion goes one step further and ranks that combined set with Reciprocal Rank Fusion, so a chunk that shows up across several rewordings outranks one that only appeared under a single phrasing, treating repeated appearance as a vote of relevance.

Neither technique is free: generating the rewrites costs one extra LLM call, and retrieving separately for each one multiplies the number of retrieval calls by however many rewrites got generated. It's a real cost paid to catch questions a single phrasing would have missed entirely.

02

Decomposition: sub-questions in sequence or in parallel

one compound question
sub-Q1
sub-Q2
sub-Q3

Some questions are really several questions stacked together, and decomposition starts by having an LLM break a compound question into ordered sub-questions instead of trying to retrieve for the whole thing at once.

When later sub-questions actually depend on earlier ones, they get solved in a chain: answer sub-question one, carry that answer forward as context into the retrieval and prompt for sub-question two, and so on, each step building on what came before.

When the sub-questions don't actually depend on each other, that chain is unnecessary overhead. Solving them independently in parallel and concatenating the results is faster and reaches the same place, the dependency between the sub-questions is what decides which approach fits.

03

Step-back prompting: asking a broader question first

original: "What is task decomposition for LLM agents?"

A narrow question, "what is task decomposition for LLM agents," can retrieve a narrow, isolated passage that answers the literal question but gives the model no surrounding context to reason with.

Step-back prompting has an LLM generate a more abstract version of the same question first, "what is the process of task decomposition," produced from a few worked examples of narrow-to-general rewrites, then retrieves for that broader question separately.

Both the narrow and the step-back contexts get retrieved independently and merged into one final prompt, so the model answers the specific question with the benefit of the general background around it, not just the one isolated fact.

Try it: which technique fits?

A retrieval scenario, and four query translation techniques to choose from. Guess which one actually fits before checking.

Same fan-out situation, several reworded queries retrieving in parallel, but now you specifically want the chunks that keep showing up across multiple rewordings pushed to the top of the final list.

Which query translation technique fits?

Waiting for your guess.