RAG Query Rewriting in Production: HyDE, Multi-Query, and Step-Back Prompting (2026)

How to run HyDE, Multi-Query Retrieval, and Step-Back Prompting in a production RAG pipeline. Working Python code, latency and cost benchmarks, an adaptive routing classifier, and a RAGAS evaluation loop for picking the right rewriter per query.

RAG Query Rewriting: HyDE + Multi-Query (2026)

Updated: September 16, 2026

RAG query rewriting is the practice of transforming a user's raw question into one or more reformulated queries before hitting the retriever, so semantic search returns passages the original phrasing would've missed. The three techniques that actually move the needle in production are HyDE (Hypothetical Document Embeddings), Multi-Query Retrieval, and Step-Back Prompting. I've shipped all three across function-calling agents at a couple of employers, and honestly, the pattern is always the same: no single rewriter wins on every query type, so you route between them and evaluate each with RAGAS-style metrics before you ship.

  • HyDE generates a hypothetical answer, embeds it, and retrieves against that vector. Best for terse or keyword-poor questions.
  • Multi-Query fans out 3–5 paraphrases in parallel and fuses results with Reciprocal Rank Fusion (RRF). Best for ambiguous or under-specified queries.
  • Step-Back Prompting asks a broader question first, retrieves grounding facts, then answers the specific query. Best for compositional or reasoning-heavy questions.
  • All three add 200–900 ms of latency and 1–5x the token cost per turn; adaptive routing with a small classifier is how you keep the bill sane.
  • Evaluate rewriting with context recall and answer relevance on a labeled set. Don't ship on "vibes."
  • Combine rewriting with a good reranker; the two are complementary, not substitutes.

What is query rewriting in RAG?

Query rewriting in RAG is any transformation of the user's input question before it reaches the retriever, with the goal of narrowing the semantic gap between how users ask and how source documents are worded. Retrieval quality collapses when the user says "Why is my checkout slow?" but the docs say "Latency degradation in the payments pipeline." Embedding models are trained to bridge some of this gap, but they're not magic. Short, colloquial, or under-specified queries embed poorly, and the top-k passages you get back are noisy.

Rewriting fixes the input side of that equation. The three families of techniques you'll see in 2026 production stacks are: generative rewriting (HyDE, hypothetical documents), paraphrase expansion (Multi-Query with RRF), and abstraction (Step-Back Prompting, subquery decomposition). Each addresses a different failure mode, and together they cover ~85% of the retrieval-recall problems I see in customer-support and internal-knowledge RAG applications. Rewriting is orthogonal to chunking strategy and reranker choice. You tune all three together, and improvements compound.

HyDE: Hypothetical Document Embeddings

HyDE, introduced in the 2022 Precise Zero-Shot Dense Retrieval Without Relevance Labels paper by Gao et al., is the simplest rewriter to implement and often the most surprising in effect. Instead of embedding the user's question, you ask an LLM to generate a plausible hypothetical answer to the question, and then embed that answer as the query vector. Because passages in your corpus are themselves answer-shaped, the hypothetical answer sits closer to the right documents in embedding space than the original question does.

Here's a minimal, working HyDE implementation in Python using the OpenAI SDK and a Qdrant vector store. I keep the hypothetical-document prompt short on purpose, since long generations dilute the query signal.

from openai import OpenAI
from qdrant_client import QdrantClient

client = OpenAI()
qdrant = QdrantClient(url="http://localhost:6333")

HYDE_PROMPT = """You are a technical writer. Given the user's question,
write a short (2-3 sentence) passage that would plausibly appear in a
documentation article as the answer. Do not hedge. Do not say "I don't
know." Just write the passage as if it were fact."""

def hyde_retrieve(question: str, collection: str, k: int = 10):
    # 1. Generate a hypothetical answer
    hypo = client.chat.completions.create(
        model="gpt-4.1-mini",
        messages=[
            {"role": "system", "content": HYDE_PROMPT},
            {"role": "user", "content": question},
        ],
        temperature=0.0,
        max_tokens=180,
    ).choices[0].message.content

    # 2. Embed the hypothetical answer, NOT the question
    vector = client.embeddings.create(
        model="text-embedding-3-large",
        input=hypo,
    ).data[0].embedding

    # 3. Retrieve against that vector
    hits = qdrant.search(
        collection_name=collection,
        query_vector=vector,
        limit=k,
    )
    return hits, hypo

When HyDE wins: short questions ("what about timeouts?"), keyword-poor questions ("why doesn't it work?"), and cross-lingual retrieval where the query language differs from the corpus. When HyDE loses: highly specific queries containing rare identifiers (SKUs, error codes, product names). The LLM tends to smooth those away, and you lose the exact-match signal. In that case, keep the raw query in a parallel BM25 lane and fuse (see the hybrid search guide for the RRF setup).

Multi-Query Retrieval with RRF fusion

Multi-Query Retrieval, popularized by the LangChain MultiQueryRetriever, asks the LLM to rewrite the user's question into N paraphrases (typically 3–5), retrieves top-k for each in parallel, and fuses the ranked lists with Reciprocal Rank Fusion. The idea is that each paraphrase covers a slightly different semantic neighborhood, so their union has higher recall than any single query.

Here's a production-shaped implementation. Note the JSON-mode call so we get structured paraphrases back, so no fragile regex parsing.

import asyncio
from openai import AsyncOpenAI
from qdrant_client import AsyncQdrantClient

client = AsyncOpenAI()
qdrant = AsyncQdrantClient(url="http://localhost:6333")

REWRITE_SYSTEM = """Rewrite the user's question as 4 different paraphrases
that a search engine could match on. Vary vocabulary, specificity, and
framing. Return JSON: {"queries": ["...", "...", "...", "..."]}"""

async def multi_query_retrieve(question: str, collection: str, k: int = 10):
    resp = await client.chat.completions.create(
        model="gpt-4.1-mini",
        messages=[
            {"role": "system", "content": REWRITE_SYSTEM},
            {"role": "user", "content": question},
        ],
        response_format={"type": "json_object"},
        temperature=0.3,
    )
    import json
    queries = json.loads(resp.choices[0].message.content)["queries"]
    queries = [question] + queries  # always include the original

    # Embed all queries in one batch
    emb_resp = await client.embeddings.create(
        model="text-embedding-3-large",
        input=queries,
    )
    vectors = [d.embedding for d in emb_resp.data]

    # Search in parallel
    async def search_one(vec):
        return await qdrant.search(collection, query_vector=vec, limit=k)
    ranked_lists = await asyncio.gather(*(search_one(v) for v in vectors))

    # Reciprocal Rank Fusion (k_rrf=60 is the standard constant)
    scores: dict[str, float] = {}
    docs: dict[str, object] = {}
    for hits in ranked_lists:
        for rank, hit in enumerate(hits):
            doc_id = hit.id
            scores[doc_id] = scores.get(doc_id, 0.0) + 1 / (60 + rank)
            docs[doc_id] = hit
    fused = sorted(scores.items(), key=lambda kv: -kv[1])[:k]
    return [docs[doc_id] for doc_id, _ in fused]

Multi-Query shines when the query is ambiguous. "How do I handle failures?" could mean network failures, tool-call failures, or partial responses, and each paraphrase pulls a different slice. It also plays well with reranking: fuse 4 lists of 20 → rerank the top 40 → return top 10. That two-stage pipeline (fusion, then cross-encoder rerank) is the current default at most teams I talk to.

Step-Back Prompting for compositional queries

Step-Back Prompting, introduced by Google DeepMind in Take a Step Back (Zheng et al., 2023), is the technique for compositional questions where the specifics obscure the underlying concept. The rewriter asks a broader "step-back" question first ("What are the retry policies of the payments service?" instead of "Why did my POST /charge fail after 30 seconds?"), retrieves grounding passages, and then the generator answers the original narrow query using those broader passages as context.

Implementation is a two-step chain. In practice I do it as two model calls plus one retrieval, and I return the step-back question in the response payload so trace inspection is easy:

STEP_BACK_PROMPT = """Given a specific question, generate a more general
question whose answer would help you answer the specific one. Keep it
one sentence. Return only the general question, no preamble."""

def step_back_retrieve(question: str, collection: str, k: int = 8):
    general = client.chat.completions.create(
        model="gpt-4.1-mini",
        messages=[
            {"role": "system", "content": STEP_BACK_PROMPT},
            {"role": "user", "content": question},
        ],
        temperature=0.0,
        max_tokens=60,
    ).choices[0].message.content.strip()

    # Retrieve BOTH: passages for the specific and the general question
    specific_vec = embed(question)
    general_vec = embed(general)
    specific_hits = qdrant.search(collection, specific_vec, limit=k)
    general_hits = qdrant.search(collection, general_vec, limit=k)

    # Deduplicate, keep highest score
    combined = {h.id: h for h in general_hits}
    for h in specific_hits:
        combined[h.id] = h  # specific wins on ties
    return list(combined.values()), general

When Step-Back wins: reasoning questions ("Why does X behave differently under Y?"), policy questions ("Is behavior Z allowed?"), and any question where the direct query would only match a leaf-level passage that lacks the surrounding context. It's essentially cheap "context expansion" without the cost of parent-document indexing. When it loses: fact-lookup questions where the specific query is already the right query ("What is the retry limit for endpoint /charge?", no step-back needed).

Subquery decomposition for multi-hop questions

The fourth technique worth knowing is subquery decomposition, which sits between multi-query and step-back. Decomposition splits a compound question into ordered subquestions, retrieves for each, and passes all context into a single generation step. Where Multi-Query fans out paraphrases of the same question, decomposition fans out different questions that each answer one part of the compound.

Example: "Which of our EU customers on the Enterprise plan had failed renewals in Q2 2026, and what was the total ARR at risk?" decomposes into "Which EU customers are on the Enterprise plan?", "Which of those had failed renewals in Q2 2026?", and "What is the ARR of those customers?" You retrieve grounding docs for each and let the generator stitch. This is GraphRAG territory when the underlying data is relational. Often, you're better off routing to a text-to-SQL or graph traversal path than trying to make dense retrieval do the join.

Decomposition is expensive (N subquery generations, N retrievals, N context inserts), so I only enable it when a classifier flags the query as multi-hop. See the adaptive-routing section below.

HyDE vs Multi-Query vs Step-Back: when to use each

The three techniques address different failure modes, so the correct answer to "which should I use?" is usually "all three, routed by query type." Here's the comparison I keep in my notes:

Dimension HyDE Multi-Query Step-Back
Best forShort / keyword-poor queriesAmbiguous queriesCompositional / reasoning queries
Extra LLM calls11 (returns N paraphrases)1
Extra retrievals1 (replaces original)N + 12
Added latency (p50)200–400 ms400–700 ms250–450 ms
Token cost multiplier1.2–1.5x2–3x1.3–1.6x
Recall lift on ambiguous QA+4–8 pts+8–15 pts+3–6 pts
Hallucination riskHigh (steers by fake doc)LowMedium
Combines well with rerankerYesYes (essential)Yes

Numbers above are from my internal benchmarks on customer-support and internal-docs corpora; your mileage will vary with domain. The one universal is that none of these techniques replace a good reranker. They change what shows up in the top-50 candidates, and the reranker still has to sort them. If you're comparing rerankers, I wrote about the Cohere vs Voyage vs Jina vs BGE tradeoffs separately.

Adaptive query rewriting with a classifier

Running all three rewriters on every query is wasteful. In production I put a small, fast classifier in front (either a fine-tuned encoder at ~30 ms, or a low-temperature call to a cheap model at ~150 ms) that labels the incoming query as one of: direct, ambiguous, compositional, or lookup. That label picks the rewriter. Here's the classifier prompt and the routing table I use:

from openai import OpenAI
client = OpenAI()

CLASSIFIER_PROMPT = """Classify the user's query into exactly one label:
- direct: specific fact-lookup with clear terms (e.g., "What is the retry limit?")
- ambiguous: short or vague query with multiple plausible interpretations
- compositional: requires combining multiple facts or reasoning steps
- lookup: contains a specific identifier (SKU, error code, name) that must match exactly
Return JSON: {"label": "..."}"""

def classify(question: str) -> str:
    resp = client.chat.completions.create(
        model="gpt-4.1-nano",
        messages=[
            {"role": "system", "content": CLASSIFIER_PROMPT},
            {"role": "user", "content": question},
        ],
        response_format={"type": "json_object"},
        temperature=0.0,
        max_tokens=20,
    )
    import json
    return json.loads(resp.choices[0].message.content)["label"]

def route_and_retrieve(question: str, collection: str):
    label = classify(question)
    if label == "direct":
        return standard_retrieve(question, collection)
    if label == "ambiguous":
        return multi_query_retrieve(question, collection)
    if label == "compositional":
        return step_back_retrieve(question, collection)
    if label == "lookup":
        return hybrid_bm25_dense_retrieve(question, collection)
    return standard_retrieve(question, collection)

This is essentially a specialized form of semantic routing, applied to retrieval strategy instead of model choice. In my last two rollouts, adaptive routing shaved 30–40% of the rewriting cost with no measurable recall regression versus running Multi-Query on every query. The classifier itself has to be evaluated: I sample 300 queries a week, hand-label them, and check drift.

Evaluating query rewriting in production

You cannot ship any of this on intuition. Query rewriting either lifts retrieval quality or it doesn't, and the only way to know is with a labeled evaluation set and repeatable metrics. This is where the schema-first, evals-before-deploy discipline pays off. Treat every rewriter as a variant behind a config flag, and let the numbers pick the winner.

The three metrics I care about are:

  • Context recall: did the retrieved passages contain the information needed to answer? Compute against a golden set of (question, expected_passages) pairs.
  • Answer relevance: did the generator's answer address the question? RAGAS scores this well.
  • Faithfulness: is the answer grounded in retrieved context, or is the model making it up? Rises and falls with retrieval quality.

I run this as a nightly job with 200–500 labeled queries and dashboard the deltas per rewriter. RAGAS, TruLens, and DeepEval all support this; I compared the tradeoffs in the RAG evaluation framework guide. Whichever framework you pick, the key is to lock the eval set and versioning around it: if you change the corpus, the labels drift and your comparisons stop being valid.

Concretely, here's the eval loop I use for a rewriter A/B in RAGAS:

from ragas import evaluate
from ragas.metrics import context_recall, answer_relevancy, faithfulness
from datasets import Dataset

def build_eval_dataset(rewriter_fn, gold_qa_pairs):
    records = []
    for pair in gold_qa_pairs:
        docs, _ = rewriter_fn(pair["question"], "docs_collection")
        contexts = [d.payload["text"] for d in docs]
        answer = generate_answer(pair["question"], contexts)
        records.append({
            "question": pair["question"],
            "answer": answer,
            "contexts": contexts,
            "ground_truth": pair["ground_truth"],
        })
    return Dataset.from_list(records)

for name, fn in [("baseline", standard_retrieve),
                 ("hyde", hyde_retrieve),
                 ("multi_query", multi_query_retrieve),
                 ("step_back", step_back_retrieve)]:
    ds = build_eval_dataset(fn, GOLD_SET)
    result = evaluate(ds, metrics=[context_recall, answer_relevancy, faithfulness])
    print(name, result)

Latency and cost tradeoffs

Rewriting is not free. Every technique adds at least one LLM call before retrieval, and Multi-Query adds N retrievals on top. For a typical customer-support agent I've measured:

  • Baseline dense retrieval: p50 90 ms, p95 220 ms, ~$0.0004 per query
  • HyDE: p50 340 ms, p95 620 ms, ~$0.0009 per query
  • Multi-Query (4 paraphrases): p50 580 ms, p95 1.1 s, ~$0.0018 per query
  • Step-Back: p50 380 ms, p95 700 ms, ~$0.0011 per query

Two levers matter for cost control. First, use the smallest reasonable model for the rewriter. GPT-4.1-mini or Claude Haiku 4.5 are more than good enough. Second, cache aggressively: rewriter outputs for repeated or near-duplicate queries can be pulled from a semantic cache. My LLM cost optimization guide covers the caching layer in depth. On top of that, prompt caching on the system prompt of the rewriter itself typically cuts 40–60% of the input token spend, because the rewriter prompt is stable across queries.

Latency-wise, run the rewriter and the embedding of the original query in parallel where you can. If the classifier picks "direct" you already have the vector ready and can skip the wait. Small optimizations like this are the difference between a 300 ms and 900 ms retrieval budget.

Frequently Asked Questions

Does query rewriting always improve RAG accuracy?

No. Rewriting helps most on ambiguous or short queries but can hurt exact-match retrieval (SKUs, error codes, specific identifiers) because the LLM smooths them away. Always evaluate against a labeled set; I've seen HyDE drop accuracy 8–12 points on domains the base model doesn't know well.

What is the difference between HyDE and multi-query retrieval?

HyDE generates a single hypothetical answer and embeds that as the query vector, replacing the original. Multi-Query generates multiple paraphrases of the question and retrieves for each in parallel, fusing results with RRF. HyDE is cheaper; Multi-Query has higher recall on ambiguous queries.

When should you use step-back prompting?

Use step-back prompting for compositional or reasoning-heavy questions where the specifics obscure a broader concept, for example, debugging questions ("Why did X fail under Y?") or policy questions ("Is Z allowed?"). Skip it for direct fact-lookups; the extra call adds latency without benefit.

How much extra does query rewriting cost per query?

HyDE typically adds 20–50% to per-query LLM cost, Step-Back adds 30–60%, and Multi-Query with 4 paraphrases roughly doubles or triples it. Prompt caching on the rewriter system prompt and a small routing classifier can bring that back down by 40–60%.

Do I still need a reranker if I use query rewriting?

Yes. Rewriting changes what shows up in the top-50 candidates but doesn't sort them well. Cross-encoder rerankers are complementary. The standard production pipeline is rewrite → fetch top-50 → rerank → top-k. Skipping the reranker leaves 5–10 recall points on the table.

Daichi Watanabe
About the Author Daichi Watanabe

LLM integration specialist with a strong opinion about function calling and an even stronger one about evaluations.