RAG for Your Own Runbooks: Retrieval That Actually Answers
Chunking, hybrid search and metadata filters decide whether your internal assistant quotes the right page or invents one — the retrieval layer matters more than the model.
Every team eventually asks for "ChatGPT but for our docs." What they usually get is a demo that answers three questions impressively and four questions with quiet confidence and the wrong page.
The model is the least interesting part of that system. If you use a capable hosted model and a mediocre retrieval layer, you will get fluent wrong answers. If you use a plain model and retrieve the right paragraph, you get a useful tool. Almost all of the quality lives in retrieval.
Start with what you already have
Before buying anything, look at the corpus. Most internal knowledge bases are a graveyard: three wikis, a Confluence nobody has edited since 2022, runbooks in a repo, and the real knowledge in Slack.
Two things matter here. First, recency — a runbook that was correct last year will actively hurt you if it outranks the current one. Second, authority — the repo's runbook and a random comment in a thread should not score the same.
If your source material is stale, retrieval will surface the staleness at scale. Fix the corpus before you fix the pipeline.
Chunk on boundaries, not on numbers
The classic advice is "chunk to ~500 tokens with 50 overlap." It produces fragments that start mid-sentence and end mid-thought, which is exactly how you get answers that quote half a paragraph and miss the caveat in the next one.
Split on structure instead: headings, sections, whole steps in a procedure, whole functions in code. A chunk should be something a human would consider a complete answer unit. When you must fall back to a fixed size, prefer breaking at paragraph boundaries and keep the parent heading attached as context.
For code and runbooks, preserve the step numbers. Retrieval that returns steps 3 and 5 of a seven-step recovery procedure is worse than no retrieval at all.
Hybrid search is the boring default that works
Pure vector search fails on the queries internal docs actually receive: error codes, service names, ticket IDs, EC2 exceeded. Embeddings are bad at exact identifiers. Pure keyword search fails on the paraphrased ones: "how do we roll back the payments service" when the page says "revert deployment."
Do both and merge the results. Keyword search nails identifiers, vector search handles phrasing, and a simple reciprocal-rank fusion is enough — you don't need a clever weighting scheme on day one.
Add metadata filters on top: product, environment, team, last-updated. "How do I rotate the production database credentials" should never return the staging page.
Metadata is the cheapest quality win
The single highest-leverage change is usually not in the model or the retriever — it's a frontmatter field.
Tag each document with owner, environment, and reviewed-on date. Then filter and sort by those fields at query time. A result set limited to documents owned by the right team and reviewed in the last six months is dramatically better than the same set ranked by cosine similarity alone.
This is also where you get an honest answer for the questions the system can't answer: if nothing relevant ranks above your threshold, say so and link to a human. Invented answers are a retrieval failure dressed up as a model failure.
Ask retrieval questions, not model questions
When quality is bad, resist the urge to swap models. Take ten real questions where it answered badly and, for each one, ask a different question: did we retrieve the right chunk into the context window?
- Right chunk present, wrong answer → prompt or model issue. Rare.
- Right chunk present but ranked fourth → ranking issue. Tune weights, add metadata.
- Right chunk absent → chunking or corpus issue. Almost always.
Log the retrieved chunk IDs alongside every answer. Without that, you are debugging by vibes — which is precisely the state an eval harness exists to end.
What I would build first
A weekly job that re-embeds changed documents, a threshold below which the system declines to answer, retrieval IDs logged with every response, and a reviewed-on field someone is actually accountable for. That covers the failures people notice.
What I'd skip: fine-tuning the embedding model on your corpus, a reranker before you've looked at your top-1 failures, and a chat interface on top of documentation that nobody could find via search either. The tool doesn't fix the corpus.
Summary
Chunk on structural boundaries, run hybrid search, attach metadata you can filter on, and log what was retrieved. Get retrieval right and a modest model will seem smart; get it wrong and no model will save you — you'll just get a more articulate wrong answer.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →