A RAG demo is easy to build and easy to be impressed by. A production system has to keep working when documents change, users have different permissions, dependencies fail and people begin to rely on the answers. This article looks at what changes between the two.
A demo shows the pipeline can work
In a demo you load a few well-behaved documents, embed them, ask questions you already know the answers to, and the results look good. Production adds varied and messy content, concurrent users, content that changes, permissions, and users who act on what they are told.
Retrieval quality is the centre of gravity
Most poor answers trace back to missing or irrelevant context. Chunk size and overlap, whether chunks respect structure such as headings and tables, the embedding model, metadata filters, the number of results retrieved and optional re-ranking or hybrid keyword-plus-semantic search all influence what the model sees. These choices are tunable, and they need to be measured rather than assumed.
Grounding and answer behaviour
The model should be instructed to answer from the supplied context, to say so when the context is insufficient, and to point back to its sources. Conversation history and retrieved content should be kept distinct, so that earlier turns do not quietly override evidence.
Security belongs in the retrieval path
The API needs authentication, but authentication alone is not enough. Authorisation has to apply to what is retrieved: a user must only receive context from content they are permitted to see, because a model cannot unsee what is placed in its prompt. Rate limiting, audit logging, and keeping secrets and environment configuration outside source control complete the picture.
Observability and caching
A correlation identifier that follows a request through retrieval and generation makes problems traceable. Logging which chunks were retrieved, rather than only the final answer, is what makes quality issues diagnosable. Caching can reduce repeated cost and latency for embeddings and frequent queries, provided that it is invalidated when the underlying content changes.
Evaluate retrieval separately from generation
Build a curated set of questions with known source passages. First check whether the right content is retrieved, then whether the answer is faithful to it. Re-run the set whenever chunking, embeddings or prompts change. Model-based judging can help scale this up, but it needs its own scrutiny.
Operations
Content changes, so re-indexing and deletion handling matter. Changing the embedding model means re-embedding the corpus. Health checks should cover dependencies such as the vector store and the model provider, and the system needs defined behaviour for timeouts and provider rate limits.
| Concern | Typical demo | Production-oriented system |
|---|---|---|
| Chunking | Fixed-size splits | Chosen and tested against the content structure |
| Retrieval | Top-k similarity | Metadata filters, tuned parameters, possibly re-ranking |
| Security | Open access | Authentication plus permission-aware retrieval |
| Observability | Final answer only | Correlated traces, retrieved sources, usage tracking |
| Evaluation | Spot checks | Repeatable question sets for retrieval and answers |
| Operations | One-off load | Re-indexing, deletion, health checks, failure handling |
Key takeaways
- Most RAG failures are retrieval failures, so measure retrieval separately.
- Authorisation must apply to retrieved content, not only to the API.
- Log what was retrieved, not just what was answered.
- Production readiness is largely operations: change, failure and re-indexing.