Case Study 01 · Enterprise AI / RAG
AI Knowledge Assistant for Enterprises
Engineering an enterprise RAG system beyond the prototype.
An enterprise RAG application that allows users to upload documents, transforms them into searchable knowledge and generates grounded answers using semantic retrieval and an LLM.
01
The Engineering Problem
Enterprise knowledge is often distributed across documents and difficult to retrieve quickly. Traditional keyword search may not understand the meaning behind a user's question, while sending entire documents directly to an LLM is inefficient and difficult to ground reliably.
The engineering challenge was to design a system that could:
- 01Ingest documents
- 02Extract and chunk content
- 03Generate embeddings
- 04Store searchable vectors
- 05Retrieve context semantically
- 06Generate answers grounded in retrieved content
- 07Support secure, production-minded API behaviour
02
Solution Overview
One shared vector store connects two pipelines: ingestion and query.
03
High-Level Architecture
04
Production Engineering
The engineering concerns that surround the RAG core once it is exposed as an API.
Access & API contract
Responsiveness
Failing predictably
Tracing requests
Accountability & usage
Each group answers a question the RAG core cannot answer on its own: who is calling it, how quickly it responds, what happens when something fails, what happened to a given request, and what was used.
05
Security & Request Lifecycle
How a request passes the security checkpoints and cross-cutting concerns before it reaches the AI and vector services.
Cyan steps are the security checkpoints. The bands above and below wrap the whole request lifecycle; they are not separate business steps.
06
Testing & Engineering Quality
How the RAG core and its API are verified: components alone, components together, secured endpoints, and an automated build gate.
- 01
Unit Tests
Individual components in isolation
- 02
Integration Tests
Components working together
- 03
Authenticated API Testing
Secured endpoints exercised with valid credentials
- 04
CI Build Validation
Builds and tests validated automatically
07
Containerisation & Deployment
How the application is packaged, configured and moved between environments.
- Multi-stage Docker build, separating build from runtime
- Environment-based configuration, so the same application can target different environments
- Secrets kept outside source control
- Automated build validation through GitHub Actions
- Azure container deployment
This is the deployment engineering implemented and explored for the project. It is not a claim of production availability or a successful production deployment.
08
Key Engineering Decisions
Retrieval before generation
Retrieve relevant document context before calling the LLM.
Answers are grounded in enterprise knowledge rather than relying only on model knowledge.
Separate orchestration from infrastructure
Keep query orchestration separate from embedding, vector database and AI response services.
Responsibilities remain clear and components can evolve independently.
Treat production concerns as architecture
Make authentication, rate limiting, caching, logging, exception handling, health checks and auditability part of the system design.
They are designed in from the start rather than added at the end.
Externalise configuration and secrets
Keep environment-specific configuration and secrets outside source control.
The same application can move safely between environments.
09
Architecture Trade-offs
Each decision has a cost. These are the trade-offs in the system as built, not claims about scale or performance.
Qdrant as the shared vector store
One purpose-built store serves both pipelines: ingestion writes embeddings into it and retrieval searches it semantically.
It is another component to run, secure and keep consistent with the documents that were ingested. Retrieval quality still depends on chunking and embeddings, which the store cannot fix.
In-memory caching
It needs no additional infrastructure and is simple to add inside the application.
Cached results live inside one application process: they are not shared across instances and disappear on restart. Anything cached must also stay consistent with its source.
SQLite for the conversation store
A file-based database needs no separate database server, which keeps the project simple to run and test.
It suits a single application instance. Multiple instances or heavier concurrent use would call for a server database; that has not been demonstrated here.
Separated service boundaries
Query orchestration is kept apart from the embedding, vector database and AI response services, so each can be tested and replaced on its own.
More interfaces and wiring to design and keep consistent, and each request passes through more layers. This separates responsibilities in the codebase; it does not by itself imply separately deployed services.
10
What I Learned
Building the project reinforced these lessons.
- 01
Retrieval quality is decided by the whole pipeline — extraction, chunking, embeddings and search together — not by the LLM alone. A weak early step limits every answer that follows.
- 02
Once security, observability, caching, failure handling and usage tracking are added, a RAG application stops being only a retrieval problem and becomes a broader software architecture problem.
- 03
Clear boundaries around embedding, vector storage and the AI response service make each easier to test and replace — at the cost of more interfaces to maintain.
- 04
Configuration, secrets and the build-to-container path need to be considered early when moving an AI prototype beyond a demo.
11
What's Next
Future improvements. These are not capabilities that exist today.
- FutureRetrieval evaluation
- FutureImproved citation/grounding experience
- FutureMore advanced document processing
- FutureProduction deployment hardening