Case Study 01 · Enterprise AI / RAG

AI Knowledge Assistant for Enterprises

Engineering an enterprise RAG system beyond the prototype.

An enterprise RAG application that allows users to upload documents, transforms them into searchable knowledge and generates grounded answers using semantic retrieval and an LLM.

.NET 8C#OpenAIRAGQdrantAzureDockerGitHub Actions

01

The Engineering Problem

Enterprise knowledge is often distributed across documents and difficult to retrieve quickly. Traditional keyword search may not understand the meaning behind a user's question, while sending entire documents directly to an LLM is inefficient and difficult to ground reliably.

The engineering challenge was to design a system that could:

  1. 01Ingest documents
  2. 02Extract and chunk content
  3. 03Generate embeddings
  4. 04Store searchable vectors
  5. 05Retrieve context semantically
  6. 06Generate answers grounded in retrieved content
  7. 07Support secure, production-minded API behaviour

02

Solution Overview

One shared vector store connects two pipelines: ingestion and query.

03

High-Level Architecture

Application / services layerExternal / dataClient / APIASP.NET Core Web APIDocument ProcessingQuery OrchestrationAI Response ServiceEmbedding ServiceVector Database ServiceConversation StoreOpenAIQdrantSQLite

04

Production Engineering

The engineering concerns that surround the RAG core once it is exposed as an API.

RAG coreIngestion · Retrieval · Generation

Access & API contract

JWT AuthenticationAPI VersioningRate Limiting

Responsiveness

Memory CachingSemantic Retrieval

Failing predictably

Health ChecksGlobal Exception HandlingProblemDetails

Tracing requests

Correlation IDsRequest LoggingStructured Logging

Accountability & usage

Audit LogsUsage Tracking

Each group answers a question the RAG core cannot answer on its own: who is calling it, how quickly it responds, what happens when something fails, what happened to a given request, and what was used.

05

Security & Request Lifecycle

How a request passes the security checkpoints and cross-cutting concerns before it reaches the AI and vector services.

Cross-cuttingCorrelation IDLogging
Client Request
Rate Limiting
JWT Authentication
API Controller
Application Service
AI / Vector Services
Response
Cross-cuttingException HandlingAudit

Cyan steps are the security checkpoints. The bands above and below wrap the whole request lifecycle; they are not separate business steps.

06

Testing & Engineering Quality

How the RAG core and its API are verified: components alone, components together, secured endpoints, and an automated build gate.

  1. 01

    Unit Tests

    Individual components in isolation

  2. 02

    Integration Tests

    Components working together

  3. 03

    Authenticated API Testing

    Secured endpoints exercised with valid credentials

  4. 04

    CI Build Validation

    Builds and tests validated automatically

CI flow
Git Push
GitHub Actions
Restore
Build
Test
Docker Build

07

Containerisation & Deployment

How the application is packaged, configured and moved between environments.

Delivery path
Application
Docker Image
CI Pipeline
Container DeploymentAzure container deployment
  • Multi-stage Docker build, separating build from runtime
  • Environment-based configuration, so the same application can target different environments
  • Secrets kept outside source control
  • Automated build validation through GitHub Actions
  • Azure container deployment

This is the deployment engineering implemented and explored for the project. It is not a claim of production availability or a successful production deployment.

08

Key Engineering Decisions

01

Retrieval before generation

Decision

Retrieve relevant document context before calling the LLM.

Why it matters

Answers are grounded in enterprise knowledge rather than relying only on model knowledge.

02

Separate orchestration from infrastructure

Decision

Keep query orchestration separate from embedding, vector database and AI response services.

Why it matters

Responsibilities remain clear and components can evolve independently.

03

Treat production concerns as architecture

Decision

Make authentication, rate limiting, caching, logging, exception handling, health checks and auditability part of the system design.

Why it matters

They are designed in from the start rather than added at the end.

04

Externalise configuration and secrets

Decision

Keep environment-specific configuration and secrets outside source control.

Why it matters

The same application can move safely between environments.

09

Architecture Trade-offs

Each decision has a cost. These are the trade-offs in the system as built, not claims about scale or performance.

01

Qdrant as the shared vector store

Decision and benefit

One purpose-built store serves both pipelines: ingestion writes embeddings into it and retrieval searches it semantically.

Cost or limitation

It is another component to run, secure and keep consistent with the documents that were ingested. Retrieval quality still depends on chunking and embeddings, which the store cannot fix.

02

In-memory caching

Decision and benefit

It needs no additional infrastructure and is simple to add inside the application.

Cost or limitation

Cached results live inside one application process: they are not shared across instances and disappear on restart. Anything cached must also stay consistent with its source.

03

SQLite for the conversation store

Decision and benefit

A file-based database needs no separate database server, which keeps the project simple to run and test.

Cost or limitation

It suits a single application instance. Multiple instances or heavier concurrent use would call for a server database; that has not been demonstrated here.

04

Separated service boundaries

Decision and benefit

Query orchestration is kept apart from the embedding, vector database and AI response services, so each can be tested and replaced on its own.

Cost or limitation

More interfaces and wiring to design and keep consistent, and each request passes through more layers. This separates responsibilities in the codebase; it does not by itself imply separately deployed services.

10

What I Learned

Building the project reinforced these lessons.

  1. 01

    Retrieval quality is decided by the whole pipeline — extraction, chunking, embeddings and search together — not by the LLM alone. A weak early step limits every answer that follows.

  2. 02

    Once security, observability, caching, failure handling and usage tracking are added, a RAG application stops being only a retrieval problem and becomes a broader software architecture problem.

  3. 03

    Clear boundaries around embedding, vector storage and the AI response service make each easier to test and replace — at the cost of more interfaces to maintain.

  4. 04

    Configuration, secrets and the build-to-container path need to be considered early when moving an AI prototype beyond a demo.

11

What's Next

Future improvements. These are not capabilities that exist today.

  • FutureRetrieval evaluation
  • FutureImproved citation/grounding experience
  • FutureMore advanced document processing
  • FutureProduction deployment hardening

Interested in how I approach intelligent system design?