Most enterprise AI conversations start with a prompt and a model. The architecture that makes those systems useful inside an organisation looks quite different. This article traces how the architecture changes as capabilities are added, namely retrieval, tools, reasoning loops, state and governance, and what each step costs in engineering terms.
The starting point: prompt and response
A basic LLM application sends a prompt and returns a completion. Everything the model knows comes from its training and from the prompt itself. For enterprise use, three limits appear quickly: the model cannot see private or recent information, it cannot check its claims against a source, and it can only produce text. Each later architectural step addresses one of these limits, and each introduces a concern of its own.
Retrieval: grounding answers in enterprise knowledge
Retrieval-augmented generation (RAG) retrieves relevant content before the model generates a response, and asks the model to answer from that content. The architecture gains an ingestion pipeline (extract, chunk, embed, index), a retrieval step and a prompt that carries the retrieved context. Answers can now be traced to sources and can reflect current data.
What becomes harder is diagnosis. Quality now depends on retrieval, so a wrong answer may come from chunking, the embedding model or ranking rather than from the LLM. Access control also moves into the retrieval path: the model will faithfully use whatever context it is given.
Tools: letting the model request actions
With tool (function) calling, the model is given descriptions of available tools and can respond with a structured request to call one. The model does not execute anything. The application does. That boundary is where input validation, permissions and logging belong.
Reasoning loops and state
Many tasks need several steps, because each result changes what should happen next. An orchestration layer runs a loop: the model proposes a step, the application executes a tool, the result is returned to the model, and this repeats until a stopping condition is met. That loop needs state: what has been tried, what evidence exists and what the current goal is.
Explicit state makes workflows resumable, auditable and testable. It also creates new failure modes: loops that never terminate, state that drifts from the real system, and cost that grows with every step. Step limits, budgets and explicit completion criteria are part of the design, not afterthoughts.
Governance: separating capability from authority
Once tools can change things, the question shifts from whether the agent can propose an action to whether it is allowed to take it. A governance layer answers that with risk assessment, deterministic policy, human approval where warranted, and an audit trail. It is deterministic on purpose: authority should not depend only on what a model happens to generate.
| Stage | What it adds | New engineering concern |
|---|---|---|
| Prompt and response | Language understanding and generation | No access to private data; no source verification |
| + Retrieval (RAG) | Grounding in enterprise knowledge | Retrieval quality, access control, chunking, evaluation |
| + Tools | Ability to request queries and actions | Validation, permissions, side effects |
| + Reasoning and state | Multi-step workflows | Termination, state integrity, cost, testing |
| + Governance | Controlled authority | Risk and policy design, approval flows, audit |
Each stage builds on the previous one. Most systems do not need every stage.
Trade-offs
Every stage adds capability and surface area. Many problems are well served by retrieval plus a fixed workflow. Reaching for autonomy where a deterministic workflow would do adds latency, cost and risk without a matching benefit. Add stages when the problem demands them, and be explicit about why.
Key takeaways
- Each architectural stage solves a specific limit and introduces a new engineering concern.
- In RAG, retrieval quality is as important as the model.
- The model proposes tool calls; the application executes them, and that boundary is where controls belong.
- Governance separates what an agent can propose from what it is allowed to do.