Case Study 02 · Agentic AI / Autonomous Systems
Autonomous Production Incident Resolution Agent
From incident detection to governed remediation — with evidence, policy and human oversight.
An Agentic AI system that investigates production incidents by gathering evidence, forming hypotheses, evaluating investigation readiness and proposing remediation — while enforcing risk, policy and human-approval boundaries before action.
01
The Engineering Problem
Production incidents require engineers to correlate multiple signals such as:
The challenge is not simply detecting that something failed. The system must gather sufficient evidence, reason about possible causes and determine whether it has enough information before recommending an action.
More importantly, an autonomous system should not automatically execute every action it proposes.
How can an AI agent investigate autonomously while keeping risky production actions governed and auditable?
02
Agent Investigation Loop
Iterative reasoning with a readiness gate before any remediation.
The loop repeats until the readiness gate is satisfied. Only then does the agent move on to remediation.
03
Evidence & Tool Architecture
The LLM reasons. Semantic Kernel orchestrates the tools that gather evidence.
04
Hypothesis-Driven Investigation
The agent keeps several candidate explanations open instead of committing to one cause immediately. Evidence strengthens or weakens each until there is enough to proceed.
Application regression
Database infrastructure saturation
Connection / resource exhaustion
Illustrative examples of candidate explanations. No probabilities, scores or metrics are implied.
05
Governance Boundary
Where autonomy ends and authority begins.
The agent determines what it wants to do.Policy determines what it is allowed to do.
- Evidence-based reasoning
- Readiness gates
- Deterministic policy controls
- Risk classification
- Human-in-the-loop approval
- Post-remediation validation
06
Incident Walkthrough — Orders API
One incident traced from first signal to validated recovery.
- Detect
- Investigate
- Hypothesise
- Readiness
- Propose
- Govern
- Approve
- Execute
- Validate
Incident Detected
DetectThe Orders API begins returning HTTP 500 errors.
Evidence Gathering
InvestigateThe agent invokes tools to gather evidence.
- Orders API is unhealthy.
- Logs show delays acquiring database connections.
- Version 2.4.1 was recently deployed.
- Database connections reached 100 / 100.
Hypothesis Evaluation
HypothesiseLeading hypothesis
Readiness Gate
ReadinessThe agent did not yet have enough evidence to safely recommend remediation, so it gathered additional database infrastructure and runbook evidence.
The agent is required to prove investigation readiness before crossing into remediation.
Proposed Remediation
Propose · Govern2.4.1 → 2.4.0Medium — Score 30
HumanApprovalRequired
Human Approval & Execution
Approve · ExecuteAfter explicit human approval, the rollback tool is permitted to execute. The LLM does not change production directly. The approved action runs through the controlled tool boundary.
Recovery Validation
ValidateThe agent validates recovery before declaring the incident resolved.
07
Key Agentic Architecture Decisions
Evidence before action
Evidence → ActionThe agent must gather and evaluate evidence before proposing remediation.
Reasoning should be grounded in observable system state rather than assumptions.
Readiness before remediation
Investigate | Gate | RemediateInvestigation and remediation are separate phases.
A readiness gate prevents the agent from moving into action until sufficient evidence has been collected.
Separate reasoning from authority
Reasoning ≠ AuthorityThe LLM can reason about what action may be appropriate, but it does not determine whether that action is authorised.
Risk and policy controls define the agent's authority.
Validate after execution
Execute → ValidateThe agent must re-check application health, error rate, infrastructure state and deployed version before closing the incident.
Successful tool execution does not automatically mean the incident is resolved.
08
What This Demonstrates
09
What I Learned
Building the project reinforced these lessons.
- 01
Agent autonomy needs clearly defined authority boundaries.
- 02
LLM reasoning and deterministic policy controls solve different problems and should remain separate.
- 03
Evidence quality is as important as reasoning quality.
- 04
Human approval is most useful when placed at meaningful risk boundaries rather than around every agent step.
- 05
Autonomous remediation needs validation and recovery checks, not just successful tool execution.
Building autonomous systems means designing both intelligence and control.
Explore more of my work across Agentic AI, enterprise AI and modern software architecture.