Case Study 02 · Agentic AI / Autonomous Systems

Autonomous Production Incident Resolution Agent

From incident detection to governed remediation — with evidence, policy and human oversight.

An Agentic AI system that investigates production incidents by gathering evidence, forming hypotheses, evaluating investigation readiness and proposing remediation — while enforcing risk, policy and human-approval boundaries before action.

.NET 8C#Semantic KernelASP.NET CoreBlazorLLMTool Calling

01

The Engineering Problem

Production incidents require engineers to correlate multiple signals such as:

Application HealthLogsDeploymentsDatabase MetricsOperational Knowledge

The challenge is not simply detecting that something failed. The system must gather sufficient evidence, reason about possible causes and determine whether it has enough information before recommending an action.

More importantly, an autonomous system should not automatically execute every action it proposes.

The engineering challenge was therefore

How can an AI agent investigate autonomously while keeping risky production actions governed and auditable?

02

Agent Investigation Loop

Iterative reasoning with a readiness gate before any remediation.

ObserveGather EvidenceForm HypothesesEvaluate EvidenceReadinessGateINVESTIGATION_INCOMPLETEGather More EvidenceRe-evaluateinsufficientiterate until evidence is sufficientsufficient evidenceREADY_FOR_REMEDIATION

The loop repeats until the readiness gate is satisfied. Only then does the agent move on to remediation.

03

Evidence & Tool Architecture

The LLM reasons. Semantic Kernel orchestrates the tools that gather evidence.

ReasonsLLM / Agent ReasoningThe model that performs the reasoning
OrchestratesSemantic Kernel / Tool OrchestrationOrchestration and tool integration. It is not the AI model.
EvidenceTools & Evidence Sources
Health
Application Logs
Deployment History
Database Metrics
Infrastructure Health
Incident Runbook / Operational Knowledge

04

Hypothesis-Driven Investigation

The agent keeps several candidate explanations open instead of committing to one cause immediately. Evidence strengthens or weakens each until there is enough to proceed.

Candidate explanations

Application regression

▲ strengthened by evidence▼ weakened by evidence

Database infrastructure saturation

▲ strengthened by evidence▼ weakened by evidence

Connection / resource exhaustion

▲ strengthened by evidence▼ weakened by evidence
Is there sufficient evidence to proceed?

Illustrative examples of candidate explanations. No probabilities, scores or metrics are implied.

05

Governance Boundary

Where autonomy ends and authority begins.

Proposal & evaluationControlled action
Proposed Remediation
Risk Assessment
Policy Evaluationdecides what is allowed
Authority boundaryHuman Approval Required
Execute
Validate Recovery
The agent determines what it wants to do.Policy determines what it is allowed to do.
  • Evidence-based reasoning
  • Readiness gates
  • Deterministic policy controls
  • Risk classification
  • Human-in-the-loop approval
  • Post-remediation validation

06

Incident Walkthrough — Orders API

One incident traced from first signal to validated recovery.

  1. Detect
  2. Investigate
  3. Hypothesise
  4. Readiness
  5. Propose
  6. Govern
  7. Approve
  8. Execute
  9. Validate
Observed Evidence
Agent Reasoning
Deterministic Policy
Human Decision
Tool Execution
01

Incident Detected

Detect

The Orders API begins returning HTTP 500 errors.

HTTP 500
37% error rate
Database connection timeout
The agent begins investigation rather than immediately proposing remediation.
02

Evidence Gathering

Investigate

The agent invokes tools to gather evidence.

Application Health
Application Logs
Recent Deployments
Database Metrics
Database Infrastructure Health
Incident Runbook
Evidence reveals
  • Orders API is unhealthy.
  • Logs show delays acquiring database connections.
  • Version 2.4.1 was recently deployed.
  • Database connections reached 100 / 100.
Agent inference, not directly observed
The recent deployment of version 2.4.1 may be related to the failure.
03

Hypothesis Evaluation

Hypothesise
H1 — Application regression introduced by version 2.4.1
Leading hypothesis
▲ Deployment correlation▲ Logs▲ Database behaviour
Supporting evidence increases as each signal is gathered.
Still under evaluation
H2 — Database infrastructure saturation
H3 — Connection / resource exhaustion
04

Readiness Gate

Readiness
INVESTIGATION_INCOMPLETE
Additional evidence required
Database Infrastructure Health
Incident Runbook
Re-evaluate readiness
READY_FOR_REMEDIATION

The agent did not yet have enough evidence to safely recommend remediation, so it gathered additional database infrastructure and runbook evidence.

The agent is required to prove investigation readiness before crossing into remediation.
05

Proposed Remediation

Propose · Govern
ROLLBACK Orders API
2.4.1 → 2.4.0
What the agent wants to do
Risk assessment
Medium — Score 30
Policy decision
HumanApprovalRequired
What the agent is allowed to do
06

Human Approval & Execution

Approve · Execute
WAITING_FOR_APPROVAL
Human Approval
APPROVED
Execute Rollback

After explicit human approval, the rollback tool is permitted to execute. The LLM does not change production directly. The approved action runs through the controlled tool boundary.

07

Recovery Validation

Validate
Orders API: HTTP 200
Error rate: 0%
Database connections: 34 / 100
Version: 2.4.0

The agent validates recovery before declaring the incident resolved.

Incident Resolved

07

Key Agentic Architecture Decisions

01

Evidence before action

Evidence → Action
Decision

The agent must gather and evaluate evidence before proposing remediation.

Why it matters

Reasoning should be grounded in observable system state rather than assumptions.

02

Readiness before remediation

Investigate | Gate | Remediate
Decision

Investigation and remediation are separate phases.

Why it matters

A readiness gate prevents the agent from moving into action until sufficient evidence has been collected.

03

Separate reasoning from authority

Reasoning ≠ Authority
Decision

The LLM can reason about what action may be appropriate, but it does not determine whether that action is authorised.

Why it matters

Risk and policy controls define the agent's authority.

04

Validate after execution

Execute → Validate
Decision

The agent must re-check application health, error rate, infrastructure state and deployed version before closing the incident.

Why it matters

Successful tool execution does not automatically mean the incident is resolved.

08

What This Demonstrates

ReasonTool-using AI agentsEvidence-grounded reasoningHypothesis-driven investigation
GovernReadiness gatesRisk-based governanceDeterministic policy controlsHuman-in-the-loop approval
Act & verifyControlled tool executionPost-action validation

09

What I Learned

Building the project reinforced these lessons.

  1. 01

    Agent autonomy needs clearly defined authority boundaries.

  2. 02

    LLM reasoning and deterministic policy controls solve different problems and should remain separate.

  3. 03

    Evidence quality is as important as reasoning quality.

  4. 04

    Human approval is most useful when placed at meaningful risk boundaries rather than around every agent step.

  5. 05

    Autonomous remediation needs validation and recovery checks, not just successful tool execution.

Building autonomous systems means designing both intelligence and control.

Explore more of my work across Agentic AI, enterprise AI and modern software architecture.