Step 1: Confirm the incident
Validate impact, severity, owners, and the active response policy.
Build an evidence-backed incident timeline and keep response owners aligned as the situation changes.
Connect alerts, deployments, tickets, runbooks, and operator actions into a live investigation record. The agent separates observations from hypotheses and routes the next approved diagnostic action. The result is a current investigation brief and coordinated response, ready for the accountable team to review and move forward. By keeping source-backed facts, open questions, and ownership together, the team can make the next decision without rebuilding context by hand.
Sequence alerts, changes, tickets, deployments, and human actions with timestamps.
Keep observed facts distinct from possible causes and missing evidence.
Route decisions, owners, and status updates without taking production authority.
Prepared output: Checkout latency · investigation timeline
Source state: Current approved records retained
A deployment correlation is visible; production rollback remains with the service owner.
Validate impact, severity, owners, and the active response policy.
Connect telemetry, changes, tickets, and operator actions.
Show evidence, possible causes, and the next approved diagnostics.
Route tasks, decisions, and updates to the accountable owners.
Return evidence, actions, current state, and follow-up work.
Put this agent to work