The people on the bridge should be making recovery decisions—not spending the outage assembling the information required to make them.
I have been thinking about Major Incident Management through an Agentic Operations lens.
Most major-incident processes already know what good looks like. Mobilize quickly. Understand impact. Bring in the right technical teams. Check recent changes. Maintain a chronology. Keep executives informed. Search known issues. Coordinate recovery. Capture enough evidence for Problem Management afterward.
None of that is new.
Yet during a real outage, humans still spend an extraordinary amount of time doing something machines are increasingly capable of doing much faster: assembling context.
A critical service goes down. The bridge opens. Someone asks which applications are affected. Another person checks the CMDB. Someone else searches recent changes. The database team says its platform looks healthy. Network is checking telemetry. The Major Incident Manager is trying to capture actions while also asking who owns the next one.
Ten minutes disappear.
Then twenty.
The organization has experts on the call. The problem is that the evidence they need is scattered across monitoring, ITSM, change records, configuration data, knowledge articles, chat, historical incidents and the heads of experienced people.
The bridge becomes the place where the organization reconstructs its own operational memory while the customer waits.
This is where Agentic Operations gets interesting.
Imagine an agent that begins working the moment a potential major incident is detected. Before the bridge is fully assembled, it correlates alerts, maps the affected business service and dependencies, checks recent changes and releases, compares the pattern with previous incidents, identifies likely resolver groups, opens the chronology, and highlights what is known, what is assumed, and what still needs validation.
As recovery continues, it listens for decisions and actions, tracks ownership, flags stalled tasks, updates the incident record, prepares stakeholder summaries and preserves the evidence Problem Management will need later.
Now the Major Incident Manager is not acting as a human clipboard.
The technical teams are not repeatedly being asked for information already available somewhere else.
The humans can concentrate on the difficult part: judgment.
Do we fail over? Roll back the change? Isolate the component? Accept temporary degradation? Invoke disaster recovery? What business risk does each option create?
Those decisions still need accountable people—especially when the action is high-impact or difficult to reverse.
That is why I would not design the agent as an unrestricted incident commander. I would design bounded authority. Let it gather, correlate, document, recommend and execute predefined reversible actions. Require human authorization when customer impact, security, contractual exposure or irreversible change crosses an agreed threshold.
The goal is not to remove people from the major-incident bridge.
It is to stop using expensive human judgment for work that is primarily searching, correlating, chasing, copying and remembering.
The bridge should remain the place where accountability lives.
It just should not have to be the brain that reconstructs the entire environment every time something breaks.

