Agentic Operations: Rethinking Major Incident Management
I've spent more than 25 years at IBM, another three years at CGI, and worked with more than 250 clients across over 40 countries.
I've been around a lot of major incidents.
And if you've spent any amount of time in IT operations, you know the scene.
A critical service goes down. The bridge opens. People start joining. The Service Desk is getting hammered. Application, database, middleware, server and network teams are trying to figure out whether the problem is theirs. Somebody asks what changed. Somebody else asks when the service will be back.
Then the executives arrive.
Now everybody wants an ETA.
The technical teams are trying to troubleshoot. The Major Incident Manager is trying to maintain control. Someone is taking notes. Someone is preparing communications. Another person is chasing a team that should have been on the bridge fifteen minutes ago.
And somewhere in all of this, we're supposed to actually restore the service.
I've seen incredibly good teams manage major incidents well. I've also seen very capable technical people struggle because the management systems around them weren't working.
Because a major incident isn't just a technical event. It's a stress test of your entire management system.
Your roles and responsibilities get tested. Your processes and procedures get tested. Your technology and tools get tested. Your meetings and governance get tested. Your reporting and measurements get tested. Your ability to analyze information gets tested. And after the service is restored, your continual improvement system gets tested.
Those happen to be the seven management systems I wrote about in The 7 Essential First-Line Management Systems.
And now we're introducing something incredibly powerful into that environment: Agentic Operations.
Don't start with automation
The temptation is going to be to ask, “What parts of Major Incident Management can AI automate?”
I think that's the wrong question.
The better question is: What would Major Incident Management look like if intelligent agents could work alongside the Major Incident Manager throughout the incident lifecycle?
Think about what actually happens during a major incident.
Before the bridge is even established, an agent could correlate monitoring events, user reports, service dependencies and recent changes to help determine impact and potential blast radius.
Once the incident is declared, another agent could maintain the timeline automatically. Another could track actions, owners and commitments. A diagnostic agent could search telemetry, known errors, previous incidents, configuration relationships and knowledge articles while technical specialists continue their investigation.
And then there's the question that seems to appear on almost every major incident bridge I've ever been on: “What changed?”
An agent doesn't have to wait for somebody to manually search through change records. It can continuously correlate recent changes against affected services and infrastructure and surface plausible relationships for the technical teams to investigate.
The Major Incident Manager doesn't disappear. The Major Incident Manager gets an intelligent operating team.
Communication becomes part of the operating model
One of the biggest pressures during a major incident is communication.
The technical teams need space to diagnose. The business needs to understand impact. Executives want to know what is happening. The Service Desk needs something useful to tell users. Customers may need updates. And everybody wants to know when the service will be restored.
An agent can help maintain a common factual picture and prepare different communications for different audiences. The technical update doesn't have to be the executive update. The executive update doesn't have to be the Service Desk message.
But this is where human accountability matters.
An agent can assemble evidence, identify patterns, draft communications and recommend actions. That doesn't mean it should be allowed to invent an ETA, execute every recovery action, or communicate externally without appropriate authority.
The human Major Incident Manager remains in command.
Recovery is not the end
This is another lesson I've learned over the years: restoring service is only one part of Major Incident Management.
Once everybody sees green again, there is enormous pressure to move on. People have been on a bridge for hours. The business is running again. Everyone is tired.
But the evidence from that incident is incredibly valuable.
What happened? What signals appeared first? What hypotheses were tested? Which actions worked? Which didn't? Where did escalation stall? Were communications effective? Did a previous incident show the same pattern?
Today, much of that has to be reconstructed after the fact.
In an agentic model, much of the evidence can already be captured and organized as the incident unfolds. That gives the Major Incident Review and Problem Management processes something far more useful than people's recollection of a stressful three-hour bridge.
Agentic Operations should not simply help us restore service faster. It should help us understand why we needed to restore it in the first place.
But there's a catch
None of these agents operate in a vacuum.
If nobody knows who has authority to declare a major incident, an agent doesn't fix that.
If your escalation process is unclear, automating it doesn't make it good.
If your service relationships and operational data are unreliable, an agent will reason from unreliable information.
If your Major Incident Reviews produce actions nobody owns, generating better actions won't create accountability.
And if Problem Management is something everybody talks about but nobody has time to do, the same incident will eventually bring everyone back onto the bridge.
This is where Agentic Operations and management systems come together.
Roles and responsibilities. Processes and procedures. Technology and tools. Meetings. Reporting and measurements. Analytics and optimization. Continual service improvement.
Agentic Operations doesn't replace those systems. It operates across them.
And that is where I want to go next.
Because before an organization asks whether it is ready for Agentic Operations, there may be a more fundamental question worth asking:
Are the management systems underneath the agents actually ready for them?



