In the first article, I talked about giving the Major Incident Manager an intelligent operating team. (www.imadlodhi.com/article/agentic-operations-rethinking-major-incident-management)

Agents that can correlate events, investigate changes, track actions, maintain timelines, draft communications and help prepare the Major Incident Review.

Sounds great.

But there’s a problem.

What exactly are we asking those agents to operate inside?

Because after more than 25 years at IBM, another three years at CGI, and working with more than 250 clients across over 40 countries, one thing I’ve learned is that technology rarely fixes a broken operating model.

Sometimes technology just helps you execute the dysfunction faster.

This is why I think the conversation about Agentic Operations has to go beyond agents, models and automation.

We have to talk about management systems.

Seven systems. One operating architecture.

In The 7 Essential First-Line Management Systems, I describe seven connected systems: Roles & Responsibilities, Processes & Procedures, Technology & Tools, Meetings, Reporting & Measurements, Analytics & Optimization, and Continual Service Improvement.

I don’t see these as seven independent checklists. They form an operating architecture.

Roles answer: who owns it?

Processes answer: how is it done?

Technology answers: what enables it?

Meetings answer: how is it governed and coordinated?

Measurements answer: what happened?

Analytics answer: why did it happen and what patterns matter?

Continual improvement answers: what changes next?

Now put Major Incident Management in the middle of that architecture and introduce intelligent agents.

This is where things get interesting.

1. Roles & Responsibilities — Who owns what?

A major incident is probably one of the worst times to discover that decision rights are unclear.

Who declares the major incident? Who is the Incident Commander? Who owns technical recovery? Who approves a risky recovery action? Who communicates with executives? Who communicates with customers? Who decides that service is stable enough to close the incident?

Now add agents.

What can an agent recommend? What can it execute? What requires human approval? Can it trigger an escalation? Can it initiate an approved runbook? Can it send an external communication?

If those boundaries are unclear for people, they will be unclear for agents.

And an unclear role does not remain an administrative problem. It becomes a delayed decision. During a major incident, a delayed decision can become a longer outage, an SLA miss and eventually customer, financial or reputational impact.

2. Processes & Procedures — How is the work done?

An agent needs more than a vague instruction to “help manage the incident.”

There has to be an operating process underneath it.

Detection. Validation. Declaration. Escalation. Diagnosis. Communication. Recovery. Service validation. Closure. Major Incident Review. Problem Management.

If every shift handles those stages differently, what exactly are we automating?

Agentic Operations can make a good process faster, more consistent and more responsive. But automating an inconsistent process can simply make inconsistency happen at machine speed.

3. Technology & Tools — What enables the work?

This is usually where the Agentic AI conversation starts.

I think it should be the third conversation, not the first.

The agent needs access to the ITSM platform, monitoring, observability, service models, configuration information, knowledge, change records, collaboration tools and automation capabilities.

But integration alone isn’t enough.

If the service relationships are unreliable, the knowledge is outdated, change records are incomplete or different tools contain conflicting versions of reality, the agent inherits those weaknesses.

The tool should enable the operating model. The operating model should not be invented around whatever the newest tool happens to do.

4. Meetings — How is it governed?

Then we get to the major incident bridge.

I’ve been on enough of these to know how quickly they can become crowded.

Forty people join. Eight are actively working the problem. Ten are waiting for somebody to ask them something. Another group is listening because their management asked them to join. Meanwhile, technical specialists are trying to diagnose a production outage while repeatedly being asked for updates.

That is not necessarily collaboration. Sometimes it is simply congestion with a conference link.

An agentic operating model can change this.

Agents can track actions, maintain the timeline, capture decisions, chase dependencies, surface unanswered questions and prepare status summaries without constantly interrupting the people restoring the service.

But the bridge still needs governance. Somebody still owns the meeting. Somebody controls the decision flow. Somebody decides who needs to be there.

The goal isn’t to automate the bridge. It is to make the bridge a better decision environment.

5. Reporting & Measurements — What happened?

MTTR matters. But it doesn’t tell the whole story.

How long did detection take? How long before the incident was declared? How long before the correct technical teams were engaged? Where did escalation stall? Were communications delivered when promised? How much business impact occurred? How often is the same service experiencing major incidents?

Agents can capture far more operational evidence than humans can reasonably record during a crisis.

That gives us the opportunity to measure the management of the incident, not just the duration of the outage.

But the measures still have to be trusted and decision-relevant. More data does not automatically create better management.

6. Analytics & Optimization — Why did it happen?

This is where Agentic Operations becomes particularly powerful.

One incident tells us what happened today. Hundreds of incidents can tell us something about the system.

Which services repeatedly fail after particular changes? Which dependencies appear across multiple incidents? Which escalation paths consistently introduce delay? Which recovery actions work fastest? Which alerts repeatedly appear before a major incident? Which incidents that looked unrelated actually share the same underlying pattern?

Humans are good at judgment and context. Agents are very good at searching large volumes of operational evidence and surfacing relationships worth investigating.

That combination matters.

7. Continual Service Improvement — What changes next?

This may be the system I worry about most.

We restore service. Everyone is exhausted. We conduct the Major Incident Review. We identify twelve actions.

Then everyone goes back to work.

Three months later, we’re on another bridge discussing something remarkably familiar.

An agent can help identify corrective actions, connect recurring incidents, track commitments and flag improvements that are aging without action.

But an agent cannot create organizational commitment where none exists.

Someone still has to own the improvement. Capacity has to be allocated. Progress has to be governed. And eventually we need evidence that the change actually reduced the risk.

A neglected improvement doesn’t disappear. It becomes tomorrow’s recurring problem.

This is the architecture underneath Agentic Operations

This is why I don’t see Agentic Operations as another management system.

I see it as an intelligence and orchestration capability operating across the seven systems.

The management systems provide structure and discipline.

The agents provide intelligence, correlation, coordination and execution.

Human leaders provide judgment, authority and accountability.

And all three matter.

Put sophisticated agents into an ad hoc operation and you still have an ad hoc operation—just with better technology.

Put them into a defined, controlled and continually improving management architecture and now we have something much more powerful.

Before asking whether your organization is ready for Agentic Operations, ask whether your management systems are ready for agents.

That may be the more important readiness assessment.