Techsplaining 101: The Anatomy of a Major Incident
We take an in-depth look at the process for declaring, managing, communicating, and resolving a Major Incident
Techsplaining 101: The Anatomy of a Major Incident
We take an in-depth look at the process for declaring, managing, communicating, and resolving a Major Incident
Have you ever wondered what happens during a technology outage at Miami? Who gets involved, who gets notified, and how does the problem get resolved?
Let’s take a look at the process for declaring, managing, communicating, and resolving a Major Incident.
Pit stop: Helpful vocabulary
In order to talk about Major Incidents, we first need to know what that is and means, and that requires a brief diversion into the world of IT Service Management. IT Services as a division at Miami is governed by both Miami-decreed policies and industry-standard practices, and one of the key industry standards we maintain is the ITIL (Information Technology Infrastructure Library) framework. This is a standardized set of practices that help us manage incidents, deliver services, and generally serve the university community effectively. The ITIL is generally accepted as a best-practice, technology-agnostic guideline for managing an IT service.
Now, let’s talk about Major Incidents. According to ITIL, “Major Incidents cause serious interruptions of business activities and must be solved with greater urgency.” In other words, something important is broken and we have to fix it quickly.
At Miami, we measure what is a “Major Incident” based on the urgency of the situation and the criticality of the business process (often taking into account how many people are impacted). Think of what it takes to make Miami tick at any given time. (For a handy cheat sheet, we have a list of “Business Critical Applications” on our IT Status page.) If Canvas goes down, we can’t serve our customers (students and faculty). If email is on the fritz, all kinds of business processes stop. If myMiami experiences issues, folks can’t navigate to the other Miami services they depend on myMiami for.
Suffice to say: Major Incidents are a big deal, and we take them seriously.
Anatomy of a Major Incident
Now that you know what a Major Incident is, let’s dive into the process that IT Services follows when a technology service we support goes down.
Start of an incident: Trouble is afoot
Someone calls IT Help to get assistance with an issue. Someone posts in Slack or Google Chat that they’re having trouble. Someone emails their favorite IT guru to ask if something is wrong with a certain program. A system administrator will notice something shifty in a log when they’re trying to make a change.
When these things happen irrespective of each other, it’s business as usual. When they happen at the same time or one after the other, and everything points to the same service – that’s when we know that something may not be copacetic.
Declaration of a Major Incident
Once enough of a red flag has been raised, someone declares a Major Incident. Any of our technology partners, IT Services employees, and help desk advisors have access to this big red MI button.
As a sidebar, we have long operated under the idea that it’s better to declare an MI and have it end up being a non-event than for a legitimate emergency to go undiscovered.
All hands on deck!
When an MI is declared, this is the trigger to solver teams that they need to set other work aside and dig into the problem. Depending on the service impacted, an MI Coordinator determines the teams that need to be involved and which manager will need to run the ship. The MI Coordinator also looks at all the tickets coming in to IT Help and sees if there are any that relate to the current outage.
Now, sometimes it’s not so simple to figure out which service is impacted. Sometimes, it’s the technologies underpinning things, so it looks like one thing is having a problem but really it’s the platform beneath that. For instance, if someone can’t log in to email, is it an email problem, or is it a problem with our single sign-on service? Or is it an issue with Duo? There are many layers to any issue, and the initial part of incident management is determining what the problem actually is before it may be solved.
Communications: Initial and ongoing
A big part of a major outage is making sure our customers are aware that a) we know there is a problem, and b) we are working on it. There are a few different audiences that receive information about outages, depending on the severity of the problem. Our technical partners across the university get the first notice, along with information about what users may experience. This is mostly because we know that users talk to their local IT person first!
Again, depending on the severity of the issue, these status updates (which are all posted to MiamiOH.edu/ITStatus) will happen on a regular cadence until the issue is resolved. Sometimes that means every 30 minutes, and sometimes it’s more infrequent.
Working the incident
While the MI Coordinator is working on tagging tickets and the comms team is putting together the notice for the university community, the solver teams get to work. They perform general troubleshooting activities and investigate where and when the problem is happening.
For example, if folks can’t log in, they may find that our single sign-on services are not functioning properly. In this case, this may require a restart of one or more servers, or it might involve merging new code (depending on which SSO service is impacted – just in case you thought that might be simple). Troubleshooting can involve one or more teams, one or more technicians, and one or more managers to help coordinate tasks and groups.
Resolution
Once solver teams are on the case, they focus on the problem until the service has been restored. In the best-case scenario, this doesn’t take very long – our teams are very good at what they do, and troubleshooting is easier when you are familiar with the systems. Sometimes, the issues persist and take longer to resolve. These longer incidents may require a changing of the guard and tagging in of other teammates, if this lasts longer than the workday.
All-clear Notice
Once the technical team thinks they have figured out the problem and implemented a resolution, an official declaration of “all clear” is next. A sufficient amount of testing goes into making sure we have a stable service, and once that is complete, the MI Coordinator and IT manager involved in the incident come to a decision about the all-clear notice. The IT Communications team sends out the notice, and the MI comes to a close.
After-action Report
After every MI, involved teams fill out reports, and for certain circumstances, an after-action meeting takes place. These meetings are opportunities for the interested parties to come together, discuss the events, and determine whether there are necessary improvements to be made to the infrastructure or processes that broke down during the MI.
Quick note about cloud-based services
Many of our most important systems are cloud-based, meaning that they are managed by other companies and do not have physical machinery on Miami’s campus. These services include things like Canvas, Workday, and Google Workspace apps. When these services are impacted, if our initial investigation finds that nothing is to be done on Miami’s side, we send notifications via the regular channels mentioned above and keep an eye on the service’s status page. (For instance, here is Canvas’s status site.)