Skip to main content

Incident Management (Outage & Degradation)

Technology services sometimes fail or stop functioning normally.

When that happens, ISS must do more than troubleshoot the technical problem. We must identify the impact, coordinate the appropriate people, communicate clearly, restore acceptable service, and learn from the event.

Incident Management is the framework ISS uses to respond to significant service interruptions, including service degradations and outages.

The ISS Incident Management framework is:

Initiate → Identify → Communicate → Resolve

The Purpose of Incident Management

Incident Management helps ISS respond consistently when a service is disrupted.

The objectives are to restore acceptable service as quickly as practical while:

  • Establishing clear ownership
  • Bringing the appropriate technical resources together
  • Coordinating the response
  • Communicating with customers and College leadership
  • Documenting what occurred
  • Identifying necessary follow-up
  • Learning from significant incidents

The framework is intended to help the response, not slow it down.

Service Degradation

A service degradation occurs when a service remains available but is not functioning normally.

Examples may include significantly slower performance, intermittent failures, limited functionality, reduced capacity, or important functions being unavailable while other portions of the service continue to operate.

Customers may still be able to work, but their normal experience or ability to perform College functions is impaired.

A degradation may still have significant impact.

Outage

An outage occurs when a service is unavailable or is so severely degraded that normal business functions cannot be performed.

An outage may affect an entire service, a large group of people, a location, multiple departments, or a critical College operation.

Within the ISS Hierarchy of Priority, an outage is 0 — Outage, the highest-priority category of work.

Degradations and Outages Can Change

The condition of a service may change during an incident.

A degradation may worsen and become an outage.

An outage may be partially restored and become a degradation while technical work continues.

The incident team should describe the current condition of the service rather than relying solely on how the incident was originally classified.

When to Initiate Incident Management

Incident Management should be initiated when an issue appears to affect more than an isolated customer or represents a meaningful interruption to a service.

Examples may include:

An incident may be identified through the Help Desk, monitoring, another ISS employee, a vendor, or another College department.

Employees do not need to know the cause before initiating the process.

Planned Maintenance Can Become an Incident

Planned maintenance begins as scheduled Build or Sustainment work.

If the maintenance unexpectedly causes a service degradation or outage, the work transitions into Incident Management.

At that point, restoring acceptable service and communicating the service condition become the immediate priorities.

The team may continue the change, pause it, roll it back, or take another recovery action depending on the circumstances.


Initiate

When an employee identifies a likely service degradation or outage, the employee should notify the ISS team by posting a message in the "Technical Outage 1" channel in Microsoft Teams.

The initial notification should contain as much information as is reasonably known, such as:

  • Service affected
  • What is happening
  • When the issue began
  • Who appears to be affected
  • Known impact
  • Actions already being taken
  • Estimated restoration time, if known
  • Related ticket number, if available
  • A tag to the relevant team and service owner, if known.

Information may be incomplete.

The purpose of the initial notification is to make the issue visible and engage the people needed to respond. The existing Incident Management framework similarly calls for early notification containing the service, issue, start time, affected population, known actions, estimated resolution, and ticket information when available.

Identify

Each service should have established responsibility that includes:

Service Owner → Technical Lead → Backup Technical Lead → Incident Coordinator → Incident Coordinator Backup

During normal business hours, the Service Owner or designated Technical Lead or Backup should take ownership of the technical response within 15 minutes of notification.

The responsible technical team should:

  • Confirm that the issue is being investigated.
  • Verify the service condition.
  • Determine the impact.
  • Identify whether the condition is a degradation or outage.
  • Engage additional resources when needed.
  • Begin restoration activities.
  • Provide technical information to the Incident Coordinator.

The source Incident Management framework establishes the same 15-minute ownership expectation during normal business hours.

Responding to an Outage

Employees responsible for an affected service should stop lower-priority work and respond to an outage when they serve as the Service Owner, Technical Lead, Backup, or another assigned incident-response resource.

Other employees may be asked to assist.

An outage involving a Critical service may require broad ISS participation until acceptable service is restored.

Managers remain responsible for coordinating resources and may deliberately assign some employees to continue other essential work while the incident team handles the outage.

Communicate

Communication is part of Incident Management.

Customers and College employees should not have to determine for themselves whether ISS knows a service is having problems.

The Incident Coordinator is responsible for coordinating incident communications. The Incident Coordinator should be identified as soon as possible. Any ISS Lead can serve as an Incident Coordinator. But, it is always best for the person already identified to assume this duty if possible.

The Incident Coordinator should serve as the primary coordination point, help technical staff obtain needed resources, prepare understandable service-status information, keep the CITO informed, coordinate communication with College leadership when appropriate, and maintain regular updates.

For significant incidents, updates should normally occur at least hourly, even when there has been no meaningful change. The established framework specifically requires hourly problem and service-status updates unless leadership requests a different frequency.

An update that says there has been no significant change can still be useful. Silence creates uncertainty.

The Incident Coordinator Is Not the Technical Resolver

The Incident Coordinator should generally not be one of the employees actively troubleshooting the technical problem.

This separation is intentional.

Technical employees need to concentrate on restoring the service.

The Incident Coordinator concentrates on communication, coordination, escalation, resource needs, and keeping others informed.

The existing framework deliberately separates these responsibilities so the Incident Coordinator can support the technical team rather than becoming another technical resolver.

Communicate What We Know

Incident communication should clearly distinguish:

  • What is known
  • What has been confirmed
  • What is being investigated
  • What is estimated
  • What remains unknown

Pressure for an answer should not cause ISS to communicate speculation as fact.

It is appropriate to say that the cause or restoration time is not yet known.

Communicate for the Audience

Customer-facing incident communication should explain what people need to know, including:

  • Service affected
  • What users are experiencing
  • Who is affected
  • Whether a workaround exists
  • What ISS is doing
  • Expected restoration time, if reasonably known
  • When another update will be provided

Avoid unnecessary technical jargon.

Security-sensitive incidents may require communication to be restricted to appropriate ISS employees and College leadership.

When public or media communication is necessary, ISS coordinates with the appropriate College Marketing and Communications representatives and College leadership.

ISS employees should not independently respond to media requests regarding an incident.

Resolve

The immediate objective of Incident Management is to restore an acceptable level of service.

Restoration may involve:

  • Correcting the underlying issue
  • Restoring a previous configuration
  • Rolling back a change
  • Failing over to another system
  • Implementing a workaround
  • Restoring from backup
  • Engaging a vendor
  • Taking another appropriate recovery action

A temporary workaround may restore service before the underlying cause is fully understood.

That is acceptable.

Incident Management focuses first on restoring service. Investigation of the underlying cause may continue after immediate service has been restored.

Verify Restoration

Before declaring an incident resolved, the responsible team should verify that:

  • The service is functioning at an acceptable level.
  • Important service functions are available.
  • Major dependencies are working.
  • Monitoring appears normal.
  • Any temporary workaround is understood.
  • Customers have been informed when appropriate.

If the service remains impaired, the incident may have moved from an outage to a degradation rather than being fully resolved.

Document the Incident

Significant incidents should have an appropriate operational record.

Documentation should capture, as appropriate:

  • What occurred
  • When the incident began
  • Services affected
  • Customer impact
  • Significant technical actions
  • Communications
  • Restoration time
  • Known or suspected cause
  • Workarounds
  • Outstanding issues
  • Required follow-up

Documentation supports future troubleshooting, Problem Management, retrospectives, and service improvement.

After Restoration

Restoration ends the immediate incident response, but it may not end the work.

Follow-up may include:

  • Additional monitoring
  • Permanent corrective work
  • Problem Management
  • Documentation updates
  • Monitoring improvements
  • A change
  • Sustainment work
  • A project
  • A retrospective

Significant outages and degradations should be scheduled for a retrospective within 30 days so the team can review what happened and identify steps that may reduce recurrence.


Connection to Our Values

People Matter because service interruptions affect people's ability to learn, teach, and work.

Communication Matters because people need useful information when the technology they depend upon is not working normally.

Integrity Matters because we communicate what we know, acknowledge what we do not know, and accurately describe what occurred.

Excellence Matters because effective Incident Management restores service while creating opportunities to make the service and our response better.