Skip to main content

Retrospectives and Lessons Learned

ISS believes that important work should produce learning.

A retrospective is a structured review of an incident, outage, project, maintenance activity, or other significant event after the immediate work is complete.

The purpose is not to assign blame.

The purpose is to understand what happened, what worked, what did not work, and what we should do differently in the future.

When a Retrospective Is Required

A retrospective should be conducted following significant service outages or degradations and all Major projects.

The Incident Management framework requires the retrospective to be scheduled within 30 days of service restoration.

Retrospectives may also be useful following:

  • Significant projects
  • High-risk changes
  • Maintenance that did not proceed as planned
  • Major security events
  • Repeated incidents
  • Complex implementations
  • Work that revealed an important process weakness
  • Work that went particularly well and should become a repeatable practice

Not every ticket or routine activity requires a retrospective.

Managers should use judgment based on impact and learning value.

Focus on the Process

Retrospectives should focus on understanding the event rather than finding someone to blame.

Useful questions include:

  • What did we expect to happen?
  • What actually happened?
  • When did we know there was a problem?
  • How was the issue detected?
  • What was the impact?
  • What worked well?
  • What slowed us down?
  • What information was missing?
  • Were responsibilities clear?
  • Did communication work?
  • Did documentation help?
  • Did monitoring work?
  • Was a rollback or recovery plan available?
  • What should we change?
  • What should we continue doing?

People make decisions using the information and conditions available to them at the time.

The retrospective should help improve those conditions.

Establish a Timeline

For significant incidents, developing a basic timeline can help the team understand the event.

The timeline may include:

  • Initial event
  • Detection
  • First internal notification
  • Technical response
  • Customer communication
  • Major troubleshooting actions
  • Escalations
  • Vendor involvement
  • Workarounds
  • Restoration
  • Final communication

The purpose is not to create a minute-by-minute transcript unless that detail is useful.

The purpose is to understand what happened and identify opportunities to improve.

Identify Contributing Factors

Not every incident has one simple root cause.

Contributing factors may include:

  • Technology failure
  • Configuration
  • Process weakness
  • Missing documentation
  • Lack of monitoring
  • Incomplete testing
  • Training gaps
  • Communication failures
  • Vendor issues
  • Capacity limitations
  • Technical debt
  • Single-person dependencies
  • Unclear ownership
  • Timing
  • Multiple factors occurring together

The goal is to understand the conditions that allowed the problem to occur or made recovery more difficult.

Turn Lessons Into Actions

A retrospective is useful only if important lessons result in action.

Corrective or improvement actions should identify:

  • What needs to be done
  • Who owns the action
  • When it should be completed

Actions may result in:

  • A ticket
  • Sustainment work
  • A change
  • Updated documentation
  • New monitoring
  • Training
  • Process changes
  • Configuration changes
  • A project
  • Service redesign
  • Vendor follow-up

Actions should enter the normal ISS work-management process rather than remain only in meeting notes.

Publish the Retrospective Record

Significant outages and degradations should have a written retrospective record or note.

The record should capture enough information for ISS to understand:

  • What happened
  • Impact
  • How service was restored
  • Important lessons
  • Corrective actions

Sensitive information should be handled appropriately.

The retrospective record should be maintained in the appropriate authoritative documentation location.

Closing the Weekly Status Item

Outages and degradations remain visible in the weekly status report until:

  • The retrospective has been completed.
  • The retrospective record has been published.
  • Appropriate corrective work has been identified.

This keeps significant service failures visible long enough for the organization to learn from them rather than simply moving on once the service returns.

Lessons Learned Beyond Incidents

The same philosophy applies to projects.

Projects already include Lessons Learned as part of the ISS Project Methodology.

The question is the same:

What did this experience teach us that should make the next one better?

Learning should feed back into Plan, Build, and Operate.


Connection to Our Values

People Matter because retrospectives should improve services without creating a culture of blame.

Inclusiveness Matters because different people involved in an event may have seen different parts of what happened.

Communication Matters because meaningful learning requires open discussion and useful documentation.

Integrity Matters because improvement requires us to examine what actually happened, including mistakes and unexpected outcomes.

Excellence Matters because we intentionally use experience to make future work better.