For a Site Reliability Engineer (SRE), making the right decisions under pressure is crucial. With multiple competing priorities, how do you decide which tasks will have the most impact on system reliability, user experience, and business outcomes?
This course, part of a comprehensive curriculum, helped SRE teams master the critical concept of rank-ordering actions, teaching them how to prioritize tasks that drive the highest value for system performance and user satisfaction.
The Challenge
SRE teams are constantly faced with a flow of incidents, maintenance tasks, and system improvements that require immediate attention. Without a clear method for prioritizing these actions, valuable engineering resources can be wasted on less impactful work.
Additionally, many SRE teams struggle with “toil” — repetitive, low-value tasks that could be automated, but often distract engineers from more meaningful, value-added work. The challenge was to equip SREs with a framework to:
- Decide which tasks, incidents, or improvements to address first based on their potential impact
- Identify and reduce toil
- Automate repetitive work to maximize time spent on higher-value engineering tasks
The Solution
To address these challenges, we developed a self-paced certification course focused on rank-ordering actions in SRE. The course helps SREs understand the importance of prioritization, recognize toil, and implement automation strategies to focus on more impactful work.
The curriculum was based on practical, real-world scenarios and decision-making frameworks, allowing learners to:
- Rank actions based on their impact on system reliability, user satisfaction, and business goals
- Apply expert insights into task prioritization and decision-making
- Maximize their time by automating repetitive, low-value tasks
The Process
We worked closely with SRE experts to extract best practices and key insights related to task prioritization. These insights were turned into decision-driven learning scenarios that SREs could apply to their own environments.
The course was built in Rise, using branching scenarios to allow learners to practice rank-ordering tasks and making decisions based on real-world case studies. The goal was to create an engaging, interactive experience that encouraged active participation in decision-making.
Early feedback highlighted a challenge in balancing theory and practical application. We refined the course through multiple iterations, focusing on scenario-based learning to help learners bridge the gap between theory and hands-on experience.



The embedded Storyline video walks learners through a scenario that allows learners to make decisions affecting backlog management,
The Outcome
The course received positive feedback from SRE teams, with learners feeling more confident in their ability to prioritize tasks effectively. They appreciated the course’s real-world relevance and the hands-on approach to rank-ordering actions. By the end of the training, SREs were better equipped to:
- Identify and prioritize toil
- Automate repetitive tasks
- Focus on high-value work that would drive significant improvements in system reliability
Experts in the field noted that the course provided clarity on prioritization within SRE, and some even reported improvements in how their teams handled incidents and allocated resources after completing the training.
Key Takeaways
- Biggest Insight: Teaching SREs how to make decisions and prioritize effectively is just as crucial as understanding technical concepts like reliability and automation.
- Best Practices for Translating SME Expertise: Structuring expert knowledge into actionable, scenario-based learning was the most effective way to engage learners and help them apply the training to real-world challenges.
- Unexpected Challenges: Early feedback revealed that SREs needed more guidance on balancing short-term fixes with long-term system improvements. We adjusted the course to address this balance more effectively.
