Course Outline
Introduction
- The intersection of SRE with traditional IT and software development.
- The critical need for automation and observability.
- Distinguishing the roles of software engineers and system administrators.
- Comparing Site Reliability Engineers with DevOps engineers.
Understanding IT System Architecture
- System design across on-premise and cloud environments.
Core SRE Principles and Practices
- Implementing Infrastructure as Code.
- The impact of containerization and orchestration (e.g., Docker, Kubernetes).
- Adopting Continuous Integration, Continuous Deployment, and Continuous Delivery.
- Achieving deep observability.
Assessing IT System Readiness
- Auditing team capabilities and organisational resources.
- Mapping existing systems and workflows.
- Forecasting the potential benefits of SRE implementation.
- Defining the role of the software engineering team.
- Clarifying the role of the operations team.
- The oversight role of management.
Sustaining System Reliability
- Defining and measuring desired service reliability levels.
- Comprehending Service Level Objectives (SLOs).
- Understanding Service Level Indicators (SLIs) and Service Level Agreements (SLAs).
- Managing Error Budgets effectively.
- Formulating a specific SLO.
Enhancing System Administration
- Configuring a development environment.
- Assessing relevant SRE tooling.
- Prioritising tasks for automation.
- Coding and developing solutions.
Implementing "Infrastructure as Code"
- Testing and refining code.
- Building anti-fragile systems.
- Learning from failures to improve resilience.
System Monitoring
- Observing and analysing system performance.
- Utilising SRE-specific tools and techniques.
The Future Trajectory of SRE
Requirements
- A foundational grasp of IT infrastructure components.
- Basic familiarity with the software development lifecycle.
- Prior experience in programming or scripting using any language.
Intended Audience
- Software Developers
- System Administrators
- Software Architects
- DevOps Engineers
- IT Managers
Testimonials (7)
How detailed subjects are explained with real world examples
Brian Hlabane - African Bank
Course - Site Reliability Engineering (SRE) Fundamentals
She is expert in area and provide really nice training. Material, training was really mix of examples , discussion and
Peter Tutka - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
View on the SRE/ DevOps from more business/ theoretical point of view. Most helpful for people who already have the practical view.
Michael Varhol - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
Approach of the training to send questionnaire before the training, so the training was planned accordingly to expectations. Brings the participants more active.
Stefan Girman - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
Sticking to the initial survey from attendees about what should be the focus of training.
Denis Majorsky - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
discussions , SRE definition
Daniel Horvath - Deutsche Telekom IT & Telecommunications Slovakia s.r.o.
Course - Site Reliability Engineering (SRE) Fundamentals
Concept of the training, keeping the people focused by asking them a questions and triggering discussions. Also group breakout sessions were great to think about things in groups and see different outcomes from other group.