Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps

  • Defining AIOps and its significance in modern operations
  • Comparing traditional monitoring with AIOps-driven observability
  • Architecture of AIOps and its essential components

Collection and Normalization of Operational Data

  • Types of observability data: metrics, logs, and traces
  • Ingesting data from diverse sources including servers, containers, and the cloud
  • Utilizing agents and exporters such as Prometheus, Beats, and Fluentd

Data Correlation and Anomaly Detection

  • Time series correlation and associated statistical methods
  • Applying ML models for effective anomaly detection
  • Identifying incidents within distributed system environments

Alerting Strategies and Noise Reduction

  • Developing intelligent alert rules and appropriate thresholds
  • Implementing suppression, deduplication, and alert grouping techniques
  • Integration with tools like Alertmanager, Slack, PagerDuty, or Opsgenie

Root Cause Analysis and Visualization

  • Leveraging dashboards to visualize metrics and identify trends
  • Investigating events and timelines for comprehensive RCA
  • Tracking issues across layers using distributed tracing tools

Automation and Remediation Processes

  • Initiating automated scripts or workflows triggered by incidents
  • Integration with ITSM platforms such as ServiceNow and Jira
  • Use cases including self-healing, auto-scaling, and traffic rerouting

Open Source and Commercial AIOps Platforms

  • Overview of key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace
  • Criteria for evaluating and selecting an appropriate AIOps platform
  • Live demonstration and hands-on practice with a selected stack

Summary and Recommended Next Steps

Requirements

  • A solid understanding of IT operations and system monitoring concepts
  • Practical experience with monitoring tools or dashboards
  • Familiarity with basic log and metric formats

Target Audience

  • Operations teams managing infrastructure and applications
  • Site Reliability Engineers (SREs)
  • IT monitoring and observability specialists

Related Categories