Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps

  • Defining AIOps and understanding its significance
  • Comparing traditional monitoring with AIOps-driven observability
  • Examining AIOps architecture and essential components

Gathering and Standardizing Operational Data

  • Understanding observability data types: metrics, logs, and traces
  • Ingesting data from diverse sources, including servers, containers, and the cloud
  • Utilizing agents and exporters such as Prometheus, Beats, and Fluentd

Correlating Data and Detecting Anomalies

  • Applying time series correlation and statistical techniques
  • Leveraging ML models for effective anomaly detection
  • Identifying incidents across distributed systems

Optimizing Alerting and Reducing Noise

  • Creating intelligent alert rules and setting appropriate thresholds
  • Implementing suppression, deduplication, and alert grouping strategies
  • Integrating with platforms like Alertmanager, Slack, PagerDuty, or Opsgenie

Performing Root Cause Analysis and Visualization

  • Using dashboards to visualize metrics and identify trends
  • Reviewing events and timelines to support RCA
  • Tracking issues across layers using distributed tracing tools

Automating Processes and Remediation

  • Initiating automated scripts or workflows triggered by incidents
  • Connecting with ITSM systems such as ServiceNow and Jira
  • Exploring use cases including self-healing, scaling, and traffic rerouting

Surveying Open Source and Commercial AIOps Platforms

  • Reviewing key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace
  • Establishing evaluation criteria for selecting an AIOps platform
  • Conducting demos and hands-on activities with a chosen stack

Recap and Future Directions

Requirements

  • A solid grasp of IT operations and system monitoring concepts
  • Practical experience with monitoring tools or dashboards
  • Basic familiarity with standard log and metric formats

Target Audience

  • Operations teams overseeing infrastructure and applications
  • Site Reliability Engineers (SREs)
  • Teams dedicated to IT monitoring and observability

Number of participants


Price per participant

Upcoming Courses

Related Categories