Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and understanding its significance
- Comparing traditional monitoring with AIOps-driven observability
- Examining AIOps architecture and essential components
Gathering and Standardizing Operational Data
- Understanding observability data types: metrics, logs, and traces
- Ingesting data from diverse sources, including servers, containers, and the cloud
- Utilizing agents and exporters such as Prometheus, Beats, and Fluentd
Correlating Data and Detecting Anomalies
- Applying time series correlation and statistical techniques
- Leveraging ML models for effective anomaly detection
- Identifying incidents across distributed systems
Optimizing Alerting and Reducing Noise
- Creating intelligent alert rules and setting appropriate thresholds
- Implementing suppression, deduplication, and alert grouping strategies
- Integrating with platforms like Alertmanager, Slack, PagerDuty, or Opsgenie
Performing Root Cause Analysis and Visualization
- Using dashboards to visualize metrics and identify trends
- Reviewing events and timelines to support RCA
- Tracking issues across layers using distributed tracing tools
Automating Processes and Remediation
- Initiating automated scripts or workflows triggered by incidents
- Connecting with ITSM systems such as ServiceNow and Jira
- Exploring use cases including self-healing, scaling, and traffic rerouting
Surveying Open Source and Commercial AIOps Platforms
- Reviewing key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace
- Establishing evaluation criteria for selecting an AIOps platform
- Conducting demos and hands-on activities with a chosen stack
Recap and Future Directions
Requirements
- A solid grasp of IT operations and system monitoring concepts
- Practical experience with monitoring tools or dashboards
- Basic familiarity with standard log and metric formats
Target Audience
- Operations teams overseeing infrastructure and applications
- Site Reliability Engineers (SREs)
- Teams dedicated to IT monitoring and observability