Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- Evolution of IT automation: moving from static runbooks to reasoning agents.
- Anatomy of an agent: the reasoning loop, tool utilization, memory, and planning.
- Determining when to automate tasks versus when to maintain human oversight.
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops.
- Multi-agent architectures: supervisor, hierarchical, and swarm models.
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent implementations.
- Constructing an initial operational agent: querying monitoring, diagnosing issues, and proposing solutions.
Tool Integration for IT Operations
- Linking agents to Prometheus, Grafana, Datadog, and PagerDuty APIs.
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk.
- Leveraging infrastructure tools: kubectl, Terraform, and Ansible via agent actions.
- Designing secure tool interfaces featuring parameter validation and idempotency.
Incident Response Automation
- Automated incident triage: severity classification and intelligent routing.
- Generating root cause hypotheses and collecting supporting evidence.
- Automated remediation strategies: restart, scaling, rollback, and failover actions.
- Developing an incident runbook agent with adjustable autonomy levels.
Safety, Guardrails, and Human-in-the-Loop
- Action categorization: read-only, low-risk, high-risk, and destructive operations.
- Establishing approval gates and escalation policies for critical tasks.
- Guardrail patterns: action allowlists, blast radius constraints, and rollback assurances.
- Maintaining audit trails and decision provenance for compliance purposes.
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialized agents: triage, diagnosis, and remediation agents.
- Managing inter-agent communication and shared context.
- Resolving conflicts when agents propose contradictory actions.
- Simulating end-to-end major incidents with a multi-agent response strategy.
Observability and Evaluation
- Tracing agent reasoning chains to facilitate debugging and auditing.
- Assessing agent decision quality: metrics for precision, recall, and resolution time.
- Implementing feedback loops to learn from operator overrides and operational outcomes.
- Tracking costs and analyzing token economics for operational agents.
Production Deployment and Operations
- Deploying agents as services: integrating APIs, webhooks, and scheduled jobs.
- Phased autonomy rollout: transitioning from shadow mode to full auto-remediation.
- Runbooks for agent failures: managing scenarios where the agent itself malfunctions.
- Developing the business case and measuring ROI for autonomous operations.
Requirements
- Practical experience with IT operations, DevOps, or SRE methodologies.
- Proficiency in Python scripting and REST API interactions.
- A foundational understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation strategies.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders evaluating agentic AI solutions for incident management.
14 Hours