Get in Touch

Course Outline

Introduction to Agentic AI in Operations

  • Evolution of IT automation: moving from static runbooks to reasoning agents.
  • Anatomy of an agent: the reasoning loop, tool utilization, memory, and planning.
  • Determining when to automate tasks versus when to maintain human oversight.

Agent Frameworks and Architectures

  • Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops.
  • Multi-agent architectures: supervisor, hierarchical, and swarm models.
  • Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent implementations.
  • Constructing an initial operational agent: querying monitoring, diagnosing issues, and proposing solutions.

Tool Integration for IT Operations

  • Linking agents to Prometheus, Grafana, Datadog, and PagerDuty APIs.
  • Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk.
  • Leveraging infrastructure tools: kubectl, Terraform, and Ansible via agent actions.
  • Designing secure tool interfaces featuring parameter validation and idempotency.

Incident Response Automation

  • Automated incident triage: severity classification and intelligent routing.
  • Generating root cause hypotheses and collecting supporting evidence.
  • Automated remediation strategies: restart, scaling, rollback, and failover actions.
  • Developing an incident runbook agent with adjustable autonomy levels.

Safety, Guardrails, and Human-in-the-Loop

  • Action categorization: read-only, low-risk, high-risk, and destructive operations.
  • Establishing approval gates and escalation policies for critical tasks.
  • Guardrail patterns: action allowlists, blast radius constraints, and rollback assurances.
  • Maintaining audit trails and decision provenance for compliance purposes.

Multi-Agent Orchestration for Complex Incidents

  • Coordinating specialized agents: triage, diagnosis, and remediation agents.
  • Managing inter-agent communication and shared context.
  • Resolving conflicts when agents propose contradictory actions.
  • Simulating end-to-end major incidents with a multi-agent response strategy.

Observability and Evaluation

  • Tracing agent reasoning chains to facilitate debugging and auditing.
  • Assessing agent decision quality: metrics for precision, recall, and resolution time.
  • Implementing feedback loops to learn from operator overrides and operational outcomes.
  • Tracking costs and analyzing token economics for operational agents.

Production Deployment and Operations

  • Deploying agents as services: integrating APIs, webhooks, and scheduled jobs.
  • Phased autonomy rollout: transitioning from shadow mode to full auto-remediation.
  • Runbooks for agent failures: managing scenarios where the agent itself malfunctions.
  • Developing the business case and measuring ROI for autonomous operations.

Requirements

  • Practical experience with IT operations, DevOps, or SRE methodologies.
  • Proficiency in Python scripting and REST API interactions.
  • A foundational understanding of LLM capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers exploring AI-driven automation strategies.
  • Platform engineers focused on building self-healing infrastructure.
  • IT operations leaders evaluating agentic AI solutions for incident management.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories