Get in Touch

Course Outline

Core Principles of Cloud Operations on AWS

  • Defining operational roles and duties in cloud environments.
  • Understanding AWS account hierarchies, organizational structures, and multi-account strategies.
  • Key operational tools: CloudWatch, CloudTrail, and AWS Config.

Infrastructure as Code and Automated Provisioning

  • IaC fundamentals and the concept of immutable infrastructure.
  • Building infrastructure using Terraform and AWS CloudFormation.
  • Handling state management, modular design, and environment promotion.

CI/CD Frameworks and Deployment Tactics

  • Architecting CI/CD pipelines for cloud-native applications.
  • Implementing blue/green, canary, and rolling deployment models.
  • Automating rollback mechanisms, health checks, and release verification.

Monitoring, Observability, and Alerting Systems

  • Managing metrics, logs, and traces: collection, storage, and analysis.
  • Leveraging CloudWatch, X-Ray, and external observability platforms.
  • Establishing SLOs/SLIs, alerting policies, and on-call protocols.

Security Operations and Identity Governance

  • IAM best practices, enforcing least privilege, and managing cross-account access.
  • Managing secrets, utilizing KMS, and securing parameter stores.
  • Operational security measures: patch management, vulnerability scanning, and audit logging.

Resilience, Backup, and Disaster Recovery

  • Building for fault tolerance and high availability.
  • Formulating backup plans, automating snapshots, and executing restore processes.
  • Developing disaster recovery strategies and creating operational runbooks.

Cost Optimization and Governance

  • Enhancing cost visibility through billing analysis, tagging, and allocation strategies.
  • Optimizing resources via rightsizing, reserved instances/savings plans, and budget controls.
  • Enforcing governance through policies, guardrails, and compliance automation.

Containers, Serverless, and Runtime Management

  • Operational nuances for ECS, EKS, and Lambda.
  • Managing service discovery, autoscaling, and resource constraints.
  • Logging, tracing, and debugging containerized workloads.

Incident Response, Playbooks, and Chaos Engineering

  • Executing runbook-based incident response and conducting postmortems.
  • Automating remediation and implementing self-healing architectures.
  • Introduction to chaos experiments for testing system resilience.

Practical Workshop: Managing a Sample Workload

  • Deploying a sample application using IaC and CI/CD pipelines.
  • Setting up monitoring, alerts, and automated remediation scripts.
  • Simulating incidents to practice runbook-driven responses.

Recap and Future Directions

Requirements

  • Foundational knowledge of cloud architecture and networking principles.
  • Proficiency in Linux command-line interfaces and scripting.
  • Working experience with source control (Git) and fundamental CI/CD concepts.

Target Audience

  • Cloud operations engineers.
  • Site Reliability Engineers (SREs) and platform engineers.
  • DevOps engineers and technical leadership.
 21 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories