Get in Touch
 Duration 14 hours

Course Outline

Readying Machine Learning Models for Deployment

  • Encapsulating models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Versioning and storage best practices

Serving Models on Kubernetes

  • Introduction to inference servers
  • Implementation of TensorFlow Serving and TorchServe
  • Establishing model endpoints

Optimizing Inference Performance

  • Strategies for request batching
  • Handling concurrent requests
  • Tuning for latency and throughput

Autoscaling ML Workloads

  • Horizontal Pod Autoscaler (HPA)
  • Vertical Pod Autoscaler (VPA)
  • Kubernetes Event-Driven Autoscaling (KEDA)

GPU Allocation and Resource Control

  • Setting up GPU nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Model Release and Rollout Strategies

  • Blue/green deployment techniques
  • Canary release patterns
  • A/B testing for model assessment

Monitoring and Observability for Production ML

  • Key metrics for inference workloads
  • Best practices for logging and tracing
  • Dashboards and alerting mechanisms

Security and Reliability Best Practices

  • Protecting model endpoints
  • Network policies and access management
  • Maintaining high availability

Wrap-up and Future Directions

Requirements

  • A grasp of workflows for containerized applications
  • Practical experience with Python-based machine learning models
  • Working knowledge of Kubernetes core concepts

Intended Audience

  • ML Engineers
  • DevOps Engineers
  • Platform Engineering Teams

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories