Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Deployment

  • Packaging models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Considerations for versioning and storage

Serving Models on Kubernetes

  • Introduction to inference servers
  • Deployment of TensorFlow Serving and TorchServe
  • Configuring model endpoints

Techniques for Optimizing Inference

  • Batching strategies
  • Handling concurrent requests
  • Tuning for latency and throughput

Autoscaling ML Workloads

  • Horizontal Pod Autoscaler (HPA)
  • Vertical Pod Autoscaler (VPA)
  • Kubernetes Event-Driven Autoscaling (KEDA)

GPU Provisioning and Resource Management

  • Configuration of GPU nodes
  • Overview of the NVIDIA device plugin
  • Setting resource requests and limits for ML workloads

Strategies for Model Rollout and Release

  • Blue/green deployments
  • Canary rollout patterns
  • A/B testing for model evaluation

Monitoring and Observability for Production ML

  • Key metrics for inference workloads
  • Best practices for logging and tracing
  • Setting up dashboards and alerts

Security and Reliability Considerations

  • Securing model endpoints
  • Implementing network policies and access control
  • Maintaining high availability

Summary and Future Steps

Requirements

  • A solid understanding of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Basic familiarity with Kubernetes concepts

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories