Get in Touch

Course Outline

Tencent Hunyuan Production Fundamentals

  • Overview of typical serving scenarios for Tencent Hunyuan models
  • Key production characteristics of large and MoE architectures
  • Common bottlenecks affecting latency, throughput, and cost
  • Establishing service-level objectives for inference workloads

Deployment Architecture and Serving Workflow

  • Essential components of a production-grade inference stack
  • Comparing containerized, on-premise, and cloud deployment models
  • Fundamentals of model loading, request routing, and GPU allocation
  • Designing for reliability and operational simplicity

Practical Latency Optimization

  • Leveraging optimized inference engines such as TensorRT where applicable
  • KV-cache concepts and strategies for practical cache tuning
  • Mitigating startup, warmup, and response overhead
  • Measuring time to first token and token generation speed

Throughput, Batching, and GPU Efficiency

  • Implementing continuous batching and request batching strategies
  • Managing concurrency and queue behavior effectively
  • Enhancing GPU utilization without compromising user experience
  • Handling long-context and mixed-workload requests

Quantization and Cost Management

  • The importance of quantization for production serving
  • Evaluating trade-offs of FP16, INT8, and other precision options
  • Balancing model quality, latency, and infrastructure expenditure
  • Developing a practical cost optimization checklist

Operations, Monitoring, and Readiness Assessment

  • Configuring autoscaling triggers for inference services
  • Monitoring latency, throughput, cache usage, and GPU health
  • Essentials of logging, alerting, and incident response
  • Reviewing a reference deployment and formulating an improvement plan

Requirements

  • Fundamental understanding of large language model deployment and inference workflows.
  • Hands-on experience with containers, cloud or on-premise infrastructure, and API-driven services.
  • Proficiency in Python or system engineering tasks.

Intended Audience

  • ML engineers responsible for integrating LLMs into production environments.
  • Platform engineers managing GPU-based inference services.
  • Solution architects designing scalable AI serving platforms.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories