Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Overview of CANN Optimization Capabilities
- Mechanisms for handling inference performance within CANN
- Key optimization objectives for edge and embedded AI systems
- Insights into AI Core utilization and memory allocation strategies
Utilizing Graph Engine for Analysis
- Introduction to the Graph Engine and its execution pipeline
- Visualization of operator graphs alongside runtime metrics
- Modifying computational graphs to drive optimization
Profiling Tools and Performance Metrics
- Employing the CANN Profiling Tool (profiler) for in-depth workload analysis
- Examining kernel execution times to identify bottlenecks
- Profiling memory access patterns and exploring tiling strategies
Custom Operator Development with TIK
- An overview of TIK and its operator programming model
- Implementing custom operators using the TIK DSL
- Testing and benchmarking operator performance metrics
Advanced Operator Optimization with TVM
- Introduction to integrating TVM with CANN
- Auto-tuning strategies applied to computational graphs
- Determining when and how to switch between TVM and TIK
Memory Optimization Techniques
- Managing memory layouts and buffer placement effectively
- Strategies to minimize on-chip memory consumption
- Best practices for asynchronous execution and data reuse
Real-World Deployment and Case Studies
- Case study: Performance tuning for smart city camera pipelines
- Case study: Optimizing inference stacks for autonomous vehicles
- Guidelines for iterative profiling and sustained improvement
Summary and Next Steps
Requirements
- A solid command of deep learning model architectures and training workflows
- Practical experience in model deployment using CANN, TensorFlow, or PyTorch
- Proficiency in Linux CLI, shell scripting, and Python programming
Target Audience
- AI performance engineers
- Inference optimization specialists
- Developers engaged with edge AI or real-time systems
14 Hours