Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Insight into the Chinese AI GPU Landscape
- Analysis of Huawei Ascend, Biren, and Cambricon MLU
- Differences between CUDA and CANN, Biren SDK, and BANGPy frameworks
- Market trends and vendor ecosystems
Migrating with Confidence
- Evaluating the structure of your CUDA codebase
- Selecting target platforms and appropriate SDK versions
- Setting up the toolchain and development environment
Techniques for Code Conversion
- Adapting CUDA memory handling and kernel logic
- Aligning compute grid and thread models
- Exploring automated versus manual translation methods
Implementation Details for Specific Platforms
- Leveraging Huawei CANN operators and custom kernels
- Navigating the Biren SDK conversion workflow
- Reconstructing models using BANGPy (Cambricon)
Testing and Optimizing Across Platforms
- Profiling runtime on each destination platform
- Comparing memory adjustments and parallel execution
- Monitoring performance and refining iteratively
Oversight of Hybrid GPU Setups
- Deployments involving multiple architectures
- Implementing fallback mechanisms and device recognition
- Utilizing abstraction layers to enhance code sustainability
Real-World Examples and Recommended Practices
- Converting vision and NLP models for Ascend or Cambricon
- Adapting inference pipelines for Biren clusters
- Resolving version discrepancies and API limitations
Recap and Future Directions
Requirements
- Proficiency in CUDA programming or GPU-centric application development
- Grasp of GPU memory architectures and compute kernels
- Awareness of AI model deployment or acceleration processes
Target Participants
- GPU developers
- System architects
- Porting experts
21 Hours