Get in Touch

Course Outline

Insight into the Chinese AI GPU Landscape

  • Analysis of Huawei Ascend, Biren, and Cambricon MLU
  • Differences between CUDA and CANN, Biren SDK, and BANGPy frameworks
  • Market trends and vendor ecosystems

Migrating with Confidence

  • Evaluating the structure of your CUDA codebase
  • Selecting target platforms and appropriate SDK versions
  • Setting up the toolchain and development environment

Techniques for Code Conversion

  • Adapting CUDA memory handling and kernel logic
  • Aligning compute grid and thread models
  • Exploring automated versus manual translation methods

Implementation Details for Specific Platforms

  • Leveraging Huawei CANN operators and custom kernels
  • Navigating the Biren SDK conversion workflow
  • Reconstructing models using BANGPy (Cambricon)

Testing and Optimizing Across Platforms

  • Profiling runtime on each destination platform
  • Comparing memory adjustments and parallel execution
  • Monitoring performance and refining iteratively

Oversight of Hybrid GPU Setups

  • Deployments involving multiple architectures
  • Implementing fallback mechanisms and device recognition
  • Utilizing abstraction layers to enhance code sustainability

Real-World Examples and Recommended Practices

  • Converting vision and NLP models for Ascend or Cambricon
  • Adapting inference pipelines for Biren clusters
  • Resolving version discrepancies and API limitations

Recap and Future Directions

Requirements

  • Proficiency in CUDA programming or GPU-centric application development
  • Grasp of GPU memory architectures and compute kernels
  • Awareness of AI model deployment or acceleration processes

Target Participants

  • GPU developers
  • System architects
  • Porting experts
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories