Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 28 hours
Course Outline
Introduction
- What is CUDA?
- CUDA vs. OpenCL vs. SYCL
- Overview of CUDA features and architecture
- Setting up the development environment
Getting Started
- Creating a new CUDA project using Visual Studio Code
- Exploring the project structure and files
- Compiling and running the program
- Displaying output using printf and fprintf
CUDA API
- Understanding the role of the CUDA API in the host program
- Using the CUDA API to query device information and capabilities
- Using the CUDA API to allocate and deallocate device memory
- Using the CUDA API to copy data between host and device
- Using the CUDA API to launch kernels and synchronize threads
- Using the CUDA API to handle errors and exceptions
CUDA C/C++
- Understanding the role of CUDA C/C++ in the device program
- Using CUDA C/C++ to write kernels that execute on the GPU and manipulate data
- Using CUDA C/C++ data types, qualifiers, operators, and expressions
- Using CUDA C/C++ built-in functions, such as math, atomic, and warp operations
- Using CUDA C/C++ built-in variables, such as threadIdx, blockIdx, and blockDim
- Using CUDA C/C++ libraries, such as cuBLAS, cuFFT, and cuRAND
CUDA Memory Model
- Understanding the differences between host and device memory models
- Using CUDA memory spaces, such as global, shared, constant, and local
- Using CUDA memory objects, such as pointers, arrays, textures, and surfaces
- Using CUDA memory access modes, such as read-only, write-only, and read-write
- Using the CUDA memory consistency model and synchronization mechanisms
CUDA Execution Model
- Understanding the differences between host and device execution models
- Using CUDA threads, blocks, and grids to define parallelism
- Using CUDA thread functions, such as threadIdx, blockIdx, and blockDim
- Using CUDA block functions, such as __syncthreads and __threadfence_block
- Using CUDA grid functions, such as gridDim, gridSync, and cooperative groups
Debugging
- Understanding common errors and bugs in CUDA programs
- Using the Visual Studio Code debugger to inspect variables, breakpoints, and the call stack
- Using CUDA-GDB to debug CUDA programs on Linux
- Using CUDA-MEMCHECK to detect memory errors and leaks
- Using NVIDIA Nsight to debug and analyze CUDA programs on Windows
Optimization
- Understanding factors that affect the performance of CUDA programs
- Using CUDA coalescing techniques to improve memory throughput
- Using CUDA caching and prefetching techniques to reduce memory latency
- Using CUDA shared and local memory techniques to optimize memory accesses and bandwidth
- Using CUDA profiling and profiling tools to measure and improve execution time and resource utilization
Summary and Next Steps
Requirements
- A solid understanding of C/C++ and parallel programming concepts
- Fundamental knowledge of computer architecture and memory hierarchy
- Proficiency with command-line tools and code editors
Audience
- Developers eager to learn how to use CUDA to program NVIDIA GPUs and exploit their parallelism
- Developers aiming to write high-performance, scalable code capable of running on various CUDA devices
- Programmers interested in exploring the low-level aspects of GPU programming to optimize code performance