Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: NVIDIA Corporation

Rating No ratings yet

This is NVIDIA's official guide for CUDA C++, a parallel computing platform that allows GPU acceleration using C++. What You’ll Learn: CUDA architecture: How threads, blocks, and warps work. Memory hierarchy: Registers, shared memory, global memory optimizations. Parallel algorithms: Implementing reductions, prefix sums, and matrix multiplication. Optimization techniques: Profiling, occupancy, and efficient memory access patterns. Example Topic: Thread divergence: Explains why branching in CUDA (e.g., if-else statements) inside a warp can slow execution.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 NVIDIA's official CUDA C++ reference teaches you to write and tune GPU-accelerated C++ by grounding you in the execution model, memory hierarchy, and the API surface you'll actually call. Best for C++ programmers with some parallel-computing curiosity who want a durable desk reference rather than a tutorial narrative. 【Book Arc】 - **Opening (~0%–10%)**: Frames why GPUs exist, then lays out the scalable programming model — kernels, thread hierarchy (including thread block clusters), memory hierarchy, heterogeneous programming, and the asynchronous SIMT model. - **Early (~10%–30%)**: Moves into the programming interface: NVCC compilation workflow (offline and JIT), binary/PTX/C++ compatibility, the CUDA runtime, device memory and L2 cache management, streams, graphs, events, multi-device systems, and graphics/interop surfaces. - **Early–Middle (~30%–50%)**: The language-extensions reference — function and variable memory-space qualifiers, built-in vector types and variables, fences and synchronization, texture/surface functions, atomics, address-space predicates and conversions, warp vote/match/reduce/shuffle, and asynchronous barriers. - **Middle (~50%–70%)**: Hardware implementation and performance guidelines: SIMT architecture, hardware multithreading, and the strategy chapters on maximizing utilization and efficient memory access. (Excerpts do not cover the detailed contents of this range.) - **Late (~70%–90%)**: Advanced and platform-specific material — cooperative groups, dynamic parallelism, virtual memory management, and multi-GPU programming. (Excerpts do not cover this range in detail.) - **Ending (~90%–100%)**: Mathematical API references, driver/runtime differences, and appendices. (Excerpts do not cover this range.) 【Key Takeaways】 - **The execution model is the mental model** (Opening): Kernels launch as grids of blocks of threads; understanding blockIdx/blockDim/threadIdx and warpSize is the prerequisite for everything else. - **Memory hierarchy drives performance** (Opening–Early): Registers, shared, constant, and global memory have different costs; the guide devotes whole sections to L2 set-aside, persistence policies, and cache-hint load/store functions. - **Compilation is a compatibility problem, not just a build step** (Early): NVCC's offline vs. JIT paths, binary/PTX/C++/64-bit compatibility, and versioning rules determine whether your binary runs on a given GPU. - **The runtime API is broad and layered** (Early): Device memory, streams, CUDA Graphs, events, multi-device peer-to-peer, unified virtual addressing, and interprocess communication are all first-class topics. - **Interop is expected, not exotic** (Early): OpenGL, Direct3D 11/12, Vulkan, SLI, and NVSCI interoperability sections show CUDA is designed to coexist with graphics and external pipelines. - **Atomics, fences, and warp primitives are the concurrency toolkit** (Middle): Memory fence functions, atomic arithmetic, warp vote/match/reduce/shuffle, and asynchronous barriers cover the synchronization patterns you'll need beyond simple `__syncthreads()`. - **Warp specialization and asynchronous barriers are advanced patterns** (Middle): The barrier chapter's discussion of temporal splitting, phase/countdown semantics, and spatial partitioning signals where serious kernel engineering heads. - **Performance work is structured, not ad hoc** (Early–Middle): The guide separates overall optimization strategy from specific tactics like maximizing utilization and memory access efficiency. 【Reading Tips】 - Treat this as a reference, not a novel: read the Opening programming-model chapters linearly, then jump via the table of contents to the API sections you need. - Deep-read the memory hierarchy and performance-guideline chapters — they pay off across every kernel you write. - Skim the long function-by-function API listings (texture, surface, atomics) on first pass; return when you actually call them. - Keep the compatibility and NVCC sections bookmarked — they answer the "why won't this run on my GPU" questions that recur in practice. - Pair the warp-level primitives chapter with hands-on reduction/prefix-sum exercises; the concepts only stick when you write them. 【Coverage Limits】 This guide is based on stratified excerpts that are heavily weighted toward the table of contents and API reference listings; the Middle, Late, and Ending ranges are largely unrepresented, so claims about those sections are inferred from chapter titles rather than read in detail.
Page 3
. 21 6.1.1 Compilation Workflow . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 6.1.1.1 Offline Compilation . . . . . . . ....
View in text
Page 4
. . . . . . . . . . . . . . . . . . . . . . . . . . 139 8.2.1 Application Level . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 6
10.9.1.13 surfCubemapLayeredread() . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184 10.9.1.14 surfCubemapLayeredwrite() . . . . . . . ....
View in text
Page 7
ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207 10.22.1 Synopsis . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Page 9
. . . . . . . . . . . . . . . . . . . . . . . . . . 281 11.2 What’s New in Cooperative Groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Excerpt 6
1 and CDP2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 332 viii 12.5.2 Compatibility and Interoperability . . . . . . . . . . . . . . . . . . ....
View in text
Page 3
. . 417 17.5.4 Pointers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 418 17.5.5 Operators . . . . . . . ....
View in text
Page 13
. . . . . . . . . . . . . . 441 17.5.25.1 Module support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 441 17.5.25.2 Corout...
View in text
Tags
AI categories
Programming LanguageC++Technology
Publisher: NVIDIA Corporation
Publish Year: 2025
Language: English
Pages: 584
File Format: PDF
File Size: 4.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…