0Why we care at all
The GPU is starving. Its math units are so fast that most kernels spend their time waiting for data from HBM (global memory). HBM is slow: a load costs hundreds of cycles.1 So the whole game is to move the bytes you need in as few and as wide trips as possible. Two tricks do that, and they attack two different parts of the problem.
Everything below is about one question: when a warp issues a load instruction, what actually travels over the wire?
1Vectorization: one thread, wide loads
A thread can load 1 element per instruction, or it can load 4, 8, or 16 bytes in a single instruction. Same latency, more data. Fewer instructions to move the same tile.
scalar: load a | load b | load c | load d 4 instructions
vectorized: load [a b c d] 1 instruction
The catch: the data has to be contiguous in memory (a, b, c, d next to each other) and aligned.2 Then one wide instruction grabs them all. This is vectorization: make each thread's load as wide as the layout allows.
LessonLatency is fixed, width is not. If a thread must make the trip anyway, carry more bytes per trip.
2Coalescing: many threads, adjacent addresses
A warp is 32 threads issuing the same load instruction at the same time.3 If those 32 threads ask for 32 addresses that are next to each other, the hardware fuses them into a few big memory transactions. If they ask for 32 scattered addresses, it does 32 separate slow transactions.
coalesced: thread0->addr0 thread1->addr1 ... thread31->addr31 (contiguous)
hardware fuses into ~1 wide transaction fast
scattered: thread0->addr0 thread1->addr500 ... (all over)
32 separate transactions slow
The rule: arrange it so adjacent threads touch adjacent addresses. This is coalescing. It is about the shape of the thread-to-data mapping, not how wide one thread's load is.
3Two different knobs, not one
vectorization = one thread reads a wide contiguous chunk
coalescing = adjacent threads read adjacent addresses
One is "how fat is a single thread's load". The other is "do the 32 threads line up on memory". You want both: fat loads and lined-up threads. Together they turn a tile move into the fewest possible wide HBM transactions.
| Property | Vectorization | Coalescing |
|---|---|---|
| Scope | one thread | the whole warp |
| What it widens | bytes per instruction | requests fused per transaction |
| Requirement | contiguous, aligned data | adjacent lane ids map to adjacent addresses |
| Failure mode | narrow scalar loads | split transactions |
| CuTe knob | num_bits_per_copy | thr_layout (the T of TV) |
Root causeMost "my kernel is bandwidth bound" surprises trace back to one of these two knobs being off, not to the algorithm. Check both before rewriting anything.
4How CuTe exposes exactly one knob each
CuTe splits these into exactly two calls.4
# 1) vectorization: how many bits one thread copies per instruction
atom = cute.make_copy_atom(cute.nvgpu.copyuniversalop(),
cutlass.float16, num_bits_per_copy=128)
# 2) coalescing: how the threads are laid out over the data (the tv layout)
tc = cute.make_tiled_copy_tv(atom, thr_layout, val_layout)
# then every thread copies its slice
thr = tc.get_slice(tidx)
cute.copy(tc, thr.partition_s(gmem), thr.partition_d(smem))
make_copy_atom(..., num_bits_per_copy=n)sets the vector width. Bump n from 16 to 128 and each thread moves 8 fp16 values in one shot instead of one. That is vectorization.make_tiled_copy_tv(atom, thr_layout, val_layout)sets the thread layout (the "t"). You choose it so consecutive threads land on consecutive memory. That is coalescing. The "v" (value layout) says how many elements each thread owns and in what shape.
So: the atom controls the fat load, the tiled copy controls the thread lineup.
LessonWhen a copy is slow, name the culprit before tuning anything: is this a width problem (fix the atom) or a lineup problem (fix the thread layout)? The fix follows from the diagnosis.
5References
- CUDA C++ Programming Guide, section on memory coalescing and warp access patterns.
- NVIDIA CUTLASS repository, CuTe layout algebra and copy atoms documentation.
- CUTLASS examples directory, tiled copy usage in working kernels.
- Local source:
bin/blogs/vectorization-coalescing.md, the original markdown note this page was rewritten from.