GPU kernels · CUTLASS CuTe · memory movement

Vectorization and coalescing: the two ways to move memory fast

A kernel is rarely slow because it computes too much. It is slow because it waits for bytes. Two independent knobs decide how fast data moves: how wide each thread's single load is, and whether adjacent threads touch adjacent addresses.

0Why we care at all

The GPU is starving. Its math units are so fast that most kernels spend their time waiting for data from HBM (global memory). HBM is slow: a load costs hundreds of cycles.1 So the whole game is to move the bytes you need in as few and as wide trips as possible. Two tricks do that, and they attack two different parts of the problem.

fn 1HBM latency is measured in hundreds of GPU clock cycles per trip. An L2 hit is cheaper, registers are almost free, but the first fetch from global memory is always the expensive one you design around.

Everything below is about one question: when a warp issues a load instruction, what actually travels over the wire?


1Vectorization: one thread, wide loads

A thread can load 1 element per instruction, or it can load 4, 8, or 16 bytes in a single instruction. Same latency, more data. Fewer instructions to move the same tile.

scalar:      load a  |  load b  |  load c  |  load d      4 instructions
vectorized:  load [a b c d]                               1 instruction

The catch: the data has to be contiguous in memory (a, b, c, d next to each other) and aligned.2 Then one wide instruction grabs them all. This is vectorization: make each thread's load as wide as the layout allows.

fn 2Alignment matters because wide loads map to machine instructions with address requirements. A 128-bit copy wants a suitably aligned base address; misaligned bases force the hardware to split or refuse the access.
Lesson

Latency is fixed, width is not. If a thread must make the trip anyway, carry more bytes per trip.


2Coalescing: many threads, adjacent addresses

A warp is 32 threads issuing the same load instruction at the same time.3 If those 32 threads ask for 32 addresses that are next to each other, the hardware fuses them into a few big memory transactions. If they ask for 32 scattered addresses, it does 32 separate slow transactions.

fn 3One warp, one instruction, 32 lanes. The instruction is issued once; what differs per lane is only the address each lane computes.
coalesced:   thread0->addr0  thread1->addr1 ... thread31->addr31  (contiguous)
             hardware fuses into ~1 wide transaction               fast

scattered:   thread0->addr0  thread1->addr500 ... (all over)
             32 separate transactions                              slow

The rule: arrange it so adjacent threads touch adjacent addresses. This is coalescing. It is about the shape of the thread-to-data mapping, not how wide one thread's load is.

COALESCED SCATTERED
transactions left1
transactions right8
bytes per trip, left32 B
bytes per trip, right4 B
contiguous lanes fuse into one wide trip; scattered lanes pay eight separate ones
Fig. 1. Eight of a warp's 32 lanes shown, fp32 elements at 4 B each. Watch contiguous requests fuse under one green bar while scattered requests each get their own red bar and their own latency bill.

3Two different knobs, not one

vectorization  = one thread reads a wide contiguous chunk
coalescing     = adjacent threads read adjacent addresses

One is "how fat is a single thread's load". The other is "do the 32 threads line up on memory". You want both: fat loads and lined-up threads. Together they turn a tile move into the fewest possible wide HBM transactions.

The two knobs compared
PropertyVectorizationCoalescing
Scopeone threadthe whole warp
What it widensbytes per instructionrequests fused per transaction
Requirementcontiguous, aligned dataadjacent lane ids map to adjacent addresses
Failure modenarrow scalar loadssplit transactions
CuTe knobnum_bits_per_copythr_layout (the T of TV)
Root cause

Most "my kernel is bandwidth bound" surprises trace back to one of these two knobs being off, not to the algorithm. Check both before rewriting anything.


4How CuTe exposes exactly one knob each

CuTe splits these into exactly two calls.4

fn 4Snippets follow the CUTLASS Python (CuTe DSL) tutorial shape. The atom sets copy width, the tiled copy sets thread partitioning.
# 1) vectorization: how many bits one thread copies per instruction
atom = cute.make_copy_atom(cute.nvgpu.copyuniversalop(),
                           cutlass.float16, num_bits_per_copy=128)

# 2) coalescing: how the threads are laid out over the data (the tv layout)
tc = cute.make_tiled_copy_tv(atom, thr_layout, val_layout)

# then every thread copies its slice
thr = tc.get_slice(tidx)
cute.copy(tc, thr.partition_s(gmem), thr.partition_d(smem))

So: the atom controls the fat load, the tiled copy controls the thread lineup.

Lesson

When a copy is slow, name the culprit before tuning anything: is this a width problem (fix the atom) or a lineup problem (fix the thread layout)? The fix follows from the diagnosis.


5References

  1. CUDA C++ Programming Guide, section on memory coalescing and warp access patterns.
  2. NVIDIA CUTLASS repository, CuTe layout algebra and copy atoms documentation.
  3. CUTLASS examples directory, tiled copy usage in working kernels.
  4. Local source: bin/blogs/vectorization-coalescing.md, the original markdown note this page was rewritten from.