pytorch · fundamentals · cheat sheet

PyTorch basics: tensors, autograd, and the training loop

Everything you type constantly in PyTorch, compressed into one page: how tensors are made and reshaped, how autograd walks its graph backwards, the five-line training loop, and the handful of mistakes that produce ninety percent of runtime errors.

0How to read this sheet

This page compresses the working core of PyTorch into one reference: tensor creation, reshaping, indexing, broadcasting, autograd, a minimal training loop, and the pitfalls that bite everyone at least once. Sections 1 through 8 cover tensor mechanics, 9 through 15 cover learning, 16 through 18 cover performance, gotchas, and copy-paste snippets.1

fn 1API names match the current stable PyTorch release. Anything version-sensitive (like torch.eig) is flagged in the docs links at the bottom.

1Creating tensors


2Inspecting tensors


3Reshape and dimensions


4Indexing, slicing, advanced indexing


5Concatenate and stack


6Broadcasting and arithmetic


7Linear algebra helpers


8Random seeds

torch.manual_seed(42)
if torch.cuda.is_available(): torch.cuda.manual_seed_all(42)

9Autograd and gradients: the core

x (x+2) (x+5) * +3 mean=z dz/dy = 1/4 dy/dx = 2x + 7 = 9 at x=1 x.grad = 2.250 forward backward
forward loss z21.000
dz/dy per element0.250
x.grad (all entries)2.250
graph size6 nodes
backward complete: every leaf that required grad received its gradient
Fig. 1. The graph from section 18's gradient snippet: z = mean((x+2)(x+5)+3) on a 2×2 tensor of ones. Analytic check: dz/dx = (2x+7)/4 = 2.25 per element, exactly what x.grad reports after z.backward().
Lesson

Autograd is bookkeeping, not magic: every operation records where it came from so backward can walk the same path in reverse. Detach when you mean it, zero grads when you loop, and never assume a gradient exists for a tensor you built with requires_grad=False.


10Common neural net building blocks

import torch.nn as nn
class MLP(nn.Module):
    def __init__(self, in_dim, hidden, out_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(in_dim, hidden),
            nn.ReLU(),
            nn.Linear(hidden, out_dim)
        )
    def forward(self, x):
        return self.net(x)

11Training loop: the minimal pattern

model.train()
for epoch in range(epochs):
    for xb, yb in dataloader:
        xb, yb = xb.to(device), yb.to(device)
        preds = model(xb)
        loss = criterion(preds, yb)
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

For evaluation:

model.eval()
with torch.no_grad():
    for xb, yb in val_loader:
        preds = model(xb)
        # compute metrics

12Optimizers and schedulers


13Save and load

model = MyModel(...)
model.load_state_dict(torch.load("m.pt", map_location=device))
model.to(device)

14Data pipeline


15Useful tensor utilities


16Memory and performance tips

Lesson

The GPU starves on transfers, not on math (see vectorization and coalescing for why). Every sync point you remove from the inner loop is free speed.


17Pitfalls and gotchas

Mistakes everyone makes once
MistakeSymptomFix
Mixing CPU and CUDA tensorsRuntimeError about device mismatchmove both operands with .to(device)
Skipping optimizer.zero_grad()gradients accumulate across batcheszero before every loss.backward()
Calling .item() on non-scalar tensorserror at runtimeindex or reduce to a single element first
In-place ops on tensors that require gradautograd errors or silent corruptionprefer out-of-place ops; know when in-place breaks the graph
Normalizing using whole-dataset statisticssilent data leakagecompute stats on the training fold only

18Short example snippets

Create a tensor, reshape, and sum:

x = torch.arange(6).view(2,3)   # [[0,1,2],[3,4,5]]
col_sum = x.sum(dim=0)          # collapse rows: [3, 5, 7]
row_sum = x.sum(dim=1)          # collapse columns: [3, 12]
fn 4The axis argument names the dimension being removed. sum(dim=0) collapses rows and leaves one value per column; sum(dim=1) collapses columns and leaves one value per row.

Gradients:

x = torch.ones(2,2, requires_grad=True)
y = (x + 2) * (x + 5) + 3
z = y.mean()
z.backward()
print(x.grad)  # gradient of z wrt x, all entries 2.25

Batch matrix multiply:

a = torch.rand(3,4,5)
b = torch.rand(3,5,4)
c = torch.bmm(a,b)   # shape (3, 4, 4)

19References

  1. PyTorch documentation: stable API reference.
  2. Official PyTorch tutorials: tensors, autograd, datasets.
  3. torch.linalg: the current linear algebra namespace.
  4. Local source: bin/blogs/pytorch-basics.md, the original markdown note this page was rewritten from.