0How to read this sheet
This page compresses the working core of PyTorch into one reference: tensor creation, reshaping, indexing, broadcasting, autograd, a minimal training loop, and the pitfalls that bite everyone at least once. Sections 1 through 8 cover tensor mechanics, 9 through 15 cover learning, 16 through 18 cover performance, gotchas, and copy-paste snippets.1
torch.eig) is flagged in the docs links at the bottom.1Creating tensors
Random:
torch.rand(shape)uniform in [0, 1);torch.randn(shape)normal.torch.Tensor(2, 3)gives an uninitialized tensor (random garbage values).Zeros / ones / full:
torch.zeros(2,3),torch.ones(2,3),torch.full((2,3), 5).From Python list or NumPy:
torch.tensor([[1,2],[3,4]]),torch.from_numpy(numpy_array). The dtype follows the source NumPy dtype (for example Float64 becomestorch.DoubleTensor).2Type-casting:
x.float(),x.long(), ortorch.tensor(..., dtype=torch.int64).In-place ops: methods ending in
_(likex.fill_(5)) modify the tensor in place.
2Inspecting tensors
x.shapeorx.size()x.dtype,x.device,x.numel()(number of elements)x.item()returns a Python scalar, but only for single-element tensors.
3Reshape and dimensions
x.view(new_shape)works if the tensor is contiguous;x.reshape(new_shape)is the safer call.x.unsqueeze(dim)adds a size-1 dim;x.squeeze(dim)removes one.x.transpose(0,1)orx.t()for 2D;x.permute(...)reorders dims.x.contiguous()produces contiguous memory before.view().
4Indexing, slicing, advanced indexing
Basic slicing:
x[:, :2],x[0,1].Fancy indexing:
torch.index_select(x, dim, idx_longtensor).Pair indexing:
x[row_idx, col_idx]where both are LongTensors.torch.nonzero(x)orx.nonzero()for indices of nonzero elements.
5Concatenate and stack
torch.cat([a,b], dim=0): concatenate along an existing dimension.torch.stack([a,b], dim=0): add a new dimension and stack.
6Broadcasting and arithmetic
Usual ops:
+,-,*,/, or functional forms:torch.add,torch.mul, and friends.Matrix multiply:
torch.mm(A,B)for 2D,torch.matmulwhen you want broadcast-aware behavior.Batch matrix multiply:
torch.bmm(a, b)for 3Daandb.Aggregations:
x.sum(dim=...),x.mean(dim=...),x.max(dim=...).
7Linear algebra helpers
torch.mm(A,B): matrix multiply (2D).torch.bmm(A,B): batch matrix multiply (3D).torch.transpose,torch.inverse(when square),torch.trace,torch.eigand friends. Check the docs for deprecation status before relying on older solvers.3
8Random seeds
torch.manual_seed(42)
if torch.cuda.is_available(): torch.cuda.manual_seed_all(42)
9Autograd and gradients: the core
Enable tracking:
x = torch.ones(2,2, requires_grad=True).Forward: compute a loss scalar,
loss = some_fn(x).Backward:
loss.backward()computes gradients; checkx.grad.Zero grads before each optimizer step:
optimizer.zero_grad()ormodel.zero_grad().Inference / disable grad: use
with torch.no_grad():ortorch.set_grad_enabled(False).Detach from graph:
y = x.detach(), orx.detach().cpu().numpy().
z = mean((x+2)(x+5)+3) on a 2×2 tensor of ones. Analytic check: dz/dx = (2x+7)/4 = 2.25 per element, exactly what x.grad reports after z.backward().LessonAutograd is bookkeeping, not magic: every operation records where it came from so backward can walk the same path in reverse. Detach when you mean it, zero grads when you loop, and never assume a gradient exists for a tensor you built with
requires_grad=False.
10Common neural net building blocks
Layers:
torch.nn.Linear,torch.nn.Conv2d,torch.nn.Embedding,torch.nn.LSTM, and friends.Losses:
torch.nn.CrossEntropyLoss(),torch.nn.MSELoss(),torch.nn.BCEWithLogitsLoss().Activations:
torch.nn.ReLU(),torch.nn.Sigmoid(),torch.nn.Softmax(dim=1), or functionalF.relu.
import torch.nn as nn
class MLP(nn.Module):
def __init__(self, in_dim, hidden, out_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, hidden),
nn.ReLU(),
nn.Linear(hidden, out_dim)
)
def forward(self, x):
return self.net(x)
11Training loop: the minimal pattern
model.train()
for epoch in range(epochs):
for xb, yb in dataloader:
xb, yb = xb.to(device), yb.to(device)
preds = model(xb)
loss = criterion(preds, yb)
optimizer.zero_grad()
loss.backward()
optimizer.step()
For evaluation:
model.eval()
with torch.no_grad():
for xb, yb in val_loader:
preds = model(xb)
# compute metrics
12Optimizers and schedulers
Common optimizers:
torch.optim.SGD(model.parameters(), lr=...),torch.optim.Adam(...).Zero grads:
optimizer.zero_grad(); step:optimizer.step().LR schedulers:
torch.optim.lr_scheduler.StepLR,ReduceLROnPlateau, and friends. For what the step actually updates, see the companion note on optimizer state.
13Save and load
Save weights:
torch.save(model.state_dict(), "m.pt").
model = MyModel(...)
model.load_state_dict(torch.load("m.pt", map_location=device))
model.to(device)
Saving the whole model (
torch.save(model, "full_model.pt")thentorch.load(...)) works but is less recommended: it pickles the module object and couples your checkpoint to source layout.
14Data pipeline
Implement
torch.utils.data.Dataset(implement__len__and__getitem__).Wrap in
DataLoader(dataset, batch_size=..., shuffle=True, num_workers=4).Use
torchvision.transformsfor images ortorchtext/ custom transforms for text.For reproducibility, be careful with
num_workers>0together with random seeds: worker processes need their own seeding strategy.
15Useful tensor utilities
torch.arange(start, end), like NumPy arange;torch.linspace(start, end, steps).torch.where(cond, x, y): elementwise condition.torch.topk(x, k),torch.argmax(x, dim),torch.argsort.x.view(-1),x.flatten().x.expand(...)vsx.repeat(...): repeat creates new memory, expand is view-like.x.clone()to copy;x.cpu().numpy()to get a NumPy array (afterdetach()if it requires grad).
16Memory and performance tips
Prefer
torch.no_grad()when evaluating: it saves memory as well as time.Move whole batches to the device once, not element by element.
Avoid frequent
.cpu()/.numpy()calls inside training loops; they force syncs.Use
pin_memory=Truein DataLoader for faster CPU-to-GPU transfers during GPU training.For multi-GPU training consider
torch.nn.DataParallelortorch.nn.parallel.DistributedDataParallel(the latter is the modern default).
LessonThe GPU starves on transfers, not on math (see vectorization and coalescing for why). Every sync point you remove from the inner loop is free speed.
17Pitfalls and gotchas
| Mistake | Symptom | Fix |
|---|---|---|
| Mixing CPU and CUDA tensors | RuntimeError about device mismatch | move both operands with .to(device) |
Skipping optimizer.zero_grad() | gradients accumulate across batches | zero before every loss.backward() |
Calling .item() on non-scalar tensors | error at runtime | index or reduce to a single element first |
| In-place ops on tensors that require grad | autograd errors or silent corruption | prefer out-of-place ops; know when in-place breaks the graph |
| Normalizing using whole-dataset statistics | silent data leakage | compute stats on the training fold only |
18Short example snippets
Create a tensor, reshape, and sum:
x = torch.arange(6).view(2,3) # [[0,1,2],[3,4,5]]
col_sum = x.sum(dim=0) # collapse rows: [3, 5, 7]
row_sum = x.sum(dim=1) # collapse columns: [3, 12]
sum(dim=0) collapses rows and leaves one value per column; sum(dim=1) collapses columns and leaves one value per row.Gradients:
x = torch.ones(2,2, requires_grad=True)
y = (x + 2) * (x + 5) + 3
z = y.mean()
z.backward()
print(x.grad) # gradient of z wrt x, all entries 2.25
Batch matrix multiply:
a = torch.rand(3,4,5)
b = torch.rand(3,5,4)
c = torch.bmm(a,b) # shape (3, 4, 4)
19References
- PyTorch documentation: stable API reference.
- Official PyTorch tutorials: tensors, autograd, datasets.
- torch.linalg: the current linear algebra namespace.
- Local source:
bin/blogs/pytorch-basics.md, the original markdown note this page was rewritten from.