PyTorch models are usually built with nn.Module. A module owns parameters, defines a forward pass, and can contain other modules. Once you understand modules, loss functions, and optimizers, most training scripts become readable.
This lesson fills the gap between tensors/autograd and full training loops. You will learn what model.parameters() returns, why losses must be scalar values, how optimizers update weights, and how parameter groups let you train different parts of a model differently.
Every layer assigned as an attribute of an nn.Module is registered automatically. That registration is why model.parameters(), model.to(device), model.train(), model.eval(), and state_dict() work across nested modules.
Classification, regression, ranking, segmentation, and language modeling use different loss functions. The optimizer then uses gradients from that loss to update model weights. A wrong loss function can make a correct architecture fail.
A readable forward pass makes debugging much easier.
import torch
from torch import nn
class TabularClassifier(nn.Module):
def __init__(self, input_dim: int, num_classes: int):
super().__init__()
self.network = nn.Sequential(
nn.Linear(input_dim, 128),
nn.ReLU(),
nn.Dropout(p=0.2),
nn.Linear(128, 64),
nn.ReLU(),
nn.Linear(64, num_classes),
)
def forward(self, x):
# x: [batch_size, input_dim]
return self.network(x) # logits: [batch_size, num_classes]
model = TabularClassifier(input_dim=20, num_classes=4)
logits = model(torch.randn(8, 20))
print(logits.shape)
Parameter groups let you apply different learning rates or weight decay to different parts of a model.
optimizer = torch.optim.AdamW([
{"params": model.network[0].parameters(), "lr": 1e-4},
{"params": model.network[3:].parameters(), "lr": 3e-4},
], weight_decay=1e-4)
loss_fn = nn.CrossEntropyLoss()
features = torch.randn(16, 20)
labels = torch.randint(0, 4, (16,))
optimizer.zero_grad(set_to_none=True)
logits = model(features)
loss = loss_fn(logits, labels)
loss.backward()
optimizer.step()
Because nn.Module tracks registered parameters and buffers recursively.
Yes. Many models return logits plus auxiliary outputs, but your loss and training loop must handle that structure explicitly.
Compare named model parameters with the tensors present in the optimizer parameter groups.
Finish the concept here, then reinforce it with hands-on coding, interview prep, or a tool that matches the topic.
Explore 500+ free tutorials across 20+ languages and frameworks.
Fresh tutorials, interview guides, and coding practice in your inbox.