and

#tags.txt B-PER B-LOC …


### Structure of the dataset

- Download the original version on the [Kaggle](https://www.kaggle.com/abhinavwalia95/entity-annotated-corpus/data) website.

- **Download the dataset:** `ner_dataset.csv` on [Kaggle](https://www.kaggle.com/abhinavwalia95/entity-annotated-corpus/data) and save it under the `nlp/data/kaggle` directory. Make sure you download the simple version `ner_dataset.csv` and NOT the full version `ner.csv`.

- **Build the dataset:** Run the following script:

```python
python build_kaggle_dataset.py
kaggle/
    train/
        sentences.txt
        labels.txt
    test/
        sentences.txt
        labels.txt
    dev/
        sentences.txt
        labels.txt
python build_vocab.py --data_dir  data/small
python build_vocab.py --data_dir data/kaggle

Loading text data

vocab = {}
with open(words_path) as f:
    for i, l in enumerate(f.read().splitlines()):
        vocab[l] = i
train_sentences = []        
train_labels = []

with open(train_sentences_file) as f:
    for sentence in f.read().splitlines():
        # replace each token by its index if it is in vocab
        # else use index of UNK
        s = [vocab[token] if token in self.vocab 
            else vocab['UNK']
            for token in sentence.split(' ')]
        train_sentences.append(s)

with open(train_labels_file) as f:
    for sentence in f.read().splitlines():
        # replace each label by its index
        l = [tag_map[label] for label in sentence.split(' ')]
        train_labels.append(l)  

Preparing a Batch

# compute length of longest sentence in batch
batch_max_len = max([len(s) for s in batch_sentences])

# prepare a numpy array with the data, initializing the data with 'PAD' 
# and all labels with -1; initializing labels to -1 differentiates tokens 
# with tags from 'PAD' tokens
batch_data = vocab['PAD']*np.ones((len(batch_sentences), batch_max_len))
batch_labels = -1*np.ones((len(batch_sentences), batch_max_len))

# copy the data to the numpy array
for j in range(len(batch_sentences)):
    cur_len = len(batch_sentences[j])
    batch_data[j][:cur_len] = batch_sentences[j]
    batch_labels[j][:cur_len] = batch_tags[j]

# since all data are indices, we convert them to torch LongTensors
batch_data, batch_labels = torch.LongTensor(batch_data), torch.LongTensor(batch_labels)

# convert Tensors to Variables
batch_data, batch_labels = Variable(batch_data), Variable(batch_labels)
# train_data contains train_sentences and train_labels
# params contains batch_size
train_iterator = data_iterator(train_data, params, shuffle=True)    

for _ in range(num_training_steps):
    batch_sentences, batch_labels = next(train_iterator)

    # pass through model, perform backpropagation and updates
    output_batch = model(train_batch)
    ...

Recurrent network model

import torch.nn as nn
import torch.nn.functional as F

class Net(nn.Module):
    def __init__(self, params):
        super(Net, self).__init__()

    # maps each token to an embedding_dim vector
    self.embedding = nn.Embedding(params.vocab_size, params.embedding_dim)

    # the LSTM takens embedded sentence
    self.lstm = nn.LSTM(params.embedding_dim, params.lstm_hidden_dim, batch_first=True)

    # FC layer transforms the output to give the final output layer
    self.fc = nn.Linear(params.lstm_hidden_dim, params.number_of_tags)
def forward(self, s):
    # apply the embedding layer that maps each token to its embedding
    s = self.embedding(s)   # dim: batch_size x batch_max_len x embedding_dim

    # run the LSTM along the sentences of length batch_max_len
    s, _ = self.lstm(s)     # dim: batch_size x batch_max_len x lstm_hidden_dim                

    # reshape the Variable so that each row contains one token
    s = s.view(-1, s.shape[2])  # dim: batch_size*batch_max_len x lstm_hidden_dim

    # apply the fully connected layer and obtain the output for each token
    s = self.fc(s)          # dim: batch_size*batch_max_len x num_tags

    return F.log_softmax(s, dim=1)   # dim: batch_size*batch_max_len x num_tags

Writing a custom loss function

def loss_fn(outputs, labels):
    # reshape labels to give a flat vector of length batch_size*seq_len
    labels = labels.view(-1)  

    # mask out 'PAD' tokens
    mask = (labels >= 0).float()

    # the number of tokens is the sum of elements in mask
    num_tokens = int(torch.sum(mask).data[0])

    # pick the values corresponding to labels and multiply by mask
    outputs = outputs[range(outputs.shape[0]), labels]*mask

    # cross entropy loss for all non 'PAD' tokens
    return -torch.sum(outputs)/num_tokens

Selected methods

Tensor shape/size

import torch

a = torch.randn(2, 3, 5)

# Get the overall shape of the tensor
a.size()   # Prints torch.Size([2, 3, 5])
a.shape    # Prints torch.Size([2, 3, 5])

# Get the size of a specific axis/dimension of the tensor
a.size(2)  # Prints 5
a.shape[2] # Prints 5

Initialization

Static

import torch.nn as nn

a = torch.empty(3, 5)
nn.init.zeros_(a)         # Initializes a with 0
nn.init.ones_(a)          # Initializes a with 1
nn.init.constant_(a, 0.3) # Initializes a with 0.3

Standard normal

\[\text{out}_{i} \sim \mathcal{N}(0, 1)\]
import torch

torch.randn(4)    # Returns 4 values from the standard normal distribution
torch.randn(2, 3) # Returns a 2x3 matrix sampled from the standard normal distribution

Xavier/Glorot

Uniform
\[a = \text{gain} \times \sqrt{\frac{6}{\text{fan_in} + \text{fan_out}}}\]
import torch.nn as nn

a = torch.empty(3, 5)
nn.init.xavier_uniform_(a, gain=nn.init.calculate_gain('relu')) # Initializes a with the Xavier uniform method
Normal
\[\text{std} = \text{gain} \times \sqrt{\frac{2}{\text{fan_in} + \text{fan_out}}}\]
import torch.nn as nn

a = torch.empty(3, 5)
nn.init.xavier_normal_(a) # Initializes a with the Xavier normal method

Kaiming/He

Uniform
\[\text{bound} = \text{gain} \times \sqrt{\frac{3}{\text{fan_mode}}}\]
import torch.nn as nn

a = torch.empty(3, 5)
nn.init.kaiming_uniform_(a, mode='fan_in', nonlinearity='relu') # Initializes a with the Kaiming uniform method 
Normal
\[\operatorname{std}=\frac{\text { gain }}{\sqrt{\text { fan_mode }}}\]
import torch.nn as nn

a = torch.empty(3, 5)
nn.init.kaiming_normal_(a, mode='fan_out', nonlinearity='relu') # Initializes a with the Kaiming uniform method 

Send Tensor to GPU

import torch

t = torch.tensor([1, 2, 3])
a = t.cuda()
type(a) # Prints <class 'numpy.ndarray'>

# Send tensor to the GPU
a = a.cuda()

# Bring the tensor back to the CPU
a = a.cpu()
if cuda_available:
    x = x.cuda()
    model.cuda()
else:
    x = x.cpu()
    model.cpu()
device = torch.device('cuda') if cuda_available else torch.device('cpu')
x = x.to(device)
model = model.to(device)

Convert to NumPy

import torch

t = torch.tensor([1, 2, 3])
a = t.numpy()               # array([1, 2, 3])
type(a)                     # Prints <class 'numpy.ndarray'>

# Send tensor to the GPU.
t = t.cuda()

b = t.cpu().numpy()          # array([1, 2, 3])
type(b)                      # <class 'numpy.ndarray'>
import torch

t = torch.tensor([1, 2, 3], requires_grad=True)
a = t.detach().numpy()       # array([1, 2, 3])
type(a)                      # Prints <class 'numpy.ndarray'>

# Send tensor to the GPU.
t = t.cuda()

# The output of the line below is a NumPy array.
b = t.detach().cpu().numpy() # array([1, 2, 3])
type(b)                      # <class 'numpy.ndarray'>

tensor.item(): Convert Single Value Tensor to Scalar

import torch

a = torch.tensor([1.0])
a.item()   # Prints 1.0

a.tolist() # Prints [1.0]

tensor.tolist(): Convert Multi Value Tensor to Scalar

a = torch.randn(2, 2)
a.tolist()      # Prints [[0.012766935862600803, 0.5415473580360413],
                #         [-0.08909505605697632, 0.7729271650314331]]
a[0,0].tolist() # Prints 0.012766935862600803

Len

import torch

a = torch.Tensor([[1, 2], [3, 4]])
print(a) # Prints tensor([[1., 2.],
         #                [3., 4.]])
len(a)   # 2

b = torch.Tensor([1, 2, 3, 4])
print(b) # Prints tensor([1., 2., 3., 4.])
len(b)   # 4

Arange

import torch

print(torch.arange(8))             # Prints tensor([0 1 2 3 4 5 6 7])
print(torch.arange(3, 8))          # Prints tensor([3 4 5 6 7])
print(torch.arange(3, 8, 2))       # Prints tensor([3 5 7])

# arange() works with floats too (but read the disclaimer below)
print(torch.arange(0.1, 0.5, 0.1)) # Prints tensor([0.1000, 0.2000, 0.3000, 0.4000])

Linspace

import torch

print(torch.linspace(1.0, 2.0, steps=5)) # Prints tensor([1.0000, 1.2500, 1.5000, 1.7500, 2.0000])

View

import torch

a = torch.arange(4).view(2, 2)

print(a.view(4, 1)) # Prints tensor([[0],
                    #                [1],
                    #                [2],
                    #                [3]])

print(a.view(1, 4)) # Prints tensor([[0, 1, 2, 3]])
import torch

a = torch.arange(4).view(2, 2)
print(a.view(-1)) # Prints tensor([0, 1, 2, 3])
import torch

a = torch.rand(4, 4)
b = a.view(2, 8)
a.storage().data_ptr() == b.storage().data_ptr() # Prints True since `a` and `b` share the same underlying data.
import torch

a = torch.rand(4, 4)
b = a.view(2, 8)
b[0][0] = 3.14

print(t[0][0]) # Prints tensor(3.14)

Transpose

import torch

a = torch.randn(2, 3, 5)
a.size()                 # Prints torch.Size([2, 3, 5])

a.transpose(0, -1).shape # Prints torch.Size([5, 3, 2])

Swapaxes

import torch

a = torch.randn(2, 3, 5)
a.size()                # Prints torch.Size([2, 3, 5])

a.swapdims(0, -1).shape # Prints torch.Size([5, 3, 2])

# swapaxes is an alias of swapdims
a.swapaxes(0, -1).shape # Prints torch.Size([5, 3, 2])

Permute

import torch

a = torch.randn(2, 3, 5)
a.size()                  # Prints torch.Size([2, 3, 5])

a.permute(2, 0, 1).size() # Prints torch.Size([5, 2, 3])
a = torch.tensor([[1, 2, 3], [4, 5, 6]])

viewed = a.view(3, 2)
perm = a.permute(1, 0)

viewed.shape   # Prints torch.Size([3, 2])
perm.shape     # Prints torch.Size([3, 2])

viewed == perm # Prints tensor([[ True, False],
               #                [False, False],
               #                [False,  True]])

viewed         # Prints tensor([[1, 2],
               #                [3, 4],
               #                [5, 6]])

perm           # Prints tensor([[1, 4],
               #                [2, 5],
               #                [3, 6]])

Movedim

import torch

a = torch.randn(2, 3, 5)
a.size()                # Prints torch.Size([2, 3, 5])

a.movedim(0, -1).shape  # Prints torch.Size([3, 5, 2])

# moveaxis is an alias of movedim
a.moveaxis(0, -1).shape # Prints torch.Size([3, 5, 2])

Randperm

import torch

torch.randperm(n=4) # Prints tensor([2, 1, 0, 3])
data[torch.randperm(data.shape[0])] # Assuming the first dimension of data is the minibatch number

Where

\[\text{out}_i = \begin{cases} \text{x}_i & \text{if } \text{condition}_i \\ \text{y}_i & \text{otherwise} \\ \end{cases}\]
import torch

a = torch.randn(3, 2) # Initializes a as a 3x2 matrix using the the standard normal distribution
b = torch.ones(3, 2)
>>> a
tensor([[-0.4620,  0.3139],
        [ 0.3898, -0.7197],
        [ 0.0478, -0.1657]])
>>> torch.where(a > 0, a, b)
tensor([[ 1.0000,  0.3139],
        [ 0.3898,  1.0000],
        [ 0.0478,  1.0000]])
>>> a = torch.randn(2, 2, dtype=torch.double)
>>> a
tensor([[ 1.0779,  0.0383],
        [-0.8785, -1.1089]], dtype=torch.float64)
>>> torch.where(a > 0, a, 0.)
tensor([[1.0779, 0.0383],
        [0.0000, 0.0000]], dtype=torch.float64)

Reshape

import torch

a = torch.arange(4*10*2).view(4, 10, 2)
b = x.permute(2, 0, 1)

# Reshape works on non-contiguous tensors (contiguous() + view())
print(b.is_contiguous())
try: 
    print(b.view(-1))
except RuntimeError as e:
    print(e)
print(b.reshape(-1))
print(b.contiguous().view(-1))

Concatenate

import torch

x = torch.randn(2, 3)
print(x) # Prints a 2x3 matrix: [[ 0.6580, -1.0969, -0.4614],
         #                       [-0.1034, -0.5790,  0.1497]]

print(torch.cat((x, x, x), 0)) # Prints a 6x3 matrix: [[ 0.6580, -1.0969, -0.4614],
                               #                       [-0.1034, -0.5790,  0.1497],
                               #                       [ 0.6580, -1.0969, -0.4614],
                               #                       [-0.1034, -0.5790,  0.1497],
                               #                       [ 0.6580, -1.0969, -0.4614],
                               #                       [-0.1034, -0.5790,  0.1497]]

print(torch.cat((x, x, x), 1)) # Prints a 2x9 matrix: [[ 0.6580, -1.0969, -0.4614,  
                               #                         0.6580, -1.0969, -0.4614,  
                               #                         0.6580, -1.0969, -0.4614],
                               #                       [-0.1034, -0.5790,  0.1497, 
                               #                        -0.1034, -0.5790,  0.1497, 
                               #                        -0.1034, -0.5790,  0.1497]]

Squeeze

import torch

a = torch.zeros(2, 1, 2, 1, 2)
print(a.size()) # Prints torch.Size([2, 1, 2, 1, 2])

b = torch.squeeze(a)
print(b.size()) # Prints torch.Size([2, 2, 2])

b = torch.squeeze(a, 0)
print(b.size()) # Prints torch.Size([2, 1, 2, 1, 2])

b = torch.squeeze(a, 1)
print(b.size()) # Prints torch.Size([2, 2, 1, 2])

Unsqueeze

import torch

a = torch.tensor([1, 2, 3, 4])
print(x.size()) # Prints torch.Size([4])

b = torch.unsqueeze(a, 0) 
print(b)        # Prints tensor([[1, 2, 3, 4]])
print(b.size()) # Prints torch.Size([1, 4])

b = torch.unsqueeze(a, 1)
print(b)        # Prints tensor([[1],
                #                [2],
                #                [3],
                #                [4]])
print(b.size()) # torch.Size([4, 1])
import torch

# 3 channels, 32 width, 32 height
a = torch.randn(3, 32, 32)

# 1 batch, 3 channels, 32 width, 32 height
a.unsqueeze(dim=0).shape
from torchvision import models
model = models.vgg16()
print(model)
VGG (
  (features): Sequential (
    (0): Conv2d(3, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): ReLU (inplace)
    (2): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (3): ReLU (inplace)
    (4): MaxPool2d (size=(2, 2), stride=(2, 2), dilation=(1, 1))
    (5): Conv2d(64, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (6): ReLU (inplace)
    (7): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (8): ReLU (inplace)
    (9): MaxPool2d (size=(2, 2), stride=(2, 2), dilation=(1, 1))
    (10): Conv2d(128, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (11): ReLU (inplace)
    (12): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (13): ReLU (inplace)
    (14): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (15): ReLU (inplace)
    (16): MaxPool2d (size=(2, 2), stride=(2, 2), dilation=(1, 1))
    (17): Conv2d(256, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (18): ReLU (inplace)
    (19): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (20): ReLU (inplace)
    (21): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (22): ReLU (inplace)
    (23): MaxPool2d (size=(2, 2), stride=(2, 2), dilation=(1, 1))
    (24): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (25): ReLU (inplace)
    (26): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (27): ReLU (inplace)
    (28): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (29): ReLU (inplace)
    (30): MaxPool2d (size=(2, 2), stride=(2, 2), dilation=(1, 1))
  )
  (classifier): Sequential (
    (0): Dropout (p = 0.5)
    (1): Linear (25088 -> 4096)
    (2): ReLU (inplace)
    (3): Dropout (p = 0.5)
    (4): Linear (4096 -> 4096)
    (5): ReLU (inplace)
    (6): Linear (4096 -> 1000)
  )
)
from torchvision import models
from torchsummary import summary

# Example for VGG16
vgg = models.vgg16()
summary(vgg, (3, 224, 224))
================================================================
        Layer (type)               Output Shape         Param #
================================================================
            Conv2d-1         [-1, 64, 224, 224]           1,792
              ReLU-2         [-1, 64, 224, 224]               0
            Conv2d-3         [-1, 64, 224, 224]          36,928
              ReLU-4         [-1, 64, 224, 224]               0
         MaxPool2d-5         [-1, 64, 112, 112]               0
            Conv2d-6        [-1, 128, 112, 112]          73,856
              ReLU-7        [-1, 128, 112, 112]               0
            Conv2d-8        [-1, 128, 112, 112]         147,584
              ReLU-9        [-1, 128, 112, 112]               0
        MaxPool2d-10          [-1, 128, 56, 56]               0
           Conv2d-11          [-1, 256, 56, 56]         295,168
             ReLU-12          [-1, 256, 56, 56]               0
           Conv2d-13          [-1, 256, 56, 56]         590,080
             ReLU-14          [-1, 256, 56, 56]               0
           Conv2d-15          [-1, 256, 56, 56]         590,080
             ReLU-16          [-1, 256, 56, 56]               0
        MaxPool2d-17          [-1, 256, 28, 28]               0
           Conv2d-18          [-1, 512, 28, 28]       1,180,160
             ReLU-19          [-1, 512, 28, 28]               0
           Conv2d-20          [-1, 512, 28, 28]       2,359,808
             ReLU-21          [-1, 512, 28, 28]               0
           Conv2d-22          [-1, 512, 28, 28]       2,359,808
             ReLU-23          [-1, 512, 28, 28]               0
        MaxPool2d-24          [-1, 512, 14, 14]               0
           Conv2d-25          [-1, 512, 14, 14]       2,359,808
             ReLU-26          [-1, 512, 14, 14]               0
           Conv2d-27          [-1, 512, 14, 14]       2,359,808
             ReLU-28          [-1, 512, 14, 14]               0
           Conv2d-29          [-1, 512, 14, 14]       2,359,808
             ReLU-30          [-1, 512, 14, 14]               0
        MaxPool2d-31            [-1, 512, 7, 7]               0
           Linear-32                 [-1, 4096]     102,764,544
             ReLU-33                 [-1, 4096]               0
          Dropout-34                 [-1, 4096]               0
           Linear-35                 [-1, 4096]      16,781,312
             ReLU-36                 [-1, 4096]               0
          Dropout-37                 [-1, 4096]               0
           Linear-38                 [-1, 1000]       4,097,000
================================================================
Total params: 138,357,544
Trainable params: 138,357,544
Non-trainable params: 0
-
Input size (MB): 0.57
Forward/backward pass size (MB): 218.59
Params size (MB): 527.79
Estimated Total Size (MB): 746.96
-

Got it. Below is the same section, but with inline end-of-line comments (aligned and indented in the same style as the rest of the primer), instead of block comments.

You can drop this directly into the document.

Keepdim (Keeping Reduced Dimensions)

Basic Syntax

torch.sum(input, dim=..., keepdim=True)
torch.mean(input, dim=..., keepdim=True)
torch.max(input, dim=..., keepdim=True)

Example 1: Sum With and Without keepdim

import torch

x = torch.tensor([[1., 2., 3.],
                  [4., 5., 6.]])       # tensor([[1., 2., 3.],
                                       #         [4., 5., 6.]])

s1 = torch.sum(x, dim=1)               # tensor([ 6., 15.])
s1.shape                               # torch.Size([2])

s2 = torch.sum(x, dim=1, keepdim=True) # tensor([[ 6.],
                                       #         [15.]])
s2.shape                               # torch.Size([2, 1])

Example 2: Mean and Broadcasting

x = torch.randn(4, 5)                   # shape: torch.Size([4, 5])

mean_no_keep = x.mean(dim=1)            # shape: torch.Size([4])
mean_keep = x.mean(dim=1, keepdim=True) # shape: torch.Size([4, 1])

y1 = x - mean_no_keep                   # RuntimeError (shape mismatch)
y2 = x - mean_keep                      # shape: torch.Size([4, 5])

Example 3: Max Values and Indices

x = torch.tensor([[1., 7., 3.],
                  [4., 2., 9.]]) # tensor([[1., 7., 3.],
                                 #         [4., 2., 9.]])

values, indices = torch.max(x, dim=1, keepdim=True)

values                           # tensor([[7.],
                                 #         [9.]])
indices                          # tensor([[1],
                                 #         [2]])

Example 4: keepdim vs. unsqueeze

a = x.sum(dim=1, keepdim=True) # shape: torch.Size([2, 1])
b = x.sum(dim=1).unsqueeze(1)  # shape: torch.Size([2, 1])

End-to-End Data to Model Pipeline

Data Pre-processing

Overview

Key Abstractions for Pre-processing

torch.utils.data.Dataset
torchtext.transforms
from torchtext import transforms
from torchtext.vocab import build_vocab_from_iterator

# Custom augmentation functions
class RandomWordDropout:
    """Randomly drops words with a given probability."""
    def __init__(self, p=0.1):
        self.p = p

    def __call__(self, tokens):
        return [tok for tok in tokens if random.random() > self.p]

class SynonymReplacement:
    """Simple synonym replacement using a lookup dictionary."""
    def __init__(self, synonym_dict, p=0.1):
        self.synonym_dict = synonym_dict
        self.p = p

    def __call__(self, tokens):
        augmented = []
        for tok in tokens:
            if tok in self.synonym_dict and random.random() < self.p:
                augmented.append(random.choice(self.synonym_dict[tok]))
            else:
                augmented.append(tok)
        return augmented

# Example synonym mapping
synonyms = {
    "good": ["great", "excellent", "nice"],
    "bad": ["terrible", "awful", "poor"]
}

# Suppose 'vocab' is built from your corpus
vocab = build_vocab_from_iterator([["this", "is", "a", "good", "sample"]], specials=["<unk>", "<pad>"])
vocab.set_default_index(vocab["<unk>"])

text_transform = transforms.Sequential(
    # Step 1: Randomly replace words with synonyms (30% probability)
    SynonymReplacement(synonyms, p=0.3),

    # Step 2: Randomly drop words from the sequence (20% probability)
    RandomWordDropout(p=0.2),

    # Step 3: Convert tokens into integer IDs using the predefined vocabulary
    transforms.VocabTransform(vocab),

    # Step 4: Truncate sequences longer than 512 tokens to a fixed maximum length
    transforms.Truncate(512),

    # Step 5: Convert the list of token IDs into a tensor and pad to uniform length
    transforms.ToTensor(padding_value=vocab['<pad>'])
)

tokens = ["this", "is", "a", "sample"]
tensorized = text_transform(tokens)
torchvision.transforms
from torchvision import transforms

image_transform = transforms.Compose([
    # Step 1: Resize the input image so the shorter side is 256 pixels
    transforms.Resize(256),

    # Step 2: Crop the central 224×224 region from the resized image
    transforms.CenterCrop(224),

    # Step 3: Randomly flip the image horizontally (helps with data augmentation)
    transforms.RandomHorizontalFlip(),

    # Step 4: Convert the PIL image to a PyTorch tensor and scale pixel values to [0, 1]
    transforms.ToTensor(),

    # Step 5: Normalize tensor using ImageNet channel means and standard deviations
    #         (this standardization helps models pretrained on ImageNet converge better)
    transforms.Normalize(mean=[0.485, 0.456, 0.406],
                         std=[0.229, 0.224, 0.225])
])
torchaudio.transforms
import torchaudio
from torchaudio import transforms as T

waveform, sample_rate = torchaudio.load("speech.wav")

audio_transform = T.Compose([
    # Step 1: Resample the raw audio waveform to a consistent 16 kHz sampling rate
    T.Resample(orig_freq=sample_rate, new_freq=16000),

    # Step 2: Convert the waveform into a Mel-spectrogram (frequency–time representation)
    # Uses 64 Mel filter banks to capture perceptually relevant frequency information
    T.MelSpectrogram(sample_rate=16000, n_mels=64),

    # Step 3: Apply frequency masking (randomly masks frequency bands)
    # Helps the model generalize to variations in spectral features
    T.FrequencyMasking(freq_mask_param=15),

    # Step 4: Apply time masking (randomly masks time segments)
    # Improves robustness to temporal distortions or missing frames
    T.TimeMasking(time_mask_param=35),

    # Step 5: Convert the Mel-spectrogram power values to decibel (dB) scale
    # Produces log-scaled features commonly used in speech and audio models
    T.AmplitudeToDB()
])

mel_spectrogram = audio_transform(waveform)
torch.utils.data.DataLoader

Pre-processing for Vision Data

Loading and Cleaning
Normalization
mean = [0.485, 0.456, 0.406]
std  = [0.229, 0.224, 0.225]
Data Augmentation
Conversion to Tensor

Pre-processing for Text Data

Tokenization
Vocabulary Building
{'<PAD>':0, '<UNK>':1, 'pytorch':2, 'is':3, 'awesome':4, '!':5}
Handling Variable-Length Sequences
Embedding Lookup
nn.Embedding(num_embeddings, embedding_dim)
Masking

Pre-processing for Audio Data

Loading and Cleaning
import torchaudio

waveform, sample_rate = torchaudio.load("speech.wav")
waveform = torchaudio.functional.resample(waveform, orig_freq=sample_rate, new_freq=16000)
Normalization (Standardization)
Feature Extraction
from torchaudio import transforms as T

audio_transform = T.MelSpectrogram(
    sample_rate=16000,
    n_mels=64,
    n_fft=1024,
    hop_length=256
)
mel_spectrogram = audio_transform(waveform)
Data Augmentation
from torchaudio import transforms as T

augment = T.Compose([
    T.FrequencyMasking(freq_mask_param=15),
    T.TimeMasking(time_mask_param=35),
    T.Vol(gain=(0.5, 1.5))
])
augmented_spectrogram = augment(mel_spectrogram)
Conversion to Tensor and Batching
Integration with Models

Data Quality Checks

Tabular Summary

Modality Key Steps Common Transforms Notes
Text Tokenize $$\rightarrow$$ Augment $$\rightarrow$$ Numericalize $$\rightarrow$$ Pad $$\rightarrow$$ Mask SynonymReplacement, RandomWordDropout, NoiseInjection, VocabTransform, Truncate, PadTransform Apply augmentation before numericalization; handle OOVs with <UNK>
Vision Resize $$\rightarrow$$ Normalize $$\rightarrow$$ Augment $$\rightarrow$$ Tensor Resize, Normalize, ToTensor, RandomCrop Use per-channel normalization
Audio Load $$\rightarrow$$ Resample $$\rightarrow$$ Feature Extract $$\rightarrow$$ Augment $$\rightarrow$$ Normalize MelSpectrogram, AmplitudeToDB, FrequencyMasking, TimeMasking Ensure consistent sampling rate and duration
Shared Cleaning, batching, shuffling Dataset, DataLoader Use multiprocessing in DataLoader

FAQs

References and Further Reading

Practical Implementation – Data Pre-processing

NLP Data Pre-processing – Text Classification Example
import torch
from torch.utils.data import Dataset, DataLoader
from torchtext.vocab import build_vocab_from_iterator
from torch.nn.utils.rnn import pad_sequence
import re

# ----------------------------------------------------------
# 1. Tokenizer
# ----------------------------------------------------------
# - Purpose: Convert raw text into a list of clean tokens (words).
# - Steps:
#     * Convert text to lowercase for uniformity.
#     * Remove all punctuation and special characters using regex.
#     * Split text into individual word tokens (by whitespace).
def tokenize(text):
    text = re.sub(r"[^a-zA-Z0-9\s]", "", text.lower())
    return text.split()

# ----------------------------------------------------------
# 2. Sample corpus
# ----------------------------------------------------------
# - Example dataset for binary classification (e.g., sentiment, relevance).
# - Each text string corresponds to one data sample, with label 0 or 1.
texts = [
    "Large Language Models are powerful AI models",
    "Deep learning is transformational",
    "Neural networks can generalize",
    "PyTorch simplifies model training"
]
labels = [1, 0, 1, 0]

# ----------------------------------------------------------
# 3. Build vocabulary
# ----------------------------------------------------------
# - The vocabulary assigns each unique token an integer index.
# - This is critical for converting tokens into numerical form that models can process.
# - The special tokens:
#     * <pad>: used for sequence padding (makes sequences same length in a batch).
#     * <unk>: represents out-of-vocabulary tokens (unknown words).
vocab = build_vocab_from_iterator(map(tokenize, texts), specials=["<pad>", "<unk>"])
vocab.set_default_index(vocab["<unk>"])  # All unseen tokens map to <unk>

# ----------------------------------------------------------
# 4. Numericalization
# ----------------------------------------------------------
# - Converts a list of string tokens into a tensor of integer IDs using the vocabulary.
# - Example:
#   Input:  ["deep", "learning", "is", "transformational"]
#   Output: tensor([5, 9, 3, 8])  # based on vocab indices
def numericalize(text):
    return torch.tensor([vocab[token] for token in tokenize(text)], dtype=torch.long)

# ----------------------------------------------------------
# 5. Dataset definition
# ----------------------------------------------------------
# - Custom Dataset class wraps the tokenized text and corresponding labels.
# - Required methods:
#     * __getitem__(self, idx): returns one sample (text tensor, label tensor).
#     * __len__(self): returns total number of samples.
class TextDataset(Dataset):
    def __init__(self, texts, labels):
        self.texts = texts
        self.labels = labels

    def __getitem__(self, idx):
        # Returns a tuple: (numericalized text tensor, label tensor)
        return numericalize(self.texts[idx]), torch.tensor(self.labels[idx])

    def __len__(self):
        return len(self.texts)

# Instantiate dataset object
dataset = TextDataset(texts, labels)

# ----------------------------------------------------------
# 6. Collate function for padding
# ----------------------------------------------------------
# - Purpose: dynamically pad variable-length text sequences in a batch.
# - pad_sequence: ensures all sequences in the batch are the same length.
# - padding_value: index corresponding to <pad> token in the vocabulary.
def collate_fn(batch):
    # Unzip the batch (list of (text, label) pairs) into separate sequences and labels
    texts, labels = zip(*batch)
    # Pad sequences so all have equal length (batch_first=True $$\rightarrow$$ shape [batch, seq_len])
    padded = pad_sequence(texts, batch_first=True, padding_value=vocab["<pad>"])
    # Stack labels into a single tensor
    return padded, torch.stack(labels)

# ----------------------------------------------------------
# 7. DataLoader
# ----------------------------------------------------------
# - Combines dataset and collate function to create mini-batches.
# - Handles:
#     * Shuffling for random sampling each epoch.
#     * Batch creation.
#     * Optional parallel data loading via num_workers.
loader = DataLoader(dataset, batch_size=2, collate_fn=collate_fn, shuffle=True)

# ----------------------------------------------------------
# 8. Inspect one mini-batch
# ----------------------------------------------------------
# - Demonstrates the final structure after pre-processing.
# - Each batch has:
#     * x_batch: tensor of shape [batch_size, sequence_length]
#     * y_batch: tensor of shape [batch_size]
for x_batch, y_batch in loader:
    print(f"Input batch shape: {x_batch.shape}")  # e.g., (2, seq_len)
    print(f"Label batch shape: {y_batch.shape}")  # e.g., (2,)
    break
    
# Can also do the following in place of the above block:
# ----------------------------------------------------------
# 8. Inspect one mini-batch (using next(iter(loader)))
# ----------------------------------------------------------
# - Demonstrates how to manually fetch a single batch from a DataLoader.
# - This approach is equivalent to running one iteration of the loop.
# - Useful for debugging or inspecting batch structure and tensor shapes using one batch only.

# Create an iterator over the DataLoader
# batch_iter = iter(loader)

# Fetch the first batch (x_batch, y_batch)
# x_batch, y_batch = next(batch_iter)

# Inspect the shapes of tensors
# print(f"Input batch shape: {x_batch.shape}")   # e.g., (2, seq_len)
# print(f"Label batch shape: {y_batch.shape}")   # e.g., (2,)    
Vision Data Pre-processing – CIFAR-10 Example
import torch
from torchvision import datasets, transforms
from torch.utils.data import DataLoader

# ----------------------------------------------------------
# 1. Define data transformations
# ----------------------------------------------------------
# These define how images are preprocessed before feeding into the model.
# Training transforms include augmentations for better generalization,
# while validation/test transforms remain deterministic for fair evaluation.

train_transforms = transforms.Compose([
    transforms.RandomHorizontalFlip(),     # Randomly flip images horizontally (helps learn invariance)
    transforms.RandomRotation(10),         # Apply small random rotations to simulate varied orientations
    transforms.ToTensor(),                 # Convert PIL image $$\rightarrow$$ Tensor with shape (C, H, W), values in [0,1]
    transforms.Normalize(mean=(0.5, 0.5, 0.5),  # Normalize RGB channels (center around 0, scale to ~[-1,1])
                         std=(0.5, 0.5, 0.5))
])

test_transforms = transforms.Compose([
    transforms.ToTensor(),                 # Convert test images to tensor (no random augmentations)
    transforms.Normalize(mean=(0.5, 0.5, 0.5),
                         std=(0.5, 0.5, 0.5))
])

# ----------------------------------------------------------
# 2. Load dataset
# ----------------------------------------------------------
# Automatically downloads and prepares CIFAR-10 dataset if not already present.
# - train=True $$\rightarrow$$ loads training split (50,000 images)
# - train=False $$\rightarrow$$ loads test split (10,000 images)
# - transform=... $$\rightarrow$$ applies preprocessing pipeline on-the-fly
train_data = datasets.CIFAR10(root='data', train=True, download=True, transform=train_transforms)
test_data = datasets.CIFAR10(root='data', train=False, download=True, transform=test_transforms)

# ----------------------------------------------------------
# 3. Prepare DataLoaders
# ----------------------------------------------------------
# DataLoader efficiently handles batching, shuffling, and multiprocessing.
# - batch_size=64 $$\rightarrow$$ each batch contains 64 images
# - shuffle=True $$\rightarrow$$ reshuffles data at each epoch to reduce overfitting
# - num_workers=4 $$\rightarrow$$ uses 4 subprocesses for parallel data loading
train_loader = DataLoader(train_data, batch_size=64, shuffle=True, num_workers=4)
test_loader = DataLoader(test_data, batch_size=64, shuffle=False, num_workers=4)

# ----------------------------------------------------------
# 4. Inspect one batch
# ----------------------------------------------------------
# Retrieve a single mini-batch using an iterator.
# - iter(train_loader) returns an iterator over batches
# - next(...) yields the first batch (images and labels)
images, labels = next(iter(train_loader))

# Print tensor shapes for verification
# CIFAR-10 images: (64 samples, 3 color channels, 32x32 resolution)
print(f"Batch shape: {images.shape}")  # Expected $$\rightarrow$$ (64, 3, 32, 32)
print(f"Label shape: {labels.shape}")  # Expected $$\rightarrow$$ (64,)
FAQs

Model Training/Fine-tuning and Evaluation Workflow

Overview

Core Components of a Training Workflow

Example: NLP Model Training Workflow

import torch
import torch.nn as nn
import torch.optim as optim

# ----------------------------------------------------------
# 1. Define RNN-based text classification model
# ----------------------------------------------------------
class RNNClassifier(nn.Module):
    def __init__(self, vocab_size, embed_dim, hidden_dim, num_classes, pad_idx):
        super().__init__()
        # Embedding layer converts token IDs into dense vector representations
        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)

        # GRU captures temporal dependencies in token sequences
        self.rnn = nn.GRU(embed_dim, hidden_dim, batch_first=True)

        # Fully connected output layer for classification logits
        self.fc = nn.Linear(hidden_dim, num_classes)

    def forward(self, x):
        # x shape: (batch_size, seq_len)
        embedded = self.embedding(x)     # $$\rightarrow$$ (batch_size, seq_len, embed_dim)
        _, h = self.rnn(embedded)        # h: (1, batch_size, hidden_dim)
        return self.fc(h.squeeze(0))     # $$\rightarrow$$ (batch_size, num_classes)


# ----------------------------------------------------------
# 2. Instantiate model, loss, and optimizer
# ----------------------------------------------------------
vocab_size = len(vocab)  # Number of tokens in vocabulary
model = RNNClassifier(
    vocab_size=vocab_size,
    embed_dim=64,
    hidden_dim=128,
    num_classes=2,
    pad_idx=vocab["<pad>"]
)

criterion = nn.CrossEntropyLoss()               # Cross-entropy for classification
optimizer = optim.Adam(model.parameters(), lr=1e-3)  # Adaptive optimizer


# ----------------------------------------------------------
# 3. Define Evaluation Function (Validation Loop)
# ----------------------------------------------------------
def evaluate_nlp_model(model, loader, criterion):
    """
    Evaluates model performance on a validation or test DataLoader.
    Returns average loss and accuracy.
    """
    model.eval()  # Disable dropout, batchnorm updates
    total_loss, correct, total = 0.0, 0, 0

    with torch.no_grad():  # Disable gradient computation for faster inference
        for x_batch, y_batch in loader:
            outputs = model(x_batch)
            loss = criterion(outputs, y_batch)
            total_loss += loss.item()

            # Get predicted labels (index of max logit)
            preds = outputs.argmax(dim=1)
            correct += (preds == y_batch).sum().item()
            total += y_batch.size(0)

    avg_loss = total_loss / len(loader)
    accuracy = correct / total
    return avg_loss, accuracy


# ----------------------------------------------------------
# 4. Define Training Loop
# ----------------------------------------------------------
def train_nlp_model(model, train_loader, val_loader, criterion, optimizer, epochs=5):
    """
    Trains the NLP model and evaluates on validation data per epoch.
    """
    best_val_loss = float('inf')

    for epoch in range(epochs):
        model.train()  # Enable training mode
        total_loss = 0.0

        # Loop over mini-batches
        for x_batch, y_batch in train_loader:
            optimizer.zero_grad()
            outputs = model(x_batch)
            loss = criterion(outputs, y_batch)
            loss.backward()
            optimizer.step()
            total_loss += loss.item()

        # Compute average training loss
        avg_train_loss = total_loss / len(train_loader)

        # Evaluate on validation data after each epoch
        val_loss, val_acc = evaluate_nlp_model(model, val_loader, criterion)

        print(f"Epoch {epoch+1}: train_loss={avg_train_loss:.3f}, "
              f"val_loss={val_loss:.3f}, val_acc={val_acc:.3f}")

        # Save model if validation loss improves
        if val_loss < best_val_loss:
            best_val_loss = val_loss
            torch.save(model.state_dict(), "best_rnn_model.pt")
            print("✅ Saved best model checkpoint.")

Example: Vision Classification Training Loop

import torch
import torch.nn as nn
import torch.optim as optim

# ----------------------------------------------------------
# 1. Define a simple CNN model for CIFAR-10 classification
# ----------------------------------------------------------
class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        # Sequential container for defining the full network in order
        self.net = nn.Sequential(
            # First convolutional block:
            #   - Input: 3 channels (RGB images)
            #   - Output: 32 feature maps
            #   - Kernel: 3x3, stride=1, padding=1 (to preserve spatial size)
            nn.Conv2d(3, 32, 3, padding=1),
            nn.ReLU(),                # Activation adds non-linearity
            nn.MaxPool2d(2),          # Downsamples 32x32 $$\rightarrow$$ 16x16

            # Second convolutional block:
            #   - Input: 32 channels
            #   - Output: 64 feature maps
            nn.Conv2d(32, 64, 3, padding=1),
            nn.ReLU(),
            nn.MaxPool2d(2),          # Downsamples 16x16 $$\rightarrow$$ 8x8

            # Flatten layer: converts 2D feature maps $$\rightarrow$$ 1D vector for dense layers
            nn.Flatten(),

            # Fully connected (dense) layer:
            #   - Input: 64 * 8 * 8 = 4096 features
            #   - Output: 128 hidden units
            nn.Linear(64 * 8 * 8, 128),
            nn.ReLU(),

            # Output layer:
            #   - Input: 128 features
            #   - Output: 10 classes (CIFAR-10 has 10 categories)
            nn.Linear(128, 10)
        )

    def forward(self, x):
        # Defines how data flows through the network
        return self.net(x)


# ----------------------------------------------------------
# 2. Initialize model, loss function, and optimizer
# ----------------------------------------------------------
model = CNN()
criterion = nn.CrossEntropyLoss()           # Suitable for multi-class classification tasks
optimizer = optim.Adam(model.parameters(),  # Adam optimizer for adaptive learning rates
                       lr=1e-3)             # Learning rate of 0.001


# ----------------------------------------------------------
# 3. Define the training loop
# ----------------------------------------------------------
def train_model(model, train_loader, val_loader, criterion, optimizer, epochs=5):
    best_val_loss = float('inf')  # Used for saving the best model checkpoint

    for epoch in range(epochs):
        model.train()             # Set model to training mode (activates dropout, batchnorm updates)
        running_loss = 0.0        # Accumulator for tracking training loss per epoch

        # Iterate through all mini-batches
        for images, labels in train_loader:
            optimizer.zero_grad()             # Reset gradients to prevent accumulation
            outputs = model(images)           # Forward pass: compute model predictions
            loss = criterion(outputs, labels) # Compute training loss
            loss.backward()                   # Backpropagation: compute gradients
            torch.nn.utils.clip_grad_norm_(   # Clip gradients to avoid exploding gradients
                model.parameters(), 1.0)
            optimizer.step()                  # Update model weights
            running_loss += loss.item()       # Add current batch loss to total

        # Compute average training loss for the epoch
        avg_train_loss = running_loss / len(train_loader)

        # Evaluate model on the validation set after each epoch
        val_loss, val_acc = evaluate_vision_model(model, val_loader, criterion)

        # Print training and validation metrics
        print(f"Epoch {epoch+1}: "
              f"train_loss={avg_train_loss:.3f}, "
              f"val_loss={val_loss:.3f}, "
              f"val_acc={val_acc:.3f}")

        # Save model checkpoint if validation loss improves
        if val_loss < best_val_loss:
            best_val_loss = val_loss
            torch.save(model.state_dict(), "best_cnn.pt")
            print("✅ Saved best model checkpoint.")


# ----------------------------------------------------------
# 4. Define validation (evaluation) loop
# ----------------------------------------------------------
def evaluate_vision_model(model, loader, criterion):
    model.eval()  # Set model to evaluation mode (turns off dropout, batchnorm updates)
    loss_total, correct, total = 0.0, 0, 0

    # Disable gradient tracking during evaluation (saves memory and compute)
    with torch.no_grad():
        for images, labels in loader:
            outputs = model(images)                     # Forward pass
            loss_total += criterion(outputs, labels).item()  # Accumulate batch loss

            # Get predicted class indices by taking the argmax along class dimension
            preds = outputs.argmax(dim=1)
            correct += (preds == labels).sum().item()   # Count correct predictions
            total += labels.size(0)                     # Count total samples processed

    # Compute average validation loss and accuracy
    avg_loss = loss_total / len(loader)
    accuracy = correct / total
    return avg_loss, accuracy

LoRA (Low-Rank Adaptation) Fine-Tuning

Example: LoRA Fine-Tuning in PyTorch
import torch
import torch.nn as nn
import torch.optim as optim

# - LoRA layer wrapper -
class LoRALayer(nn.Module):
    """
    Implements a Low-Rank Adaptation (LoRA) layer that wraps an existing Linear layer.
    Instead of updating all model weights, this layer learns two low-rank matrices (A and B)
    that approximate the weight updates, reducing memory and computation cost.
    """
    def __init__(self, linear_layer, rank=4, alpha=1.0):
        super().__init__()
        self.linear = linear_layer  # Original (frozen) linear layer
        self.rank = rank            # Rank for low-rank decomposition
        self.alpha = alpha          # Scaling factor for update strength

        # Freeze base layer parameters — LoRA does NOT modify pretrained weights
        for param in self.linear.parameters():
            param.requires_grad = False

        # Extract dimensions from the wrapped linear layer
        in_features = self.linear.in_features   # e.g., 768
        out_features = self.linear.out_features # e.g., 3072

        # A: projects input down to rank dimension  $$\rightarrow$$ shape (rank, in_features)
        # B: projects it back up to output dimension $$\rightarrow$$ shape (out_features, rank)
        # Multiplying by 0.01 ensures small initial weights so that the LoRA update
        # starts near zero — this prevents disrupting the pretrained model’s outputs
        # at the beginning of training.
        # Use nn.Parameter so that A and B are registered as learnable parameters of the module.
        # This ensures they appear in model.parameters() and are updated by the optimizer
        # during backpropagation, unlike regular tensors which would not receive gradients.        
        self.A = nn.Parameter(torch.randn(rank, in_features) * 0.01)    # (r, in_features)
        self.B = nn.Parameter(torch.randn(out_features, rank) * 0.01)   # (out_features, r)
        self.scaling = self.alpha / self.rank  # Normalization factor to scale LoRA output

    def forward(self, x):
        # x: (batch_size, in_features)
        base_out = self.linear(x)  # (batch_size, out_features)

        # --- LoRA update computation ---
        # self.A: (rank, in_features)
        # self.B: (out_features, rank)

        # Step 1: (self.B @ self.A)
        # (out_features, rank) @ (rank, in_features) $$\rightarrow$$ (out_features, in_features)
        #
        # Step 2: ((self.B @ self.A) @ x.T)
        # (out_features, in_features) @ (in_features, batch_size) $$\rightarrow$$ (out_features, batch_size)
        #
        # Result: delta_out now has shape (out_features, batch_size)
        delta_out = (self.B @ self.A) @ x.T

        # Step 3: transpose back to (batch_size, out_features)
        delta_out = delta_out.T * self.scaling  # (batch_size, out_features)

        # Final output = frozen base layer output + scaled low-rank update
        return base_out + delta_out  # (batch_size, out_features)


# Example usage with a simple feed-forward model
class SimpleLoRAModel(nn.Module):
    """
    A simple feed-forward network with one LoRA-adapted linear layer.
    Demonstrates how LoRA can be integrated into a standard PyTorch model.
    """
    def __init__(self, input_dim=768, hidden_dim=128, num_classes=2):
        super().__init__()
        # First linear layer wrapped with LoRA for fine-tuning
        self.fc1 = nn.Linear(input_dim, hidden_dim)
        self.fc1 = LoRALayer(self.fc1, rank=8, alpha=2.0)

        # Non-linear activation
        self.relu = nn.ReLU()

        # Output layer — trains normally
        self.fc2 = nn.Linear(hidden_dim, num_classes)

    def forward(self, x):
        # Forward pass through LoRA-adapted and normal layers
        x = self.relu(self.fc1(x))
        return self.fc2(x)


# Simulated training process
# Only LoRA parameters (A and B) are trainable, base model weights are frozen
model = SimpleLoRAModel()

# Filter optimizer to update only trainable (LoRA) parameters
optimizer = optim.Adam(filter(lambda p: p.requires_grad, model.parameters()), lr=1e-3)
criterion = nn.CrossEntropyLoss()

# Dummy training loop for demonstration
for epoch in range(3):
    model.train()

    # Generate synthetic data (batch_size=16, input_dim=768)
    inputs = torch.randn(16, 768)
    labels = torch.randint(0, 2, (16,))

    # Zero gradients before each step
    optimizer.zero_grad()

    # Forward pass through the model
    outputs = model(inputs)

    # Compute classification loss
    loss = criterion(outputs, labels)

    # Backward pass — computes gradients only for LoRA matrices A and B
    loss.backward()

    # Update LoRA parameters
    optimizer.step()

    print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}")

Evaluation and Metrics

Classification Metrics
\[\text{Accuracy} = \frac{\text{Correct Predictions}}{\text{Total Samples}}\] \[\text{Precision} = \frac{TP}{TP + FP}\] \[\text{Recall} = \frac{TP}{TP + FN}\] \[F1 = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}\]

Early Stopping and Checkpointing

# Define early stopping parameters
patience = 3                      # Number of epochs to wait for improvement before stopping
best_loss = float('inf')          # Initialize best validation loss as infinity (no best yet)
patience_counter = 0              # Counts epochs with no improvement

for epoch in range(epochs):
    ...
    # After each epoch, evaluate the model on the validation set
    val_loss, _ = evaluate_model(model, val_loader, criterion)
    
    # Check if validation loss improved
    if val_loss < best_loss:
        best_loss = val_loss               # Update best recorded loss
        patience_counter = 0               # Reset counter since improvement occurred
        torch.save(model.state_dict(), "best_model.pt")  # Save best model checkpoint
        print(f"Validation improved. Saving model with val_loss={val_loss:.4f}")
    else:
        patience_counter += 1              # Increment counter (no improvement)
        print(f"No improvement for {patience_counter} epoch(s).")

        # If model hasn’t improved for 'patience' epochs $$\rightarrow$$ stop training
        if patience_counter >= patience:
            print("Early stopping triggered. Training halted.")
            break

Reproducibility Techniques

import torch, random, numpy as np

# 1. Set the random seed for PyTorch operations (CPU and GPU)
# Ensures all torch-level randomness (e.g., weight initialization, dropout) is reproducible.
torch.manual_seed(42)

# 2. Set the random seed for Python's built-in random module
# Controls functions like random.shuffle(), random.sample(), etc.
random.seed(42)

# 3. Set the random seed for NumPy
# Makes NumPy-generated random numbers (e.g., np.random.rand()) deterministic.
np.random.seed(42)

# 4. Make CuDNN deterministic
# Forces PyTorch to use deterministic algorithms for operations like convolutions.
# This avoids slight variations in results between runs.
torch.backends.cudnn.deterministic = True

# 5. Disable CuDNN benchmarking
# CuDNN usually selects the fastest algorithm for the hardware, which can introduce randomness.
# Setting this to False ensures consistency at the cost of minor speed reductions.
torch.backends.cudnn.benchmark = False

Logging and Monitoring

from torch.utils.tensorboard import SummaryWriter

# 1. Initialize TensorBoard writer
# Creates a log directory where TensorBoard will store metrics for visualization.
# Each run (experiment) can have its own directory for tracking progress.
writer = SummaryWriter(log_dir='runs/exp1')

# 2. Training loop
# For each training epoch, record metrics (e.g., training and validation loss).
for epoch in range(epochs):
    train_loss = ...  # Compute or retrieve the average training loss for this epoch
    val_loss = ...    # Compute or retrieve the average validation loss for this epoch

    # 3. Log both training and validation losses to TensorBoard
    # The 'Loss' tag groups related metrics together for easy comparison.
    # Each scalar value is associated with the current epoch number.
    writer.add_scalars('Loss', {'train': train_loss, 'val': val_loss}, epoch)

# 4. Close the writer
# Flushes and saves all pending events to disk to ensure they appear in TensorBoard.
writer.close()
tensorboard --logdir=runs

FAQs

Practical Implementation – Model Training/Fine-tuning and Evaluation Workflow

Example 1: Vision Use-Case (Image Classification on CIFAR-10)
Step 1: Setup and Data Loading
import torch
import torch.nn as nn
import torch.optim as optim
from torchvision import datasets, transforms
from torch.utils.data import DataLoader

# ----------------------------------------------------------
# 1. Define data transformations (normalization + augmentation)
# ----------------------------------------------------------
# For training: include random flips and crops for data augmentation
#   - RandomHorizontalFlip(): randomly flip images to improve generalization
#   - RandomCrop(): crop randomly to simulate viewpoint variation
#   - ToTensor(): convert PIL image $$\rightarrow$$ PyTorch tensor (scales pixels to [0,1])
#   - Normalize(): normalize with dataset-specific mean/std per channel
transform_train = transforms.Compose([
    transforms.RandomHorizontalFlip(),
    transforms.RandomCrop(32, padding=4),
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465),  # mean (R, G, B)
                         (0.2023, 0.1994, 0.2010))  # std (R, G, B)
])

# For testing/validation: use deterministic preprocessing (no augmentation)
#   - This ensures consistent evaluation conditions
transform_test = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465),
                         (0.2023, 0.1994, 0.2010))
])

# ----------------------------------------------------------
# 2. Load CIFAR-10 dataset
# ----------------------------------------------------------
#   - root="./data": location to store the dataset
#   - train=True: load training set
#   - transform=...: apply defined preprocessing pipeline
#   - download=True: automatically download if not present
train_dataset = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform_train)
test_dataset = datasets.CIFAR10(root="./data", train=False, download=True, transform=transform_test)

# ----------------------------------------------------------
# 3. Split validation set from training data
# ----------------------------------------------------------
#   - Reserve 5,000 samples for validation
#   - random_split ensures samples are split randomly while maintaining reproducibility if seed is set
train_set, val_set = torch.utils.data.random_split(train_dataset, [45000, 5000])

# ----------------------------------------------------------
# 4. Create DataLoaders for batching and shuffling
# ----------------------------------------------------------
# DataLoader wraps datasets for efficient batching, shuffling, and multiprocessing
#   - batch_size: number of samples per batch
#   - shuffle=True: randomize order each epoch (important for training)
#   - num_workers: number of subprocesses to load data in parallel
train_loader = DataLoader(train_set, batch_size=64, shuffle=True, num_workers=2)
val_loader = DataLoader(val_set, batch_size=64, shuffle=False)
test_loader = DataLoader(test_dataset, batch_size=64, shuffle=False)

# ----------------------------------------------------------
# 5. (Optional) Inspect batch shapes
# ----------------------------------------------------------
# You can quickly verify data shapes:
# images, labels = next(iter(train_loader))
# print(images.shape)  # e.g., torch.Size([64, 3, 32, 32])
# print(labels.shape)  # e.g., torch.Size([64])
Step 2: Define the Model
import torch
import torch.nn as nn

# ----------------------------------------------------------
# CNN Classifier for CIFAR-10 (3x32x32 input images)
# ----------------------------------------------------------
class CNNClassifier(nn.Module):
    def __init__(self):
        super().__init__()

        # -------------------------------
        # 1. Convolutional feature extractor
        # -------------------------------
        # This block extracts spatial features from images using convolution,
        # nonlinearity, and pooling operations.
        self.conv_block = nn.Sequential(
            # First convolution layer:
            #   - Input: 3 input channels (RGB)
            #   - Output: 32 filters $$\rightarrow$$ 32 output channels $$\rightarrow$$ 32 feature maps
            #   - Kernel: 3x3 convolution filters 
            #   - Padding: 1 to preserve spatial dimensions
            #   - Stride: controls how far the kernel moves each step (default = 1)
            nn.Conv2d(3, 32, 3, padding=1),
            nn.ReLU(),          # Non-linear activation
            nn.MaxPool2d(2),    # Downsample feature maps by a factor of 2 (32x32 $$\rightarrow$$ 16x16)

            # Second convolution layer:
            #   - Input: 32 feature maps
            #   - Output: 64 feature maps
            nn.Conv2d(32, 64, 3, padding=1),
            nn.ReLU(),
            nn.MaxPool2d(2)     # Downsample again (16x16 $$\rightarrow$$ 8x8)
        )

        # -------------------------------
        # 2. Fully connected classifier
        # -------------------------------
        # This block flattens the feature maps and predicts class logits.
        self.fc_block = nn.Sequential(
            nn.Flatten(),                       # Flatten from (batch, 64, 8, 8) $$\rightarrow$$ (batch, 64*8*8)
            nn.Linear(64 * 8 * 8, 128),         # Fully connected layer with 128 hidden units
            nn.ReLU(),                          # Non-linear activation
            nn.Dropout(0.3),                    # Dropout (30%) to reduce overfitting
            nn.Linear(128, 10)                  # Output layer for 10 CIFAR-10 classes
        )

    # -------------------------------
    # 3. Forward pass
    # -------------------------------
    # Defines the data flow: input $$\rightarrow$$ conv layers $$\rightarrow$$ fully connected layers $$\rightarrow$$ output
    def forward(self, x):
        # Pass input through convolutional block, then classification block
        return self.fc_block(self.conv_block(x))
Step 3: Training and Evaluation Loops
def train_vision_model(model, train_loader, val_loader, criterion, optimizer, epochs=5):
    # ----------------------------------------------------------
    # 1. Setup device and model
    # ----------------------------------------------------------
    # Select GPU if available, otherwise fall back to CPU
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.to(device)  # Move model parameters to the chosen device

    # Initialize the best validation loss to a large number (for checkpointing)
    best_val_loss = float('inf')

    # ----------------------------------------------------------
    # 2. Main training loop (iterate over epochs)
    # ----------------------------------------------------------
    for epoch in range(epochs):
        model.train()  # Set model to training mode (enables dropout, batchnorm updates)
        running_loss = 0.0  # Accumulates total training loss per epoch

        # ----------------------------------------------------------
        # 3. Iterate through mini-batches in the training set
        # ----------------------------------------------------------
        for images, labels in train_loader:
            # Move data and labels to the same device as the model
            images, labels = images.to(device), labels.to(device)

            optimizer.zero_grad()        # Clear previous gradients
            outputs = model(images)      # Forward pass through the model
            loss = criterion(outputs, labels)  # Compute loss (e.g., cross-entropy)
            loss.backward()              # Backpropagate gradients
            torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)  # Prevent exploding gradients
            optimizer.step()             # Update model weights
            running_loss += loss.item()  # Track cumulative batch loss

        # ----------------------------------------------------------
        # 4. Compute average training loss for the epoch
        # ----------------------------------------------------------
        avg_train_loss = running_loss / len(train_loader)

        # ----------------------------------------------------------
        # 5. Evaluate model on the validation set
        # ----------------------------------------------------------
        val_loss, val_acc = evaluate_vision_model(model, val_loader, criterion, device)

        # Display training progress
        print(f"Epoch {epoch+1}: train_loss={avg_train_loss:.3f}, val_loss={val_loss:.3f}, val_acc={val_acc:.3f}")

        # ----------------------------------------------------------
        # 6. Checkpoint the best model (based on validation loss)
        # ----------------------------------------------------------
        if val_loss < best_val_loss:
            best_val_loss = val_loss
            torch.save(model.state_dict(), "best_cifar10_model.pt")  # Save model weights
            print("✅ Best model updated and saved.")


# ----------------------------------------------------------
# Validation / Evaluation Function
# ----------------------------------------------------------
def evaluate_vision_model(model, loader, criterion, device):
    model.eval()  # Set model to evaluation mode (disables dropout, batchnorm updates)
    total_loss, correct, total = 0.0, 0, 0  # Initialize counters

    # Disable gradient computation for inference (saves memory and time)
    with torch.no_grad():
        # Iterate through the validation or test data
        for images, labels in loader:
            images, labels = images.to(device), labels.to(device)
            outputs = model(images)                   # Forward pass
            loss = criterion(outputs, labels)         # Compute batch loss
            total_loss += loss.item()                 # Accumulate total loss
            preds = outputs.argmax(dim=1)             # [batch_size, num_classes] -> Convert logits to class predictions
                                                      # Return index of the max logit across classes (dim=1) for each sample
            correct += (preds == labels).sum().item() # Count correct predictions
            total += labels.size(0)                   # Count total samples

    # ----------------------------------------------------------
    # 7. Return average loss and accuracy for the validation set
    # ----------------------------------------------------------
    avg_loss = total_loss / len(loader)
    accuracy = correct / total
    return avg_loss, accuracy
Step 4: Run Training
# ----------------------------------------------------------
# 1. Initialize the model
# ----------------------------------------------------------
# Create an instance of the CNNClassifier defined earlier.
# This model will be trained on the CIFAR-10 dataset.
model = CNNClassifier()

# ----------------------------------------------------------
# 2. Define the loss function
# ----------------------------------------------------------
# CrossEntropyLoss is commonly used for multi-class classification problems.
# It combines LogSoftmax + Negative Log Likelihood into a single step.
criterion = nn.CrossEntropyLoss()

# ----------------------------------------------------------
# 3. Define the optimizer
# ----------------------------------------------------------
# Adam optimizer is chosen for its adaptive learning rate and momentum properties.
# It adjusts individual learning rates for each parameter, making convergence faster.
optimizer = optim.Adam(model.parameters(), lr=1e-3)

# ----------------------------------------------------------
# 4. Train the model
# ----------------------------------------------------------
# Call the training loop function defined earlier.
# Arguments:
#   - model: the CNN model to be trained
#   - train_loader: batches of training data
#   - val_loader: batches of validation data for monitoring performance
#   - criterion: the loss function used to compute training error
#   - optimizer: updates the model weights based on computed gradients
#   - epochs: number of full passes through the training dataset
train_vision_model(model, train_loader, val_loader, criterion, optimizer, epochs=10)
Example 2: NLP Use-Case (Sentiment Classification on IMDb Dataset)
Step 1: Data Preparation
from torchtext.datasets import IMDB
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator
from torch.utils.data import DataLoader
from torch.nn.utils.rnn import pad_sequence
import torch

# ----------------------------------------------------------
# 1. Tokenizer
# ----------------------------------------------------------
# The IMDB dataset consists of raw text reviews and sentiment labels ("pos"/"neg").
# We define a simple tokenizer to split text into lowercase tokens using basic English rules.
tokenizer = get_tokenizer("basic_english")

# ----------------------------------------------------------
# 2. Vocabulary Building
# ----------------------------------------------------------
# The vocabulary maps unique tokens $$\rightarrow$$ integer IDs.
# This enables converting tokenized words into numeric tensors for embedding lookup.
def yield_tokens(data_iter):
    """Generator function that yields token lists from each text sample."""
    for label, line in data_iter:
        yield tokenizer(line)

# Load the training split of IMDB dataset (only used here for vocabulary construction)
train_iter = IMDB(split='train')

# Build vocabulary from training text tokens
# - specials: add special tokens for unknown (<unk>) and padding (<pad>)
vocab = build_vocab_from_iterator(yield_tokens(train_iter), specials=["<unk>", "<pad>"])

# Set default index for out-of-vocabulary words
vocab.set_default_index(vocab["<unk>"])
pad_idx = vocab["<pad>"]  # Store padding index for later use

# ----------------------------------------------------------
# 3. Collate Function for Batching
# ----------------------------------------------------------
# The collate function defines how individual dataset items are combined into a batch.
# It handles tokenization, numericalization, and dynamic padding.
def collate_batch(batch):
    labels, texts = [], []
    for label, text in batch:
        # Convert labels: "pos" $$\rightarrow$$ 1, "neg" $$\rightarrow$$ 0
        labels.append(1 if label == "pos" else 0)

        # Tokenize and numericalize text
        tokens = vocab(tokenizer(text))
        texts.append(torch.tensor(tokens, dtype=torch.long))

    # Pad sequences in the batch to the same length for tensor batching
    padded_texts = pad_sequence(texts, batch_first=True, padding_value=pad_idx)
    label_tensor = torch.tensor(labels)
    return padded_texts, label_tensor

# ----------------------------------------------------------
# 4. Load IMDB Data Splits
# ----------------------------------------------------------
# Reload IMDB data for training/validation/testing after building the vocab.
# Each sample is a tuple: (label, text)
train_iter, test_iter = IMDB()

# Convert iterators to lists (so they can be indexed and split)
# NOTE: For demonstration, we use a subset (first 5,000 samples)
train_list = list(train_iter)[:4000]
val_list = list(train_iter)[4000:5000]

# ----------------------------------------------------------
# 5. Create DataLoaders
# ----------------------------------------------------------
# DataLoader wraps dataset lists and applies batching + collate function.
# - collate_fn: applies our tokenization, numericalization, and padding logic
# - shuffle=True: randomizes sample order each epoch
train_loader = DataLoader(train_list, batch_size=32, collate_fn=collate_batch, shuffle=True)
val_loader = DataLoader(val_list, batch_size=32, collate_fn=collate_batch)

# ----------------------------------------------------------
# 6. (Optional) Inspect one batch
# ----------------------------------------------------------
# for x_batch, y_batch in train_loader:
#     print("Input batch shape:", x_batch.shape)  # (batch_size, seq_len)
#     print("Label batch shape:", y_batch.shape)  # (batch_size,)
#     break
Step 2: Define the LSTM Model
class SentimentRNN(nn.Module):
    def __init__(self, vocab_size, embed_dim, hidden_dim, output_dim, pad_idx):
        super().__init__()
        # ----------------------------------------------------------
        # 1. Embedding layer
        # ----------------------------------------------------------
        # Converts token indices into dense vector representations.
        # Each word index maps to an 'embed_dim'-dimensional vector.
        # 'padding_idx' ensures that <pad> tokens are ignored during training.
        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)

        # ----------------------------------------------------------
        # 2. LSTM layer
        # ----------------------------------------------------------
        # Processes the embedded sequence to capture temporal dependencies.
        # 'hidden_dim' determines the dimensionality of the hidden state.
        # 'batch_first=True' ensures input/output tensors are of shape (batch, seq, feature).
        self.lstm = nn.LSTM(embed_dim, hidden_dim, batch_first=True)

        # ----------------------------------------------------------
        # 3. Fully connected (output) layer
        # ----------------------------------------------------------
        # Maps the final hidden state from LSTM to class logits.
        # For binary classification, output_dim=2; for sentiment (pos/neg).
        self.fc = nn.Linear(hidden_dim, output_dim)

        # ----------------------------------------------------------
        # 4. Dropout layer
        # ----------------------------------------------------------
        # Randomly zeros out some activations to prevent overfitting.
        self.dropout = nn.Dropout(0.3)

    def forward(self, x):
        # ----------------------------------------------------------
        # Forward pass
        # ----------------------------------------------------------
        # x: input tensor of token indices with shape (batch_size, seq_len)

        # Step 1: Convert tokens to embeddings
        embedded = self.embedding(x)  # shape $$\rightarrow$$ (batch_size, seq_len, embed_dim)

        # Step 2: Pass through LSTM
        # 'hidden' captures the final hidden state from the last time step
        _, (hidden, _) = self.lstm(embedded)

        # Step 3: Apply dropout to the final hidden state
        dropped = self.dropout(hidden.squeeze(0))  # remove extra LSTM dimension

        # Step 4: Compute class logits
        output = self.fc(dropped)  # shape $$\rightarrow$$ (batch_size, output_dim)

        return output
Step 3: Training and Validation Loop
def train_text_model(model, train_loader, val_loader, criterion, optimizer, epochs=5):
    # ----------------------------------------------------------
    # 1. Device setup
    # ----------------------------------------------------------
    # Automatically select GPU if available; otherwise use CPU
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.to(device)  # Move model to selected device (GPU/CPU)

    best_val_loss = float('inf')  # Initialize best validation loss for checkpointing

    # ----------------------------------------------------------
    # 2. Training loop
    # ----------------------------------------------------------
    for epoch in range(epochs):
        model.train()  # Set model to training mode (enables dropout, etc.)
        total_loss = 0  # Accumulate training loss per epoch

        # Iterate through all mini-batches
        for x_batch, y_batch in train_loader:
            # Move input and labels to device (GPU/CPU)
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)

            optimizer.zero_grad()          # Reset gradients before backward pass
            outputs = model(x_batch)       # Forward pass: compute predictions
            loss = criterion(outputs, y_batch)  # Compute loss (e.g., CrossEntropy)
            loss.backward()                # Backward pass: compute gradients
            optimizer.step()               # Update model parameters
            total_loss += loss.item()      # Track cumulative batch loss

        # Compute average training loss for the epoch
        avg_train_loss = total_loss / len(train_loader)

        # ----------------------------------------------------------
        # 3. Validation phase
        # ----------------------------------------------------------
        # Evaluate model on validation set without gradient computation
        val_loss, val_acc = evaluate_text_model(model, val_loader, criterion, device)

        # Print training progress
        print(f"Epoch {epoch+1}: train_loss={avg_train_loss:.3f}, val_loss={val_loss:.3f}, val_acc={val_acc:.3f}")

        # ----------------------------------------------------------
        # 4. Checkpointing
        # ----------------------------------------------------------
        # Save the model if validation loss improves
        if val_loss < best_val_loss:
            best_val_loss = val_loss
            torch.save(model.state_dict(), "best_imdb_model.pt")
            print("✅ Saved new best model checkpoint.")


# ----------------------------------------------------------
# Evaluation function
# ----------------------------------------------------------
def evaluate_text_model(model, loader, criterion, device):
    model.eval()  # Set model to evaluation mode (disables dropout, etc.)
    total_loss, correct, total = 0.0, 0, 0

    # Disable gradient tracking for faster evaluation and lower memory usage
    with torch.no_grad():
        for x_batch, y_batch in loader:
            # Move inputs and labels to correct device
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)

            # Forward pass only
            outputs = model(x_batch)
            loss = criterion(outputs, y_batch)  # Compute loss
            total_loss += loss.item()           # Accumulate total loss

            # Compute predictions and accuracy
            preds = outputs.argmax(dim=1)       # Get class with highest probability
            correct += (preds == y_batch).sum().item()
            total += y_batch.size(0)

    # Return mean validation loss and overall accuracy
    return total_loss / len(loader), correct / total
Step 4: Run Training
# ----------------------------------------------------------
# 1. Define model input dimensions
# ----------------------------------------------------------
# vocab_size: total number of unique tokens in the vocabulary
#   - determines the size of the embedding layer
vocab_size = len(vocab)

# ----------------------------------------------------------
# 2. Initialize the SentimentRNN model
# ----------------------------------------------------------
# Model parameters:
#   - embed_dim: dimensionality of word embeddings (dense representations)
#   - hidden_dim: number of hidden units in the RNN
#   - output_dim: number of target classes (e.g., 2 for positive/negative sentiment)
#   - pad_idx: index of the <pad> token, ensures padding doesn’t affect learning
model = SentimentRNN(
    vocab_size=vocab_size,
    embed_dim=64,
    hidden_dim=128,
    output_dim=2,
    pad_idx=pad_idx
)

# ----------------------------------------------------------
# 3. Define loss function
# ----------------------------------------------------------
# CrossEntropyLoss combines LogSoftmax and NLLLoss
#   - suitable for multi-class classification problems
#   - expects raw logits as model outputs
criterion = nn.CrossEntropyLoss()

# ----------------------------------------------------------
# 4. Define optimizer
# ----------------------------------------------------------
# Adam optimizer adapts learning rates per parameter
#   - lr=1e-3 is a common starting learning rate for NLP models
optimizer = optim.Adam(model.parameters(), lr=1e-3)

# ----------------------------------------------------------
# 5. Train the model
# ----------------------------------------------------------
# The training function handles:
#   - forward and backward passes
#   - gradient updates
#   - validation loss/accuracy computation
#   - checkpoint saving for the best model
train_text_model(model, train_loader, val_loader, criterion, optimizer, epochs=5)
End Result

Model Experimentation and Hyperparameter Tuning

Overview

Core Principles of Experimentation

Isolate One Variable at a Time
Track Every Experiment
Automate Where Possible

Example: Manual Hyperparameter Tuning Loop

import torch
from torch.utils.data import DataLoader
import itertools  # Used for creating combinations of hyperparameters

# ----------------------------------------------------------
# 1. Define hyperparameter search space
# ----------------------------------------------------------
# We will run experiments for every combination of learning rate and batch size.
learning_rates = [1e-2, 1e-3, 1e-4]  # Candidate learning rates
batch_sizes = [32, 64, 128]          # Candidate batch sizes

# To store results of each experiment (lr, batch_size, val_loss, val_acc)
results = []

# ----------------------------------------------------------
# 2. Loop through all hyperparameter combinations
# ----------------------------------------------------------
# itertools.product() generates all possible (lr, bs) pairs.
for lr, bs in itertools.product(learning_rates, batch_sizes):
    print(f"Running experiment: lr={lr}, batch_size={bs}")

    # ------------------------------------------------------
    # 3. Reinitialize model, optimizer, and dataloaders
    # ------------------------------------------------------
    # Ensure each experiment starts with a fresh model and optimizer.
    # This avoids parameter carryover from previous runs.
    model = CNN()  # Recreate model instance
    optimizer = torch.optim.Adam(model.parameters(), lr=lr)

    # Create DataLoaders with the current batch size
    # shuffle=True ensures randomization of samples per epoch
    train_loader = DataLoader(train_data, batch_size=bs, shuffle=True)
    val_loader = DataLoader(test_data, batch_size=bs)

    # ------------------------------------------------------
    # 4. Train model and evaluate performance
    # ------------------------------------------------------
    # Train for a small number of epochs (e.g., 3) for rapid prototyping.
    train_vision_model(model, train_loader, val_loader, criterion, optimizer, epochs=3)

    # Evaluate on validation/test data
    val_loss, val_acc = evaluate_vision_model(model, val_loader, criterion)

    # Store experiment results
    results.append((lr, bs, val_loss, val_acc))

# ----------------------------------------------------------
# 5. Summarize results
# ----------------------------------------------------------
# After all runs complete, print or log experiment outcomes.
print("All experiments complete.")
# You can also sort or analyze results later:
# best_run = min(results, key=lambda x: x[2])  # e.g., best by lowest val_loss

Example: Automated Tuning with Optuna

import optuna

# ----------------------------------------------------------
# 1. Define the objective function for optimization
# ----------------------------------------------------------
# The objective function trains and evaluates a model for each trial.
# Optuna will call this function multiple times with different hyperparameters.
def objective(trial):
    # ------------------------------------------------------
    # 1a. Suggest hyperparameters to tune
    # ------------------------------------------------------
    # 'suggest_loguniform' samples the learning rate from a log-uniform distribution
    #   $$\rightarrow$$ allows exploration of several orders of magnitude (1e-5 to 1e-2).
    lr = trial.suggest_loguniform("lr", 1e-5, 1e-2)

    # 'suggest_int' samples an integer value (here, hidden dimension size)
    #   $$\rightarrow$$ tested in increments of 64 (from 64 to 512).
    hidden_dim = trial.suggest_int("hidden_dim", 64, 512, step=64)

    # ------------------------------------------------------
    # 1b. Initialize model, optimizer, and loss function
    # ------------------------------------------------------
    # Create an instance of the model using the sampled hyperparameters.
    model = RNNClassifier(
        vocab_size, 
        embed_dim=64, 
        hidden_dim=hidden_dim, 
        num_classes=2, 
        pad_idx=vocab["<pad>"]
    )

    # Define optimizer with sampled learning rate
    optimizer = torch.optim.Adam(model.parameters(), lr=lr)

    # Define loss function for classification
    criterion = nn.CrossEntropyLoss()

    # ------------------------------------------------------
    # 1c. Train and evaluate model
    # ------------------------------------------------------
    # Train for a few epochs (short training to keep tuning fast)
    train_nlp_model(model, loader, criterion, optimizer, epochs=3)

    # Evaluate model performance on validation data
    val_loss, val_acc = evaluate_nlp_model(model, loader, criterion)

    # ------------------------------------------------------
    # 1d. Return metric to minimize (validation loss)
    # ------------------------------------------------------
    # Optuna will minimize this value across trials to find best hyperparameters.
    return val_loss


# ----------------------------------------------------------
# 2. Create a study and specify the optimization direction
# ----------------------------------------------------------
# direction="minimize" $$\rightarrow$$ Optuna tries to find the lowest validation loss.
study = optuna.create_study(direction="minimize")

# ----------------------------------------------------------
# 3. Run the optimization
# ----------------------------------------------------------
# study.optimize() runs the objective function multiple times (n_trials).
# Each trial uses a new set of hyperparameters suggested by Optuna's sampler.
study.optimize(objective, n_trials=10)

# ----------------------------------------------------------
# 4. Inspect best results
# ----------------------------------------------------------
# 'best_params' gives the hyperparameter set that achieved the lowest validation loss.
print(study.best_params)

Configuration Management with Hydra

conf/
  config.yaml
  model.yaml
  data.yaml
train.py
defaults:
  - model: rnn
  - data: text_dataset

trainer:
  epochs: 5
  lr: 1e-3
  batch_size: 64
import hydra
from omegaconf import DictConfig

# ----------------------------------------------------------
# 1. Hydra main decorator
# ----------------------------------------------------------
# @hydra.main() is the entry point for a Hydra-powered application.
# It automatically:
#   - Loads configuration files from the specified path.
#   - Composes them into a single hierarchical config object.
#   - Passes this config (as DictConfig) into your main function.
#
# Parameters:
#   config_path="conf"   $$\rightarrow$$ folder containing configuration YAML files
#   config_name="config" $$\rightarrow$$ main config file (e.g., conf/config.yaml)
@hydra.main(config_path="conf", config_name="config")
def train(cfg: DictConfig):
    # ----------------------------------------------------------
    # 2. Access configuration parameters
    # ----------------------------------------------------------
    # The config (cfg) behaves like a nested dictionary.
    # For example, given conf/config.yaml:
    # trainer:
    #   lr: 0.001
    #   epochs: 10
    #   batch_size: 32
    #
    # The following line prints:
    #   0.001 10 32
    print(cfg.trainer.lr, cfg.trainer.epochs, cfg.trainer.batch_size)

    # ----------------------------------------------------------
    # 3. Placeholder for model setup and training logic
    # ----------------------------------------------------------
    # This is where you'd typically:
    #   - Initialize your model (e.g., CNN, Transformer)
    #   - Set up optimizer and loss function
    #   - Implement your training and validation loops
    #   - Possibly log results to TensorBoard or W&B
    #
    # Example:
    # model = MyModel(cfg.model)
    # optimizer = torch.optim.Adam(model.parameters(), lr=cfg.trainer.lr)
    # train_vision_model(model, optimizer, cfg.trainer.epochs)
    # ----------------------------------------------------------
    # Currently, it only prints configuration values as a demo.
    pass

# ----------------------------------------------------------
# 4. Hydra entry point
# ----------------------------------------------------------
# The standard Python entry point ensures this script can be run directly:
#   python train.py
#
# Hydra automatically changes the working directory for each run
# (e.g., outputs/2025-10-19/10-30-12) to keep results organized.
if __name__ == "__main__":
    train()

Experiment Tracking with Weights & Biases

import wandb

# ----------------------------------------------------------
# 1. Initialize Weights & Biases (W&B) run
# ----------------------------------------------------------
#   - project: the W&B project name where runs will be grouped
#   - config: dictionary storing hyperparameters and metadata
#   - wandb.init() starts a new run and tracks all subsequent logs
wandb.init(project="pytorch-experiments", config={
    "learning_rate": 1e-3,
    "epochs": 5,
    "batch_size": 64
})

# Retrieve parameters from config for clarity (optional)
config = wandb.config
epochs = config.epochs

# ----------------------------------------------------------
# 2. Training and validation loop
# ----------------------------------------------------------
#   - Logs metrics (train/validation loss, accuracy) to the W&B dashboard
#   - Allows real-time monitoring and comparison across runs
for epoch in range(epochs):
    train_loss = ...  # (Placeholder) Compute average training loss for this epoch
    val_loss, val_acc = evaluate_model(model, val_loader, criterion)  # Evaluate on validation data

    # Log key metrics for the current epoch
    #   - Each call to wandb.log() records a single set of metrics
    #   - Automatically associates them with the current run
    wandb.log({
        "epoch": epoch,
        "train_loss": train_loss,
        "val_loss": val_loss,
        "val_acc": val_acc
    })

# ----------------------------------------------------------
# 3. Finalize the W&B run
# ----------------------------------------------------------
#   - Ensures all logs, metrics, and artifacts are properly synced
#   - Closes the active tracking session
wandb.finish()

Visualization of Hyperparameter Results

import pandas as pd
import matplotlib.pyplot as plt

# ----------------------------------------------------------
# 1. Create a DataFrame from experimental results
# ----------------------------------------------------------
# 'results' is assumed to be a list of tuples or lists containing:
#   (learning_rate, batch_size, validation_loss, validation_accuracy)
# Example: results = [(0.001, 32, 0.45, 0.86), (0.001, 64, 0.42, 0.88), ...]
# The DataFrame makes it easier to analyze and visualize model performance.
df = pd.DataFrame(results, columns=["lr", "batch_size", "val_loss", "val_acc"])

# ----------------------------------------------------------
# 2. Initialize a new matplotlib figure
# ----------------------------------------------------------
# This creates a blank plotting canvas for the visualization.
plt.figure()

# ----------------------------------------------------------
# 3. Plot validation accuracy vs. learning rate for each batch size
# ----------------------------------------------------------
# Iterate over each unique batch size value in the DataFrame.
for bs in df["batch_size"].unique():
    # Filter rows corresponding to the current batch size
    subset = df[df["batch_size"] == bs]
    
    # Plot validation accuracy (y-axis) vs learning rate (x-axis)
    # Each line corresponds to a specific batch size
    plt.plot(subset["lr"], subset["val_acc"], label=f"Batch={bs}")

# ----------------------------------------------------------
# 4. Format and label the plot
# ----------------------------------------------------------
# Use logarithmic scale for learning rate — helps visualize small values clearly
plt.xscale('log')

# Label axes for clarity
plt.xlabel("Learning Rate")
plt.ylabel("Validation Accuracy")

# Add legend to distinguish lines by batch size
plt.legend()

# ----------------------------------------------------------
# 5. Display the plot
# ----------------------------------------------------------
# Renders the figure showing how validation accuracy changes
# across learning rates and batch sizes.
plt.show()

Reproducibility and Randomness Control

FAQs

Practical Implementation – Model Experimentation and Hyperparameter Tuning

Example 1: Vision Use-Case (CIFAR-10 CNN Hyperparameter Tuning)
Step 1: Setup Experiment Function
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
import itertools

# ----------------------------------------------------------
# 1. Define preprocessing transformations
# ----------------------------------------------------------
# The CIFAR-10 dataset contains RGB images of size 32x32.
# Transforms are used to convert PIL images to tensors and normalize pixel values.
#   - ToTensor(): converts the image to a PyTorch tensor and scales pixel values from [0, 255] $$\rightarrow$$ [0.0, 1.0].
#   - Normalize(): standardizes each channel (R, G, B) using dataset-specific mean and std.
#     This helps the model converge faster and more stably during training.
transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465),   # mean for CIFAR-10 (R, G, B)
                         (0.2023, 0.1994, 0.2010))   # std for CIFAR-10 (R, G, B)
])

# ----------------------------------------------------------
# 2. Load CIFAR-10 dataset
# ----------------------------------------------------------
# torchvision.datasets provides built-in access to CIFAR-10.
#   - root="data": directory to store or load the dataset.
#   - train=True: loads the training split (50,000 images).
#   - train=False: loads the test split (10,000 images).
#   - download=True: automatically downloads if not already present.
#   - transform=transform: applies the preprocessing pipeline defined above.
train_data = datasets.CIFAR10(root="data", train=True, download=True, transform=transform)
test_data = datasets.CIFAR10(root="data", train=False, download=True, transform=transform)

# ----------------------------------------------------------
# 3. Split training data into training and validation sets
# ----------------------------------------------------------
# It’s common to reserve a subset of the training data for validation.
# This allows monitoring model performance on unseen data during training.
# random_split() randomly partitions the dataset into:
#   - train_set: 45,000 samples used for learning
#   - val_set:   5,000 samples used for validation
train_set, val_set = torch.utils.data.random_split(train_data, [45000, 5000])

# ----------------------------------------------------------
# (Optional) Inspect sample data
# ----------------------------------------------------------
# You can visualize or check one sample to verify shape and normalization:
# image, label = train_set[0]
# print(image.shape)   # Expected: torch.Size([3, 32, 32])
# print(label)         # Integer class label (0–9)
Step 2: Experiment Function
def run_experiment(lr, batch_size, dropout_rate):
    """
    Run a single training experiment with specified hyperparameters:
      - lr: learning rate
      - batch_size: number of samples per batch
      - dropout_rate: dropout probability used in the model
    """

    # ----------------------------------------------------------
    # 1. Initialize model and set dropout rate dynamically
    # ----------------------------------------------------------
    model = CNNClassifier()  # Instantiate the CNN model (assumed predefined)
    for module in model.modules():
        # Find all dropout layers and update their probability (p)
        if isinstance(module, nn.Dropout):
            module.p = dropout_rate

    # ----------------------------------------------------------
    # 2. Define data loaders and optimizer
    # ----------------------------------------------------------
    # Create DataLoader objects for training and validation sets
    #   - batch_size is variable, allowing tuning
    #   - shuffle=True randomizes the order of samples each epoch
    train_loader = DataLoader(train_set, batch_size=batch_size, shuffle=True, num_workers=2)
    val_loader = DataLoader(val_set, batch_size=batch_size)

    # Define optimizer and loss function
    #   - Adam: adaptive learning rate optimizer
    #   - lr: variable learning rate for experimentation
    optimizer = optim.Adam(model.parameters(), lr=lr)
    criterion = nn.CrossEntropyLoss()  # Suitable for classification tasks

    # ----------------------------------------------------------
    # 3. Setup device and training configuration
    # ----------------------------------------------------------
    # Automatically use GPU if available, otherwise fallback to CPU
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.to(device)  # Move model parameters to chosen device
    best_val_acc = 0.0  # Track best validation accuracy during training

    # ----------------------------------------------------------
    # 4. Training loop (short runs for hyperparameter tuning)
    # ----------------------------------------------------------
    for epoch in range(3):  # Fewer epochs for quick experiments
        model.train()  # Enable training mode (dropout + batchnorm active)

        for images, labels in train_loader:
            # Move data to GPU (if available)
            images, labels = images.to(device), labels.to(device)

            optimizer.zero_grad()          # Reset gradients from previous step
            outputs = model(images)        # Forward pass through model
            loss = criterion(outputs, labels)  # Compute loss
            loss.backward()                # Backpropagate to compute gradients
            optimizer.step()               # Update model parameters

        # ----------------------------------------------------------
        # 5. Validation phase
        # ----------------------------------------------------------
        # Evaluate model on validation set after each epoch
        val_loss, val_acc = evaluate_vision_model(model, val_loader, criterion, device)

        # Track best validation accuracy achieved so far
        best_val_acc = max(best_val_acc, val_acc)

    # ----------------------------------------------------------
    # 6. Return best performance metric
    # ----------------------------------------------------------
    return best_val_acc  # Useful for hyperparameter tuning results
import itertools  # Used for generating all possible parameter combinations

# ----------------------------------------------------------
# 1. Define hyperparameter search space
# ----------------------------------------------------------
# These lists specify the values to try for each hyperparameter.
#   - learning_rates: controls step size for optimizer updates
#   - batch_sizes: number of samples per training step
#   - dropout_rates: regularization strength to prevent overfitting
learning_rates = [1e-2, 1e-3, 1e-4]
batch_sizes = [32, 64]
dropout_rates = [0.2, 0.4, 0.6]

# ----------------------------------------------------------
# 2. Run grid search across all parameter combinations
# ----------------------------------------------------------
# itertools.product() generates the Cartesian product of all parameter lists.
# Example: (lr, bs, dr) = (0.01, 32, 0.2), (0.01, 32, 0.4), ...
results = []
for lr, bs, dr in itertools.product(learning_rates, batch_sizes, dropout_rates):
    print(f"Running: lr={lr}, batch_size={bs}, dropout={dr}")

    # Run a single experiment using the current hyperparameters.
    # The run_experiment() function is assumed to:
    #   1. Build a model with dropout=dr
    #   2. Train using learning_rate=lr and batch_size=bs
    #   3. Return validation/test accuracy
    acc = run_experiment(lr, bs, dr)

    # Store (learning_rate, batch_size, dropout_rate, accuracy)
    results.append((lr, bs, dr, acc))

# ----------------------------------------------------------
# 3. Sort and display results
# ----------------------------------------------------------
# Sort the results list in descending order of accuracy (best model first)
results.sort(key=lambda x: x[3], reverse=True)

# Print all configurations and their corresponding accuracies
for r in results:
    print(f"lr={r[0]}, bs={r[1]}, dr={r[2]} -> acc={r[3]:.3f}")
Step 4: Visualizing Results
import pandas as pd
import matplotlib.pyplot as plt

# ----------------------------------------------------------
# 1. Create a DataFrame to organize hyperparameter tuning results
# ----------------------------------------------------------
# 'results' is expected to be a list of tuples or lists,
# where each entry corresponds to (learning_rate, batch_size, dropout, val_acc)
# Example:
# results = [
#     [0.001, 32, 0.3, 0.82],
#     [0.001, 64, 0.3, 0.85],
#     [0.01,  32, 0.3, 0.78],
#     ...
# ]
df = pd.DataFrame(results, columns=["lr", "batch_size", "dropout", "val_acc"])

# ----------------------------------------------------------
# 2. Initialize the plot
# ----------------------------------------------------------
# Create a new figure with defined size for better readability
plt.figure(figsize=(7, 5))

# ----------------------------------------------------------
# 3. Plot validation accuracy vs. learning rate for each batch size
# ----------------------------------------------------------
# Loop through each unique batch size in the DataFrame
for bs in df["batch_size"].unique():
    # Filter rows corresponding to the current batch size
    subset = df[df["batch_size"] == bs]
    # Plot learning rate (x-axis) vs validation accuracy (y-axis)
    # 'marker="o"' adds circle markers at each data point
    plt.plot(subset["lr"], subset["val_acc"], marker="o", label=f"Batch={bs}")

# ----------------------------------------------------------
# 4. Customize axes and scale
# ----------------------------------------------------------
# Use a logarithmic scale on the x-axis since learning rates typically vary exponentially
plt.xscale("log")

# Label the axes for clarity
plt.xlabel("Learning Rate (log scale)")
plt.ylabel("Validation Accuracy")

# ----------------------------------------------------------
# 5. Add legend and title
# ----------------------------------------------------------
# Legend helps distinguish different batch size curves
plt.legend()
plt.title("CIFAR-10 Hyperparameter Tuning")

# ----------------------------------------------------------
# 6. Display the plot
# ----------------------------------------------------------
plt.show()
Example 2: NLP Use-Case (IMDb Sentiment Classification Tuning)
Step 1: Setup Study Environment
import optuna
import torch
import torch.nn as nn
import torch.optim as optim

# Assume SentimentRNN model class, vocabulary (vocab, pad_idx),
# and data loaders (train_loader, val_loader) are already defined elsewhere.
# Optuna will use these for hyperparameter tuning.
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")  # Use GPU if available

# ----------------------------------------------------------
# Objective function that Optuna will repeatedly call
# Each call corresponds to a single hyperparameter trial.
# ----------------------------------------------------------
def objective(trial):
    # ------------------------------------------------------
    # 1. Define hyperparameter search space
    # ------------------------------------------------------
    # Optuna samples hyperparameters automatically from given ranges/distributions.
    embed_dim = trial.suggest_categorical("embed_dim", [32, 64, 128])   # Embedding dimension
    hidden_dim = trial.suggest_int("hidden_dim", 64, 256, step=64)      # Hidden size for RNN
    lr = trial.suggest_loguniform("lr", 1e-4, 1e-2)                     # Learning rate (log-scale sampling)

    # ------------------------------------------------------
    # 2. Initialize model, loss function, and optimizer
    # ------------------------------------------------------
    # Create a new model instance for this trial with sampled parameters
    model = SentimentRNN(
        vocab_size=len(vocab),
        embed_dim=embed_dim,
        hidden_dim=hidden_dim,
        output_dim=2,
        pad_idx=pad_idx
    )
    model.to(device)  # Move model to GPU or CPU

    optimizer = optim.Adam(model.parameters(), lr=lr)  # Optimizer for parameter updates
    criterion = nn.CrossEntropyLoss()                  # Loss function for classification

    # ------------------------------------------------------
    # 3. Training loop (shortened for faster experimentation)
    # ------------------------------------------------------
    # We only train for a few epochs to quickly assess performance.
    model.train()
    for epoch in range(2):
        total_loss = 0.0
        for x_batch, y_batch in train_loader:
            # Move batch data to the correct device
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)

            optimizer.zero_grad()      # Reset gradients before each batch
            outputs = model(x_batch)   # Forward pass through the model
            loss = criterion(outputs, y_batch)  # Compute batch loss
            loss.backward()            # Backpropagate errors
            optimizer.step()           # Update weights using optimizer
            total_loss += loss.item()  # Accumulate batch loss for reporting

        # Optional: print or log average epoch loss
        # avg_loss = total_loss / len(train_loader)
        # print(f"Epoch {epoch+1}, Loss: {avg_loss:.3f}")

    # ------------------------------------------------------
    # 4. Validation step
    # ------------------------------------------------------
    # Evaluate model on validation data to measure performance.
    # 'evaluate_text_model' should return (loss, accuracy)
    val_loss, val_acc = evaluate_text_model(model, val_loader, criterion, device)

    # Report metric to Optuna (used for pruning and progress tracking)
    trial.report(val_acc, epoch)

    # ------------------------------------------------------
    # 5. Return objective value
    # ------------------------------------------------------
    # Optuna minimizes the objective, so we negate accuracy
    # to maximize it effectively.
    return -val_acc
Step 2: Run the Optimization
# ----------------------------------------------------------
# 1. Create an Optuna study object
# ----------------------------------------------------------
# Optuna is an automatic hyperparameter optimization framework.
# 'direction="minimize"' means the objective function's goal 
# is to minimize the evaluation metric (e.g., validation loss).
# For maximization tasks (e.g., accuracy), use direction="maximize".
study = optuna.create_study(direction="minimize")

# ----------------------------------------------------------
# 2. Run the optimization process
# ----------------------------------------------------------
# study.optimize():
#   - Runs the user-defined 'objective' function multiple times (n_trials)
#   - Each trial corresponds to one set of hyperparameters suggested by Optuna
#   - The objective function must return a scalar score (e.g., validation loss)
#   - Optuna internally searches for the best hyperparameters using Bayesian optimization
study.optimize(objective, n_trials=10)

# ----------------------------------------------------------
# 3. Display the best hyperparameters found
# ----------------------------------------------------------
# study.best_params returns a dictionary of parameter names and their best values
# according to the optimization results (i.e., the lowest objective value).
print("Best Parameters:", study.best_params)
Step 3: Track Experiments with Weights & Biases (Optional)
import wandb

# ----------------------------------------------------------
# 1. Initialize a new W&B run
# ----------------------------------------------------------
#   - project="imdb-tuning": specifies the W&B project where results are stored.
#   - config=study.best_params: logs the best hyperparameters found from Optuna (or another tuner).
#   - Each call to wandb.init() starts a new run on the W&B dashboard.
wandb.init(project="imdb-tuning", config=study.best_params)

# ----------------------------------------------------------
# 2. Log metrics and hyperparameters
# ----------------------------------------------------------
#   - wandb.log(): records key metrics and hyperparameter values to W&B.
#   - "best_val_acc": best validation accuracy achieved (note: -study.best_value is used because Optuna minimizes loss, so we negate it).
#   - "embed_dim", "hidden_dim", "lr": model hyperparameters from the tuning study.
wandb.log({
    "best_val_acc": -study.best_value,                  # Best validation accuracy (negated loss)
    "embed_dim": study.best_params["embed_dim"],        # Embedding dimension used in the model
    "hidden_dim": study.best_params["hidden_dim"],      # Hidden layer size in the RNN or MLP
    "lr": study.best_params["lr"]                       # Learning rate used for training
})

# ----------------------------------------------------------
# 3. Finalize the W&B run
# ----------------------------------------------------------
#   - Ensures that all metrics and configuration data are properly synced.
#   - Closes the active run gracefully to prevent logging overlap in subsequent runs.
wandb.finish()
Summary of Both Pipelines
Pipeline Tuning Approach Parameters Explored Tool Used Outcome
Vision (CIFAR-10 CNN) Manual Grid Search lr, batch size, dropout itertools + pandas Top configurations ranked and visualized
NLP (IMDb LSTM) Automated Bayesian Search lr, embed_dim, hidden_dim Optuna Optimal configuration discovered efficiently
Key Takeaways

Model Evaluation, Benchmarking, and Reporting

Overview

Dataset Splitting and Evaluation Protocols

Standard Split
Cross-Validation
Stratified Sampling

Quantitative Evaluation Metrics

Classification Metrics
Regression Metrics
NLP and Generation Metrics

Practical Evaluation Code Example

Classification Example (Vision or NLP)
from sklearn.metrics import accuracy_score, f1_score, confusion_matrix, classification_report
import seaborn as sns
import matplotlib.pyplot as plt

# ----------------------------------------------------------
# 1. Define evaluation function for model performance metrics
# ----------------------------------------------------------
def evaluate_metrics(model, loader):
    model.eval()  # Set model to evaluation mode (disable dropout/batchnorm)
    all_preds, all_labels = [], []  # Lists to store predictions and ground truth labels

    # Disable gradient computation for faster inference and lower memory usage
    with torch.no_grad():
        # Iterate over all batches in the provided DataLoader
        for x_batch, y_batch in loader:
            outputs = model(x_batch)          # Forward pass through the model
            preds = outputs.argmax(dim=1)     # Get predicted class indices (max logit)
            # Move predictions and labels to CPU and convert to NumPy arrays
            all_preds.extend(preds.cpu().numpy())
            all_labels.extend(y_batch.cpu().numpy())

    # ------------------------------------------------------
    # 2. Compute quantitative evaluation metrics
    # ------------------------------------------------------
    # Accuracy: proportion of correct predictions
    acc = accuracy_score(all_labels, all_preds)
    # Weighted F1-score: harmonic mean of precision and recall, weighted by class frequency
    f1 = f1_score(all_labels, all_preds, average='weighted')

    # Print metrics summary
    print(f"Accuracy: {acc:.3f}, F1-score: {f1:.3f}")

    # Return predictions and true labels for further analysis
    return all_preds, all_labels


# ----------------------------------------------------------
# 3. Generate predictions on the test set and visualize confusion matrix
# ----------------------------------------------------------
preds, labels = evaluate_metrics(model, test_loader)  # Evaluate model and get outputs
cm = confusion_matrix(labels, preds)  # Compute confusion matrix (true vs. predicted)

# ----------------------------------------------------------
# 4. Visualize confusion matrix using seaborn heatmap
# ----------------------------------------------------------
plt.figure(figsize=(6,5))
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')  # Display counts with annotations
plt.xlabel('Predicted')  # X-axis label for predicted classes
plt.ylabel('True')       # Y-axis label for true classes
plt.title('Confusion Matrix')  # Add title for clarity
plt.show()

Benchmarking Across Models

Example
Model Params Val Accuracy Test Accuracy F1-score
CNN Baseline 2.5M 84.2% 83.5% 0.835
CNN + Augmentation 2.5M 87.8% 86.9% 0.869
ResNet18 11.2M 90.1% 89.7% 0.897
Vision Transformer 86.4M 92.3% 91.9% 0.919

Qualitative Evaluation and Error Analysis

Error Bucketing
Visualization for Vision Models
Visualization for Text Models
# Loop over the first 3 samples in the dataset (for inspection or debugging)
for i in range(3):
    # Print the original text input (raw sentence)
    print(f"Text: {texts[i]}")
    
    # Print both the true label and the model’s predicted label for comparison
    print(f"True label: {labels[i]}, Predicted: {preds[i]}")

Reporting and Documentation

Statistical Significance Testing

FAQs

Practical Implementation – Model Evaluation, Benchmarking, and Reporting

Example 1: Vision Use-Case (Evaluating CIFAR-10 CNN)
Step 1: Load Model and Prepare Test Data
import torch
import torch.nn as nn
from torchvision import datasets, transforms
from torch.utils.data import DataLoader

# ----------------------------------------------------------
# 1. Define test data transformations
# ----------------------------------------------------------
# Use the same normalization statistics (mean and std) as used during training.
# This ensures consistency between training and evaluation pipelines.
#   - ToTensor(): converts image from [0, 255] $$\rightarrow$$ [0, 1] tensor
#   - Normalize(): standardizes pixel values using precomputed CIFAR-10 stats
transform_test = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465),
                         (0.2023, 0.1994, 0.2010))
])

# ----------------------------------------------------------
# 2. Load CIFAR-10 test set
# ----------------------------------------------------------
#   - train=False ensures we’re loading only the test split
#   - download=True automatically downloads dataset if missing
#   - transform applies preprocessing pipeline defined above
test_dataset = datasets.CIFAR10(root="./data", train=False, download=True, transform=transform_test)

# Wrap dataset with DataLoader for batch iteration
#   - batch_size=64 controls number of images per evaluation batch
#   - shuffle=False ensures deterministic order for evaluation
test_loader = DataLoader(test_dataset, batch_size=64, shuffle=False)

# ----------------------------------------------------------
# 3. Load the trained model
# ----------------------------------------------------------
# Instantiate your CNN architecture (same class definition used in training).
# It must match the model structure saved in "best_cifar10_model.pt".
model = CNNClassifier()

# Load model weights from checkpoint file.
# torch.load() returns a dictionary of parameter tensors.
model.load_state_dict(torch.load("best_cifar10_model.pt"))

# Switch model to evaluation mode.
# This disables dropout, batch normalization updates, etc.
model.eval()

# ----------------------------------------------------------
# 4. Configure computation device
# ----------------------------------------------------------
#   - Use GPU (CUDA) if available; otherwise, fall back to CPU
#   - Move model to the selected device for inference
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

# ----------------------------------------------------------
# (Optional) Evaluate on test data
# ----------------------------------------------------------
# You could now loop through test_loader and compute accuracy:
# correct, total = 0, 0
# with torch.no_grad():
#     for images, labels in test_loader:
#         images, labels = images.to(device), labels.to(device)
#         outputs = model(images)
#         preds = outputs.argmax(dim=1)
#         correct += (preds == labels).sum().item()
#         total += labels.size(0)
# print(f"Test Accuracy: {100 * correct / total:.2f}%")
Step 2: Generate Predictions and Compute Metrics
from sklearn.metrics import accuracy_score, f1_score, classification_report, confusion_matrix
import numpy as np

# ----------------------------------------------------------
# 1. Initialize containers for predictions and labels
# ----------------------------------------------------------
# We'll collect all model predictions and true labels
# from the test set to compute evaluation metrics later.
all_preds, all_labels = [], []

# ----------------------------------------------------------
# 2. Run inference on the test dataset
# ----------------------------------------------------------
# torch.no_grad() disables gradient computation:
#   - reduces memory usage
#   - speeds up inference
# since we don't need to backpropagate during evaluation.
with torch.no_grad():
    for images, labels in test_loader:
        # Move inputs and labels to GPU/CPU device as appropriate
        images, labels = images.to(device), labels.to(device)
        
        # Forward pass through the trained model
        outputs = model(images)
        
        # Get the predicted class (index of the highest logit)
        preds = outputs.argmax(dim=1)
        
        # preds and labels are PyTorch tensors that currently live on the GPU (if device='cuda')
        # .cpu() moves them from GPU memory to CPU memory so that NumPy (and scikit-learn) can use them.
        # .numpy() converts the PyTorch tensor into a NumPy array — sklearn functions expect NumPy arrays, not PyTorch tensors.
        # .extend() adds all elements of that NumPy array to the Python list 'all_preds' (flattening it rather than appending as a nested array).
        all_preds.extend(preds.cpu().numpy())
        all_labels.extend(labels.cpu().numpy())

# ----------------------------------------------------------
# 3. Compute evaluation metrics
# ----------------------------------------------------------
# Accuracy: overall proportion of correct predictions
acc = accuracy_score(all_labels, all_preds)

# F1-score: harmonic mean of precision and recall
# "weighted" accounts for class imbalance
f1 = f1_score(all_labels, all_preds, average="weighted")

# Display summarized results
print(f"Test Accuracy: {acc:.3f}, Weighted F1-score: {f1:.3f}")

# ----------------------------------------------------------
# 4. (Optional) Detailed reporting
# ----------------------------------------------------------
# For deeper analysis, you can uncomment these lines:
# print(classification_report(all_labels, all_preds))
# print("Confusion Matrix:\n", confusion_matrix(all_labels, all_preds))
Step 3: Visualize Confusion Matrix
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.metrics import confusion_matrix  # Ensure you import this if not already

# ----------------------------------------------------------
# 1. Compute the confusion matrix
# ----------------------------------------------------------
# - all_labels: ground-truth labels collected from test set
# - all_preds: model predictions from test set
# - confusion_matrix() returns a 2D array (num_classes x num_classes)
#   where entry (i, j) represents the number of samples with
#   true label i and predicted label j
cm = confusion_matrix(all_labels, all_preds)

# ----------------------------------------------------------
# 2. Get class names from the dataset
# ----------------------------------------------------------
# test_dataset.classes contains human-readable class names for CIFAR-10
# e.g., ['airplane', 'automobile', 'bird', 'cat', 'deer', 'dog', 'frog', 'horse', 'ship', 'truck']
classes = test_dataset.classes

# ----------------------------------------------------------
# 3. Create the confusion matrix heatmap
# ----------------------------------------------------------
plt.figure(figsize=(8, 6))  # Define figure size for better readability

# Use seaborn heatmap for visualization
# - annot=True: show numeric values in each cell
# - fmt='d': format annotation as integers
# - cmap="Blues": use a blue color gradient
# - xticklabels / yticklabels: set axis labels to class names
sns.heatmap(cm, annot=True, fmt='d', cmap="Blues", 
            xticklabels=classes, yticklabels=classes)

# ----------------------------------------------------------
# 4. Label axes and title
# ----------------------------------------------------------
plt.xlabel("Predicted Labels")  # X-axis represents model predictions
plt.ylabel("True Labels")       # Y-axis represents actual ground-truth labels
plt.title("Confusion Matrix - CIFAR-10 CNN")  # Add descriptive title

# ----------------------------------------------------------
# 5. Display the plot
# ----------------------------------------------------------
plt.show()  # Render the confusion matrix visualization
Step 4: Benchmark Report Across Model Variants
import pandas as pd  # Import pandas for tabular data handling

# ----------------------------------------------------------
# 1. Create a DataFrame to store benchmark results
# ----------------------------------------------------------
# Each dictionary in the list represents the results of one experiment/model.
# The keys correspond to column names, and values are the metrics being tracked:
#   - "Model": model architecture name
#   - "Params (M)": number of trainable parameters (in millions)
#   - "Test Acc": test set accuracy
#   - "F1": F1-score, a balanced measure of precision and recall
benchmark_results = pd.DataFrame([
    {"Model": "SimpleCNN", "Params (M)": 1.2, "Test Acc": 0.84, "F1": 0.835},
    {"Model": "ResNet18", "Params (M)": 11.2, "Test Acc": 0.91, "F1": 0.908},
    {"Model": "ViT-Tiny", "Params (M)": 5.6, "Test Acc": 0.89, "F1": 0.885}
])

# ----------------------------------------------------------
# 2. Display the benchmark table
# ----------------------------------------------------------
# Printing the DataFrame shows a formatted comparison of all models
# Useful for summarizing experiments and comparing performance trade-offs
print(benchmark_results)
Model Params (M) Test Acc F1
SimpleCNN 1.2 0.84 0.835
ResNet18 11.2 0.91 0.908
ViT-Tiny 5.6 0.89 0.885
Example 2: NLP Use-Case (Evaluating IMDb Sentiment Model)
Step 1: Load Model and Prepare Test Loader
from torchtext.datasets import IMDB
from torch.utils.data import DataLoader
from torch.nn.utils.rnn import pad_sequence

# ----------------------------------------------------------
# 1. Load IMDB test dataset
# ----------------------------------------------------------
#   - The IMDB dataset is a sentiment classification dataset (pos/neg reviews)
#   - 'split="test"' loads only the test partition for evaluation
test_iter = IMDB(split='test')

# Convert the test iterator into a list for indexing and batching
# Each element is a tuple: (label, text)
test_list = list(test_iter)

# ----------------------------------------------------------
# 2. Create DataLoader for batching
# ----------------------------------------------------------
#   - batch_size=32: evaluate 32 samples at a time
#   - collate_fn=collate_batch: custom function handles tokenization,
#     numericalization, and dynamic padding to equal sequence lengths
test_loader = DataLoader(test_list, batch_size=32, collate_fn=collate_batch)

# ----------------------------------------------------------
# 3. Initialize model for inference
# ----------------------------------------------------------
#   - SentimentRNN: previously defined model (e.g., GRU or LSTM-based)
#   - vocab_size: size of vocabulary used during training
#   - embed_dim: dimensionality of word embeddings
#   - hidden_dim: hidden layer size in RNN
#   - output_dim: number of output classes (2 $$\rightarrow$$ positive / negative)
#   - pad_idx: index of padding token, ensures embeddings ignore padding
model = SentimentRNN(vocab_size=len(vocab), embed_dim=64, hidden_dim=128, output_dim=2, pad_idx=pad_idx)

# ----------------------------------------------------------
# 4. Load the best saved model weights
# ----------------------------------------------------------
#   - The model was saved earlier during training using torch.save()
#   - Restoring ensures consistent evaluation with the best checkpoint
model.load_state_dict(torch.load("best_imdb_model.pt"))

# ----------------------------------------------------------
# 5. Set model to evaluation mode
# ----------------------------------------------------------
#   - Disables dropout and batch normalization updates
#   - Ensures deterministic inference behavior
model.eval()

# ----------------------------------------------------------
# 6. Move model to appropriate device (CPU or GPU)
# ----------------------------------------------------------
#   - 'device' is typically defined as torch.device("cuda" if available)
#   - Ensures data and model reside on the same device during inference
model.to(device)
Step 2: Compute Predictions and Metrics
from sklearn.metrics import classification_report, confusion_matrix

# ----------------------------------------------------------
# 1. Initialize lists to store predictions and true labels
# ----------------------------------------------------------
# We'll collect all predictions and labels from the test set
# to compute overall metrics after the loop.
all_preds, all_labels = [], []

# ----------------------------------------------------------
# 2. Disable gradient computation for evaluation
# ----------------------------------------------------------
# We don't need gradients during inference, so this speeds up
# computation and reduces memory usage.
with torch.no_grad():
    # Iterate over all test batches
    for x_batch, y_batch in test_loader:
        # Move data to the same device as the model (CPU or GPU)
        x_batch, y_batch = x_batch.to(device), y_batch.to(device)

        # Forward pass: compute model outputs (logits)
        outputs = model(x_batch)

        # Get predicted class indices (highest logit per sample)
        preds = outputs.argmax(dim=1)

        # preds and labels are PyTorch tensors that currently live on the GPU (if device='cuda')
        # .cpu() moves them from GPU memory to CPU memory so that NumPy (and scikit-learn) can use them.
        # .numpy() converts the PyTorch tensor into a NumPy array — sklearn functions expect NumPy arrays, not PyTorch tensors.
        # .extend() adds all elements of that NumPy array to the Python list 'all_preds' (flattening it rather than appending as a nested array).
        all_preds.extend(preds.cpu().numpy())
        all_labels.extend(y_batch.cpu().numpy())

# ----------------------------------------------------------
# 3. Generate a detailed classification report
# ----------------------------------------------------------
# classification_report provides precision, recall, F1-score,
# and support for each class (here: negative/positive).
print(classification_report(all_labels, all_preds, target_names=["negative", "positive"]))
              precision    recall  f1-score   support

    negative       0.89      0.86      0.87      12500
    positive       0.87      0.90      0.89      12500

    accuracy                           0.88      25000
   macro avg       0.88      0.88      0.88      25000
weighted avg       0.88      0.88      0.88      25000
Step 3: Inspect Misclassified Examples
# ----------------------------------------------------------
# 1. Select a small subset of test samples for inspection
# ----------------------------------------------------------
# Extract only the text portions from the first 10 test samples
texts = [text for _, text in test_list[:10]]
# Extract their corresponding ground-truth sentiment labels
true_labels = [label for label, _ in test_list[:10]]

# Instead of separately extracting the texts and labels (like above), can do:
# true_labels, texts = map(list, zip(*test_list[:10]))
# The * operator unpacks the list of tuples so that each (label, text) pair
# is passed as a separate argument to zip().
# zip(*test_list[:10]) then groups the first elements (all labels) together
# and the second elements (all texts) together into two tuples.
# This avoids writing two separate list comprehensions and keeps the code concise.

# ----------------------------------------------------------
# 2. Switch model to evaluation mode
# ----------------------------------------------------------
# Disables dropout and batch normalization updates for deterministic inference
model.eval()

# ----------------------------------------------------------
# 3. Run inference for each text sample
# ----------------------------------------------------------
for text, true_label in zip(texts, true_labels):

    # Tokenize the input text and map tokens to integer IDs using the vocabulary
    # 'tokenizer(text)' $$\rightarrow$$ list of tokens
    # 'vocab(token_list)' $$\rightarrow$$ list of token IDs
    # Convert to a tensor and add batch dimension using unsqueeze(0)
    # Move tensor to the same device as the model (CPU or GPU)
    tokens = torch.tensor(vocab(tokenizer(text)), dtype=torch.long).unsqueeze(0).to(device)

    # Perform forward pass through the model and get class prediction
    # argmax(dim=1) returns the predicted class index (0 or 1)
    pred = model(tokens).argmax(dim=1).item()

    # Convert numerical prediction to human-readable label
    pred_label = "pos" if pred == 1 else "neg"

    # ----------------------------------------------------------
    # 4. Display prediction results
    # ----------------------------------------------------------
    # Print true label, predicted label, and a short snippet of the review text
    print(f"True: {true_label}, Predicted: {pred_label}")
    print(f"Snippet: {text[:120]}...\n")

Explanation: Manually inspecting errors often reveals patterns like negations (“not good”) or sarcasm that models misinterpret — guiding future dataset improvements or model changes.

Step 4: Report Comparison Across Experiments
import pandas as pd

# ----------------------------------------------------------
# 1. Create a DataFrame for NLP model benchmark comparison
# ----------------------------------------------------------
# Each row corresponds to a model evaluated on a text classification task.
# Columns capture:
#   - Model: model architecture name
#   - Params (M): number of parameters in millions (model size)
#   - Val Acc: validation accuracy (on held-out data)
#   - Test Acc: test accuracy (on unseen data)
#   - F1: F1-score (harmonic mean of precision and recall)
nlp_benchmarks = pd.DataFrame([
    {"Model": "LSTM", "Params (M)": 0.8, "Val Acc": 0.87, "Test Acc": 0.88, "F1": 0.88},
    {"Model": "BiLSTM", "Params (M)": 1.5, "Val Acc": 0.89, "Test Acc": 0.89, "F1": 0.89},
    {"Model": "DistilBERT", "Params (M)": 66, "Val Acc": 0.92, "Test Acc": 0.91, "F1": 0.91}
])

# ----------------------------------------------------------
# 2. Display benchmark table
# ----------------------------------------------------------
# Printing the DataFrame shows a structured table comparing model sizes
# and performance metrics side by side.
print(nlp_benchmarks)
Model Params (M) Val Acc Test Acc F1
LSTM 0.8 0.87 0.88 0.88
BiLSTM 1.5 0.89 0.89 0.89
DistilBERT 66 0.92 0.91 0.91
Summary of the Evaluation Workflow
Step Vision Task NLP Task
Metrics Computed Accuracy, F1, Confusion Matrix Precision, Recall, F1, Report
Visualization Heatmap of confusion Error sample inspection
Benchmarking Table comparing CNNs and Transformers Table comparing LSTM vs Transformer models
Goal Interpret model strengths and weaknesses Diagnose misclassification patterns
Key Takeaways

Model Deployment, Monitoring, and Continuous Evaluation

Overview

Deployment Modes

Deployment Mode Description Typical Use Case
Batch Inference Run predictions on large datasets periodically Analytics, risk scoring, nightly reports
Online Inference (API) Serve predictions via REST/gRPC endpoint Chatbots, real-time recommendations
Edge Deployment Deploy on mobile or embedded devices Offline applications, IoT
Serverless / Cloud Functions Stateless, on-demand inference Event-driven systems
Hybrid Deployment Combination (e.g., preprocess offline, predict online) Large-scale production pipelines

Model Serialization and Export

TorchScript Export
import torch

# Example model
model = CNN()
traced_model = torch.jit.trace(model, torch.randn(1, 3, 32, 32))
torch.jit.save(traced_model, "cnn_model.pt")
ONNX Export
dummy_input = torch.randn(1, 3, 32, 32)
torch.onnx.export(model, dummy_input, "cnn_model.onnx", input_names=['input'], output_names=['output'])

Serving Infrastructure

Option 1: TorchServe
torch-model-archiver --model-name cnn --version 1.0 --serialized-file cnn_model.pt --handler image_classifier
Option 2: FastAPI or Flask
from fastapi import FastAPI, UploadFile
import torch

# ----------------------------------------------------------
# 1. Initialize FastAPI application
# ----------------------------------------------------------
# FastAPI is used to create an HTTP API for model inference.
# It automatically generates interactive Swagger docs at /docs.
app = FastAPI()

# ----------------------------------------------------------
# 2. Load the trained model
# ----------------------------------------------------------
# TorchScript model is loaded from file for deployment.
# TorchScript allows the model to be portable and run without the full Python source.
model = torch.jit.load("cnn_model.pt")

# Set model to evaluation mode to disable dropout/batchnorm updates
model.eval()

# ----------------------------------------------------------
# 3. Define inference endpoint
# ----------------------------------------------------------
# This route handles POST requests at /predict.
# Users upload an image file, and the server returns a predicted label.
@app.post("/predict")
async def predict(file: UploadFile):
    # ------------------------------------------------------
    # Step 1: Preprocess the uploaded image
    # ------------------------------------------------------
    # The `preprocess_image()` function (to be implemented)
    # should handle:
    #   - Reading file bytes
    #   - Converting to PIL image or tensor
    #   - Applying transforms (resize, normalize, etc.)
    image = preprocess_image(file)

    # ------------------------------------------------------
    # Step 2: Run inference
    # ------------------------------------------------------
    # Disable gradient tracking for faster, memory-efficient inference.
    with torch.no_grad():
        output = model(image)  # Forward pass through the model

    # ------------------------------------------------------
    # Step 3: Format and return prediction
    # ------------------------------------------------------
    # The model output is a tensor of class logits.
    # `argmax()` selects the index of the highest-scoring class.
    # `.item()` converts it to a Python integer for JSON serialization.
    return {"prediction": output.argmax().item()}
Option 3: Cloud Deployments

Inference Optimization

Quantization
quantized_model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)
Pruning
Batch Inference and Caching

Continuous Monitoring and Feedback Loops

System Metrics
Model Metrics
Example: Drift Detection
\[D_{KL}(P \mid \mid Q) = \sum_i P(i) \log\frac{P(i)}{Q(i)}\]

Model Versioning and Retraining

Version Control
Retraining Workflow
  1. Detect drift or performance degradation.
  2. Trigger retraining with fresh data.
  3. Validate model against existing baselines.
  4. Roll out gradually (shadow or A/B testing).
  5. Promote new model if performance improves.

Continuous Evaluation

Shadow Mode
A/B Testing
Feedback Integration

Logging and Observability

Deployment Safety and Compliance

FAQs

Practical Implementation – Model Deployment, Monitoring, and Continuous Evaluation

Example 1: Vision Use-Case (Deploying CIFAR-10 CNN as an API)
Step 1: Export the Trained Model
import torch

# ----------------------------------------------------------
# 1. Load the trained model checkpoint
# ----------------------------------------------------------
# Instantiate the CNN model architecture (must match the training definition)
model = CNNClassifier()

# Load the saved model weights from checkpoint
#   - torch.load() loads serialized tensors from disk
#   - load_state_dict() restores weights into the model
model.load_state_dict(torch.load("best_cifar10_model.pt"))

# Set the model to evaluation mode
#   - Disables dropout and batch normalization updates
#   - Ensures deterministic behavior during inference
model.eval()

# ----------------------------------------------------------
# 2. Export the model to TorchScript
# ----------------------------------------------------------
# TorchScript allows saving a static, optimized version of the model
# that can run independently of Python (e.g., in C++ or mobile environments)

# Create an example input tensor matching the model’s expected input shape:
#   - batch_size=1, channels=3 (RGB), height=32, width=32
example_input = torch.randn(1, 3, 32, 32)

# Use torch.jit.trace() to record operations from a single forward pass
#   - Generates a TorchScript graph (computational graph representation)
traced_model = torch.jit.trace(model, example_input)

# Save the scripted (traced) model to disk for deployment
#   - This .pt file can be loaded directly in production environments
torch.jit.save(traced_model, "cnn_cifar10_scripted.pt")

# ----------------------------------------------------------
# 3. Confirmation message
# ----------------------------------------------------------
print("TorchScript model saved as cnn_cifar10_scripted.pt")
Step 2: Create a FastAPI Inference Server
from fastapi import FastAPI, UploadFile
from PIL import Image
import io
import torchvision.transforms as transforms
import time
import torch  # Added import for model inference

# ----------------------------------------------------------
# 1. Initialize FastAPI app
# ----------------------------------------------------------
# FastAPI provides an easy way to build REST APIs for ML model deployment.
app = FastAPI()

# ----------------------------------------------------------
# 2. Load TorchScript model
# ----------------------------------------------------------
# TorchScript models are serialized PyTorch models that can be loaded
# without needing the original Python class definitions.
#   - cnn_cifar10_scripted.pt is a TorchScript version of your trained CNN.
#   - model.eval() puts the model into inference mode (disables dropout, BN updates).
model = torch.jit.load("cnn_cifar10_scripted.pt")
model.eval()

# ----------------------------------------------------------
# 3. Define preprocessing (same as training normalization)
# ----------------------------------------------------------
# The preprocessing pipeline ensures input images match the distribution
# of images used during training (same normalization and resizing).
transform = transforms.Compose([
    transforms.Resize((32, 32)),                        # Resize input to CIFAR-10 dimensions
    transforms.ToTensor(),                              # Convert PIL Image $$\rightarrow$$ Tensor (C, H, W)
    transforms.Normalize((0.4914, 0.4822, 0.4465),      # Normalize per CIFAR-10 channel mean
                         (0.2023, 0.1994, 0.2010))      # Normalize per CIFAR-10 channel std
])

# ----------------------------------------------------------
# 4. Define the inference endpoint
# ----------------------------------------------------------
# Endpoint: POST /predict
# Accepts an uploaded image file and returns the predicted class index.
@app.post("/predict")
async def predict(file: UploadFile):
    start_time = time.time()  # Measure inference latency

    # ------------------------------------------------------
    # 4.1 Load and preprocess the input image
    # ------------------------------------------------------
    #   - Read uploaded file bytes from the request.
    #   - Convert bytes into a PIL image.
    #   - Ensure RGB format (some inputs may be grayscale or RGBA).
    #   - Apply preprocessing and add batch dimension.
    image_bytes = await file.read()
    image = Image.open(io.BytesIO(image_bytes)).convert("RGB")
    input_tensor = transform(image).unsqueeze(0)  # Add batch dimension (1, 3, 32, 32)

    # ------------------------------------------------------
    # 4.2 Run model inference
    # ------------------------------------------------------
    #   - torch.no_grad(): disables gradient tracking for faster inference.
    #   - model(input_tensor): forward pass.
    #   - outputs.argmax(dim=1): returns index of the highest-probability class.
    with torch.no_grad():
        outputs = model(input_tensor)
        pred = outputs.argmax(dim=1).item()

    latency = time.time() - start_time  # Compute total latency

    # ------------------------------------------------------
    # 4.3 Log and return prediction
    # ------------------------------------------------------
    #   - Prints result and latency for server monitoring.
    #   - Returns JSON response to the client.
    print(f"Prediction: {pred}, Latency: {latency:.3f}s")
    return {"predicted_class": int(pred), "latency_seconds": latency}
uvicorn app:app --reload
curl -X POST -F "file=@test_image.jpg" http://localhost:8000/predict
Step 3: Add Basic Monitoring (Latency & Drift Tracking)
import numpy as np
from collections import deque

# ----------------------------------------------------------
# 1. Rolling metric buffers for latency and confidence
# ----------------------------------------------------------
# Using deques to maintain a fixed-size window (maxlen=100) for recent predictions.
# These help track average inference latency and confidence over the last 100 requests.
latency_buffer = deque(maxlen=100)
confidence_buffer = deque(maxlen=100)

# ----------------------------------------------------------
# 2. Define API endpoint for model inference
# ----------------------------------------------------------
# This asynchronous FastAPI endpoint handles file uploads (e.g., image input)
@app.post("/predict")
async def predict(file: UploadFile):
    start = time.time()  # Record start time to measure inference latency

    # ------------------------------------------------------
    # 3. Read and preprocess the uploaded image
    # ------------------------------------------------------
    # - Read the uploaded file bytes asynchronously
    # - Open as a PIL image and convert to RGB format
    # - Apply pre-defined transform (resize, normalize, tensor conversion, etc.)
    # - Add batch dimension with unsqueeze(0)
    image_bytes = await file.read()
    image = Image.open(io.BytesIO(image_bytes)).convert("RGB")
    input_tensor = transform(image).unsqueeze(0)

    # ------------------------------------------------------
    # 4. Perform model inference
    # ------------------------------------------------------
    # - Disable gradient computation with torch.no_grad() to reduce memory usage
    # - Pass the input through the model
    # - Apply softmax to get class probabilities
    # - Extract predicted class (argmax) and its confidence (max probability)
    with torch.no_grad():
        outputs = model(input_tensor)
        probs = torch.softmax(outputs, dim=1).cpu().numpy()[0]
        pred = int(np.argmax(probs))
        confidence = float(np.max(probs))

    # ------------------------------------------------------
    # 5. Compute and record inference metrics
    # ------------------------------------------------------
    # - Calculate latency for this request
    # - Append latency and confidence to rolling buffers
    latency = time.time() - start
    latency_buffer.append(latency)
    confidence_buffer.append(confidence)

    # ------------------------------------------------------
    # 6. Compute rolling averages (for live monitoring)
    # ------------------------------------------------------
    # - Calculate moving averages for latency and confidence
    # - Print metrics to logs for tracking model serving performance
    avg_latency = np.mean(latency_buffer)
    avg_confidence = np.mean(confidence_buffer)
    print(f"Pred={pred}, Conf={confidence:.3f}, Avg Lat={avg_latency:.3f}s, Avg Conf={avg_confidence:.3f}")

    # ------------------------------------------------------
    # 7. Return JSON response
    # ------------------------------------------------------
    # Send prediction result, current confidence, and rolling average latency
    return {
        "prediction": pred,
        "confidence": confidence,
        "avg_latency": avg_latency
    }
Step 4: Automate Retraining via Drift Detection
from scipy.stats import entropy
import numpy as np

# ----------------------------------------------------------
# 1. Define baseline (reference) class distribution
# ----------------------------------------------------------
# Represents the original class probability distribution from the training dataset.
# Here it's uniform (10 classes, each 10%) — meaning no class imbalance initially.
train_distribution = np.array([0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1])


# ----------------------------------------------------------
# 2. Drift detection function
# ----------------------------------------------------------
def detect_drift(predictions):
    """
    Compare the current prediction class distribution with the baseline
    using Kullback–Leibler (KL) divergence (a.k.a. relative entropy).
    
    Args:
        predictions (list or np.array): recent model predictions (class indices)
    Returns:
        drift_score (float): KL divergence value — higher means more drift.
    """
    # Compute normalized histogram of current predictions (as probability distribution)
    current_distribution = np.bincount(predictions, minlength=10) / len(predictions)
    
    # Compute KL divergence between current and baseline distributions
    # KL divergence quantifies how much the current distribution differs from the baseline.
    drift_score = entropy(current_distribution, train_distribution)
    
    return drift_score


# ----------------------------------------------------------
# 3. Example usage of drift detection
# ----------------------------------------------------------
# Recent predictions simulate new model outputs — possibly from recent data.
recent_preds = [0, 0, 0, 0, 0, 3, 3, 3, 8, 8, 8]

# Compute drift score
drift = detect_drift(recent_preds)

# ----------------------------------------------------------
# 4. Threshold-based alerting
# ----------------------------------------------------------
# If drift score exceeds a small threshold, it signals significant distributional shift.
if drift > 0.05:
    print("Warning: Prediction drift detected, consider retraining.")
else:
    print("Model predictions are stable.")
Example 2: NLP Use-Case (Deploying IMDb Sentiment Classifier)
Step 1: Prepare Model for Inference
import torch
from fastapi import FastAPI
import time

# ----------------------------------------------------------
# 1. Initialize FastAPI app (optional)
# ----------------------------------------------------------
# This line would typically be used to set up an API endpoint
# for serving the model via HTTP requests (e.g., sentiment predictions).
# app = FastAPI()

# ----------------------------------------------------------
# 2. Load pretrained LSTM sentiment model
# ----------------------------------------------------------
# 'SentimentRNN' is assumed to be a custom-defined PyTorch model class
# used for binary sentiment classification (e.g., IMDb dataset).

# Initialize the model with the same architecture and parameters
# that were used during training.
model = SentimentRNN(
    vocab_size=len(vocab),    # Vocabulary size from tokenizer
    embed_dim=64,             # Embedding dimension
    hidden_dim=128,           # Hidden layer size in LSTM
    output_dim=2,             # Number of classes (e.g., positive/negative)
    pad_idx=pad_idx           # Padding index for <pad> tokens
)

# ----------------------------------------------------------
# 3. Load saved model weights
# ----------------------------------------------------------
# Load the trained parameters (weights and biases) from the checkpoint file.
# The file 'best_imdb_model.pt' contains the model state saved after training.
model.load_state_dict(torch.load("best_imdb_model.pt"))

# ----------------------------------------------------------
# 4. Set model to evaluation mode
# ----------------------------------------------------------
# This disables dropout and batch normalization updates,
# ensuring deterministic and consistent outputs during inference.
model.eval()

# ----------------------------------------------------------
# 5. Move model to device (CPU or GPU)
# ----------------------------------------------------------
# If CUDA is available, 'device' would typically be set to 'cuda';
# otherwise, it defaults to 'cpu'. Moving the model to the device
# ensures that both inputs and model parameters are on the same hardware.
model.to(device)
Step 2: Define API for Sentiment Prediction
from torchtext.data.utils import get_tokenizer
tokenizer = get_tokenizer("basic_english")  # Load a basic English tokenizer (splits on spaces/punctuation)

app = FastAPI()  # Initialize the FastAPI web application

# ------------------------------------------------------------
# Define an API endpoint for sentiment analysis
# ------------------------------------------------------------
@app.post("/analyze")
async def analyze_sentiment(text: str):
    # Record start time for latency measurement
    start = time.time()

    # --------------------------------------------------------
    # 1. Text preprocessing
    # --------------------------------------------------------
    # Tokenize the input text using the tokenizer
    # Convert tokens to vocabulary indices, wrap in a tensor, 
    # and add a batch dimension using unsqueeze(0)
    tokens = torch.tensor(vocab(tokenizer(text)), dtype=torch.long).unsqueeze(0).to(device)

    # --------------------------------------------------------
    # 2. Model inference
    # --------------------------------------------------------
    # Disable gradient tracking since we’re in inference mode
    with torch.no_grad():
        outputs = model(tokens)                     # Forward pass through the model
        probs = torch.softmax(outputs, dim=1).cpu().numpy()[0]  # Convert logits $$\rightarrow$$ probabilities
        pred = int(np.argmax(probs))                # Get predicted class index (0=neg, 1=pos)
        confidence = float(np.max(probs))           # Extract confidence score of the prediction

    # --------------------------------------------------------
    # 3. Post-processing
    # --------------------------------------------------------
    latency = time.time() - start                   # Compute total inference time
    label = "positive" if pred == 1 else "negative" # Map class index to human-readable label

    # --------------------------------------------------------
    # 4. Logging and response
    # --------------------------------------------------------
    # Print results to console for monitoring
    print(f"Prediction: {label}, Confidence: {confidence:.2f}, Latency: {latency:.3f}s")

    # Return results as JSON response
    return {
        "sentiment": label,
        "confidence": confidence,
        "latency_seconds": latency
    }
curl -X POST "http://localhost:8000/analyze?text=This+movie+was+excellent+and+moving."
Step 3: Confidence Drift Monitoring
import numpy as np
from collections import deque

# ----------------------------------------------------------
# 1. Initialize rolling logs for tracking confidence and predictions
# ----------------------------------------------------------
# Deques (fixed-length queues) automatically discard the oldest entries
# when new items are appended after reaching max length.
# Used here to maintain a sliding window of the last 100 predictions.
confidence_log = deque(maxlen=100)
prediction_log = deque(maxlen=100)

# ----------------------------------------------------------
# 2. Define FastAPI endpoint for sentiment analysis
# ----------------------------------------------------------
# This function handles POST requests to the /analyze route.
# It accepts a text input, tokenizes it, runs model inference,
# and returns the predicted sentiment and confidence values.
@app.post("/analyze")
async def analyze_sentiment(text: str):
    start = time.time()  # Start latency timer

    # ------------------------------------------------------
    # 3. Tokenize input text and convert to tensor
    # ------------------------------------------------------
    # - The tokenizer converts text $$\rightarrow$$ list of token IDs
    # - vocab() maps each token to an integer index
    # - unsqueeze(0) adds a batch dimension (shape: [1, seq_len])
    # - .to(device) moves tensor to CPU or GPU for inference
    tokens = torch.tensor(vocab(tokenizer(text)), dtype=torch.long).unsqueeze(0).to(device)

    # ------------------------------------------------------
    # 4. Run model inference (disable gradient computation)
    # ------------------------------------------------------
    # - model(tokens) outputs raw logits
    # - softmax converts logits $$\rightarrow$$ probabilities
    # - argmax picks the most likely sentiment label
    # - max(probabilities) gives the model's confidence score
    with torch.no_grad():
        outputs = model(tokens)
        probs = torch.softmax(outputs, dim=1).cpu().numpy()[0]
        pred = int(np.argmax(probs))     # predicted class index
        conf = float(np.max(probs))      # model confidence

    # ------------------------------------------------------
    # 5. Compute and log inference metrics
    # ------------------------------------------------------
    latency = time.time() - start  # Measure response time

    # Append latest prediction and confidence to rolling history
    confidence_log.append(conf)
    prediction_log.append(pred)

    # Compute moving averages for monitoring
    avg_conf = np.mean(confidence_log)        # Average confidence (over last 100)
    pos_ratio = np.mean(np.array(prediction_log) == 1)  # Ratio of positive predictions

    # Print live monitoring metrics for drift detection
    print(f"Avg Conf={avg_conf:.3f}, Pos Ratio={pos_ratio:.3f}")

    # ------------------------------------------------------
    # 6. Return structured JSON response
    # ------------------------------------------------------
    # - "sentiment": model's categorical output
    # - "confidence": model’s current prediction confidence
    # - "avg_confidence": rolling average for monitoring drift or calibration
    return {
        "sentiment": "positive" if pred == 1 else "negative",
        "confidence": conf,
        "avg_confidence": avg_conf
    }
Step 4: Continuous Evaluation via Feedback Integration
# ----------------------------------------------------------
# 1. Initialize feedback storage
# ----------------------------------------------------------
# This list acts as an in-memory "database" to collect user feedback.
# In a production setup, this would typically be replaced by a
# persistent store (e.g., database, message queue, or cloud storage).
feedback_store = []

# ----------------------------------------------------------
# 2. Define API endpoint for feedback collection
# ----------------------------------------------------------
# The FastAPI decorator defines a POST endpoint at "/feedback".
# It receives text and its correct label from the client (user feedback)
# and appends the data to the feedback_store for later retraining.
@app.post("/feedback")
async def record_feedback(text: str, true_label: str):
    # Append the feedback data as a dictionary to the store
    feedback_store.append({"text": text, "label": true_label})
    
    # Log feedback receipt on the server console for monitoring
    print(f"Received feedback: {true_label}")
    
    # Return a simple acknowledgment response to the client
    return {"status": "stored"}

# ----------------------------------------------------------
# 3. Prepare dataset for retraining
# ----------------------------------------------------------
# This function processes collected feedback into a format suitable
# for reusing in fine-tuning or retraining a text classification model.
def prepare_retraining_dataset():
    # Extract the text inputs from collected feedback
    texts = [f["text"] for f in feedback_store]
    
    # Convert textual labels into numeric format (e.g., 1 = positive, 0 = negative)
    labels = [1 if f["label"] == "positive" else 0 for f in feedback_store]
    
    # Print summary of feedback samples to be used for retraining
    print(f"Retraining on {len(texts)} feedback samples")

    # (In a real-world pipeline, you would now tokenize, batch,
    # and feed this data into your model retraining workflow.)
Summary of the Deployment and Monitoring Workflow
Stage Vision (CNN) NLP (LSTM)
Model Export TorchScript Standard PyTorch checkpoint
Serving API FastAPI + file uploads FastAPI + text endpoint
Monitoring Latency, confidence, drift detection Confidence and sentiment ratio monitoring
Feedback Loop Retrain on drift triggers Retrain using user feedback
Key Takeaways

Practical Implementation – End-to-End Example: From Data to Deployment

Example 1: End-to-End Vision Pipeline (CIFAR-10 CNN)

Step 1: Project Structure
vision_pipeline/
├── data/
├── models/
│   ├── cnn_model.py
│   └── best_model.pt
├── train.py
├── evaluate.py
├── serve.py
└── utils.py
Step 2: Data Preparation (utils.py)
from torchvision import datasets, transforms
from torch.utils.data import DataLoader, random_split

def get_data_loaders(batch_size=64):
    """
    Creates DataLoaders for the CIFAR-10 dataset with standard preprocessing,
    including data augmentation for training and normalization for all splits.
    """

    # ----------------------------------------------------------
    # 1. Define training data transformations
    # ----------------------------------------------------------
    # The training transform includes random augmentations to improve model generalization.
    # - RandomHorizontalFlip(): randomly flips images horizontally.
    # - RandomCrop(): crops the image with padding to simulate spatial variation.
    # - ToTensor(): converts a PIL image to a PyTorch tensor (scales pixel values to [0,1]).
    # - Normalize(): standardizes pixel intensities using CIFAR-10 mean and std per channel.
    transform_train = transforms.Compose([
        transforms.RandomHorizontalFlip(),
        transforms.RandomCrop(32, padding=4),
        transforms.ToTensor(),
        transforms.Normalize((0.4914, 0.4822, 0.4465),   # mean for R, G, B channels
                             (0.2023, 0.1994, 0.2010))   # std deviation for R, G, B channels
    ])

    # ----------------------------------------------------------
    # 2. Define test/validation transformations
    # ----------------------------------------------------------
    # No augmentation is applied here to keep evaluation consistent and reproducible.
    transform_test = transforms.Compose([
        transforms.ToTensor(),
        transforms.Normalize((0.4914, 0.4822, 0.4465),
                             (0.2023, 0.1994, 0.2010))
    ])

    # ----------------------------------------------------------
    # 3. Load CIFAR-10 datasets
    # ----------------------------------------------------------
    # - train=True loads the training set (50,000 images)
    # - train=False loads the test set (10,000 images)
    # - transform applies preprocessing to each image dynamically during access.
    # - download=True automatically downloads if not already present in ./data
    train_dataset = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform_train)
    test_dataset = datasets.CIFAR10(root="./data", train=False, download=True, transform=transform_test)

    # ----------------------------------------------------------
    # 4. Split training data into train/validation subsets
    # ----------------------------------------------------------
    #  - The training dataset is split into 45,000 training and 5,000 validation samples.
    #  - random_split ensures a random division of samples each time (for reproducibility, set a manual seed).
    train_set, val_set = random_split(train_dataset, [45000, 5000])

    # ----------------------------------------------------------
    # 5. Create DataLoaders for efficient batching and shuffling
    # ----------------------------------------------------------
    # DataLoaders wrap datasets to:
    #  - batch samples together for efficient GPU processing,
    #  - shuffle the order of samples each epoch (for training),
    #  - use parallel workers to speed up data loading.
    train_loader = DataLoader(train_set, batch_size=batch_size, shuffle=True, num_workers=2)
    val_loader = DataLoader(val_set, batch_size=batch_size)
    test_loader = DataLoader(test_dataset, batch_size=batch_size)

    # ----------------------------------------------------------
    # 6. Return all DataLoaders
    # ----------------------------------------------------------
    # The returned loaders are ready to be used in a model training pipeline.
    return train_loader, val_loader, test_loader
Step 3: Model Definition (models/cnn_model.py)
import torch.nn as nn

# Define a simple Convolutional Neural Network (CNN) for image classification
class CNNClassifier(nn.Module):
    def __init__(self, dropout=0.3):
        super().__init__()

        # ----------------------------------------------------------
        # 1. Convolutional feature extractor block
        # ----------------------------------------------------------
        # This block extracts spatial features from the input image
        # - Conv2d: learns filters to capture visual patterns (edges, textures)
        # - ReLU: introduces non-linearity
        # - MaxPool2d: downsamples spatial dimensions (reduces feature map size)
        self.conv_block = nn.Sequential(
            nn.Conv2d(3, 32, 3, padding=1), nn.ReLU(),   # Input: (3, 32, 32) $$\rightarrow$$ Output: (32, 32, 32)
            nn.MaxPool2d(2),                            # Downsample $$\rightarrow$$ (32, 16, 16)
            nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), # Output: (64, 16, 16)
            nn.MaxPool2d(2)                             # Downsample $$\rightarrow$$ (64, 8, 8)
        )

        # ----------------------------------------------------------
        # 2. Fully connected classification block
        # ----------------------------------------------------------
        # Converts flattened feature maps into class scores
        # - Flatten: reshapes 3D feature maps into 1D vectors
        # - Linear: dense layers learn higher-level combinations of features
        # - Dropout: regularization to prevent overfitting
        # - Output layer: maps to 10 logits (for CIFAR-10’s 10 classes)
        self.fc_block = nn.Sequential(
            nn.Flatten(),                 # (64, 8, 8) $$\rightarrow$$ (4096)
            nn.Linear(64 * 8 * 8, 128),   # Hidden layer
            nn.ReLU(),                    # Non-linearity
            nn.Dropout(dropout),          # Randomly zero out activations
            nn.Linear(128, 10)            # Output layer (10 classes)
        )

    # ----------------------------------------------------------
    # 3. Forward pass
    # ----------------------------------------------------------
    # Defines how data flows through the network layers.
    # Input x passes through conv_block $$\rightarrow$$ fc_block.
    def forward(self, x):
        return self.fc_block(self.conv_block(x))
Step 4: Training Script (train.py)
import torch
import torch.nn as nn
import torch.optim as optim
from utils import get_data_loaders
from models.cnn_model import CNNClassifier

# ----------------------------------------------------------
# Function: train_vision_model
# Trains a CNN-based image classifier on vision data (e.g., CIFAR-10)
# ----------------------------------------------------------
def train_vision_model(lr=1e-3, dropout=0.3, epochs=5):
    # ------------------------------------------------------
    # 1. Get data loaders for training, validation, and test sets
    # ------------------------------------------------------
    # get_data_loaders() is assumed to return three DataLoaders
    # (train_loader, val_loader, test_loader)
    train_loader, val_loader, _ = get_data_loaders()

    # ------------------------------------------------------
    # 2. Initialize model, loss function, and optimizer
    # ------------------------------------------------------
    model = CNNClassifier(dropout=dropout)           # Custom CNN model with dropout
    criterion = nn.CrossEntropyLoss()                # Loss function for multi-class classification
    optimizer = optim.Adam(model.parameters(), lr=lr) # Adam optimizer with learning rate lr

    # ------------------------------------------------------
    # 3. Configure computation device (GPU if available)
    # ------------------------------------------------------
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.to(device)  # Move model to GPU (if available)

    # ------------------------------------------------------
    # 4. Initialize training loop variables
    # ------------------------------------------------------
    best_val_loss = float('inf')  # Track best validation loss for checkpointing

    # ------------------------------------------------------
    # 5. Training loop over epochs
    # ------------------------------------------------------
    for epoch in range(epochs):
        model.train()              # Set model to training mode
        running_loss = 0.0         # Accumulate total loss across batches

        # --------------------------------------------------
        # Iterate over all training batches
        # --------------------------------------------------
        for images, labels in train_loader:
            images, labels = images.to(device), labels.to(device)  # Move batch to GPU if available
            optimizer.zero_grad()           # Reset gradients before each batch
            outputs = model(images)         # Forward pass
            loss = criterion(outputs, labels) # Compute batch loss
            loss.backward()                 # Backward pass (compute gradients)
            optimizer.step()                # Update weights
            running_loss += loss.item()     # Track loss for this batch

        # --------------------------------------------------
        # Compute average training loss for the epoch
        # --------------------------------------------------
        avg_train_loss = running_loss / len(train_loader)

        # --------------------------------------------------
        # 6. Evaluate model on validation set
        # --------------------------------------------------
        val_loss, val_acc = evaluate_vision_model(model, val_loader, criterion, device)

        # --------------------------------------------------
        # 7. Print training and validation progress
        # --------------------------------------------------
        print(f"Epoch {epoch+1}: train_loss={avg_train_loss:.3f}, val_loss={val_loss:.3f}, val_acc={val_acc:.3f}")

        # --------------------------------------------------
        # 8. Save model checkpoint if validation improves
        # --------------------------------------------------
        if val_loss < best_val_loss:
            best_val_loss = val_loss
            torch.save(model.state_dict(), "models/best_model.pt")
            print("✅ Saved new best model checkpoint.")


# ----------------------------------------------------------
# Function: evaluate_vision_model
# Evaluates the model’s loss and accuracy on a given dataset
# ----------------------------------------------------------
def evaluate_vision_model(model, loader, criterion, device):
    model.eval()  # Set model to evaluation mode (disable dropout/batchnorm updates)
    total_loss, correct, total = 0, 0, 0

    # ------------------------------------------------------
    # Disable gradient computation for faster inference
    # ------------------------------------------------------
    with torch.no_grad():
        for images, labels in loader:
            images, labels = images.to(device), labels.to(device)  # Move data to GPU if available
            outputs = model(images)                 # Forward pass
            loss = criterion(outputs, labels)       # Compute loss for batch
            total_loss += loss.item()               # Accumulate loss

            preds = outputs.argmax(dim=1)           # Get predicted class indices
            correct += (preds == labels).sum().item() # Count correctly predicted samples
            total += labels.size(0)                 # Count total samples processed

    # ------------------------------------------------------
    # Return average loss and overall accuracy
    # ------------------------------------------------------
    return total_loss / len(loader), correct / total
python train.py
Step 5: Evaluation Script (evaluate.py)
import torch
from sklearn.metrics import classification_report
from models.cnn_model import CNNClassifier
from utils import get_data_loaders

# ----------------------------------------------------------
# 1. Define test_model() — evaluate a trained CNN on the test set
# ----------------------------------------------------------
def test_model():
    # ------------------------------------------------------
    # Load data
    # ------------------------------------------------------
    # get_data_loaders() is a utility that returns (train_loader, val_loader, test_loader)
    # Here, we only need the test_loader for final evaluation.
    _, _, test_loader = get_data_loaders()

    # ------------------------------------------------------
    # Load the trained model
    # ------------------------------------------------------
    # Initialize model architecture (must match the saved model’s structure)
    model = CNNClassifier()
    # Load the best model weights from checkpoint
    model.load_state_dict(torch.load("models/best_model.pt"))
    model.eval()  # Set to evaluation mode (disables dropout, batchnorm updates)

    # ------------------------------------------------------
    # Set up device for computation
    # ------------------------------------------------------
    # Automatically use GPU if available; otherwise, use CPU.
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.to(device)

    # ------------------------------------------------------
    # Initialize storage for predictions and labels
    # ------------------------------------------------------
    all_preds, all_labels = [], []

    # Disable gradient tracking — faster inference and lower memory usage
    with torch.no_grad():
        # Iterate through the test dataset in batches
        for images, labels in test_loader:
            # Move data to the selected device (GPU/CPU)
            images, labels = images.to(device), labels.to(device)

            # Forward pass through the model to get class logits
            outputs = model(images)

            # Get predicted class indices by selecting the max logit per sample
            preds = outputs.argmax(dim=1)

            # Collect predictions and true labels for later evaluation
            all_preds.extend(preds.cpu().numpy())
            all_labels.extend(labels.cpu().numpy())

    # ------------------------------------------------------
    # Generate classification metrics
    # ------------------------------------------------------
    # classification_report computes precision, recall, F1-score, and support per class
    print(classification_report(all_labels, all_preds))
python evaluate.py
Step 6: Deployment API (serve.py)
from fastapi import FastAPI, UploadFile
import torch
from PIL import Image
import io
from torchvision import transforms
from models.cnn_model import CNNClassifier

# ----------------------------------------------------------
# 1. Initialize FastAPI app
# ----------------------------------------------------------
# FastAPI creates a lightweight, high-performance web server for model inference.
app = FastAPI()

# ----------------------------------------------------------
# 2. Load trained model
# ----------------------------------------------------------
# Instantiate your CNN model architecture.
# Load pretrained weights from checkpoint and set it to evaluation mode.
model = CNNClassifier()
model.load_state_dict(torch.load("models/best_model.pt"))
model.eval()  # Disable dropout, batchnorm updates for inference

# ----------------------------------------------------------
# 3. Define image preprocessing pipeline
# ----------------------------------------------------------
# This transform chain must match the preprocessing used during training.
#   - Resize: scales input to (32x32), matching CIFAR-10 dimensions
#   - ToTensor: converts PIL image $$\rightarrow$$ PyTorch tensor (C,H,W) in [0,1]
#   - Normalize: standardizes each channel using dataset mean & std
transform = transforms.Compose([
    transforms.Resize((32, 32)),
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465),
                         (0.2023, 0.1994, 0.2010))
])

# ----------------------------------------------------------
# 4. Define prediction endpoint
# ----------------------------------------------------------
# This endpoint accepts an uploaded image (as multipart/form-data)
# and returns the model’s predicted class index.
@app.post("/predict")
async def predict(file: UploadFile):
    # Read the raw image bytes from the uploaded file asynchronously.
    image_bytes = await file.read()

    # Open the image from bytes and ensure it's in RGB mode.
    image = Image.open(io.BytesIO(image_bytes)).convert("RGB")

    # Apply preprocessing transform and add a batch dimension (1, C, H, W).
    input_tensor = transform(image).unsqueeze(0)

    # Perform inference in no-grad context (disables autograd for speed & memory).
    with torch.no_grad():
        outputs = model(input_tensor)             # Forward pass
        pred = outputs.argmax(dim=1).item()       # Get class with highest score

    # Return prediction as a JSON response.
    return {"prediction": int(pred)}
uvicorn serve:app --reload

Pipeline Summary

Step Vision (CIFAR-10 CNN) NLP (IMDb Sentiment)
Data Preprocessing Transforms + augmentation Tokenization + padding
Model CNN with dropout LSTM with embeddings
Training CrossEntropy + Adam CrossEntropy + Adam
Evaluation F1, Accuracy, Confusion Matrix Precision, Recall, F1
Deployment FastAPI with TorchScript FastAPI with text input
Monitoring Latency + drift Confidence + feedback loop

Key Takeaways

Example 1: End-to-End NLP Pipeline (IMDb Sentiment Analysis)

Step 1: Data Preparation
from torchtext.datasets import IMDB
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator
from torch.utils.data import DataLoader
from torch.nn.utils.rnn import pad_sequence
import torch

# ----------------------------------------------------------
# 1. Initialize a tokenizer
# ----------------------------------------------------------
#   - The 'basic_english' tokenizer splits text into lowercase words and handles punctuation spacing.
#   - Used for converting raw strings into lists of tokens for further processing.
tokenizer = get_tokenizer("basic_english")


# ----------------------------------------------------------
# 2. Helper function to yield tokenized text from the dataset
# ----------------------------------------------------------
#   - Takes in an iterable of (label, text) pairs.
#   - Tokenizes each text sample and yields the token list.
#   - This generator function is used when building the vocabulary.
def yield_tokens(data_iter):
    for label, text in data_iter:
        yield tokenizer(text)


# ----------------------------------------------------------
# 3. Load IMDB dataset and build vocabulary
# ----------------------------------------------------------
#   - IMDB dataset: 50,000 movie reviews labeled as positive or negative.
#   - split='train' loads only the training portion.
train_iter = IMDB(split='train')

#   - Build vocabulary using tokens from the training data.
#   - 'specials' adds reserved tokens for unknown words and padding.
vocab = build_vocab_from_iterator(yield_tokens(train_iter), specials=["<unk>", "<pad>"])

#   - Set default index for out-of-vocabulary words to the <unk> token index.
vocab.set_default_index(vocab["<unk>"])

#   - Retrieve padding token index for use later in batching.
pad_idx = vocab["<pad>"]


# ----------------------------------------------------------
# 4. Define collate function for DataLoader
# ----------------------------------------------------------
#   - This function prepares each mini-batch before feeding it into the model.
#   - It performs:
#       (a) Label conversion (pos$$\rightarrow$$1, neg$$\rightarrow$$0)
#       (b) Tokenization and numericalization (tokens $$\rightarrow$$ integers via vocab)
#       (c) Padding sequences to the same length within a batch
def collate_batch(batch):
    labels, texts = [], []

    for label, text in batch:
        # Convert text labels to binary (1 = positive, 0 = negative)
        labels.append(1 if label == "pos" else 0)

        # Tokenize text and map tokens to integer IDs
        tokens = vocab(tokenizer(text))
        texts.append(torch.tensor(tokens, dtype=torch.long))

    # Pad all sequences in the batch to the same length with <pad> token index
    padded_texts = pad_sequence(texts, batch_first=True, padding_value=pad_idx)

    # Convert labels list to a tensor
    label_tensor = torch.tensor(labels)

    return padded_texts, label_tensor


# ----------------------------------------------------------
# 5. Example usage (optional)
# ----------------------------------------------------------
#   You can create a DataLoader to batch IMDB samples:
#   from torch.utils.data import DataLoader
#   train_iter = IMDB(split='train')
#   train_loader = DataLoader(list(train_iter), batch_size=8, collate_fn=collate_batch)
#   x_batch, y_batch = next(iter(train_loader))
#   print(x_batch.shape, y_batch)
Step 2: Model Definition
import torch.nn as nn

# ----------------------------------------------------------
# Define a simple LSTM-based sentiment classification model
# ----------------------------------------------------------
class SentimentRNN(nn.Module):
    def __init__(self, vocab_size, embed_dim, hidden_dim, output_dim, pad_idx):
        super().__init__()

        # Embedding layer:
        #   - Converts token indices into dense vector representations
        #   - vocab_size: number of unique tokens in the vocabulary
        #   - embed_dim: dimensionality of each embedding vector
        #   - padding_idx: ensures the <pad> token has zero embedding (not learned)
        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)

        # LSTM layer:
        #   - Processes the embedded token sequence
        #   - hidden_dim: size of the LSTM’s hidden state
        #   - batch_first=True: input/output tensors use (batch, seq, feature) format
        self.lstm = nn.LSTM(embed_dim, hidden_dim, batch_first=True)

        # Fully connected (dense) output layer:
        #   - Maps the final hidden state to output classes (e.g., positive/negative)
        self.fc = nn.Linear(hidden_dim, output_dim)

        # Dropout for regularization:
        #   - Randomly zeroes some elements to prevent overfitting
        self.dropout = nn.Dropout(0.3)

    def forward(self, x):
        # Forward pass:
        # x: (batch_size, seq_len)
        
        # Step 1: Look up embeddings for each token in the batch
        embedded = self.embedding(x)  # (batch_size, seq_len, embed_dim)
        
        # Step 2: Pass the embeddings through the LSTM
        # lstm output shapes:
        #   - output: (batch_size, seq_len, hidden_dim)
        #   - hidden: (num_layers * num_directions, batch_size, hidden_dim)
        #   - cell:   (num_layers * num_directions, batch_size, hidden_dim)
        _, (hidden, _) = self.lstm(embedded)  # hidden: (1, batch_size, hidden_dim)
        
        # Step 3: Apply dropout to the hidden state for regularization
        # hidden.squeeze(0): (batch_size, hidden_dim)
        dropped = self.dropout(hidden.squeeze(0))  # (batch_size, hidden_dim)
        
        # Step 4: Pass through the linear layer to get class logits
        logits = self.fc(dropped)  # (batch_size, output_dim)
        
        return logits  # (batch_size, output_dim)
Step 3: Train and Evaluate
def train_sentiment_model():
    # ----------------------------------------------------------
    # 1. Load and split the IMDb dataset
    # ----------------------------------------------------------
    # IMDB(split='train') loads the IMDb training set of (label, text) pairs.
    # Convert the iterator into a list for indexing/slicing.
    train_iter = IMDB(split='train')
    train_list = list(train_iter)[:4000]   # Use first 4000 samples for training
    val_list = list(train_iter)[4000:5000] # Next 1000 samples for validation

    # Create DataLoaders for batching and shuffling.
    # collate_fn handles tokenization, padding, and tensor conversion per batch.
    train_loader = DataLoader(train_list, batch_size=32, collate_fn=collate_batch, shuffle=True)
    val_loader = DataLoader(val_list, batch_size=32, collate_fn=collate_batch)

    # ----------------------------------------------------------
    # 2. Initialize model, loss function, and optimizer
    # ----------------------------------------------------------
    # SentimentRNN: a simple RNN-based text classifier (embedding + GRU/LSTM + FC)
    model = SentimentRNN(len(vocab), 64, 128, 2, pad_idx)  # vocab size, embed dim, hidden dim, output classes, pad index

    # CrossEntropyLoss: suitable for multi-class classification problems
    criterion = nn.CrossEntropyLoss()

    # Adam optimizer: adaptive learning rate for efficient convergence
    optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

    # Choose device (GPU if available, else CPU)
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.to(device)  # Move model parameters to the chosen device

    # ----------------------------------------------------------
    # 3. Training loop
    # ----------------------------------------------------------
    # Run for a fixed number of epochs
    for epoch in range(3):
        model.train()        # Set model to training mode (activates dropout, etc.)
        total_loss = 0       # Accumulate total training loss per epoch

        # Iterate over mini-batches from the DataLoader
        for x_batch, y_batch in train_loader:
            # Move data to device (GPU/CPU)
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)

            # Reset optimizer gradients
            optimizer.zero_grad()

            # Forward pass: compute predictions
            outputs = model(x_batch)

            # Compute loss between predictions and ground truth
            loss = criterion(outputs, y_batch)

            # Backward pass: compute gradients
            loss.backward()

            # Update model parameters based on gradients
            optimizer.step()

            # Accumulate loss for reporting
            total_loss += loss.item()

        # ----------------------------------------------------------
        # 4. Log progress per epoch
        # ----------------------------------------------------------
        print(f"Epoch {epoch+1}, Loss={total_loss/len(train_loader):.3f}")

    # ----------------------------------------------------------
    # 5. Save the trained model
    # ----------------------------------------------------------
    # Save model weights for later evaluation or inference
    torch.save(model.state_dict(), "sentiment_model.pt")
    print("✅ Training complete. Model saved to 'sentiment_model.pt'.")
Step 4: Deploy as Text API
from fastapi import FastAPI
import torch
import numpy as np

# ----------------------------------------------------------
# 1. Initialize FastAPI application
# ----------------------------------------------------------
# FastAPI is a lightweight web framework for serving ML models via REST APIs.
app = FastAPI()

# ----------------------------------------------------------
# 2. Load pre-trained PyTorch model
# ----------------------------------------------------------
# Instantiate model using same architecture and parameters as during training.
# 'SentimentRNN' should be defined elsewhere (same structure as training phase).
model = SentimentRNN(len(vocab), 64, 128, 2, pad_idx)

# Load saved model weights from checkpoint file
model.load_state_dict(torch.load("sentiment_model.pt"))

# Switch model to evaluation mode:
# disables dropout, batchnorm updates, and gradient tracking
model.eval()

# Move model to appropriate device (CPU or GPU)
model.to(device)

# ----------------------------------------------------------
# 3. Define REST API endpoint for sentiment analysis
# ----------------------------------------------------------
# The endpoint accepts a POST request at `/analyze` with a text input.
# Example request: POST /analyze { "text": "this movie was great" }
@app.post("/analyze")
async def analyze_sentiment(text: str):
    # ------------------------------------------------------
    # (a) Text tokenization and numericalization
    # ------------------------------------------------------
    # Convert input string $$\rightarrow$$ tokens $$\rightarrow$$ numerical indices using vocabulary.
    # `tokenizer(text)` splits text into tokens.
    # `vocab(tokenizer(text))` maps tokens to integer IDs.
    tokens = torch.tensor(vocab(tokenizer(text)), dtype=torch.long).unsqueeze(0).to(device)
    # unsqueeze(0) adds a batch dimension $$\rightarrow$$ shape becomes (1, seq_len)

    # ------------------------------------------------------
    # (b) Forward pass through model (inference mode)
    # ------------------------------------------------------
    with torch.no_grad():  # Disable gradient computation for efficiency
        outputs = model(tokens)                 # Raw model logits
        probs = torch.softmax(outputs, dim=1).cpu().numpy()[0]  # Convert to probabilities

        # Predicted class index (0 = negative, 1 = positive)
        pred = int(np.argmax(probs))

    # ------------------------------------------------------
    # (c) Format prediction and return JSON response
    # ------------------------------------------------------
    # Map predicted index to human-readable label
    label = "positive" if pred == 1 else "negative"

    # Return sentiment label and confidence score as JSON
    return {"sentiment": label, "confidence": float(np.max(probs))}

End-to-End Orchestration with Prefect or Airflow

Comparison: Prefect vs. Airflow

Feature Prefect Airflow
Syntax Pure Python (imperative) DAG-based (declarative)
Execution Dynamic and reactive Static, schedule-driven
Best For Research, prototyping, flexible ML flows Enterprise pipelines, ETL, recurring jobs
Monitoring Built-in UI or Prefect Cloud Airflow UI + logs
Retraining Trigger Native conditional branching Requires Sensors or ExternalTaskTrigger
Setup Complexity Simple Heavier setup (scheduler, DB, UI)

Orchestrating a Vision (CIFAR-10 Image Classification) Pipeline with Prefect or Airflow

Prefect
Step 1: Install Dependencies
pip install prefect torch torchvision scikit-learn
Step 2: Define Pipeline Script (vision_pipeline.py)
from prefect import task, Flow
import torch
import torch.nn as nn
import torch.optim as optim
from torchvision import datasets, transforms
from torch.utils.data import DataLoader, random_split
from sklearn.metrics import accuracy_score
import numpy as np
import os
import time

# ----------------------------------------------------------
# TASK 1: DATA PREPARATION
# ----------------------------------------------------------
@task
def prepare_vision_data(batch_size=64):
    # Define data augmentation and normalization for training images
    transform = transforms.Compose([
        transforms.RandomHorizontalFlip(),   # Randomly flip images horizontally
        transforms.RandomCrop(32, padding=4),  # Randomly crop image with padding for augmentation
        transforms.ToTensor(),               # Convert PIL Image $$\rightarrow$$ PyTorch tensor (0–1 range)
        transforms.Normalize((0.4914, 0.4822, 0.4465),  # CIFAR-10 mean (per channel)
                             (0.2023, 0.1994, 0.2010))  # CIFAR-10 std (per channel)
    ])

    # Download and load CIFAR-10 dataset with defined transformations
    dataset = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform)

    # Split into training and validation sets (45k train / 5k validation)
    train_set, val_set = random_split(dataset, [45000, 5000])

    # Create DataLoaders for batching, shuffling, and parallel loading
    train_loader = DataLoader(train_set, batch_size=batch_size, shuffle=True, num_workers=2)
    val_loader = DataLoader(val_set, batch_size=batch_size, num_workers=2)

    print("✅ Data preparation complete.")
    return train_loader, val_loader


# ----------------------------------------------------------
# MODEL DEFINITION
# ----------------------------------------------------------
class CNNClassifier(nn.Module):
    def __init__(self, dropout=0.3):
        super().__init__()

        # Convolutional feature extractor (two conv + pooling blocks)
        self.conv_block = nn.Sequential(
            nn.Conv2d(3, 32, 3, padding=1), nn.ReLU(),   # Conv layer 1
            nn.MaxPool2d(2),                            # Downsample by 2
            nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), # Conv layer 2
            nn.MaxPool2d(2)                             # Downsample by 2
        )

        # Fully connected classification head
        self.fc_block = nn.Sequential(
            nn.Flatten(),                # Flatten feature maps into a vector
            nn.Linear(64 * 8 * 8, 128),  # Dense layer
            nn.ReLU(),
            nn.Dropout(dropout),         # Regularization to prevent overfitting
            nn.Linear(128, 10)           # Output layer (10 classes for CIFAR-10)
        )

    def forward(self, x):
        # Define forward pass through convolutional and fully connected blocks
        return self.fc_block(self.conv_block(x))


# ----------------------------------------------------------
# TASK 2: MODEL TRAINING
# ----------------------------------------------------------
@task
def train_cnn_model(train_loader, val_loader, lr=1e-3, epochs=5):
    # Use GPU if available for faster training
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model = CNNClassifier().to(device)
    optimizer = optim.Adam(model.parameters(), lr=lr)
    criterion = nn.CrossEntropyLoss()  # Suitable for multi-class classification
    best_val_acc = 0.0  # Track best validation accuracy for checkpointing

    # Main training loop
    for epoch in range(epochs):
        model.train()  # Enable training mode (activates dropout, batchnorm updates)

        # Iterate over all batches in training data
        for x_batch, y_batch in train_loader:
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)
            optimizer.zero_grad()       # Reset accumulated gradients
            outputs = model(x_batch)    # Forward pass
            loss = criterion(outputs, y_batch)  # Compute loss
            loss.backward()             # Backpropagate loss to compute gradients
            optimizer.step()            # Update model weights

        # Evaluate model on validation data after each epoch
        val_acc = evaluate_cnn_model(model, val_loader)
        print(f"Epoch {epoch+1}, Validation Accuracy: {val_acc:.3f}")

        # Save checkpoint if validation performance improves
        if val_acc > best_val_acc:
            best_val_acc = val_acc
            torch.save(model.state_dict(), "best_cnn_model.pt")
            print("✅ Saved improved model checkpoint.")

    return best_val_acc


# ----------------------------------------------------------
# TASK 3: MODEL EVALUATION
# ----------------------------------------------------------
@task
def evaluate_cnn_model(model, loader):
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.eval()  # Set model to evaluation mode (disable dropout, batchnorm updates)

    preds, labels = [], []
    with torch.no_grad():  # Disable autograd for inference efficiency
        for x_batch, y_batch in loader:
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)
            outputs = model(x_batch)  # Forward pass
            preds.extend(outputs.argmax(dim=1).cpu().numpy())  # Get predicted classes
            labels.extend(y_batch.cpu().numpy())               # Store true labels

    # Compute overall accuracy using sklearn
    return accuracy_score(labels, preds)


# ----------------------------------------------------------
# TASK 4: MODEL DEPLOYMENT
# ----------------------------------------------------------
@task
def deploy_cnn_model():
    # Load best model weights from training
    model = CNNClassifier()
    model.load_state_dict(torch.load("best_cnn_model.pt"))
    model.eval()

    # Export to TorchScript for production deployment
    torch.jit.save(torch.jit.script(model), "deployed_cnn_model.pt")
    print("✅ CNN model deployed successfully.")
    return "deployed_cnn_model.pt"


# ----------------------------------------------------------
# TASK 5: MODEL MONITORING
# ----------------------------------------------------------
@task
def monitor_cnn_model():
    # Simulate monitoring: model’s live accuracy fluctuates around 0.9
    metrics = np.random.normal(loc=0.9, scale=0.05, size=10)
    avg_acc = np.mean(metrics)
    print(f"Average live accuracy: {avg_acc:.3f}")

    # Trigger retraining if performance drops below threshold
    if avg_acc < 0.85:
        print("⚠️ Drift detected! Retraining required.")
        return True
    return False


# ----------------------------------------------------------
# PREFECT FLOW DEFINITION
# ----------------------------------------------------------
with Flow("Vision-CNN-Pipeline") as vision_flow:
    # Step 1: Data loading and preprocessing
    train_loader, val_loader = prepare_vision_data()

    # Step 2: Model training
    acc = train_cnn_model(train_loader, val_loader)

    # Step 3: Deployment
    deploy = deploy_cnn_model()

    # Step 4: Monitoring and drift detection
    drift_flag = monitor_cnn_model()

    # Step 5: Conditional retraining logic (retrain if drift detected)
    if drift_flag:
        train_cnn_model(train_loader, val_loader)
Step 3: Run the Flow
prefect run -p vision_pipeline.py
Step 4: Schedule and Automate
prefect deployment build vision_pipeline.py:vision_flow -n "CIFAR10_Retrain"
prefect deployment apply vision_pipeline-deployment.yaml
prefect agent start
Step 5: Optional Integrations
@task
def send_alert(message):
    print(f"📢 ALERT: {message}")
if drift_flag:
    send_alert("Retraining triggered for CIFAR-10 CNN model.")
Summary of the Vision Prefect Pipeline
Stage Task Purpose
Data Preparation Loading, augmentation, split Prepares training/validation data
Model Training CNN training loop Produces best checkpoint
Evaluation Validation accuracy computation Tracks performance improvement
Deployment TorchScript export Enables serving and portability
Monitoring Accuracy drift simulation Auto-triggers retraining
Key Takeaways
Airflow
Airflow Setup
pip install apache-airflow
airflow db init
airflow users create --username admin --firstname admin --lastname user --role Admin --email admin@example.com
airflow webserver --port 8080
airflow scheduler
Define the Airflow DAG
from airflow import DAG
from airflow.operators.python import PythonOperator
from datetime import datetime, timedelta
import torch
import torch.nn as nn
import torch.optim as optim
from torchvision import datasets, transforms
from torch.utils.data import DataLoader, random_split
import numpy as np

# ----------------------------------------------------------
# 1. Default DAG arguments
# ----------------------------------------------------------
# These settings define task-level behavior and retry policies.
default_args = {
    "owner": "airflow",                 # DAG owner
    "depends_on_past": False,           # Run each task independently of previous runs
    "email": ["alerts@example.com"],    # Alert recipient
    "email_on_failure": True,           # Send alert if any task fails
    "retries": 1,                       # Retry once if a task fails
    "retry_delay": timedelta(minutes=5) # Wait 5 minutes before retrying
}

# ----------------------------------------------------------
# 2. DAG definition
# ----------------------------------------------------------
# DAG = Directed Acyclic Graph — defines task workflow structure
dag = DAG(
    "vision_cnn_dag",
    default_args=default_args,
    description="Vision (CIFAR-10) Training and Deployment Pipeline",
    schedule_interval="@daily",         # Run every day
    start_date=datetime(2025, 1, 1),    # DAG starts from this date
    catchup=False,                      # Do not backfill missed runs
)

# ----------------------------------------------------------
# 3. CNN Model Definition
# ----------------------------------------------------------
# Simple CNN model for CIFAR-10 with two conv layers + FC classifier
class CNNClassifier(nn.Module):
    def __init__(self, dropout=0.3):
        super().__init__()
        # Convolutional feature extraction block
        self.conv_block = nn.Sequential(
            nn.Conv2d(3, 32, 3, padding=1), nn.ReLU(),   # Conv layer 1
            nn.MaxPool2d(2),                             # Downsample to 16x16
            nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(),  # Conv layer 2
            nn.MaxPool2d(2)                              # Downsample to 8x8
        )
        # Fully connected classification head
        self.fc_block = nn.Sequential(
            nn.Flatten(),                                # Flatten 64×8×8 feature map
            nn.Linear(64 * 8 * 8, 128), nn.ReLU(),       # Hidden dense layer
            nn.Dropout(dropout),                         # Regularization
            nn.Linear(128, 10)                           # 10 output classes
        )

    def forward(self, x):
        # Forward pass through conv and FC blocks
        return self.fc_block(self.conv_block(x))

# ----------------------------------------------------------
# 4. Data Preparation Function
# ----------------------------------------------------------
def prepare_data():
    # Define data transformations with augmentation
    transform = transforms.Compose([
        transforms.RandomHorizontalFlip(),                      # Random flip for diversity
        transforms.ToTensor(),                                  # Convert PIL $$\rightarrow$$ Tensor
        transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))  # Normalize per channel
    ])

    # Load CIFAR-10 dataset
    dataset = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform)

    # Split into training and validation subsets (45K / 5K)
    train_set, val_set = random_split(dataset, [45000, 5000])

    # Save dataset metadata for traceability
    torch.save({"train_set": len(train_set), "val_set": len(val_set)}, "data_info.pt")
    print("✅ Data prepared and saved.")

# ----------------------------------------------------------
# 5. Model Training Function
# ----------------------------------------------------------
def train_vision_model():
    # Select GPU if available
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

    # Initialize model, optimizer, and loss function
    model = CNNClassifier().to(device)
    optimizer = optim.Adam(model.parameters(), lr=1e-3)
    criterion = nn.CrossEntropyLoss()

    # Define preprocessing for training
    transform = transforms.Compose([
        transforms.ToTensor(),
        transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))
    ])

    # Load dataset and create DataLoader
    dataset = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform)
    train_loader = DataLoader(dataset, batch_size=64, shuffle=True)

    # Single-epoch demo training loop
    for epoch in range(1):
        model.train()
        for images, labels in train_loader:
            images, labels = images.to(device), labels.to(device)
            optimizer.zero_grad()               # Reset gradients
            outputs = model(images)             # Forward pass
            loss = criterion(outputs, labels)   # Compute loss
            loss.backward()                     # Backpropagate gradients
            optimizer.step()                    # Update weights

    # Save trained weights for deployment
    torch.save(model.state_dict(), "vision_model.pt")
    print("✅ CNN model trained and saved.")

# ----------------------------------------------------------
# 6. Model Deployment Function
# ----------------------------------------------------------
def deploy_model():
    # Load trained weights into model
    model = CNNClassifier()
    model.load_state_dict(torch.load("vision_model.pt"))

    # Convert to TorchScript for optimized deployment
    torch.jit.save(torch.jit.script(model), "deployed_vision_model.pt")
    print("✅ CNN model deployed as TorchScript.")

# ----------------------------------------------------------
# 7. Model Monitoring Function
# ----------------------------------------------------------
def monitor_model():
    # Simulate monitoring process with random validation accuracy
    metrics = np.random.normal(loc=0.9, scale=0.05, size=20)
    avg = np.mean(metrics)

    print(f"Average validation accuracy: {avg:.3f}")
    if avg < 0.85:
        print("⚠️ Drift detected, retraining needed.")
    else:
        print("✅ Model stable.")
    return avg

# ----------------------------------------------------------
# 8. Airflow Operators (Tasks)
# ----------------------------------------------------------
# Each Python function is wrapped in a PythonOperator, which Airflow executes as a DAG node.
prepare_task = PythonOperator(
    task_id="prepare_data", 
    python_callable=prepare_data, 
    dag=dag
)

train_task = PythonOperator(
    task_id="train_model", 
    python_callable=train_vision_model, 
    dag=dag
)

deploy_task = PythonOperator(
    task_id="deploy_model", 
    python_callable=deploy_model, 
    dag=dag
)

monitor_task = PythonOperator(
    task_id="monitor_model", 
    python_callable=monitor_model, 
    dag=dag
)

# ----------------------------------------------------------
# 9. DAG Dependencies
# ----------------------------------------------------------
# Defines execution order:
#   prepare_data $$\rightarrow$$ train_model $$\rightarrow$$ deploy_model $$\rightarrow$$ monitor_model
prepare_task >> train_task >> deploy_task >> monitor_task
How It Works
  1. prepare_data loads and preprocesses CIFAR-10 images.
  2. train_model trains a CNN and saves the best weights.
  3. deploy_model exports a TorchScript model for serving.
  4. monitor_model checks accuracy drift and logs performance daily.

Orchestrating an NLP (Sentiment Prediction) Pipeline with Prefect or Airflow

Prefect
Step 1: Install Dependencies
pip install prefect torch torchtext scikit-learn
Step 2: Define Pipeline Script (nlp_pipeline.py)
from prefect import task, Flow
import torch
import torch.nn as nn
import torch.optim as optim
from torchtext.datasets import IMDB
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator
from torch.utils.data import DataLoader
from torch.nn.utils.rnn import pad_sequence
from sklearn.metrics import accuracy_score
import numpy as np
import os
import time

# ----------------------------------------------------------
# TASK 1: DATA PREPARATION
# ----------------------------------------------------------
@task
def prepare_nlp_data(batch_size=32):
    # Tokenizer converts raw text into lists of tokens
    tokenizer = get_tokenizer("basic_english")

    # Helper function to yield tokens for building vocabulary
    def yield_tokens(data_iter):
        for label, text in data_iter:
            yield tokenizer(text)

    # Build a vocabulary from the training data
    train_iter = IMDB(split="train")
    vocab = build_vocab_from_iterator(yield_tokens(train_iter), specials=["<unk>", "<pad>"])
    vocab.set_default_index(vocab["<unk>"])  # Handle unseen tokens as <unk>
    pad_idx = vocab["<pad>"]  # Padding index for sequence alignment

    # Function to convert raw text + labels $$\rightarrow$$ padded tensors
    def collate_batch(batch):
        labels, texts = [], []
        for label, text in batch:
            labels.append(1 if label == "pos" else 0)  # Encode labels (pos$$\rightarrow$$1, neg$$\rightarrow$$0)
            tokens = vocab(tokenizer(text))  # Tokenize and numericalize
            texts.append(torch.tensor(tokens, dtype=torch.long))
        # Pad variable-length sequences to equal length for batching
        return pad_sequence(texts, batch_first=True, padding_value=pad_idx), torch.tensor(labels)

    # Reload train/test sets since the iterator is exhausted after vocab building
    train_iter, test_iter = IMDB(split=("train", "test"))

    # Subset the data for demonstration (faster training)
    train_list = list(train_iter)[:4000]
    val_list = list(train_iter)[4000:5000]

    # Create DataLoaders for batching and shuffling
    train_loader = DataLoader(train_list, batch_size=batch_size, collate_fn=collate_batch, shuffle=True)
    val_loader = DataLoader(val_list, batch_size=batch_size, collate_fn=collate_batch)
    test_loader = DataLoader(list(test_iter)[:1000], batch_size=batch_size, collate_fn=collate_batch)

    print("Data preparation complete.")
    return vocab, pad_idx, train_loader, val_loader, test_loader


# ----------------------------------------------------------
# MODEL DEFINITION
# ----------------------------------------------------------
class SentimentRNN(nn.Module):
    def __init__(self, vocab_size, embed_dim, hidden_dim, output_dim, pad_idx):
        super().__init__()
        # Embedding layer converts token IDs into dense vectors
        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)
        # LSTM captures sequential dependencies in text
        self.lstm = nn.LSTM(embed_dim, hidden_dim, batch_first=True)
        # Fully connected layer maps hidden state $$\rightarrow$$ class logits
        self.fc = nn.Linear(hidden_dim, output_dim)
        # Dropout regularizes the model to prevent overfitting
        self.dropout = nn.Dropout(0.3)

    def forward(self, x):
        embedded = self.embedding(x)         # (batch, seq_len, embed_dim)
        _, (hidden, _) = self.lstm(embedded) # Get final hidden state from LSTM
        return self.fc(self.dropout(hidden.squeeze(0)))  # Output class scores


# ----------------------------------------------------------
# TASK 2: MODEL TRAINING
# ----------------------------------------------------------
@task
def train_nlp_model(vocab, pad_idx, train_loader, val_loader, lr=1e-3, epochs=3):
    # Select GPU if available
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

    # Initialize model, optimizer, and loss
    model = SentimentRNN(len(vocab), 64, 128, 2, pad_idx).to(device)
    optimizer = optim.Adam(model.parameters(), lr=lr)
    criterion = nn.CrossEntropyLoss()
    best_val_acc = 0.0

    # Main training loop
    for epoch in range(epochs):
        model.train()  # Enable dropout, gradient tracking
        for x_batch, y_batch in train_loader:
            # Move data to device
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)

            optimizer.zero_grad()  # Reset gradients
            outputs = model(x_batch)  # Forward pass
            loss = criterion(outputs, y_batch)  # Compute loss
            loss.backward()  # Backpropagation
            optimizer.step()  # Update weights

        # Evaluate after each epoch
        val_acc = evaluate_nlp_model(model, val_loader)
        print(f"Epoch {epoch+1}, Val Acc: {val_acc:.3f}")

        # Save best-performing model
        if val_acc > best_val_acc:
            best_val_acc = val_acc
            torch.save(model.state_dict(), "best_nlp_model.pt")

    return best_val_acc


# ----------------------------------------------------------
# TASK 3: EVALUATION
# ----------------------------------------------------------
@task
def evaluate_nlp_model(model, loader):
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    model.eval()  # Disable dropout and batch norm updates
    preds, labels = [], []

    # Disable gradient tracking for faster inference
    with torch.no_grad():
        for x_batch, y_batch in loader:
            x_batch, y_batch = x_batch.to(device), y_batch.to(device)
            outputs = model(x_batch)
            preds.extend(outputs.argmax(dim=1).cpu().numpy())  # Get predicted class
            labels.extend(y_batch.cpu().numpy())               # True labels

    # Compute classification accuracy
    return accuracy_score(labels, preds)


# ----------------------------------------------------------
# TASK 4: MODEL DEPLOYMENT
# ----------------------------------------------------------
@task
def deploy_nlp_model():
    # Recreate model structure (with placeholder vocab size for demonstration)
    model = SentimentRNN(50000, 64, 128, 2, 1)
    model.load_state_dict(torch.load("best_nlp_model.pt"))  # Load trained weights
    model.eval()

    # Export model to TorchScript for production deployment
    torch.jit.save(torch.jit.script(model), "deployed_nlp_model.pt")
    print("✅ Sentiment model deployed successfully.")
    return "deployed_nlp_model.pt"


# ----------------------------------------------------------
# TASK 5: MONITORING
# ----------------------------------------------------------
@task
def monitor_nlp_model():
    # Simulate model performance drift over time using random confidence scores
    conf_history = np.random.normal(loc=0.85, scale=0.05, size=100)
    avg_conf = np.mean(conf_history)
    print(f"Average confidence: {avg_conf:.3f}")

    # If average confidence drops below threshold $$\rightarrow$$ retraining trigger
    if avg_conf < 0.8:
        print("⚠️ Confidence drift detected! Triggering retraining.")
        return True
    return False


# ----------------------------------------------------------
# FLOW DEFINITION (PIPELINE)
# ----------------------------------------------------------
# Prefect flow orchestrates all tasks end-to-end
with Flow("NLP-Sentiment-Pipeline") as nlp_flow:
    # 1. Prepare data
    vocab, pad_idx, train_loader, val_loader, test_loader = prepare_nlp_data()

    # 2. Train model
    acc = train_nlp_model(vocab, pad_idx, train_loader, val_loader)

    # 3. Deploy trained model
    deploy = deploy_nlp_model()

    # 4. Monitor performance drift
    drift_flag = monitor_nlp_model()

    # 5. Conditional retraining when drift detected
    if drift_flag:
        train_nlp_model(vocab, pad_idx, train_loader, val_loader)
Step 3: Run the Flow
prefect run -p nlp_pipeline.py
Step 4: Schedule and Automate
prefect deployment build nlp_pipeline.py:nlp_flow -n "IMDB_Sentiment_Retrain"
prefect deployment apply nlp_pipeline-deployment.yaml
prefect agent start
Step 5: Optional Integrations
@task
def send_alert(message):
    print(f"📢 ALERT: {message}")
if drift_flag:
    send_alert("Retraining triggered for NLP sentiment model.")
Summary of the NLP Prefect Pipeline
Stage Task Purpose
Data Preparation Tokenization, vocab building, padding Converts text to tensors
Model Training LSTM training loop Produces best checkpoint
Evaluation Validation accuracy computation Monitors overfitting
Deployment TorchScript model export Enables portable serving
Monitoring Confidence drift simulation Auto-triggers retraining
Key Takeaways
Airflow
Airflow Setup
pip install apache-airflow
airflow db init
airflow users create --username admin --firstname admin --lastname user --role Admin --email admin@example.com
airflow webserver --port 8080
airflow scheduler
Define the Airflow DAG
from airflow import DAG
from airflow.operators.python import PythonOperator
from datetime import datetime, timedelta
import torch
import torch.nn as nn
import torch.optim as optim
from torchtext.datasets import IMDB
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator
from torch.utils.data import DataLoader
from torch.nn.utils.rnn import pad_sequence
import numpy as np

# ----------------------------------------------------------
# 1. DEFAULT DAG ARGUMENTS
# ----------------------------------------------------------
# These parameters define DAG-wide behavior:
#   - owner: identifies the DAG owner
#   - retries: number of retry attempts upon failure
#   - retry_delay: wait time between retries
#   - email_on_failure: send alerts if any task fails
default_args = {
    "owner": "airflow",
    "depends_on_past": False,
    "email": ["alerts@example.com"],
    "email_on_failure": True,
    "email_on_retry": False,
    "retries": 1,
    "retry_delay": timedelta(minutes=5),
}

# ----------------------------------------------------------
# 2. DAG DEFINITION
# ----------------------------------------------------------
# The DAG (Directed Acyclic Graph) defines task structure and scheduling.
#   - schedule_interval="@daily": runs every day
#   - start_date: earliest start date for DAG runs
#   - catchup=False: skip retroactive runs for missed dates
dag = DAG(
    "nlp_sentiment_dag",
    default_args=default_args,
    description="NLP Sentiment Training and Deployment Pipeline",
    schedule_interval="@daily",
    start_date=datetime(2025, 1, 1),
    catchup=False,
)

# ----------------------------------------------------------
# 3. MODEL DEFINITION – SENTIMENT CLASSIFIER (LSTM)
# ----------------------------------------------------------
# A lightweight RNN-based text classifier for sentiment analysis.
# It includes:
#   - Embedding layer: converts word indices into dense vectors
#   - LSTM: captures sequential dependencies in text
#   - Dropout: regularization to reduce overfitting
#   - Linear layer: outputs logits for binary classification
class SentimentRNN(nn.Module):
    def __init__(self, vocab_size, embed_dim, hidden_dim, output_dim, pad_idx):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, embed_dim, padding_idx=pad_idx)
        self.lstm = nn.LSTM(embed_dim, hidden_dim, batch_first=True)
        self.fc = nn.Linear(hidden_dim, output_dim)
        self.dropout = nn.Dropout(0.3)

    def forward(self, x):
        embedded = self.embedding(x)
        _, (hidden, _) = self.lstm(embedded)
        return self.fc(self.dropout(hidden.squeeze(0)))

# ----------------------------------------------------------
# 4. TASK FUNCTION: DATA PREPARATION
# ----------------------------------------------------------
# - Loads IMDB dataset
# - Builds vocabulary and saves it for later use
# - Tokenizes text using a basic English tokenizer
def prepare_data():
    tokenizer = get_tokenizer("basic_english")

    def yield_tokens(data_iter):
        for label, text in data_iter:
            yield tokenizer(text)

    # Build vocabulary from training data
    train_iter = IMDB(split="train")
    vocab = build_vocab_from_iterator(yield_tokens(train_iter), specials=["<unk>", "<pad>"])
    vocab.set_default_index(vocab["<unk>"])
    pad_idx = vocab["<pad>"]

    # Save vocab and pad index for reuse in other tasks
    torch.save({"vocab": vocab, "pad_idx": pad_idx}, "vocab_info.pt")
    print("✅ Vocabulary prepared and saved.")

# ----------------------------------------------------------
# 5. TASK FUNCTION: MODEL TRAINING
# ----------------------------------------------------------
# - Loads vocabulary and builds dataloader
# - Defines and trains LSTM-based classifier
# - Saves trained model to disk
def train_model():
    # Load previously saved vocab
    info = torch.load("vocab_info.pt")
    vocab, pad_idx = info["vocab"], info["pad_idx"]

    # Custom batch collation (handles tokenization and padding)
    def collate_batch(batch):
        labels, texts = [], []
        tokenizer = get_tokenizer("basic_english")
        for label, text in batch:
            labels.append(1 if label == "pos" else 0)
            tokens = vocab(tokenizer(text))
            texts.append(torch.tensor(tokens, dtype=torch.long))
        # Pad sequences to equal length
        return pad_sequence(texts, batch_first=True, padding_value=pad_idx), torch.tensor(labels)

    # Prepare data loader with a subset of the IMDB dataset
    train_iter = IMDB(split="train")
    train_list = list(train_iter)[:2000]  # sample small subset for demonstration
    loader = DataLoader(train_list, batch_size=32, collate_fn=collate_batch, shuffle=True)

    # Initialize model, optimizer, and loss function
    model = SentimentRNN(len(vocab), 64, 128, 2, pad_idx)
    optimizer = optim.Adam(model.parameters(), lr=1e-3)
    criterion = nn.CrossEntropyLoss()

    # Training loop for 2 epochs
    for epoch in range(2):
        model.train()
        for x_batch, y_batch in loader:
            optimizer.zero_grad()       # Reset gradients
            outputs = model(x_batch)    # Forward pass
            loss = criterion(outputs, y_batch)
            loss.backward()             # Backpropagation
            optimizer.step()            # Update weights

    # Save the trained model parameters
    torch.save(model.state_dict(), "nlp_model.pt")
    print("✅ Model trained and saved.")

# ----------------------------------------------------------
# 6. TASK FUNCTION: MODEL DEPLOYMENT
# ----------------------------------------------------------
# - Loads trained model
# - Converts it to TorchScript for deployment
def deploy_model():
    # Initialize dummy model with expected structure
    model = SentimentRNN(50000, 64, 128, 2, 1)
    # Load trained weights
    model.load_state_dict(torch.load("nlp_model.pt"))
    # Convert model to TorchScript (for optimized serving)
    torch.jit.save(torch.jit.script(model), "deployed_nlp_model.pt")
    print("✅ Model deployed as TorchScript.")

# ----------------------------------------------------------
# 7. TASK FUNCTION: MONITORING
# ----------------------------------------------------------
# - Simulates confidence drift detection
# - Checks model stability over time
def monitor_model():
    # Simulated confidence scores for predictions
    conf = np.random.normal(loc=0.85, scale=0.05, size=100)
    avg_conf = np.mean(conf)
    print(f"Average confidence: {avg_conf:.3f}")

    # If model confidence drops below threshold, retraining is recommended
    if avg_conf < 0.8:
        print("⚠️ Confidence drift detected! Retraining required.")
    else:
        print("✅ Model stable.")
    return avg_conf

# ----------------------------------------------------------
# 8. DEFINE AIRFLOW TASKS (OPERATORS)
# ----------------------------------------------------------
# Each function above becomes a PythonOperator task.
# Airflow runs them as isolated, trackable units in the DAG.
prepare_task = PythonOperator(task_id="prepare_data", python_callable=prepare_data, dag=dag)
train_task = PythonOperator(task_id="train_model", python_callable=train_model, dag=dag)
deploy_task = PythonOperator(task_id="deploy_model", python_callable=deploy_model, dag=dag)
monitor_task = PythonOperator(task_id="monitor_model", python_callable=monitor_model, dag=dag)

# ----------------------------------------------------------
# 9. SET TASK DEPENDENCIES (EXECUTION ORDER)
# ----------------------------------------------------------
# DAG flow:
#   1. prepare_data $$\rightarrow$$ 2. train_model $$\rightarrow$$ 3. deploy_model $$\rightarrow$$ 4. monitor_model
# The ">>" operator defines directional dependencies between tasks.
prepare_task >> train_task >> deploy_task >> monitor_task
How It Works
  1. prepare_data builds the vocabulary and stores it.
  2. train_model trains a small LSTM and saves weights.
  3. deploy_model converts and saves the TorchScript model for serving.
  4. monitor_model runs daily drift checks and prints metrics.

References

Model Training and Evaluation

Hyperparameter Tuning

Model Evaluation and Benchmarking

Model Deployment

Monitoring and Continuous Evaluation

Continuous Integration and Orchestration

Advanced Topics

Citation

If you found our work useful, please cite it as:

@article{Chadha2020PyTorchPrimer,
  title   = {PyTorch Primer},
  author  = {Chadha, Aman},
  journal = {Distilled AI},
  year    = {2020},
  note    = {\url{https://aman.ai}}
}