Mixed precision training with basic PyTorch
You will use low precision floating point data types to speed up training for your language translation model. For example, 16-bit floating point data types (float16) are only half the size of their 32-bit counterparts (float32). This accelerates computations of matrix multiplications and convolutions. Recall that this involves scaling gradients and casting operations to 16 bit floating point.
Some objects have been preloaded: dataset, model, dataloader, and optimizer.
本练习是课程的一部分
Efficient AI Model Training with PyTorch
练习说明
- Before the loop, define a scaler for the gradients using
torch.amp.GradScaler. - In the loop, cast operations to the 16-bit floating point data type using
torch.autocastas a context manager. - In the loop, scale the loss and call
.backward()to create scaled gradients.
交互式实操练习
通过完成这段示例代码来试试这个练习。
# Define a scaler for the gradients
scaler = torch.amp.____()
for batch in train_dataloader:
inputs, targets = batch["input_ids"], batch["labels"]
# Casts operations to mixed precision
with torch.____(device_type="cpu", dtype=torch.____):
outputs = model(inputs, labels=targets)
loss = outputs.loss
# Compute scaled gradients
scaler.____(loss).backward()
scaler.step(optimizer)
scaler.update()
optimizer.zero_grad()