开始使用免费开始使用

Gradient accumulation with Accelerator

You're training a language model to simplify translations by paraphrasing complex sentences, but your GPU is running out of memory. Gradient accumulation allows the model to effectively train on larger batches by using small batches that fit into memory. You prefer to write the training loop explicitly to see its structure, so you're using Accelerator. Note that this exercise actually runs on the CPU, but the code remains the same for the GPU.

The model, train_dataloader, optimizer, and lr_scheduler have been pre-defined.

本练习是课程的一部分

Efficient AI Model Training with PyTorch

查看课程

练习说明

  • Configure Accelerator() to use gradient accumulation with two steps.
  • Set up an Accelerator context manager to enable gradient accumulation for the model.

交互式实操练习

通过完成这段示例代码来试试这个练习。

# Configure Accelerator
accelerator = ____(____=____)
model, optimizer, train_dataloader, lr_scheduler = accelerator.prepare(model, optimizer, train_dataloader, lr_scheduler)
for batch in train_dataloader:
    # Set up an Accelerator context manager
    with ____.____(____):
        inputs, targets = batch["input_ids"], batch["labels"]
        outputs = model(inputs, labels=targets)
        loss = outputs.loss
        accelerator.backward(loss)
        optimizer.step()
        lr_scheduler.step()
        optimizer.zero_grad()
        print(f"Loss = {loss}")
编辑并运行代码