Fine-tune bằng PPO
Sau khi đã khởi tạo trainer, giờ bạn cần khởi tạo vòng lặp để fine-tune mô hình.
Reward trainer ppo_trainer đã được khởi tạo bằng lớp PPOTrainer từ thư viện Python trl.
Bài tập này là một phần của khóa học
Reinforcement Learning from Human Feedback (RLHF)
Hướng dẫn bài tập
- Tạo các tensor phản hồi bằng cách dùng input ids và trainer trong vòng lặp PPO.
- Hoàn thiện bước trong vòng lặp PPO sử dụng dữ liệu truy vấn, phản hồi và phần thưởng để tối ưu mô hình PPO.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
for batch in tqdm(ppo_trainer.dataloader):
# Generate responses for the given queries using the trainer
response_tensors = ____(batch["input_ids"])
batch["response"] = [tokenizer.decode(r.squeeze()) for r in response_tensors]
texts = [q + r for q, r in zip(batch["query"], batch["response"])]
rewards = reward_model(texts)
# Training PPO step with the query, responses ids, and rewards
stats = ____(batch["input_ids"], response_tensors, rewards)
ppo_trainer.log_stats(stats, batch, rewards)