Bắt đầu ngayBắt đầu miễn phí

Lấy mẫu từ bộ đệm PER

Trước khi bạn có thể dùng lớp Prioritized Experience Buffer để huấn luyện agent, bạn vẫn cần triển khai phương thức .sample(). Phương thức này nhận vào kích thước mẫu bạn muốn rút và trả về các chuyển tiếp đã lấy mẫu dưới dạng tensors, kèm theo chỉ số của chúng trong bộ đệm nhớ và trọng số tầm quan trọng.

Một bộ đệm với sức chứa 10 đã được nạp sẵn trong môi trường của bạn để bạn thực hiện lấy mẫu.

Bài tập này là một phần của khóa học

Deep Reinforcement Learning bằng Python

Xem khóa học

Hướng dẫn bài tập

  • Tính xác suất lấy mẫu tương ứng với mỗi chuyển tiếp.
  • Rút các chỉ số tương ứng với các chuyển tiếp trong mẫu; np.random.choice(a, s, p=p) lấy một mẫu kích thước s có hoàn lại từ mảng a, dựa trên mảng xác suất p.
  • Tính trọng số tầm quan trọng tương ứng với mỗi chuyển tiếp.

Bài tập tương tác thực hành trực tiếp

Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.

def sample(self, batch_size):
    priorities = np.array(self.priorities)
    # Calculate the sampling probabilities
    probabilities = ____ / np.sum(____)
    # Draw the indices for the sample
    indices = np.random.choice(____)
    # Calculate the importance weights
    weights = (1 / (len(self.memory) * ____)) ** ____
    weights /= np.max(weights)
    states, actions, rewards, next_states, dones = zip(*[self.memory[idx] for idx in indices])
    weights = [weights[idx] for idx in indices]
    states_tensor = torch.tensor(states, dtype=torch.float32)
    rewards_tensor = torch.tensor(rewards, dtype=torch.float32)
    next_states_tensor = torch.tensor(next_states, dtype=torch.float32)
    dones_tensor = torch.tensor(dones, dtype=torch.float32)
    weights_tensor = torch.tensor(weights, dtype=torch.float32)
    actions_tensor = torch.tensor(actions, dtype=torch.long).unsqueeze(1)
    return (states_tensor, actions_tensor, rewards_tensor, next_states_tensor,
            dones_tensor, indices, weights_tensor)

PrioritizedReplayBuffer.sample = sample
print("Sampled transitions:\n", buffer.sample(3))
Chỉnh sửa và Chạy Mã