Lấy mẫu từ bộ đệm PER
Trước khi bạn có thể dùng lớp Prioritized Experience Buffer để huấn luyện agent, bạn vẫn cần triển khai phương thức .sample(). Phương thức này nhận vào kích thước mẫu bạn muốn rút và trả về các chuyển tiếp đã lấy mẫu dưới dạng tensors, kèm theo chỉ số của chúng trong bộ đệm nhớ và trọng số tầm quan trọng.
Một bộ đệm với sức chứa 10 đã được nạp sẵn trong môi trường của bạn để bạn thực hiện lấy mẫu.
Bài tập này là một phần của khóa học
Deep Reinforcement Learning bằng Python
Hướng dẫn bài tập
- Tính xác suất lấy mẫu tương ứng với mỗi chuyển tiếp.
- Rút các chỉ số tương ứng với các chuyển tiếp trong mẫu;
np.random.choice(a, s, p=p)lấy một mẫu kích thướcscó hoàn lại từ mảnga, dựa trên mảng xác suấtp. - Tính trọng số tầm quan trọng tương ứng với mỗi chuyển tiếp.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
def sample(self, batch_size):
priorities = np.array(self.priorities)
# Calculate the sampling probabilities
probabilities = ____ / np.sum(____)
# Draw the indices for the sample
indices = np.random.choice(____)
# Calculate the importance weights
weights = (1 / (len(self.memory) * ____)) ** ____
weights /= np.max(weights)
states, actions, rewards, next_states, dones = zip(*[self.memory[idx] for idx in indices])
weights = [weights[idx] for idx in indices]
states_tensor = torch.tensor(states, dtype=torch.float32)
rewards_tensor = torch.tensor(rewards, dtype=torch.float32)
next_states_tensor = torch.tensor(next_states, dtype=torch.float32)
dones_tensor = torch.tensor(dones, dtype=torch.float32)
weights_tensor = torch.tensor(weights, dtype=torch.float32)
actions_tensor = torch.tensor(actions, dtype=torch.long).unsqueeze(1)
return (states_tensor, actions_tensor, rewards_tensor, next_states_tensor,
dones_tensor, indices, weights_tensor)
PrioritizedReplayBuffer.sample = sample
print("Sampled transitions:\n", buffer.sample(3))