Gefixeerde Q-targets

Je gaat je Lunar Lander trainen met gefixeerde Q-targets. Als voorbereiding moet je zowel het online netwerk (dat de actie kiest) als het target-netwerk (gebruikt voor de TD-targetberekening) instantieren.

Je moet ook een functie update_target_network implementeren die je bij elke trainingsstap kunt gebruiken. Het target-netwerk wordt niet met gradient descent geüpdatet; in plaats daarvan duwt update_target_network de gewichten een klein stukje richting het Q-netwerk, zodat het over tijd stabiel blijft.

Let op: alleen voor deze oefening gebruik je een heel klein netwerk zodat we de state dictionary eenvoudig kunnen afdrukken en bekijken. Het heeft slechts één verborgen laag van grootte twee; de actieruimte en toestandsruimte hebben ook dimensie 2.

De functie print_state_dict() is beschikbaar in je omgeving om de state dict af te drukken.

Deze oefening maakt deel uit van de cursus

Deep Reinforcement Learning in Python

Cursus bekijken

Oefeninstructies

Haal de .state_dict() op voor zowel het target- als het online netwerk.
Werk de state dict voor het target-netwerk bij door het gewogen gemiddelde te nemen van de parameters van het online netwerk en het target-netwerk, waarbij je tau gebruikt als gewicht voor het online netwerk.
Laad de bijgewerkte state dict terug in het target-netwerk.

Praktische interactieve oefening

Probeer deze oefening eens door deze voorbeeldcode in te vullen.

def update_target_network(target_network, online_network, tau):
    # Obtain the state dicts for both networks
    target_net_state_dict = ____
    online_net_state_dict = ____
    for key in online_net_state_dict:
        # Calculate the updated state dict for the target network
        target_net_state_dict[key] = (online_net_state_dict[____] * ____ + target_net_state_dict[____] * ____)
        # Load the updated state dict into the target network
        target_network.____
    return None
  
print("online network weights:", print_state_dict(online_network))
print("target network weights (pre-update):", print_state_dict(target_network))
update_target_network(target_network, online_network, .001)
print("target network weights (post-update):", print_state_dict(target_network))

Code bewerken en uitvoeren

Deze oefening maakt deel uit van de cursus

Deep Reinforcement Learning in Python

SkillTag.level.advancedSkillTag.label

4.8+

Begin de cursus gratis

Discover how deep reinforcement learning improves upon traditional Reinforcement Learning while studying and implementing your first Deep Q Learning algorithm.

Exercise 1: Introduction to deep reinforcement learning Exercise 2: Environment and neural network setup Exercise 3: DRL training loop Exercise 4: Introduction to deep Q learning Exercise 5: Deep learning and DQN Exercise 6: The Q-Network architecture Exercise 7: Instantiating the Q-Network Exercise 8: The barebone DQN algorithm Exercise 9: Barebone DQN action selection Exercise 10: Barebone DQN loss function Exercise 11: Training the barebone DQN

Dive into Deep Q-learning by implementing the original DQN algorithm, featuring Experience Replay, epsilon-greediness and fixed Q-targets. Beyond DQN, you will then explore two fascinating extensions that improve the performance and stability of Deep Q-learning: Double DQN and Prioritized Experience Replay.

Exercise 1: DQN met experience replay Exercise 2: De double-ended queue Exercise 3: Experience replay-buffer Exercise 4: DQN met experience replay Exercise 5: Het complete DQN-algoritme Exercise 6: Epsilon-greediness Exercise 7: Gefixeerde Q-targets

Huidige oefening

Exercise 8: Het complete DQN-algoritme implementeren Exercise 9: Double DQN Exercise 10: Online netwerk en targetnetwerk in DDQN Exercise 11: De Double DQN trainen Exercise 12: Prioritized experience replay Exercise 13: Prioritized experience replay-buffer Exercise 14: Steekproeven uit de PER-buffer Exercise 15: DQN met prioritaire experience replay

Learn about the foundational concepts of policy gradient methods found in DRL. You will begin with the policy gradient theorem, which forms the basis for these methods. Then, you will implement the REINFORCE algorithm, a powerful approach to learning policies. The chapter will then guide you through Actor-Critic methods, focusing on the Advantage Actor-Critic (A2C) algorithm, which combines the strengths of both policy gradient and value-based methods to enhance learning efficiency and stability.

Exercise 1: Introduction to policy gradient Exercise 2: The policy network architecture Exercise 3: Working with discrete distributions Exercise 4: Policy gradient and REINFORCE Exercise 5: Action selection in REINFORCE Exercise 6: Training the REINFORCE algorithm Exercise 7: Advantage Actor Critic Exercise 8: Critic network Exercise 9: Actor Critic loss calculations Exercise 10: Training the A2C algorithm

Explore Proximal Policy Optimization (PPO) for robust DRL performance. Next, you will examine using an entropy bonus in PPO, which encourages exploration by preventing premature convergence to deterministic policies. You'll also learn about batch updates in policy gradient methods. Finally, you will learn about hyperparameter optimization with Optuna, a powerful tool for optimizing performance in your DRL models.

Exercise 1: Proximal policy optimization Exercise 2: The clipped probability ratio Exercise 3: The clipped surrogate objective function Exercise 4: Entropy bonus and PPO Exercise 5: Entropy playground Exercise 6: Training the PPO algorithm Exercise 7: Batch updates in policy gradient Exercise 8: Minibatch and DRL Exercise 9: A2C with batch updates Exercise 10: Hyperparameter optimization with Optuna Exercise 11: Hyperparameter or not?Exercise 12: Hands-on with Optuna Exercise 13: Congratulations!