Value-iteratie implementeren

Value-iteratie is een kernmethode in RL om het optimale beleid te vinden. Je verbetert iteratief de waardefunctie voor elke toestand totdat deze convergeert, waardoor je het optimale beleid ontdekt. Je begint met een geïnitialiseerde waardefunctie V en policy, allebei al voor je geladen. Daarna werk je ze bij in een lus totdat de waardefunctie convergeert en zie je het beleid in actie.

De functie get_max_action_and_value(state, V) is alvast voor je geladen.

Deze oefening maakt deel uit van de cursus

Reinforcement Learning met Gymnasium in Python

Cursus bekijken

Oefeninstructies

Zoek voor elke toestand de actie met de maximale Q-waarde (max_action) en de bijbehorende waarde (max_q_value).
Werk het new_V-woordenboek en de policy bij op basis van max_action en max_q_value.
Controleer op convergentie door na te gaan of het verschil tussen new_v en V voor elke toestand kleiner is dan threshold.

Praktische interactieve oefening

Probeer deze oefening eens door deze voorbeeldcode in te vullen.

threshold = 0.001
while True:
  new_V = {}
  for state in range(num_states-1):
    # Get action with maximum Q-value and its value 
    max_action, max_q_value = ____
    # Update the value function and policy
    new_V[state] = ____
    policy[state] = ____
  # Test if change in state values is negligeable
  if all(abs(____ - ____) < ____ for state in ____):
    break
  V = new_V
render_policy(policy)

Code bewerken en uitvoeren

Deze oefening maakt deel uit van de cursus

Reinforcement Learning met Gymnasium in Python

SkillTag.level.advancedSkillTag.label

4.8+

Begin de cursus gratis

Dive into the exciting world of Reinforcement Learning (RL) by exploring its foundational concepts, roles, and applications. Navigate through the RL framework, uncovering the agent-environment interaction. You'll also learn how to use the Gymnasium library to create environments, visualize states, and perform actions, thus gaining a practical foundation in RL concepts and applications.

Exercise 1: Fundamentals of reinforcement learning Exercise 2: What is Reinforcement Learning?Exercise 3: RL vs. other ML sub-domains Exercise 4: Scenarios for applying RL Exercise 5: Navigating the RL framework Exercise 6: RL interaction loop Exercise 7: Episodic and continuous RL tasks Exercise 8: Calculating discounted returns for agent strategies Exercise 9: Interacting with Gymnasium environments Exercise 10: Setting up a Mountain Car environment Exercise 11: Visualizing the Mountain Car Environment Exercise 12: Interacting with the Frozen Lake environment

Delve deeper into the world of RL focusing on model-based learning. Unravel the complexities of Markov Decision Processes (MDPs), understanding their essential components. Enhance your skill set by learning about policies and value functions. Gain expertise in policy optimization with policy iteration and value Iteration techniques.

Exercise 1: Markov-beslissingsprocessen Exercise 2: Aangepaste Frozen Lake-MDP-componenten Exercise 3: Verkennen van toestand- en actieruimtes Exercise 4: Overgangswaarschijnlijkheden en beloningen Exercise 5: Policies en toestandswaardefuncties Exercise 6: Een deterministisch beleid definiëren Exercise 7: Toestandwaardes voor een policy berekenen Exercise 8: Beleid vergelijken Exercise 9: Actiewaardefuncties Exercise 10: Q-waarden berekenen Exercise 11: Een beleid verbeteren Exercise 12: Policy-iteratie en value-iteratie Exercise 13: Policy-iteratie toepassen voor een optimale policy Exercise 14: Value-iteratie implementeren

Huidige oefening

Embark on a journey through the dynamic realm of Model-Free Learning in RL. Get introduced to to the foundational Monte Carlo methods, and apply first-visit and every-visit Monte Carlo prediction algorithms. Transition into the world of Temporal Difference Learning, exploring the SARSA algorithm. Finally, dive into the depths of Q-Learning, and analyze its convergence in challenging environments.

Exercise 1: Monte Carlo methods Exercise 2: Episode generation for Monte Carlo methods Exercise 3: Implementing first-visit Monte Carlo Exercise 4: Implementing every-visit Monte Carlo Exercise 5: Temporal difference learning Exercise 6: Implementing the SARSA update rule Exercise 7: Solving 8x8 Frozen Lake with SARSA Exercise 8: Q-learning Exercise 9: Implementing Q-learning update rule Exercise 10: Solving 8x8 Frozen Lake with Q-learning Exercise 11: Evaluating policy on a slippery Frozen Lake

Dive into advanced strategies in Model-Free RL, focusing on enhancing decision-making algorithms. Learn about Expected SARSA for more accurate policy updates and Double Q-learning to mitigate overestimation bias. Explore the Exploration-Exploitation Tradeoff, mastering epsilon-greedy and epsilon-decay strategies for optimal action selection. Tackle the Multi-Armed Bandit Problem, applying strategies to solve decision-making challenges under uncertainty.

Exercise 1: Expected SARSA Exercise 2: Expected SARSA update rule Exercise 3: Applying Expected SARSA Exercise 4: Double Q-learning Exercise 5: Implementing double Q-learning update rule Exercise 6: Applying double Q-learning Exercise 7: Balancing exploration and exploitation Exercise 8: Defining epsilon-greedy function Exercise 9: Solving CliffWalking with epsilon greedy strategy Exercise 10: Solving CliffWalking with decayed epsilon-greedy strategy Exercise 11: Multi-armed bandits Exercise 12: Creating a multi-armed bandit Exercise 13: Solving a multi-armed bandit Exercise 14: Assessing convergence in a multi-armed bandit Exercise 15: Congratulations!