Implementare la regola di aggiornamento del Q-learning

Il Q-learning è un algoritmo off-policy nel reinforcement learning (RL) che cerca la migliore azione da compiere dato lo stato attuale. A differenza di SARSA, che considera la prossima azione effettivamente eseguita, il Q-learning aggiorna i suoi valori Q usando il massimo premio futuro indipendentemente dall’azione intrapresa. Questa differenza permette al Q-learning di apprendere la politica ottimale pur seguendo una politica esplorativa o persino casuale. Ecco il compito: implementare una funzione che aggiorni una Q-table in base alla regola del Q-learning. Qui sotto trovi la regola di aggiornamento del Q-learning; il tuo compito è implementare una funzione che aggiorni la Q-table seguendo questa regola.

La libreria NumPy è stata importata come np.

Image showing the mathematical formula of the Q-learning update rule.

Questo esercizio fa parte del corso

Reinforcement Learning con Gymnasium in Python

Visualizza il corso

Istruzioni dell'esercizio

Recupera l’attuale valore Q per la coppia stato-azione fornita.
Determina il valore Q massimo per lo stato successivo tra tutte le azioni possibili in actions.
Aggiorna il valore Q per la coppia stato-azione corrente usando la formula del Q-learning.
Aggiorna la Q-table Q, dato che un agente esegue l’azione 0 nello stato 0, riceve una ricompensa di 5 e si sposta allo stato 1.

Esercizio pratico interattivo

Prova a risolvere questo esercizio completando il codice di esempio.

actions = ['action1', 'action2'] 
def update_q_table(state, action, reward, next_state):
  	# Get the old value of the current state-action pair
    old_value = ____
    # Determine the maximum Q-value for the next state
    next_max = ____
    # Compute the new value of the current state-action pair
    Q[state, action] = ____

alpha = 0.1
gamma = 0.95
Q = np.array([[10, 8], [20, 15]], dtype='float32')
# Update the Q-table
____
print(Q)

Modifica ed esegui il codice

Questo esercizio fa parte del corso

Reinforcement Learning con Gymnasium in Python

AvançadoNível de habilidade

4.8+

Inizia il corso gratis

Dive into the exciting world of Reinforcement Learning (RL) by exploring its foundational concepts, roles, and applications. Navigate through the RL framework, uncovering the agent-environment interaction. You'll also learn how to use the Gymnasium library to create environments, visualize states, and perform actions, thus gaining a practical foundation in RL concepts and applications.

Exercise 1: Fundamentals of reinforcement learning Exercise 2: What is Reinforcement Learning?Exercise 3: RL vs. other ML sub-domains Exercise 4: Scenarios for applying RL Exercise 5: Navigating the RL framework Exercise 6: RL interaction loop Exercise 7: Episodic and continuous RL tasks Exercise 8: Calculating discounted returns for agent strategies Exercise 9: Interacting with Gymnasium environments Exercise 10: Setting up a Mountain Car environment Exercise 11: Visualizing the Mountain Car Environment Exercise 12: Interacting with the Frozen Lake environment

Delve deeper into the world of RL focusing on model-based learning. Unravel the complexities of Markov Decision Processes (MDPs), understanding their essential components. Enhance your skill set by learning about policies and value functions. Gain expertise in policy optimization with policy iteration and value Iteration techniques.

Exercise 1: Markov Decision Processes Exercise 2: Custom Frozen Lake MDP components Exercise 3: Exploring state and action spaces Exercise 4: Transition probabilities and rewards Exercise 5: Policies and state-value functions Exercise 6: Defining a deterministic policy Exercise 7: Computing state-values for a policy Exercise 8: Comparing policies Exercise 9: Action-value functions Exercise 10: Computing Q-values Exercise 11: Improving a policy Exercise 12: Policy iteration and value iteration Exercise 13: Applying policy iteration for optimal policy Exercise 14: Implementing value iteration

Embark on a journey through the dynamic realm of Model-Free Learning in RL. Get introduced to to the foundational Monte Carlo methods, and apply first-visit and every-visit Monte Carlo prediction algorithms. Transition into the world of Temporal Difference Learning, exploring the SARSA algorithm. Finally, dive into the depths of Q-Learning, and analyze its convergence in challenging environments.

Exercise 1: Metodi Monte Carlo Exercise 2: Generazione di episodi per i metodi Monte Carlo Exercise 3: Implementare Monte Carlo a prima visita Exercise 4: Implementare Every-Visit Monte Carlo Exercise 5: Apprendimento a differenze temporali Exercise 6: Implementare la regola di aggiornamento SARSA Exercise 7: Risolvi Frozen Lake 8x8 con SARSA Exercise 8: Q-learning Exercise 9: Implementare la regola di aggiornamento del Q-learning

Esercizio in corso

Exercise 10: Risolvi Frozen Lake 8x8 con Q-learning Exercise 11: Valutare una policy su un Frozen Lake scivoloso

Dive into advanced strategies in Model-Free RL, focusing on enhancing decision-making algorithms. Learn about Expected SARSA for more accurate policy updates and Double Q-learning to mitigate overestimation bias. Explore the Exploration-Exploitation Tradeoff, mastering epsilon-greedy and epsilon-decay strategies for optimal action selection. Tackle the Multi-Armed Bandit Problem, applying strategies to solve decision-making challenges under uncertainty.

Exercise 1: Expected SARSA Exercise 2: Expected SARSA update rule Exercise 3: Applying Expected SARSA Exercise 4: Double Q-learning Exercise 5: Implementing double Q-learning update rule Exercise 6: Applying double Q-learning Exercise 7: Balancing exploration and exploitation Exercise 8: Defining epsilon-greedy function Exercise 9: Solving CliffWalking with epsilon greedy strategy Exercise 10: Solving CliffWalking with decayed epsilon-greedy strategy Exercise 11: Multi-armed bandits Exercise 12: Creating a multi-armed bandit Exercise 13: Solving a multi-armed bandit Exercise 14: Assessing convergence in a multi-armed bandit Exercise 15: Congratulations!