Risolvere CliffWalking con strategia epsilon-greedy a decadimento

Per potenziare la strategia epsilon-greedy, si introduce un fattore di decadimento per ridurre gradualmente il tasso di esplorazione, epsilon, man mano che l'agente impara di più sull'ambiente. Questo approccio favorisce l'esplorazione nelle prime fasi dell'apprendimento e lo sfruttamento delle conoscenze acquisite quando l'agente diventa più familiare con l'ambiente. Ora applicherai questa strategia per risolvere l'ambiente CliffWalking.

L'ambiente è stato inizializzato ed è accessibile tramite la variabile env. Le variabili epsilon, min_epsilon ed epsilon_decay sono già state definite per te. Le funzioni epsilon_greedy() e update_q_table() sono state importate.

Questo esercizio fa parte del corso

Reinforcement Learning con Gymnasium in Python

Visualizza il corso

Istruzioni dell'esercizio

Implementa l'intero ciclo di training scegliendo un action, eseguendola, accumulando il reward ricevuto in episode_reward e aggiornando la Q-table.
Riduci epsilon usando il tasso epsilon_decay, assicurandoti che non scenda sotto min_epsilon.

Esercizio pratico interattivo

Prova a risolvere questo esercizio completando il codice di esempio.

rewards_decay_eps_greedy = []
for episode in range(total_episodes):
    state, info = env.reset()
    episode_reward = 0
    for i in range(max_steps):
      	# Implement the training loop
        action = ____
        new_state, reward, terminated, truncated, info = ____
        episode_reward += ____       
        ____      
        state = new_state
    rewards_decay_eps_greedy.append(episode_reward)
    # Update epsilon
    epsilon = ____
print("Average reward per episode: ", np.mean(rewards_decay_eps_greedy))

Modifica ed esegui il codice

Questo esercizio fa parte del corso

Reinforcement Learning con Gymnasium in Python

AvançadoNível de habilidade

4.8+

Inizia il corso gratis

Dive into the exciting world of Reinforcement Learning (RL) by exploring its foundational concepts, roles, and applications. Navigate through the RL framework, uncovering the agent-environment interaction. You'll also learn how to use the Gymnasium library to create environments, visualize states, and perform actions, thus gaining a practical foundation in RL concepts and applications.

Exercise 1: Fundamentals of reinforcement learning Exercise 2: What is Reinforcement Learning?Exercise 3: RL vs. other ML sub-domains Exercise 4: Scenarios for applying RL Exercise 5: Navigating the RL framework Exercise 6: RL interaction loop Exercise 7: Episodic and continuous RL tasks Exercise 8: Calculating discounted returns for agent strategies Exercise 9: Interacting with Gymnasium environments Exercise 10: Setting up a Mountain Car environment Exercise 11: Visualizing the Mountain Car Environment Exercise 12: Interacting with the Frozen Lake environment

Delve deeper into the world of RL focusing on model-based learning. Unravel the complexities of Markov Decision Processes (MDPs), understanding their essential components. Enhance your skill set by learning about policies and value functions. Gain expertise in policy optimization with policy iteration and value Iteration techniques.

Exercise 1: Markov Decision Processes Exercise 2: Custom Frozen Lake MDP components Exercise 3: Exploring state and action spaces Exercise 4: Transition probabilities and rewards Exercise 5: Policies and state-value functions Exercise 6: Defining a deterministic policy Exercise 7: Computing state-values for a policy Exercise 8: Comparing policies Exercise 9: Action-value functions Exercise 10: Computing Q-values Exercise 11: Improving a policy Exercise 12: Policy iteration and value iteration Exercise 13: Applying policy iteration for optimal policy Exercise 14: Implementing value iteration

Embark on a journey through the dynamic realm of Model-Free Learning in RL. Get introduced to to the foundational Monte Carlo methods, and apply first-visit and every-visit Monte Carlo prediction algorithms. Transition into the world of Temporal Difference Learning, exploring the SARSA algorithm. Finally, dive into the depths of Q-Learning, and analyze its convergence in challenging environments.

Exercise 1: Monte Carlo methods Exercise 2: Episode generation for Monte Carlo methods Exercise 3: Implementing first-visit Monte Carlo Exercise 4: Implementing every-visit Monte Carlo Exercise 5: Temporal difference learning Exercise 6: Implementing the SARSA update rule Exercise 7: Solving 8x8 Frozen Lake with SARSA Exercise 8: Q-learning Exercise 9: Implementing Q-learning update rule Exercise 10: Solving 8x8 Frozen Lake with Q-learning Exercise 11: Evaluating policy on a slippery Frozen Lake

Dive into advanced strategies in Model-Free RL, focusing on enhancing decision-making algorithms. Learn about Expected SARSA for more accurate policy updates and Double Q-learning to mitigate overestimation bias. Explore the Exploration-Exploitation Tradeoff, mastering epsilon-greedy and epsilon-decay strategies for optimal action selection. Tackle the Multi-Armed Bandit Problem, applying strategies to solve decision-making challenges under uncertainty.

Exercise 1: Expected SARSA Exercise 2: Regola di aggiornamento di Expected SARSA Exercise 3: Applicare Expected SARSA Exercise 4: Double Q-learning Exercise 5: Implementare la regola di aggiornamento del Double Q-learning Exercise 6: Applicare il Double Q-learning Exercise 7: Bilanciare esplorazione e sfruttamento Exercise 8: Definire la funzione epsilon-greedy Exercise 9: Risolvi CliffWalking con la strategia epsilon-greedy Exercise 10: Risolvere CliffWalking con strategia epsilon-greedy a decadimento

Esercizio in corso

Exercise 11: Banditi a più braccia Exercise 12: Creare un multi-armed bandit Exercise 13: Risolvi un multi-armed bandit Exercise 14: Valutare la convergenza in un multi-armed bandit Exercise 15: Congratulazioni!