Extraer un parámetro de Random Forest

Ahora vas a adaptar el trabajo que hiciste con el modelo de regresión logística a un modelo de random forest. Un parámetro de este modelo es, para un árbol concreto, cómo decide dividir en cada nivel.

Este análisis no resulta tan útil como los coeficientes de la regresión logística, ya que es poco probable que explores cada división y cada árbol de un random forest. Aun así, es un ejercicio muy útil para mirar “bajo el capó” y entender qué hace el modelo.

En este ejercicio extraeremos un único árbol de nuestro modelo de random forest, lo visualizaremos y extraeremos por código una de sus divisiones.

Tienes disponible:

Un objeto de modelo de random forest, rf_clf
Una imagen de la parte superior del árbol de decisión seleccionado, tree_viz_image
El DataFrame X_train y la lista original_variables

Este ejercicio forma parte del curso

Ajuste de hiperparámetros en Python

Instrucciones del ejercicio

Extrae el séptimo árbol (índice 6) del modelo de random forest.
Visualiza este árbol (tree_viz_image) para ver las decisiones de división.
Extrae la variable y el umbral de la división superior.
Imprime juntos la variable y el umbral.

ejercicio interactivo práctico

Prueba este ejercicio completando este código de ejemplo.

# Extract the 7th (index 6) tree from the random forest
chosen_tree = rf_clf.estimators_[____]

# Visualize the graph using the provided image
imgplot = plt.imshow(____)
plt.show()

# Extract the parameters and level of the top (index 0) node
split_column = chosen_tree.tree_.feature[____]
split_column_name = X_train.columns[split_column]
split_value = chosen_tree.tree_.threshold[____]

# Print out the feature and level
print("This node split on feature {}, at a value of {}".format(split_column_name, ____))

Editar y ejecutar código

Este ejercicio forma parte del curso

Ajuste de hiperparámetros en Python

IntermedioNivel de habilidad

4.9+

Empieza el curso gratis

In this introductory chapter you will learn the difference between hyperparameters and parameters. You will practice extracting and analyzing parameters, setting hyperparameter values for several popular machine learning algorithms. Along the way you will learn some best practice tips & tricks for choosing which hyperparameters to tune and what values to set & build learning curves to analyze your hyperparameter choices.

Exercise 1: Introducción y «parámetros»Exercise 2: Parámetros en la regresión logística Exercise 3: Extraer un parámetro de Logistic Regression Exercise 4: Extraer un parámetro de Random Forest

Ejercicio actual

Exercise 5: Introducción a los hiperparámetros Exercise 6: Hiperparámetros en Random Forests Exercise 7: Explorando los hiperparámetros de Random Forest Exercise 8: Hiperparámetros de KNN Exercise 9: Definir y analizar valores de hiperparámetros Exercise 10: Automatizar la elección de hiperparámetros Exercise 11: Construir curvas de aprendizaje

This chapter introduces you to a popular automated hyperparameter tuning methodology called Grid Search. You will learn what it is, how it works and practice undertaking a Grid Search using Scikit Learn. You will then learn how to analyze the output of a Grid Search & gain practical experience doing this.

Exercise 1: Introducing Grid Search Exercise 2: Build Grid Search functions Exercise 3: Iteratively tune multiple hyperparameters Exercise 4: How Many Models?Exercise 5: Grid Search with Scikit Learn Exercise 6: GridSearchCV inputs Exercise 7: GridSearchCV with Scikit Learn Exercise 8: Understanding a grid search output Exercise 9: Using the best outputs Exercise 10: Exploring the grid search results Exercise 11: Analyzing the best results Exercise 12: Using the best results

In this chapter you will be introduced to another popular automated hyperparameter tuning methodology called Random Search. You will learn what it is, how it works and importantly how it differs from grid search. You will learn some advantages and disadvantages of this method and when to choose this method compared to Grid Search. You will practice undertaking a Random Search with Scikit Learn as well as visualizing & interpreting the output.

Exercise 1: Introducing Random Search Exercise 2: Randomly Sample Hyperparameters Exercise 3: Randomly Search with Random Forest Exercise 4: Visualizing a Random Search Exercise 5: Random Search in Scikit Learn Exercise 6: RandomSearchCV inputs Exercise 7: The RandomizedSearchCV Object Exercise 8: RandomSearchCV in Scikit Learn Exercise 9: Comparing Grid and Random Search Exercise 10: Comparing Random & Grid Search Exercise 11: Grid and Random Search Side by Side

In this final chapter you will be given a taste of more advanced hyperparameter tuning methodologies known as ''informed search''. This includes a methodology known as Coarse To Fine as well as Bayesian & Genetic hyperparameter tuning algorithms. You will learn how informed search differs from uninformed search and gain practical skills with each of the mentioned methodologies, comparing and contrasting them as you go.

Exercise 1: Informed Search: Coarse to Fine Exercise 2: Visualizing Coarse to Fine Exercise 3: Coarse to Fine Iterations Exercise 4: Informed Search: Bayesian Statistics Exercise 5: Bayes Rule in Python Exercise 6: Bayesian Hyperparameter tuning with Hyperopt Exercise 7: Informed Search: Genetic Algorithms Exercise 8: Genetic Hyperparameter Tuning with TPOT Exercise 9: Analysing TPOT's stability Exercise 10: Congratulations!