1. Learn
  2. /
  3. Courses
  4. /
  5. Feature Engineering with PySpark

Exercise

Verifying DataTypes

In the age of data we have access to more attributes than we ever had before. To handle all of them we will build a lot of automation but at a minimum requires that their datatypes be correct. In this exercise we will validate a dictionary of attributes and their datatypes to see if they are correct. This dictionary is stored in the variable validation_dict and is available in your workspace.

Instructions

100 XP
  • Using df create a list of attribute and datatype tuples with dtypes called actual_dtypes_list.
  • Iterate through actual_dtypes_list, checking if the column names exist in the dictionary of expected dtypes validation_dict.
  • For the keys that exist in the dictionary, check their dtypes and print those that match.