Datentypen prüfen
Im Datenzeitalter haben wir Zugriff auf mehr Attribute als je zuvor. Um all diese zu handhaben, bauen wir viel Automatisierung auf – mindestens genauso wichtig ist jedoch, dass ihre Datentypen korrekt sind. In dieser Übung validieren wir ein Dictionary mit Attributen und ihren Datentypen, um zu prüfen, ob sie stimmen. Dieses Dictionary ist in der Variablen validation_dict gespeichert und steht dir in deinem Workspace zur Verfügung.
Diese Übung ist Teil des Kurses
<Kurs>Feature Engineering mit PySpark</Kurs>Übungsanweisungen
- Erzeuge aus
dfmitdtypeseine Liste von Tupeln aus Attribut und Datentyp namensactual_dtypes_list. - Iteriere über
actual_dtypes_listund prüfe, ob die Spaltennamen im Dictionary der erwarteten Datentypenvalidation_dictexistieren. - Für die Schlüssel, die im Dictionary vorhanden sind, prüfe die Datentypen und gib diejenigen aus, die übereinstimmen.
Interaktive praktische Übung
Versuche dich an dieser Übung, indem du diesen Beispielcode vervollständigst.
# create list of actual dtypes to check
actual_dtypes_list = df.____
print(actual_dtypes_list)
# Iterate through the list of actual dtypes tuples
for attribute_tuple in ____:
# Check if column name is dictionary of expected dtypes
col_name = attribute_tuple[____]
if col_name in ____:
# Compare attribute types
col_type = attribute_tuple[____]
if col_type == validation_dict[____]:
print(col_name + ' has expected dtype.')