การตัดแบ่งและทำความสะอาด PySpark DataFrame
หลังจากตรวจสอบข้อมูลแล้ว มักจำเป็นต้องทำความสะอาดข้อมูล ซึ่งได้แก่ การตัดแบ่งคอลัมน์ การเปลี่ยนชื่อคอลัมน์ และการลบแถวที่ซ้ำกัน เป็นต้น PySpark DataFrame API มี operator หลายตัวที่ช่วยดำเนินการเหล่านี้ได้ ในแบบฝึกหัดนี้ ให้เลือกคอลัมน์ 'name', 'sex' และ 'date of birth' จาก DataFrame people_df จากนั้นลบแถวที่ซ้ำกันออก และนับจำนวนแถวก่อนและหลังการลบข้อมูลซ้ำ
โดย SparkSession spark และ DataFrame people_df ถูกสร้างไว้ใน workspace ของคุณแล้ว
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Big Data Fundamentals with PySpark
คำแนะนำการฝึกหัด
- เลือกคอลัมน์ 'name', 'sex' และ 'date of birth' จาก
people_dfแล้วสร้าง DataFrame ชื่อpeople_df_sub - แสดง 10 แถวแรกของ DataFrame
people_df_sub - ลบแถวที่ซ้ำกันออกจาก
people_df_subแล้วสร้าง DataFrame ชื่อpeople_df_sub_nodup - จำนวนแถวก่อนและหลังการลบข้อมูลซ้ำแตกต่างกันอย่างไร?
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Select name, sex and date of birth columns
people_df_sub = people_df.____('name', ____, ____)
# Print the first 10 observations from people_df_sub
people_df_sub.____(____)
# Remove duplicate entries from people_df_sub
people_df_sub_nodup = people_df_sub.____()
# Count the number of rows
print("There were {} rows before removing duplicates, and {} rows after removing duplicates".format(people_df_sub.____(), people_df_sub_nodup.____()))