Chọn lọc và làm sạch DataFrame trong PySpark
Sau khi kiểm tra dữ liệu, bạn thường cần làm sạch dữ liệu, chủ yếu bao gồm chọn lọc (subsetting), đổi tên cột, loại bỏ các hàng trùng lặp, v.v. PySpark DataFrame API cung cấp nhiều toán tử để thực hiện điều này. Trong bài tập này, nhiệm vụ của bạn là chọn các cột 'name', 'sex' và 'date of birth' từ DataFrame people_df, loại bỏ mọi hàng trùng lặp khỏi tập dữ liệu đó và đếm số hàng trước và sau bước loại bỏ trùng lặp.
Lưu ý: Bạn đã có sẵn SparkSession spark và DataFrame people_df trong không gian làm việc.
Bài tập này là một phần của khóa học
Nền tảng Big Data với PySpark
Hướng dẫn bài tập
- Chọn các cột 'name', 'sex' và 'date of birth' từ
people_dfvà tạo DataFramepeople_df_sub. - In 10 quan sát đầu tiên trong DataFrame
people_df_sub. - Loại bỏ các bản ghi trùng lặp từ DataFrame
people_df_subvà tạo DataFramepeople_df_sub_nodup. - Có bao nhiêu hàng trước và sau khi loại bỏ trùng lặp?
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Select name, sex and date of birth columns
people_df_sub = people_df.____('name', ____, ____)
# Print the first 10 observations from people_df_sub
people_df_sub.____(____)
# Remove duplicate entries from people_df_sub
people_df_sub_nodup = people_df_sub.____()
# Count the number of rows
print("There were {} rows before removing duplicates, and {} rows after removing duplicates".format(people_df_sub.____(), people_df_sub_nodup.____()))