การเข้ารหัสคอลัมน์ประเภท Categorical II: OneHotEncoder
ตอนนี้เข้ารหัสคอลัมน์ประเภท categorical เป็นตัวเลขแล้ว แต่ยังไม่พร้อมใช้งานกับ pipeline และ XGBoost เลยทีเดียว เพราะในคอลัมน์ categorical ของชุดข้อมูลนี้ ค่าต่าง ๆ ไม่ได้มีลำดับที่มีความหมายในตัวเอง ตัวอย่างเช่น เมื่อใช้ LabelEncoder ย่าน CollgCr ถูกเข้ารหัสเป็น 5 ย่าน Veenker เป็น 24 และย่าน Crawfor เป็น 6 แต่จะบอกว่า Veenker "มากกว่า" Crawfor และ CollgCr ได้จริงหรือ? ไม่ใช่เลย และถ้าปล่อยให้โมเดลสมมติว่ามีลำดับแบบนี้ อาจส่งผลให้ประสิทธิภาพลดลงได้
ดังนั้นจึงต้องมีขั้นตอนเพิ่มเติม คือการใช้ one-hot encoding เพื่อสร้างตัวแปรแบบ binary หรือที่เรียกว่า "dummy" variables โดยใช้ OneHotEncoder ของ scikit-learn
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Extreme Gradient Boosting with XGBoost
คำแนะนำการฝึกหัด
- นำเข้า
OneHotEncoderจากsklearn.preprocessing - สร้างออบเจกต์
OneHotEncoderชื่อoheโดยระบุอาร์กิวเมนต์คีย์เวิร์ดsparse=False - ใช้เมธอด
.fit_transform()เพื่อนำOneHotEncoderไปใช้กับdfและบันทึกผลลัพธ์ไว้ในตัวแปรdf_encodedซึ่งผลลัพธ์จะเป็น NumPy array - แสดง 5 แถวแรกของ
df_encodedจากนั้นแสดง shape ของทั้งdfและdf_encodedเพื่อเปรียบเทียบความแตกต่าง
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Import OneHotEncoder
____
# Create OneHotEncoder: ohe
ohe = ____
# Apply OneHotEncoder to categorical columns - output is no longer a dataframe: df_encoded
df_encoded = ____
# Print first 5 rows of the resulting dataset - again, this will no longer be a pandas dataframe
print(df_encoded[:5, :])
# Print the shape of the original DataFrame
print(df.shape)
# Print the shape of the transformed array
print(df_encoded.shape)