เริ่มต้นใช้งานเริ่มต้นใช้งานได้ฟรี

การเข้ารหัสคอลัมน์ประเภท Categorical II: OneHotEncoder

ตอนนี้เข้ารหัสคอลัมน์ประเภท categorical เป็นตัวเลขแล้ว แต่ยังไม่พร้อมใช้งานกับ pipeline และ XGBoost เลยทีเดียว เพราะในคอลัมน์ categorical ของชุดข้อมูลนี้ ค่าต่าง ๆ ไม่ได้มีลำดับที่มีความหมายในตัวเอง ตัวอย่างเช่น เมื่อใช้ LabelEncoder ย่าน CollgCr ถูกเข้ารหัสเป็น 5 ย่าน Veenker เป็น 24 และย่าน Crawfor เป็น 6 แต่จะบอกว่า Veenker "มากกว่า" Crawfor และ CollgCr ได้จริงหรือ? ไม่ใช่เลย และถ้าปล่อยให้โมเดลสมมติว่ามีลำดับแบบนี้ อาจส่งผลให้ประสิทธิภาพลดลงได้

ดังนั้นจึงต้องมีขั้นตอนเพิ่มเติม คือการใช้ one-hot encoding เพื่อสร้างตัวแปรแบบ binary หรือที่เรียกว่า "dummy" variables โดยใช้ OneHotEncoder ของ scikit-learn

แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร

Extreme Gradient Boosting with XGBoost

ดูคอร์ส

คำแนะนำการฝึกหัด

  • นำเข้า OneHotEncoder จาก sklearn.preprocessing
  • สร้างออบเจกต์ OneHotEncoder ชื่อ ohe โดยระบุอาร์กิวเมนต์คีย์เวิร์ด sparse=False
  • ใช้เมธอด .fit_transform() เพื่อนำ OneHotEncoder ไปใช้กับ df และบันทึกผลลัพธ์ไว้ในตัวแปร df_encoded ซึ่งผลลัพธ์จะเป็น NumPy array
  • แสดง 5 แถวแรกของ df_encoded จากนั้นแสดง shape ของทั้ง df และ df_encoded เพื่อเปรียบเทียบความแตกต่าง

แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ

ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์

# Import OneHotEncoder
____

# Create OneHotEncoder: ohe
ohe = ____

# Apply OneHotEncoder to categorical columns - output is no longer a dataframe: df_encoded
df_encoded = ____

# Print first 5 rows of the resulting dataset - again, this will no longer be a pandas dataframe
print(df_encoded[:5, :])

# Print the shape of the original DataFrame
print(df.shape)

# Print the shape of the transformed array
print(df_encoded.shape)
แก้ไขและรันโค้ด