การสร้าง UDF สำหรับข้อมูลเวกเตอร์
DataFrame df พร้อมใช้งานแล้ว โดยมีคอลัมน์ output ที่มีชนิดข้อมูลเป็น vector ซึ่งแสดง 5 แถวแรกไว้ใน console
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Introduction to Spark SQL in Python
คำแนะนำการฝึกหัด
- สร้าง UDF ชื่อ
first_udfเพื่อดึง element แรกออกจากคอลัมน์เวกเตอร์ กำหนดค่า default เป็น 0.0 สำหรับ item ที่ไม่ใช่เวกเตอร์หรือมีจำนวน element น้อยกว่าหนึ่ง และแปลงผลลัพธ์เป็น float - ใช้ operation
selectบนdfเพื่อนำfirst_udfไปใช้กับคอลัมน์output
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Selects the first element of a vector column
first_udf = ____(lambda x:
____(x.indices[0])
if (x and hasattr(x, "toArray") and x.____())
else 0.0,
FloatType())
# Apply first_udf to the output column
df.select(____("output").alias("result")).show(5)