สถิติสรุปของ MSD
มาทำความคุ้นเคยกับชุดข้อมูลย่อย Million Songs Echo Nest Taste Profile กัน สำหรับคอร์สนี้ เราจะเรียกชุดข้อมูลนี้ว่า Million Songs dataset หรือ msd เริ่มด้วยการดูจำนวนผู้ใช้และจำนวนเพลง รวมถึงดูว่าเพลงใดในชุดข้อมูลนี้มียอดเล่นสูงที่สุด
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
การสร้าง Recommendation Engines ด้วย PySpark
คำแนะนำการฝึกหัด
- ใช้เมธอด
.show()เพื่อดูว่าข้อมูลมีหน้าตาเป็นอย่างไร - เติมโค้ดให้สมบูรณ์เพื่อนับจำนวน
userIdที่ไม่ซ้ำกัน โดย select คอลัมน์userIdจากนั้นเรียก.distinct()และ.count() - ทำแบบเดียวกันสำหรับ
songIdโดย select คอลัมน์songIdแล้วเรียก.distinct()และ.count()บนคอลัมน์นั้น
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Look at the data
msd.____()
# Count the number of distinct userIds
user_count = msd.select("____").____().count()
print("Number of users: ", user_count)
# Count the number of distinct songIds
song_count = msd.select("____").____().count()
print("Number of songs: ", song_count)