การคำนวณบน pivot table
Pivot table เต็มไปด้วยสถิติสรุปผล แต่นั่นเป็นเพียงจุดเริ่มต้นในการค้นหาข้อมูลเชิงลึก บ่อยครั้งจำเป็นต้องคำนวณเพิ่มเติมจาก pivot table เช่น การหาแถวหรือคอลัมน์ที่มีค่าสูงสุดหรือต่ำสุด
จำได้จากบทที่ 1 ว่าสามารถ subset Series หรือ DataFrame เพื่อกรองแถวที่ต้องการได้ง่ายๆ โดยใช้เงื่อนไขเชิงตรรกะภายในวงเล็บเหลี่ยม เช่น series[series > value]
pandas ถูกโหลดไว้เป็น pd และ DataFrame temp_by_country_city_vs_year พร้อมใช้งานแล้ว
ด้านล่างแสดงผลลัพธ์จาก .head() ของ DataFrame นี้ โดยแสดงเฉพาะบางคอลัมน์ปี:
| country | city | 2000 | 2001 | 2002 | … | 2013 |
|---|---|---|---|---|---|---|
| Afghanistan | Kabul | 15.823 | 15.848 | 15.715 | … | 16.206 |
| Angola | Luanda | 24.410 | 24.427 | 24.791 | … | 24.554 |
| Australia | Melbourne | 14.320 | 14.180 | 14.076 | … | 14.742 |
| Sydney | 17.567 | 17.854 | 17.734 | … | 18.090 | |
| Bangladesh | span translate="no">Dhaka | 25.905 | 25.931 | 26.095 | … | 26.587 |
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
การจัดการข้อมูลด้วย pandas
คำแนะนำการฝึกหัด
- คำนวณอุณหภูมิเฉลี่ยในแต่ละปี แล้วกำหนดให้กับ
mean_temp_by_year - กรอง
mean_temp_by_yearเพื่อหาปีที่มีอุณหภูมิเฉลี่ยสูงสุด - คำนวณอุณหภูมิเฉลี่ยของแต่ละเมือง (ตามคอลัมน์) แล้วกำหนดให้กับ
mean_temp_by_city - กรอง
mean_temp_by_cityเพื่อหาเมืองที่มีอุณหภูมิเฉลี่ยต่ำสุด
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Get the worldwide mean temp by year
mean_temp_by_year = temp_by_country_city_vs_year.____
# Filter for the year that had the highest mean temp
print(mean_temp_by_year[____])
# Get the mean temp by city
mean_temp_by_city = temp_by_country_city_vs_year.____
# Filter for the city that had the lowest mean temp
print(mean_temp_by_city[____])