เริ่มต้นใช้งานเริ่มต้นใช้งานได้ฟรี

การหา Stem จากทวีต

ในแบบฝึกหัดนี้ จะได้ทำงานกับอาร์เรย์ชื่อ tweets ซึ่งเก็บข้อความจากข้อมูลความรู้สึก (sentiment) ของสายการบินที่รวบรวมจาก Twitter

โจทย์คือแปลงอาร์เรย์นี้ให้เป็นลิสต์ของ token โดยใช้ list comprehension จากนั้นวนซ้ำผ่านลิสต์ของ token แล้วสร้าง stem จาก token แต่ละตัว ทั้งนี้ list comprehension คือทางเลือกแบบบรรทัดเดียวแทนการใช้ลูป for

แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร

Sentiment Analysis ด้วย Python

ดูคอร์ส

คำแนะนำการฝึกหัด

  • Import ฟังก์ชันที่ใช้แปลงสตริงให้เป็น stem
  • เรียกใช้ฟังก์ชัน Porter stemmer ที่เพิ่ง import มา
  • สร้างลิสต์ tokens โดยใช้ list comprehension โดยให้ลิสต์นี้เก็บ word token ทั้งหมดจากอาร์เรย์ tweets
  • วนซ้ำผ่านลิสต์ tokens แล้วนำฟังก์ชัน stemming ไปใช้กับแต่ละรายการในลิสต์

แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ

ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์

# Import the function to perform stemming
____
from nltk import word_tokenize

# Call the stemmer
porter = ____()

# Transform the array of tweets to tokens
tokens = [____]
# Stem the list of tokens
stemmed_tokens = [[____.____(word) for word in tweet] for tweet in tokens] 
# Print the first element of the list
print(stemmed_tokens[0])
แก้ไขและรันโค้ด