การวิเคราะห์ความถี่ของคำ
ยินดีด้วย! คุณเพิ่งเข้าร่วมทีม PyBooks ซึ่งกำลังพัฒนาระบบแนะนำหนังสือและต้องการค้นหารูปแบบและแนวโน้มในข้อความเพื่อปรับปรุงคุณภาพการแนะนำ
ในขั้นต้น จะเริ่มด้วยการทำความเข้าใจความถี่ของคำในข้อความที่กำหนด และลบคำที่พบน้อยครั้งออก
โปรดทราบว่าชุดข้อมูลในสถานการณ์จริงมักจะมีขนาดใหญ่กว่าตัวอย่างนี้
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Deep Learning สำหรับข้อความด้วย PyTorch
คำแนะนำการฝึกหัด
- นำเข้า
get_tokenizerจากtorchtextและFreqDistจากไลบรารีnltk - กำหนดค่าเริ่มต้นให้ตัว tokenizer สำหรับภาษาอังกฤษ จากนั้น tokenize ข้อความ
textที่กำหนดให้ - คำนวณการกระจายความถี่ของ
tokensและลบคำที่พบน้อยครั้งโดยใช้ list comprehension
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Import the necessary functions
from torchtext.data.utils import ____
from nltk.probability import ____
text = "In the city of Dataville, a data analyst named Alex explores hidden insights within vast data. With determination, Alex uncovers patterns, cleanses the data, and unlocks innovation. Join this adventure to unleash the power of data-driven decisions."
# Initialize the tokenizer and tokenize the text
tokenizer = ____("basic_english")
tokens = tokenizer(____)
threshold = 1
# Remove rare words and print common tokens
freq_dist = ____(____)
common_tokens = [token for token in tokens if ____[token] > ____]
print(common_tokens)