Phân tích tần suất từ
Chúc mừng bạn! Bạn vừa gia nhập PyBooks. PyBooks đang phát triển hệ thống gợi ý sách và muốn tìm các mẫu và xu hướng trong văn bản để cải thiện đề xuất.
Để bắt đầu, bạn cần hiểu tần suất xuất hiện của các từ trong một văn bản nhất định và loại bỏ các từ hiếm.
Lưu ý rằng các bộ dữ liệu thực tế thường lớn hơn ví dụ này.
Bài tập này là một phần của khóa học
Deep Learning cho Văn bản với PyTorch
Hướng dẫn bài tập
- Import
get_tokenizertừtorchtextvàFreqDisttừ thư việnnltk. - Khởi tạo bộ tách từ cho tiếng Anh và tách token cho
textđã cho. - Tính phân bố tần suất của
tokensvà dùng list comprehension để loại bỏ các từ hiếm.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Import the necessary functions
from torchtext.data.utils import ____
from nltk.probability import ____
text = "In the city of Dataville, a data analyst named Alex explores hidden insights within vast data. With determination, Alex uncovers patterns, cleanses the data, and unlocks innovation. Join this adventure to unleash the power of data-driven decisions."
# Initialize the tokenizer and tokenize the text
tokenizer = ____("basic_english")
tokens = tokenizer(____)
threshold = 1
# Remove rare words and print common tokens
freq_dist = ____(____)
common_tokens = [token for token in tokens if ____[token] > ____]
print(common_tokens)