Làm sạch dữ liệu văn bản
Giờ bạn đã xác định được stopwords và dấu câu, hãy dùng chúng để làm sạch thêm email enron trong dataframe df. Các danh sách chứa stopwords và dấu câu có sẵn trong stop và exclude. Vẫn còn vài bước nữa trước khi dữ liệu thực sự sạch, như "lemmatization" từ và stemming động từ. Các động từ trong dữ liệu email đã được stemming, và việc lemmatization cũng đã được thực hiện sẵn cho bạn trong bài tập này.
Bài tập này là một phần của khóa học
Phát hiện gian lận với Python
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Import the lemmatizer from nltk
from nltk.stem.wordnet import WordNetLemmatizer
lemma = WordNetLemmatizer()
# Define word cleaning function
def clean(text, stop):
text = text.____()
# Remove stopwords
stop_free = " ".join([word for word in text.lower().split() if ((___ not in ___) and (not word.isdigit()))])
# Remove punctuations
punc_free = ''.join(word for word in stop_free if ___ not in ____)
# Lemmatize all words
normalized = " ".join(____.____(word) for word in punc_free.split())
return normalized