Bắt đầu ngayBắt đầu miễn phí

Làm sạch dữ liệu văn bản

Giờ bạn đã xác định được stopwords và dấu câu, hãy dùng chúng để làm sạch thêm email enron trong dataframe df. Các danh sách chứa stopwords và dấu câu có sẵn trong stopexclude. Vẫn còn vài bước nữa trước khi dữ liệu thực sự sạch, như "lemmatization" từ và stemming động từ. Các động từ trong dữ liệu email đã được stemming, và việc lemmatization cũng đã được thực hiện sẵn cho bạn trong bài tập này.

Bài tập này là một phần của khóa học

Phát hiện gian lận với Python

Xem khóa học

Bài tập tương tác thực hành trực tiếp

Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.

# Import the lemmatizer from nltk
from nltk.stem.wordnet import WordNetLemmatizer
lemma = WordNetLemmatizer()

# Define word cleaning function
def clean(text, stop):
    text = text.____()
	# Remove stopwords
    stop_free = " ".join([word for word in text.lower().split() if ((___ not in ___) and (not word.isdigit()))])
	# Remove punctuations
    punc_free = ''.join(word for word in stop_free if ___ not in ____)
	# Lemmatize all words
    normalized = " ".join(____.____(word) for word in punc_free.split())      
    return normalized
Chỉnh sửa và Chạy Mã