Khám phá vector văn bản, phần 1
Hãy mở rộng phương pháp khám phá vector văn bản vừa học, sử dụng các vector tf/idf của cột title trong tập dữ liệu volunteer. Ở phần đầu của bài khám phá vector văn bản này, chúng ta sẽ bổ sung vào hàm đã học trên các slide. Hàm sẽ trả về một danh sách số. Ở bài tập tiếp theo, bạn sẽ viết một hàm khác để thu thập các từ xuất hiện nhiều nhất trên tất cả tài liệu, trích xuất chúng, rồi dùng danh sách đó để lọc vector text_tfidf.
Bài tập này là một phần của khóa học
Tiền xử lý cho Machine Learning bằng Python
Hướng dẫn bài tập
- Thêm các tham số
original_vocab(ứng vớitfidf_vec.vocabulary_) vàtop_n. - Gọi
pd.Series()trên đối tượng dictionary đã zip. Cách này giúp bạn thao tác dễ hơn. - Dùng hàm
.sort_values()để sắp xếp Series và cắt phần index đếntop_ntừ. - Gọi hàm, đặt
original_vocab=tfidf_vec.vocabulary_, đặtvector_index=8để lấy hàng thứ 9, và đặttop_n=3để lấy 3 từ có trọng số cao nhất.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Add in the rest of the arguments
def return_weights(vocab, ____, vector, vector_index, ____):
zipped = dict(zip(vector[vector_index].indices, vector[vector_index].data))
# Transform that zipped dict into a series
zipped_series = ____({vocab[i]:zipped[i] for i in vector[vector_index].indices})
# Sort the series to pull out the top n weighted words
zipped_index = zipped_series.____(ascending=False)[:____].index
return [original_vocab[i] for i in zipped_index]
# Print out the weighted words
print(return_weights(vocab, ____, text_tfidf, ____, ____))