Bắt đầu ngayBắt đầu miễn phí

Khám phá vector văn bản, phần 1

Hãy mở rộng phương pháp khám phá vector văn bản vừa học, sử dụng các vector tf/idf của cột title trong tập dữ liệu volunteer. Ở phần đầu của bài khám phá vector văn bản này, chúng ta sẽ bổ sung vào hàm đã học trên các slide. Hàm sẽ trả về một danh sách số. Ở bài tập tiếp theo, bạn sẽ viết một hàm khác để thu thập các từ xuất hiện nhiều nhất trên tất cả tài liệu, trích xuất chúng, rồi dùng danh sách đó để lọc vector text_tfidf.

Bài tập này là một phần của khóa học

Tiền xử lý cho Machine Learning bằng Python

Xem khóa học

Hướng dẫn bài tập

  • Thêm các tham số original_vocab (ứng với tfidf_vec.vocabulary_) và top_n.
  • Gọi pd.Series() trên đối tượng dictionary đã zip. Cách này giúp bạn thao tác dễ hơn.
  • Dùng hàm .sort_values() để sắp xếp Series và cắt phần index đến top_n từ.
  • Gọi hàm, đặt original_vocab=tfidf_vec.vocabulary_, đặt vector_index=8 để lấy hàng thứ 9, và đặt top_n=3 để lấy 3 từ có trọng số cao nhất.

Bài tập tương tác thực hành trực tiếp

Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.

# Add in the rest of the arguments
def return_weights(vocab, ____, vector, vector_index, ____):
    zipped = dict(zip(vector[vector_index].indices, vector[vector_index].data))
    
    # Transform that zipped dict into a series
    zipped_series = ____({vocab[i]:zipped[i] for i in vector[vector_index].indices})
    
    # Sort the series to pull out the top n weighted words
    zipped_index = zipped_series.____(ascending=False)[:____].index
    return [original_vocab[i] for i in zipped_index]

# Print out the weighted words
print(return_weights(vocab, ____, text_tfidf, ____, ____))
Chỉnh sửa và Chạy Mã