Cosine Similarity Matrix ของ Corpus
ในแบบฝึกหัดนี้ มี corpus ให้พร้อมแล้ว ซึ่งเป็นลิสต์ที่ประกอบด้วยประโยค 5 ประโยค โดย corpus จะแสดงผลในคอนโซล ให้คำนวณ cosine similarity matrix ที่เก็บค่า cosine similarity แบบ pairwise สำหรับทุกคู่ประโยค (โดยแปลงเป็นเวกเตอร์ด้วย tf-idf)
ขอให้จำไว้ว่า ค่าในแถวที่ i และคอลัมน์ที่ j ของ similarity matrix แทนค่าคะแนนความคล้ายคลึงระหว่างเวกเตอร์ที่ i และเวกเตอร์ที่ j
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Feature Engineering for NLP in Python
คำแนะนำการฝึกหัด
- สร้าง instance ของ
TfidfVectorizerแล้วตั้งชื่อว่าtfidf_vectorizer - ใช้
fit_transform()เพื่อสร้างเวกเตอร์ tf-idf สำหรับcorpusแล้วตั้งชื่อว่าtfidf_matrix - ใช้
cosine_similarity()โดยส่งtfidf_matrixเข้าไปเพื่อคำนวณ cosine similarity matrix ชื่อcosine_sim
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Initialize an instance of tf-idf Vectorizer
tfidf_vectorizer = ____
# Generate the tf-idf vectors for the corpus
tfidf_matrix = tfidf_vectorizer.fit_transform(____)
# Compute and print the cosine similarity matrix
cosine_sim = ____(____, tfidf_matrix)
print(cosine_sim)