เริ่มต้นใช้งานเริ่มต้นใช้งานได้ฟรี

Cosine Similarity Matrix ของ Corpus

ในแบบฝึกหัดนี้ มี corpus ให้พร้อมแล้ว ซึ่งเป็นลิสต์ที่ประกอบด้วยประโยค 5 ประโยค โดย corpus จะแสดงผลในคอนโซล ให้คำนวณ cosine similarity matrix ที่เก็บค่า cosine similarity แบบ pairwise สำหรับทุกคู่ประโยค (โดยแปลงเป็นเวกเตอร์ด้วย tf-idf)

ขอให้จำไว้ว่า ค่าในแถวที่ i และคอลัมน์ที่ j ของ similarity matrix แทนค่าคะแนนความคล้ายคลึงระหว่างเวกเตอร์ที่ i และเวกเตอร์ที่ j

แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร

Feature Engineering for NLP in Python

ดูคอร์ส

คำแนะนำการฝึกหัด

  • สร้าง instance ของ TfidfVectorizer แล้วตั้งชื่อว่า tfidf_vectorizer
  • ใช้ fit_transform() เพื่อสร้างเวกเตอร์ tf-idf สำหรับ corpus แล้วตั้งชื่อว่า tfidf_matrix
  • ใช้ cosine_similarity() โดยส่ง tfidf_matrix เข้าไปเพื่อคำนวณ cosine similarity matrix ชื่อ cosine_sim

แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ

ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์

# Initialize an instance of tf-idf Vectorizer
tfidf_vectorizer = ____

# Generate the tf-idf vectors for the corpus
tfidf_matrix = tfidf_vectorizer.fit_transform(____)

# Compute and print the cosine similarity matrix
cosine_sim = ____(____, tfidf_matrix)
print(cosine_sim)
แก้ไขและรันโค้ด