Bag-of-words สำหรับชื่อหนังสือ
ขณะนี้ PyBooks มีรายการชื่อหนังสือที่ต้องเข้ารหัสเพื่อนำไปวิเคราะห์ต่อ ทีมข้อมูลเชื่อว่าโมเดล Bag of Words (BoW) น่าจะเป็นแนวทางที่เหมาะสมที่สุด
ได้นำเข้าแพ็กเกจต่อไปนี้ให้แล้ว: torch, torchtext
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Deep Learning สำหรับข้อความด้วย PyTorch
คำแนะนำการฝึกหัด
- นำเข้าคลาส
CountVectorizerเพื่อใช้งาน bag-of-words - สร้างออบเจกต์จากคลาสที่นำเข้ามา จากนั้นใช้ออบเจกต์นั้นแปลง
titlesให้เป็น matrix representation - ดึงข้อมูลและแสดงชื่อฟีเจอร์ 5 รายการแรกพร้อมชื่อหนังสือที่เข้ารหัสแล้วด้วยเมธอด
get_feature_names_out()
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Import from sklearn
from sklearn.feature_extraction.text import ____
titles = ['The Great Gatsby','To Kill a Mockingbird','1984','The Catcher in the Rye','The Hobbit', 'Great Expectations']
# Initialize Bag-of-words with the list of book titles
vectorizer = ____()
bow_encoded_titles = ____.fit_transform(____)
# Extract and print the first five features
print(vectorizer.____[:5])
print(bow_encoded_titles.toarray()[0, :5])