การเพิ่ม pipe ใน spaCy
โดยทั่วไปเราจะใช้โมเดล spaCy ที่มีอยู่แล้วสำหรับงาน NLP ต่าง ๆ อย่างไรก็ตาม ในบางกรณี pipeline component สำเร็จรูป เช่น การแบ่งประโยค (sentence segmentation) อาจใช้เวลานานกว่าจะได้ผลลัพธ์ตามที่ต้องการ แบบฝึกหัดนี้จะฝึกการเพิ่ม pipeline component เข้าไปในโมเดล spaCy (text processing pipeline)
จะใช้รีวิวห้ารายการแรกจากชุดข้อมูล Amazon Fine Food Reviews โดยเข้าถึงรีวิวเหล่านี้ผ่านสตริง texts
แพ็กเกจ spaCy ถูก import ไว้ให้แล้ว
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
การประมวลผลภาษาธรรมชาติด้วย spaCy
คำแนะนำการฝึกหัด
- โหลดโมเดล
spaCyภาษาอังกฤษแบบว่างเปล่า แล้วเพิ่ม componentsentencizerเข้าไปในโมเดล - สร้าง
Doccontainer สำหรับtextsจากนั้นสร้างลิสต์เพื่อเก็บsentencesของเอกสารนั้น และแสดงจำนวนประโยคทั้งหมด - แสดงรายการ token ในประโยคที่สองจากลิสต์
sentences
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Load a blank spaCy English model and add a sentencizer component
nlp = spacy.____("en")
nlp.____("sentencizer")
# Create Doc containers, store sentences and print its number of sentences
doc = ____
sentences = [____ for s in ____]
print("Number of sentences: ", len(____), "\n")
# Print the list of tokens in the second sentence
print("Second sentence tokens: ", [____ for ____ in sentences[1]])