NER ด้วย spaCy
Named entity recognition (NER) ช่วยให้ระบุองค์ประกอบสำคัญในเอกสารได้อย่างรวดเร็ว เช่น ชื่อบุคคลและสถานที่ ฟีเจอร์นี้ช่วยจัดระเบียบข้อมูลที่ไม่มีโครงสร้าง และดึงข้อมูลที่สำคัญออกมา ซึ่งมีประโยชน์มากเมื่อต้องจัดการกับชุดข้อมูลขนาดใหญ่ ในแบบฝึกหัดนี้ จะได้ฝึกใช้งาน Named Entity Recognition
โหลด en_core_web_sm ให้เป็น nlp ไว้แล้ว และมีคอมเมนต์จากชุดข้อมูล Airline Travel Information System (ATIS) จำนวน 3 รายการอยู่ใน list ชื่อ texts
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
การประมวลผลภาษาธรรมชาติด้วย spaCy
คำแนะนำการฝึกหัด
- สร้าง
documentsซึ่งเป็น list ของDoccontainer สำหรับแต่ละข้อความในtextsโดยใช้ list comprehension - วนลูปผ่าน
doc.entsเพื่อพิมพ์ข้อความและ label ของแต่ละ entity ในdoccontainer แต่ละตัว - พิมพ์ข้อความของ token ตัวที่ 6 และประเภท entity ของ
Doccontainer ตัวที่ 2
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Compile a list of all Doc containers of texts
documents = [____ for text in texts]
# Print the entity text and label for the entities in each document
for doc in documents:
print([(____, ____) for ent in ____])
# Print the 6th token's text and entity type of the second document
print("\nText:", documents[1][5].____, "| Entity type: ", documents[1][5].____)