การจับคู่ด้วย syntax แบบขยายใน spaCy
การดึงข้อมูลแบบอิงกฎเป็นส่วนสำคัญของ NLP pipeline ทุกประเภท คลาส Matcher ช่วยให้กำหนด pattern ได้ยืดหยุ่นยิ่งขึ้น โดยรองรับ operator ภายในวงเล็บปีกกา ซึ่งใช้สำหรับการเปรียบเทียบแบบขยาย และมีรูปแบบคล้ายกับ operator in, not in และ operator เปรียบเทียบของ Python ในแบบฝึกหัดนี้ จะได้ฝึกใช้ฟังก์ชัน Matcher ของ spaCy เพื่อค้นหาคำที่ตรงกับเงื่อนไขที่กำหนดจากข้อความตัวอย่าง
คลาส Matcher ถูก import จากไลบรารี spacy.matcher ไว้เรียบร้อยแล้ว และในแบบฝึกหัดนี้จะมีคอนเทนเนอร์ Doc ของข้อความตัวอย่างให้เรียกใช้ผ่านตัวแปร doc พร้อมกับโมเดล spaCy ที่โหลดไว้ล่วงหน้าผ่านตัวแปร nlp
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
การประมวลผลภาษาธรรมชาติด้วย spaCy
คำแนะนำการฝึกหัด
- สร้าง matcher object โดยใช้
Matcherและnlp - ใช้ operator
INเพื่อกำหนด pattern สำหรับจับคู่กับtiny squaresและtiny mouthful - ใช้ pattern ดังกล่าวเพื่อค้นหาการจับคู่จาก
doc - แสดงค่า index ของ token เริ่มต้นและสิ้นสุด พร้อมกับ text span ของผลลัพธ์ที่จับคู่ได้
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
nlp = spacy.load("en_core_web_sm")
doc = nlp(example_text)
# Define a matcher object
matcher = Matcher(nlp.____)
# Define a pattern to match tiny squares and tiny mouthful
pattern = [{"lower": ____}, {"lower": {____: ["squares", "mouthful"]}}]
# Add the pattern to matcher object and find matches
matcher.____("CustomMatcher", [____])
matches = ____(____)
# Print out start and end token indices and the matched text span per match
for match_id, start, end in matches:
print("Start token: ", ____, " | End token: ", ____, "| Matched text: ", doc[____:____].text)