การจับคู่คำเดี่ยวใน spaCy
รูปแบบ RegEx อ่าน เขียน และดีบักได้ยากพอสมควร แต่ spaCy มีทางเลือกที่อ่านง่ายและพร้อมใช้งานในระดับ production อย่าง Matcher class ซึ่งสามารถจับคู่กฎที่กำหนดไว้ล่วงหน้ากับลำดับ token ใน Doc container ได้ ในแบบฝึกหัดนี้จะได้ฝึกใช้ Matcher เพื่อค้นหาคำเดียว
สามารถเข้าถึงข้อความตัวอย่างได้จาก example_text และใช้ nlp กับ doc เพื่อเข้าถึงโมเดล spaCy และ Doc container ของ example_text ตามลำดับ
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
การประมวลผลภาษาธรรมชาติด้วย spaCy
คำแนะนำการฝึกหัด
- สร้าง
Matcherclass - กำหนด pattern เพื่อจับคู่คำว่า
witchในรูปตัวพิมพ์เล็กในexample_text - เพิ่ม pattern เข้าสู่
Matcherclass และค้นหาคำที่ตรงกัน - วนซ้ำผ่านผลลัพธ์ที่ตรงกัน แล้วพิมพ์ index token เริ่มต้นและสิ้นสุด รวมถึง span ของข้อความที่จับคู่ได้
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
nlp = spacy.load("en_core_web_sm")
doc = nlp(example_text)
# Initialize a Matcher object
matcher = Matcher(nlp.____)
# Define a pattern to match lower cased word witch
pattern = [{"lower" : ____}]
# Add the pattern to matcher object and find matches
matcher.add("CustomMatcher", [____])
matches = matcher(____)
# Print start and end token indices and span of the matched text
for match_id, start, end in matches:
print("Start token: ", ____, " | End token: ", ____, "| Matched text: ", doc[____:____].text)