เริ่มต้นใช้งานเริ่มต้นใช้งานได้ฟรี

การจับคู่ด้วย syntax แบบขยายใน spaCy

การดึงข้อมูลแบบอิงกฎเป็นส่วนสำคัญของ NLP pipeline ทุกประเภท คลาส Matcher ช่วยให้กำหนด pattern ได้ยืดหยุ่นยิ่งขึ้น โดยรองรับ operator ภายในวงเล็บปีกกา ซึ่งใช้สำหรับการเปรียบเทียบแบบขยาย และมีรูปแบบคล้ายกับ operator in, not in และ operator เปรียบเทียบของ Python ในแบบฝึกหัดนี้ จะได้ฝึกใช้ฟังก์ชัน Matcher ของ spaCy เพื่อค้นหาคำที่ตรงกับเงื่อนไขที่กำหนดจากข้อความตัวอย่าง

คลาส Matcher ถูก import จากไลบรารี spacy.matcher ไว้เรียบร้อยแล้ว และในแบบฝึกหัดนี้จะมีคอนเทนเนอร์ Doc ของข้อความตัวอย่างให้เรียกใช้ผ่านตัวแปร doc พร้อมกับโมเดล spaCy ที่โหลดไว้ล่วงหน้าผ่านตัวแปร nlp

แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร

การประมวลผลภาษาธรรมชาติด้วย spaCy

ดูคอร์ส

คำแนะนำการฝึกหัด

  • สร้าง matcher object โดยใช้ Matcher และ nlp
  • ใช้ operator IN เพื่อกำหนด pattern สำหรับจับคู่กับ tiny squares และ tiny mouthful
  • ใช้ pattern ดังกล่าวเพื่อค้นหาการจับคู่จาก doc
  • แสดงค่า index ของ token เริ่มต้นและสิ้นสุด พร้อมกับ text span ของผลลัพธ์ที่จับคู่ได้

แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ

ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์

nlp = spacy.load("en_core_web_sm")
doc = nlp(example_text)

# Define a matcher object
matcher = Matcher(nlp.____)
# Define a pattern to match tiny squares and tiny mouthful
pattern = [{"lower": ____}, {"lower": {____: ["squares", "mouthful"]}}]

# Add the pattern to matcher object and find matches
matcher.____("CustomMatcher", [____])
matches = ____(____)

# Print out start and end token indices and the matched text span per match
for match_id, start, end in matches:
    print("Start token: ", ____, " | End token: ", ____, "| Matched text: ", doc[____:____].text)
แก้ไขและรันโค้ด