EntityRuler với nhiều pattern trong spaCy
EntityRuler cho phép bạn thêm thực thể vào doc.ents và cải thiện hiệu suất nhận diện thực thể có tên. Trong bài tập này, bạn sẽ luyện cách thêm một thành phần EntityRuler vào pipeline nlp hiện có để đảm bảo nhiều thực thể được phân loại chính xác.
Mô hình en_core_web_sm đã được nạp sẵn và có sẵn dưới tên nlp. Bạn có thể truy cập văn bản ví dụ trong example_text và dùng nlp và doc để lần lượt truy cập mô hình spaCy và đối tượng Doc chứa example_text.
Bài tập này là một phần của khóa học
Xử lý ngôn ngữ tự nhiên với spaCy
Hướng dẫn bài tập
- In danh sách các tuple gồm văn bản thực thể và kiểu của chúng trong
example_textbằng mô hìnhnlp. - Định nghĩa nhiều pattern để khớp
brothervàsisters(chữ thường) với nhãnPERSON. - Thêm một thành phần
EntityRulervào pipelinenlpvà thêmpatternsvàoEntityRuler. - In các tuple gồm văn bản và kiểu của thực thể cho
example_textbằng mô hìnhnlp.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
nlp = spacy.load("en_core_web_md")
# Print a list of tuples of entities text and types in the example_text
print("Before EntityRuler: ", [____ for ____ in nlp(____).____], "\n")
# Define pattern to add a label PERSON for lower cased sisters and brother entities
patterns = [{"label": ____, "pattern": [{"lower": ____}]},
{"label": ____, "pattern": [{"lower": ____}]}]
# Add an EntityRuler component and add the patterns to the ruler
ruler = nlp.____("entity_ruler")
ruler.____(____)
# Print a list of tuples of entities text and types
print("After EntityRuler: ", [(ent.____, ent.____) for ent in nlp(example_text).____])