Tokenization with spaCy

In this exercise, you'll practice tokenizing text. You'll use the first review from the Amazon Fine Food Reviews dataset for this exercise. You can access this review by using the text object provided.

The en_core_web_sm model is already loaded for you. You can access it by calling nlp(). You can use list comprehension to compile output lists.

This exercise is part of the course

Natural Language Processing with spaCy

Exercise instructions

Store Doc container for the pre-loaded review in a document object.
Store and review texts of all the tokens of the document in the variable first_text_tokens.

Hands-on interactive exercise

Have a go at this exercise by completing this sample code.

# Create a Doc container of the given text
document = ____(____)
    
# Store and review the token text values of tokens for the Doc container
first_text_tokens = [____ for ____ in ____]
print("First text tokens:\n", first_text_tokens, "\n")

Edit and Run Code

This exercise is part of the course

Natural Language Processing with spaCy

IntermediateSkill Level

4.8+

Start Course for Free

This chapter will introduce you to NLP, some of its use cases such as named-entity recognition and AI-powered chatbots. You’ll learn how to use the powerful spaCy library to perform various natural language processing tasks such as tokenization, sentence segmentation, POS tagging, and named entity recognition.

Exercise 1: Natural Language Processing (NLP) basics Exercise 2: Doc container in spaCy Exercise 3: NER use case Exercise 4: Tokenization with spaCy

Current Exercise

Exercise 5: spaCy basics Exercise 6: Running a spaCy pipeline Exercise 7: Lemmatization with spaCy Exercise 8: Sentence segmentation with spaCy Exercise 9: Linguistic features in spaCy Exercise 10: POS tagging with spaCy Exercise 11: NER with spaCy Exercise 12: Text processing with spaCy

Learn about linguistic features, word vectors, semantic similarity, analogies, and word vector operations. In this chapter you’ll discover how to use spaCy to extract word vectors, categorize texts that are relevant to a given topic and find semantically similar terms to given words from a corpus or from a spaCy model vocabulary.

Exercise 1: Linguistic features Exercise 2: Linguistic annotations in spaCy Exercise 3: Word-sense disambiguation with spaCy Exercise 4: Dependency parsing with spaCy Exercise 5: Introduction to word vectors Exercise 6: spaCy vocabulary Exercise 7: Word vectors in spaCy vocabulary Exercise 8: Word vectors and spaCy Exercise 9: Analogies and vector operations Exercise 10: Word vectors projection Exercise 11: Similar words in a vocabulary Exercise 12: Measuring semantic similarity with spaCy Exercise 13: Doc similarity with spaCy Exercise 14: Span similarity with spaCy Exercise 15: Semantic similarity for categorizing text

Get familiar with spaCy pipeline components, how to add a pipeline component, and analyze the NLP pipeline. You will also learn about multiple approaches for rule-based information extraction using EntityRuler, Matcher, and PhraseMatcher classes in spaCy and RegEx Python package.

Exercise 1: spaCy pipelines Exercise 2: Adding pipes in spaCy Exercise 3: Analyzing pipelines in spaCy Exercise 4: spaCy EntityRuler Exercise 5: EntityRuler with blank spaCy model Exercise 6: EntityRuler for NER Exercise 7: EntityRuler with multi-patterns in spaCy Exercise 8: RegEx with spaCy Exercise 9: RegEx in Python Exercise 10: RegEx with EntityRuler in spaCy Exercise 11: spaCy Matcher and PhraseMatcher Exercise 12: Matching a single term in spaCy Exercise 13: PhraseMatcher in spaCy Exercise 14: Matching with extended syntax in spaCy

Explore multiple real-world use cases where spaCy models may fail and learn how to train them further to improve model performance. You’ll be introduced to spaCy training steps and understand how to train an existing spaCy model or from scratch, and evaluate the model at the inference time.

Exercise 1: Customizing spaCy models Exercise 2: Training spaCy models Exercise 3: Model performance on your data Exercise 4: spaCy training data format Exercise 5: Training steps Exercise 6: Annotation and preparing training data Exercise 7: Compatible training data Exercise 8: Training with spaCy Exercise 9: Training preparation steps Exercise 10: Train an existing NER model Exercise 11: Training a spaCy model from scratch Exercise 12: Wrap-up