Langere n-grams gebruiken

Tot nu toe heb je features gemaakt op basis van losse woorden in elke tekst. Dat kan behoorlijk krachtig zijn in een Machine Learning-model, maar misschien maak je je zorgen dat door woorden afzonderlijk te bekijken veel context verloren gaat. Om dit aan te pakken kun je bij het bouwen van modellen n-grams gebruiken: reeksen van n woorden die samen gegroepeerd zijn. Bijvoorbeeld:

bigrams: reeksen van twee opeenvolgende woorden
trigrams: reeksen van drie opeenvolgende woorden

Deze kun je automatisch in je gegevensset laten maken door het argument ngram_range op te geven als een tuple (n1, n2), waarbij alle n-grams in het bereik van n1 tot en met n2 worden meegenomen.

Deze oefening maakt deel uit van de cursus

Feature engineering voor Machine Learning in Python

Cursus bekijken

Oefeninstructies

Importeer CountVectorizer uit sklearn.feature_extraction.text.
Maak een instantie van CountVectorizer die alleen trigrams meeneemt.
Fit de vectorizer en pas hem in één stap toe op de kolom text_clean.
Print de featurenamen die door de vectorizer zijn gegenereerd.

Praktische interactieve oefening

Probeer deze oefening eens door deze voorbeeldcode in te vullen.

# Import CountVectorizer
from sklearn.feature_extraction.text import ____

# Instantiate a trigram vectorizer
cv_trigram_vec = CountVectorizer(max_features=100, 
                                 stop_words='english', 
                                 ____)

# Fit and apply trigram vectorizer
cv_trigram = ____(speech_df['text_clean'])

# Print the trigram features
print(cv_trigram_vec.____)

Code bewerken en uitvoeren

Deze oefening maakt deel uit van de cursus

Feature engineering voor Machine Learning in Python

SkillTag.level.intermediateSkillTag.label

4.8+

Begin de cursus gratis

In this chapter, you will explore what feature engineering is and how to get started with applying it to real-world data. You will load, explore and visualize a survey response dataset, and in doing so you will learn about its underlying data types and why they have an influence on how you should engineer your features. Using the pandas package you will create new features from both categorical and continuous columns.

Exercise 1: Why generate features?Exercise 2: Getting to know your data Exercise 3: Selecting specific data types Exercise 4: Dealing with categorical features Exercise 5: One-hot encoding and dummy variables Exercise 6: Dealing with uncommon categories Exercise 7: Numeric variables Exercise 8: Binarizing columns Exercise 9: Binning values

This chapter introduces you to the reality of messy and incomplete data. You will learn how to find where your data has missing values and explore multiple approaches on how to deal with them. You will also use string manipulation techniques to deal with unwanted characters in your dataset.

Exercise 1: Why do missing values exist?Exercise 2: How sparse is my data?Exercise 3: Finding the missing values Exercise 4: Dealing with missing values (I)Exercise 5: Listwise deletion Exercise 6: Replacing missing values with constants Exercise 7: Dealing with missing values (II)Exercise 8: Filling continuous missing values Exercise 9: Imputing values in predictive models Exercise 10: Dealing with other data issues Exercise 11: Dealing with stray characters (I)Exercise 12: Dealing with stray characters (II)Exercise 13: Method chaining

In this chapter, you will focus on analyzing the underlying distribution of your data and whether it will impact your machine learning pipeline. You will learn how to deal with skewed data and situations where outliers may be negatively impacting your analysis.

Exercise 1: Data distributions Exercise 2: What does your data look like? (I)Exercise 3: What does your data look like? (II)Exercise 4: When don't you have to transform your data?Exercise 5: Scaling and transformations Exercise 6: Normalization Exercise 7: Standardization Exercise 8: Log transformation Exercise 9: When can you use normalization?Exercise 10: Removing outliers Exercise 11: Percentage based outlier removal Exercise 12: Statistical outlier removal Exercise 13: Scaling and transforming new data Exercise 14: Train and testing transformations (I)Exercise 15: Train and testing transformations (II)

Finally, in this chapter, you will work with unstructured text data, understanding ways in which you can engineer columnar features out of a text corpus. You will compare how different approaches may impact how much context is being extracted from a text, and how to balance the need for context, without too many features being created.

Exercise 1: Tekst encoderen Exercise 2: Je tekst opschonen Exercise 3: Hoogwaardige tekstkenmerken Exercise 4: Woordtellingen Exercise 5: Woorden tellen (I)Exercise 6: Woorden tellen (II)Exercise 7: Je features beperken Exercise 8: Tekst naar DataFrame Exercise 9: Term frequency-inverse document frequency Exercise 10: Tf-idf Exercise 11: Tf-idf-waarden inspecteren Exercise 12: Ongeziene data transformeren Exercise 13: N-grammen Exercise 14: Langere n-grams gebruiken

Huidige oefening

Exercise 15: De meest voorkomende woorden vinden Exercise 16: Afronding