String operators with the Twitter data
You continue working with the tweets
data where the text
column stores the content of each tweet.
Your task is to turn the text
column into a list of tokens. Then, using string operators, remove all non-alphabetic characters from the created list of tokens.
Este ejercicio forma parte del curso
Sentiment Analysis in Python
Instrucciones del ejercicio
- Import the word tokenizing function.
- Create word tokens from each tweet.
- Filter out all non-alphabetic characters from the created list, i.e. retain only letters.
Ejercicio interactivo práctico
Prueba este ejercicio y completa el código de muestra.
# Import the word tokenizing package
____
# Tokenize the text column
word_tokens = [____(review) for review in tweets.text]
print('Original tokens: ', word_tokens[0])
# Filter out non-letter characters
cleaned_tokens = [[word for word in item if ____.____] for item in word_tokens]
print('Cleaned tokens: ', cleaned_tokens[0])