Everything clean

Back to your Twitter sentiment analysis project! There are several types of strings that increase your sentiment analysis complexity. But these strings do not provide any useful sentiment. Among them, we can have links and user mentions.

In order to clean the tweets, you want to extract some examples first. You know that most of the times links start with http and do not contain any whitespace, e.g. https://www.datacamp.com. User mentions start with @ and can have letters and numbers only, e.g. @johnsmith3.

You write down some helpful quantifiers to help you: * zero or more times, + once or more, ? zero or once.

The list sentiment_analysis containing the text of three tweets are already loaded in your session. You can use print() to view the data in the IPython Shell.

This exercise is part of the course

Regular Expressions in Python

Exercise instructions

Import the re module.
Write a regex to find all the matches of http links appearing in each tweet in sentiment_analysis. Print out the result.
Write a regex to find all the matches of user mentions appearing in each tweet in sentiment_analysis. Print out the result.

Hands-on interactive exercise

Have a go at this exercise by completing this sample code.

# Import re module
____

for tweet in sentiment_analysis:
	# Write regex to match http links and print out result
	print(re.____(____"____", ____))

	# Write regex to match user mentions and print out result
	print(re.____(____"____", ____))

Edit and Run Code

This exercise is part of the course

Regular Expressions in Python

BeginnerSkill Level

4.8+

Start Course for Free

Start your journey into the regular expression world! From slicing and concatenating, adjusting the case, removing spaces, to finding and replacing strings. You will learn how to master basic operation for string manipulation using a movie review dataset.

Exercise 1: Introduction to string manipulation Exercise 2: First day!Exercise 3: Artificial reviews Exercise 4: Palindromes Exercise 5: String operations Exercise 6: Normalizing reviews Exercise 7: Time to join!Exercise 8: Split lines or split the line?Exercise 9: Finding and replacing Exercise 10: Finding a substring Exercise 11: Where's the word?Exercise 12: Replacing negations

Following your journey, you will learn the main approaches that can be used to format or interpolate strings in python using a dataset containing information scraped from the web. You will explore the advantages and disadvantages of using positional formatting, embedding expressing inside string constants, and using the Template class.

Exercise 1: Positional formatting Exercise 2: Put it in order!Exercise 3: Calling by its name Exercise 4: What day is today?Exercise 5: Formatted string literal Exercise 6: Literally formatting Exercise 7: Make this function Exercise 8: On time Exercise 9: Template method Exercise 10: Preparing a report Exercise 11: Identifying prices Exercise 12: Playing safe

Time to discover the fundamental concepts of regular expressions! In this key chapter, you will learn to understand the basic concepts of regular expression syntax. Using a real dataset with tweets meant for sentiment analysis, you will learn how to apply pattern matching using normal and special characters, and greedy and lazy quantifiers.

Exercise 1: Introduction to regular expressions Exercise 2: Are they bots?Exercise 3: Find the numbers Exercise 4: Match and split Exercise 5: Repetitions Exercise 6: Everything clean

Current Exercise

Exercise 7: Some time ago Exercise 8: Getting tokens Exercise 9: Regex metacharacters Exercise 10: Finding files Exercise 11: Give me your email Exercise 12: Invalid password Exercise 13: Greedy vs. non-greedy matching Exercise 14: Understanding the difference Exercise 15: Greedy matching Exercise 16: Lazy approach

In the last step of your journey, you will learn more complex methods of pattern matching using parentheses to group strings together or to match the same text as matched previously. Also, you will get an idea of how you can look around expressions.

Exercise 1: Capturing groups Exercise 2: Try another name Exercise 3: Flying home Exercise 4: Alternation and non-capturing groups Exercise 5: Love it!Exercise 6: Ugh! Not for me!Exercise 7: Backreferences Exercise 8: Parsing PDF files Exercise 9: Close the tag, please!Exercise 10: Reeepeated characters Exercise 11: Lookaround Exercise 12: Surrounding words Exercise 13: Filtering phone numbers Exercise 14: Finishing line