Reeepeated characters

Back to your sentiment analysis! Your next task is to replace elongated words that appear in the tweets. We define an elongated word as a word that contains a repeating character twice or more times. e.g. "Awesoooome".

Replacing those words is very important since a classifier will treat them as a different term from the source words lowering their frequency.

To find them, you will use capturing groups and reference them back using numbers. E.g \4.

If you want to find a match for Awesoooome. You first need to capture Awes. Then, match o and reference the same character back, and then, me.

The list sentiment_analysis, containing the text of three tweets, and the re module are loaded in your session. You can use print() to view the data in the IPython Shell.

Deze oefening maakt deel uit van de cursus

Regular Expressions in Python

Cursus bekijken

Oefeninstructies

Complete the regular expression to match an elongated word as described.
Search the elements in sentiment_analysis list to find out if they contain elongated words. Assign the result to match_elongated.
Assign the captured group number zero to the variable elongated_word.
Print the result contained in the variable elongated_word.

Praktische interactieve oefening

Probeer deze oefening eens door deze voorbeeldcode in te vullen.

# Complete the regex to match an elongated word
regex_elongated = r"____(____)____\w*"

for tweet in sentiment_analysis:
	# Find if there is a match in each tweet 
	match_elongated = re.____(____, ____)
    
	if match_elongated:
		# Assign the captured group zero 
		elongated_word = match_elongated.____(____)
        
		# Complete the format method to print the word
		print("Elongated word found: {____}".format(word=____))
	else:
		print("No elongated word found")

Code bewerken en uitvoeren

Deze oefening maakt deel uit van de cursus

Regular Expressions in Python

SkillTag.level.beginnerSkillTag.label

4.8+

Begin de cursus gratis

Start your journey into the regular expression world! From slicing and concatenating, adjusting the case, removing spaces, to finding and replacing strings. You will learn how to master basic operation for string manipulation using a movie review dataset.

Exercise 1: Introduction to string manipulation Exercise 2: First day!Exercise 3: Artificial reviews Exercise 4: Palindromes Exercise 5: String operations Exercise 6: Normalizing reviews Exercise 7: Time to join!Exercise 8: Split lines or split the line?Exercise 9: Finding and replacing Exercise 10: Finding a substring Exercise 11: Where's the word?Exercise 12: Replacing negations

Following your journey, you will learn the main approaches that can be used to format or interpolate strings in python using a dataset containing information scraped from the web. You will explore the advantages and disadvantages of using positional formatting, embedding expressing inside string constants, and using the Template class.

Exercise 1: Positional formatting Exercise 2: Put it in order!Exercise 3: Calling by its name Exercise 4: What day is today?Exercise 5: Formatted string literal Exercise 6: Literally formatting Exercise 7: Make this function Exercise 8: On time Exercise 9: Template method Exercise 10: Preparing a report Exercise 11: Identifying prices Exercise 12: Playing safe

Time to discover the fundamental concepts of regular expressions! In this key chapter, you will learn to understand the basic concepts of regular expression syntax. Using a real dataset with tweets meant for sentiment analysis, you will learn how to apply pattern matching using normal and special characters, and greedy and lazy quantifiers.

Exercise 1: Introduction to regular expressions Exercise 2: Are they bots?Exercise 3: Find the numbers Exercise 4: Match and split Exercise 5: Repetitions Exercise 6: Everything clean Exercise 7: Some time ago Exercise 8: Getting tokens Exercise 9: Regex metacharacters Exercise 10: Finding files Exercise 11: Give me your email Exercise 12: Invalid password Exercise 13: Greedy vs. non-greedy matching Exercise 14: Understanding the difference Exercise 15: Greedy matching Exercise 16: Lazy approach

In the last step of your journey, you will learn more complex methods of pattern matching using parentheses to group strings together or to match the same text as matched previously. Also, you will get an idea of how you can look around expressions.

Exercise 1: Capturing groups Exercise 2: Try another name Exercise 3: Flying home Exercise 4: Alternation and non-capturing groups Exercise 5: Love it!Exercise 6: Ugh! Not for me!Exercise 7: Backreferences Exercise 8: Parsing PDF files Exercise 9: Close the tag, please!Exercise 10: Reeepeated characters

Huidige oefening

Exercise 11: Lookaround Exercise 12: Surrounding words Exercise 13: Filtering phone numbers Exercise 14: Finishing line