Understanding the difference

You need to keep working and cleaning your tweets dataset. You realize that there are some HTML tags present. You need to remove them but keep the inside content as they are useful for analysis.

Let's take a look at this sentence containing an HTML tag:

I want to see that <strong>amazing show</strong> again!.

You know that to get the HTML tag you need to match anything that sits inside angle brackets < >. But the biggest problem is that the closing tag has the same structure. If you match too much, you will end up removing key information. So you need to decide whether to use a greedy or a lazy quantifier.

The string is already loaded as string to your session.

Diese Übung ist Teil des Kurses

Regular Expressions in Python

Anleitung zur Übung

Import the re module.
Write a regex expression to replace HTML tags with an empty string.
Print out the result.

Interaktive Übung

Vervollständige den Beispielcode, um diese Übung erfolgreich abzuschließen.

# Import re
____

# Write a regex to eliminate tags
string_notags = re.____(r"____", "____", ____)

# Print out the result
____

Code bearbeiten und ausführen

Diese Übung ist Teil des Kurses

Regular Expressions in Python

Geringe SchwierigkeitSchwierigkeitsgrad

4.8+

Kurs kostenlos starten

Start your journey into the regular expression world! From slicing and concatenating, adjusting the case, removing spaces, to finding and replacing strings. You will learn how to master basic operation for string manipulation using a movie review dataset.

Exercise 1: Introduction to string manipulation Exercise 2: First day!Exercise 3: Artificial reviews Exercise 4: Palindromes Exercise 5: String operations Exercise 6: Normalizing reviews Exercise 7: Time to join!Exercise 8: Split lines or split the line?Exercise 9: Finding and replacing Exercise 10: Finding a substring Exercise 11: Where's the word?Exercise 12: Replacing negations

Following your journey, you will learn the main approaches that can be used to format or interpolate strings in python using a dataset containing information scraped from the web. You will explore the advantages and disadvantages of using positional formatting, embedding expressing inside string constants, and using the Template class.

Exercise 1: Positional formatting Exercise 2: Put it in order!Exercise 3: Calling by its name Exercise 4: What day is today?Exercise 5: Formatted string literal Exercise 6: Literally formatting Exercise 7: Make this function Exercise 8: On time Exercise 9: Template method Exercise 10: Preparing a report Exercise 11: Identifying prices Exercise 12: Playing safe

Time to discover the fundamental concepts of regular expressions! In this key chapter, you will learn to understand the basic concepts of regular expression syntax. Using a real dataset with tweets meant for sentiment analysis, you will learn how to apply pattern matching using normal and special characters, and greedy and lazy quantifiers.

Exercise 1: Introduction to regular expressions Exercise 2: Are they bots?Exercise 3: Find the numbers Exercise 4: Match and split Exercise 5: Repetitions Exercise 6: Everything clean Exercise 7: Some time ago Exercise 8: Getting tokens Exercise 9: Regex metacharacters Exercise 10: Finding files Exercise 11: Give me your email Exercise 12: Invalid password Exercise 13: Greedy vs. non-greedy matching Exercise 14: Understanding the difference

Aktuelle Übung

Exercise 15: Greedy matching Exercise 16: Lazy approach

In the last step of your journey, you will learn more complex methods of pattern matching using parentheses to group strings together or to match the same text as matched previously. Also, you will get an idea of how you can look around expressions.

Exercise 1: Capturing groups Exercise 2: Try another name Exercise 3: Flying home Exercise 4: Alternation and non-capturing groups Exercise 5: Love it!Exercise 6: Ugh! Not for me!Exercise 7: Backreferences Exercise 8: Parsing PDF files Exercise 9: Close the tag, please!Exercise 10: Reeepeated characters Exercise 11: Lookaround Exercise 12: Surrounding words Exercise 13: Filtering phone numbers Exercise 14: Finishing line