Chapter 1 of Introduction to Natural Langauge Processing prepares you for running your first analysis on text. You will explore regular expressions and tokenization, two of the most common components of most analysis tasks. With regular expressions, you can search for any pattern you can think of, and with tokenization, you can prepare and clean text for more sophisticated analysis. This chapter is necessary for tackling the techniques we will learn in the remaining chapters of this course.

Regular expression basics

Practicing syntax with grep

Exploring regular expression functions.

Tokenization

tidytext functions

Tokenization: sentences

Text cleaning basics

Text preprocessing: remove stop words

Text preprocessing: Stemming

True Fundamentals

In this chapter, you will learn the most common and studied ways to analyze text. You will look at creating a text corpus, expanding a bag-of-words representation into a TFIDF matrix, and use cosine-similarity metrics to determine how similar two pieces of text are to each other.  You build on your foundations for practicing NLP before you dive into applications of NLP in chapters 3 and  4. 

Understanding an R corpus

Explore an R corpus

Creating a tibble from a corpus

Creating a corpus

The bag-of-words representation

Practice BoW

BoW Example

Sparse matrices

The TFIDF

Manual calculations

TFIDF Practice

Cosine Similarity

An example of failing at text analysis

Cosine similarity example

Representations of Text

Chapter 3 focuses on two common text analysis approaches, classification modeling, and topic modeling. If you are working on text analysis projects, you will inevitably use one or both of these methods. This chapter teaches you how to perform both techniques and provides insight into how to approach these techniques from a practical point of you.

Preparing text for modeling

Data preparation

Removing sparse terms

Classification modeling

Classification modeling example

Confusion matrices

TFIDF tibble vs dtm

Introduction to topic modeling

LDA practice

Assigning topics to documents

LDA in practice

Testing perplexity

Reviewing LDA results

Applications: Classification and Topic Modeling

In chapter 4 we cover two staples of natural language processing, sentiment analysis, and word embeddings. These are two analysis techniques that are a must for anyone learning the fundamentals of text analysis. Furthermore, you will briefly learn about BERT, part-of-speech tagging, and named entity recognition. Almost 15 different analysis techniques were covered in this course, so chapter 4 ends by recapping all of the great techniques you will learn about in this course. 

Sentiment analysis

tidytext lexicons

Sentiment scores

Sentiment and emotion

Word embeddings

h2o practice

word2vec

Additional NLP analysis

Reviewing methods #1

Review methods #2

Conclusion

Advanced Techniques

Animal Farm

Russian Troll tweets

As with any fundamentals course, Introduction to Natural Language Processing in R is designed to equip you with the necessary tools to begin your adventures in analyzing text. Natural language processing (NLP) is a constantly growing field in data science, with some very exciting advancements over the last decade. This course will cover the basics of these topics and prepare you for expanding your analysis capabilities. We dive into regular expressions, topic modeling, named entity recognition, and others, all while providing thorough examples that can be used to kick start your future analysis.

Intermediate R

Introduction to the Tidyverse

Discover basics skills and tools for natural language processing (NLP) in R, like regular expressions, topic modeling, named entity recognition and more.

Introduction to Natural Language Processing in R

Gain an overview of all the skills and tools needed to excel in Natural Language Processing in R.

Research Data Scientist

Bag-of-word pitfalls

Sharing common words

Tacos matter

TFIDF

IDF Equation

TF + IDF

Calculating the TFIDF matrix

bind_tf_idf output

https://s3.amazonaws.com/assets.datacamp.com/production/course_19730/subtitles/course_19730_8395c87cacf280760694a0caac2a0cd1.vtt

Representations of Text - The TFIDF

The TFIDF

Create Your Free Account