Intro to word stemming and stem completion
Still, another useful preprocessing step involves word-stemming and stem completion. Word stemming reduces words to unify across documents. For example, the stem of "computational", "computers" and "computation" is "comput". But because "comput" isn't a real word, we want to reconstruct the words so that "computational", "computers", and "computation" all refer to a recognizable word, such as "computer". The reconstruction step is called stem completion.
The tm package provides the stemDocument() function to get to a word's root. This function either takes in a character vector and returns a character vector, or takes in a PlainTextDocument and returns a PlainTextDocument.
For example,
stemDocument(c("computational", "computers", "computation"))
returns "comput" "comput" "comput".
You will use stemCompletion() to reconstruct these word roots back into a known term. stemCompletion() accepts a character vector and a completion dictionary. The completion dictionary can be a character vector or a Corpus object. Either way, the completion dictionary for our example would need to contain the word "computer," so all instances of "comput" can be reconstructed.
이 연습은 강의의 일부입니다
Text Mining with Bag-of-Words in R
연습 안내
- Create a vector called
complicateconsisting of the words "complicated", "complication", and "complicatedly" in that order. - Store the stemmed version of
complicateto an object calledstem_doc. - Create
comp_dictthat contains one word, "complicate". - Create
complete_textby applyingstemCompletion()tostem_doc. Re-complete the words usingcomp_dictas the reference corpus. - Print
complete_textto the console.
실습형 인터랙티브 연습
이 예제를 이 샘플 코드를 완성하여 풀어보세요.
# Create complicate
complicate <- ___
# Perform word stemming: stem_doc
stem_doc <- ___
# Create the completion dictionary: comp_dict
comp_dict <- ___
# Perform stem completion: complete_text
complete_text <- ___
# Print complete_text
complete_text