벡터를 VCorpus 객체로 만들기 (2)
이제 벡터를 Source 객체로 변환했으니, 이를 또 다른 tm 함수인 VCorpus()에 전달해 휘발성 코퍼스(volatile corpus)를 만들 거예요. 꽤 간단하죠?
VCorpus 객체는 중첩 리스트, 즉 리스트의 리스트입니다. VCorpus 객체의 각 인덱스에는 실제 텍스트 데이터(content)와 해당 메타데이터(meta)를 담은 리스트인 PlainTextDocument 객체가 들어 있어요. 전체 구조를 이해하는 데에는 VCorpus 객체를 시각화해 보는 것도 도움이 됩니다.
단일 문서 객체(10번째)를 확인하려면 이중 대괄호로 서브셋팅하세요.
coffee_corpus[[10]]
실제 ‘텍스트’를 보려면 리스트에 두 번 인덱싱합니다. 타임스탬프 같은 문서의 메타데이터에 접근하려면 [1]을 [2]로 바꾸면 됩니다. 또 다른 방법으로는 두 번째 대괄호가 필요 없는 content() 함수를 사용하는 것입니다.
coffee_corpus[[10]][1]
content(coffee_corpus[[10]])
이 연습은 강의의 일부입니다
R로 배우는 Bag-of-Words 텍스트 마이닝
연습 안내
coffee_source객체에VCorpus()함수를 호출해coffee_corpus를 생성하세요.- 콘솔에 출력해
coffee_corpus가VCorpus객체인지 확인하세요. - 이중 대괄호 서브셋팅을 사용해
coffee_corpus의 15번째 요소를 콘솔에 출력하고, 이 요소가 15번째 트윗의 콘텐츠와 메타데이터를 담은PlainTextDocument인지 확인하세요. coffee_corpus에서 15번째 트윗의 콘텐츠를 출력하세요. 해당 트윗을 이중 대괄호로 선택한 뒤, 단일 대괄호로 그 트윗의 콘텐츠를 추출하세요.coffee_corpus의 10번째 트윗에 대해content()결과를 출력하세요.
실습형 인터랙티브 연습
이 예제를 이 샘플 코드를 완성하여 풀어보세요.
## coffee_source is already in your workspace
# Make a volatile corpus from coffee_source
coffee_corpus <- ___
# Print out coffee_corpus
___
# Print the 15th tweet in coffee_corpus
___
# Print the contents of the 15th tweet in coffee_corpus
___
# Now use content to review the plain text of the 10th tweet
___(___[[___]])