Xem lại dữ liệu bán buôn: Khám phá
Từ phân tích trước đó, bạn đã thấy rằng k = 2 có độ rộng silhouette trung bình cao nhất. Trong bài tập này, bạn sẽ tiếp tục phân tích dữ liệu khách hàng bán buôn bằng cách xây dựng và khám phá một mô hình kmeans với 2 cụm.
Bài tập này là một phần của khóa học
Phân cụm bằng R
Hướng dẫn bài tập
- Xây dựng mô hình k-means tên
model_customerscho dữ liệucustomers_spendbằng hàmkmeans()vớicenters = 2. - Trích xuất vector gán cụm từ mô hình
model_customers$clustervà lưu vào biếnclust_customers. - Gắn các nhãn cụm như một cột
clustervào data framecustomers_spendvà lưu kết quả thành một data frame mới tênsegment_customers. - Tính kích thước của mỗi cụm bằng
count().
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
set.seed(42)
# Build a k-means model for the customers_spend with a k of 2
model_customers <- ___
# Extract the vector of cluster assignments from the model
clust_customers <- ___
# Build the segment_customers data frame
segment_customers <- mutate(___, cluster = ___)
# Calculate the size of each cluster
count(___, ___)
# Calculate the mean for each category
segment_customers %>%
group_by(cluster) %>%
summarise_all(list(mean))