Sử dụng pipeline
Giờ bạn đã có pipeline được định nghĩa, tức là kết hợp logistic regression với phương pháp SMOTE, hãy chạy nó trên dữ liệu. Bạn có thể coi pipeline như một mô hình machine learning duy nhất. Dữ liệu X và y đã được chuẩn bị, và pipeline được định nghĩa ở bài trước. Bạn có tò mò muốn biết kết quả mô hình không? Hãy thử xem nhé!
Bài tập này là một phần của khóa học
Phát hiện gian lận với Python
Hướng dẫn bài tập
- Chia dữ liệu 'X' và 'y' thành tập huấn luyện và tập kiểm tra. Dành 30% dữ liệu cho tập kiểm tra, và đặt
random_statebằng 0. - Huấn luyện pipeline trên dữ liệu huấn luyện và lấy dự đoán bằng cách chạy hàm
pipeline.predict()trên tập dữ liệuX_test.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Split your data X and y, into a training and a test set and fit the pipeline onto the training data
X_train, X_test, y_train, y_test = ____
# Fit your pipeline onto your training set and obtain predictions by fitting the model onto the test data
pipeline.fit(____, ____)
predicted = pipeline.____(____)
# Obtain the results from the classification report and confusion matrix
print('Classifcation report:\n', classification_report(y_test, predicted))
conf_mat = confusion_matrix(y_true=y_test, y_pred=predicted)
print('Confusion matrix:\n', conf_mat)