Đánh giá & So sánh thuật toán
Bây giờ khi chúng ta đã tạo một mô hình mới với GBTRegressor, đã đến lúc so sánh nó với mô hình chuẩn RandomForestRegressor. Để làm điều này, chúng ta sẽ so sánh dự đoán của cả hai mô hình với dữ liệu thực tế và tính RMSE và R^2.
Bài tập này là một phần của khóa học
Feature Engineering với PySpark
Hướng dẫn bài tập
- Import
RegressionEvaluatortừpyspark.ml.evaluationđể dùng sau. - Khởi tạo
RegressionEvaluatorbằng cách đặtlabelCollà dữ liệu thực tếSALESCLOSEPRICEvàpredictionCollà dữ liệu dự đoánPrediction_Price. - Để tính các chỉ số, gọi
evaluatetrênevaluatorvới giá trị dự đoánpredsvà tạo một dictionary với khóaevaluator.metricNamevà giá trịrmse; làm tương tự cho chỉ sốr2.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
from ____ import ____
# Select columns to compute test error
evaluator = ____(____=____,
____=____)
# Dictionary of model predictions to loop over
models = {'Gradient Boosted Trees': gbt_predictions, 'Random Forest Regression': rfr_predictions}
for key, preds in models.items():
# Create evaluation metrics
rmse = evaluator.____(____, {____: ____})
r2 = evaluator.____(____, {____: ____})
# Print Model Metrics
print(key + ' RMSE: ' + str(rmse))
print(key + ' R^2: ' + str(r2))