Belmont Stakes 的結果符合常態分配嗎?
自 1926 年起,Belmont Stakes 是一場長度 1.5 英里的 3 歲純種馬賽事。1973 年,Secretariat 創下史上最快的 Belmont Stakes 紀錄。雖然那是最快的一年,但 1970 年因為天候異常潮濕、賽道泥濘而是最慢的一年。將這兩個離群值自資料集中移除後,請計算 Belmont 冠軍時間的平均數與標準差。接著用 rng.normal(),以這個平均數與標準差自常態分配取樣,並繪製 CDF。再把冠軍時間的 ECDF 疊加在同一張圖上。這些結果是否接近常態分配?
(註)Justin 是從 Belmont 的維基百科頁面擷取與 Belmont Stakes 有關的資料。
本練習屬於課程
Statistical Thinking in Python (Part 1)
練習說明
- 使用移除兩個離群值後的 Belmont 冠軍時間來計算平均數與標準差。這些資料已在 NumPy 陣列
belmont_no_outliers中。 - 使用
rng.normal(),以該平均數與標準差自常態分配取 10,000 個樣本。 - 分別計算理論樣本的 CDF 與 Belmont 冠軍資料的 ECDF,並將結果分別指定給
x_theor, y_theor與x, y。 - 按下 Submit 以將你的樣本 CDF 與 ECDF 一起繪圖,替座標軸加上標籤並顯示圖表。
動手互動練習
試著完成這個範例程式碼,體驗一下這個練習。
# Compute mean and standard deviation: mu, sigma
# Sample out of a normal distribution with this mu and sigma: samples
# Get the CDF of the samples and of the data
# Plot the CDFs and show the plot
_ = plt.plot(x_theor, y_theor)
_ = plt.plot(x, y, marker='.', linestyle='none')
_ = plt.xlabel('Belmont winning time (sec.)')
_ = plt.ylabel('CDF')
plt.show()