शुरू करेंमुफ़्त में शुरू करें

क्या Belmont Stakes के नतीजे Normal रूप से distributed हैं?

1926 से, Belmont Stakes 3-वर्षीय thoroughbred घोड़ों की 1.5 माइल लंबी रेस है. 1973 में Secretariat ने इतिहास का सबसे तेज़ Belmont Stakes दौड़ा. वह सबसे तेज़ साल था, जबकि 1970 सबसे धीमा रहा क्योंकि हालात असामान्य रूप से गीले और फिसलन भरे थे. इन दो outliers को डेटा सेट से हटाकर, Belmont winners के समयों का mean और standard deviation निकालें. इसी mean और standard deviation के साथ rng.normal() फंक्शन से Normal distribution से sample लें और एक CDF प्लॉट करें. जीतने वाले Belmont समयों की ECDF को उसी पर overlay करें. क्या ये लगभग Normal रूप से distributed लगते हैं?

नोट: Justin ने Belmont Stakes से संबंधित डेटा Belmont Wikipedia पेज से स्क्रैप किया था.

यह अभ्यास पाठ्यक्रम का हिस्सा है

Statistical Thinking in Python (Part 1)

पाठ्यक्रम देखें

अभ्यास निर्देश

  • दो outliers हटाने के बाद Belmont winners के समयों का mean और standard deviation निकालें. ये डेटा NumPy array belmont_no_outliers में है.
  • rng.normal() का उपयोग करके इसी mean और standard deviation के साथ 10,000 samples Normal distribution से लें.
  • Theoretical samples की CDF और Belmont winners के डेटा की ECDF निकालें, और नतीजों को क्रमशः x_theor, y_theor और x, y में असाइन करें.
  • CDF और ECDF को प्लॉट करने, axes को label करने और प्लॉट दिखाने के लिए "Submit Answer" दबाएँ.

इंटरैक्टिव व्यावहारिक अभ्यास

इस अभ्यास को इस नमूना कोड को पूरा करके आज़माएँ।

# Compute mean and standard deviation: mu, sigma



# Sample out of a normal distribution with this mu and sigma: samples


# Get the CDF of the samples and of the data



# Plot the CDFs and show the plot
_ = plt.plot(x_theor, y_theor)
_ = plt.plot(x, y, marker='.', linestyle='none')
_ = plt.xlabel('Belmont winning time (sec.)')
_ = plt.ylabel('CDF')
plt.show()
कोड संपादित करें और चलाएँ