Skip to main content
Workforce LibreTexts

12: The Mean

  • Page ID
    50857
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    In the previous chapter we saw how a sample can be used to estimate a population parameter. In this chapter we will learn how to describe data using measurements of center and spread, and how concepts such as the normal distribution and the Central Limit Theorem help us understand why estimates from different samples can differ.

    Properties of the Mean

    Calculation of the Mean

    The mean is also commonly called the average. It is found by summing all the data values in the dataset and dividing by the total number of data values.

    import numpy as np
    import pandas as pd
    import matplotlib.pyplot as plt
    
    array = np.array([12, 3, 23, 4, 15])
    print(sum(array)/len(array))
    print(np.mean(array))
    Output?
    11.4
    11.4
    

    When Data is 1 and 0, or True and False

    When the data sequence contains only 1's and 0's, then:

    • The sum is the number of 1's. Since adding 0 does not change the total, each 1 increases the sum by one. The final sum is simply the count of the 1's.
    • The mean is the proportion of the 1's. The mean is the sum divided by the total number of values. Since the sum is the number of 1's, the mean equals the fraction or proportion of values that are 1.
    array = np.array([1,0,1,1,1,0])
    print("Count of 1's:", sum(array))
    print("Proportion of 1's:", sum(array)/len(array))
    print("The mean:", np.mean(array))
    Output?
    Count of 1's: 4
    Proportion of 1's: 0.6666666666666666
    The mean: 0.6666666666666666
    

    For computers, True means 1 and False means 0. Therefore, when the data sequence contains True and False, we can calculate the count of True values and proportion of True values the same way.

    array = np.array([True, False, True, True, True, False])
    print("Count of True:", sum(array))
    print("Proportion of True:", sum(array)/len(array))
    print("The mean:", np.mean(array))
    Output?
    Count of True: 4
    Proportion of True: 0.6666666666666666
    The mean: 0.6666666666666666
    

    The Mean and the Distribution

    The mean can be calculated by multiplying each unique data value by its proportion in the dataset, and then adding the results.

    In an array with four values: 1, 2, 2, 6, the unique data values are 1, 2, and 6. Their proportions in the dataset are:

    • 1 appears once, so its proportion is 1/4.
    • 2 appears twice, so its proportion is 2/4.
    • 6 appears once, so its proportion is 1/4.
    array1 = np.array([1,2,2,6])
    print("The mean:", 1*1/4 + 2*2/4 + 6*1/4)
    print("The mean:", np.mean(array1))
    Output?
    The mean: 2.75
    The mean: 2.75
    

    In the following array2, the unique data values are also 1, 2, 6, and each value's proportion is the same as in array1:

    array2 = np.array([1,1,2,2,2,2,6,6])
    print("The mean:", 1*2/8 + 2*4/8 + 6*2/8)
    print("The mean:", np.mean(array2))
    Output?
    The mean: 2.75
    The mean: 2.75
    

    The above examples show that if two arrays have the same distribution (same proportion for each unique data value), then they have the same mean.

    We recall from Chapter 5 that a density histogram scales the bars so that they represent proportions instead of counts. This makes it easy to compare the distributions of datasets even when they contain different numbers of data values.

    We plot the density distribution of array1 and array2:

    plt.figure(figsize=(4,3))
    plt.hist(array2, bins=np.arange(1,7), density=True, edgecolor='black', alpha=0.5)
    plt.hist(array1, bins=np.arange(1,7), density=True, edgecolor='black', alpha=0.5)
    plt.title("Density Distribution of array1 and array2")
    plt.xlabel("Value")
    
    # You don't need to write the following code.
    # The code draws a triangle at the mean.
    triangle_x = [2.75 - 0.1, 2.75 + 0.1, 2.75]
    triangle_y = [-0.03,-0.03, 0]
    plt.fill(triangle_x, triangle_y, color='black', zorder=3)
    plt.ylim(bottom=-0.04)
    
    plt.show()
    Output?

    It looks like only one distribution was plotted, but actually both array1 and array2 distributions were plotted. Since they're the same distribution, the two histograms overlap each other perfectly. You can turn one of the plt.hist() lines of code into a comment so it doesn't run, and you'll see that the plot above is made of a blue histogram and an orange histogram that overlap each other.

    For the histogram above, imagine that the x-axis is a straight board or a plank, and the three bars rest on top of the plank. As shown, the plank is balanced on a fulcrum, which is the black triangle.

    • If the fulcrum is in the middle of the plank, which is at 3.5, the plank will tilt left because there's more "weight" on the left side.
    • If the fulcrum is at 1.5, then the plank will tilt right.
    • If the fulcrum is at 2.75 as shown, which is where the mean is, then the plank will be balanced.

    The mean is the balance point of the histogram.

    The Mean and the Median

    Recall that the median is the midpoint of the sorted data values. If a distribution is symmetrical, then the mean and the median are the same.

    Below we create a symmetrical distribution, where the left and right side are balanced.

    array = np.array([1,2,2,3,3,3,4,4,5])
    print("The mean:", np.mean(array))
    print("The median:", np.median(array))
    Output?
    The mean: 3.0
    The median: 3.0
    
    plt.figure(figsize=(4,3))
    plt.hist(array, bins=np.arange(1,7), density=True, edgecolor='black')
    plt.title("Distribution of a Symmetrical Array")
    plt.grid()
    plt.plot()
    Output?
    []

    When the distribution is not symmetric, or is skewed, so that one side extends farther than the other side (also called having a tail), then the mean is pulled away from the median towards the tail.

    Below we create an array with a right hand tail.

    array = np.array([1,2,2,3,3,3,4,4,4,4,5,5,6,6,7,8,12,15,16,20])
    print("The mean:", np.mean(array))
    print("The median:", np.median(array))
    Output?
    The mean: 6.5
    The median: 4.5
    
    plt.figure(figsize=(4,3))
    plt.hist(array, bins=np.arange(1,22), density=True, edgecolor='black')
    plt.title("Distribution of Skewed Array")
    plt.scatter(np.mean(array), 0, color='red', s=40, zorder=2)
    plt.scatter(np.median(array), 0, color='yellow', s=40, zorder=2)
    plt.grid()
    plt.plot()
    Output?
    []

    The histogram is skewed right and has a right hand tail, therefore the red mean is to the right of the yellow median.

    The mean is affected by the tail. As values in the right tail become larger, they pull the mean to the right, making the mean larger. In contrast, the median is determined by the middle value, so it is not as affected by a few extreme values in the tail.

    For skewed distributions with a long tail, we often use the median instead of the mean to describe the center of the data. In the distribution above, most of the data values are between 2 and 6, so the median of 4.5 is a better description of the typical value than the mean of 6.5, which is pulled to the right by the long tail.

    Measures of Spread

    We saw earlier that the mean is the balance point of the histogram, and that data spread out from the mean on both the left and right sides of the mean. In this section we learn to measure the spread of the data.

    We begin by finding the deviation of each data value from the mean. A deviation is simply the distance between a data value and the mean. The following steps show how to calculate the deviations.

    The Deviation

    We find the difference between the mean and each data point. This difference is called the deviation from the mean. The deviation is positive if the data value is greater than the mean, and negative if the data is less than the mean.

    # create an array of 4 numbers and find the mean
    
    array = np.array([12, 5, 9, 3])
    # find the mean
    the_mean = np.mean(array)
    print("The mean:", the_mean)
    Output?
    The mean: 7.25
    
    # find the deviations
    
    # subtracts the mean from each value in array, 
    # creating an array of differences
    deviations = array - np.mean(array)
    
    print("The deviations:", deviations)
    Output?
    The deviations: [ 4.75 -2.25  1.75 -4.25]
    

    Note that if we add all the deviations, the result is always 0. This makes sense because the mean is the balance point for the histogram. The sum of distances on the left of the mean should be equal to the sum of distances on the right of the mean.

    print(sum(deviations))
    Output?
    0.0
    

    The Variance

    Our goal is to measure the spread of the data, or how far on average the data values are from the mean. To calculate the average, we need to add all the distances of data from the mean and divide by the number of data values. However, we see from above that adding all the deviations results in 0, so this approach isn't useful to measure the spread.

    Instead, we square the deviations and then find the average of the squared deviations. We square the deviations to remove the negative signs so the sum would not be 0.

    The variance is the mean of the squared deviations.

    # square each deviation and findsthe mean of the squares
    variance = np.mean(deviations**2)
    
    print("The variance:", variance)
    Output?
    The variance: 12.1875
    

    We also note that squaring the deviations gives more weight to the data that are far from the mean. For example, a dataset of deviations [2, -2, 2, -2] will have a variance of (4 + 4 + 4 + 4)/4 = 4. And a dataset of deviations [-2, -2, -2, 6] will have a variance of (4 + 4 + 4 + 36)/4 = 12. One deviation increase from 2 to 6 causes a much larger increase in variance.

    As a result, a dataset with some values far from the center has a larger variance than a dataset where the values are all clustered close to the mean.

    The Standard Deviation

    Since we need to square the deviations in order to find the variance, we now take the square root to "undo" the squaring. The result is the standard deviation.

    sd = np.sqrt(variance)
    print("The standard deviation:", round(sd, 2))
    Output?
    The standard deviation: 3.49
    

    Lucky for us, numpy also has the std() function to calculate the standard deviation of an array, so we don't have to go through all three steps above when we want to find the standard deviation.

    print("The standard deviation:", round(np.std(array), 2))
    Output?
    The standard deviation: 3.49
    

    The standard deviation of 3.49 means that the data values are typically about 3.49 units away from the mean. The larger the standard deviation, the more spread out the data are.

    In the example above the data values are from 3 to 12. Since the standard deviation is 3.49, we can say that the data are fairly spread out.

    Typically a standard deviation is useful when comparing the spread of two datasets. A dataset with a standard deviation of 1.2 is less spread out than a dataset with a standard deviation of 3.9.

    Standard Deviation Example

    We use the CDC survey of people's health habits and status. We will look at the standard deviation, or SD, of the age column

    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/cdc.csv"
    health = pd.read_csv(url)
    health.head()
    Output?
      genhlth exerany hlthplan smoke100 height weight wtdesire age gender
    0 good 0 1 0 70 175 175 77 m
    1 good 0 1 1 64 125 115 33 f
    2 good 1 1 1 60 105 105 49 f
    3 good 1 1 0 66 132 124 42 f
    4 very good 0 1 0 61 150 130 55 f

    We now find the mean and standard deviation (SD) of the age.

    mean_age = np.mean(health["age"])
    sd_age = np.std(health["age"])
    print("Mean Age:", round(mean_age,2))
    print("Standard deviation:", round(sd_age,2))
    Output?
    Mean Age: 45.07
    Standard deviation: 17.19
    

    Then we look at the oldest age and compare it against the SD.

    # find the oldest age
    oldest = health["age"].max()
    print("Oldest age:", oldest)
    print("Number of standard deviations:", round((oldest - mean_age)/sd_age,2))
    Output?
    Oldest age: 99
    Number of standard deviations: 3.14
    

    The oldest age of 99 is 3 standard deviations from the mean age.

    We look at the age distribution and the 1 SD, +2 SD, and +3 SD lines.

    plt.figure(figsize=(6,3))
    plt.hist(health["age"], edgecolor='black')
    plt.axvline(mean_age, color='yellow', linestyle='-')
    plt.axvline(mean_age + sd_age, color='orange', linestyle='--')
    plt.axvline(mean_age - sd_age, color='orange', linestyle='--')
    plt.axvline(mean_age + 2*sd_age, color='red', linestyle='--')
    plt.axvline(mean_age + 3*sd_age, color='green', linestyle='--')
    plt.xlim(18, 100)
    plt.title("Age Distribution with Mean and SD Lines")
    plt.xlabel("Age")
    plt.ylabel("Count")
    plt.grid()
    plt.show()
    Output?

    We make the following observations from the histogram:

    • The age range is from 18 to 99.
    • The mean, or balance point of the distribution, is the yellow line.
    • The +/ − 1 SD lines are the orange lines.
    • The +2 SD line is red, the +3 SD line is green. We don't show the negative SD lines since they're outside of the range of data.
    • The oldest age of 99 is just outside of the +3 SD line.
    • Most ages appear to be within 2 standard deviations of the mean.

    Chebyshev's Theorem

    From the histogram, we observe that most of the ages are within 2 standard deviations of the mean. This is not a coincidence. Chebyshev's Theorem tells us that, for any dataset, most of the data values lie within a few standard deviations of the mean, regardless of the shape of the distribution.

    More specifically, Chebyshev's Theorem states that:

    • At least 75% of the data lie within 2 standard deviations of the mean.
    • At least 89% of the data lie within 3 standard deviations of the mean.

    Since the oldest age of 99 is just beyond 3 standard deviations from the mean, it is an unusually high age compared with the rest of the dataset.

    Standard Units or Z-Scores

    So far, we have described how far a data value is from the mean by saying it is 1 standard deviation, 2 standard deviations, or 3 standard deviations away. But there's a simpler way to describe the number of standard deviations as a single number. This number is the standard unit or z-score.

    A z-score tells us how many standard deviations a data value is above or below the mean. Positive z-scores represent values above the mean, while negative z-scores represent values below the mean.

    Examples of z-score values:

    • A value at the mean has a z-score of 0.
    • A value one standard deviation above the mean has a z-score of +1, or is 1 standard unit above the mean.
    • A value two standard deviations below the mean has a z-score of -2, or is 2 standard units below the mean.
    • The oldest age in our dataset is just over 3 standard deviations above the mean, so its z-score is just over +3 or 3 standard units above the mean.

    To find the number of standard units or z-score for a dataset:

    def standard_units(array):
        return (array - np.mean(array))/np.std(array)
    
    # in the function above, the mean is subtracted from each data value,
    # then the differences are divided by the SD to get the number of SDs,
    # and an array of number of SDs (z-score) is returned
    Output?

    Using the ages in the survey dataset above, we find the standard units for age:

    age_s_units = standard_units(health["age"])
    age_s_units.head()
    Output?
    0    1.857333
    1   -0.701958
    2    0.228693
    3   -0.178467
    4    0.577687
    Name: age, dtype: float64

    Then we add the age_s_units as a new column in the health DataFrame.

    health["age_su"] = age_s_units
    health.head()
    Output?
      genhlth exerany hlthplan smoke100 height weight wtdesire age gender age_su
    0 good 0 1 0 70 175 175 77 m 1.857333
    1 good 0 1 1 64 125 115 33 f -0.701958
    2 good 1 1 1 60 105 105 49 f 0.228693
    3 good 1 1 0 66 132 124 42 f -0.178467
    4 very good 0 1 0 61 150 130 55 f 0.577687

    Next we sort in descending order the DataFrame by the age_su column so we can see the longest delay times.

    print("10 oldest ages:")
    health.sort_values("age_su", ascending=False).head(10)
    Output?
    10 oldest ages:
    
      genhlth exerany hlthplan smoke100 height weight wtdesire age gender age_su
    899 very good 0 1 0 70 150 150 99 f 3.136979
    6709 excellent 1 1 1 62 120 115 99 f 3.136979
    10349 good 1 1 1 60 130 130 97 f 3.020647
    17050 good 1 1 1 64 128 128 96 f 2.962481
    11911 very good 0 1 0 62 110 110 95 f 2.904316
    16083 good 1 1 0 62 127 127 95 f 2.904316
    7085 very good 1 1 0 66 125 130 94 m 2.846150
    19502 fair 0 1 0 66 140 140 94 m 2.846150
    2527 very good 1 1 0 61 118 118 94 f 2.846150
    1443 poor 1 1 0 63 102 102 93 f 2.787984

    We see that there are probably quite a few people whose age is more than 2 standard units.

    Since Chebyshev's Theorem says that at least 75% of the data should be 2 standard units or less, we want to check if Chebyshev's bound holds true for this dataset.

    We find the proportion of people whose age is 2 standard units or less, then we find the proportion of people whose age is 3 standard units or less.

    proportion_2_su = len(health[np.abs(health["age_su"]) <= 2]) / len(health)
    print("Proportion of ages at 2 standard units or less:", round(proportion_2_su, 4))
    
    proportion_3_su = len(health[np.abs(health["age_su"]) <= 3]) / len(health)
    print("Proportion of ages at 3 standard units or less:", round(proportion_3_su, 4))
    Output?
    Proportion of ages at 2 standard units or less: 0.9693
    Proportion of ages at 3 standard units or less: 0.9999
    

    Chebyshev's Theorem does hold for the age data. In fact the theorem gives a conservative lower bound.

    • 97% of the ages are within 2 standard units, compared to Chebyshev's boundary of at least 75%.
    • Almost 100% of the ages are within 3 standard units, compared to Chebyshev's boundary of at least 89%.
      Looking at the first 10 rows of the sorted DataFrame, we see that only three out of 20,000 people (ages 99, 99, and 97) are above 3 standard units above the mean.

    Because the age distribution is concentrated near the mean, a much larger percentage of the data falls within 2 and 3 standard units than Chebyshev's Theorem guarantees. In the next section, we will see that for data that are approximately normally distributed, these percentages can be predicted even more precisely.

    The Normal Distribution and Standard Units

    The normal distribution is also known as the Gaussian distribution or the bell curve. It has a symmetric bell-shaped curve, with a peak at the mean value and tapering off equally on both left and right sides. The normal distribution has a special relationship with the standard deviation.

    We first create a normal distribution and fill in the areas under the normal curve that are within 1 standard deviation and within 2 standard deviations.

    # You don't need to study this code.
    # The code is to demo the standard deviations of the normal distribution.
    
    from scipy.stats import norm
    
    mean = 0
    sd = 1
    
    # create 1000 x values centered around 0
    x = np.linspace(mean - 3*sd, mean + 3*sd, 1000)
    # create y values that form the normal distribution
    y = norm.pdf(x, mean, sd)
    
    # create 2 side-by-side plots
    fig, axes = plt.subplots(1, 2, figsize=(12, 3))
    
    # plot normal distribution with filled area within 1 SD
    axes[0].plot(x, y)
    x_fill_1 = np.linspace(mean - sd, mean + sd, 1000)
    y_fill_1 = norm.pdf(x_fill_1, mean, sd)
    axes[0].fill_between(x_fill_1, y_fill_1, alpha=0.4)
    axes[0].set_title('Area Under Normal Curve Within 1 SD')
    axes[0].grid(True)
    
    # plot normal distribution with filled area within 2 SD
    axes[1].plot(x, y)
    x_fill_2 = np.linspace(mean - 2*sd, mean + 2*sd, 1000)
    y_fill_2 = norm.pdf(x_fill_2, mean, sd)
    axes[1].fill_between(x_fill_2, y_fill_2, alpha=0.4)
    axes[1].set_title('Area Under Normal Curve within 2 SDs')
    axes[1].grid(True)
    Output?

    We can see that the area under the curve that's within 2 SDs covers a substantial part of the entire area under the curve.

    To calculate the area under the curve, we use the cdf function of the norm module. The cdf or Cumulative Distribution Function returns the proportion of $$ \frac{ \text{area under the curve up to an x value}} { \text{total area under the curve}} $$

    To call the cdf function we use the format:
    area_up_to_x = norm.cdf(x)

    Using this function, we find the proportion of the area under the curve within 1 SD and within 2 SDs compared to the total area under the curve.

    # Area within 1 SD is between -1 and 1
    print("Proportion of area within 1 SD:", round(norm.cdf(1) - norm.cdf(-1), 3))
    
    # Area within 2 SD is between -2 and 2
    print("Proportion of area within 2 SDs:", round(norm.cdf(2) - norm.cdf(-2), 3))
    
    # Area within 3 SD is between -3 and 3
    print("Proportion of area within 3 SDs:", round(norm.cdf(3) - norm.cdf(-3), 3))
    Output?
    Proportion of area within 1 SD: 0.683
    Proportion of area within 2 SDs: 0.954
    Proportion of area within 3 SDs: 0.997
    

    We note that for a normal distribution:

    • About 95% of the area under the curve is within 2 standard deviations of the mean, and about 99.7% is within 3 standard deviations of the mean.
    • These percentages are much higher than Chebyshev's lower bounds of 75% and 89%. This is because Chebyshev's Theorem applies to any distribution, regardless of its shape, while the percentages above apply only to the normal distribution.

    Because the normal distribution has a specific bell shape, we can make much stronger statements about the proportion of data within 2 or 3 standard deviations than we can for an arbitrary distribution.

    The Central Limit Theorem

    The central limit theorem says that if we take a large number of random samples from a population and find the mean of each sample, then the distribution of the means will approach a normal distribution in shape, regardless of the original population's distribution.

    Just like with Chebyshev's Theorem, we will see how the central limit theorem works by investigating the food delivery dataset.

    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/delivery.csv"
    delivery = pd.read_csv(url)
    delivery.head()
    Output?
      Order_ID Distance_km Weather Traffic_Level Time_of_Day Vehicle_Type Preparation_Time_min Courier_Experience_yrs Delivery_Time_min
    0 522 7.93 Windy Low Afternoon Scooter 12 1.0 43
    1 738 16.42 Clear Medium Evening Bike 20 2.0 84
    2 741 9.52 Foggy Low Night Scooter 28 1.0 59
    3 661 7.44 Rainy Medium Afternoon Scooter 5 1.0 37
    4 412 19.03 Clear Low Morning Bike 16 5.0 68

    We look at the delivery distance Distance_km distribution.

    plt.figure(figsize=(6,4))
    plt.hist(delivery["Distance_km"], bins=30, edgecolor='black')
    plt.title("Distribution of Delivery Distances (km)")
    plt.xlabel("Distance_km")
    plt.ylabel("Count")
    plt.grid()
    plt.show()
    Output?

    The distribution of delivery distances is not a bell curve or normal distribution. The distances are spread fairly evenly from about 1 to 20 kilometers.

    We will repeatedly take random samples from this dataset and calculate the mean distance of each sample. According to the Central Limit Theorem, the distribution of these sample means will become approximately a normal distribution.

    We find the mean and standard deviation of the delivery distance.

    delivery_mean = np.mean(delivery["Distance_km"])
    delivery_standard_dev = np.std(delivery["Distance_km"])
    print("The mean:", round(delivery_mean,2))
    print("The SD:", round(delivery_standard_dev,2))
    Output?
    The mean: 10.06
    The SD: 5.69
    

    Next we take a large number of samples of the delivery times and plot the mean of the samples.

    We write a function to select a sample of num_of_data values with replacement and find the mean of the sample.

    def get_sample_mean(num_of_data):
        return np.mean(delivery["Distance_km"].sample(num_of_data, replace=True ))
    Output?

    We call the get_sample_mean 10,000 times to get 10,000 means of delivery time

    means = []
    for i in range(10000):
        # get the mean of a sample of 300 delivery times
        # and append to the means list
        means.append(get_sample_mean(300))
    # convert to a numpy array
    mean_array = np.array(means)
    Output?

    We plot the distribution of the sample means.

    plt.figure(figsize=(4,3))
    plt.hist(mean_array, edgecolor='black')
    plt.xlabel('Mean Distance')
    plt.title('Distribution of Sample Means')
    plt.grid()
    plt.show()
    Output?

    We see that the distribution of the sample means is approximately bell-shaped, as predicted by the central limit theorem. We also see that the distribution is centered around the mean, which we calculated above to be 10.06.

    We now repeat the same sampling steps as above, but our sample size is 1,000 instead of 300.

    means = []
    for i in range(10000):
        means.append(get_sample_mean(1000))
    mean_array = np.array(means)
    
    plt.figure(figsize=(4,3))
    plt.hist(mean_array, edgecolor='black')
    plt.xlabel('Mean Distance')
    plt.title('Distribution of Sample Means')
    plt.grid()
    plt.show()
    Output?

    We notice that once again the distribution of sample means is a normal distribution, just as predicted by the central limit theorem, and that the distribution is centered around the mean of 10.06.

    But we also notice that for samples of 1000 values, the distribution is narrower than the distribution for samples of 300 values. This makes sense because as the sample size increases, each individual data value (including the unusually high value or low value) has less influence on the sample mean. As a result, the sample mean varies less from sample to sample, and the means from repeated samples are closer together around the mean. This results in a narrower sampling distribution.

    The Standard Deviation of the Sampling Distribution

    The distribution of the sample means is also called the sampling distribution. The spread of the sampling distribution tells us how much the sample mean will vary from sample to sample.

    The standard deviation of the sampling distribution is: $$\text{SD of sampling distribution} = \frac{ \text{population SD}}{ \sqrt{ \text{sample size}}}$$

    The stanard deviation of the sampling distribution is often called the standard error of the mean, or simply the standard error.

    print("SD of sampling distibution for sample of 300:", round(np.std(delivery["Distance_km"]) / np.sqrt(300), 2))
    print("SD of sampling distribution for sample 1000:", round(np.std(delivery["Distance_km"]) / np.sqrt(1000),2))
    Output?
    SD of sampling distibution for sample of 300: 0.33
    SD of sampling distribution for sample 1000: 0.18
    

    The SD of the larger sample (1000) is smaller than the SD of the smaller sample (300). A smaller SD means the sampling distribution is narrower so the sample mean is more likely to be close to the population mean.

    As the equation above shows, the SD of the sampling distribution decreases in reverse proportion with $\sqrt{ \text{sample size}}$. To reduce the SD by a factor of 10, the sample size must increase by a factor of 100.

    Summary

    In this chapter we learned how to describe the center and spread of a dataset using the mean, median, variance, and standard deviation. We also learned how z-scores or standard units, Chebyshev's Theorem, the normal distribution, and the Central Limit Theorem help us understand the distribution of data and why larger samples produce more reliable estimates of the population mean.


      This page titled 12: The Mean was last modified on Fri, 25 Sep 2026 01:21:37 GMT and is shared under a CC BY 4.0 license and was authored, remixed, and/or curated by Clare Nguyen.

      • Was this article helpful?