12: The Mean
- Page ID
- 50857
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)In the previous chapter we saw how a sample can be used to estimate a population parameter. In this chapter we will learn how to describe data using measurements of center and spread, and how concepts such as the normal distribution and the Central Limit Theorem help us understand why estimates from different samples can differ.
Properties of the Mean
Calculation of the Mean
The mean is also commonly called the average. It is found by summing all the data values in the dataset and dividing by the total number of data values.
11.4
11.4
When Data is 1 and 0, or True and False
When the data sequence contains only 1's and 0's, then:
- The sum is the number of 1's. Since adding 0 does not change the total, each 1 increases the sum by one. The final sum is simply the count of the 1's.
- The mean is the proportion of the 1's. The mean is the sum divided by the total number of values. Since the sum is the number of 1's, the mean equals the fraction or proportion of values that are 1.
Count of 1's: 4
Proportion of 1's: 0.6666666666666666
The mean: 0.6666666666666666
For computers, True means 1 and False means 0. Therefore, when the data sequence contains True and False, we can calculate the count of True values and proportion of True values the same way.
Count of True: 4
Proportion of True: 0.6666666666666666
The mean: 0.6666666666666666
The Mean and the Distribution
The mean can be calculated by multiplying each unique data value by its proportion in the dataset, and then adding the results.
In an array with four values: 1, 2, 2, 6, the unique data values are 1, 2, and 6. Their proportions in the dataset are:
- 1 appears once, so its proportion is 1/4.
- 2 appears twice, so its proportion is 2/4.
- 6 appears once, so its proportion is 1/4.
The mean: 2.75
The mean: 2.75
In the following array2, the unique data values are also 1, 2, 6, and each value's proportion is the same as in array1:
The mean: 2.75
The mean: 2.75
The above examples show that if two arrays have the same distribution (same proportion for each unique data value), then they have the same mean.
We recall from Chapter 5 that a density histogram scales the bars so that they represent proportions instead of counts. This makes it easy to compare the distributions of datasets even when they contain different numbers of data values.
We plot the density distribution of array1 and array2:
It looks like only one distribution was plotted, but actually both array1 and array2 distributions were plotted. Since they're the same distribution, the two histograms overlap each other perfectly. You can turn one of the plt.hist() lines of code into a comment so it doesn't run, and you'll see that the plot above is made of a blue histogram and an orange histogram that overlap each other.
For the histogram above, imagine that the x-axis is a straight board or a plank, and the three bars rest on top of the plank. As shown, the plank is balanced on a fulcrum, which is the black triangle.
- If the fulcrum is in the middle of the plank, which is at 3.5, the plank will tilt left because there's more "weight" on the left side.
- If the fulcrum is at 1.5, then the plank will tilt right.
- If the fulcrum is at 2.75 as shown, which is where the mean is, then the plank will be balanced.
The mean is the balance point of the histogram.
The Mean and the Median
Recall that the median is the midpoint of the sorted data values. If a distribution is symmetrical, then the mean and the median are the same.
Below we create a symmetrical distribution, where the left and right side are balanced.
The mean: 3.0
The median: 3.0
[]
When the distribution is not symmetric, or is skewed, so that one side extends farther than the other side (also called having a tail), then the mean is pulled away from the median towards the tail.
Below we create an array with a right hand tail.
The mean: 6.5
The median: 4.5
[]
The histogram is skewed right and has a right hand tail, therefore the red mean is to the right of the yellow median.
The mean is affected by the tail. As values in the right tail become larger, they pull the mean to the right, making the mean larger. In contrast, the median is determined by the middle value, so it is not as affected by a few extreme values in the tail.
For skewed distributions with a long tail, we often use the median instead of the mean to describe the center of the data. In the distribution above, most of the data values are between 2 and 6, so the median of 4.5 is a better description of the typical value than the mean of 6.5, which is pulled to the right by the long tail.
Measures of Spread
We saw earlier that the mean is the balance point of the histogram, and that data spread out from the mean on both the left and right sides of the mean. In this section we learn to measure the spread of the data.
We begin by finding the deviation of each data value from the mean. A deviation is simply the distance between a data value and the mean. The following steps show how to calculate the deviations.
The Deviation
We find the difference between the mean and each data point. This difference is called the deviation from the mean. The deviation is positive if the data value is greater than the mean, and negative if the data is less than the mean.
The mean: 7.25
The deviations: [ 4.75 -2.25 1.75 -4.25]
Note that if we add all the deviations, the result is always 0. This makes sense because the mean is the balance point for the histogram. The sum of distances on the left of the mean should be equal to the sum of distances on the right of the mean.
0.0
The Variance
Our goal is to measure the spread of the data, or how far on average the data values are from the mean. To calculate the average, we need to add all the distances of data from the mean and divide by the number of data values. However, we see from above that adding all the deviations results in 0, so this approach isn't useful to measure the spread.
Instead, we square the deviations and then find the average of the squared deviations. We square the deviations to remove the negative signs so the sum would not be 0.
The variance is the mean of the squared deviations.
The variance: 12.1875
We also note that squaring the deviations gives more weight to the data that are far from the mean. For example, a dataset of deviations [2, -2, 2, -2] will have a variance of (4 + 4 + 4 + 4)/4 = 4. And a dataset of deviations [-2, -2, -2, 6] will have a variance of (4 + 4 + 4 + 36)/4 = 12. One deviation increase from 2 to 6 causes a much larger increase in variance.
As a result, a dataset with some values far from the center has a larger variance than a dataset where the values are all clustered close to the mean.
The Standard Deviation
Since we need to square the deviations in order to find the variance, we now take the square root to "undo" the squaring. The result is the standard deviation.
The standard deviation: 3.49
Lucky for us, numpy also has the std() function to calculate the standard deviation of an array, so we don't have to go through all three steps above when we want to find the standard deviation.
The standard deviation: 3.49
The standard deviation of 3.49 means that the data values are typically about 3.49 units away from the mean. The larger the standard deviation, the more spread out the data are.
In the example above the data values are from 3 to 12. Since the standard deviation is 3.49, we can say that the data are fairly spread out.
Typically a standard deviation is useful when comparing the spread of two datasets. A dataset with a standard deviation of 1.2 is less spread out than a dataset with a standard deviation of 3.9.
Standard Deviation Example
We use the CDC survey of people's health habits and status. We will look at the standard deviation, or SD, of the age column
| genhlth | exerany | hlthplan | smoke100 | height | weight | wtdesire | age | gender | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | good | 0 | 1 | 0 | 70 | 175 | 175 | 77 | m |
| 1 | good | 0 | 1 | 1 | 64 | 125 | 115 | 33 | f |
| 2 | good | 1 | 1 | 1 | 60 | 105 | 105 | 49 | f |
| 3 | good | 1 | 1 | 0 | 66 | 132 | 124 | 42 | f |
| 4 | very good | 0 | 1 | 0 | 61 | 150 | 130 | 55 | f |
We now find the mean and standard deviation (SD) of the age.
Mean Age: 45.07
Standard deviation: 17.19
Then we look at the oldest age and compare it against the SD.
Oldest age: 99
Number of standard deviations: 3.14
The oldest age of 99 is 3 standard deviations from the mean age.
We look at the age distribution and the 1 SD, +2 SD, and +3 SD lines.
We make the following observations from the histogram:
- The age range is from 18 to 99.
- The mean, or balance point of the distribution, is the yellow line.
- The +/ − 1 SD lines are the orange lines.
- The +2 SD line is red, the +3 SD line is green. We don't show the negative SD lines since they're outside of the range of data.
- The oldest age of 99 is just outside of the +3 SD line.
- Most ages appear to be within 2 standard deviations of the mean.
Chebyshev's Theorem
From the histogram, we observe that most of the ages are within 2 standard deviations of the mean. This is not a coincidence. Chebyshev's Theorem tells us that, for any dataset, most of the data values lie within a few standard deviations of the mean, regardless of the shape of the distribution.
More specifically, Chebyshev's Theorem states that:
- At least 75% of the data lie within 2 standard deviations of the mean.
- At least 89% of the data lie within 3 standard deviations of the mean.
Since the oldest age of 99 is just beyond 3 standard deviations from the mean, it is an unusually high age compared with the rest of the dataset.
Standard Units or Z-Scores
So far, we have described how far a data value is from the mean by saying it is 1 standard deviation, 2 standard deviations, or 3 standard deviations away. But there's a simpler way to describe the number of standard deviations as a single number. This number is the standard unit or z-score.
A z-score tells us how many standard deviations a data value is above or below the mean. Positive z-scores represent values above the mean, while negative z-scores represent values below the mean.
Examples of z-score values:
- A value at the mean has a z-score of 0.
- A value one standard deviation above the mean has a z-score of +1, or is 1 standard unit above the mean.
- A value two standard deviations below the mean has a z-score of -2, or is 2 standard units below the mean.
- The oldest age in our dataset is just over 3 standard deviations above the mean, so its z-score is just over +3 or 3 standard units above the mean.
To find the number of standard units or z-score for a dataset:
Using the ages in the survey dataset above, we find the standard units for age:
0 1.857333
1 -0.701958
2 0.228693
3 -0.178467
4 0.577687
Name: age, dtype: float64
Then we add the age_s_units as a new column in the health DataFrame.
| genhlth | exerany | hlthplan | smoke100 | height | weight | wtdesire | age | gender | age_su | |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | good | 0 | 1 | 0 | 70 | 175 | 175 | 77 | m | 1.857333 |
| 1 | good | 0 | 1 | 1 | 64 | 125 | 115 | 33 | f | -0.701958 |
| 2 | good | 1 | 1 | 1 | 60 | 105 | 105 | 49 | f | 0.228693 |
| 3 | good | 1 | 1 | 0 | 66 | 132 | 124 | 42 | f | -0.178467 |
| 4 | very good | 0 | 1 | 0 | 61 | 150 | 130 | 55 | f | 0.577687 |
Next we sort in descending order the DataFrame by the age_su column so we can see the longest delay times.
10 oldest ages:
| genhlth | exerany | hlthplan | smoke100 | height | weight | wtdesire | age | gender | age_su | |
|---|---|---|---|---|---|---|---|---|---|---|
| 899 | very good | 0 | 1 | 0 | 70 | 150 | 150 | 99 | f | 3.136979 |
| 6709 | excellent | 1 | 1 | 1 | 62 | 120 | 115 | 99 | f | 3.136979 |
| 10349 | good | 1 | 1 | 1 | 60 | 130 | 130 | 97 | f | 3.020647 |
| 17050 | good | 1 | 1 | 1 | 64 | 128 | 128 | 96 | f | 2.962481 |
| 11911 | very good | 0 | 1 | 0 | 62 | 110 | 110 | 95 | f | 2.904316 |
| 16083 | good | 1 | 1 | 0 | 62 | 127 | 127 | 95 | f | 2.904316 |
| 7085 | very good | 1 | 1 | 0 | 66 | 125 | 130 | 94 | m | 2.846150 |
| 19502 | fair | 0 | 1 | 0 | 66 | 140 | 140 | 94 | m | 2.846150 |
| 2527 | very good | 1 | 1 | 0 | 61 | 118 | 118 | 94 | f | 2.846150 |
| 1443 | poor | 1 | 1 | 0 | 63 | 102 | 102 | 93 | f | 2.787984 |
We see that there are probably quite a few people whose age is more than 2 standard units.
Since Chebyshev's Theorem says that at least 75% of the data should be 2 standard units or less, we want to check if Chebyshev's bound holds true for this dataset.
We find the proportion of people whose age is 2 standard units or less, then we find the proportion of people whose age is 3 standard units or less.
Proportion of ages at 2 standard units or less: 0.9693
Proportion of ages at 3 standard units or less: 0.9999
Chebyshev's Theorem does hold for the age data. In fact the theorem gives a conservative lower bound.
- 97% of the ages are within 2 standard units, compared to Chebyshev's boundary of at least 75%.
- Almost 100% of the ages are within 3 standard units, compared to Chebyshev's boundary of at least 89%.
Looking at the first 10 rows of the sorted DataFrame, we see that only three out of 20,000 people (ages 99, 99, and 97) are above 3 standard units above the mean.
Because the age distribution is concentrated near the mean, a much larger percentage of the data falls within 2 and 3 standard units than Chebyshev's Theorem guarantees. In the next section, we will see that for data that are approximately normally distributed, these percentages can be predicted even more precisely.
The Normal Distribution and Standard Units
The normal distribution is also known as the Gaussian distribution or the bell curve. It has a symmetric bell-shaped curve, with a peak at the mean value and tapering off equally on both left and right sides. The normal distribution has a special relationship with the standard deviation.
We first create a normal distribution and fill in the areas under the normal curve that are within 1 standard deviation and within 2 standard deviations.
We can see that the area under the curve that's within 2 SDs covers a substantial part of the entire area under the curve.
To calculate the area under the curve, we use the cdf function of the norm module. The cdf or Cumulative Distribution Function returns the proportion of $$ \frac{ \text{area under the curve up to an x value}} { \text{total area under the curve}} $$
To call the cdf function we use the format:
area_up_to_x = norm.cdf(x)
Using this function, we find the proportion of the area under the curve within 1 SD and within 2 SDs compared to the total area under the curve.
Proportion of area within 1 SD: 0.683
Proportion of area within 2 SDs: 0.954
Proportion of area within 3 SDs: 0.997
We note that for a normal distribution:
- About 95% of the area under the curve is within 2 standard deviations of the mean, and about 99.7% is within 3 standard deviations of the mean.
- These percentages are much higher than Chebyshev's lower bounds of 75% and 89%. This is because Chebyshev's Theorem applies to any distribution, regardless of its shape, while the percentages above apply only to the normal distribution.
Because the normal distribution has a specific bell shape, we can make much stronger statements about the proportion of data within 2 or 3 standard deviations than we can for an arbitrary distribution.
The Central Limit Theorem
The central limit theorem says that if we take a large number of random samples from a population and find the mean of each sample, then the distribution of the means will approach a normal distribution in shape, regardless of the original population's distribution.
Just like with Chebyshev's Theorem, we will see how the central limit theorem works by investigating the food delivery dataset.
| Order_ID | Distance_km | Weather | Traffic_Level | Time_of_Day | Vehicle_Type | Preparation_Time_min | Courier_Experience_yrs | Delivery_Time_min | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 522 | 7.93 | Windy | Low | Afternoon | Scooter | 12 | 1.0 | 43 |
| 1 | 738 | 16.42 | Clear | Medium | Evening | Bike | 20 | 2.0 | 84 |
| 2 | 741 | 9.52 | Foggy | Low | Night | Scooter | 28 | 1.0 | 59 |
| 3 | 661 | 7.44 | Rainy | Medium | Afternoon | Scooter | 5 | 1.0 | 37 |
| 4 | 412 | 19.03 | Clear | Low | Morning | Bike | 16 | 5.0 | 68 |
We look at the delivery distance Distance_km distribution.
The distribution of delivery distances is not a bell curve or normal distribution. The distances are spread fairly evenly from about 1 to 20 kilometers.
We will repeatedly take random samples from this dataset and calculate the mean distance of each sample. According to the Central Limit Theorem, the distribution of these sample means will become approximately a normal distribution.
We find the mean and standard deviation of the delivery distance.
The mean: 10.06
The SD: 5.69
Next we take a large number of samples of the delivery times and plot the mean of the samples.
We write a function to select a sample of num_of_data values with replacement and find the mean of the sample.
We call the get_sample_mean 10,000 times to get 10,000 means of delivery time
We plot the distribution of the sample means.
We see that the distribution of the sample means is approximately bell-shaped, as predicted by the central limit theorem. We also see that the distribution is centered around the mean, which we calculated above to be 10.06.
We now repeat the same sampling steps as above, but our sample size is 1,000 instead of 300.
We notice that once again the distribution of sample means is a normal distribution, just as predicted by the central limit theorem, and that the distribution is centered around the mean of 10.06.
But we also notice that for samples of 1000 values, the distribution is narrower than the distribution for samples of 300 values. This makes sense because as the sample size increases, each individual data value (including the unusually high value or low value) has less influence on the sample mean. As a result, the sample mean varies less from sample to sample, and the means from repeated samples are closer together around the mean. This results in a narrower sampling distribution.
The Standard Deviation of the Sampling Distribution
The distribution of the sample means is also called the sampling distribution. The spread of the sampling distribution tells us how much the sample mean will vary from sample to sample.
The standard deviation of the sampling distribution is: $$\text{SD of sampling distribution} = \frac{ \text{population SD}}{ \sqrt{ \text{sample size}}}$$
The stanard deviation of the sampling distribution is often called the standard error of the mean, or simply the standard error.
SD of sampling distibution for sample of 300: 0.33
SD of sampling distribution for sample 1000: 0.18
The SD of the larger sample (1000) is smaller than the SD of the smaller sample (300). A smaller SD means the sampling distribution is narrower so the sample mean is more likely to be close to the population mean.
As the equation above shows, the SD of the sampling distribution decreases in reverse proportion with $\sqrt{ \text{sample size}}$. To reduce the SD by a factor of 10, the sample size must increase by a factor of 100.
Summary
In this chapter we learned how to describe the center and spread of a dataset using the mean, median, variance, and standard deviation. We also learned how z-scores or standard units, Chebyshev's Theorem, the normal distribution, and the Central Limit Theorem help us understand the distribution of data and why larger samples produce more reliable estimates of the population mean.


