11: Estimation
- Page ID
- 50856
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)In the previous two chapters we used hypothesis testing to draw conclusions about a parameter of a dataset. In this chapter we learn how to use the dataset to estimate a parameter for the entire population.
An example of estimation is when we survey a sample of voters on their potential voting choice for a candidate, and then we use the sample's data to estimate the percentage of voters nationwide that will vote for that candidate. During the estimation we need to consider that the voters in the sample are randomly chosen. Different random samples can produce slightly different results simply due to chance. We therefore need to account for this sampling variability when making an estimate.
We start by getting an understanding of percentiles: what it is and what does it mean when a value is in the 75th percentile, for example. Then we continue with the concept of bootstrap through resampling to make an estimation.
Percentile
A dataset that contains numerical data can be sorted in increasing or decreasing order. After sorting, each number in the dataset has a particular position or a rank, and a percentile is the value at a particular rank. A percentile, as the name implies, is a value that is between 0 and 100.
A simple definition of a percentile:
The 80th percentile is the smallest value in the dataset that is greater than or equal to 80% of the data.
If we have a sequence of n numbers, to find the number that's at the pth percentile:
- Sort the numbers in increasing order.
- k = (p / 100) * n
- If k is an integer, take the kth element of the sorted numbers.
- If k is not an integer, round it up to the next integer, and take that element of the sorted numbers.
Example:
-
A sorted sequence is 5, 8, 20, 24, 30
-
The 80th percentile is:
k = (80 / 100) * 5 = 4, the 4th element is 24 -
The 35th percentile is:
k = (35/100) * 5 = 1.75, rounding 1.75 to 2, the 2nd element is 8
Now that we have a general idea of percentile, we apply it to the larger dataset, which is the delivery dataset for a food delivery company that we used in an earlier chapter.
First 5 rows:
| Order_ID | Distance_km | Weather | Traffic_Level | Time_of_Day | Vehicle_Type | Preparation_Time_min | Courier_Experience_yrs | Delivery_Time_min | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 522 | 7.93 | Windy | Low | Afternoon | Scooter | 12 | 1.0 | 43 |
| 1 | 738 | 16.42 | Clear | Medium | Evening | Bike | 20 | 2.0 | 84 |
| 2 | 741 | 9.52 | Foggy | Low | Night | Scooter | 28 | 1.0 | 59 |
| 3 | 661 | 7.44 | Rainy | Medium | Afternoon | Scooter | 5 | 1.0 | 37 |
| 4 | 412 | 19.03 | Clear | Low | Morning | Bike | 16 | 5.0 | 68 |
Since we're interested in the delivery distance only, we create a variable for the distance and plot the distance distribution.
0 7.93
1 16.42
2 9.52
3 7.44
4 19.03
Name: Distance_km, dtype: float64
We see that the distances range from 1 km to 20 km, and there's no clear mode, which means no delivery distance that occurs much more frequently than the others.
To find the 85th percentile, we can use the same calculation steps as above, or we can use the percentile function of numpy:
p = np.percentile(data_sequence, percentile)
85th percentile: 17.1375
The delivery distance of 17.14km is at or above 85% of the delivery distances.
Quartile
The range of percentiles, which is between 0 and 100, can be divided into four sections where each section is called a quartile.
- The first quartile is the 25th percentile
- The second quartile is the 50th percentile. The 50th percentile is also called the median, the midpoint of the sorted data.
- The third quartile is the 75th percentile.
In a distribution, values above the 25th percentile and below the 75th percentile is in the midddle 50% interval.
For the delivery distances, we calculate:
First quartile: 5.105
Second quartile: 10.19
Third quartile: 15.0175
Middle 50% is between: 5.105 and 15.0175
Gather Data
We use data from the US Bureau of Labor and Statistics (BLS) which contains the 2024 unemployment rate for each of the 50 states. The data are in the file state_unemployment.csv.
| state | unemployment_rate | |
|---|---|---|
| 0 | Alabama | 3.1 |
| 1 | Alaska | 4.6 |
| 2 | Arizona | 3.6 |
| 3 | Arkansas | 3.5 |
| 4 | California | 5.3 |
| 5 | Colorado | 4.3 |
| 6 | Connecticut | 3.2 |
| 7 | Delaware | 3.7 |
| 8 | Florida | 3.4 |
| 9 | Georgia | 3.5 |
Number of rows: 50
0 Alabama
1 Alaska
2 Arizona
3 Arkansas
4 California
5 Colorado
6 Connecticut
7 Delaware
8 Florida
9 Georgia
10 Hawaii
11 Idaho
12 Illinois
13 Indiana
14 Iowa
15 Kansas
16 Kentucky
17 Louisiana
18 Maine
19 Maryland
20 Massachusetts
21 Michigan
22 Minnesota
23 Mississippi
24 Missouri
25 Montana
26 Nebraska
27 Nevada
28 New Hampshire
29 New Jersey
30 New Mexico
31 New York
32 North Carolina
33 North Dakota
34 Ohio
35 Oklahoma
36 Oregon
37 Pennsylvania
38 Rhode Island
39 South Carolina
40 South Dakota
41 Tennessee
42 Texas
43 Utah
44 Vermont
45 Virginia
46 Washington
47 West Virginia
48 Wisconsin
49 Wyoming
Name: state, dtype: object
We call the DataFrame us because it contains the unemployment rates of all 50 US states. Since our data includes all 50 states, we have data for the entire population.
We look at the distribution of unemployment rates and find the mean unemployment rate of the population.
Mean unemployment: 3.68
Our goal is to see if a sample of 10 states can be used to estimate the mean unemployment rate of the entire population.
Take a Sample
Suppose we don't have the data for all 50 states, and instead we only have the unemployment data for 10 states that had been randomly chosen out of the 50 states.
Here is our sample of 10 states and the mean unemployment rate of the sample.
| state | unemployment_rate | |
|---|---|---|
| 0 | Michigan | 4.7 |
| 1 | Washington | 4.5 |
| 2 | Massachusetts | 4.0 |
| 3 | Rhode Island | 4.3 |
| 4 | Alabama | 3.1 |
| 5 | Oregon | 4.2 |
| 6 | Virginia | 2.9 |
| 7 | Montana | 3.0 |
| 8 | Kansas | 3.6 |
| 9 | Idaho | 3.7 |
Mean unemployment rate for sample: 3.8
We see that the mean unemployment rate of the sample is different from the mean unemployment rate of the population (3.68%). This shows that a statistic calculated from a random sample may not exactly match the population statistic.
How confident can we be that our sample gives us a reasonable estimate of the population?
Bootstrap the Sample
We've discovered that the Law of Large Numbers says that as the number of observations increases, the average of those observations tends to get closer to the population average. But we only have one sample of 10 states, how do we get the large number of observations to use to estimate the population average?
The answer is to bootstrap the sample. The term bootstrap comes from the saying that someone pull themselves up from their bootstrap, which is when the person uses what they have to reach their goal, without external help.
We will resample our original sample of 10 states to get many new samples to do our estimation. They key point in resampling is that we sample with replacement to create new samples that are different from the original.
Sample with Replacement
If we resample the original sample without replacement, we'll end up with the same sample
| state | unemployment_rate | |
|---|---|---|
| 0 | Michigan | 4.7 |
| 1 | Washington | 4.5 |
| 2 | Massachusetts | 4.0 |
| 3 | Rhode Island | 4.3 |
| 4 | Alabama | 3.1 |
| 5 | Oregon | 4.2 |
| 6 | Virginia | 2.9 |
| 7 | Montana | 3.0 |
| 8 | Kansas | 3.6 |
| 9 | Idaho | 3.7 |
| state | unemployment_rate | |
|---|---|---|
| 1 | Washington | 4.5 |
| 7 | Montana | 3.0 |
| 2 | Massachusetts | 4.0 |
| 3 | Rhode Island | 4.3 |
| 5 | Oregon | 4.2 |
| 4 | Alabama | 3.1 |
| 0 | Michigan | 4.7 |
| 6 | Virginia | 2.9 |
| 8 | Kansas | 3.6 |
| 9 | Idaho | 3.7 |
We see that the resample data are the same 10 states, just in a different order. This is the reshuffling that we used for a permutation in hypothesis testing in the previous chapter.
Now we sample with replacement. As we recall from Chapter 8, this means when a state is randomly selected and copied to the new resample, it remains available to be selected again. As a result, a state can appear more than once in the sample. This also means that some of the 10 original states may not appear in the resample at all.
| state | unemployment_rate | |
|---|---|---|
| 7 | Montana | 3.0 |
| 2 | Massachusetts | 4.0 |
| 0 | Michigan | 4.7 |
| 6 | Virginia | 2.9 |
| 4 | Alabama | 3.1 |
| 3 | Rhode Island | 4.3 |
| 4 | Alabama | 3.1 |
| 4 | Alabama | 3.1 |
| 9 | Idaho | 3.7 |
| 7 | Montana | 3.0 |
If we run the Code cell above a few times, we observe that the resamples are slightly different each time.
The Bootstrap
The bootstrap consists of two steps:
- Build one bootstrap sample by resampling the original sample, and finding the statistic of this bootstrap sample.
- Repeat the process to create many bootstrap samples and calculate the statistic to each one. These statistics form the bootstrap distribution.
The bootstrap distribution can then be used to estimate the population statistic.
One Bootstrap Sample
We write a function named resample that will resample the original sample and calculate and return the mean unemployment rate of the new sample.
Mean unemployment rate of bootstrap sample: 3.68
When we read in data above, we calculated the mean_us_unemployment or actual population mean of 3.68%.
We see that this population mean is pretty close to the bootstrap sample mean shown in the output above. However, the bootstrap sample mean will vary each time we create a new bootstrap sample, because the resampling creates a different new sample each time. If we run the Code cell above several times, we can see that the bootstrap sample mean changes with every resampling.
Therefore we want to repeat the process to create many bootstrap samples and calculate the mean of each one. The variation in these means will help us with estimating the population mean.
Repeat with Many Samples
We run the resample function 1000 times. Each time we create a bootstrap sample and find its mean, then we find and store the boostrap mean in an array. Last we plot the array, which is the bootstrap distribution.
We observe that the bootstrap distribution is centered around the yellow circle, which represents the original sample mean, as expected. And we also notice that the red circle, the population mean, falls within the main range of the bootstrap distribution.
Next we discuss this main range in more detail.
Confidence Interval
The bootstrap distribution shows a range of possible means from our resampling. Most of the bootstrap means are clustered near the center, while fewer are at the two ends of the distribution.
Rather than using the entire distribution to estimate the population mean, we can focus on the values in the middle of the distribution. We want a range of values that is wide enough to give us confidence that our estimate will be within the range, but not so wide that the estimation is not useful. A commonly used range is the middle 95% of the bootstrap distribution, called a 95% confidence interval. 95% is common because it includes most of the bootstrap distribution while still leaving out the more extreme values at both ends.
To find the middle 95%, we leave out 2.5% at each end of the bootstrap distribution. Therefore, the 95% confidence interval extends from the 2.5th percentile to the 97.5th percentile. In the histogram of the bootstrap distribution below, the 95% confidence interval is between the two dashed lines.
The plot shows the 95% confidence interval between the red lines. The red circle is the population mean, and it is inside the 95% confidence interval.
The 95% confidence interval is between 3.45 and 4.17
The red circle or population mean is 3.68
The yellow circle or original sample mean is 3.8
Our conclusion is:
The estimate mean unemployment rate for the 50 US states is 3.8
The actual mean unemployment rate is estimated to be between 3.45 and 4.17
Studies of the bootstrap method have shown that, if we repeatedly took random samples from the population and constructed a 95% confidence interval for each sample, then about 95 out of every 100 confidence intervals would contain the actual population mean.
Estimation Example
We now use the bootstrap to work with a sample where we don't know the actual population parameter. We'll use the sample to estimate the population parameter and construct a 95% confidence interval.
Gather Data
We will use the dataset from the CDC survey of people's health habits and conditions.
| genhlth | exerany | hlthplan | smoke100 | height | weight | wtdesire | age | gender | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | good | 0 | 1 | 0 | 70 | 175 | 175 | 77 | m |
| 1 | good | 0 | 1 | 1 | 64 | 125 | 115 | 33 | f |
| 2 | good | 1 | 1 | 1 | 60 | 105 | 105 | 49 | f |
| 3 | good | 1 | 1 | 0 | 66 | 132 | 124 | 42 | f |
| 4 | very good | 0 | 1 | 0 | 61 | 150 | 130 | 55 | f |
Number of rows: 20000
We see that the sample is 2000 people. We want to estimate the age of the population, so we first find the range of ages in the sample.
Youngest age: 18
Oldest age: 99
The Parameter
The age range shows that we have a sample of adults. We would like to estimate the mean age of all adults in the population.
The Bootstrap
First we find the mean of the original sample.
Original sample mean: 45.07
As in the previous example, we write a function to resample. Then we repeat the resample 1000 times to create the bootstrap distribution of sample means.
We plot the bootstrap distribution, the 95% confidence interval, and our estimate.
The 95% confidence interval is between 44.84 and 45.31
The yellow circle or original sample mean is 45.07
Our conclusion:
The estimate mean age for the population is 45.07
The actual mean age is estimated to be between 44.84 and 45.31
Relationship between Confidence Interval and Hypothesis Testing
In the previous chapter we learned that hypothesis testing determines whether there is enough evidence to reject a hypothesis about a population parameter. A confidence interval provides another way to make the same decision.
For the estimation of the mean age example, suppose the null hypothesis is: The population mean age is 40 years, and the alternative hypothesis is: The population mean age is not 40 years.
Our 95% confidence interval for the population mean age is 44.83 to 45.31 years. Because the null hypothesis mean age of 40 years is outside this interval, we reject the null hypothesis. This is the same conclusion we would reach from a hypothesis test with a p-value less than 0.05. In general, if the null hypothesis value lies outside a 95% confidence interval, the corresponding hypothesis test will have a p-value less than 0.05.
Guidelines for Using the Bootstrap
We have seen that the bootstrap is a simple way to estimate a population parameter and construct a confidence interval. To get reliable results, there are a few points to keep in mind:
- The sample should be representative of the population. Since the bootstrap repeatedly resamples from the original sample, any bias in the original sample will also appear in the boostrap samples.
- The sample should be reasonably large, A larger sample will likely produce more reliable bootstrap estimates because it represents the population better. Note that in the estimation of the mean unemployment of the 50 states, the 95% confidence interval is larger that the 95% confidence interval of the estimation of the mean age. This is because the sample size of the second estimation is large.
- Use many bootstrap resamples. Repeating the resampling 1000 times is sufficient for simple applications. Typically 10,000 resamples are used in general applications.
- The bootstrap works best for common statistics such as the mean, median, and proportion. The parameter that we want to estimate must not be at the edge of the population distribution such as the max, min, or very high or very low percentiles.
Summary
In this chapter, we learn how to estimate an unknown population parameter with a sample of the population. We repeatedly resample from the original sample to get the bootstrap distribution, then we construct a 95% confidence interval to estimate a range for the population parameter. We also learn when the bootstrap works well and how a 95% confidence interval is connected to hypothesis testing.


