Skip to main content
Workforce LibreTexts

11: Estimation

  • Page ID
    50856
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    In the previous two chapters we used hypothesis testing to draw conclusions about a parameter of a dataset. In this chapter we learn how to use the dataset to estimate a parameter for the entire population.

    An example of estimation is when we survey a sample of voters on their potential voting choice for a candidate, and then we use the sample's data to estimate the percentage of voters nationwide that will vote for that candidate. During the estimation we need to consider that the voters in the sample are randomly chosen. Different random samples can produce slightly different results simply due to chance. We therefore need to account for this sampling variability when making an estimate.

    We start by getting an understanding of percentiles: what it is and what does it mean when a value is in the 75th percentile, for example. Then we continue with the concept of bootstrap through resampling to make an estimation.

    Percentile

    A dataset that contains numerical data can be sorted in increasing or decreasing order. After sorting, each number in the dataset has a particular position or a rank, and a percentile is the value at a particular rank. A percentile, as the name implies, is a value that is between 0 and 100.

    A simple definition of a percentile:
    The 80th percentile is the smallest value in the dataset that is greater than or equal to 80% of the data.

    If we have a sequence of n numbers, to find the number that's at the pth percentile:

    • Sort the numbers in increasing order.
    • k = (p / 100) * n
    • If k is an integer, take the kth element of the sorted numbers.
    • If k is not an integer, round it up to the next integer, and take that element of the sorted numbers.

    Example:

    • A sorted sequence is 5, 8, 20, 24, 30

    • The 80th percentile is:
      k = (80 / 100) * 5 = 4, the 4th element is 24

    • The 35th percentile is:
      k = (35/100) * 5 = 1.75, rounding 1.75 to 2, the 2nd element is 8

    Now that we have a general idea of percentile, we apply it to the larger dataset, which is the delivery dataset for a food delivery company that we used in an earlier chapter.

    import numpy as np
    import pandas as pd
    import matplotlib.pyplot as plt
    
    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/delivery.csv"
    delivery = pd.read_csv(url)
    print("First 5 rows:")
    delivery.head()
    Output?
    First 5 rows:
    
      Order_ID Distance_km Weather Traffic_Level Time_of_Day Vehicle_Type Preparation_Time_min Courier_Experience_yrs Delivery_Time_min
    0 522 7.93 Windy Low Afternoon Scooter 12 1.0 43
    1 738 16.42 Clear Medium Evening Bike 20 2.0 84
    2 741 9.52 Foggy Low Night Scooter 28 1.0 59
    3 661 7.44 Rainy Medium Afternoon Scooter 5 1.0 37
    4 412 19.03 Clear Low Morning Bike 16 5.0 68

    Since we're interested in the delivery distance only, we create a variable for the distance and plot the distance distribution.

    distance = delivery['Distance_km']
    distance.head()
    Output?
    0     7.93
    1    16.42
    2     9.52
    3     7.44
    4    19.03
    Name: Distance_km, dtype: float64
    plt.figure(figsize=(4,3))
    plt.hist(distance, bins=20, edgecolor='black')
    plt.xlabel('Distance (km)')
    plt.title('Distribution of Delivery Distance')
    plt.grid()
    plt.show()
    Output?

    We see that the distances range from 1 km to 20 km, and there's no clear mode, which means no delivery distance that occurs much more frequently than the others.

    To find the 85th percentile, we can use the same calculation steps as above, or we can use the percentile function of numpy:
    p = np.percentile(data_sequence, percentile)

    print("85th percentile:", np.percentile(distance, 85))
    Output?
    85th percentile: 17.1375
    

    The delivery distance of 17.14km is at or above 85% of the delivery distances.

    Quartile

    The range of percentiles, which is between 0 and 100, can be divided into four sections where each section is called a quartile.

    • The first quartile is the 25th percentile
    • The second quartile is the 50th percentile. The 50th percentile is also called the median, the midpoint of the sorted data.
    • The third quartile is the 75th percentile.

    In a distribution, values above the 25th percentile and below the 75th percentile is in the midddle 50% interval.

    For the delivery distances, we calculate:

    print("First quartile:", np.percentile(distance, 25))
    print("Second quartile:", np.percentile(distance, 50))
    print("Third quartile:", np.percentile(distance, 75))
    print("Middle 50% is between:", np.percentile(distance, 25), "and", np.percentile(distance, 75))
    Output?
    First quartile: 5.105
    Second quartile: 10.19
    Third quartile: 15.0175
    Middle 50% is between: 5.105 and 15.0175
    

    Gather Data

    We use data from the US Bureau of Labor and Statistics (BLS) which contains the 2024 unemployment rate for each of the 50 states. The data are in the file state_unemployment.csv.

    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/state_unemployment.csv"
    us = pd.read_csv(url)
    display(us.head(10))
    print("Number of rows:", us.shape[0])
    Output?
      state unemployment_rate
    0 Alabama 3.1
    1 Alaska 4.6
    2 Arizona 3.6
    3 Arkansas 3.5
    4 California 5.3
    5 Colorado 4.3
    6 Connecticut 3.2
    7 Delaware 3.7
    8 Florida 3.4
    9 Georgia 3.5
    Number of rows: 50
    
    # print all the states to check
    us.state
    Output?
    0            Alabama
    1             Alaska
    2            Arizona
    3           Arkansas
    4         California
    5           Colorado
    6        Connecticut
    7           Delaware
    8            Florida
    9            Georgia
    10            Hawaii
    11             Idaho
    12          Illinois
    13           Indiana
    14              Iowa
    15            Kansas
    16          Kentucky
    17         Louisiana
    18             Maine
    19          Maryland
    20     Massachusetts
    21          Michigan
    22         Minnesota
    23       Mississippi
    24          Missouri
    25           Montana
    26          Nebraska
    27            Nevada
    28     New Hampshire
    29        New Jersey
    30        New Mexico
    31          New York
    32    North Carolina
    33      North Dakota
    34              Ohio
    35          Oklahoma
    36            Oregon
    37      Pennsylvania
    38      Rhode Island
    39    South Carolina
    40      South Dakota
    41         Tennessee
    42             Texas
    43              Utah
    44           Vermont
    45          Virginia
    46        Washington
    47     West Virginia
    48         Wisconsin
    49           Wyoming
    Name: state, dtype: object

    We call the DataFrame us because it contains the unemployment rates of all 50 US states. Since our data includes all 50 states, we have data for the entire population.

    We look at the distribution of unemployment rates and find the mean unemployment rate of the population.

    plt.figure(figsize=(4, 3))
    plt.hist(us['unemployment_rate'], edgecolor='black')
    plt.xlabel('Unemployment Rate (%)')
    plt.title('Distribution of Unemployment Rates by State')
    plt.grid()
    plt.show()
    Output?

    mean_us_unemployment = us["unemployment_rate"].mean()
    print("Mean unemployment:", round(mean_us_unemployment,2))
    Output?
    Mean unemployment: 3.68
    

    Our goal is to see if a sample of 10 states can be used to estimate the mean unemployment rate of the entire population.

    Take a Sample

    Suppose we don't have the data for all 50 states, and instead we only have the unemployment data for 10 states that had been randomly chosen out of the 50 states.

    Here is our sample of 10 states and the mean unemployment rate of the sample.

    # take a random sample of 10, and reset the index of these states 
    # so we don't use their old index in the 50 states
    
    original_sample = us.sample(n=10).reset_index(drop=True)
    original_sample
    Output?
      state unemployment_rate
    0 Michigan 4.7
    1 Washington 4.5
    2 Massachusetts 4.0
    3 Rhode Island 4.3
    4 Alabama 3.1
    5 Oregon 4.2
    6 Virginia 2.9
    7 Montana 3.0
    8 Kansas 3.6
    9 Idaho 3.7
    original_mean = original_sample["unemployment_rate"].mean()
    print("Mean unemployment rate for sample:", round(original_mean, 2))
    Output?
    Mean unemployment rate for sample: 3.8
    

    We see that the mean unemployment rate of the sample is different from the mean unemployment rate of the population (3.68%). This shows that a statistic calculated from a random sample may not exactly match the population statistic.

    How confident can we be that our sample gives us a reasonable estimate of the population?

    Bootstrap the Sample

    We've discovered that the Law of Large Numbers says that as the number of observations increases, the average of those observations tends to get closer to the population average. But we only have one sample of 10 states, how do we get the large number of observations to use to estimate the population average?

    The answer is to bootstrap the sample. The term bootstrap comes from the saying that someone pull themselves up from their bootstrap, which is when the person uses what they have to reach their goal, without external help.

    We will resample our original sample of 10 states to get many new samples to do our estimation. They key point in resampling is that we sample with replacement to create new samples that are different from the original.

    Sample with Replacement

    If we resample the original sample without replacement, we'll end up with the same sample

    # The original sample
    
    display(original_sample)
    Output?
      state unemployment_rate
    0 Michigan 4.7
    1 Washington 4.5
    2 Massachusetts 4.0
    3 Rhode Island 4.3
    4 Alabama 3.1
    5 Oregon 4.2
    6 Virginia 2.9
    7 Montana 3.0
    8 Kansas 3.6
    9 Idaho 3.7
    # resample without replacement
    
    new_sample = original_sample.sample(n=10, replace=False)
    new_sample
    Output?
      state unemployment_rate
    1 Washington 4.5
    7 Montana 3.0
    2 Massachusetts 4.0
    3 Rhode Island 4.3
    5 Oregon 4.2
    4 Alabama 3.1
    0 Michigan 4.7
    6 Virginia 2.9
    8 Kansas 3.6
    9 Idaho 3.7

    We see that the resample data are the same 10 states, just in a different order. This is the reshuffling that we used for a permutation in hypothesis testing in the previous chapter.

    Now we sample with replacement. As we recall from Chapter 8, this means when a state is randomly selected and copied to the new resample, it remains available to be selected again. As a result, a state can appear more than once in the sample. This also means that some of the 10 original states may not appear in the resample at all.

    # resample with replacement
    
    new_sample = original_sample.sample(n=10, replace=True)
    new_sample
    Output?
      state unemployment_rate
    7 Montana 3.0
    2 Massachusetts 4.0
    0 Michigan 4.7
    6 Virginia 2.9
    4 Alabama 3.1
    3 Rhode Island 4.3
    4 Alabama 3.1
    4 Alabama 3.1
    9 Idaho 3.7
    7 Montana 3.0

    If we run the Code cell above a few times, we observe that the resamples are slightly different each time.

    The Bootstrap

    The bootstrap consists of two steps:

    • Build one bootstrap sample by resampling the original sample, and finding the statistic of this bootstrap sample.
    • Repeat the process to create many bootstrap samples and calculate the statistic to each one. These statistics form the bootstrap distribution.

    The bootstrap distribution can then be used to estimate the population statistic.

    One Bootstrap Sample

    We write a function named resample that will resample the original sample and calculate and return the mean unemployment rate of the new sample.

    def resample():
      new_sample = original_sample.sample(n=10, replace=True)
      return np.mean(new_sample["unemployment_rate"])
    
    # show the mean unemployment rate for the new sample
    print("Mean unemployment rate of bootstrap sample:", round(resample(), 2))
    Output?
    Mean unemployment rate of bootstrap sample: 3.68
    

    When we read in data above, we calculated the mean_us_unemployment or actual population mean of 3.68%.

    We see that this population mean is pretty close to the bootstrap sample mean shown in the output above. However, the bootstrap sample mean will vary each time we create a new bootstrap sample, because the resampling creates a different new sample each time. If we run the Code cell above several times, we can see that the bootstrap sample mean changes with every resampling.

    Therefore we want to repeat the process to create many bootstrap samples and calculate the mean of each one. The variation in these means will help us with estimating the population mean.

    Repeat with Many Samples

    We run the resample function 1000 times. Each time we create a bootstrap sample and find its mean, then we find and store the boostrap mean in an array. Last we plot the array, which is the bootstrap distribution.

    # create a list to store the bootstrap means
    means = []
    # run resample 1000 times and store the mean unemployment rate of each sample
    for i in range(1000):
        means.append(resample())
    # convert the list to an array
    mean_rate_array = np.array(means)
    Output?
    # show the distribution of the means 
    
    plt.figure(figsize=(4,4))
    plt.hist(mean_rate_array, edgecolor='black')
    plt.xlabel('Mean Unemployment Rate (%)')
    plt.title('Bootstrap Distribution')
    # add a red circle for the population mean
    plt.scatter(mean_us_unemployment, 0, color='red', s=100, zorder=3)
    # add a yellow circle for the boostrap mean
    plt.scatter(original_mean, 0, color='yellow', s=100, zorder=3)
    plt.grid()
    plt.show()
    Output?

    We observe that the bootstrap distribution is centered around the yellow circle, which represents the original sample mean, as expected. And we also notice that the red circle, the population mean, falls within the main range of the bootstrap distribution.

    Next we discuss this main range in more detail.

    Confidence Interval

    The bootstrap distribution shows a range of possible means from our resampling. Most of the bootstrap means are clustered near the center, while fewer are at the two ends of the distribution.

    Rather than using the entire distribution to estimate the population mean, we can focus on the values in the middle of the distribution. We want a range of values that is wide enough to give us confidence that our estimate will be within the range, but not so wide that the estimation is not useful. A commonly used range is the middle 95% of the bootstrap distribution, called a 95% confidence interval. 95% is common because it includes most of the bootstrap distribution while still leaving out the more extreme values at both ends.

    To find the middle 95%, we leave out 2.5% at each end of the bootstrap distribution. Therefore, the 95% confidence interval extends from the 2.5th percentile to the 97.5th percentile. In the histogram of the bootstrap distribution below, the 95% confidence interval is between the two dashed lines.

    plt.figure(figsize=(4,4))
    # plot the bootstrap means
    plt.hist(mean_rate_array, edgecolor='black')
    # find the 2.5th and 97.5th percentiles
    lower_percentile = np.percentile(mean_rate_array, 2.5)
    upper_percentile = np.percentile(mean_rate_array, 97.5)
    # plot the 2.5th and 97.5th percentile lines
    plt.axvline(lower_percentile, color='red', linestyle='--', linewidth=2)
    plt.axvline(upper_percentile, color='red', linestyle='--', linewidth=2)
    # plot the population mean
    plt.scatter(mean_us_unemployment, 0, color='red', s=100, zorder=3)
    # plot the sample mean
    plt.scatter(original_mean, 0, color='yellow', s=100, zorder=3)
    # show labels
    plt.xlabel('Mean Unemployment Rate (%)')
    plt.ylabel('Frequency')
    plt.title('Bootstrap Distribution with 95% Confidence Interval')
    plt.grid()
    plt.show()
    Output?

    The plot shows the 95% confidence interval between the red lines. The red circle is the population mean, and it is inside the 95% confidence interval.

    print("The 95% confidence interval is between", round(lower_percentile, 2), "and", round(upper_percentile,2))
    print("The red circle or population mean is", round(mean_us_unemployment, 2))
    print("The yellow circle or original sample mean is", round(original_mean,2))
    Output?
    The 95% confidence interval is between 3.45 and 4.17
    The red circle or population mean is 3.68
    The yellow circle or original sample mean is 3.8
    

    Our conclusion is:

    print("The estimate mean unemployment rate for the 50 US states is", round(original_mean, 2))
    print("The actual mean unemployment rate is estimated to be between", round(lower_percentile, 2), "and", round(upper_percentile,2))
    Output?
    The estimate mean unemployment rate for the 50 US states is 3.8
    The actual mean unemployment rate is estimated to be between 3.45 and 4.17
    

    Studies of the bootstrap method have shown that, if we repeatedly took random samples from the population and constructed a 95% confidence interval for each sample, then about 95 out of every 100 confidence intervals would contain the actual population mean.

    Estimation Example

    We now use the bootstrap to work with a sample where we don't know the actual population parameter. We'll use the sample to estimate the population parameter and construct a 95% confidence interval.

    Gather Data

    We will use the dataset from the CDC survey of people's health habits and conditions.

    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/cdc.csv"
    health = pd.read_csv(url)
    display(health.head())
    print("Number of rows:", health.shape[0])
    Output?
      genhlth exerany hlthplan smoke100 height weight wtdesire age gender
    0 good 0 1 0 70 175 175 77 m
    1 good 0 1 1 64 125 115 33 f
    2 good 1 1 1 60 105 105 49 f
    3 good 1 1 0 66 132 124 42 f
    4 very good 0 1 0 61 150 130 55 f
    Number of rows: 20000
    

    We see that the sample is 2000 people. We want to estimate the age of the population, so we first find the range of ages in the sample.

    print("Youngest age:", health.age.min())
    print("Oldest age:", health.age.max())
    Output?
    Youngest age: 18
    Oldest age: 99
    

    The Parameter

    The age range shows that we have a sample of adults. We would like to estimate the mean age of all adults in the population.

    The Bootstrap

    First we find the mean of the original sample.

    original_mean = health.age.mean()
    print("Original sample mean:", round(original_mean, 2))
    Output?
    Original sample mean: 45.07
    

    As in the previous example, we write a function to resample. Then we repeat the resample 1000 times to create the bootstrap distribution of sample means.

    def resample():
        new_sample = health.age.sample(n=len(health), replace=True)
        return np.mean(new_sample)
    
    means = []
    for i in range(1000):
        means.append(resample())
    mean_array = np.array(means)
    Output?

    We plot the bootstrap distribution, the 95% confidence interval, and our estimate.

    plt.figure(figsize=(4,4))
    # plot the bootstrap means
    plt.hist(mean_array, edgecolor='black')
    # find the 2.5th and 97.5th percentiles
    lower_percentile = np.percentile(mean_array, 2.5)
    upper_percentile = np.percentile(mean_array, 97.5)
    # plot the 2.5th and 97.5th percentile lines
    plt.axvline(lower_percentile, color='red', linestyle='--', linewidth=2)
    plt.axvline(upper_percentile, color='red', linestyle='--', linewidth=2)
    # plot the sample mean
    plt.scatter(original_mean, 0, color='yellow', s=100, zorder=3)
    # show labels
    plt.xlabel('Mean Age')
    plt.ylabel('Frequency')
    plt.title('Bootstrap Distribution with 95% Confidence Interval')
    plt.grid()
    plt.show()
    Output?

    print("The 95% confidence interval is between", round(lower_percentile, 2), "and", round(upper_percentile,2))
    print("The yellow circle or original sample mean is", round(original_mean,2))
    Output?
    The 95% confidence interval is between 44.84 and 45.31
    The yellow circle or original sample mean is 45.07
    

    Our conclusion:

    print("The estimate mean age for the population is", round(original_mean, 2))
    print("The actual mean age is estimated to be between", round(lower_percentile, 2), "and", round(upper_percentile,2))
    Output?
    The estimate mean age for the population is 45.07
    The actual mean age is estimated to be between 44.84 and 45.31
    

    Relationship between Confidence Interval and Hypothesis Testing

    In the previous chapter we learned that hypothesis testing determines whether there is enough evidence to reject a hypothesis about a population parameter. A confidence interval provides another way to make the same decision.

    For the estimation of the mean age example, suppose the null hypothesis is: The population mean age is 40 years, and the alternative hypothesis is: The population mean age is not 40 years.

    Our 95% confidence interval for the population mean age is 44.83 to 45.31 years. Because the null hypothesis mean age of 40 years is outside this interval, we reject the null hypothesis. This is the same conclusion we would reach from a hypothesis test with a p-value less than 0.05. In general, if the null hypothesis value lies outside a 95% confidence interval, the corresponding hypothesis test will have a p-value less than 0.05.

    Guidelines for Using the Bootstrap

    We have seen that the bootstrap is a simple way to estimate a population parameter and construct a confidence interval. To get reliable results, there are a few points to keep in mind:

    1. The sample should be representative of the population. Since the bootstrap repeatedly resamples from the original sample, any bias in the original sample will also appear in the boostrap samples.
    2. The sample should be reasonably large, A larger sample will likely produce more reliable bootstrap estimates because it represents the population better. Note that in the estimation of the mean unemployment of the 50 states, the 95% confidence interval is larger that the 95% confidence interval of the estimation of the mean age. This is because the sample size of the second estimation is large.
    3. Use many bootstrap resamples. Repeating the resampling 1000 times is sufficient for simple applications. Typically 10,000 resamples are used in general applications.
    4. The bootstrap works best for common statistics such as the mean, median, and proportion. The parameter that we want to estimate must not be at the edge of the population distribution such as the max, min, or very high or very low percentiles.

    Summary

    In this chapter, we learn how to estimate an unknown population parameter with a sample of the population. We repeatedly resample from the original sample to get the bootstrap distribution, then we construct a 95% confidence interval to estimate a range for the population parameter. We also learn when the bootstrap works well and how a 95% confidence interval is connected to hypothesis testing.


      This page titled 11: Estimation was last modified on Fri, 25 Sep 2026 01:21:13 GMT and is shared under a CC BY 4.0 license and was authored, remixed, and/or curated by Clare Nguyen.

      • Was this article helpful?