8: Sampling
- Page ID
- 50784
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)To study or analyze a phenomenon in a large population, it is often not possible or practical to gather data from all of the population. For example, if we want to study the sleeping habits of people living in the US, it's difficult to get the sleeping data from every person living in the US. Instead, data scientists will choose a smaller number of people to represent the entire US population and obtain sleeping data from these people only. This smaller group of people is a sample of the US population.
In this chapter we learn the different types of samples and their properties.
Types of Samples
In a DataFrame, each row is the data record for one entity in the population. To sample the data and get a subset of the data, we select certain rows of the DataFrame. The way we select the rows determines the type of samples.
To illustrate the different sampling methods, we use the cars dataset that we've used in a previous chapter.
Number of rows: 392
First 10 rows:
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 18.0 | 8 | 307.0 | 130.0 | 3504.0 | 12.0 | 70 | 1 | chevrolet chevelle malibu |
| 1 | 15.0 | 8 | 350.0 | 165.0 | 3693.0 | 11.5 | 70 | 1 | buick skylark 320 |
| 2 | 18.0 | 8 | 318.0 | 150.0 | 3436.0 | 11.0 | 70 | 1 | plymouth satellite |
| 3 | 16.0 | 8 | 304.0 | 150.0 | 3433.0 | 12.0 | 70 | 1 | amc rebel sst |
| 4 | 17.0 | 8 | 302.0 | 140.0 | 3449.0 | 10.5 | 70 | 1 | ford torino |
| 5 | 15.0 | 8 | 429.0 | 198.0 | 4341.0 | 10.0 | 70 | 1 | ford galaxie 500 |
| 6 | 14.0 | 8 | 454.0 | 220.0 | 4354.0 | 9.0 | 70 | 1 | chevrolet impala |
| 7 | 14.0 | 8 | 440.0 | 215.0 | 4312.0 | 8.5 | 70 | 1 | plymouth fury iii |
| 8 | 14.0 | 8 | 455.0 | 225.0 | 4425.0 | 10.0 | 70 | 1 | pontiac catalina |
| 9 | 15.0 | 8 | 390.0 | 190.0 | 3850.0 | 8.5 | 70 | 1 | amc ambassador dpl |
Deterministic Sample
A deterministic sample is the result of identifying and selecting specific rows, without using any random chance.
One way to specify certain rows is by their row numbers or index numbers:
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 15.0 | 8 | 350.0 | 165.0 | 3693.0 | 11.5 | 70 | 1 | buick skylark 320 |
| 2 | 18.0 | 8 | 318.0 | 150.0 | 3436.0 | 11.0 | 70 | 1 | plymouth satellite |
| 5 | 15.0 | 8 | 429.0 | 198.0 | 4341.0 | 10.0 | 70 | 1 | ford galaxie 500 |
Another way is to specify a certain characteristic of the data:
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 70 | 19.0 | 3 | 70.0 | 97.0 | 2330.0 | 13.5 | 72 | 3 | mazda rx2 coupe |
| 110 | 18.0 | 3 | 70.0 | 90.0 | 2124.0 | 13.5 | 73 | 3 | maxda rx3 |
| 241 | 21.5 | 3 | 80.0 | 110.0 | 2720.0 | 13.5 | 77 | 3 | mazda rx-4 |
| 331 | 23.7 | 3 | 70.0 | 100.0 | 2420.0 | 12.5 | 80 | 3 | mazda rx-7 gs |
Probability Sample
A probability sample is one where each member of the population has a known and non-zero chance of being selected to be in the sample.
The cars DataFrame has 398 rows, as seen in the first Code cell. If we randomly select 10 rows to create a sample, then each car has the same chance of being chosen, which is:
P(being selected) = 10/398
To randomly choose N rows from a DataFrame, we use the sample() method.
sample_DataFrame = a_DataFrame.sample(n = N)
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 175 | 23.0 | 4 | 115.0 | 95.0 | 2694.0 | 15.0 | 75 | 2 | audi 100ls |
| 129 | 32.0 | 4 | 71.0 | 65.0 | 1836.0 | 21.0 | 74 | 3 | toyota corolla 1200 |
| 80 | 28.0 | 4 | 97.0 | 92.0 | 2288.0 | 17.0 | 72 | 3 | datsun 510 (sw) |
| 198 | 18.0 | 6 | 250.0 | 78.0 | 3574.0 | 21.0 | 76 | 1 | ford granada ghia |
| 301 | 31.8 | 4 | 85.0 | 65.0 | 2020.0 | 19.2 | 79 | 3 | datsun 210 |
To choose N random values from a column of a DataFrame or an array, we use the Numpy random.choice() method.
random_sample = np.random.choice(array, N)
['toyota corolla' 'pontiac phoenix' 'dodge aries wagon (sw)' 'amc matador']
Systematic Sample
A special type of probability sample is a systematic sample, where we select every k-th value in the population.
Steps for systematic sampling cars to get a sample size of 10:
- Find the sampling intervals based on the sample size. The larger the sample size, the smaller the interval.
- Find a random starting point in the first interval.
- Select every k-th index from the starting point.
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 32 | 19.0 | 6 | 232.0 | 100.0 | 2634.0 | 13.0 | 71 | 1 | amc gremlin |
| 71 | 15.0 | 8 | 304.0 | 150.0 | 3892.0 | 12.5 | 72 | 1 | amc matador (sw) |
| 110 | 18.0 | 3 | 70.0 | 90.0 | 2124.0 | 13.5 | 73 | 3 | maxda rx3 |
| 149 | 31.0 | 4 | 79.0 | 67.0 | 2000.0 | 16.0 | 74 | 2 | fiat x1.9 |
| 188 | 14.5 | 8 | 351.0 | 152.0 | 4215.0 | 12.8 | 76 | 1 | ford gran torino |
| 227 | 16.0 | 8 | 400.0 | 180.0 | 4220.0 | 11.1 | 77 | 1 | pontiac grand prix lj |
| 266 | 27.2 | 4 | 119.0 | 97.0 | 2300.0 | 14.7 | 78 | 3 | datsun 510 |
| 305 | 26.8 | 6 | 173.0 | 115.0 | 2700.0 | 12.9 | 79 | 1 | oldsmobile omega brougham |
| 344 | 37.7 | 4 | 89.0 | 62.0 | 2050.0 | 17.3 | 81 | 3 | toyota tercel |
| 383 | 22.0 | 6 | 232.0 | 112.0 | 2835.0 | 14.7 | 82 | 1 | ford granada l |
In step 3, [start::interval] means select the indices starting from the start starting point, all the way to the last index, and skipping every interval index positions.
For example: [2::5] means start at index 2, go to the last index, skipping every 5 indices, which means we get rows at indices: 2, 7, 12, 17, etc.
Sampling With and Without Replacement
Sampling with replacement means that each time we select an item from a population, we put it back before drawing the next one. This allows the same item to be selected more than once in the sample.
An example of sampling with replacement would be if we select a card out of a 52-card deck, look at the card face value, then put it back in the deck. The next time we select another card, we could end up with the same card we selected before.
Sampling without replacement means that each time we select an item from a population, it is removed from the population. This means no item will appear more than once in the sample.
An example of sampling without replacement is if we select a card out of a deck of cards and put it aside. When we select another card, we cannot select the same card again.
If we roll a die three times and record the three output numbers, would that be sampling with replacement or sampling without replacement?
Empirical Distribution
In data science, the word empirical means observed. So an empirical distribution is the histogram of observed data, such as data from a sample, since we measure and observe only the sample and not the entire population.
In the previous Chapter 7, we discussed that when a fair coin is tossed, P(heads) is 0.5 and P(tails) is 0.5. If we plotted the theoretical probability distribution, we would get two equal bars.
Now we use the same coin toss simulation function from Chapter 7 to get the observed numbers of heads and tails for N coin tosses:
Then we call the function to simulate 10 coin tosses, and plot the distribution of the resulting number of heads and tails
Next we repeat the same simulation for 100 coin tosses and then 1000 coin tosses.
If we run the Code cells for 10, 100, 1000 coin tosses several times, we can consistently see that as the number of coin tosses gets higher, the empirical histogram approaches to the calculated theoretical histogram. This is known as the convergence of empirical distribution, which happens because of the Law of Large Numbers.
The Law of Large Numbers states that as we repeat an experiment more times (tossing the coin more times), the average experiment result will get closer and closer to the average of the theoretical expected value.
Empirical Distribution and Sampling
The Law of Large Numbers also applies to sampling. The larger the sample size, the more closely the empirical distribution (or the distribution of the sample) resembles the population distribution.
To observe this, we continue with the delivery time dataset.
First 5 rows
| Order_ID | Distance_km | Weather | Traffic_Level | Time_of_Day | Vehicle_Type | Preparation_Time_min | Courier_Experience_yrs | Delivery_Time_min | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 522 | 7.93 | Windy | Low | Afternoon | Scooter | 12 | 1.0 | 43 |
| 1 | 738 | 16.42 | Clear | Medium | Evening | Bike | 20 | 2.0 | 84 |
| 2 | 741 | 9.52 | Foggy | Low | Night | Scooter | 28 | 1.0 | 59 |
| 3 | 661 | 7.44 | Rainy | Medium | Afternoon | Scooter | 5 | 1.0 | 37 |
| 4 | 412 | 19.03 | Clear | Low | Morning | Bike | 16 | 5.0 | 68 |
First we plot the distribution of the delivery times for the entire population.
Note that we use a density histogram (from Chapter 5) since we will be comparing the distributions of different size datasets.
Observing the histogram, we see that the data is close to a bell curve, but has a longer right tail, where the excessively long delivery times are. We check for the number of these long delivery times.
Delivery_Time_min
False 997
True 3
Name: count, dtype: int64
Since there are only three delivery times that are excessively long, we will drop them and only consider the delivery times that are in the bulk of the data.
997
We plot the histogram of the normal delivery times.
This histogram is our goal. We want the samples below to have a similar distribution.
First we randomly draw 20 delivery times out of normal_delivery_times and plot their distribution.
Clearly the sample is too small to have a distribution that's similar to the population distribution.
Next we draw 100 and 250 delivery times and plot their distributions.
We can see that the larger the sample size, the more similar its distribution is to the distribution of the population.
Empirical Distribution of a Statistic
Often data scientists are interested in a numerical measurement of a population, for example, what is the mean height of 18-year-olds in California. The mean height is an unknown measurement of the 18-year-old population and is called a parameter of the population.
Since it is not practical and is expensive to measure the height of each 18-year-old in California, data scientists will create a sample of 18-year-olds, measure their heights, and find the mean of these heights. The resulting mean is an estimate of the mean height parameter and it is called a statistic.
Note that the statistic or estimate of the parameter is singular, it doesn't have the same meaning as the plural statistics, which is a field of mathematics.
The steps to find and study the statistic of a sample:
- Decide what statistic to simulate.
- Draw a random sample and calculate the statistic.
- Repeat step 2 a number of times to get many readings of the statistic.
- Plot the readings to observe the statistic.
Step 1. Determine the statistic to simulate
Using the delivery time dataset, we choose the median delivery time as the statistic we want to study.
Recall that the median is the midpoint of the dataset, where 50% of the data is higher and 50% of the data is lower.
We start by calculating the median for all delivery times in the dataset. This value is the value we hope to estimate with our samples.
Population median delivery time: 55.0
Step 2. Sample and calculate the statistic
We write a function to take a sample of 250 delivery times and find the median delivery time. The function returns the median delivery time.
Step 3. Repeat step 2 to get many readings of the statistic
Then we write code to repeatedly call the function above 5000 times and save the returned median delivery times.
Generally the number of repetition is between 1,000 and 10,000 times.
5000
Step 4. Plot the readings
We plot the delivery_medians distribution.
We see that the median delivery time estimates are between 50 and 60 minutes, and the majority of the readings are between 54 and 57 minutes, which match the actual median delivery times of 55 minutes for the population.
Summary
In this chapter we learn the power of sampling and simulation. Since studying an entire population is rarely practical, we rely on population sampling and simulation techniques to accurately estimate real-world statistics and distributions.


