10: Two-Sample Tests
- Page ID
- 50855
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\ket}[1]{\left| #1 \right>}\)
\(\newcommand{\bra}[1]{\left< #1 \right|}\)
\(\newcommand{\braket}[2]{\left< #1 \vphantom{#2} \right| \left. #2 \vphantom{#1} \right>}\)
\(\newcommand{\braopket}[3]{\left< #1 \vphantom{#2}\vphantom{#3} \right| #2 \vphantom{#1}\vphantom{#3} \left| #3 \vphantom{#1}\vphantom{#2} \right>}\)
\(\newcommand{\qmvec}[1]{\mathbf{\vec{#1}}}\)
\(\newcommand{\op}[1]{\hat{\mathbf{#1}}}\)
\(\newcommand{\expect}[1]{\langle #1 \rangle}\)
\(\newcommand{\dfn}[1]{\emph{\textbf{#1}}}\)
Up to now we've done statistical analysis with one data sample, such as the number of heads in a coin toss experiment. But often data scientists also need to make decisions based on comparing two data samples with each other, such as when comparing two groups of students who studied with background music and with no background music.
In this chapter we use hypothesis testing with two distinct groups of data to determine if the differences between them are just due to random chance or if the difference is meaningful.
Gather Data
The data are from the US Center for Disease Control and Prevention (CDC) and is stored at OpenIntro. The CDC annually surveys the health status and habits of 350,000 people in the US, and the dataset is a random sample of 20,000 people from the survey conducted in 2000.
| genhlth | exerany | hlthplan | smoke100 | height | weight | wtdesire | age | gender | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | good | 0 | 1 | 0 | 70 | 175 | 175 | 77 | m |
| 1 | good | 0 | 1 | 1 | 64 | 125 | 115 | 33 | f |
| 2 | good | 1 | 1 | 1 | 60 | 105 | 105 | 49 | f |
| 3 | good | 1 | 1 | 0 | 66 | 132 | 124 | 42 | f |
| 4 | very good | 0 | 1 | 0 | 61 | 150 | 130 | 55 | f |
Number of rows: 20000
The rows from left to right are: general health, any exercise in the past month, health plan coverage, smoke at least 100 cigarettes in respondent's life, height, weight, desired weight, age, and gender.
We will look at two features: exercise habit or exerany (0 for no exercise in past month, 1 for exercise in past month), and weight (a quantitative variable). We want to know whether people who exercise have a lower body weight distribution than people who don't exercise.
Exploratory Data Analysis (EDA)
It might seem reasonable that people who exercise have lower body weight than those who don't exercise. However, a person's weight is influenced by their height. Taller people naturally tend to weigh more than shorter people, no matter how much they exercise.
Assessing the Model
If we only look at the weight, then when we build the model for the null hypothesis, if one of our sample groups happen to have slightly higher average height, then their weight would tend to be higher regardless of their exercise habit. We also need to take into account the differences in height when comparing the two groups.
In this case the height is a confounding variable. A confounding variable is an extra factor that can bias the results when studying only weight and exercise.
To take into account both the height and weight, we will use the Body Mass Index (BMI) instead. The formula for BMI is \[\text{BMI} = \dfrac{\text{weight}}{\text{height}^2} \times 703\] With the BMI value, the weight is scaled relative to the height.
We add another column to the health DataFrame to store the BMI.
| genhlth | exerany | hlthplan | smoke100 | height | weight | wtdesire | age | gender | bmi | |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | good | 0 | 1 | 0 | 70 | 175 | 175 | 77 | m | 25.107143 |
| 1 | good | 0 | 1 | 1 | 64 | 125 | 115 | 33 | f | 21.453857 |
| 2 | good | 1 | 1 | 1 | 60 | 105 | 105 | 49 | f | 20.504167 |
| 3 | good | 1 | 1 | 0 | 66 | 132 | 124 | 42 | f | 21.303030 |
| 4 | very good | 0 | 1 | 0 | 61 | 150 | 130 | 55 | f | 28.339156 |
| 5 | very good | 1 | 1 | 0 | 64 | 114 | 114 | 55 | f | 19.565918 |
| 6 | very good | 1 | 1 | 0 | 71 | 194 | 185 | 31 | m | 27.054553 |
| 7 | very good | 0 | 1 | 0 | 67 | 170 | 160 | 45 | m | 26.622856 |
| 8 | good | 0 | 1 | 1 | 65 | 150 | 130 | 27 | f | 24.958580 |
| 9 | good | 1 | 1 | 0 | 70 | 180 | 170 | 44 | m | 25.824490 |
We see that on row index 0 someone weighs 175 lbs, which is heavier than the person on row index 8, whose weight is 150. But because of their height difference, they have very similar BMI. Now we can compare standardized body mass (BMI) against exercise habit.
Prepare Data
Next we concentrate on the two features, exercise habit and BMI, by creating a new DataFrame exercise that only has the exerany and bmi columns.
| exerany | bmi | |
|---|---|---|
| 0 | 0 | 25.107143 |
| 1 | 0 | 21.453857 |
| 2 | 1 | 20.504167 |
| 3 | 1 | 21.303030 |
| 4 | 0 | 28.339156 |
Number of rows: 20000
We find the count and proportion of exercise and non-exercise groups.
No exercise (0): count = 14914 proportion = 0.75
Exercise (1): count = 5086 proportion = 0.25
Then we plot the distribution of the BMI of the two groups.
It looks like the BMI distributions overlap quite a bit for the exercise and non-exercise groups. We also see that there seems to be a right tail in the distribution, which means that there are some people with BMI values that are much higher than those of most people.
The Test Statistic
To compare the exercise and no exercise group's BMI, we find the difference in the mean BMI of the two groups.
Mean BMI for exercise group: 25.98
Mean BMI for no exercise group: 27.27
Difference (exercise - no exercise): -1.29
There is a small difference: the mean BMI of the exercise group is lower than the mean BMI of the no exercise group by 1.29.
Is the 1.29 difference due to random chance or is there evidence of a difference between the two groups?
We will use hypothesis testing to find the answer.
Hypothesis Testing
We'll use the same steps as in Chapter 9, but this time we're working with two data samples so our model for the null hypothesis will be built differently.
State the Hypothesis
The null hypothesis: There is no difference between the mean BMI of the exercise and no exercise groups. Any difference is due to random chance.
The alternative hypothesis: There is a difference between the mean BMI of the exercise and no exercise groups.
Build the Model
Recall from Chapter 9 that we want to build a model for the null hypothesis, which means the model should show us what the test statistic looks like if there is no difference in the mean BMI of the exercise and no exercise groups.
If there is no difference, then whether a person is labeled as "exercise" or "no exercise" should not be associated with their BMI. We can model this by randomly shuffling the exerany (the exercise habit) column among the people in the dataset. This is called random permutation.
Random Permutation Example
To see how random permutation works, we copy the first seven rows of the exercise_bmi DataFrame into an example Dataframe so we can randomly shuffle the exerany column.
| exerany | bmi | |
|---|---|---|
| 0 | 0 | 25.107143 |
| 1 | 0 | 21.453857 |
| 2 | 1 | 20.504167 |
| 3 | 1 | 21.303030 |
| 4 | 0 | 28.339156 |
| 5 | 1 | 19.565918 |
| 6 | 1 | 27.054553 |
Then we shuffle the exerany column by using the same sample function discusssed in Chapter 8:
sample_DataFrame = a_DataFrame.sample(n = N)
But this time we use replace=False for sampling with no replacement, and we use frac=1 so 100% of the data are selected.
5 1
2 1
4 0
0 0
6 1
3 1
1 0
Name: exerany, dtype: int64
Note that the index values show us that the rows of exerany had been shuffled randomly.
We combine the shuffled_exerany column with the original example DataFrame.
| exerany | bmi | shuffled_exerany | |
|---|---|---|---|
| 0 | 0 | 25.107143 | 1 |
| 1 | 0 | 21.453857 | 1 |
| 2 | 1 | 20.504167 | 0 |
| 3 | 1 | 21.303030 | 0 |
| 4 | 0 | 28.339156 | 1 |
| 5 | 1 | 19.565918 | 1 |
| 6 | 1 | 27.054553 | 0 |
As the example DataFrame shows, the BMI of each person is the same, but after shuffling, the person has been randomly assigned to the exercise or no exercise group. Note also that the shuffle guarantees that we will end up with the same number of people in each group. The original exerany feature has four people in the exercise group and three people in the no exercise group, and the shuffled_exerany feature has the same counts.
Build the Null Distribution
After shuffling the exerany feature of the dataset, we calculate the mean BMI of the two shuffled groups and find the difference. This difference is due to random chance, as assumed by the null hypothesis.
Then we repeat the two steps: 1) shuffle and 2) find the mean BMI difference between the two groups, so we can have a large number of differences. These differences will show us the range of values we would expect from random chance.
To get ready to repeat the two steps of shuffling and finding the difference, we write a function that does the two steps.
We repeat the permutation 1000 times to get enough simulated differences that can occur by random chance.
We plot the distribution of the differences.
As expected, if the null hypothesis is true and the difference in mean BMI of the two groups is due to random chance, then the difference tends to be close to 0. This can be seen int the distribution above, with most of the differences between -0.15 and 0.15.
Compare Observed Data and the Model's Data
When we calculated the test statistic with the observed data in an earlier step, we found that
Exercise group mean BMI - No exercise group mean BMI = -1.29
Given the null distribution above, -1.29 is very far outside the range of what the 1000 permutations produced by random chance. Therefore, it's unlikely that the difference in the observed data is due to random chance.
p-value Calculation
We can calculate the p-value to see if it agrees with our conclusion above. We use the same calculation of p-value as in Chapter 8.
0.0
The calculated p-value of 0.0 is expected. With the 1000 permutations that we did, none of them had a distance larger than 1.29. But this doesn't mean that the p-value is exactly 0.0. It appears 0.0 for 1000 simulations, so we can say that the p-value is smaller than 1/1000 or smaller than 0.001.
With a p-value value that's less than 0.001, the p-value is statistically significant. This agrees with our conclusion from observing the null distribution, therefore we reject the null hypothesis and conclude that there is evidence of a difference in mean BMI between the exercise and no exercise groups.
Observational Data
It is important to note that our hypothesis testing above provides evidence that there is an association between BMI value and exercise habit. There seems to be a difference in BMI values among people who exercise and those who don't, but we cannot say that exercise causes the difference in BMI. This is because people were not randomly assigned to the exercise or no exercise group.
If people are randomly assigned to a group, the people who have healthy diets can be assigned to either group, and older and younger people can be assigned to either group. Random assignment tends to balance these other characteristics between the groups, making exercise habit one of the primary difference between the groups. This makes it possible to infer that exercise may cause a difference in mean BMI.
On the other hand, when people are surveyed for their health condition, as in our dataset, then the two groups can be different in many ways. For example, people who exercise may also have healthy diets or are younger. These factors, the confounding variables, may also influence BMI so it's not possible to say that exercise itself causes the difference in BMI.
The dataset for the exercise habit and mean BMI is observational data. They are data gathered by measuring or recording behaviors, events, or phenomena as they naturally occur, without manipulating any chracteristics of the data. Observational data can only show evidence of association between two factors, they cannot be used to determine cause and effect.
Causality and A/B Testing
We now look at an example of experimental data, which is the opposite of observational data. In experimental data people are randomly assigned to two groups, and this makes it possible to infer a cause and effect relationship as discussed above.
The Criteo Uplift Prediction dataset contains data from online advertising experiments conducted by Criteo, a digital advertising company. Users were randomly assigned to either an advertising treatment group or a control group. Criteo could show advertisements to users in the treatment group but withheld its advertising from users in the control group. Then the company recorded whether the user visited the advertiser's website.
This type of randomized experiment is commonly called an A/B test. A/B testing is a type of randomized experiment used to compare two alternatives. Participants are randomly assigned to one of two groups:
- Group A (control): receives the existing or control condition.
- Group B (treatment): receives the new condition being tested.
We then compare an outcome from the two groups to determine whether the treatment made a difference. Because participants are randomly assigned, A/B testing can provide evidence that the treatment caused a difference in the outcome.
A/B testing is widely used in business, technology, and marketing to determine whether a change causes an improvement in a desired outcome.
Gather Data
The original dataset contains millions of records and many features. For this example, we use a random sample of 20,000 users and focus on two features:
- treatment: whether the user was randomly assigned to the control group (0) or advertising treatment group (1)
- visit: whether the user did not visit (0) or visited (1) the advertiser's website
Last 10 rows
| treatment | visit | |
|---|---|---|
| 19990 | 1 | 0 |
| 19991 | 1 | 0 |
| 19992 | 1 | 1 |
| 19993 | 1 | 0 |
| 19994 | 1 | 0 |
| 19995 | 1 | 0 |
| 19996 | 1 | 0 |
| 19997 | 1 | 0 |
| 19998 | 1 | 0 |
| 19999 | 1 | 0 |
Number of rows: 20000
In the last 10 rows, the users were all in the treatment group, and only one of them visited the website.
The Test Statistic
We first see how many users from each group visited or not visited the website.
treatment visit
1 0 16136
0 0 2916
1 1 828
0 1 120
Name: count, dtype: int64
We see that for the treatment group, 16,136 people did not visit the website and 828 visited.
For the control group, 2916 people did not visit the website and 120 people visited.
For each group we find the rate of visit, which is the proportion \[ \frac{\text{number of visits}}{\text{total count of visits and no visits}} \nonumber\]
We observe that since the data is 1 or 0 only, if we find the sum of the visits (all the 1 values + all the 0 values), the result equals the sum of all the 1 values, which is also the number of visits. This means the fraction above can be written as \[\frac{\text{sum of visits}}{\text{total count of visits and no visits}} = \text{the mean of visits} \nonumber\]
The test statistic is the difference in the proportion of website visits for each group. Does the treatment group have a higher proportion of website visits than that of the control group?
Proportion of control group that visited the website 0.0395
Proportion of treatment group that visited the website 0.0488
Observed difference: 0.0093
We see that the observed difference between the two groups is only 0.0093 or 0.93%. This looks like a small percentage, but we don't know yet if it's due to random chance. We will use hypothesis testing to investigate.
Hypothesis Testing
The question we ask about the treatment and control groups:
Is the difference in visit proportions likely due to random chance, or is there evidence that advertising increases the visit proportion?
State the Hypothesis
The null hypothesis: There is no difference in the proportion of website visits between the treatment group and control group. Any difference is due to random chance.
The alternative hypothesis: There is a difference in the proportion of the website visits between the treatment and control groups.
Build the Model
Just as in the previous example with the mean BMI difference, we will randomly shuffle the treatment column while leaving the visit column unchanged for one permutation, then we find the difference in visit proportion between the two groups. Then we repeat the permutation 1000 times to get the distribution of differences.
We repeat the permutation to get 1000 differences and plot them.
Compare the Observed Data and the Model's Data
If the null hypothesis is true, then the difference in visit proportion between the two groups would be centered around 0. The distribution shows that most of the differences are in the range of -0.005 and 0.005. Our observed difference is about 0.009, which is near the right tail of the distribution. This means the observed difference is likely rare.
We calculate the p-value to confirm that it is rare.
0.022
Conclusion
The p-value of 0.036 is smaller than the significance level of 0.05. Therefore it is statistically significant and we reject the null hypothesis. We conclude that the data provide evidence that the treatment group has a higher website visit rate than the control group, and because users were randomly assigned to the treatment and control groups, there is evidence that the advertising treatment caused an increase in website visits.
However, the distribution shows that there are differences that are as large as or even larger than the observed difference of 0.009. This means that they can occasionally occur due to random chance. Our p-value of 0.036 means that, if the null hypothesis was true, differences at least as extreme as the observed difference occurred in about 3.6% of our simulations.
Therefore, rejecting the null hypothesis does not guarantee that it is false. There is still a possibility that we observed an unusual result due to random chance. Rejecting a null hypothesis that is actually true is called a Type I error.
Total Variation Distance
In the previous example we worked with a dataset with two categories of results: visit or no visit. With two categories, the statistic that we want to measure is simple: it's the difference in visit proportion between the two groups. We don't have to separately consider the no-visit proportion because it is determined by the visit proportion.
But suppose there are three categories of results: website visit, no visit, and purchase. Now each group has three proportions, and finding the difference in the proportions of only one category is not sufficient. This is where the Total Variation Distance or TVD comes in. TVD is a way to measure the difference between two distributions while taking into account the proportion of all categories.
To calculate the TVD:
- for each category, find the absolute value of the difference in proportion
- add the absolute values together
- divide the sum by 2
To see how TVD is calculated, we create an example of the three categories of result: visit, no visit, and purchase. We also have their proportions for both treatment and control groups, and the difference between the proportions:
| control | treatment | difference | |
|---|---|---|---|
| no visit | 0.80 | 0.7 | -0.10 |
| visit | 0.14 | 0.2 | 0.06 |
| purchase | 0.06 | 0.1 | 0.04 |
We note that for the control and treatment groups, each proportion adds up to 1.0.
We also note that the decrease in "no visit" (-0.10) is redistributed to "visit" and "purchase" which together increase by the same amount (0.06 and 0.04). Therefore 0.10 or 10% of the proportion has been redistributed among the categories. We could think of this as 10% of the people changing categories to turn the control distribution into the treatment distribution.
If we simply add the difference of the two groups, we would get: −0.10 + 0.06 + 0.04 = 0.0
The differences cancel out because the proportions in both groups have to add up to 1.0.
Instead we add the absolute value of the difference so the decreases and increases do not cancel each other, but this results in a value that's doubled what it should be: 0.10 + 0.06 + 0.04 = 0.20
This counts the redistribution twice: once when the 0.10 leaves no visit, and again when the same 0.10 appears as 0.06 in visit and 0.04 in purchase.
Therefore we divide by 2 to get the actual difference between the two groups.
TVD: 0.1
The TVD gives us one number that summarizes the differences across all the categories.
A TVD of 0 means the distributions are identical. The farther the TVD is from 0, the more different the two distributions are.
Summary
In this chapter we apply our sampling and simulation skills to two-sample hypothesis testing, which is used to compare two groups of data. We also see the importance of random assignment to groups, which allows us to determine whether a treatment may have caused a difference in the outcome.


