Skip to main content
Workforce LibreTexts

10: Two-Sample Tests

  • Page ID
    50855
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\ket}[1]{\left| #1 \right>}\)
    \(\newcommand{\bra}[1]{\left< #1 \right|}\)
    \(\newcommand{\braket}[2]{\left< #1 \vphantom{#2} \right| \left. #2 \vphantom{#1} \right>}\)
    \(\newcommand{\braopket}[3]{\left< #1 \vphantom{#2}\vphantom{#3} \right| #2 \vphantom{#1}\vphantom{#3} \left| #3 \vphantom{#1}\vphantom{#2} \right>}\)
    \(\newcommand{\qmvec}[1]{\mathbf{\vec{#1}}}\)
    \(\newcommand{\op}[1]{\hat{\mathbf{#1}}}\)
    \(\newcommand{\expect}[1]{\langle #1 \rangle}\)
    \(\newcommand{\dfn}[1]{\emph{\textbf{#1}}}\)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    Up to now we've done statistical analysis with one data sample, such as the number of heads in a coin toss experiment. But often data scientists also need to make decisions based on comparing two data samples with each other, such as when comparing two groups of students who studied with background music and with no background music.

    In this chapter we use hypothesis testing with two distinct groups of data to determine if the differences between them are just due to random chance or if the difference is meaningful.

    Gather Data

    The data are from the US Center for Disease Control and Prevention (CDC) and is stored at OpenIntro. The CDC annually surveys the health status and habits of 350,000 people in the US, and the dataset is a random sample of 20,000 people from the survey conducted in 2000.

    import pandas as pd
    import matplotlib.pyplot as plt
    import numpy as np
    
    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/cdc.csv"
    health = pd.read_csv(url)
    display(health.head())
    print("Number of rows:", health.shape[0])
    Output?
      genhlth exerany hlthplan smoke100 height weight wtdesire age gender
    0 good 0 1 0 70 175 175 77 m
    1 good 0 1 1 64 125 115 33 f
    2 good 1 1 1 60 105 105 49 f
    3 good 1 1 0 66 132 124 42 f
    4 very good 0 1 0 61 150 130 55 f
    Number of rows: 20000
    

    The rows from left to right are: general health, any exercise in the past month, health plan coverage, smoke at least 100 cigarettes in respondent's life, height, weight, desired weight, age, and gender.

    We will look at two features: exercise habit or exerany (0 for no exercise in past month, 1 for exercise in past month), and weight (a quantitative variable). We want to know whether people who exercise have a lower body weight distribution than people who don't exercise.

    Exploratory Data Analysis (EDA)

    It might seem reasonable that people who exercise have lower body weight than those who don't exercise. However, a person's weight is influenced by their height. Taller people naturally tend to weigh more than shorter people, no matter how much they exercise.

    Assessing the Model

    If we only look at the weight, then when we build the model for the null hypothesis, if one of our sample groups happen to have slightly higher average height, then their weight would tend to be higher regardless of their exercise habit. We also need to take into account the differences in height when comparing the two groups.

    In this case the height is a confounding variable. A confounding variable is an extra factor that can bias the results when studying only weight and exercise.

    To take into account both the height and weight, we will use the Body Mass Index (BMI) instead. The formula for BMI is \[\text{BMI} = \dfrac{\text{weight}}{\text{height}^2} \times 703\] With the BMI value, the weight is scaled relative to the height.

    We add another column to the health DataFrame to store the BMI.

    # find the BMI from the weight and height
    BMI = health.weight / (health.height ** 2) * 703
    
    # add to the health DataFrame
    health['bmi'] = BMI
    display(health.head(10))
    Output?
      genhlth exerany hlthplan smoke100 height weight wtdesire age gender bmi
    0 good 0 1 0 70 175 175 77 m 25.107143
    1 good 0 1 1 64 125 115 33 f 21.453857
    2 good 1 1 1 60 105 105 49 f 20.504167
    3 good 1 1 0 66 132 124 42 f 21.303030
    4 very good 0 1 0 61 150 130 55 f 28.339156
    5 very good 1 1 0 64 114 114 55 f 19.565918
    6 very good 1 1 0 71 194 185 31 m 27.054553
    7 very good 0 1 0 67 170 160 45 m 26.622856
    8 good 0 1 1 65 150 130 27 f 24.958580
    9 good 1 1 0 70 180 170 44 m 25.824490

    We see that on row index 0 someone weighs 175 lbs, which is heavier than the person on row index 8, whose weight is 150. But because of their height difference, they have very similar BMI. Now we can compare standardized body mass (BMI) against exercise habit.

    Prepare Data

    Next we concentrate on the two features, exercise habit and BMI, by creating a new DataFrame exercise that only has the exerany and bmi columns.

    exercise_bmi = health[['exerany', 'bmi']]
    display(exercise_bmi.head())
    print("Number of rows:", exercise_bmi.shape[0])
    Output?
      exerany bmi
    0 0 25.107143
    1 0 21.453857
    2 1 20.504167
    3 1 21.303030
    4 0 28.339156

    Number of rows: 20000

     

    We find the count and proportion of exercise and non-exercise groups.

    # Find the exercise and no exercise groups
    exercise = exercise_bmi[exercise_bmi["exerany"] == 1]
    no_exercise = exercise_bmi[exercise_bmi["exerany"] == 0]
    
    print("No exercise (0): count =", len(exercise), "  proportion =", round(len(exercise)/len(exercise_bmi),2))
    print("Exercise (1): count =", len(no_exercise), "  proportion = ", round(len(no_exercise)/len(exercise_bmi),2))
    Output?
    No exercise (0): count = 14914   proportion = 0.75
    Exercise (1): count = 5086   proportion =  0.25
    

    Then we plot the distribution of the BMI of the two groups.

    plt.figure(figsize=(6, 4))
    plt.hist(no_exercise["bmi"], bins=30, density=True, alpha=0.5, label="No exercise")
    plt.hist(exercise["bmi"], bins=30, density=True, alpha=0.5, label="Exercise")
    plt.legend()
    plt.title("BMI Distribution by Exercise Habit")
    plt.xlabel("BMI")
    plt.ylabel("Density")
    plt.grid()
    plt.show()
    Output?

    It looks like the BMI distributions overlap quite a bit for the exercise and non-exercise groups. We also see that there seems to be a right tail in the distribution, which means that there are some people with BMI values that are much higher than those of most people.

    The Test Statistic

    To compare the exercise and no exercise group's BMI, we find the difference in the mean BMI of the two groups.

    # Calculate the mean BMI for each group
    mean_bmi_exercise = exercise["bmi"].mean()
    mean_bmi_no_exercise = no_exercise["bmi"].mean()
    observed_bmi_difference = mean_bmi_exercise - mean_bmi_no_exercise
    
    print("Mean BMI for exercise group:", round(mean_bmi_exercise, 2))
    print("Mean BMI for no exercise group:", round(mean_bmi_no_exercise, 2))
    print("Difference (exercise - no exercise):", round(observed_bmi_difference, 2))
    Output?
    Mean BMI for exercise group: 25.98
    Mean BMI for no exercise group: 27.27
    Difference (exercise - no exercise): -1.29
    

    There is a small difference: the mean BMI of the exercise group is lower than the mean BMI of the no exercise group by 1.29.

    Is the 1.29 difference due to random chance or is there evidence of a difference between the two groups?
    We will use hypothesis testing to find the answer.

    Hypothesis Testing

    We'll use the same steps as in Chapter 9, but this time we're working with two data samples so our model for the null hypothesis will be built differently.

    State the Hypothesis

    The null hypothesis: There is no difference between the mean BMI of the exercise and no exercise groups. Any difference is due to random chance.
    The alternative hypothesis: There is a difference between the mean BMI of the exercise and no exercise groups.

    Build the Model

    Recall from Chapter 9 that we want to build a model for the null hypothesis, which means the model should show us what the test statistic looks like if there is no difference in the mean BMI of the exercise and no exercise groups.

    If there is no difference, then whether a person is labeled as "exercise" or "no exercise" should not be associated with their BMI. We can model this by randomly shuffling the exerany (the exercise habit) column among the people in the dataset. This is called random permutation.

    Random Permutation Example

    To see how random permutation works, we copy the first seven rows of the exercise_bmi DataFrame into an example Dataframe so we can randomly shuffle the exerany column.

    example = pd.DataFrame(exercise_bmi.head(7))
    display(example)
    Output?
      exerany bmi
    0 0 25.107143
    1 0 21.453857
    2 1 20.504167
    3 1 21.303030
    4 0 28.339156
    5 1 19.565918
    6 1 27.054553

    Then we shuffle the exerany column by using the same sample function discusssed in Chapter 8:
    sample_DataFrame = a_DataFrame.sample(n = N)
    But this time we use replace=False for sampling with no replacement, and we use frac=1 so 100% of the data are selected.

    shuffled_exerany = example["exerany"].sample(frac=1, replace=False)
    display(shuffled_exerany)
    Output?
    5    1
    2    1
    4    0
    0    0
    6    1
    3    1
    1    0
    Name: exerany, dtype: int64

    Note that the index values show us that the rows of exerany had been shuffled randomly.

    We combine the shuffled_exerany column with the original example DataFrame.

    # add the shuffled column to example,
    # and reset the index so it matches the index of the DataFrame
    example["shuffled_exerany"] = shuffled_exerany.reset_index(drop=True)
    display(example)
    Output?
      exerany bmi shuffled_exerany
    0 0 25.107143 1
    1 0 21.453857 1
    2 1 20.504167 0
    3 1 21.303030 0
    4 0 28.339156 1
    5 1 19.565918 1
    6 1 27.054553 0

    As the example DataFrame shows, the BMI of each person is the same, but after shuffling, the person has been randomly assigned to the exercise or no exercise group. Note also that the shuffle guarantees that we will end up with the same number of people in each group. The original exerany feature has four people in the exercise group and three people in the no exercise group, and the shuffled_exerany feature has the same counts.

    Build the Null Distribution

    After shuffling the exerany feature of the dataset, we calculate the mean BMI of the two shuffled groups and find the difference. This difference is due to random chance, as assumed by the null hypothesis.

    Then we repeat the two steps: 1) shuffle and 2) find the mean BMI difference between the two groups, so we can have a large number of differences. These differences will show us the range of values we would expect from random chance.

    To get ready to repeat the two steps of shuffling and finding the difference, we write a function that does the two steps.

    def one_permutation(dataframe) :
      # create a copy of the original dataframe to add a new shuffled_treatment column
      dataframe_permutation = dataframe.copy()
    
      # create a shuffled_exerany column
      shuffled_exerany = dataframe_permutation["exerany"].sample(frac=1, replace=False)
      dataframe_permutation["shuffled_exerany"] = shuffled_exerany.reset_index(drop=True)
    
      # separate the data into two shuffled groups
      exercise = dataframe_permutation[dataframe_permutation["shuffled_exerany"] == 1]
      no_exercise = dataframe_permutation[dataframe_permutation["shuffled_exerany"] == 0]
    
      # calculate the mean BMI of each group
      exercise_mean_bmi = exercise["bmi"].mean()
      no_exercise_mean_bmi = no_exercise["bmi"].mean()
    
      # return the difference between the mean BMI values
      return exercise_mean_bmi - no_exercise_mean_bmi
    Output?

    We repeat the permutation 1000 times to get enough simulated differences that can occur by random chance.

    # create a list to store the 1000 differences
    differences = []
    
    for i in range(1000):
        one_difference = one_permutation(exercise_bmi)
        differences.append(one_difference)
    mean_bmi_difference_array = np.array(differences)
    Output?

    We plot the distribution of the differences.

    plt.figure(figsize=(4, 4))
    plt.hist(mean_bmi_difference_array, density=True, edgecolor='black')
    plt.title("Distribution of Permutation Differences")
    plt.xlabel("Difference in Mean BMI")
    plt.ylabel("Density")
    plt.grid()
    plt.show()
    Output?

    As expected, if the null hypothesis is true and the difference in mean BMI of the two groups is due to random chance, then the difference tends to be close to 0. This can be seen int the distribution above, with most of the differences between -0.15 and 0.15.

    Compare Observed Data and the Model's Data

    When we calculated the test statistic with the observed data in an earlier step, we found that
    Exercise group mean BMI - No exercise group mean BMI = -1.29

    Given the null distribution above, -1.29 is very far outside the range of what the 1000 permutations produced by random chance. Therefore, it's unlikely that the difference in the observed data is due to random chance.

    p-value Calculation

    We can calculate the p-value to see if it agrees with our conclusion above. We use the same calculation of p-value as in Chapter 8.

    observed_distance = abs(observed_bmi_difference)
    simulated_distance = np.abs(mean_bmi_difference_array)
    p_value = np.mean(simulated_distance >= observed_distance)
    print(p_value)
    Output?
    0.0
    

    The calculated p-value of 0.0 is expected. With the 1000 permutations that we did, none of them had a distance larger than 1.29. But this doesn't mean that the p-value is exactly 0.0. It appears 0.0 for 1000 simulations, so we can say that the p-value is smaller than 1/1000 or smaller than 0.001.

    With a p-value value that's less than 0.001, the p-value is statistically significant. This agrees with our conclusion from observing the null distribution, therefore we reject the null hypothesis and conclude that there is evidence of a difference in mean BMI between the exercise and no exercise groups.

    Observational Data

    It is important to note that our hypothesis testing above provides evidence that there is an association between BMI value and exercise habit. There seems to be a difference in BMI values among people who exercise and those who don't, but we cannot say that exercise causes the difference in BMI. This is because people were not randomly assigned to the exercise or no exercise group.

    If people are randomly assigned to a group, the people who have healthy diets can be assigned to either group, and older and younger people can be assigned to either group. Random assignment tends to balance these other characteristics between the groups, making exercise habit one of the primary difference between the groups. This makes it possible to infer that exercise may cause a difference in mean BMI.

    On the other hand, when people are surveyed for their health condition, as in our dataset, then the two groups can be different in many ways. For example, people who exercise may also have healthy diets or are younger. These factors, the confounding variables, may also influence BMI so it's not possible to say that exercise itself causes the difference in BMI.

    The dataset for the exercise habit and mean BMI is observational data. They are data gathered by measuring or recording behaviors, events, or phenomena as they naturally occur, without manipulating any chracteristics of the data. Observational data can only show evidence of association between two factors, they cannot be used to determine cause and effect.

    Causality and A/B Testing

    We now look at an example of experimental data, which is the opposite of observational data. In experimental data people are randomly assigned to two groups, and this makes it possible to infer a cause and effect relationship as discussed above.

    The Criteo Uplift Prediction dataset contains data from online advertising experiments conducted by Criteo, a digital advertising company. Users were randomly assigned to either an advertising treatment group or a control group. Criteo could show advertisements to users in the treatment group but withheld its advertising from users in the control group. Then the company recorded whether the user visited the advertiser's website.

    This type of randomized experiment is commonly called an A/B test. A/B testing is a type of randomized experiment used to compare two alternatives. Participants are randomly assigned to one of two groups:

    • Group A (control): receives the existing or control condition.
    • Group B (treatment): receives the new condition being tested.

    We then compare an outcome from the two groups to determine whether the treatment made a difference. Because participants are randomly assigned, A/B testing can provide evidence that the treatment caused a difference in the outcome.

    A/B testing is widely used in business, technology, and marketing to determine whether a change causes an improvement in a desired outcome.

    Gather Data

    The original dataset contains millions of records and many features. For this example, we use a random sample of 20,000 users and focus on two features:

    • treatment: whether the user was randomly assigned to the control group (0) or advertising treatment group (1)
    • visit: whether the user did not visit (0) or visited (1) the advertiser's website
    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/ad.csv"
    ad = pd.read_csv(url)
    print("Last 10 rows")
    display(ad.tail(10))
    print("Number of rows:", ad.shape[0])
    Output?
    Last 10 rows
    
      treatment visit
    19990 1 0
    19991 1 0
    19992 1 1
    19993 1 0
    19994 1 0
    19995 1 0
    19996 1 0
    19997 1 0
    19998 1 0
    19999 1 0
    Number of rows: 20000
    

    In the last 10 rows, the users were all in the treatment group, and only one of them visited the website.

    The Test Statistic

    We first see how many users from each group visited or not visited the website.

    ad.value_counts()
    Output?
    treatment  visit
    1          0        16136
    0          0         2916
    1          1          828
    0          1          120
    Name: count, dtype: int64

    We see that for the treatment group, 16,136 people did not visit the website and 828 visited.
    For the control group, 2916 people did not visit the website and 120 people visited.

    For each group we find the rate of visit, which is the proportion \[ \frac{\text{number of visits}}{\text{total count of visits and no visits}} \nonumber\]

    We observe that since the data is 1 or 0 only, if we find the sum of the visits (all the 1 values + all the 0 values), the result equals the sum of all the 1 values, which is also the number of visits. This means the fraction above can be written as \[\frac{\text{sum of visits}}{\text{total count of visits and no visits}} = \text{the mean of visits} \nonumber\]

    The test statistic is the difference in the proportion of website visits for each group. Does the treatment group have a higher proportion of website visits than that of the control group?

    # select a group's visit column and find the mean
    control_visit_proportion = round(ad[ad["treatment"] == 0]["visit"].mean(), 4)
    treatment_visit_proportion = round(ad[ad["treatment"] == 1]["visit"].mean(), 4)
    observed_visit_difference = round(treatment_visit_proportion - control_visit_proportion, 4)
    print("Proportion of control group that visited the website", control_visit_proportion)
    print("Proportion of treatment group that visited the website", treatment_visit_proportion)
    print("Observed difference:", observed_visit_difference)
    Output?
    Proportion of control group that visited the website 0.0395
    Proportion of treatment group that visited the website 0.0488
    Observed difference: 0.0093
    

    We see that the observed difference between the two groups is only 0.0093 or 0.93%. This looks like a small percentage, but we don't know yet if it's due to random chance. We will use hypothesis testing to investigate.

    Hypothesis Testing

    The question we ask about the treatment and control groups:
    Is the difference in visit proportions likely due to random chance, or is there evidence that advertising increases the visit proportion?

    State the Hypothesis

    The null hypothesis: There is no difference in the proportion of website visits between the treatment group and control group. Any difference is due to random chance.
    The alternative hypothesis: There is a difference in the proportion of the website visits between the treatment and control groups.

    Build the Model

    Just as in the previous example with the mean BMI difference, we will randomly shuffle the treatment column while leaving the visit column unchanged for one permutation, then we find the difference in visit proportion between the two groups. Then we repeat the permutation 1000 times to get the distribution of differences.

    def one_permutation(dataframe) :
      # create a copy of the original dataframe to add a new shuffled_treatment column
      dataframe_permutation = dataframe.copy()
    
      # create a shuffled_treatment column
      shuffled_treatment = dataframe_permutation["treatment"].sample(frac=1, replace=False)
      dataframe_permutation["shuffled_treatment"] = shuffled_treatment.reset_index(drop=True)
    
      # separate the data into two shuffled groups
      treatment = dataframe_permutation[dataframe_permutation["shuffled_treatment"] == 1]
      control = dataframe_permutation[dataframe_permutation["shuffled_treatment"] == 0]
    
      # calculate the mean BMI of each group
      treatment_proportion = treatment["visit"].mean()
      control_proportion = control["visit"].mean()
    
      # return the difference between the mean BMI values
      return treatment_proportion - control_proportion
    Output?

    We repeat the permutation to get 1000 differences and plot them.

    # create a list to store the 1000 differences
    differences = []
    
    for i in range(1000):
        one_difference = one_permutation(ad)
        differences.append(one_difference)
    visit_difference_array = np.array(differences)
    Output?
    plt.figure(figsize=(6, 4))
    plt.hist(visit_difference_array, density=True, edgecolor='black')
    plt.title("Distribution of Visit Proportion Differences")
    plt.xlabel("Difference in Visit Proportion")
    plt.ylabel("Density")
    plt.grid()
    plt.show()
    Output?

    Compare the Observed Data and the Model's Data

    If the null hypothesis is true, then the difference in visit proportion between the two groups would be centered around 0. The distribution shows that most of the differences are in the range of -0.005 and 0.005. Our observed difference is about 0.009, which is near the right tail of the distribution. This means the observed difference is likely rare.

    We calculate the p-value to confirm that it is rare.

    observed_distance = abs(observed_visit_difference)
    simulated_distance = np.abs(visit_difference_array)
    p_value = np.mean(simulated_distance >= observed_distance)
    print(p_value)
    Output?
    0.022
    

    Conclusion

    The p-value of 0.036 is smaller than the significance level of 0.05. Therefore it is statistically significant and we reject the null hypothesis. We conclude that the data provide evidence that the treatment group has a higher website visit rate than the control group, and because users were randomly assigned to the treatment and control groups, there is evidence that the advertising treatment caused an increase in website visits.

    However, the distribution shows that there are differences that are as large as or even larger than the observed difference of 0.009. This means that they can occasionally occur due to random chance. Our p-value of 0.036 means that, if the null hypothesis was true, differences at least as extreme as the observed difference occurred in about 3.6% of our simulations.

    Therefore, rejecting the null hypothesis does not guarantee that it is false. There is still a possibility that we observed an unusual result due to random chance. Rejecting a null hypothesis that is actually true is called a Type I error.

    Total Variation Distance

    In the previous example we worked with a dataset with two categories of results: visit or no visit. With two categories, the statistic that we want to measure is simple: it's the difference in visit proportion between the two groups. We don't have to separately consider the no-visit proportion because it is determined by the visit proportion.

    But suppose there are three categories of results: website visit, no visit, and purchase. Now each group has three proportions, and finding the difference in the proportions of only one category is not sufficient. This is where the Total Variation Distance or TVD comes in. TVD is a way to measure the difference between two distributions while taking into account the proportion of all categories.

    To calculate the TVD:

    • for each category, find the absolute value of the difference in proportion
    • add the absolute values together
    • divide the sum by 2

    To see how TVD is calculated, we create an example of the three categories of result: visit, no visit, and purchase. We also have their proportions for both treatment and control groups, and the difference between the proportions:

    example = pd.DataFrame(
        {
            "control": [0.8, 0.14, 0.06],
            "treatment": [0.7, 0.2, 0.1],
            "difference":[-0.10, 0.06, 0.04]
        },
        index=["no visit", "visit", "purchase"],
    )
    display(example)
    Output?
      control treatment difference
    no visit 0.80 0.7 -0.10
    visit 0.14 0.2 0.06
    purchase 0.06 0.1 0.04

    We note that for the control and treatment groups, each proportion adds up to 1.0.
    We also note that the decrease in "no visit" (-0.10) is redistributed to "visit" and "purchase" which together increase by the same amount (0.06 and 0.04). Therefore 0.10 or 10% of the proportion has been redistributed among the categories. We could think of this as 10% of the people changing categories to turn the control distribution into the treatment distribution.

    If we simply add the difference of the two groups, we would get: −0.10 + 0.06 + 0.04 = 0.0
    The differences cancel out because the proportions in both groups have to add up to 1.0.

    Instead we add the absolute value of the difference so the decreases and increases do not cancel each other, but this results in a value that's doubled what it should be: 0.10 + 0.06 + 0.04 = 0.20
    This counts the redistribution twice: once when the 0.10 leaves no visit, and again when the same 0.10 appears as 0.06 in visit and 0.04 in purchase.

    Therefore we divide by 2 to get the actual difference between the two groups.

    # calculate the TVD
    
    absolute_difference = np.abs(example["treatment"] - example["control"])
    total = np.sum(absolute_difference)
    tvd = total / 2
    print("TVD:", round(tvd, 2))
    Output?
    TVD: 0.1
    

    The TVD gives us one number that summarizes the differences across all the categories.

    A TVD of 0 means the distributions are identical. The farther the TVD is from 0, the more different the two distributions are.

    Summary

    In this chapter we apply our sampling and simulation skills to two-sample hypothesis testing, which is used to compare two groups of data. We also see the importance of random assignment to groups, which allows us to determine whether a treatment may have caused a difference in the outcome.


    This page titled 10: Two-Sample Tests was last modified on Sat, 26 Sep 2026 07:55:57 GMT and is shared under a CC BY 4.0 license and was authored, remixed, and/or curated by Clare Nguyen.