9: Hypothesis Testing
- Page ID
- 50854
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)In this chapter we will apply our understanding of random samples and empirical distributions to test a hypothesis about a population. A hypothesis is an assumption that we make about a dataset's parameter. An example hypothesis is: people who exercise regularly report lower stress levels.
In hypothesis testing, we use a statistical process to decide whether observed data from a sample provide enough evidence to support a hypothesis about a larger population.
Steps of Hypothesis Testing
Hypothesis testing begins when we have a yes/no question to answer about a parameter of a dataset. Some examples include:
- Do people who exercise regularly have lower stress levels than those who don't exercise?
- Do students who study with background music remember the material better than those who study with no background music?
- In sports, do the home teams win more often?
The question involves two different results and is about whether the difference is due to random chance or due to an influence.
Hypothesis testing involves these steps:
- State the hypothesis: the null hypothesis and alternative hypothesis
- Build a model of the data for the null hypothesis
- Obtain the observed data
- Compare the observed data to the model's data and calculate the p-value
- Draw a conclusion
We will explore how hypothesis testing works with a coin toss experiment.
Suppose we have a coin and want to know if it's a fair coin, which means that it has an equal chance of landing heads or tails. We'll use the steps of hypothesis testing to see if the coin is fair.
State the Hypothesis
Since the yes/no question we try to answer has two results, such as win/loss data when a sport team plays at home and win/loss data when the team plays away from home, hypothesis testing includes two hypotheses:
- The null hypothesis says that there is no significant difference between the two results. Any difference is due to random chance rather than due to the influence of a real phenomenon.
The term null means that any difference is due to nothing but chance. - The alternative hypothesis says that any difference between the two results is due to some influence rather than chance.
For the coin toss experiment:
- The null hypothesis: The coin is fair, there is no influence making heads or tails more likely. P(heads) is 0.5.
- The alternative hypothesis: The coin is not fair and something influences it to land on one side more often. P(heads) is not 0.5.
Build a Model
The model shows what the data would look like if the null hypothesis is true.
For the coin toss experiment, we build a model that shows the likely result if the coin is fair. But before we use the model for hypothesis testing, it is important to choose a model and assess that it is valid for our data.
Assess the Model
We want a model that truly reflects a fair coin being tossed. In Chapter 8 we see that as the sample size increases, the empirical distribution approaches the theoretical distribution. Therefore, our model will simulate a coin being tossed 500 times and records the number of heads. Then we repeat the simulation 10,000 times to get 10,000 numbers of heads. Because we repeat the simulation a large number of times, the overall distribution of our results will be similar to the theoretical distribution for a fair coin.
Use the Model
In the code below we re-use the coin toss simulation function from Chapter 8.
model_heads_arr contains 10,000 numbers, each number is the number of heads from one coin toss simulation.
Next we plot the model's distribution of the number of heads.
Because the coin is tossed many times, the distribution of the number of heads is a bell curve, as can be seen above (why it's a bell curve is discussed in Chapter 12). The bell curve shows that the majority of the simulations produce between 240 and 260 heads, which corresponds with the expected value of 250 heads out of 500 coin tosses.
Obtain the Observed Data
To truly test a real coin, we would take the coin and toss it 500 times, recording each result.
But for this exercise we'll simulate the coin toss and pretend this is our actual coin toss results with the real coin.
Number of heads: 261
Number of tails: 239
We bring back the model's distribution of the number of heads, and plot our observed number of heads from tossing the actual coin being tested.
In the plot we can clearly see that the observed result, the red circle, is in the range of 240-260 heads expected for a fair coin, so it looks like a fair coin.
Compare Data and Calculate p-value
We can also measure and compare the observed data and the model's data. To do this, we need to identify what the test statistic is. Recall from Chapter 8 that a statistic is a measurement of the data that summarizes an aspect of the data, such as the mean height of people in the dataset. In hypothesis testing the test statistic measures of how far the data is from what we expect in the null hypothesis.
The null hypothesis says that there's no influence that makes the observed data different from the model's data. This means the observed data (from tossing the tested coin) should be similar to the model's data (from tossing the fair coin). And the model's data from tossing the fair coin says that out of 500 coin tosses, we expect 250 will be heads.
Given that 250 is what we expect in the null hypothesis, the test statistic is how far the number of heads is from 250. This means we find the difference between the number of heads and 250.
Comparing the data involves 3 steps:
- find the difference between the observed number of heads and 250. This is the observed distance.
- find the difference between each of the 10,000 simulated numbers of heads and 250. This is an array of simulated distances.
- then we find what proportion of the simulated distances is larger (or farther from 250) than the observed distance.
If the proportion is large (>= 0.05), then the observed number is well within the range of what the null hypothesis says, since the simulated result is the data of the null hypothesis.
If the percentage is small (< 0.05), then the observed number is outside the range of what the null hypothesis says.
This proportion is called the p-value. If the p-value is large, then we can reject the relative hypothesis.
And if the p-value is small, it is statistically significant and we can reject the null hypothesis.
p-value: 0.3487
In the calculation above:
- We use the absolute value of the difference because we want to find the distance from 250 only, we don't need to know if it's on the right or left of 250.
- To calculate the p-value, since
simulated_distancesis a numpy array, the comparisonsimulated_distances >= observed_distanceresults in a numpy array of 0s and 1s. Adding all the 1s (or the counts of when the condition is true) and dividing by 10,000 gives us the proportion of when thesimulated_distancesare larger than or equal to theobserved_distance. - Since the numpy array contains 0s and 1s, adding all the data in the array is the same as adding all the 1s. And if we add all the data in the array and divide the total by the count of data in the array, we effectively calculate the mean of the data in the array.
To see how the observed_distance is compared to the simulated_distances of a fair coin, we show where the observed_distance is in the distribution of the simulated distances
We can see that the observed_distance is in the typical range simulated_distances and not in the right tail end. This is why the p-value is large: there are many simulated_distances that are larger than the observed_distance.
And with the large p-value, we can reject the alternative hypothesis, which states that something influences the coin to land on one side more often. Therefore there isn't enough evidence to conclude that the coin is biased.
Statistically Significant p-value
We now investigate what will happen if we have a biased coin.
We simulate the biased coin by making it likely to land heads 55% of the time, and keep everything else the same.
Number of heads: 272
Number of tails: 228
p-value: 0.0551
Since the coin has a small bias such that the P(heads) is 0.55 instead of 0.5, it's a good idea to run the code above a few times and observe the output. The p-value can change with each simulation of the biased coin toss, but generally it's below the 0.05 threshold for p-value.
When the p-value is less than 0.05, we look at where the observed distance is compared to the distribution of the simulated distances.
When the observed distance is at the tail end of the distribution of simulated distances, we can see that very few simulated distances are larger than the observed distance, causing the p-value to be small and statistically significant. This means we reject the null hypothesis, and the coin is biased.
Error Probability
Because the coin is tossed many times, the distribution of the number of heads is a bell curve, as discussed above and shown again here:
When the data has the bell curve shape or is a normal distribution, then 95% of the data is within 2 standard deviations from the mean of the data (discussed in Chapter 12). This means the remaining 5% of the data occurs much less frequently and that is why the p-value cutoff is 0.05.
For our simulated data, 2 standard deviations is:
2 standard deviations or p-value cut off: 22.373348608556565
Using our p-value cutoff of 0.05, this means if a fair coin is tossed 500 times and the number of heads is above 250+22 or 272, or below 250-22 or 228, we will conclude incorrectly that the coin is not a fair coin.
The histogram of the simulated distances from the expected value also shows that we can incorrectly conclude that it's not a fair coin.
With the dashed red line as the p-value cut off point, we see that a fair coin still has a small chance that its distance from the expected value would be large enough that we would conclude incorrectly that the coin is biased.
Therefore the p-value cut off is also called the error probability, or the probability that we could be in error with our conclusion.
Being aware of the error probability is important. For example, it means that any conclusion made from breakthrough scientific research should not rely on a single experiment. The experiment needs to be replicated, in case the error probability leads to the wrong conclusion the first time.
Summary
In this chapter we applied sampling and simulation to build a model to do hypothesis testing. The simulated data in our model can reject or fail to reject the null hypothesis, where the hypothesis is our assumption about a parameter of the data. When we draw a conclusion about the hypothesis, we need to keep in mind that there is an error probability caused by the nature of random variation in our data.


