13: Linear Regression
- Page ID
- 50858
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)In the previous chapter, we learned how to describe a single variable using measurements such as the mean and standard deviation, and how variation affects the accuracy of estimates. In this chapter, we turn our attention to the relationship between two variables. We will learn how to measure the strength of that relationship with correlation and describe it with a regression line.
Scatterplots
In a dataset with multiple features, or columns describing the data, often we want to know whether two features are related. For example, with the car dataset that we first worked with in Chapter 4, we want to know whether the weight of the car and the fuel efficiency (MPG) are related. Does a heavier car tend to have lower MPG?
A scatterplot is used to help us visually determine if two features are related. Scatterplots were first introduced in Chapter 5, and in this chapter we use them to investigate whether there is an association between two features.
For the car dataset, a scatterplot displays each car as a point, with one feature plotted on the horizontal axis and the other on the vertical axis. When all the cars are plotted together, the points can form a pattern that reveals the direction, form, and strength of the relationship between the two features.
We start with the car dataset. The two features we investigate are mpg and weight.
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 18.0 | 8 | 307.0 | 130.0 | 3504.0 | 12.0 | 70 | 1 | chevrolet chevelle malibu |
| 1 | 15.0 | 8 | 350.0 | 165.0 | 3693.0 | 11.5 | 70 | 1 | buick skylark 320 |
| 2 | 18.0 | 8 | 318.0 | 150.0 | 3436.0 | 11.0 | 70 | 1 | plymouth satellite |
| 3 | 16.0 | 8 | 304.0 | 150.0 | 3433.0 | 12.0 | 70 | 1 | amc rebel sst |
| 4 | 17.0 | 8 | 302.0 | 140.0 | 3449.0 | 10.5 | 70 | 1 | ford torino |
The scatterplot is created by the function
plt.scatter(x_axis_variable, y_axis_variable)
It doesn't matter which variable is placed on the horizontal axis and which is placed on the vertical axis. Swapping the axes will rotate the scatterplot, but the same relationship between the two variables is still visible.
In the scatterplot above, we can see that as the weight increases, the mpg tends to decrease. Although the points do not fall exactly on a straight line, they form a downward trend.
Correlation
The points in the scatterplot above form a pattern with a downward trend: as the weight increases, the mpg tends to decrease. This type of relationship between two features or two variables is called correlation. Correlation describes how closely the two variables are related.
When one variable tends to decrease when the other variable increases, they have a negative correlation. When two variables tend to increase together, they have a positive correlation. If there is no clear relationship between the variables, they have little or no correlation.
The weight and mpg variables have a negative correlation.
We look at the weight and horsepower of the cars.
We see that as the weight increases, the horsepower also tends to increase. They have a positive correlation.
Lastly we look at model_year and acceleration.
The scatterplot above looks different than the first two scatterplots.
Each vertical column of points represents one model year. Since the cars were made only in whole years (1970, 1971, 1972, ...), there are no cars with model years between them, so the points appear in vertical columns instead of being spread evenly across the horizontal axis.
Although there is a very slight upward trend, the points have a wide spread in each year. This means knowing the model year tells us very little about a car's acceleration. We say they have little or no correlation.
Scatterplots in Standard Units
In Chapter 12 we learned about standard units or the z-score. A value in standard units tells us how many standard deviations it is above or below the mean of the dataset.
We wrote the standard_units function to convert every value in an array to standard units:
We will now apply this function to both variables in a scatterplot. Although the numbers on the axes will change, the overall pattern of the scatterplot will remain the same.
We work with the weight and horsepower of the cars.
We convert the weight and horsepower data into standard units and create a new DataFrame to store them.
| weight (z) | horsepower (z) | |
|---|---|---|
| 0 | 0.620540 | 0.664133 |
| 1 | 0.843334 | 1.574594 |
| 2 | 0.540382 | 1.184397 |
| 3 | 0.536845 | 1.184397 |
| 4 | 0.555706 | 0.924265 |
| ... | ... | ... |
| 387 | -0.221125 | -0.480448 |
| 388 | -0.999134 | -1.364896 |
| 389 | -0.804632 | -0.532474 |
| 390 | -0.415627 | -0.662540 |
| 391 | -0.303641 | -0.584501 |
392 rows × 2 columns
Note that the data for both weight (z) and horsepower (z) are now centered around 0. The positive z values come from actual weight or horsepower values that are greater than the mean, and the negative z values come from actual weight or horsepower values that are less than the mean. The numbers have changed, but the relative position of the data values remains the same.
We now display the weight (z) vs. horsepower (z) in a scatterplot to show that the plot in standard units still has the same shape as the plot with actual units in pounds and horsepower shown above.
Although the axes are now measured in standard units instead of in actual units of pounds and horsepower, the scatterplot has the same overall shape as before. This shows that converting data to standard units does not change the relationship between the two variables.
The Correlation Coefficient
A scatterplot visually shows us the relationship between two variables. To measure the strength of that relationship, we calculate the correlation coefficient.
The correlation coefficient is denoted by r and has the following characteristics:
ris a value between 1 and -1.- When
r= 1, all the points lie on a straight line with a positive correlation. - When
r= 0, there is no linear correlation. - When
r= -1, all the points lie on a straight line with a negative correlation.
We observe the shape of the scatterplot and its relationship to the r value below.
The correlation coefficient r is the average of the products of the two variables, after both variables have been converted to standard units, or have been standardized.
Multiplying the two standardized variables captures two aspects of their relationship:
- Direction. If both variables are above their means or both are below their means, their product is positive. If one variable is above its mean and the other is below its mean, their product is negative. When we average these products, a positive average indicates a positive correlation, while a negative average indicates a negative correlation.
- Strength. A large product means that both variables are far from their means at the same time. A small product means that one or both variables are close to their means. Therefore, for the correlation coefficient r to have a large magnitude, many data points must move far from their means together in a consistent direction, which means that they have a strong linear correlation.
We define a function correlation that calculates the r value for 2 sequences of data:
We test the correlation function, with the x1 and y1 data from subplot 1 above, and with the x3 and y3 data from subplot 3 above.
r for subplot 1: 0.94
r for subplot 3: -0.02
The calculated r results correspond with the scatterplots' shapes.
Generally, the closer the value of r is to 1 or −1, the stronger the linear relationship between the variables. The closer r is to 0, the weaker the linear relationship.
There is no universal cutoff for deciding what is considered "strong" or "weak." The scatterplot should always be examined together with the value of r.
The Regression Line
We now look more closely at the weight and mpg of the cars. Here's the scatterplot again:
In the scatterplot we can see that heavier cars generally have lower mpg or lower fuel efficiency.
The points do not fall exactly on a straight line, but they form a clear downward pattern. We can imagine drawing a line that passes through the middle of the points. The line doesn't have to touch all the data points, but it gets as close as possible on average to all the data points.
If we follow such a line, we can predict a car's mpg based on its weight. For example, if the car is 4000 lbs, we can estimate that the car's mpg will be about 15. And if the car is 3000 lbs, we can estimate that the mpg will be about 22.
The line we used above to make these predictions is called the regression line, or the line of best fit. It summarizes the overall linear relationship between the two variables. For each value of weight on the horizontal axis, the regression line gives a predicted value of mpg on the vertical axis.
Calculate the Regression Line
In the previous example, we estimated the regression line visually. However, different people might draw slightly different lines. Linear regression gives us a mathematical method for calculating the line that best fits the data.
In this case weight is the predictor variable because it is used to make the prediction, and mpg is the response variable because it is the value being predicted. The predictor variable is placed on the x-axis, and the response variable is placed on the y-axis.
weight= predictor, on x-axismpg= response, on y-axis
Use Standard Units
Just as the correlation coefficient is calculated using standard units, the regression line is also calculated using standard units. We will later convert the regression line back to the original units.
We convert the weight and mpg into standard units.
When the two data sequences x and y have standard units, then the regression line is described as: ŷ = r × x where r is the correlation coefficient,
and ŷ (called y hat) is the predicted value of yon the regression line for a given value of x.
The equation says that for every value x, we can find the corresponding ŷ on the regression line by multiplying x by r.
If x and y had a perfect positive linear relationship, then r = 1 and the regression line is simply: ŷ = x
However, two variables are usually not perfectly correlated, so the r value adjusts the slope of the line to reflect the direction and strength of the relationship.
For the weight and mpg, the r value calculated with our correlation function from earlier:
r: -0.8322442148315758
We note that r is negative as expected from the downward trend in the scatterplot. It's also reasonably close to -1, which agrees with the strong linear correlation that we observe from the scatterplot.
Now that we've found the r value, given any x value, we can multiply it by r to find the ŷ value of the regression line, which is our predicted y value.
We display the same scatterplot and add a calculated regression line in red.
For the red regression line above, which is describe as: ŷ = r × x
- the slope of the line is
r. The sign ofrdetermines the direction of the line. Sinceris negative, the line slopes downward: as x increases, the predicted value ŷ decreases. The closerris to −1, the steeper the downward slope. - the intercept of the line is at (0, 0). The intercept is where the line crosses the y-axis. Because both x and y are measured in standard units, their means are 0. Therefore, when x=0 (or the mean of x), the predicted value is also ŷ=0 (or the mean of y).
Convert to Original Units
The regression line: ŷ = r × x predicts values only in standard units. To make predictions with the original data, we must convert the regression line back to the original units.
Recall the definition of the standard units: $$ \text{standard units} = \frac{ data - mean } { \text{standard deviation}} $$
Substitute the definition of standard_units for ŷ and x in the equation for the regression line: $$ \frac{\hat{y}- mean_{\hat{y}}}{SD_{y}} = r \times \frac{x- mean_{x}}{SD_{x}} $$
Solve for ŷ: $$ \hat{y} = r \times \frac{SD_{y}}{SD_{x}} \times (x-mean_{x}) + mean_{\hat{y}} $$
Distribute (x − meanx) $$ \hat{y} = (r \times \frac{SD_{y}}{SD_{x}} \times x) - (r \times \frac{SD_{y}}{SD_{x}} \times mean_{x}) + mean_{\hat{y}} $$
Rearrange into the standard form of a line: y = mx + b $$ \hat{y} = (r \times \frac{SD_{y}}{SD_{x}}) \times x + (mean_{\hat{y}} - (r \times \frac{SD_{y}}{SD_{x}} \times mean_{x})) $$
In the equation of the line: y = mx + b
m is the slope and b is the intercept.
Therefore the slope of the regression line is: $$ slope = r \times \frac{SD_{y}}{SD_{x}} $$
The intercept is: $$ intercept = mean_{\hat{y}} - (r \times \frac{SD_{y}}{SD_{x}} \times mean_{x}) $$ Or: intercept = meanŷ − slope × meanx
We can now write a function to find the slope and the intercept:
We now apply the functions to the cars weight and mpg data. First we review the dataset.
Displaying the scatterplot for weight vs mpg, along with the regression line.
With the linear regression line, we now have a consistent way to predict the MPG or ŷ, given a car's weight, which is x.
Predicted MPG for a car weighing 4000 lbs: 15.6
Meaning of the Slope
With the weight vs mpg data, the slope is:
-0.01
The slope of the regression line is −0.01. This means that, on average, for every additional pound of weight, the predicted MPG decreases by about 0.01 MPG. Equivalently, a car that is 100 pounds heavier is predicted to get about 1 MPG less.
In general, the slope of a regression line represents the average change in the predicted value of y for each one-unit increase in x. A positive slope means the predicted value of y increases as x increases, while a negative slope means the predicted value of y decreases as x increases.
Summary
In this chapter, we learned how to investigate the relationship between two variables using scatterplots, the correlation coefficient, and the regression line. We saw how scatterplots reveal patterns in the data, how the correlation coefficient measures the direction and strength of a linear relationship, and how the regression line can be used to predict one variable from another.


