16: Classifier
- Page ID
- 50861
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)In the previous chapter we learned the basic ideas behind classification and the k-nearest neighbors (KNN) algorithm.
In this chapter, we will see how to separate data into training and testing sets. Then, using the same automated steps from the previous chapter, we implement a KNN classifier that can make predictions on a variety of datasets, including ones with more than two features.
Training and Testing
Two important steps in working with an algorithm to make predictions are:
- Training the algorithm by providing it with labeled data that it will use to make predictions. The labeled dataset is called the training dataset.
- Testing the algorithm on new data to measure how accurately it makes predictions. The new dataset is called the testing dataset.
For KNN, the algoritm is the sequence of steps to find the k nearest neighbors of a new data point and use the majority class among those neighbors to make the prediction.
Before we train and test an algorithm, we begin with a single dataset that has been gathered and prepared. Then we divide this input dataset into two parts: a training dataset and a testing dataset.
Gather Data
We use the same cdc dataset as in the previous chapter. The original dataset contains 20,000 rows. Because KNN compares every testing row with every training row, classifying all 20,000 rows takes a long time. We will randomly select 3,000 rows so that the classifier runs more quickly while still using a large and varied sample of the data.
Prepare Data
We take the same steps to prepare data as in Chapter 15: convert gender values to numbers, change height and weight to standard units.
Create Training and Testing Data
We use random sampling to create the two datasets. First we randomly shuffle all the rows of the dataset, in the same way that we randomly shuffle data for hypothesis testing in Chapter 10. Then we assign some of the rows to the training dataset and the remaining rows to the testing dataset. For this chapter we will use half of the rows for training and half for testing.
We confirm the training dataset by plotting the height vs weight.
The scatterplot shows only the training data. When KNN makes a prediction, it searches only the training dataset for the nearest neighbors. The testing data are not used to make predictions because we want them to represent new, unseen data. After the predictions have been made, we compare them with the actual classes in the testing dataset to measure the classifier's accuracy.
Implement a Classifier
We now have the necessary datasets to create and test a simple classifier. We use the training and testing datasets above to train and test the classifier.
To write the code for the classification algorithm, we follow the steps that we've used in Chapter 15:
- Find the distance between the new data point and all the existing data points in the dataset.
- Use these distances to find the k points that are closest (the shortest distance) to the new data point.
- Find the class that occurs the most in the k nearest data points. This will be our predicted class for the new data point.
We bring in the distance_from_point function and find_distances function from Chapter 15.
We also add one new function find_k_nearest_neighbors and bring in the majority function from Chapter 15.
We have now automated each step of the KNN algorithm. We've written the functions to calculate distances, find the k nearest neighbors, and determine the majority class. The final function, classify, is the classifier. When we call the classifier, it calls the other functions in turn to carry out each step of the KNN algorithm and return the predicted class.
Test the Classifier
We start by testing the classifier with a new data point.
Predicted class: 0
Visually checking the location (2,2) in the scatterplot, we see that a prediction of class 0 is expected.
Next we use the testing data with our classifier.
First we save the _class (gender) data from testing_set, then we remove the _class data from the testing_set. The remaining columns contain only the feature values that the classifier will use to make its predictions.
Then we test the classifier with every row of the test_set. Each prediction for a data point in test_set is saved in a list of predicted values.
1500 predictions made
Find Prediction Accuracy
We store the actual data and the predicted data in a DataFrame and print the first five rows.
| Actual | Predicted | |
|---|---|---|
| 18280 | 1 | 1 |
| 10229 | 1 | 1 |
| 17305 | 0 | 0 |
| 9053 | 0 | 0 |
| 4971 | 1 | 1 |
Since the data is randomly shuffled before being divided into the training and testing sets, the results above may differ slightly with each run of this example. The predicted classes for individual data points may be different, but the overall accuracy should be similar.
We calculate the percent accuracy by finding the ratio of correct predictions over the total predictions.
Accuracy: 84 %
The classifier correctly predicted the gender for a high percentage of the people in the testing dataset.
Looking at the female and male ratios in the starting dataset of 3000 people, we find the count and the percentage of f and m values:
gender
f 1570
m 1430
Name: count, dtype: int64
gender
f 52.333333
m 47.666667
Name: proportion, dtype: float64
We see that if a classfier always predicted female, it would be correct about 52% of the time. Our classifier achieves a much higher accuracy, showing that it has learned patterns in the data rather than simply predicting the most common class.
The classifier correctly predicts the gender for a high percentage of the people in the testing dataset, but it is not perfect because height and weight alone do not completely distinguish males from females. Many people have similar heights and weights, so some data points are difficult to classify correctly.
Test the Classifier with a Different Dataset
In the previous example we use the classifier to predict gender by using height and weight. One advantage of KNN is that the algorithm does not depend on a particular dataset. As long as we provide a training dataset with feature values and a class label, we can use the same classifier to make predictions on completely different data.
To demonstrate this, we will use the cars dataset that we've used in previous chapters. Instead of predicting gender, we will predict whether a car is fuel efficient. The features will be horsepower, weight, and cylinders, and the classifier will indicate whether the car is fuel efficient or not. Note that we do not need to change the classifier itself, we simply provide it with a different dataset.
| mpg | cylinders | displacement | horsepower | weight | acceleration | model_year | origin | car_name | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 18.0 | 8 | 307.0 | 130.0 | 3504.0 | 12.0 | 70 | 1 | chevrolet chevelle malibu |
| 1 | 15.0 | 8 | 350.0 | 165.0 | 3693.0 | 11.5 | 70 | 1 | buick skylark 320 |
| 2 | 18.0 | 8 | 318.0 | 150.0 | 3436.0 | 11.0 | 70 | 1 | plymouth satellite |
| 3 | 16.0 | 8 | 304.0 | 150.0 | 3433.0 | 12.0 | 70 | 1 | amc rebel sst |
| 4 | 17.0 | 8 | 302.0 | 140.0 | 3449.0 | 10.5 | 70 | 1 | ford torino |
Number of rows: 392
We create a DataFrame with only the features and target we want.
We want to predict whether a car is fuel efficient. Since fuel efficiency is measured by mpg, we will use mpg to define our classes. However, a classifier can only predict categorical data, and mpg is continuous quantitative data, so we convert mpg into two categories:
- 0: mpg < 25, not fuel efficient
- 1: mpg >= 25, fuel efficient
| horsepower | weight | cylinders | mpg | |
|---|---|---|---|---|
| 0 | 130.0 | 3504.0 | 8 | 0 |
| 1 | 165.0 | 3693.0 | 8 | 0 |
| 2 | 150.0 | 3436.0 | 8 | 0 |
| 3 | 150.0 | 3433.0 | 8 | 0 |
| 4 | 140.0 | 3449.0 | 8 | 0 |
In the previous example we used two features, height and weight, to predict gender. In this example, we will use three features: horsepower, weight, and cylinders to predict whether a car is fuel efficient. Even though the dataset and the number of features are different, the KNN algorithm does not change. The same classifier works without modifications on completely different datasets, even if they have a different number of features.
Although we first developed the algorithm using two features, the code is not limited to two dimensions. The distance_from_point function works with arrays of any length, so the same classifier can be used with datasets that have two features, three features, or many more.
We continue preparing the data:
- change the
mpglabel to_classas required by our classifier - convert all the features to standard units
| horsepower | weight | cylinders | _class | |
|---|---|---|---|---|
| 0 | 0.663285 | 0.619748 | 1.482053 | 0 |
| 1 | 1.572585 | 0.842258 | 1.482053 | 0 |
| 2 | 1.182885 | 0.539692 | 1.482053 | 0 |
| 3 | 1.182885 | 0.536160 | 1.482053 | 0 |
| 4 | 0.923085 | 0.554997 | 1.482053 | 0 |
We divide the data into the training and testing sets.
We save the _class column from the testing_set. Then we remove the _class column from the testing data.
| horsepower | weight | cylinders | |
|---|---|---|---|
| 277 | -0.947474 | -0.991973 | -0.862911 |
| 377 | -0.973454 | -1.192113 | -0.862911 |
| 139 | -0.557775 | -0.893080 | -0.862911 |
| 317 | -0.765614 | -0.512812 | -0.862911 |
| 330 | 0.715245 | -0.079567 | 0.309571 |
Now that the data has been prepared, we use the classifier to predict the fuel efficiency of the cars.
Accuracy: 84
We see that the classifier has high accuracy on a completely different dataset. This demonstrates that the KNN algorithm is not tied to a particular dataset and can be reused to solve different classification problems.
Summary
In this chapter, we built a KNN classifier step by step by combining several functions into a single classifier that predicts the class of new data. We used training and testing datasets to measure its accuracy and saw that the same classifier can be applied to completely different datasets without changing the algorithm.
KNN is one of many machine learning algorithms. Machine learning is a foundation of modern AI because it enables computers to learn patterns from data and make predictions.


