Skip to main content
Workforce LibreTexts

16: Classifier

  • Page ID
    50861
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    In the previous chapter we learned the basic ideas behind classification and the k-nearest neighbors (KNN) algorithm.

    In this chapter, we will see how to separate data into training and testing sets. Then, using the same automated steps from the previous chapter, we implement a KNN classifier that can make predictions on a variety of datasets, including ones with more than two features.

    Training and Testing

    Two important steps in working with an algorithm to make predictions are:

    • Training the algorithm by providing it with labeled data that it will use to make predictions. The labeled dataset is called the training dataset.
    • Testing the algorithm on new data to measure how accurately it makes predictions. The new dataset is called the testing dataset.

    For KNN, the algoritm is the sequence of steps to find the k nearest neighbors of a new data point and use the majority class among those neighbors to make the prediction.

    Before we train and test an algorithm, we begin with a single dataset that has been gathered and prepared. Then we divide this input dataset into two parts: a training dataset and a testing dataset.

    Gather Data

    We use the same cdc dataset as in the previous chapter. The original dataset contains 20,000 rows. Because KNN compares every testing row with every training row, classifying all 20,000 rows takes a long time. We will randomly select 3,000 rows so that the classifier runs more quickly while still using a large and varied sample of the data.

    import numpy as np
    import pandas as pd
    import matplotlib.pyplot as plt
    
    # read input data from file
    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/cdc.csv"
    health = pd.read_csv(url)
    # random sample of 3000 rows
    health = health.sample(n=3000, random_state=1)
    Output?

    Prepare Data

    We take the same steps to prepare data as in Chapter 15: convert gender values to numbers, change height and weight to standard units.

    # select only the features we want to use for classification and form a new dataset
    data = health[['height', 'weight', 'gender']].copy()
    # convert gender to numeric values
    data['gender'] = data['gender'].map({'m': 0, 'f': 1})
    # convert height and weight to standard units
    data['height'] = (data['height'] - data['height'].mean()) / data['height'].std()
    data['weight'] = (data['weight'] - data['weight'].mean()) / data['weight'].std()
        
    Output?

    Create Training and Testing Data

    We use random sampling to create the two datasets. First we randomly shuffle all the rows of the dataset, in the same way that we randomly shuffle data for hypothesis testing in Chapter 10. Then we assign some of the rows to the training dataset and the remaining rows to the testing dataset. For this chapter we will use half of the rows for training and half for testing.

    # shuffle all the rows
    shuffled_data = data.sample(frac=1)
    # take first half for training set
    training_set = shuffled_data.iloc[:len(shuffled_data)//2].copy()
    # rename the gender column to _class, a requirement for our classifier
    training_set = training_set.rename(columns={'gender':'_class'})
    # take second half for testing set
    testing_set = shuffled_data.iloc[len(shuffled_data)//2:].copy()
    # rename the gender column to _class, a requirement for our classifier
    testing_set = testing_set.rename(columns={'gender':'_class'})
    Output?

    We confirm the training dataset by plotting the height vs weight.

    plt.figure(figsize=(4,4))
    
    groups = training_set.groupby('_class')
    for name, group in groups:
        plt.scatter(group['height'], group['weight'], alpha = 0.4, label=name)
    plt.xlabel('Height')
    plt.ylabel('Weight')
    plt.legend()
    plt.title('Training Data')
    plt.grid()
    plt.show()
    Output?

    The scatterplot shows only the training data. When KNN makes a prediction, it searches only the training dataset for the nearest neighbors. The testing data are not used to make predictions because we want them to represent new, unseen data. After the predictions have been made, we compare them with the actual classes in the testing dataset to measure the classifier's accuracy.

    Implement a Classifier

    We now have the necessary datasets to create and test a simple classifier. We use the training and testing datasets above to train and test the classifier.

    To write the code for the classification algorithm, we follow the steps that we've used in Chapter 15:

    1. Find the distance between the new data point and all the existing data points in the dataset.
    2. Use these distances to find the k points that are closest (the shortest distance) to the new data point.
    3. Find the class that occurs the most in the k nearest data points. This will be our predicted class for the new data point.

    We bring in the distance_from_point function and find_distances function from Chapter 15.

    # calculate the distance between a new data point and one data point in the dataset
    # - new_point is a new data point with x, y coordinates on the scatterplot
    # - row is a row or one data point in the dataset, with x, y coordinates on the scatterplot
    # return the distance between the two data points
    def distance_from_point(new_point, row):
      # convert each data point into an array
      new_point = np.array(new_point)   # x1, y1
      row = np.array(row)               # x2, y2
      # subtract the 2 arrays
      difference = new_point - row      # x1 - x2, y1 - y2
      # square each difference
      squared_difference = difference ** 2      # (x1 - x2)^2, (y1 - y2)^2
      # add the squared differences and take the square root
      D = np.sqrt(np.sum(squared_difference))   
      return D
    
    # find the distances between a new data point and all the data points in the dataset
    # - new_point is a new data point with x and y values 
    # - dataset has rows of data features, with x and y values
    # return the list of distances 
    def find_distances(new_point, dataset):
        # create a list to store the distances
        distances = []
        for row in dataset:
            # calculate the distance between the new point and the current point
            a_distance = distance_from_point(new_point, row)
            distances.append(a_distance)
        return distances
    Output?

    We also add one new function find_k_nearest_neighbors and bring in the majority function from Chapter 15.

    # find the k nearest neighbors of a new data point in the dataset
    # - new_point is a new data point with x and y values (height and weight)
    # - dataset is the DataFrame containing the training data
    # - k is the number of nearest neighbors to find
    # return a DataFrame containing the k nearest neighbors and their distances
    def find_k_nearest_neighbors(new_point, dataset, k):
        # extract the features of each row of data
        features = dataset.drop(columns='_class')
        # find all distances
        distances = find_distances(new_point, features.values)
        # create a new DataFrame and add the distances to it
        results = dataset.copy()
        results["Distance"] = distances
    
        # sort the dataset by 'Distance', then take the top k rows
        k_nearest_neighbors = results.sort_values(by='Distance').head(k)
        return k_nearest_neighbors
    
    # find the majority class of the k nearest neighbors
    # - k_nearest is a DataFrame containing the k nearest neighbors and their distances
    # return the predicted class 
    def majority(k_nearest):
        # find the count of 1's
        ones = np.sum(k_nearest._class == 1)
        # find the count of 0's
        zeros = np.sum(k_nearest._class == 0)
        # return the larger of the two, or the majority class
        if ones > zeros:
            return 1
        else:
            return 0
    Output?

    We have now automated each step of the KNN algorithm. We've written the functions to calculate distances, find the k nearest neighbors, and determine the majority class. The final function, classify, is the classifier. When we call the classifier, it calls the other functions in turn to carry out each step of the KNN algorithm and return the predicted class.

    # k-Nearest Neighbors Classifier. A function that accepts as input:
    # - the new data point,
    # - the DataFrame of existing data,
    # - the number of nearest neighbors
    def classify(new_point, dataset, k):
        k_nearest = find_k_nearest_neighbors(new_point, dataset, k)
        return majority(k_nearest)
    Output?

    Test the Classifier

    We start by testing the classifier with a new data point.

    new_point = [2, 2]  # x,y coordinate of the new data point in the scatterplot
    
    # run the classifier on the new data point and print the predicted class
    print("Predicted class:", classify(new_point, training_set, 5))
    Output?
    Predicted class: 0
    

    Visually checking the location (2,2) in the scatterplot, we see that a prediction of class 0 is expected.

    Next we use the testing data with our classifier.
    First we save the _class (gender) data from testing_set, then we remove the _class data from the testing_set. The remaining columns contain only the feature values that the classifier will use to make its predictions.

    actual = testing_set["_class"]
    test_set = testing_set.drop(columns=['_class'])
    Output?

    Then we test the classifier with every row of the test_set. Each prediction for a data point in test_set is saved in a list of predicted values.

    # create a list to store the predicted classes
    predicted = []
    # repeat for each row in the test set
    for index, row in test_set.iterrows():   # iterrows() returns the index and row of test_set
        # run the classifier on the row and append the predicted class to the list
        predicted.append(classify(row, training_set, 5))
    print(len(predicted), "predictions made")
    Output?
    1500 predictions made
    

    Find Prediction Accuracy

    We store the actual data and the predicted data in a DataFrame and print the first five rows.

    results = pd.DataFrame({'Actual': actual, 'Predicted': predicted})
    results.head(5)
    Output?
      Actual Predicted
    18280 1 1
    10229 1 1
    17305 0 0
    9053 0 0
    4971 1 1

    Since the data is randomly shuffled before being divided into the training and testing sets, the results above may differ slightly with each run of this example. The predicted classes for individual data points may be different, but the overall accuracy should be similar.

    We calculate the percent accuracy by finding the ratio of correct predictions over the total predictions.

    # add all the rows where Actual is the same as Predicted
    correct = (results.Actual == results.Predicted).sum()
    # find the ratio of correct predictions over total predictions as a percentage
    print("Accuracy:", round(correct / len(results) * 100), "%")
    Output?
    Accuracy: 84 %
    

    The classifier correctly predicted the gender for a high percentage of the people in the testing dataset.

    Looking at the female and male ratios in the starting dataset of 3000 people, we find the count and the percentage of f and m values:

    display(health["gender"].value_counts())
    display(health["gender"].value_counts(normalize=True) * 100)
    Output?
    gender
    f    1570
    m    1430
    Name: count, dtype: int64
    gender
    f    52.333333
    m    47.666667
    Name: proportion, dtype: float64

    We see that if a classfier always predicted female, it would be correct about 52% of the time. Our classifier achieves a much higher accuracy, showing that it has learned patterns in the data rather than simply predicting the most common class.

    The classifier correctly predicts the gender for a high percentage of the people in the testing dataset, but it is not perfect because height and weight alone do not completely distinguish males from females. Many people have similar heights and weights, so some data points are difficult to classify correctly.

    Test the Classifier with a Different Dataset

    In the previous example we use the classifier to predict gender by using height and weight. One advantage of KNN is that the algorithm does not depend on a particular dataset. As long as we provide a training dataset with feature values and a class label, we can use the same classifier to make predictions on completely different data.

    To demonstrate this, we will use the cars dataset that we've used in previous chapters. Instead of predicting gender, we will predict whether a car is fuel efficient. The features will be horsepower, weight, and cylinders, and the classifier will indicate whether the car is fuel efficient or not. Note that we do not need to change the classifier itself, we simply provide it with a different dataset.

    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/cars.csv"
    cars = pd.read_csv(url)
    display(cars.head())
    print("Number of rows:", cars.shape[0])
    Output?
      mpg cylinders displacement horsepower weight acceleration model_year origin car_name
    0 18.0 8 307.0 130.0 3504.0 12.0 70 1 chevrolet chevelle malibu
    1 15.0 8 350.0 165.0 3693.0 11.5 70 1 buick skylark 320
    2 18.0 8 318.0 150.0 3436.0 11.0 70 1 plymouth satellite
    3 16.0 8 304.0 150.0 3433.0 12.0 70 1 amc rebel sst
    4 17.0 8 302.0 140.0 3449.0 10.5 70 1 ford torino
    Number of rows: 392
    

    We create a DataFrame with only the features and target we want.

    car_data = cars[["horsepower", "weight", "cylinders", "mpg"]].copy()
    Output?

    We want to predict whether a car is fuel efficient. Since fuel efficiency is measured by mpg, we will use mpg to define our classes. However, a classifier can only predict categorical data, and mpg is continuous quantitative data, so we convert mpg into two categories:

    • 0: mpg < 25, not fuel efficient
    • 1: mpg >= 25, fuel efficient
    # Convert mpg to binary classification: 0 if < 25, 1 if >= 25
    car_data['mpg'] = (car_data['mpg'] >= 25).astype(int)
    car_data.head()
    Output?
      horsepower weight cylinders mpg
    0 130.0 3504.0 8 0
    1 165.0 3693.0 8 0
    2 150.0 3436.0 8 0
    3 150.0 3433.0 8 0
    4 140.0 3449.0 8 0

    In the previous example we used two features, height and weight, to predict gender. In this example, we will use three features: horsepower, weight, and cylinders to predict whether a car is fuel efficient. Even though the dataset and the number of features are different, the KNN algorithm does not change. The same classifier works without modifications on completely different datasets, even if they have a different number of features.

    Although we first developed the algorithm using two features, the code is not limited to two dimensions. The distance_from_point function works with arrays of any length, so the same classifier can be used with datasets that have two features, three features, or many more.

    We continue preparing the data:

    • change the mpg label to _class as required by our classifier
    • convert all the features to standard units
    # rename the mpg column to _class
    car_data = car_data.rename(columns={'mpg': '_class'})
    Output?
    # convert all the features to standard units
    car_data['horsepower'] = (car_data['horsepower'] - car_data['horsepower'].mean()) / car_data['horsepower'].std()
    car_data['weight'] = (car_data['weight'] - car_data['weight'].mean()) / car_data['weight'].std()
    car_data['cylinders'] = (car_data['cylinders'] - car_data['cylinders'].mean()) / car_data['cylinders'].std()  
         
    Output?
    car_data.head()
    Output?
      horsepower weight cylinders _class
    0 0.663285 0.619748 1.482053 0
    1 1.572585 0.842258 1.482053 0
    2 1.182885 0.539692 1.482053 0
    3 1.182885 0.536160 1.482053 0
    4 0.923085 0.554997 1.482053 0

    We divide the data into the training and testing sets.

    shuffled_data = car_data.sample(frac=1)
    training_set = shuffled_data.iloc[:len(shuffled_data)//2].copy()
    testing_set = shuffled_data.iloc[len(shuffled_data)//2:].copy()
    Output?

    We save the _class column from the testing_set. Then we remove the _class column from the testing data.

    # save the actual target or `_class`, then drop `_class` from the testing_set
    actual = testing_set._class
    test_set = testing_set.drop(columns='_class')
    test_set.head()
    Output?
      horsepower weight cylinders
    277 -0.947474 -0.991973 -0.862911
    377 -0.973454 -1.192113 -0.862911
    139 -0.557775 -0.893080 -0.862911
    317 -0.765614 -0.512812 -0.862911
    330 0.715245 -0.079567 0.309571

    Now that the data has been prepared, we use the classifier to predict the fuel efficiency of the cars.

    predicted = []
    for index, row in test_set.iterrows():   
        predicted.append(classify(row, training_set, 5))
    
    # combine the actual and predicted into a DataFrame, 
    # and find the number of correct predictions
    correct = (actual == predicted).sum()
    print("Accuracy:", round(correct / len(actual) * 100))
    Output?
    Accuracy: 84
    

    We see that the classifier has high accuracy on a completely different dataset. This demonstrates that the KNN algorithm is not tied to a particular dataset and can be reused to solve different classification problems.

    Summary

    In this chapter, we built a KNN classifier step by step by combining several functions into a single classifier that predicts the class of new data. We used training and testing datasets to measure its accuracy and saw that the same classifier can be applied to completely different datasets without changing the algorithm.

    KNN is one of many machine learning algorithms. Machine learning is a foundation of modern AI because it enables computers to learn patterns from data and make predictions.


      This page titled 16: Classifier was last modified on Fri, 25 Sep 2026 01:22:37 GMT and is shared under a CC BY 4.0 license and was authored, remixed, and/or curated by Clare Nguyen.

      • Was this article helpful?