Skip to main content
Workforce LibreTexts

5: Plots

  • Page ID
    47979
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

    In the previous chapter we saw that an array and a DataFrame are convenient ways to store a dataset in an organized way that makes it easy to select columns or rows of data and analyze them. However, it's not easy to make observations about the data by looking at multiple rows and columns of many numbers. In this chapter we learn to show the data visually so that it's easy to spot the patterns and trends in the data.

    This chapter covers four different types of plots that will display the data visually: scatterplots, line plots, bar charts, and histograms. We will use the same dataset about cars that we explored in the last chapters.

    Gather Data

    First we read the same cars dataset into a DataFrame and display the first five rows.

    # import pandas to use the DataFrame
    import pandas as pd
    
    # 1. Store the URL of the movie data webpage
    url = "https://raw.githubusercontent.com/DeAnzaDataScience/CIS11/refs/heads/main/datasets_notes/cars.csv"
    
    # 2. Read the CSV file into a DataFrame
    cars = pd.read_csv("cars.csv")
    
    # 3. Show the number of rows and columns
    rows, columns = cars.shape
    print("Number of rows:", rows)
    print("Number of columns:", columns)
    
    # 4. Print the first 10 lines of the DataFrame
    print("First 10 rows:")
    cars.head(10)
    Output?
    Number of rows: 392
    Number of columns: 9
    First 10 rows:
    
      mpg cylinders displacement horsepower weight acceleration model_year origin car_name
    0 18.0 8 307.0 130.0 3504.0 12.0 70 1 chevrolet chevelle malibu
    1 15.0 8 350.0 165.0 3693.0 11.5 70 1 buick skylark 320
    2 18.0 8 318.0 150.0 3436.0 11.0 70 1 plymouth satellite
    3 16.0 8 304.0 150.0 3433.0 12.0 70 1 amc rebel sst
    4 17.0 8 302.0 140.0 3449.0 10.5 70 1 ford torino
    5 15.0 8 429.0 198.0 4341.0 10.0 70 1 ford galaxie 500
    6 14.0 8 454.0 220.0 4354.0 9.0 70 1 chevrolet impala
    7 14.0 8 440.0 215.0 4312.0 8.5 70 1 plymouth fury iii
    8 14.0 8 455.0 225.0 4425.0 10.0 70 1 pontiac catalina
    9 15.0 8 390.0 190.0 3850.0 8.5 70 1 amc ambassador dpl
     
    cars.describe()
    Output?
      mpg cylinders displacement horsepower weight acceleration model_year origin
    count 392.000000 392.000000 392.000000 392.000000 392.000000 392.000000 392.000000 392.000000
    mean 23.445918 5.471939 194.411990 104.469388 2977.584184 15.541327 75.979592 1.576531
    std 7.805007 1.705783 104.644004 38.491160 849.402560 2.758864 3.683737 0.805518
    min 9.000000 3.000000 68.000000 46.000000 1613.000000 8.000000 70.000000 1.000000
    25% 17.000000 4.000000 105.000000 75.000000 2225.250000 13.775000 73.000000 1.000000
    50% 22.750000 4.000000 151.000000 93.500000 2803.500000 15.500000 76.000000 1.000000
    75% 29.000000 8.000000 275.750000 126.000000 3614.750000 17.025000 79.000000 2.000000
    max 46.600000 8.000000 455.000000 230.000000 5140.000000 24.800000 82.000000 3.000000
     

    We now use different plots to visualize the data.

    To plot data, we use another Python module called matplotlib.pyplot. This module has all the plotting functions that will plot data in specific ways.

    import matplotlib.pyplot as plt  # shorten the module name to plt
    Output?

    Scatterplots

    Use Case for a Scatterplot

    A scatterplot shows the relationship between 2 features or 2 columns of the data. If one feature increases in value, does the other feature tend to increase value as well? Or does it decrease instead? Or is there simply no pattern between the two features?

    Scatterplot Application

    For the car dataset we want to see whether a car's weight is related to the mpg (or fuel efficiency). We suspect that the heavier the car is, the lower its fuel efficiency, since it takes more fuel to move a larger mass. But is it true?

    To answer the question, we use a scatterplot to see whether there's a pattern or association between these two features. The mpg feature will be along the x-axis or horizontal axis, and the weight feature will be along the y-axis or the vertical axis. The choice of which feature goes on the x-axis vs the y-axis is arbitrary. We can also easily swap the mpg and weight on the x-axis and y-axis.

    [View the color choices for plotting data in matplotlib]

    # set the size of the plot
    plt.figure(figsize=(4,4))
    
    # create a scatterplot of mpg vs weight
    plt.scatter(cars['mpg'], cars['weight'], color='green') 
    # x-axis data first, y-axis data second
    
    # label (describe) the x-axis or horizontal axis
    plt.xlabel('Miles per Gallon (mpg)')
    
    # label (describe) the y-axis or vertical axis
    plt.ylabel('Weight')
    
    # title the plot
    plt.title('MPG vs Weight')
    
    # put a grid on the plot and show the plot
    plt.grid()
    plt.show()
    Output?

     

    In the scatterplot above each green point is for one car in the dataset. See if you can locate the car with the highest mpg in the dataset, close to 50, which weighs a little above 2000 lbs.

    The scatterplot shows a general pattern: the heavier the car is, the lower its fuel efficiency. This is just as we suspected, but the scatterplot makes it easy to see the relationship in the data.

    When one feature generally increases or decreases with another feature, we say that there's an association between the two features. In this case, there is a negative association because as one feature increases, the other feature decreases. If both features increase together or decrease together, then it is a positive association.

    We also note that the negative association in the plot above is a general trend and doesn't apply to all cars. For example, looking at the green dots above, we can find one car that weighs around 3000 lbs and is around 38 mpg, which is much higher than the 18 mpg of another car that weighs only around 2000 lbs.

    Line Plots

    Use Case for a Line Plot

    A line plot is used to show the trend or the change in data over a range of measurements, typically a range of time. If you've ever seen a plot that shows how the price of a company stock fluctuates during a trading day, then you've seen a line plot.

    Line Plot Application

    For the car dataset we want to see whether the fuel efficiency of cars has increased over time. We hope that later model cars have better mileage than earlier model cars, and we'll use a line plot to see whether it's true.

    For the line plot:

    • The years represent our the time range that the data span. It is the independent variable of the line plot and is on the x-axis.
    • The mpg changes over time. It is the dependent variable since it depends on the years so it is used on the y-axis.

    Before we can plot the mpg, we need to find the average mpg for each year. This is because in each year there can be cars with high fuel efficiency as well as low fuel efficient cars, and we want the average fuel efficiency for that year.

    Fortunately the DataFrame has a method that makes it easy to calculate the average mpg per year: the groupby method:

    yearly_mpg = cars.groupby('model_year')['mpg'].mean().reset_index()
    Output?

    In the calculation above, we ask the DataFrame to group the cars by year or model_year. For each year group, we choose the mpg column to calculate the mean or the average, resulting in one average mpg per year. Since we end up with a small number of average mpg compared to the individual mpg of all the cars, we use the reset_index function to reset the index to the count of years, starting at 0.

    display(yearly_mpg)
    Output?
      model_year mpg
    0 70 17.689655
    1 71 21.111111
    2 72 18.714286
    3 73 17.100000
    4 74 22.769231
    5 75 20.266667
    6 76 21.573529
    7 77 23.375000
    8 78 24.061111
    9 79 25.093103
    10 80 33.803704
    11 81 30.185714
    12 82 32.000000
     
    # line plot of the average mpg over the years
    
    plt.figure(figsize=(4,4))
    plt.plot(yearly_mpg['model_year'], yearly_mpg['mpg'], color='purple')
    # x-axis data are the years, y-axis data are the average mpg
    
    plt.title('Average Fuel Efficiency (MPG) Over Years')
    plt.xlabel('Year')
    plt.ylabel('Average Miles Per Gallon')
    plt.grid()
    plt.show()
    Output?

     

    The line plot shows that there was a big increase in the mpg between 1976 and 1980, and the overall trend is that fuel efficiency did indeed increase over time.

    Bar Charts

    Quantitative Data vs Categorical Data

    Sometimes we work with data that are in a continuous range, such as the mpg of the cars. Data that are in a continuous range are quantitative data.

    At other times we may work with data that are in categories instead of in a range. An example of categorical data would be whether a person is a full-time student, part-time student, or not a student. In the car dataset, one categorical feature is the number of cylinders.

    print("Example of quantitative data: mpg")
    print("Numbers of unique mpg values:", cars['mpg'].unique())
    print()
    print("Example of categorical data: cylinders")
    print("Numbers of unique cylinders values:", cars['cylinders'].unique())
    Output?
    Example of quantitative data: mpg
    Numbers of unique mpg values: [18.  15.  16.  17.  14.  24.  22.  21.  27.  26.  25.  10.  11.   9.
     28.  19.  12.  13.  23.  30.  31.  35.  20.  29.  32.  33.  17.5 15.5
     14.5 22.5 24.5 18.5 29.5 26.5 16.5 31.5 36.  25.5 33.5 20.5 30.5 21.5
     43.1 36.1 32.8 39.4 19.9 19.4 20.2 19.2 25.1 20.6 20.8 18.6 18.1 17.7
     27.5 27.2 30.9 21.1 23.2 23.8 23.9 20.3 21.6 16.2 19.8 22.3 17.6 18.2
     16.9 31.9 34.1 35.7 27.4 25.4 34.2 34.5 31.8 37.3 28.4 28.8 26.8 41.5
     38.1 32.1 37.2 26.4 24.3 19.1 34.3 29.8 31.3 37.  32.2 46.6 27.9 40.8
     44.3 43.4 36.4 44.6 33.8 32.7 23.7 32.4 26.6 25.8 23.5 39.1 39.  35.1
     32.3 37.7 34.7 34.4 29.9 33.7 32.9 31.6 28.1 30.7 24.2 22.4 34.  38.
     44. ]
    
    Example of categorical data: cylinders
    Numbers of unique cylinders values: [8 4 6 3 5]
    
     

    The quantitative mpg data are within a range from 9 to 44.
    The categorical cylinders data are in five distinct values or categories.

    Use Case for a Bar Chart

    When we work with categorical data, a bar chart is used to show the count of data for each category. Each category is represented by a bar, and the length of the bar corresponds to the count of data in that category.

    Bar Chart Application

    For the car dataset we want to see how many cars there are for each number of cylinder category.

    First we need to group the data by number of cylinders and count the number of cars in each group.

    car_count = cars.groupby('cylinders').cylinders.count()
    car_count
    Output?
    cylinders
    3      4
    4    199
    5      3
    6     83
    8    103
    Name: cylinders, dtype: int64
     

    Now we use a bar chart to visually compare the count of each group.

    plt.figure(figsize=(4,4))
    plt.bar(car_count.index, car_count.values)
    plt.title('Number of Cars by Cylinder Count')
    plt.xlabel('Cylinder Count')
    plt.ylabel('Number of Cars')
    plt.grid()
    plt.show()
    Output?

     

    From the output above, we can easily see that the dataset has the highest number of 4-cylinder cars, followed by 8-cylinder cars. There are very few 3-cylinder and 5-cylinder cars.

    It is also possible to plot a horizontal bar chart instead of a vertical bar chart as above. We use most of the same code for plotting except we use barh() for horizontal bar chart, and we swap the x-axis and y-axis labels.

    plt.figure(figsize=(5,4))
    
    # barh for horizontal bar chart
    plt.barh(car_count.index, car_count.values)
    plt.title('Number of Cars by Cylinder Count')
    plt.xlabel('Number of Cars')
    plt.ylabel('Cylinder Count')
    plt.grid()
    plt.show()
    Output?

     

    Generally, when the x-ticks are long strings, the bar chart is easier to read if it's a horizontal bar chart. In the first vertical bar chart, the cylinder counts along the x-axis are small numbers so they don't overlap each other in the plot, so both a vertical and horizontal bar chart can be used.

    Histograms

    In the previous sections we learned how to visualize the count of categorical data, an important fact about the data. From the bar chart output we can immediately see which categories have the most data points or the fewest data points, and this can help us draw conclusions on the data.

    For quantitative data, which are data points that belong on a continuous range of values instead of in separate categories, it can also be helpful to know the count of data. However, finding the count of quantitative data is slightly different than finding the count of categorical data:

    • For categorical data, a data point can be in one of several categories (such as full-time student, part-time student, or not a student), so it's easy to find the count of the three categories in the dataset. From the output of the bar chart, we can easily see what type of students is most common in the dataset.

    • For quantitative data, most data points will have different values. For example, let's say the cost of a new car in the US is in a range from 15,000 to 125,000 dollars. If we have a dataset of 1000 new cars, most of the "out the door" price (including taxes, fees, licensing, add-ons, etc.) for these cars will be slightly different. If we used a bar chart to display the count of each "category" or each price, we will have many categories, each with a count of one or at most two or three, and this doesn't give us much information about which price ranges are the most common in the dataset.

    Binning Data

    To solve the problem of how to find the count of quantitative data, we use a statistical method called binning the data. When we bin the data, we take the entire range of data and divide it into equal intervals, each called a bin. Each bin is a small part of the entire range of data. Then we put each data point into a bin whose range includes the data point.

    For the new car price example above, we can divide the range of $15,000 to $125,000 into bins of $10,000 each:

    • first bin: $15,000 to $24,999
    • second bin: $25,000 to $34,999
    • third bin: $35,000 to $44,999
    • ...
    • last bin: $115,000 to $125,000 (note the different end for the range here)

    If a car has a price of $27,895, then it goes in the second bin.

    After all the data points have been put into bins, we can plot the bins as if they're "categories", except that the bins form a continuous range. The resulting plot is called a histogram.

    A histogram is similar to a bar chart, except that the bars represent ranges of numerical values rather than separate categories.

    Use Case for a Histogram

    A histogram shows the shape of the data sequence. With a histogram we can see the minimum and maximum values, what range of values is the most common in the dataset, whether the data range is skewed to the right or left, where the average is, whether there is a gap, and whether there are outliers, which are extreme values compared to the rest of the data.

    Histogram Application

    Using the same car dataset, we display the first 10 rows of the cars DataFrame:

    cars.head(10)
    Output?
      mpg cylinders displacement horsepower weight acceleration model_year origin car_name
    0 18.0 8 307.0 130.0 3504.0 12.0 70 1 chevrolet chevelle malibu
    1 15.0 8 350.0 165.0 3693.0 11.5 70 1 buick skylark 320
    2 18.0 8 318.0 150.0 3436.0 11.0 70 1 plymouth satellite
    3 16.0 8 304.0 150.0 3433.0 12.0 70 1 amc rebel sst
    4 17.0 8 302.0 140.0 3449.0 10.5 70 1 ford torino
    5 15.0 8 429.0 198.0 4341.0 10.0 70 1 ford galaxie 500
    6 14.0 8 454.0 220.0 4354.0 9.0 70 1 chevrolet impala
    7 14.0 8 440.0 215.0 4312.0 8.5 70 1 plymouth fury iii
    8 14.0 8 455.0 225.0 4425.0 10.0 70 1 pontiac catalina
    9 15.0 8 390.0 190.0 3850.0 8.5 70 1 amc ambassador dpl
     

    We now look at the acceleration column, which is the number of seconds it takes to accelerate from 0 to 60 mph. We plot the acceleration values in a histogram.

    plt.figure(figsize=(4,4))
    
    # create a histogram, the default number of bins is 10:
    plt.hist(cars['acceleration'], color='brown', edgecolor='black')
    
    # add title, labels, grid, and show the plot
    plt.title('Distribution of Acceleration')
    plt.xlabel('Acceleration')
    plt.ylabel('Count of cars')
    plt.grid()
    plt.show()
    Output?

     

    The histogram above has 10 bars or 10 bins that cover the continuous range of acceleration data along the x-axis.

    We see that the fastest acceleration is just under 5 seconds, and the slowest acceleration is almost 25 seconds.

    Since the range of each bin is 1/10 of the entire data range, the width of each bar is the same. Therefore the height of the bar corresponds to the count of data points in a particular bin, and we see that the tallest bar or the largest count is for an acceleration of around 15-16 seconds. This tallest bin is also around the middle of the data range, so we expect the mean or average acceleration of all cars is also between 15-16 seconds.

    The histogram shows us the shape of the acceleration data. Looking at the three tallest bars in the middle, we can say that the majority of the cars in the dataset has an acceleration of around 13 to 18 seconds, and the faster and slower accelerations slope down symmetrically on the left and right of center. The shape of this data distribution is shaped like a bell, or is a bell curve: it is symmetric, with no skew on the left or right.

    We also notice that there's no gap or no empty bin, and no outliers, which are data values that are far on the left or right on the x-axis from the rest of the data.

    To check the count of data points in each bin, we call the same hist() function to get the counts:

    plt.figure(figsize=(4,4))
    
    (counts, edges, patches) = plt.hist(cars['acceleration'], color='brown', edgecolor='black')
    # the hist() function returns:
    # the count of each bin,
    # the value at the max edge of each bin,
    # and the patches, the standard name from hist() for the plot figure
    
    print("Count of each bin:", counts)
    print("Max value of bin edges:", edges)
    plt.xlabel('Acceleration')
    plt.ylabel('Count of cars')
    plt.grid()
    plt.show()
    Output?
    Count of each bin: [ 6. 15. 50. 85. 91. 78. 44. 12.  7.  4.]
    Max value of bin edges: [ 8.    9.68 11.36 13.04 14.72 16.4  18.08 19.76 21.44 23.12 24.8 ]
    

     

    The printed counts from the output above match our expectations from viewing the plot. The middle bin has the highest count of 92 cars, and the counts of bins on both sides slope down fairly symmetrically .

    We can change the bin range to smaller intervals by requesting a larger number of bins. This lets us see the histogram with more granularity:

    plt.figure(figsize=(4,4))
    
    # asking for twice the number of bins as the default:
    (counts, edges, patches) = plt.hist(cars['acceleration'], color='brown', edgecolor='black', bins=20)
    
    print("Count of each bin:", counts)
    plt.xlabel('Acceleration')
    plt.ylabel('Count of cars')
    plt.grid()
    plt.show()
    Output?
    Count of each bin: [ 3.  3.  5. 10. 21. 29. 29. 56. 57. 34. 50. 28. 19. 25.  6.  6.  7.  0.
      2.  2.]
    

     

    There are now 20 bars for the 20 bins. The shape of the 20 bars are not exactly the same as with 10 bins. The bar heights are generally lower since the bar widths are half as wide. This is because there are twice as many bins so each bin has fewer data points. We also see there is a gap, the third bin from the right has no data values. But overall the distribution or shape of the data are similar in both histograms.

    Density Scaling

    Last but not least, it's important to note that in each of the two histograms above, the height of the bars corresponds to the count of data in the bins only because the bars all have the same width. But there are two common cases when the count of data in the bins should not directly determine the height of the bars.

    Case 1

    Sometimes the bars in a histogram don't have equal widths due to the nature of the data, and this can cause misleading histogram plots. As an example:

    • Suppose in a histogram, the bar for bin A happens to be twice as wide as the bar for bin B, and the two bins have the same count of data.
    • If we use the count to determine the height of the bars, then bin A and bin B will be the same height.
    • But since the bar for bin A is twice as wide than the bar for bin B, the bar for bin A will appear twice as large as the bar for bin B, misleading the viewer into concluding that bin A has twice the count of data as bin B.

    Case 2

    If the count of data in the bins determines the height of the bars, then it can be difficult to compare the distribution of two datasets of different sizes. As an example:

    • Suppose dataset A has twice as many data points as dataset B, and we need to compare the distribution of the two datasets.
    • Since the total count of data in dataset A is larger, it means that the count in each bin will generally be larger, and the bars will be generally be higher than the bars in dataset B.
    • When comparing the histograms, it would be difficult to determine whether the difference in bar heights between the 2 histograms is due to one dataset simply having more data points to go in the bins, or if it's due to the different chracteristics of the datasets.

    The solution for the two cases above is to:

    1. Scale the count in each bin to be a ratio of the count / total data points.
      This turns the count of each bin into a proportion of the dataset that falls into each bin.
      It normalizes the data so we can compare the distribution of two datasets that differ greatly in size (10,000 data points vs 100 data points).
    2. The area of the bar represents this ratio.
      This means the ratio = height of the bar * width of the bar
      Or: height of the bar = ratio / width of the bar.
      Since the area of each bar is a ratio of count / total data points, when we add all the areas of all the bars, the total area is 1. Again, this normalizes the distribution so that different distributions can be compared together.

    When the height of the bars is determined this way, it represents the density or concentration of data at a bin, and the histogram is called a density histogram.

    Fortunately we don't have to do all the calculations above to plot the bar height as density instead of as counts.
    The hist() function has a variable named density that determines whether the height of the bar is determined by the count in the bins or the density of the bins. The default is to use the count in the bins, but when we set density to True, then the height of the bar is determined by the density.

    Below we compare the frequency histogram (y-axis as count) and the density histogram (y-axis as density).

    plt.figure(figsize=(10, 4))
    
    # Left plot: frequency histogram
    plt.subplot(1, 2, 1)
    plt.hist(cars['acceleration'], color='brown', edgecolor='black')
    plt.xlabel('Acceleration')
    plt.ylabel('Count of cars')
    plt.title('Frequency Histogram')
    plt.grid()
    
    # Right plot: density histogram
    plt.subplot(1, 2, 2)
    plt.hist(cars['acceleration'], color='brown', edgecolor='black', density=True)
    plt.xlabel('Acceleration')
    plt.ylabel('Density')
    plt.title('Density Histogram')
    plt.grid()
    
    plt.tight_layout()
    plt.show()
    Output?

     

    Note that the two distribution shapes are identical. In the frequency histogram the y-axis is the actual count of cars in a bin, while in the density histogram, the y-axis is the density or how crowded with data a bin is.

    Overlaid Plots

    Sometimes in data analysis it can be helpful to compare the plots of two sequences of data to visually see how the two datasets relate to each other over the same time period or under the same conditions. To facilitate this visual comparison, each of the four plotting functions we've discussed can plot more than one sequence of data in the same plot. This is to overlay the plots in the same figure.

    To plot two or more sequences of data in the same plot, the sequences should:

    • be related and have the same number of data points
    • have the same scale

    An example is to plot the homework scores between 2 students in a class during a semester. We work with a small dataset below:

    scores = pd.DataFrame(data=[[20,18,18,19,20,19,18,18,19,19],
                                [18,17,16,19,18,20,19,19,19,20]],
                          columns=["HW1","HW2","HW3","HW4","HW5",
                                   "HW6","HW7","HW8","HW9","HW10"],
                          index=["Student A","Student B"])
    scores
    Output?
      HW1 HW2 HW3 HW4 HW5 HW6 HW7 HW8 HW9 HW10
    Student A 20 18 18 19 20 19 18 18 19 19
    Student B 18 17 16 19 18 20 19 19 19 20
     

    We note that the sequences of scores for each student meet the conditions to be plotted together in one plot:

    • The sequences of scores are related: each sequence has the scores for one student, and since the students are in the same class, the scores are for the same work.
    • The sequences of scores have the same number of data points, there are 8 scores for each sequence.
    • The sequences have the same scale, they are all out of 20 points.

    To compare how well the 2 students did with their homework during the semester, we can use a line plot to show the trend of scores over time

    plt.figure(figsize=(5,4))
    
    # plot each student score, and add a label to identify the student
    plt.plot(scores.columns, scores.loc['Student A'], label='Student A')
    plt.plot(scores.columns, scores.loc['Student B'], label='Student B')
    
    plt.ylabel('Score')
    plt.legend()  # the legend indifies the label for each plot line
    
    plt.grid()
    plt.show()
    Output?

     

    The plot makes it easier to see that Student A is more consistent with the homework scores, while Student B scored lower at the beginning of the semester, but improved steadily and surpassed Student A at the end of the semester.

    When overlaying two plots, matplotlib.pyplot automatically chooses different colors for the two plots. But without a legend to explain what the colors represent, it's not easy to understand the plot. In the overlaid plots above, the legend shows that Student A's data are in blue and Student B's data are in orange. This makes it easy to compare the trend in scores of the students.

    In the example above we see how two line plots can be overlaid in the same plot figure to make it easier to compare them. We can also use the same steps to overlay two scatterplots (by calling plt.scatter twice) or two bar charts (by calling plt.bar twice) or two histograms (by calling plt.hist twice).

    Summary

    In this chapter we learn that data visualization or plotting helps us observe more easily trends and patterns in a dataset. By selecting the appropriate plot to display data, we can see whether two features of a dataset have an association with each other, or see the trend of a feature over time, or see the category or range that contains the most number of the data points in the dataset.


    This page titled 5: Plots was last modified on Thu, 24 Sep 2026 23:53:14 GMT and is shared under a CC BY 4.0 license and was authored, remixed, and/or curated by Clare Nguyen.

    • Was this article helpful?