Ctrl+k

Describe data distributions

A data distribution describes how numerical values are spread and organized, including its center, variability, overall shape, clusters, gaps, and possible outliers, as shown in dot plots, histograms, or box plots. Its features can be summarized with measures such as mean or median and range or interquartile range, while recognizing that a single measure does not fully describe the data; formal probability distributions and advanced statistical modeling are not included.

Detailed Explanation: Describe data distributions

A data distribution tells how the values in a data set are spread out. To describe it, look for:

  • Center: a typical value, such as the mean or median
  • Variability: how much the values differ, such as the range or interquartile range
  • Shape: whether the data are balanced, skewed, or have a long tail
  • Clusters: groups of values close together
  • Gaps: intervals with no values
  • Outliers: values far from most of the others

Worked example

A teacher records the number of minutes 12 students spend reading:

2, 3, 3, 4, 4, 5, 5, 5, 6, 6, 7, 152,\ 3,\ 3,\ 4,\ 4,\ 5,\ 5,\ 5,\ 6,\ 6,\ 7,\ 15

1. Describe the center

The median is the middle value. There are 12 values, so use the average of the 6th and 7th values:

Median=5+52=5\text{Median}=\frac{5+5}{2}=5

The typical reading time is about 5 minutes.

The mean is:

Mean=6512≈5.4\text{Mean}=\frac{65}{12}\approx 5.4

The mean is a little larger than the median because the value (15)(15) pulls the mean upward.

2. Describe the variability

The range is the greatest value minus the least value:

Range=15−2=13\text{Range}=15-2=13

So, the data spread across 13 minutes.

The middle half of the data goes from (3.5)(3.5) to 66, so the interquartile range is:

IQR=6−3.5=2.5\text{IQR}=6-3.5=2.5

This shows that most values are fairly close together, even though the full range is large.

3. Describe the shape, clusters, gaps, and outliers

  • There is a cluster from 22 to 77, where most values occur.
  • There is a gap from 88 to (14)(14), because no students have times in that interval.
  • The value (15)(15) is a possible outlier because it is far from the other values.
  • The distribution is skewed right because of the long stretch toward the larger value (15)(15).

A complete description is:

The data are centered near 55 minutes, with a median of 55 and a mean of about (5.4)(5.4). Most values are clustered between 22 and 77. The range is (13)(13), but the middle half varies by only (2.5)(2.5). There is a gap from 88 to (14)(14), and (15)(15) is a possible outlier, making the distribution skewed right.

Remember that one number does not tell the whole story. The median, range, and shape together give a clearer description of the distribution.

Learn by doing: Describe data distributions

Click a topic below to practice the foundational skills you'll need, learn the steps, or master this skill

Practice with unlimited practice problems

Statistics - Histograms - Histogram to Distribution Type


    ?