Reference articles on history, science, culture and more
Encyclopedia

Statistical dispersion

Statistical property quantifying how much a collection of data is spread out

Image credit is listed at the end of this article.

In statistics, dispersion (also called variability, scatter, or spread) is the extent to which a distribution is stretched or squeezed. Common examples of measures of statistical dispersion are the variance, standard deviation, and interquartile range. For instance, when the variance of data in a set is large, the data is widely scattered. On the other hand, when the variance is small, the data in the set is clustered.

Dispersion is contrasted with location or central tendency, and together they are the most used properties of distributions.

01Measures of statistical dispersion

A measure of statistical dispersion is a nonnegative real number that is zero if all the data are the same and increases as the data become more diverse.

Most measures of dispersion have the same units as the quantity being measured. In other words, if the measurements are in metres or seconds, so is the measure of dispersion. Examples of dispersion measures include:

These are frequently used (together with scale factors) as estimators of scale parameters, in which capacity they are called estimates of scale. Robust measures of scale are those unaffected by a small number of outliers, and include the IQR and MAD.

All the above measures of statistical dispersion have the useful property that they are location-invariant and linear in scale. This means that if a random variable X has a dispersion of S_{X} then a linear transformation Y=aX+b for real a and b should have dispersion S_{Y}=|a|S_{X}, where |a| is the absolute value of a, that is, ignores a preceding negative sign -.

Other measures of dispersion are dimensionless. In other words, they have no units even if the variable itself has units. These include:

There are other measures of dispersion:

Some measures of dispersion have specialized purposes. The Allan variance can be used for applications where the noise disrupts convergence. The Hadamard variance can be used to counteract linear frequency drift sensitivity.

For categorical variables, it is less common to measure dispersion by a single number; see qualitative variation. One measure that does so is the discrete entropy.

02A partial ordering of dispersion

A mean-preserving spread (MPS) is a change from one probability distribution A to another probability distribution B, where B is formed by spreading out one or more portions of A's probability density function while leaving the mean (the expected value) unchanged. The concept of a mean-preserving spread provides a partial ordering of probability distributions according to their dispersions: of two probability distributions, one may be ranked as having more dispersion than the other, or alternatively neither may be ranked as having more dispersion.

Watch videos about Statistical dispersionExplainers and documentaries on YouTube (opens in a new tab)

Sources and credits

This article is adapted from the Wikipedia article Statistical dispersion, written by its contributors and licensed under CC BY-SA 4.0. Fathomly has changed the layout, removed citation markers, navigation and maintenance notices, and adjusted punctuation. This adapted version is shared under the same license. For references, see the original article.

Images, from Wikimedia Commons:

Fathomly is not affiliated with or endorsed by the Wikimedia Foundation. Spotted a problem? Tell us.