Noisy data
Data with additional meaningless information in it
Noisy data are data that are corrupted, distorted, or have a low signal-to-noise ratio. Improper procedures (or improperly documented procedures) to subtract out the noise in data can lead to a false sense of accuracy or false conclusions.
Noisy data are data with a large amount of additional meaningless information in them, known as noise. This includes data corruption and the term is often used as a synonym for corrupt data. It also includes any data that a user system cannot understand and interpret correctly. Many systems, for example, cannot use unstructured text.
Noisy data can adversely affect the results of any data analysis and skew conclusions if not handled properly. Statistical analysis is sometimes used to weed the noise out of noisy data.
01Sources of noise
Differences in real-world measured data from the true values come about from multiple factors affecting the measurement.
Random noise is often a large component of the noise in data. Random noise in a signal is quantified as the signal-to-noise ratio. Random noise contains a wide range of frequencies, and is also called white noise (as wide range of colors of light combine to make white).
Random noise affects the data collection and data preparation processes, where errors commonly occur. Noise has two main sources: errors introduced by measurement tools and random errors introduced by processing or by experts when the data is gathered.
Improper filtering can add noise if the filtered signal is treated as if it were a directly measured signal. As an example, Convolution-type digital filters such a moving average can have side effects such as lags or truncation of peaks. Differentiating digital filters amplifies random noise in the original data.
Outlier data are data that appear to not belong in the data set. It can be caused by human error such as transposing numerals, mislabeling, programming bugs, etc. If actual outliers are not removed from the data set, they corrupt the results to a small or large degree, depending on circumstances. If valid data is identified as an outlier and is mistakenly removed, that also corrupts results.
Individuals may deliberately skew data to influence the results toward a desired conclusion. Data that looks good with few outliers reflects well on the individual collecting it, and so there may be incentive to remove more data as outliers or make the data look smoother than it is.


Sources and credits
This article is adapted from the Wikipedia article “Noisy data”, written by its contributors and licensed under CC BY-SA 4.0. Fathomly has changed the layout, removed citation markers, navigation and maintenance notices, and adjusted punctuation. This adapted version is shared under the same license. For references, see the original article.
Images, from Wikimedia Commons:
- Contextual Outlier.png by Osrecki, CC BY-SA 4.0
- Moving Average Types comparison - Simple and Exponential.png by Alex Kofman, Public domain
Fathomly is not affiliated with or endorsed by the Wikimedia Foundation. Spotted a problem? Tell us.