The range is a measure of dispersion that has largely been replaced by the standard deviation. The range is the maximum value of the data set minus the minimum value of the data set.
The standard deviation assumes that the data is evenly distributed
However, in real-world situations, data is rarely evenly distributed. he
For example, if you were to measure the height of every person in a large crowd, the standard deviation would assume that all of these people are equally dispersed across the average population height. he
But we all know that some people are shorter or taller than the average population height. The standard deviation would not take this into account and would therefore underestimate the spread of people in the crowd. he
Because of this issue, some researchers use the range (the highest value – lowest value) as an alternative measure of dispersion. The range assumes that all values in the sample population are different from each other and that none are repeated.
The range does not assume that the data is evenly distributed
As mentioned before, the standard deviation assumes that the data is evenly distributed and that half of the data is above the mean and half of the data is below the mean.
This is not always true and can be a problem when using this measure of dispersion. For example, if you are measuring the height of people, then the standard deviation would say that there is an even chance that someone is taller or shorter than your measured mean person.
However, if there was a sudden decrease in average height due to a new diet fad, then the standard deviation would still assume that half of people are taller or shorter than your measured mean person. This would be incorrect!
The range does not make this assumption and instead simply counts how many numbers are above any given number and below it. The number of numbers on either side equal the range.
The standard deviation takes into account only one variable
While the range does the same, it also takes into account the number of occurrences in a set range. The range does not take into account how close these occurrences are to each other.
The standard deviation is a measure of dispersion that defines the average distance of the data points from the mean. It is essentially a measurement of how spread out the data is.
The standard deviation is typically used instead of the range because it takes into account both the distance of occurrences from one another and the average magnitude of those occurrences.
As such, it provides a better picture of dispersion than simply stating how far apart two instances are. It can be tricky, however, to interpret what the standard deviation means in terms of real-world applications.
The range takes into account multiple variables
While the standard deviation accounts for most variables, it does not take into account the level of the individual variable.
For example, let’s say that the average height of all men is 5’10” and the standard deviation of all men is 2’10”. This means that 95% of all men are within 2’10” of the average height (5’10”) and 5% of all men are above or below this number.
However, if we took a sample size of all men and found that some were 6’0″, some were 4’2″, and some were 5’6″, then the range would be 6′ – 4′ = 2′, which is less than 2′10″. The average would still be 5‘10″ so does not change.
It is easier to determine the bounds of a dataset using the standard deviation
As mentioned before, the standard deviation is the average distance of the data points from the mean.
The range does not take into account how far away the largest and smallest data points are from each other. This makes it harder to determine the bounds of a dataset using the range.
Because we know the standard deviation is a measure of dispersion, we can use that to our advantage. We can determine the bounds of a dataset using both the standard deviation and the range, then compare them to each other to see which is larger.
As with any measurement, both have their pros and cons. It depends on what you are looking for to determine which is more useful to you.
It is easier to determine the bounds of a dataset using the range
The range of a dataset is the total lowest value plus the total highest value of all of the data points. The range does not take into account how dispersed the data is, only how large or small the values are.
Because calculating the range of a dataset is relatively simple, it is a common first step when analyzing data. For example, if we have a dataset that includes people’s ages, then the range of their ages would be youngest age + oldest age.
The problem with using only the range to describe dispersion is that it does not tell you how far apart these values are related to each other. Imagine two datasets: one contains the ages 0–100, and another contains the ages 0–1 billion. Both have an enormous range, but there is virtually no dispersion between these values.
Leave a Reply