If the data are skewed, which option of central tendency measure is the most unreliable indicator?
Mean
Measures of central tendency help us find a single value that represents the center of a dataset. The most common measures are the mean, median, and mode. However, the shape of the data distribution significantly impacts which measure is the most appropriate or reliable.
Skewness in data refers to the asymmetry of the distribution. In a perfectly symmetrical distribution (like a normal distribution), the mean, median, and mode are often equal or very close. However, in a skewed distribution, one tail is longer than the other, pulling the central tendency measures apart.
Let's look at how each measure of central tendency behaves when the data is skewed:
The mean is calculated by summing all the values in a dataset and dividing by the number of values. It takes into account every single value. The formula for the mean ($\bar{x}$) for a sample is:
$\bar{x} = \frac{\sum x_i}{n}$
Because the mean is influenced by every value, including extreme values or outliers in the tails of a skewed distribution, it gets pulled towards the longer tail. This makes the mean a poor indicator of the typical or central value in a skewed dataset, as it doesn't accurately represent where most of the data points cluster.
The median is the middle value in a dataset that has been ordered from least to greatest. If there's an even number of data points, the median is the average of the two middle values. The median is a positional average; its value depends on the order of the data, not the exact value of every data point.
Since the median is based on the position of values, it is not heavily affected by extreme values in the tails. It remains a good indicator of the center point, dividing the data into two equal halves (50% of values are below the median, and 50% are above). Therefore, the median is generally a reliable measure of central tendency for skewed data.
The mode is the value that appears most frequently in a dataset. A dataset can have one mode (unimodal), multiple modes (multimodal), or no mode. Like the median, the mode is not affected by extreme values.
The mode represents the peak or highest frequency point in the distribution. In a skewed distribution, the mode often remains near the highest concentration of data points. While it tells us the most common value, it might not always be representative of the overall center, especially in distributions with multiple peaks or flat distributions. However, compared to the mean, it is less sensitive to skewness.
The range is the difference between the highest and lowest values in a dataset (Range = Maximum Value - Minimum Value). The range is a measure of data spread or dispersion, not a measure of central tendency. It describes the total spread of the data but doesn't indicate the center point. Therefore, its reliability as a measure of central tendency in skewed data is irrelevant because it isn't a measure of central tendency in the first place.
Based on the analysis, the mean is the measure of central tendency that is most sensitive to extreme values caused by skewness. It gets pulled away from the bulk of the data towards the long tail. The median and mode, being less affected by outliers, are more robust indicators of the center in skewed distributions.
Thus, when data is skewed, the mean is the most unreliable indicator of the typical value or center of the distribution.
| Measure | Description | Affected by Skewness? | Reliability in Skewed Data |
|---|---|---|---|
| Mean | Average of all values | Strongly affected by extreme values | Unreliable; pulled towards the tail |
| Median | Middle value when ordered | Slightly affected, but remains central | Reliable; resistant to outliers |
| Mode | Most frequent value | Generally less affected than mean | More reliable than mean; represents peak |
| Range | Difference between max and min | Highly affected by extreme values | Not a measure of central tendency |
| Term | Definition | Relevance to Skewness |
|---|---|---|
| Central Tendency | A single value representing the center of a dataset. | Skewness impacts which measure best represents the center. |
| Skewness | Asymmetry in a data distribution. | Causes mean, median, and mode to differ. |
| Mean | Arithmetic average. | Most sensitive measure to skewness. |
| Median | Middle value. | Robust measure for skewed data. |
| Mode | Most frequent value. | Less sensitive than mean to skewness. |
Selecting the appropriate measure of central tendency is crucial for accurately describing a dataset. The choice depends primarily on the shape of the distribution and the presence of outliers.
Understanding the shape of your data distribution is the first step in choosing the most reliable measure of central tendency for analysis and interpretation.
In a negatively skewed distribution
If the distribution is negatively skewed, then the:
The first four moments about the mean of distribution are 0, μ 2, 0.7 and 18.75. If the distribution is mesokurtic, the value of μ 2, is
If Mean > Median > Mode, the distribution is:
The following measures were computed for a moderately symmetrical frequency distribution: mean = 50, coefficient of variation = 35% and Karl Pearson's Coefficient of Skewness = - 0.25. The value of the median of the distribution is: