All Exams Test series for 1 year @ ₹349 only
Question

Which of the following statements are true?

(A) A good measure of dispersion is not duly affected by extreme observations

(B) Standard deviation incorporates the effect of extreme values thus making it most reliable measure of dispersion

(C) Co-efficient of determination (R 2) is always positive and less than one

(D) Non parametric tests do not require us to make restrictive assumptions about shape of population distribution

(E) Z-test can not be applied when population is normal

Choose the correct answer from the options given below:

The correct answer is (A), (C), (D) only

Evaluating Statements on Key Statistical Concepts

The question asks us to identify the true statements among the given options related to various statistical concepts like measures of dispersion, coefficient of determination, and statistical tests.

Let's analyze each statement carefully:

Analysis of Statement (A): Measure of Dispersion and Extreme Observations

Statement (A) says: A good measure of dispersion is not duly affected by extreme observations.

  • Measures of dispersion quantify the spread or variability of a data set.
  • Extreme observations (outliers) can heavily influence some measures of dispersion.
  • A measure of dispersion that is robust to outliers is less affected by them. Examples of such measures include the Interquartile Range (IQR).
  • Measures like variance and standard deviation are sensitive to outliers because they involve squaring deviations from the mean, giving more weight to larger deviations caused by extremes.
  • A measure of dispersion being "good" can depend on the context. In some cases, sensitivity to extremes is desired, while in others, robustness is preferred. However, often, a measure less affected by extremes is considered "good" or robust for describing the typical spread of the majority of the data.

Based on the concept of robustness, statement (A) is considered true, as robust measures of dispersion exist and are often considered "good" in the presence of outliers.

Analysis of Statement (B): Standard Deviation and Reliability

Statement (B) says: Standard deviation incorporates the effect of extreme values thus making it most reliable measure of dispersion.

  • Standard deviation (\(\sigma\) or \(s\)) is calculated using every data point, including extreme values. This means it is affected by outliers.
  • Saying this sensitivity "makes it most reliable" is debatable. While standard deviation is widely used and has desirable mathematical properties, its sensitivity to outliers can make it less representative of the dispersion in the bulk of the data when outliers are present.
  • Measures like the Interquartile Range (IQR) are often considered more reliable measures of spread when the data is skewed or contains outliers, precisely because they are less affected by these extreme values.
  • Therefore, the claim that sensitivity to extreme values makes it the "most reliable" is not universally true and depends on the data distribution and the goal of the analysis.

Statement (B) is considered false because sensitivity to outliers does not necessarily make standard deviation the "most reliable" measure in all situations, especially compared to robust alternatives like the IQR.

Analysis of Statement (C): Coefficient of Determination (R²)

Statement (C) says: Co-efficient of determination (R²) is always positive and less than one.

  • The coefficient of determination, \(R^2\), represents the proportion of the variance in the dependent variable that is predictable from the independent variable(s) in a regression model.
  • For standard linear regression models with an intercept, \(R^2\) is defined as: \[ R^2 = 1 - \frac{\text{Sum of Squares of Residuals (SSR)}}{\text{Total Sum of Squares (SST)}} \]
  • \(SST = \sum (y_i - \bar{y})^2\) and \(SSR = \sum (y_i - \hat{y}_i)^2\). Both are sums of squared values, so they are non-negative.
  • In a standard model with an intercept, the model fitting process ensures that \(SSR \le SST\).
  • Thus, \(\frac{SSR}{SST} \le 1\), which implies \(R^2 = 1 - \frac{SSR}{SST} \ge 0\).
  • Also, \(SSR \ge 0\), so \(R^2 = 1 - \frac{SSR}{SST} \le 1 - 0 = 1\).
  • Therefore, for a standard linear regression model with an intercept, \(R^2\) is always between 0 and 1, inclusive (\(0 \le R^2 \le 1\)).
  • The statement "always positive and less than one" implies \(0 < R^2 < 1\). However, \(R^2\) can be exactly 0 (if the model explains none of the variance) or exactly 1 (if the model explains all the variance perfectly).
  • In the context of typical MCQ options, "positive" often implies non-negative, and "less than one" often implies less than or equal to one. If interpreted as \(0 \le R^2 \le 1\), the statement is true for standard models. Given the options, this interpretation is likely intended.

Interpreting "positive" as non-negative (\(\ge 0\)) and "less than one" as less than or equal to one (\(\le 1\)), statement (C) is considered true for standard linear regression.

Analysis of Statement (D): Non-parametric Tests

Statement (D) says: Non parametric tests do not require us to make restrictive assumptions about shape of population distribution.

  • Parametric tests (like t-tests, Z-tests, ANOVA) typically assume that the data comes from a specific type of distribution, most commonly the normal distribution. They also make assumptions about parameters like means and variances.
  • Non-parametric tests (like Mann-Whitney U, Wilcoxon rank-sum, Kruskal-Wallis, chi-squared) do not rely on these restrictive assumptions about the exact shape (e.g., normality) or parameters of the population distribution.
  • They are often based on ranks or signs of the data rather than the raw values themselves.

Statement (D) accurately describes a key characteristic of non-parametric tests and is considered true.

Analysis of Statement (E): Z-test Application

Statement (E) says: Z-test can not be applied when population is normal.

  • The Z-test is a parametric statistical test.
  • A primary condition for applying a Z-test (especially for small sample sizes) is that the population from which the sample is drawn is normally distributed, and the population standard deviation is known.
  • Even if the population is not normal, the Z-test can be applied with large sample sizes due to the Central Limit Theorem, which states that the distribution of sample means will be approximately normal regardless of the population distribution shape, provided the sample size is large enough.
  • The statement says the Z-test *cannot* be applied when the population is normal. This is the opposite of the truth; normality is a condition under which the Z-test is valid.

Statement (E) is considered false because the Z-test *can* be applied when the population is normal.

Summary of Statement Analysis

Based on the analysis:

  • Statement (A): True
  • Statement (B): False
  • Statement (C): True (assuming \(0 \le R^2 \le 1\) interpretation)
  • Statement (D): True
  • Statement (E): False

The true statements are (A), (C), and (D).

Let's check the given options:

  1. (A), (C), (D) only
  2. (B), (D), (E) only
  3. (D), (E) only
  4. (A), (D), (E) only

The option that lists only statements (A), (C), and (D) as true is the correct one.

Therefore, the correct answer is option 1.

Revision Table: Statistical Statements

Statement Evaluation Reasoning
(A) A good measure of dispersion is not duly affected by extreme observations True Describes robust measures like IQR, which are often considered "good".
(B) Standard deviation incorporates the effect of extreme values thus making it most reliable measure of dispersion False Sensitivity to outliers doesn't make it "most reliable" in all contexts; robustness is often preferred.
(C) Co-efficient of determination (R²) is always positive and less than one True For standard linear regression, \(0 \le R^2 \le 1\). (Interpreting "positive" as non-negative and "less than one" as \(\le 1\)).
(D) Non parametric tests do not require us to make restrictive assumptions about shape of population distribution True This is a defining characteristic of non-parametric tests.
(E) Z-test can not be applied when population is normal False Z-test is valid when the population is normal.

Additional Information: Deep Dive into Concepts

Measures of Dispersion

Measures of dispersion tell us how spread out the data points are in a distribution. Common measures include:

  • Range: The difference between the maximum and minimum values. Highly sensitive to outliers.
  • Variance (\(\sigma^2\) or \(s^2\)): The average of the squared deviations from the mean. Sensitive to outliers.
  • Standard Deviation (\(\sigma\) or \(s\)): The square root of the variance. Represents the typical distance of data points from the mean. Also sensitive to outliers.
  • Interquartile Range (IQR): The range of the middle 50% of the data (\(Q_3 - Q_1\)). Robust to outliers.

The choice of dispersion measure depends on the nature of the data and the research question. For data with outliers or skewed distributions, IQR might be more appropriate than standard deviation.

Coefficient of Determination (\(R^2\))

In simple linear regression, \(R^2\) is the square of the correlation coefficient (\(r\)). It quantifies the proportion of the variance in the dependent variable explained by the independent variable(s).

  • \(R^2 = 0\) means the model explains none of the variance in the dependent variable.
  • \(R^2 = 1\) means the model explains all of the variance in the dependent variable (a perfect fit).
  • \(0 < R^2 < 1\) means the model explains some, but not all, of the variance.

While \(R^2\) is typically between 0 and 1 for standard regression with an intercept, it can be negative in specific cases, such as when a model without an intercept is forced, or when comparing non-nested models using certain software defaults. However, for the vast majority of practical applications taught at an introductory level, \(R^2\) falls within the [0, 1] range.

Parametric vs. Non-parametric Tests

  • Parametric Tests: Make assumptions about the population parameters and often the distribution shape (e.g., normality). Examples: Z-test, t-test, ANOVA, Pearson correlation. They are generally more powerful than non-parametric tests if their assumptions are met.
  • Non-parametric Tests: Do not make strong assumptions about the population distribution shape. They are often used when assumptions for parametric tests are violated or with ordinal/ranked data. Examples: Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis, Spearman rank correlation, Chi-squared test. They are less powerful than parametric tests when parametric assumptions hold but are more robust to violations of assumptions.

Z-test

The Z-test is used to test hypotheses about a population mean or compare two population means, typically when the population standard deviation is known. Key conditions for its valid application include:

  • The population is normally distributed, OR
  • The sample size is large (typically \(n \ge 30\)) so that the Central Limit Theorem ensures the sampling distribution of the mean is approximately normal.
  • The population standard deviation (\(\sigma\)) is known. If \(\sigma\) is unknown and estimated from the sample standard deviation (s), the t-test is generally more appropriate, especially for small sample sizes.

Therefore, the Z-test is indeed applicable, and often preferred (if \(\sigma\) is known), when the population distribution is normal.

Was this answer helpful?

Important Questions from Measurement and Analysis of Data - Teaching

  1. Given below are two statements

    Statement I: The qualitative data are powerful because they are collected from very sensitive social, historical and temporal context.

    Statement II: Context sensitivity cannot be completely removed from the qualitative data.

    In light of the above statements, choose the correct answer from the options given below

  2. Given below is a summary of ANOVA for four groups of students tested in a research project:

    Source of varianceSS (Sum of squares)df (Degree of freedom)MS (Mean sum of squares)
    Between groups76323.33
    Within groups122167.62

    What will be the value of 'F' for the above data?

  3. An investigator used ANOVA to compare four groups of students on numerical ability on the basis of a test. After analysis of raw scores, the following results were obtained:

    Source of variationdfSum of Squares
    Between Groups3625.00
    Within Groups362128.00

    The value of F-ratio would be approximate:

  4. In randomly constituted two groups-experimental and control, a researcher obtains the following results after using a parametric 't' test:

    Value of t = 3 for N = 300

    On the basis of this evidence which decision in respect of substantive research hypothesis and the null hypothesis will be justified?

  5. Given below are two statements, one labelled as Assertion (A) and the other labelled as Reason (R). Read the statements and choose the correct answer using the code given below.

    Assertion (A): Homogenous tests have low reliability.

    Reason (R): The range of test scores affects reliability.

Need Expert Advice?

Start Your Preparation with Prepp Mobile App

Download the app from Google Play & App Store
Download the app from Google Play & App Store
Prepp Mobile App