A multiple choice test consists of 50 test items. The difficulty values of items vary in a narrow range with an average value of 0.48. The test gave a variance of 81 when administered to a group of students. The reliability co efficient of the test would be :
0.86
Test reliability is a crucial concept in educational and psychological measurement. It refers to the consistency of a test score over time or across different forms of the test. A reliable test provides stable and dependable results. One common method for estimating the internal consistency reliability of a test with dichotomous items (like multiple-choice questions scored as right or wrong) is the Kuder-Richardson Formula 20 (KR-20).
The question provides us with the following information about a multiple choice test:
We need to find the reliability coefficient of the test.
The Kuder-Richardson Formula 20 (KR-20) is given by:
$$r_{KR-20} = \left( \frac{n}{n-1} \right) \left( 1 - \frac{\sum pq}{S_t^2} \right)$$
Where:
The problem states that the difficulty values vary in a "narrow range with an average value of 0.48". While we don't have the individual item difficulties ($p_i$) to calculate the exact $\sum p_i q_i$, a common approximation when only the average difficulty ($\bar{p}$) is provided is to assume that the average of the item variances is approximately the variance of an item with the average difficulty. Thus, $\sum p_i q_i \approx n \times \bar{p}(1 - \bar{p})$.
Using this approximation:
Now we can plug the values into the KR-20 formula:
$$r_{KR-20} \approx \left( \frac{50}{49} \right) \left( 1 - \frac{12.48}{81} \right)$$
First, calculate the fraction term:
$$\frac{50}{49} \approx 1.0204$$
Next, calculate the term inside the parenthesis:
$$\frac{12.48}{81} \approx 0.15407$$
$$1 - \frac{12.48}{81} \approx 1 - 0.15407 = 0.84593$$
Finally, multiply the two parts:
$$r_{KR-20} \approx 1.0204 \times 0.84593 \approx 0.8631$$
The calculated reliability coefficient is approximately 0.8631. Let's compare this value with the given options:
| Option | Value |
|---|---|
| 1 | 0.91 |
| 2 | 0.86 |
| 3 | 0.78 |
| 4 | 0.67 |
The calculated value 0.8631 is closest to 0.86.
Based on the Kuder-Richardson Formula 20 and the provided test statistics (number of items, average difficulty, and test variance), the estimated reliability coefficient of the test is approximately 0.86. A reliability coefficient of 0.86 indicates a reasonably high level of internal consistency for the test scores, suggesting that the items are measuring a similar construct and the scores are quite consistent.
| Term | Definition/Meaning |
|---|---|
| Test Reliability | Consistency of test scores. |
| Kuder-Richardson Formula 20 (KR-20) | A formula to estimate internal consistency reliability for tests with dichotomous items. |
| Number of Items ($n$) | Total count of questions in the test. |
| Item Difficulty ($p_i$) | Proportion of test-takers who answered item $i$ correctly. |
| Item Variance ($p_i q_i$) | Variance for item $i$, where $q_i = 1 - p_i$. A measure of how much responses vary for that item. |
| Test Variance ($S_t^2$) | Variance of the total scores on the test. A measure of the spread of scores. |
| Internal Consistency | The degree to which items within a test measure the same construct. |
In test theory, understanding variance is key to calculating reliability. The total variance of a test ($S_t^2$) reflects the overall spread of scores among test-takers. It's influenced by both the true differences in ability (true variance) and random errors (error variance).
For dichotomous items, the item variance ($p_i q_i$) is highest when item difficulty ($p_i$) is 0.50 ($0.50 \times 0.50 = 0.25$) and decreases as difficulty approaches 0 or 1. Items with difficulty near 0.50 contribute more to the variance of the total test score (if they discriminate well) than very easy or very hard items. The sum of item variances ($\sum pq$) represents the variance attributed to the individual items, assuming their scores are independent.
The KR-20 formula essentially compares the sum of item variances ($\sum pq$) to the total test variance ($S_t^2$). If the sum of item variances is much smaller than the total test variance, it implies that the items are related to each other and contribute to a larger, common variance (the true score variance), leading to higher reliability. Conversely, if the sum of item variances is large relative to the total test variance, it suggests more random error or that items are not measuring the same thing consistently, resulting in lower reliability.
The term $\left( \frac{n}{n-1} \right)$ is a correction factor applied in the formula.
Given below are two statements
Statement I: The qualitative data are powerful because they are collected from very sensitive social, historical and temporal context.
Statement II: Context sensitivity cannot be completely removed from the qualitative data.
In light of the above statements, choose the correct answer from the options given below
Given below is a summary of ANOVA for four groups of students tested in a research project:
| Source of variance | SS (Sum of squares) | df (Degree of freedom) | MS (Mean sum of squares) |
| Between groups | 76 | 3 | 23.33 |
| Within groups | 122 | 16 | 7.62 |
What will be the value of 'F' for the above data?
An investigator used ANOVA to compare four groups of students on numerical ability on the basis of a test. After analysis of raw scores, the following results were obtained:
| Source of variation | df | Sum of Squares |
| Between Groups | 3 | 625.00 |
| Within Groups | 36 | 2128.00 |
The value of F-ratio would be approximate:
In randomly constituted two groups-experimental and control, a researcher obtains the following results after using a parametric 't' test:
Value of t = 3 for N = 300
On the basis of this evidence which decision in respect of substantive research hypothesis and the null hypothesis will be justified?
Given below are two statements, one labelled as Assertion (A) and the other labelled as Reason (R). Read the statements and choose the correct answer using the code given below.
Assertion (A): Homogenous tests have low reliability.
Reason (R): The range of test scores affects reliability.