All Exams Test series for 1 year @ ₹349 only
Question

Consider a two-class classification problem, where the densities of the two competing classes are given by 

$f_1(x) = \begin{cases} 1 & \text{if } 0 \leq x \leq 1 \\ 0 & \text{otherwise} \end{cases}$ 

and 

$f_2(x) = \begin{cases} 2x & \text{if } 0 \leq x \leq 1 \\ 0 & \text{otherwise} \end{cases}$. 

Let $\pi_1$ and $\pi_2$ be the prior probabilities of these two classes. Now consider a classifier $\delta$, which classifies an observation $x$ to class $1$ if $x < 1/2$ and to class $2$ if $x \geq 1/2$.

To determine when the classifier \( \delta \) is the Bayes classifier and to find the average probability of misclassification, let's first understand the problem:

We have two classes with probability densities:

  • Class 1 density: \(f_1(x) = \begin{cases} 1 & \text{if } 0 \leq x \leq 1 \\ 0 & \text{otherwise} \end{cases}\)
  • Class 2 density: \(f_2(x) = \begin{cases} 2x & \text{if } 0 \leq x \leq 1 \\ 0 & \text{otherwise} \end{cases}\)

Priors are given as \( \pi_1 \) and \( \pi_2 \), representing the prior probabilities of the classes. A Bayes classifier \(\delta\) minimizes the probability of misclassification.

The given classifier \(\delta\) classifies:

  • To class 1 if \( x < 1/2 \)
  • To class 2 if \( x \geq 1/2 \)

According to the Bayes rule, \(\delta\) is optimal when it classifies by maximizing posterior probabilities. For \(\pi_1 = \pi_2\):

  • When \(0 \leq x < 1/2\): \(f_1(x) > f_2(x) = 2x\) (since \(f_1(x) = 1\)), correctly classified to class 1.
  • When \(1/2 \leq x \leq 1\): \(f_2(x) > f_1(x) = 1\) doesn't hold, but because densities are equal at \(1/2\) and for simplicity equal prior probabilities align decision boundary at \(1/2\).

Hence, when \(\pi_1 = \pi_2\), the classifier \(\delta\) is indeed a Bayes classifier.

Next, for the average probability of misclassification calculation:

  • Misclassification in the region \(0 \leq x < 1/2\) happens when a class 2 point is chosen: \(\int_0^{1/2} 2x \, dx = [x^2]_0^{1/2} = \left(\frac{1}{2}\right)^2 = \frac{1}{4}\).
  • Misclassification in the region \(1/2 \leq x \leq 1\) happens when a class 1 point is chosen: \(\int_{1/2}^1 1 \, dx = [x]_{1/2}^1 = \frac{1}{2}\).
  • The total probability of misclassification: \(\pi_1 \frac{1}{4} + \pi_2 \frac{1}{2}\\).
  • If \(\pi_1 = \pi_2 = \frac{1}{2}\): \(\frac{1}{2} \left(\frac{1}{4}\right) + \frac{1}{2} \left(\frac{1}{2}\right) = \frac{1}{8} + \frac{2}{8} = \frac{3}{8}\).

Conclusively, if \(\pi_1 = \pi_2\), then \(\delta\) is the Bayes classifier, and the average probability of misclassification for \(\delta\) is \(\frac{3}{8}\).

Was this answer helpful?

Important Questions from Elementary Statistics (Notes)

  1. Let $X_1, X_2, X_3$ be a random sample of size 3 from an absolutely continuous distribution that is symmetric about 0. For $i=1,2,3$, let $R_i$ denote the rank of $|X_i|$ among $|X_1|, |X_2|$ and $|X_3|$. 

    If $T^+ = \sum_{i=1, X_i>0}^3 R_i$

     is the Willcoxon signed-rank statistic, then which of the following statements are true?, 

  2. What is the geometric mean of 2, 4 and 8?
  3. In correlation analysis, the two variables

    1. Are treated with distinction.
    2. Are treated differently based on individual characteristics.
    3. Are treated symmetrically.
    4. Are regressed.
  4. In statistics, standard error measures the

    1. Specification error of the model.
    2. Autocorrelation in the regression model.
    3. Correlation between dependent and independent variables.
    4. Precision of an estimate.
  5. Linear regression model is

    1. linear in explanatory variables but may not be linear in parameters
    2. non-linear in parameters and must be linear in variables
    3. linear in parameters and must be linear in variables
    4. linear in parameters and may be linear in variables
Need Expert Advice?

Start Your Preparation with Prepp Mobile App

Download the app from Google Play & App Store
Download the app from Google Play & App Store
Prepp Mobile App