Exam 2

Show all relevant work.

Write either True or False for each of the following statements. No justification is required.

Suppose X_1, \dotsc, X_n are any random variables and S = (X_1 + X_2 + X_3 + \dotsb + X_n)^2. Then S is approximately normally distributed.

False.

Let Z be the standard normal distribution. For any x, \Pr(Z \geq x) = 1 - \Phi(x).

True.

If we're using a set of data to find a confidence interval for a parameter, then the 95% confidence interval will be wider than the 90% confidence interval.

True.

In hypothesis testing, a type I error is the probability of accepting the null hypothesis.

False.

For a particular set of collected data, the p-value of a 2-tailed test will be as large or larger than the p-value of a 1-tailed test.

True.

In a \chi^2 test, if \chi^2 \lt 0.05, then we can reject the null hypothesis.

False.

In each of the following scenarios, determine whether we should use a 1-tailed test or a 2-tailed test. Indicate clearly which words or phrase in the description of the scenario would lead you to draw your conclusion.

We wonder whether the population of mice in a city have a different average weight to the population of mice in a rural town.

2-tailed, because we wonder whether the cities have "different average weight" without suspecting a direction of difference.

A new fertilizer is designed to increase the yield of a particular species of fruit, and we wonder whether the fertilizer is effective.

1-tailed, because the new fertilizer is "designed to increase the yield", so we suspect the change happens in a particular direction.

We find a weighted 6-sided die and suspect that it will roll a 6 more often than a fair die would.

1-tailed, because we suspect the die will roll 6 "more often", indicating difference in a particular direction.

A coin has a probability \theta = 0.7 of coming up heads. We flip the coin 180 times, and write S for the number of heads. Estimate \Pr(120 \leq S \leq 130). (Use a continuity correction if appropriate.)

S is binomially distributed with \E(S) = (180)(0.7) = 126 and \Var(S) = (180)(0.7)(0.3) = 37.8. Therefore: \Pr(120 \leq S \leq 130) \amp \approx \Pr\left( \frac{119.5 - 126}{\sqrt{37.8}} \leq Z \leq \frac{130.5 - 126}{\sqrt{37.8}}\right) \amp \approx \Pr(-1.06 \leq Z \leq 0.73) \amp = \Phi(0.73) - \Phi(-1.06) \amp = \Phi(0.73) - \Phi(-1.06) \amp \approx 0.7673 - 0.1446 \amp = 0.6227.

The weights of five apples are measured and recorded below. Find the sample mean, sample variance, and a 95% confidence interval around the sample mean for the weights of the apples. (Pretend that 5 measurements is enough for the CLT to apply. Do not use the t-distribution.)

Apple i 1 2 3 4 5 Weight W_i (g) 140 120 165 168 137

A = \text{Avg weight} \amp = \frac{\sum W_i}{5} = 146 \sum W_i^2 \amp = 108218 s^2 \amp = \frac{\sum W_i^2 - n(A^2)}{n -1} = \frac{108218 - 5(146^2)}{4} = 409.5 \sqrt{\frac{s^2}{n}} \amp = \sqrt{\frac{409.5}{5}} \approx 9.05 \mu_{\ell} \amp = 146 - 1.96(9.05) = 128.262 \mu_{h} \amp = 146 + 1.96(9.05) = 163.738

Suppose we're told that a proportion of \theta = 0.4 of a population has a particular trait, but we suspect that this information is incorrect. We sample 160 people from the population and see 76 people with the trait. Is this enough evidence to reject the null hypothesis at the 0.05 significance level?

"We suspect the information is incorrect" indicates a 2-tailed test. Let S be the number of people in the sample with the trait. Then, under H_0, S \sim \Bin(160, 0.4), so \E(S) = (160)(0.4) = 64 and \Var(S) = (160)(0.4)(0.6) = 38.4. The observed 76 people is 12 more than expected. Equally extreme in the other direction would be 12 fewer, so 52 people. Therefore the p-value is: \Pr(S \geq 76) + \Pr(S \leq 52) \amp \approx \Pr\left( Z \geq \frac{75.5 - 64}{\sqrt{38.4}} \right) + \Pr\left(Z \leq \frac{52.5 - 64}{\sqrt{38.4}}\right) \amp \approx \Pr(Z \geq 1.86) + \Pr(Z \leq -1.86) \amp = 2 \Phi(-1.86) \amp = 2 (0.0314) \amp = 0.0628 \gt 0.05, so we do not reject H_0.

Suppose we have a coin that we're told has probability \theta = 0.4 of coming up heads, but we suspect it comes up heads more often than that. We flip the coin 160 times.

What is the minimum number of heads we would need to see to reject the null hypothesis?

"More often" suggests we should use a 1-tailed test. Let S be the number of heads. Then, under H_0, S \sim \Bin(160, 0.4), so \E(S) = (160)(0.4) = 64 and \Var(S) = (160)(0.4)(0.6) = 38.4. Therefore, the p-value when observing k heads would be: \Pr(S \geq k) \amp \approx \Pr\left( Z \geq \underbrace{\frac{k - 0.5 - 64}{\sqrt{38.4}}}_{z} \right) \Pr(Z \geq z) \amp = 0.05 1 - \Phi(z) \amp = 0.05 \quad \Rightarrow \quad \Phi(z) = 0.95. We would need the z-score to be at least 1.65 to reject H_0, so: \frac{k - 0.5 - 64}{\sqrt{38.4}} \amp = 1.65 k = 64.5 + 1.65 \sqrt{38.4} \amp \approx 74.72 Therefore, it would take at least 75 heads to reject H_0.

If the coin is actually fair, what is the power of this test?

If the coin is fair, then \E(S) = (160)(0.5) = 80 and \Var(S) = (160)(0.5)(0.5) = 40. The probability of rejecting H_0 is then: \Pr(S \geq 75) \amp \approx \Pr\left(Z \geq \frac{74.5 - 80}{\sqrt{40}}\right) \amp \approx \Pr(Z \geq -0.87) \amp \approx 1 - \Phi(-0.87) \amp \approx 1 - 0.1922 \amp = 0.8078.

A store owner assumes an equal number of people on average come into the store each day of the week. Are the following counts of 400 observed shoppers consistent with this hypothesis?

Day Mon Tue Wed Thu Fri Observed Shoppers 83 61 65 90 101

Under the null hypothesis, we would expect 400/5 = 80 people each day. So: \chi^2 \amp = \frac{(83 - 80)^2}{80} + \dotsb + \frac{(101 - 80)^2}{80} = 14.2. The number of degrees of freedom is (5 - 1) - (0) = 4, so the critical value is 9.488. Since 14.2 \gt 9.488, we reject the null hypothesis.