One Sample Tests

Suppose 10% of the general population is left-handed. In a sample of 100 patients with carpal tunnel syndrome, 16 are found to be left-handed. Is this evidence that the proportion of left-handedness is higher among carpal tunnel syndrome patients than among the general population?

We would like to develop a framework through which we can analyze whether data we've collected provides evidence for a particular hypothesis. In fact, we will generally consider two hypotheses:

The null hypothesis, often denoted H_0, is the assumption that an effect being studied or proposed does not exist. The alternative hypothesis, H_a, is the claim that the effect does exist.

In the previous example, if we write \theta for the true proportion of left-handed people among the population with carpal tunnel syndrome, the null hypothesis would be that this proportion is the same as the general population. The alternative hypothesis might be, for example, that the proportion is higher than among the general population: H_0: \theta \amp = 0.1 H_a: \theta \amp \gt 0.1

After performing an experiment, we will either accept or reject the null hypothesis based on the data we collect. Consider the following scenarios:

accept H_0 reject H_0 H_0 true correct type I error H_0 false type II error correct

We'll use the notation \alpha for the probability of making a type I error, also called the significance level, and \beta for the probability of making a type II error. The power of a test is the probability of correctly rejecting H_0, i.e., 1 - \beta.

The p-value is the probability of observing a result at least as extreme as measured if H_0 is true.

Continuing the previous example, under H_0 that \theta = 0.1, the probability of seeing at least 16 left-handed people in a sample of 100 people would be: \sum_{k = 16}^{100} {100 \choose k} (0.1)^k (1 - 0.1)^{100 - k} \approx 0.04, which is our p-value.

The goal of our calculation is to control the chance of making a type I error by choosing a significance level cutoff, often 0.05. If our calculated p-value is below the cutoff, then reject H_0. Otherwise, accept H_0. In the previous example, we would reject H_0, because it appears that the data we collected is pretty unlikely to see if the null hypothesis were true. We would accept that, about 4% of the time if the null hypothesis is true, we would see data at least this extreme, and therefore make a mistake by rejecting H_0.

The test we just applied is called 1-tailed. The alternative hypothesis \theta \gt 0.1 proposed that \theta was different from 0.1 in a specific direction. For a 2-tailed test, we could use the alternative hypothesis that \theta \neq 0.1. In this setting, the idea of data "at least as extreme" as what was measured is reframed. We measured 16 left-handed people in a sample of 100, which is 6 more than we would expect under the null hypothesis. So we should also include the possibility of seeing at least 6 fewer left-handed people in the sample than expected: p\text{-value} \amp = \Pr(\geq 16 \text{ left-handed people}) + \Pr(\leq 4 \text{ left-handed people}) \amp = \left(\sum_{k = 16}^{100} {100 \choose k} (0.1)^k (0.9)^{100 - k}\right) + \left(\sum_{k = 0}^{4} {100 \choose k} (0.1)^k (0.9)^{100 - k}\right) \amp \approx 0.064 \gt 0.05, so, using a 2-tailed test, we would fail to reject the null hypothesis.

In suitable situations, we can also use a normal approximation via to calculate the p-value.

We find a coin on the street and wonder if it's a fair coin. We flip it 100 times and see 62 heads. Is this strong evidence to reject the null hypothesis of a fair coin at a 0.05 significance level?

Let S be the random variable which counts the number of heads in 100 flips. The null hypothesis H_0 of a fair coin would mean the parameter \theta = \Pr(\text{heads}) = 0.5. (We'll avoid the letter p for the parameter to prevent confusion with the new term, p-value.) So: \E(S) \amp = n\theta = 100(0.5) = 50 \Var(S) \amp = n\theta(1-\theta) = 100(0.5)(0.5) = 25 Then, by , S \approx \Norm(50, 25). So, using a 2-tailed test, the p-value is: \Pr(S \geq 62) + \Pr(S \leq 38) \amp \approx \Pr\left(Z \geq \frac{61.5 - 50}{\sqrt{25}}\right) + \Pr\left(Z \leq \frac{38.5 - 50}{\sqrt{25}}\right) \amp = (1 - \Phi(2.3)) + \Phi(-2.3) \amp \approx (1 - 0.9893) + 0.0107 \amp = 0.0214 \lt 0.05, so this is strong enough evidence to reject H_0.

In each of the following scenarios, determine whether we should use a 1-tailed test or a 2-tailed test.

We have a coin which we've flipped many times, seeing an above-average number of heads. We suspect the coin comes up heads more often than a fair coin would.

We find a coin on the street and wonder whether or not it's a fair coin.

We suspect there will be a difference in average weight of mice caught during the summer versus during the winter.

In a medical study, a group of patients are gathered and the proportion experiencing particular symptoms is measured. A new drug intended to eliminate these symptoms is administered, after which the proportion experiencing symptoms is measured again.

Suppose we find a coin and wonder whether it's fair. As a first test, we decide to flip the coin 200 times and count the number of heads. If we see 112 heads, should we accept or reject the null hypothesis of a fair coin at a significance level of 0.05?

Suppose we have a coin which we suspect comes up heads more often than a fair coin would. As a first test, we decide to flip the coin 200 times and count the number of heads. If we see 112 heads, should we accept or reject the null hypothesis of a fair coin at a significance level of 0.05?

Suppose we find a six-sided die and wonder whether it's fair. As a first test, we decide to roll the die 100 times and count the number of times it comes up 1. If we roll 22 1's, should we accept or reject the null hypothesis of a fair die at a significance level of 0.05?

Suppose a particular plant when grown outdoors has an average height of 39 in with a variance of 20 in^2. We suspect that growing this plant in a greenhouse will increase its height. A sample of 50 plants grown in a greenhouse has an average height of 40 in. Is this significant enough data to reject the null hypothesis of equal means at the p = 0.05 significance level?