Skip to main content

Section Tuesday, Feb 24

This is an outline of the topics we covered in class. These notes are not a substitute for your own note-taking. I highly recommend that you take your own notes during class. If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.

Subsection More Central Limit Theorem

Example 119.

Suppose we roll a fair D6 100 times, and let \(m\) be the average of the rolls. Estimate \(\Pr(3.45 \leq m \leq 3.55)\text{.}\)
Let \(R_1, \dotsc, R_{100}\) be each roll’s result. Then we have previously calculated \(\E(R_i) = 3.5, \Var(R_i) = \frac{35}{12}\text{.}\) We have:
\begin{align*} m \amp = \frac{R_1 + \dotsb + R_{100}}{100}. \\ \E(m) \amp = \E\left(\frac{R_1 + \dotsb + R_{100}}{100}\right) \\ \amp = \frac{1}{100}\left[\E(R_1) + \dotsb + \E(R_{100})\right] \\ \amp = \frac{1}{100}(100)(3.5) \\ \amp = 3.5 \\ \Var(m) \amp = \Var\left(\frac{R_1 + \dotsb + R_{100}}{100}\right) \\ \amp = \frac{1}{100^2} \left[\Var(R_1) + \dotsb + \Var(R_{100}) \right] \\ \amp = \frac{1}{100^2} (100)\left(\frac{35}{12}\right) \\ \amp = \frac{35}{1200} \end{align*}
So, by the CLT, \(m \approx \Norm\left(3.5, \frac{35}{1200}\right)\text{.}\) Even though \(m\) takes on decimal values, it’s still a discrete random variable. The sum of the rolls can only take on integer values, so \(m\) will take on the values \(1.00, 1.01, \dotsc, 5.99, 6.00\text{.}\) Last time, we described the continuity correction as "extending the range by half a unit’s width in each direction". Here, the width of one unit is \(0.01\text{.}\) So:
\begin{align*} \Pr(3.45 \leq m \leq 3.55) \amp \approx \Pr\left( \frac{3.4445 - 3.5}{\sqrt{35/1200}} \leq Z \leq \frac{3.555 - 3.5}{\sqrt{35/1200}}\right) \\ \amp \approx \Pr(-0.32 \leq Z \leq 0.32) \\ \amp = \Phi(0.32) - \Phi(-0.32) \\ \amp \approx 0.6255 - 0.3745 \\ \amp = 0.2510. \end{align*}

Subsection Confidence Intervals

Idea: suppose we collected (i.i.d.) measurements \(X_1, \dotsc, X_n\) which have some population mean \(\mu\) and variance \(\sigma^2\text{.}\) We calculate the average:
\begin{gather*} A_n = \frac{A_1 + \dotsb + A_n}{n} \end{gather*}
and we want to estimate the value of \(\mu\) from the collected data. Previously, we’ve discussed the MLE, the single most likely estimate of the parameter. Now, we’d like to give a range \([\mu_{\ell}, \mu_h]\) where we can say something like: "we are 95% confident that \(\mu \in [\mu_{\ell}, \mu_h]\text{.}\)"
From the CLT, if \(n\) is large enough, we approximate \(A_n \approx \Norm\left(\mu, \frac{\sigma^2}{n}\right)\text{.}\) We can think of an interval centered at \(\mu\) as a collection of numbers within a certain distance \(d\) of \(\mu\text{.}\) We’d like to choose a distance so that our measured \(A_n\) is within \(d\) of \(\mu\text{.}\) We can shift our point of view here: if \(A_n\) is within \(d\) of \(\mu\text{,}\) then \(\mu\) is within \(d\) of \(A_n\text{.}\)
TODO: image
To pick the distance \(d\text{,}\) we want to make our interval large enough so that there’s only a small amount of area under the normal curve outside of the interval. So:
  1. Pick some value \(\alpha \in (0, 1)\) (representing the area under the curve outside the interval)
  2. Let \(c_{\alpha} = \Phi^{-1}\left(1 - \frac{\alpha}{2}\right)\) (this is the standardized \(z\)-score of the right-hand side of the interval)
  3. Then, the interval will be:
    \begin{gather*} \left[A_n - c_{\alpha} \sqrt{\frac{\sigma^2}{n}}, A_n + c_{\alpha} \sqrt{\frac{\sigma^2}{n}}\right]. \end{gather*}
For \(\alpha = 0.05\text{,}\) we’ll have \(c_{\alpha} = \Phi^{-1}\left(1 - \frac{0.05}{2}\right) = \Phi^{-1}(0.975) = 1.96\text{.}\) So:

Definition 121.

The term \(\sqrt{\frac{\sigma^2}{n}}\) is called the standard error of the mean. (It’s the standard deviation of \(A_n\text{.}\))
We don’t know the true value of \(\sigma^2\text{,}\) just like we don’t know the true value of \(\mu\text{.}\) So we’ll have to use our sample of measurements to estimate \(\sigma^2\text{,}\) and then use that estimate to give us our range of values for \(\mu\text{.}\)

Definition 122.

Let \(X_1, \dotsc, X_n\) be (i.i.d.) measurements. Then the sample mean is:
\begin{gather*} A_n = \frac{X_1 + \dotsb + X_n}{n} \end{gather*}
and the sample variance is:
\begin{gather*} s^2 = \frac{\sum (X_i - A_n)^2}{n - 1} = \frac{\left(\sum X_i^2\right) - n(A_n^2)}{n - 1} \end{gather*}
Why \(n - 1\) here? It turns out that, as defined above, \(s^2\) is an unbiased estimator for \(\sigma^2\text{,}\) whereas it wouldn’t be if we divided by \(n\) instead. (We’ll ommit both the calculation justifying this and a more conceptual explanation for now.)
So, we can adjust the confidence limits from the theorem:
\begin{align*} \mu_{\ell} \amp = A_n - 1.96 \sqrt{\frac{s^2}{n}} \\ \mu_{h} \amp = A_n + 1.96 \sqrt{\frac{s^2}{n}} \end{align*}

Example 123.

Suppose we’re given measurements \(X_1, X_2, X_3, X_4, X_5\) below.
\(i\) \(X_i\)
\(1\) \(19.2\)
\(2\) \(20.1\)
\(3\) \(21.3\)
\(4\) \(20.7\)
\(5\) \(19.8\)
We’re just practicing the computation here, so we’ll pretend that 5 measurements is large enough for the CLT to apply.
\begin{align*} A_n \amp = \frac{19.2 + \dotsb + 19.8}{5} = 20.22 \\ \sum X_i^2 \amp = 19.2^2 + \dotsb + 19.8^2 = 2046.87 \\ s^2 \amp = \frac{2046.87 - 5(20.22^2)}{4} = 0.657 \\ \sqrt{\frac{s^2}{n}} \amp = \sqrt{\frac{0.657}{5}} \approx 0.362 \\ \mu_{\ell} \amp = 20.22 - 1.96(0.362) \approx 19.51 \\ \mu_{h} \amp = 20.22 + 1.96(0.362) \approx 20.93 \end{align*}
This way of computing confidence limits only applies if we can think of the parameter we’re estimating as a mean.

Example 124.

Suppose we flip a coin 100 times and see 40 heads. Find a 98% confidence interval for the bias \(p\text{.}\)
Let \(H_1, \dotsc, H_{100}\) indicate heads on each flip. Let \(S_{100} = H_1 + \dotsb + H_{100}\) and \(A_n = \frac{S_n}{n}\text{.}\) Then:
\begin{align*} \E(A_n) \amp = \E\left(\frac{S_n}{n}\right) = \frac{1}{n}\E(S_n) = \frac{1}{n} np = p \end{align*}
So a 98% confidence interval for \(p\) is also a 98% confidence interval for \(p\text{.}\)
In this setting, we can simplify the computation of \(s^2\) and \(\sqrt{\frac{s^2}{n}}\text{,}\) using our knowledge of the binomial distribution \(S_n\text{.}\)
\begin{align*} \Var(S_n) \amp = np(1 - p) \\ \Var(A_n) \amp = \Var\left(\frac{S_n}{n}\right) = \frac{1}{n^2} \Var(S_n) = \frac{1}{n^2} np(1-p) = \frac{p(1-p)}{n} \end{align*}
So \(s^2 = p(1 - p)\) and \(\sqrt{\frac{s^2}{n}} = \sqrt{\frac{p(1-p)}{n}}\text{.}\) We don’t know the value of \(p\text{,}\) so we’ll use the MLE \(\widehat{p} = \frac{40}{100} = 0.4\text{.}\) So:
\begin{gather*} \sqrt{\frac{s^2}{n}} = \sqrt{\frac{(0.4)(0.6)}{100}} \approx 0.049. \end{gather*}
For a 98% confidence interval, we have \(\alpha = 1 - 0.98 = 0.02\text{,}\) so:
\begin{gather*} c_{\alpha} = \Phi^{-1}\left(1 - \frac{\alpha}{2}\right) = \Phi^{-1}(0.99) \approx 2.33. \end{gather*}
Then, the confidence limits are:
\begin{align*} p_{\ell} \amp = 0.4 - 2.33(0.049) \approx 0.286 \\ p_{h} \amp = 0.4 + 2.33(0.049) \approx 0.514 \end{align*}
For comparison, the 95% confidence limits are:
\begin{align*} p_{\ell} \amp = 0.4 - 1.96(0.049) \approx 0.304 \\ p_{h} \amp = 0.4 + 1.96(0.049) \approx 0.496 \end{align*}