diff --git a/source/docinfo.ptx b/source/docinfo.ptx index 3447d98..d00332c 100644 --- a/source/docinfo.ptx +++ b/source/docinfo.ptx @@ -19,6 +19,7 @@ \DeclareMathOperator{\Geom}{Geom} \DeclareMathOperator{\Poiss}{Poiss} \DeclareMathOperator{\Exp}{Exp} + \DeclareMathOperator{\Norm}{N} \DeclareMathOperator{\E}{E} diff --git a/source/main.ptx b/source/main.ptx index 0fa5d9a..759ca1a 100644 --- a/source/main.ptx +++ b/source/main.ptx @@ -31,6 +31,8 @@ + + diff --git a/source/notes/2-19.ptx b/source/notes/2-19.ptx new file mode 100644 index 0000000..080dd48 --- /dev/null +++ b/source/notes/2-19.ptx @@ -0,0 +1,174 @@ + + +
+ Thursday, Feb 19 + + +

+ This is an outline of the topics we covered in class. + These notes are not a substitute for your own note-taking. + I highly recommend that you take your own notes during class. + If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes. +

+
+ + + + Central Limit Theorem + + + +

+ A continuous random variable X has the normal distribution with mean \mu and variance \sigma^2 if it has pdf: + + f(x; \mu, \sigma^2) = \frac{1}{\sqrt{2\pi \sigma^2}} e^{\frac{-(x - \mu)^2}{2\sigma^2}} + + We'll write X \sim \Norm(\mu, \sigma^2). +

+ +

+ The specific distribution \Norm(0, 1) is called the standard normal distribution, and its pdf has the notation: + + f(x; 0, 1) = \phi(x). + +

+
+
+ +

+ If X \sim \Norm(\mu, \sigma^2), then: + + \Pr(a \leq X \leq b) = \int_a^b f(x; \mu, \sigma^2)\ dx. + + Unfortunately, e^{-x^2} has no elementary antiderivative. + But we know \int_a^b f(x)\ dx represents the area under the graph y = f(x; \mu, \sigma^2), and this area can be approximated to arbitrary precision. +

+ + + +

+ Suppose X_1, \dotsc, X_n are independent and identically distributed (or i.i.d.) with finite expected value \mu and finite variance \sigma^2. + Define: + + S_n \amp = X_1 + X_2 + \dotsb + X_n + A_n \amp = \frac{X_1 + X_2 + \dotsb + X_n}{n} = \frac{S_n}{n} + + Then, for large enough n: + + S_n \amp \approx \Norm(n\mu, n\sigma^2) + A_n \amp \approx \Norm\left(\mu, \frac{\sigma^2}{n}\right) + +

+
+
+ +

+ We'll use two general rules of thumb for determining whether the number of measurements n is large enough for the approximation to be a good one: +

    +
  1. +

    + n \geq 30 +

    +
  2. + +
  3. +

    + If S\sim \Bin(n, p), then n should be large enough so that there are at least 5 heads and 5 tails. +

    +
  4. +
+

+ +

+ We'll omit a proof of , but we can at least check that the expected values and variances of S_n, A_n are correct. + Given \E(X_i) = \mu, \Var(X_i) = \sigma^2: + + \E(S_n) \amp = \E(X_1 + \dotsb + X_n) + \amp = \E(X_1) + \dotsb + \E(X_n) + \amp = \mu + \dotsb + \mu + \amp = n \mu. \checkmark + \Var(S_n) \amp = \Var(X_1 + \dotsb + X_n) + \amp = \Var(X_1) + \dotsb + \Var(X_n) + \amp = \sigma^2 + \dotsb + \sigma^2 + \amp = n \sigma^2. \checkmark + \E(A_n) \amp = \E\left(\frac{S_n}{n}\right) + \amp = \frac{1}{n} \E(S_n) + \amp = \frac{1}{n} n \mu + \amp = \frac{1}{n} n \mu + \amp = \mu. \checkmark + \Var(A_n) \amp = \Var\left(\frac{S_n}{n}\right) + \amp = \frac{1}{n^2} \Var(S_n) + \amp = \frac{1}{n^2} n \sigma^2 + \amp = \frac{\sigma^2}{n}. \checkmark + +

+ + + +

+ Suppose we flip a fair coin 100 times, and let S be the number of heads. + Estimate \Pr(40 \leq S \leq 60). +

+ +

+ Since S \sim \Bin(n, p), we know: + + \E(S) \amp = np = (100)(0.5) = 50 + \Var(S) \amp = np(1-p) = (100)(0.5)(1 - 0.5) = 25 + + So S \approx \Norm(50, 25). + Then: + + \Pr(40 \leq S \leq 60) \approx \int_{40}^{60} f(x; 50, 25)\ dx, + + but we can't compute this integral. +

+ +

+ Instead, we can standardize by shifting and scaling: + + Z = \frac{S - 50}{\sqrt{25}} = \frac{S - 50}{5}. + + Then Z \sim \Norm(0, 1). + Values of Z are called z-scores, sometimes indicated by an asterisk. + (I.e., if a is a value of S, then a^* is the corresponding z-score.) Now: + + \Pr(40 \leq S \leq 60) \amp \approx \Pr\left(\frac{40 - 50}{5} \leq Z \leq \frac{60 - 50}{5}\right) + \amp = \Pr(-2 \leq Z \leq 2) + \amp = \int_{-2}^2 \phi(x)\ dx. + + Now, we still can't compute the value of the integral. + However, there's a huge benefit to translating to the standard normal distribution, regardless of which normal distribution we used to approximate S. + We can approximate integrals like \int_a^b \phi(x)\ dx to abritrary precision and record the results in a table. + Then, we can look up the values when needed. + We don't need to redo our approximations for different normal distributions, we just standardize whatever normal distribution we come across. +

+ +

+ So: + + \Pr(40 \leq S \leq 60) \amp \approx \Pr(-2 \leq Z \leq 2) + \amp = \Phi(2) - \Phi(-2) + \amp \approx 0.9772 - 0.0228 + \amp \approx 0.9544. + +

+
+
+ +

+ TODO: complex image introducing the idea of the continuity correction. +

+ +

+ We can make the approximation more accurate by extending the range of S-values by half a unit in each direction: + + \Pr(40 \leq S \leq 60) \amp \approx \Pr\left(\frac{39.5 - 50}{5} \leq Z \leq \frac{60.5 - 50}{5}\right) + \amp = \Pr(-2.1 \leq Z \leq 2.1) + \amp = \Phi(2.1) - \Phi(-2.1) + \amp \approx 0.9821 - 0.0179 + \amp \approx 0.9642. + +

+
+
\ No newline at end of file diff --git a/source/notes/2-24.ptx b/source/notes/2-24.ptx new file mode 100644 index 0000000..bdb3740 --- /dev/null +++ b/source/notes/2-24.ptx @@ -0,0 +1,278 @@ + + +
+ Tuesday, Feb 24 + + +

+ This is an outline of the topics we covered in class. + These notes are not a substitute for your own note-taking. + I highly recommend that you take your own notes during class. + If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes. +

+
+ + + + More Central Limit Theorem + + + +

+ Suppose we roll a fair D6 100 times, and let m be the average of the rolls. + Estimate \Pr(3.45 \leq m \leq 3.55). +

+ +

+ Let R_1, \dotsc, R_{100} be each roll's result. + Then we have previously calculated \E(R_i) = 3.5, \Var(R_i) = \frac{35}{12}. + We have: + + m \amp = \frac{R_1 + \dotsb + R_{100}}{100}. + \E(m) \amp = \E\left(\frac{R_1 + \dotsb + R_{100}}{100}\right) + \amp = \frac{1}{100}\left[\E(R_1) + \dotsb + \E(R_{100})\right] + \amp = \frac{1}{100}(100)(3.5) + \amp = 3.5 + \Var(m) \amp = \Var\left(\frac{R_1 + \dotsb + R_{100}}{100}\right) + \amp = \frac{1}{100^2} \left[\Var(R_1) + \dotsb + \Var(R_{100}) \right] + \amp = \frac{1}{100^2} (100)\left(\frac{35}{12}\right) + \amp = \frac{35}{1200} + + So, by the CLT, m \approx \Norm\left(3.5, \frac{35}{1200}\right). + Even though m takes on decimal values, it's still a discrete random variable. + The sum of the rolls can only take on integer values, so m will take on the values 1.00, 1.01, \dotsc, 5.99, 6.00. + Last time, we described the continuity correction as "extending the range by half a unit's width in each direction". + Here, the width of one unit is 0.01. + So: + + \Pr(3.45 \leq m \leq 3.55) \amp \approx \Pr\left( \frac{3.4445 - 3.5}{\sqrt{35/1200}} \leq Z \leq \frac{3.555 - 3.5}{\sqrt{35/1200}}\right) + \amp \approx \Pr(-0.32 \leq Z \leq 0.32) + \amp = \Phi(0.32) - \Phi(-0.32) + \amp \approx 0.6255 - 0.3745 + \amp = 0.2510. + +

+
+
+
+ + + + Confidence Intervals + +

+ Idea: suppose we collected (i.i.d.) measurements X_1, \dotsc, X_n which have some population mean \mu and variance \sigma^2. + We calculate the average: + + A_n = \frac{A_1 + \dotsb + A_n}{n} + + and we want to estimate the value of \mu from the collected data. + Previously, we've discussed the MLE, the single most likely estimate of the parameter. + Now, we'd like to give a range [\mu_{\ell}, \mu_h] where we can say something like: "we are 95% confident that \mu \in [\mu_{\ell}, \mu_h]." +

+ +

+ From the CLT, if n is large enough, we approximate A_n \approx \Norm\left(\mu, \frac{\sigma^2}{n}\right). + We can think of an interval centered at \mu as a collection of numbers within a certain distance d of \mu. + We'd like to choose a distance so that our measured A_n is within d of \mu. + We can shift our point of view here: if A_n is within d of \mu, then \mu is within d of A_n. +

+ +

+ TODO: image +

+ +

+ To pick the distance d, we want to make our interval large enough so that there's only a small amount of area under the normal curve outside of the interval. + So: +

    +
  1. +

    + Pick some value \alpha \in (0, 1) (representing the area under the curve outside the interval) +

    +
  2. + +
  3. +

    + Let c_{\alpha} = \Phi^{-1}\left(1 - \frac{\alpha}{2}\right) (this is the standardized z-score of the right-hand side of the interval) +

    +
  4. + +
  5. +

    + Then, the interval will be: + + \left[A_n - c_{\alpha} \sqrt{\frac{\sigma^2}{n}}, A_n + c_{\alpha} \sqrt{\frac{\sigma^2}{n}}\right]. + +

    +
  6. +
+

+ +

+ For \alpha = 0.05, we'll have c_{\alpha} = \Phi^{-1}\left(1 - \frac{0.05}{2}\right) = \Phi^{-1}(0.975) = 1.96. + So: +

+ + + +

+ The 95% confidence limits are given by: + + \mu_{\ell} \amp = A_n - 1.96 \sqrt{\frac{\sigma^2}{n}} + \mu_{h} \amp = A_n + 1.96 \sqrt{\frac{\sigma^2}{n}} + +

+
+
+ + + +

+ The term \sqrt{\frac{\sigma^2}{n}} is called the standard error of the mean. + (It's the standard deviation of A_n.) +

+
+
+ +

+ We don't know the true value of \sigma^2, just like we don't know the true value of \mu. + So we'll have to use our sample of measurements to estimate \sigma^2, and then use that estimate to give us our range of values for \mu. +

+ + + +

+ Let X_1, \dotsc, X_n be (i.i.d.) measurements. + Then the sample mean is: + + A_n = \frac{X_1 + \dotsb + X_n}{n} + + and the sample variance is: + + s^2 = \frac{\sum (X_i - A_n)^2}{n - 1} = \frac{\left(\sum X_i^2\right) - n(A_n^2)}{n - 1} + +

+
+
+ +

+ Why n - 1 here? It turns out that, as defined above, s^2 is an unbiased estimator for \sigma^2, whereas it wouldn't be if we divided by n instead. + (We'll ommit both the calculation justifying this and a more conceptual explanation for now.) +

+ +

+ So, we can adjust the confidence limits from the theorem: + + \mu_{\ell} \amp = A_n - 1.96 \sqrt{\frac{s^2}{n}} + \mu_{h} \amp = A_n + 1.96 \sqrt{\frac{s^2}{n}} + +

+ + + +

+ Suppose we're given measurements X_1, X_2, X_3, X_4, X_5 below. +

+ + + + i + X_i + + + + 1 + 19.2 + + + + 2 + 20.1 + + + + 3 + 21.3 + + + + 4 + 20.7 + + + + 5 + 19.8 + + + +

+ We're just practicing the computation here, so we'll pretend that 5 measurements is large enough for the CLT to apply. + + A_n \amp = \frac{19.2 + \dotsb + 19.8}{5} = 20.22 + \sum X_i^2 \amp = 19.2^2 + \dotsb + 19.8^2 = 2046.87 + s^2 \amp = \frac{2046.87 - 5(20.22^2)}{4} = 0.657 + \sqrt{\frac{s^2}{n}} \amp = \sqrt{\frac{0.657}{5}} \approx 0.362 + \mu_{\ell} \amp = 20.22 - 1.96(0.362) \approx 19.51 + \mu_{h} \amp = 20.22 + 1.96(0.362) \approx 20.93 + +

+
+
+ +

+ This way of computing confidence limits only applies if we can think of the parameter we're estimating as a mean. +

+ + + +

+ Suppose we flip a coin 100 times and see 40 heads. + Find a 98% confidence interval for the bias p. +

+ +

+ Let H_1, \dotsc, H_{100} indicate heads on each flip. + Let S_{100} = H_1 + \dotsb + H_{100} and A_n = \frac{S_n}{n}. + Then: + + \E(A_n) \amp = \E\left(\frac{S_n}{n}\right) = \frac{1}{n}\E(S_n) = \frac{1}{n} np = p + + So a 98% confidence interval for p is also a 98% confidence interval for p. +

+ +

+ In this setting, we can simplify the computation of s^2 and \sqrt{\frac{s^2}{n}}, using our knowledge of the binomial distribution S_n. + + \Var(S_n) \amp = np(1 - p) + \Var(A_n) \amp = \Var\left(\frac{S_n}{n}\right) = \frac{1}{n^2} \Var(S_n) = \frac{1}{n^2} np(1-p) = \frac{p(1-p)}{n} + + So s^2 = p(1 - p) and \sqrt{\frac{s^2}{n}} = \sqrt{\frac{p(1-p)}{n}}. + We don't know the value of p, so we'll use the MLE \widehat{p} = \frac{40}{100} = 0.4. + So: + + \sqrt{\frac{s^2}{n}} = \sqrt{\frac{(0.4)(0.6)}{100}} \approx 0.049. + +

+ +

+ For a 98% confidence interval, we have \alpha = 1 - 0.98 = 0.02, so: + + c_{\alpha} = \Phi^{-1}\left(1 - \frac{\alpha}{2}\right) = \Phi^{-1}(0.99) \approx 2.33. + + Then, the confidence limits are: + + p_{\ell} \amp = 0.4 - 2.33(0.049) \approx 0.286 + p_{h} \amp = 0.4 + 2.33(0.049) \approx 0.514 + + For comparison, the 95% confidence limits are: + + p_{\ell} \amp = 0.4 - 1.96(0.049) \approx 0.304 + p_{h} \amp = 0.4 + 1.96(0.049) \approx 0.496 + +

+
+
+
+
\ No newline at end of file