diff --git a/source/docinfo.ptx b/source/docinfo.ptx
index 3447d98..d00332c 100644
--- a/source/docinfo.ptx
+++ b/source/docinfo.ptx
@@ -19,6 +19,7 @@
\DeclareMathOperator{\Geom}{Geom}
\DeclareMathOperator{\Poiss}{Poiss}
\DeclareMathOperator{\Exp}{Exp}
+ \DeclareMathOperator{\Norm}{N}
\DeclareMathOperator{\E}{E}
diff --git a/source/main.ptx b/source/main.ptx
index 0fa5d9a..759ca1a 100644
--- a/source/main.ptx
+++ b/source/main.ptx
@@ -31,6 +31,8 @@
+
+
diff --git a/source/notes/2-19.ptx b/source/notes/2-19.ptx
new file mode 100644
index 0000000..080dd48
--- /dev/null
+++ b/source/notes/2-19.ptx
@@ -0,0 +1,174 @@
+
+
+
+ Thursday, Feb 19
+
+
+
+ This is an outline of the topics we covered in class.
+ These notes are not a substitute for your own note-taking.
+ I highly recommend that you take your own notes during class.
+ If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
+
+
+
+
+
+ Central Limit Theorem
+
+
+
+
+ A continuous random variable X has the normal distribution with mean \mu and variance \sigma^2 if it has pdf:
+
+ f(x; \mu, \sigma^2) = \frac{1}{\sqrt{2\pi \sigma^2}} e^{\frac{-(x - \mu)^2}{2\sigma^2}}
+
+ We'll write X \sim \Norm(\mu, \sigma^2).
+
+
+
+ The specific distribution \Norm(0, 1) is called the standard normal distribution, and its pdf has the notation:
+
+ f(x; 0, 1) = \phi(x).
+
+
+
+
+
+
+ If X \sim \Norm(\mu, \sigma^2), then:
+
+ \Pr(a \leq X \leq b) = \int_a^b f(x; \mu, \sigma^2)\ dx.
+
+ Unfortunately, e^{-x^2} has no elementary antiderivative.
+ But we know \int_a^b f(x)\ dx represents the area under the graph y = f(x; \mu, \sigma^2), and this area can be approximated to arbitrary precision.
+
+
+
+
+
+ Suppose X_1, \dotsc, X_n are independent and identically distributed (or i.i.d.) with finite expected value \mu and finite variance \sigma^2.
+ Define:
+
+ S_n \amp = X_1 + X_2 + \dotsb + X_n
+ A_n \amp = \frac{X_1 + X_2 + \dotsb + X_n}{n} = \frac{S_n}{n}
+
+ Then, for large enough n:
+
+ S_n \amp \approx \Norm(n\mu, n\sigma^2)
+ A_n \amp \approx \Norm\left(\mu, \frac{\sigma^2}{n}\right)
+
+
+
+
+
+
+ We'll use two general rules of thumb for determining whether the number of measurements n is large enough for the approximation to be a good one:
+
+ -
+
+ n \geq 30
+
+
+
+ -
+
+ If S\sim \Bin(n, p), then n should be large enough so that there are at least 5 heads and 5 tails.
+
+
+
+
+
+
+ We'll omit a proof of , but we can at least check that the expected values and variances of S_n, A_n are correct.
+ Given \E(X_i) = \mu, \Var(X_i) = \sigma^2:
+
+ \E(S_n) \amp = \E(X_1 + \dotsb + X_n)
+ \amp = \E(X_1) + \dotsb + \E(X_n)
+ \amp = \mu + \dotsb + \mu
+ \amp = n \mu. \checkmark
+ \Var(S_n) \amp = \Var(X_1 + \dotsb + X_n)
+ \amp = \Var(X_1) + \dotsb + \Var(X_n)
+ \amp = \sigma^2 + \dotsb + \sigma^2
+ \amp = n \sigma^2. \checkmark
+ \E(A_n) \amp = \E\left(\frac{S_n}{n}\right)
+ \amp = \frac{1}{n} \E(S_n)
+ \amp = \frac{1}{n} n \mu
+ \amp = \frac{1}{n} n \mu
+ \amp = \mu. \checkmark
+ \Var(A_n) \amp = \Var\left(\frac{S_n}{n}\right)
+ \amp = \frac{1}{n^2} \Var(S_n)
+ \amp = \frac{1}{n^2} n \sigma^2
+ \amp = \frac{\sigma^2}{n}. \checkmark
+
+
+
+
+
+
+ Suppose we flip a fair coin 100 times, and let S be the number of heads.
+ Estimate \Pr(40 \leq S \leq 60).
+
+
+
+ Since S \sim \Bin(n, p), we know:
+
+ \E(S) \amp = np = (100)(0.5) = 50
+ \Var(S) \amp = np(1-p) = (100)(0.5)(1 - 0.5) = 25
+
+ So S \approx \Norm(50, 25).
+ Then:
+
+ \Pr(40 \leq S \leq 60) \approx \int_{40}^{60} f(x; 50, 25)\ dx,
+
+ but we can't compute this integral.
+
+
+
+ Instead, we can standardize by shifting and scaling:
+
+ Z = \frac{S - 50}{\sqrt{25}} = \frac{S - 50}{5}.
+
+ Then Z \sim \Norm(0, 1).
+ Values of Z are called z-scores, sometimes indicated by an asterisk.
+ (I.e., if a is a value of S, then a^* is the corresponding z-score.) Now:
+
+ \Pr(40 \leq S \leq 60) \amp \approx \Pr\left(\frac{40 - 50}{5} \leq Z \leq \frac{60 - 50}{5}\right)
+ \amp = \Pr(-2 \leq Z \leq 2)
+ \amp = \int_{-2}^2 \phi(x)\ dx.
+
+ Now, we still can't compute the value of the integral.
+ However, there's a huge benefit to translating to the standard normal distribution, regardless of which normal distribution we used to approximate S.
+ We can approximate integrals like \int_a^b \phi(x)\ dx to abritrary precision and record the results in a table.
+ Then, we can look up the values when needed.
+ We don't need to redo our approximations for different normal distributions, we just standardize whatever normal distribution we come across.
+
+
+
+ So:
+
+ \Pr(40 \leq S \leq 60) \amp \approx \Pr(-2 \leq Z \leq 2)
+ \amp = \Phi(2) - \Phi(-2)
+ \amp \approx 0.9772 - 0.0228
+ \amp \approx 0.9544.
+
+
+
+
+
+
+ TODO: complex image introducing the idea of the continuity correction.
+
+
+
+ We can make the approximation more accurate by extending the range of S-values by half a unit in each direction:
+
+ \Pr(40 \leq S \leq 60) \amp \approx \Pr\left(\frac{39.5 - 50}{5} \leq Z \leq \frac{60.5 - 50}{5}\right)
+ \amp = \Pr(-2.1 \leq Z \leq 2.1)
+ \amp = \Phi(2.1) - \Phi(-2.1)
+ \amp \approx 0.9821 - 0.0179
+ \amp \approx 0.9642.
+
+
+
+
\ No newline at end of file
diff --git a/source/notes/2-24.ptx b/source/notes/2-24.ptx
new file mode 100644
index 0000000..bdb3740
--- /dev/null
+++ b/source/notes/2-24.ptx
@@ -0,0 +1,278 @@
+
+
+
+ Tuesday, Feb 24
+
+
+
+ This is an outline of the topics we covered in class.
+ These notes are not a substitute for your own note-taking.
+ I highly recommend that you take your own notes during class.
+ If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
+
+
+
+
+
+ More Central Limit Theorem
+
+
+
+
+ Suppose we roll a fair D6 100 times, and let m be the average of the rolls.
+ Estimate \Pr(3.45 \leq m \leq 3.55).
+
+
+
+ Let R_1, \dotsc, R_{100} be each roll's result.
+ Then we have previously calculated \E(R_i) = 3.5, \Var(R_i) = \frac{35}{12}.
+ We have:
+
+ m \amp = \frac{R_1 + \dotsb + R_{100}}{100}.
+ \E(m) \amp = \E\left(\frac{R_1 + \dotsb + R_{100}}{100}\right)
+ \amp = \frac{1}{100}\left[\E(R_1) + \dotsb + \E(R_{100})\right]
+ \amp = \frac{1}{100}(100)(3.5)
+ \amp = 3.5
+ \Var(m) \amp = \Var\left(\frac{R_1 + \dotsb + R_{100}}{100}\right)
+ \amp = \frac{1}{100^2} \left[\Var(R_1) + \dotsb + \Var(R_{100}) \right]
+ \amp = \frac{1}{100^2} (100)\left(\frac{35}{12}\right)
+ \amp = \frac{35}{1200}
+
+ So, by the CLT, m \approx \Norm\left(3.5, \frac{35}{1200}\right).
+ Even though m takes on decimal values, it's still a discrete random variable.
+ The sum of the rolls can only take on integer values, so m will take on the values 1.00, 1.01, \dotsc, 5.99, 6.00.
+ Last time, we described the continuity correction as "extending the range by half a unit's width in each direction".
+ Here, the width of one unit is 0.01.
+ So:
+
+ \Pr(3.45 \leq m \leq 3.55) \amp \approx \Pr\left( \frac{3.4445 - 3.5}{\sqrt{35/1200}} \leq Z \leq \frac{3.555 - 3.5}{\sqrt{35/1200}}\right)
+ \amp \approx \Pr(-0.32 \leq Z \leq 0.32)
+ \amp = \Phi(0.32) - \Phi(-0.32)
+ \amp \approx 0.6255 - 0.3745
+ \amp = 0.2510.
+
+
+
+
+
+
+
+
+ Confidence Intervals
+
+
+ Idea: suppose we collected (i.i.d.) measurements X_1, \dotsc, X_n which have some population mean \mu and variance \sigma^2.
+ We calculate the average:
+
+ A_n = \frac{A_1 + \dotsb + A_n}{n}
+
+ and we want to estimate the value of \mu from the collected data.
+ Previously, we've discussed the MLE, the single most likely estimate of the parameter.
+ Now, we'd like to give a range [\mu_{\ell}, \mu_h] where we can say something like: "we are 95% confident that \mu \in [\mu_{\ell}, \mu_h]."
+
+
+
+ From the CLT, if n is large enough, we approximate A_n \approx \Norm\left(\mu, \frac{\sigma^2}{n}\right).
+ We can think of an interval centered at \mu as a collection of numbers within a certain distance d of \mu.
+ We'd like to choose a distance so that our measured A_n is within d of \mu.
+ We can shift our point of view here: if A_n is within d of \mu, then \mu is within d of A_n.
+
+
+
+ TODO: image
+
+
+
+ To pick the distance d, we want to make our interval large enough so that there's only a small amount of area under the normal curve outside of the interval.
+ So:
+
+ -
+
+ Pick some value \alpha \in (0, 1) (representing the area under the curve outside the interval)
+
+
+
+ -
+
+ Let c_{\alpha} = \Phi^{-1}\left(1 - \frac{\alpha}{2}\right) (this is the standardized z-score of the right-hand side of the interval)
+
+
+
+ -
+
+ Then, the interval will be:
+
+ \left[A_n - c_{\alpha} \sqrt{\frac{\sigma^2}{n}}, A_n + c_{\alpha} \sqrt{\frac{\sigma^2}{n}}\right].
+
+
+
+
+
+
+
+ For \alpha = 0.05, we'll have c_{\alpha} = \Phi^{-1}\left(1 - \frac{0.05}{2}\right) = \Phi^{-1}(0.975) = 1.96.
+ So:
+
+
+
+
+
+ The 95% confidence limits are given by:
+
+ \mu_{\ell} \amp = A_n - 1.96 \sqrt{\frac{\sigma^2}{n}}
+ \mu_{h} \amp = A_n + 1.96 \sqrt{\frac{\sigma^2}{n}}
+
+
+
+
+
+
+
+
+ The term \sqrt{\frac{\sigma^2}{n}} is called the standard error of the mean.
+ (It's the standard deviation of A_n.)
+
+
+
+
+
+ We don't know the true value of \sigma^2, just like we don't know the true value of \mu.
+ So we'll have to use our sample of measurements to estimate \sigma^2, and then use that estimate to give us our range of values for \mu.
+
+
+
+
+
+ Let X_1, \dotsc, X_n be (i.i.d.) measurements.
+ Then the sample mean is:
+
+ A_n = \frac{X_1 + \dotsb + X_n}{n}
+
+ and the sample variance is:
+
+ s^2 = \frac{\sum (X_i - A_n)^2}{n - 1} = \frac{\left(\sum X_i^2\right) - n(A_n^2)}{n - 1}
+
+
+
+
+
+
+ Why n - 1 here? It turns out that, as defined above, s^2 is an unbiased estimator for \sigma^2, whereas it wouldn't be if we divided by n instead.
+ (We'll ommit both the calculation justifying this and a more conceptual explanation for now.)
+
+
+
+ So, we can adjust the confidence limits from the theorem:
+
+ \mu_{\ell} \amp = A_n - 1.96 \sqrt{\frac{s^2}{n}}
+ \mu_{h} \amp = A_n + 1.96 \sqrt{\frac{s^2}{n}}
+
+
+
+
+
+
+ Suppose we're given measurements X_1, X_2, X_3, X_4, X_5 below.
+
+
+
+
+ | i |
+ X_i |
+
+
+
+ | 1 |
+ 19.2 |
+
+
+
+ | 2 |
+ 20.1 |
+
+
+
+ | 3 |
+ 21.3 |
+
+
+
+ | 4 |
+ 20.7 |
+
+
+
+ | 5 |
+ 19.8 |
+
+
+
+
+ We're just practicing the computation here, so we'll pretend that 5 measurements is large enough for the CLT to apply.
+
+ A_n \amp = \frac{19.2 + \dotsb + 19.8}{5} = 20.22
+ \sum X_i^2 \amp = 19.2^2 + \dotsb + 19.8^2 = 2046.87
+ s^2 \amp = \frac{2046.87 - 5(20.22^2)}{4} = 0.657
+ \sqrt{\frac{s^2}{n}} \amp = \sqrt{\frac{0.657}{5}} \approx 0.362
+ \mu_{\ell} \amp = 20.22 - 1.96(0.362) \approx 19.51
+ \mu_{h} \amp = 20.22 + 1.96(0.362) \approx 20.93
+
+
+
+
+
+
+ This way of computing confidence limits only applies if we can think of the parameter we're estimating as a mean.
+
+
+
+
+
+ Suppose we flip a coin 100 times and see 40 heads.
+ Find a 98% confidence interval for the bias p.
+
+
+
+ Let H_1, \dotsc, H_{100} indicate heads on each flip.
+ Let S_{100} = H_1 + \dotsb + H_{100} and A_n = \frac{S_n}{n}.
+ Then:
+
+ \E(A_n) \amp = \E\left(\frac{S_n}{n}\right) = \frac{1}{n}\E(S_n) = \frac{1}{n} np = p
+
+ So a 98% confidence interval for p is also a 98% confidence interval for p.
+
+
+
+ In this setting, we can simplify the computation of s^2 and \sqrt{\frac{s^2}{n}}, using our knowledge of the binomial distribution S_n.
+
+ \Var(S_n) \amp = np(1 - p)
+ \Var(A_n) \amp = \Var\left(\frac{S_n}{n}\right) = \frac{1}{n^2} \Var(S_n) = \frac{1}{n^2} np(1-p) = \frac{p(1-p)}{n}
+
+ So s^2 = p(1 - p) and \sqrt{\frac{s^2}{n}} = \sqrt{\frac{p(1-p)}{n}}.
+ We don't know the value of p, so we'll use the MLE \widehat{p} = \frac{40}{100} = 0.4.
+ So:
+
+ \sqrt{\frac{s^2}{n}} = \sqrt{\frac{(0.4)(0.6)}{100}} \approx 0.049.
+
+
+
+
+ For a 98% confidence interval, we have \alpha = 1 - 0.98 = 0.02, so:
+
+ c_{\alpha} = \Phi^{-1}\left(1 - \frac{\alpha}{2}\right) = \Phi^{-1}(0.99) \approx 2.33.
+
+ Then, the confidence limits are:
+
+ p_{\ell} \amp = 0.4 - 2.33(0.049) \approx 0.286
+ p_{h} \amp = 0.4 + 2.33(0.049) \approx 0.514
+
+ For comparison, the 95% confidence limits are:
+
+ p_{\ell} \amp = 0.4 - 1.96(0.049) \approx 0.304
+ p_{h} \amp = 0.4 + 1.96(0.049) \approx 0.496
+
+
+
+
+
+
\ No newline at end of file