From e5d3538af5776b19f1b615c4cbf1160b0dca2cea Mon Sep 17 00:00:00 2001 From: andyeisenberg Date: Thu, 12 Feb 2026 07:09:15 -0500 Subject: [PATCH] Notes up through 2-10 --- source/docinfo.ptx | 1 + source/main.ptx | 2 + source/notes/2-10.ptx | 259 ++++++++++++++++++++++++++++++++++++++++++ source/notes/2-3.ptx | 235 ++++++++++++++++++++++++++++++++++++++ source/notes/2-5.ptx | 207 +++++++++++++++++++++++++++++++++ 5 files changed, 704 insertions(+) create mode 100644 source/notes/2-10.ptx create mode 100644 source/notes/2-5.ptx diff --git a/source/docinfo.ptx b/source/docinfo.ptx index 95d059c..3447d98 100644 --- a/source/docinfo.ptx +++ b/source/docinfo.ptx @@ -24,6 +24,7 @@ \DeclareMathOperator{\E}{E} \DeclareMathOperator{\Var}{Var} \DeclareMathOperator{\Cov}{Cov} + \newcommand{\L}{\mathcal{L}} diff --git a/source/main.ptx b/source/main.ptx index 0ea7f5f..4c1af32 100644 --- a/source/main.ptx +++ b/source/main.ptx @@ -29,6 +29,8 @@ + + diff --git a/source/notes/2-10.ptx b/source/notes/2-10.ptx new file mode 100644 index 0000000..c170eaf --- /dev/null +++ b/source/notes/2-10.ptx @@ -0,0 +1,259 @@ + + +
+ Tuesday, Feb 10 + + +

+ This is an outline of the topics we covered in class. + These notes are not a substitute for your own note-taking. + I highly recommend that you take your own notes during class. + If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes. +

+
+ + + + Likelihood + +

+ In probability theory, we start with some probability model and its parameter values, and we try to find the probabilities of seeing certain types of data. + In statistics, we start with the collected data, and we try to find the most likely values of parameters for some underlying probability model. +

+ + + +

+ An estimator is a way of estimating a parameter value based on data collected. +

+
+
+ + + +

+ Suppose we flip a coin 100 times and see 52 heads. + Let p be the bias of the coin. + Then we might estimate: + + \widehat{p} = \frac{52}{100} + + (The notation \widehat{p} is sometimes used to indicate an estimator for p.) +

+
+
+ + + +

+ If we see k heads in n flips, then the estimator \widehat{p} = \frac{k}{n} is called the common sense estimator for the binomial distribution parameter p. +

+
+
+ + + +

+ An estimator \widehat{p} is unbiased if \E(\widehat{p}) = p. +

+
+
+ + + +

+ Flip a coin 100 times, and let S be the number of heads. + Then S \sim \Bin(100, p). + Let \widehat{p} = \frac{S}{100} = \frac{S}{n}. + Then: + + \E(\widehat{p}) \amp = \E\left(\frac{S}{n}\right) + \amp = \frac{1}{n} \E\left(S\right) + \amp = \frac{1}{n} (np) + \amp = p. + + So \E(\widehat{p} = p, i.e., the common sense estimator is unbiased. +

+
+
+ + + +

+ We perform an experiment and collect data. + Let p be an unknown parameter value. + The likelihood function is: + + \L(p) = \Pr(\text{data} \mid \text{paramater value is } p). + +

+
+
+ + + +

+ Let S \sim \Bin(100, p). + Suppose we see 52 heads. + Then: + + \Pr(S = k) \amp = b(k; 100, p) = {100 \choose k}p^k (1 - p)^{100 - k} + \L(p) \amp = {100 \choose 52}p^{52} (1 - p)^{48} + + In the first line, the variable k represents the data. + In the second line, the variable p represents the parameter value. + Our goal, given the collected data, is to find the maximum likelihood estimation (MLE) for the parameter value. +

+ +

+ \L(p) is a continuous function over a closed interval p \in [0, 1], so we use the Closed Interval Method. + + \L'(p) \amp = {100 \choose 52}\left[ 52 p^{51} (1 - p)^{48} + p^{52}48 (1-p)^{47}(-1)\right] + \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 (1 - p) - 48p \right] + \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 - 100p \right] + + Now, we look for critical numbers in the interior of the interval: + + \L'(p) = 0 \text{ when } p = \frac{52}{100} + + Finally, we test the critical numbers and the endpoints of the interval to find the max: +

+ + + Check Candidate Locations for Max + + + + p + \L(p) + + + + 0 + 0 + + + + 52/100 + \gt 0 + + + + 1 + 0 + + +
+ +

+ So \widehat{p} = \frac{52}{100} is the MLE. +

+ +

+ More generally, a similar calculation will show that, with k heads in n flips, the MLE will be \widehat{p} = \frac{k}{n}. +

+
+
+ + + +

+ Suppose we observe a cell, measuring the time T until a toxin molecule leaves the cell. + Then T \sim \Exp(\lambda) for some \lambda, with pdf + + f(t) =\lambda e^{-\lambda t}, \lambda \gt 0 + + If we see a toxin molecule leave at 0.3 min, what's the MLE for \lambda? + + \L(\lambda) = \lambda e^{-0.3 \lambda} \quad \text{(density, not probability)} + + We want to maximize \L(\lambda) over \lambda \in (0, \infty), an open interval. + So we'll use the "Open Interval Method". + + \L'(\lambda) \amp = e^{-0.3\lambda} + \lambda e^{-0.3\lambda}(-0.3) + \amp = e^{-0.3\lambda}\left( 1 - 0.3 e^{-0.3\lambda}\right) + + Then \L'(\lambda) = 0 when \lambda = \frac{1}{0.3}\approx 3.33. + Checking \lambda values to the left and the right: + + \L'(1) \amp = (+)(+) = (+) + \L'(10) \amp = (+)(-) = (-) + + So \L'(\lambda) \gt 0 (and therefore \L(\lambda) is increasing) on (0, 1/0.3), and \L'(\lambda) \lt 0 (and therefore \L(\lambda) is decreasing) on (1/0.3, \infty). + Now we can conclude that \widehat{\lambda} = 1/0.3 \approx 3.33 is the location of a global (and not just local) maximum value. +

+ +

+ More generally, if the observed time is t, then the MLE will be \widehat{\lambda} = \frac{1}{t}. +

+ +

+ What if we had more data points? For example, suppose two toxin molecules leave the cell at t_1 = 0.3 min and t_2 = 0.5 min? +

+ + + Waiting Times + + + + Molecule + Time + Rate Estimate + + + + 1 + 0.3 + 1/0.3 \approx 3.33 + + + + 2 + 0.5 + 1/0.5 = 2 + + +
+ +

+ How do we combine these data points? We could take the average of the rate estimates: + + \frac{3.33 + 2}{2} \approx 2.67 + + Alternatively, we could average the times first, then create a new rate estimate from the average time: + + \frac{0.3 + 0.5}{2} \amp = 0.4 + \frac{1}{0.4} \amp = 2.5 + + Both of these make some sense, but let's do a careful computation to be certain which way is correct (if either of them is!). + + \L(\lambda) \amp = \left( \lambda e^{-0.3 \lambda} \right)\left( \lambda e^{-0.5 \lambda} \right) + \amp = \lambda^2 e^{-0.3 \lambda - 0.5 \lambda} + \amp = \lambda^2 e^{- 0.8 \lambda} + + Using the Open Interval Method: + + \L'(\lambda) \amp = 2\lambda e^{-0.8 \lambda} + \lambda^2 e^{-0.8 \lambda} (-0.8) + \amp = \lambda e^{-0.8 \lambda} \left[ 2 - 0.8\lambda \right] + + So \L'(\lambda) = 0 when \lambda = \frac{2}{0.8} = 2.5. + Testing points to the left and right: + + \L'(1) \amp = (+)(+)(+) = (+) + \L'(5) \amp = (+)(+)(-) = (-) + + Now \L'(\lambda) \gt 0 (\L(\lambda) is increasing) on (0, 2.5), and \L'(\lambda) \lt 0 (\L(\lambda) is decreasing) on (2.5, \infty). + Therefore, the MLE is \widehat{\lambda} = \frac{2}{0.8} = 2.5. +

+ +

+ Tracing the values 2 and 0.8 throughout the calculation, we can see that the value 2 will generally match the number of waiting times collected, and the value 0.8 will be the sum of the waiting times. + So, generally, with collected waiting times of t_1, \dotsc, t_n, the MLE will be: + + \widehat{\lambda} = \frac{n}{t_1 + \dotsb + t_n} = \frac{1}{\text{avg time}}. + +

+
+
+
+
\ No newline at end of file diff --git a/source/notes/2-3.ptx b/source/notes/2-3.ptx index 288fe9c..2f64c38 100644 --- a/source/notes/2-3.ptx +++ b/source/notes/2-3.ptx @@ -37,5 +37,240 @@

+ +

+ We can write an alternative formula here: + + \Var(X) \amp = \E\left[ (X - \mu)^2 \right] + \amp = \E\left[ X^2 - 2\mu X + \mu^2 \right] + \amp = \E(X^2) - 2\mu \E(X) + \E(\mu^2) + \amp = \E(X^2) - 2\mu^2 + \mu^2 + \amp = \E(X^2) - \mu^2. + + This formula is generally more useful for performing computations. +

+ + + +

+ Let X indicate event A with \Pr(A) = p. +

+ + + + Distribution for <m>X</m> + + + + k + \Pr(X = k) + + + + 0 + 1 - p + + + + 1 + p + + +
+ + + Distribution for <m>X^2</m> + + + + k + \Pr(X^2 = k) + + + + 0^2 + 1 - p + + + + 1^2 + p + + +
+
+ +

+ Then \E(X) = p and \E(X^2) = p, so: + + \Var(X) = \E(X^2) - \left(\E(X)\right)^2 = p - p^2 = p(1 - p). + +

+
+
+ + + +

+ Let R be the roll of a fair D6. +

+ + + + Distribution for <m>R</m> + + + + k + \Pr(R = k) + + + + 1 + 1/6 + + + + 2 + 1/6 + + + + 3 + 1/6 + + + + 4 + 1/6 + + + + 5 + 1/6 + + + + 6 + 1/6 + + +
+ + + Distribution for <m>R^2</m> + + + + k + \Pr(R^2 = k) + + + + 1^2 + 1/6 + + + + 2^2 + 1/6 + + + + 3^2 + 1/6 + + + + 4^2 + 1/6 + + + + 5^2 + 1/6 + + + + 6^2 + 1/6 + + +
+
+ +

+ Then \E(R) = \frac{7}{2} and \E(R^2) = \frac{91}{6}, so: + + \Var(R) = \E(R^2) - \left(\E(R)\right)^2 = \frac{91}{6} - \frac{49}{4} = \frac{35}{12}. + +

+
+
+ + + +

+ Let X \in [0, 1] with pdf f(x) = 2x. + + \E(X) \amp = \int_0^1 x \cdot 2x\ dx + \amp = \int_0^1 2x^2\ dx + \amp = \frac{2x^3}{3}\bigg|_0^1 + \amp = \frac{2}{3} - 0 + \amp = \frac{2}{3} + \E(X^2) \amp = \int_0^1 x^2\cdot 2x\ dx + \amp = \int_0^1 2x^3\ dx + \amp = \frac{2x^4}{4}\bigg|_0^1 + \amp = \frac{1}{2} - 0 + \amp = \frac{1}{2} + \Rightarrow \quad \Var(X) \amp = \E(X^2) - \left(\E(X)\right)^2 + \amp = \frac{1}{2} - \frac{4}{9} + \amp = \frac{1}{18}. + +

+
+
+ +

+ WARNING!!! In general, variance is not linear! That is, + + \Var(X + Y) \amp \neq \Var(X) + \Var(Y) + \Var(kX) \amp \neq k \Var(X) + + But there are still relevant properties we can state here: +

+ + + +

+ Let X be a random variable and k\in \R. + Then: + + \Var(X + k) \amp = \Var(X) \amp \Var(kX) \amp = k^2\Var(X). + + Moreover, if Y is another random variable and X, Y are independent, then: + + \Var(X + Y) = \Var(X) + \Var(Y). + +

+
+
+ + + +

+ Let H_1, H_2, \dotsc, H_n be indicators with parameter p. + Let S = H_1 + \dotsb + H_n, so S \sim \Bin(n, p). + We know \Var(H_i) = p(1-p), and H_1, \dotsc, H_n are independent. + So: + + \Var(S) \amp = \Var(H_1 + \dotsb + H_n) + \amp = \Var(H_1) + \dotsb + \Var(H_n) + \amp = p(1-p) + \dotsb + p(1-p) + \amp = np(1-p) + +

+
+
\ No newline at end of file diff --git a/source/notes/2-5.ptx b/source/notes/2-5.ptx new file mode 100644 index 0000000..b0c99b4 --- /dev/null +++ b/source/notes/2-5.ptx @@ -0,0 +1,207 @@ + + +
+ Thursday, Feb 5 + + +

+ This is an outline of the topics we covered in class. + These notes are not a substitute for your own note-taking. + I highly recommend that you take your own notes during class. + If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes. +

+
+ + + + Summary + +

+ Here are the expected value and variance formulas for common distributions. + Some of these, we've shown justification for. + Others requires techniques beyond the scope of the class to justify. +

+ + + Expected Value and Variance Formulas + + + + Distribution + Parameters + Expected Value + Variance + + + + Indicator + p + p + p(1-p) + + + + Binomial + n, p + np + np(1-p) + + + + Geometric + p + \frac{1}{p} + \frac{1-p}{p^2} + + + + Poisson + \lambda + \lambda + \lambda + + + + Exponential + \lambda + \frac{1}{\lambda} + \frac{1}{\lambda^2} + + +
+
+ + + + Covariance + + + +

+ If X, Y are random variables with expected values of \mu_X, \mu_Y, then the covariance of X and Y is: + + \Cov(X, Y) \amp = \E\left[ (X - \mu_X)(Y - \mu_Y)\right]. + +

+
+
+ +

+ Observe that: + + \Cov(X, X) \amp = \E\left[(X - \mu_X)(X - \mu_X)\right] + \amp = \E\left[(X - \mu_X)^2\right] + \amp = \Var(X), + + so covariance generalizes the variance formula to two variables. + As with variance, there's an alternative formula more suited to doing computations: + + \Var(X) \amp = \E\left[(X - \mu_X)(X - \mu_X)\right] = \E(X^2) - \mu_X^2 + \Cov(X, Y) \amp = \E\left[(X - \mu_X)(Y - \mu_Y)\right] = \E(XY) - \mu_X\mu_Y + +

+ + + +

+ Consider the joint distribution table: +

+ + + Joint Distribution + + + + + X = 0 + X = 1 + + + + Y = 0 + 0.2 + 0.1 + + + + Y = 1 + 0.05 + 0.65 + + +
+ +

+ From the table, we can calculate the marginal distributions: + + \Pr(X = 0) \amp = 0.25 \amp \Pr(Y = 0) \amp = 0.3 + \Pr(X = 1) \amp = 0.75 \amp \Pr(Y = 1) \amp = 0.7 + + So \E(X) = \mu_X = 0.75 and \E(Y) = \mu_Y = 0.7. + Then: + + \E(XY) \amp = (0)(0)(0.2) + (1)(0)(0.1) + (0)(1)(0.05) + (1)(1)(0.65) + \amp = 0.65 + \Cov(X, Y) \amp = \E(XY) - \mu_X \mu_Y + \amp = 0.65 - (0.75)(0.7) + \amp = 0.125. + +

+
+
+ +

+ Question: the formula \Var(X) = \E\left[(X - \mu_X)^2\right] makes it clear that variance cannot be negative, since squares are nonnegative. + What about \Cov(X, Y)? +

+ + + +

+ In the previous example, since X, Y were both indicator random variables, the variances for each were simply equal to the sum of the second row/column. + Similarly, \E(XY) was equal to the X = 1, Y = 1 entry in the table. + Using two indicator random variables, significantly simplifies the covariance calculation, so we can vary the table and recalculate covariance quickly. + Consider the following joint distribution: +

+ + + Joint Distribution + + + + + X = 0 + X = 1 + + + + Y = 0 + 0.1 + 0.4 + + + + Y = 1 + 0.3 + 0.2 + + +
+ +

+ Then: + + \Cov(X, Y) = 0.2 - (0.6)(0.5) = -0.1 \lt 0. + +

+
+
+ +

+ Question: how do we interpret \Cov(X, Y)? X - \mu_X is positive when X \gt \mu_X and negative when X \lt \mu_X. + Y - \mu_Y is positive when Y \gt \mu_Y and negative when Y \lt \mu_Y. + So the product (X - \mu_X)(Y - \mu_Y) is positive when X, Y are both larger or both smaller than their expected values, and negative when one is larger and one is smaller. + That is, covariance tries to quantify the tendency of X, Y to get big/small at the same time. +

+
+
\ No newline at end of file