From e5d3538af5776b19f1b615c4cbf1160b0dca2cea Mon Sep 17 00:00:00 2001
From: andyeisenberg
Date: Thu, 12 Feb 2026 07:09:15 -0500
Subject: [PATCH] Notes up through 2-10
---
source/docinfo.ptx | 1 +
source/main.ptx | 2 +
source/notes/2-10.ptx | 259 ++++++++++++++++++++++++++++++++++++++++++
source/notes/2-3.ptx | 235 ++++++++++++++++++++++++++++++++++++++
source/notes/2-5.ptx | 207 +++++++++++++++++++++++++++++++++
5 files changed, 704 insertions(+)
create mode 100644 source/notes/2-10.ptx
create mode 100644 source/notes/2-5.ptx
diff --git a/source/docinfo.ptx b/source/docinfo.ptx
index 95d059c..3447d98 100644
--- a/source/docinfo.ptx
+++ b/source/docinfo.ptx
@@ -24,6 +24,7 @@
\DeclareMathOperator{\E}{E}
\DeclareMathOperator{\Var}{Var}
\DeclareMathOperator{\Cov}{Cov}
+ \newcommand{\L}{\mathcal{L}}
diff --git a/source/main.ptx b/source/main.ptx
index 0ea7f5f..4c1af32 100644
--- a/source/main.ptx
+++ b/source/main.ptx
@@ -29,6 +29,8 @@
+
+
diff --git a/source/notes/2-10.ptx b/source/notes/2-10.ptx
new file mode 100644
index 0000000..c170eaf
--- /dev/null
+++ b/source/notes/2-10.ptx
@@ -0,0 +1,259 @@
+
+
+
+ Tuesday, Feb 10
+
+
+
+ This is an outline of the topics we covered in class.
+ These notes are not a substitute for your own note-taking.
+ I highly recommend that you take your own notes during class.
+ If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
+
+
+
+
+
+ Likelihood
+
+
+ In probability theory, we start with some probability model and its parameter values, and we try to find the probabilities of seeing certain types of data.
+ In statistics, we start with the collected data, and we try to find the most likely values of parameters for some underlying probability model.
+
+
+
+
+
+ An estimator is a way of estimating a parameter value based on data collected.
+
+
+
+
+
+
+
+ Suppose we flip a coin 100 times and see 52 heads.
+ Let p be the bias of the coin.
+ Then we might estimate:
+
+ \widehat{p} = \frac{52}{100}
+
+ (The notation \widehat{p} is sometimes used to indicate an estimator for p.)
+
+
+
+
+
+
+
+ If we see k heads in n flips, then the estimator \widehat{p} = \frac{k}{n} is called the common sense estimator for the binomial distribution parameter p.
+
+
+
+
+
+
+
+ An estimator \widehat{p} is unbiased if \E(\widehat{p}) = p.
+
+
+
+
+
+
+
+ Flip a coin 100 times, and let S be the number of heads.
+ Then S \sim \Bin(100, p).
+ Let \widehat{p} = \frac{S}{100} = \frac{S}{n}.
+ Then:
+
+ \E(\widehat{p}) \amp = \E\left(\frac{S}{n}\right)
+ \amp = \frac{1}{n} \E\left(S\right)
+ \amp = \frac{1}{n} (np)
+ \amp = p.
+
+ So \E(\widehat{p} = p, i.e., the common sense estimator is unbiased.
+
+
+
+
+
+
+
+ We perform an experiment and collect data.
+ Let p be an unknown parameter value.
+ The likelihood function is:
+
+ \L(p) = \Pr(\text{data} \mid \text{paramater value is } p).
+
+
+
+
+
+
+
+
+ Let S \sim \Bin(100, p).
+ Suppose we see 52 heads.
+ Then:
+
+ \Pr(S = k) \amp = b(k; 100, p) = {100 \choose k}p^k (1 - p)^{100 - k}
+ \L(p) \amp = {100 \choose 52}p^{52} (1 - p)^{48}
+
+ In the first line, the variable k represents the data.
+ In the second line, the variable p represents the parameter value.
+ Our goal, given the collected data, is to find the maximum likelihood estimation (MLE) for the parameter value.
+
+
+
+ \L(p) is a continuous function over a closed interval p \in [0, 1], so we use the Closed Interval Method.
+
+ \L'(p) \amp = {100 \choose 52}\left[ 52 p^{51} (1 - p)^{48} + p^{52}48 (1-p)^{47}(-1)\right]
+ \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 (1 - p) - 48p \right]
+ \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 - 100p \right]
+
+ Now, we look for critical numbers in the interior of the interval:
+
+ \L'(p) = 0 \text{ when } p = \frac{52}{100}
+
+ Finally, we test the critical numbers and the endpoints of the interval to find the max:
+
+
+
+ Check Candidate Locations for Max
+
+
+
+ | p |
+ \L(p) |
+
+
+
+ | 0 |
+ 0 |
+
+
+
+ | 52/100 |
+ \gt 0 |
+
+
+
+ | 1 |
+ 0 |
+
+
+
+
+
+ So \widehat{p} = \frac{52}{100} is the MLE.
+
+
+
+ More generally, a similar calculation will show that, with k heads in n flips, the MLE will be \widehat{p} = \frac{k}{n}.
+
+
+
+
+
+
+
+ Suppose we observe a cell, measuring the time T until a toxin molecule leaves the cell.
+ Then T \sim \Exp(\lambda) for some \lambda, with pdf
+
+ f(t) =\lambda e^{-\lambda t}, \lambda \gt 0
+
+ If we see a toxin molecule leave at 0.3 min, what's the MLE for \lambda?
+
+ \L(\lambda) = \lambda e^{-0.3 \lambda} \quad \text{(density, not probability)}
+
+ We want to maximize \L(\lambda) over \lambda \in (0, \infty), an open interval.
+ So we'll use the "Open Interval Method".
+
+ \L'(\lambda) \amp = e^{-0.3\lambda} + \lambda e^{-0.3\lambda}(-0.3)
+ \amp = e^{-0.3\lambda}\left( 1 - 0.3 e^{-0.3\lambda}\right)
+
+ Then \L'(\lambda) = 0 when \lambda = \frac{1}{0.3}\approx 3.33.
+ Checking \lambda values to the left and the right:
+
+ \L'(1) \amp = (+)(+) = (+)
+ \L'(10) \amp = (+)(-) = (-)
+
+ So \L'(\lambda) \gt 0 (and therefore \L(\lambda) is increasing) on (0, 1/0.3), and \L'(\lambda) \lt 0 (and therefore \L(\lambda) is decreasing) on (1/0.3, \infty).
+ Now we can conclude that \widehat{\lambda} = 1/0.3 \approx 3.33 is the location of a global (and not just local) maximum value.
+
+
+
+ More generally, if the observed time is t, then the MLE will be \widehat{\lambda} = \frac{1}{t}.
+
+
+
+ What if we had more data points? For example, suppose two toxin molecules leave the cell at t_1 = 0.3 min and t_2 = 0.5 min?
+
+
+
+ Waiting Times
+
+
+
+ | Molecule |
+ Time |
+ Rate Estimate |
+
+
+
+ | 1 |
+ 0.3 |
+ 1/0.3 \approx 3.33 |
+
+
+
+ | 2 |
+ 0.5 |
+ 1/0.5 = 2 |
+
+
+
+
+
+ How do we combine these data points? We could take the average of the rate estimates:
+
+ \frac{3.33 + 2}{2} \approx 2.67
+
+ Alternatively, we could average the times first, then create a new rate estimate from the average time:
+
+ \frac{0.3 + 0.5}{2} \amp = 0.4
+ \frac{1}{0.4} \amp = 2.5
+
+ Both of these make some sense, but let's do a careful computation to be certain which way is correct (if either of them is!).
+
+ \L(\lambda) \amp = \left( \lambda e^{-0.3 \lambda} \right)\left( \lambda e^{-0.5 \lambda} \right)
+ \amp = \lambda^2 e^{-0.3 \lambda - 0.5 \lambda}
+ \amp = \lambda^2 e^{- 0.8 \lambda}
+
+ Using the Open Interval Method:
+
+ \L'(\lambda) \amp = 2\lambda e^{-0.8 \lambda} + \lambda^2 e^{-0.8 \lambda} (-0.8)
+ \amp = \lambda e^{-0.8 \lambda} \left[ 2 - 0.8\lambda \right]
+
+ So \L'(\lambda) = 0 when \lambda = \frac{2}{0.8} = 2.5.
+ Testing points to the left and right:
+
+ \L'(1) \amp = (+)(+)(+) = (+)
+ \L'(5) \amp = (+)(+)(-) = (-)
+
+ Now \L'(\lambda) \gt 0 (\L(\lambda) is increasing) on (0, 2.5), and \L'(\lambda) \lt 0 (\L(\lambda) is decreasing) on (2.5, \infty).
+ Therefore, the MLE is \widehat{\lambda} = \frac{2}{0.8} = 2.5.
+
+
+
+ Tracing the values 2 and 0.8 throughout the calculation, we can see that the value 2 will generally match the number of waiting times collected, and the value 0.8 will be the sum of the waiting times.
+ So, generally, with collected waiting times of t_1, \dotsc, t_n, the MLE will be:
+
+ \widehat{\lambda} = \frac{n}{t_1 + \dotsb + t_n} = \frac{1}{\text{avg time}}.
+
+
+
+
+
+
\ No newline at end of file
diff --git a/source/notes/2-3.ptx b/source/notes/2-3.ptx
index 288fe9c..2f64c38 100644
--- a/source/notes/2-3.ptx
+++ b/source/notes/2-3.ptx
@@ -37,5 +37,240 @@
+
+
+ We can write an alternative formula here:
+
+ \Var(X) \amp = \E\left[ (X - \mu)^2 \right]
+ \amp = \E\left[ X^2 - 2\mu X + \mu^2 \right]
+ \amp = \E(X^2) - 2\mu \E(X) + \E(\mu^2)
+ \amp = \E(X^2) - 2\mu^2 + \mu^2
+ \amp = \E(X^2) - \mu^2.
+
+ This formula is generally more useful for performing computations.
+
+
+
+
+
+ Let X indicate event A with \Pr(A) = p.
+
+
+
+
+ Distribution for X
+
+
+
+ | k |
+ \Pr(X = k) |
+
+
+
+ | 0 |
+ 1 - p |
+
+
+
+ | 1 |
+ p |
+
+
+
+
+
+ Distribution for X^2
+
+
+
+ | k |
+ \Pr(X^2 = k) |
+
+
+
+ | 0^2 |
+ 1 - p |
+
+
+
+ | 1^2 |
+ p |
+
+
+
+
+
+
+ Then \E(X) = p and \E(X^2) = p, so:
+
+ \Var(X) = \E(X^2) - \left(\E(X)\right)^2 = p - p^2 = p(1 - p).
+
+
+
+
+
+
+
+
+ Let R be the roll of a fair D6.
+
+
+
+
+ Distribution for R
+
+
+
+ | k |
+ \Pr(R = k) |
+
+
+
+ | 1 |
+ 1/6 |
+
+
+
+ | 2 |
+ 1/6 |
+
+
+
+ | 3 |
+ 1/6 |
+
+
+
+ | 4 |
+ 1/6 |
+
+
+
+ | 5 |
+ 1/6 |
+
+
+
+ | 6 |
+ 1/6 |
+
+
+
+
+
+ Distribution for R^2
+
+
+
+ | k |
+ \Pr(R^2 = k) |
+
+
+
+ | 1^2 |
+ 1/6 |
+
+
+
+ | 2^2 |
+ 1/6 |
+
+
+
+ | 3^2 |
+ 1/6 |
+
+
+
+ | 4^2 |
+ 1/6 |
+
+
+
+ | 5^2 |
+ 1/6 |
+
+
+
+ | 6^2 |
+ 1/6 |
+
+
+
+
+
+
+ Then \E(R) = \frac{7}{2} and \E(R^2) = \frac{91}{6}, so:
+
+ \Var(R) = \E(R^2) - \left(\E(R)\right)^2 = \frac{91}{6} - \frac{49}{4} = \frac{35}{12}.
+
+
+
+
+
+
+
+
+ Let X \in [0, 1] with pdf f(x) = 2x.
+
+ \E(X) \amp = \int_0^1 x \cdot 2x\ dx
+ \amp = \int_0^1 2x^2\ dx
+ \amp = \frac{2x^3}{3}\bigg|_0^1
+ \amp = \frac{2}{3} - 0
+ \amp = \frac{2}{3}
+ \E(X^2) \amp = \int_0^1 x^2\cdot 2x\ dx
+ \amp = \int_0^1 2x^3\ dx
+ \amp = \frac{2x^4}{4}\bigg|_0^1
+ \amp = \frac{1}{2} - 0
+ \amp = \frac{1}{2}
+ \Rightarrow \quad \Var(X) \amp = \E(X^2) - \left(\E(X)\right)^2
+ \amp = \frac{1}{2} - \frac{4}{9}
+ \amp = \frac{1}{18}.
+
+
+
+
+
+
+ WARNING!!! In general, variance is not linear! That is,
+
+ \Var(X + Y) \amp \neq \Var(X) + \Var(Y)
+ \Var(kX) \amp \neq k \Var(X)
+
+ But there are still relevant properties we can state here:
+
+
+
+
+
+ Let X be a random variable and k\in \R.
+ Then:
+
+ \Var(X + k) \amp = \Var(X) \amp \Var(kX) \amp = k^2\Var(X).
+
+ Moreover, if Y is another random variable and X, Y are independent, then:
+
+ \Var(X + Y) = \Var(X) + \Var(Y).
+
+
+
+
+
+
+
+
+ Let H_1, H_2, \dotsc, H_n be indicators with parameter p.
+ Let S = H_1 + \dotsb + H_n, so S \sim \Bin(n, p).
+ We know \Var(H_i) = p(1-p), and H_1, \dotsc, H_n are independent.
+ So:
+
+ \Var(S) \amp = \Var(H_1 + \dotsb + H_n)
+ \amp = \Var(H_1) + \dotsb + \Var(H_n)
+ \amp = p(1-p) + \dotsb + p(1-p)
+ \amp = np(1-p)
+
+
+
+
\ No newline at end of file
diff --git a/source/notes/2-5.ptx b/source/notes/2-5.ptx
new file mode 100644
index 0000000..b0c99b4
--- /dev/null
+++ b/source/notes/2-5.ptx
@@ -0,0 +1,207 @@
+
+
+
+ Thursday, Feb 5
+
+
+
+ This is an outline of the topics we covered in class.
+ These notes are not a substitute for your own note-taking.
+ I highly recommend that you take your own notes during class.
+ If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
+
+
+
+
+
+ Summary
+
+
+ Here are the expected value and variance formulas for common distributions.
+ Some of these, we've shown justification for.
+ Others requires techniques beyond the scope of the class to justify.
+
+
+
+ Expected Value and Variance Formulas
+
+
+
+ | Distribution |
+ Parameters |
+ Expected Value |
+ Variance |
+
+
+
+ | Indicator |
+ p |
+ p |
+ p(1-p) |
+
+
+
+ | Binomial |
+ n, p |
+ np |
+ np(1-p) |
+
+
+
+ | Geometric |
+ p |
+ \frac{1}{p} |
+ \frac{1-p}{p^2} |
+
+
+
+ | Poisson |
+ \lambda |
+ \lambda |
+ \lambda |
+
+
+
+ | Exponential |
+ \lambda |
+ \frac{1}{\lambda} |
+ \frac{1}{\lambda^2} |
+
+
+
+
+
+
+
+ Covariance
+
+
+
+
+ If X, Y are random variables with expected values of \mu_X, \mu_Y, then the covariance of X and Y is:
+
+ \Cov(X, Y) \amp = \E\left[ (X - \mu_X)(Y - \mu_Y)\right].
+
+
+
+
+
+
+ Observe that:
+
+ \Cov(X, X) \amp = \E\left[(X - \mu_X)(X - \mu_X)\right]
+ \amp = \E\left[(X - \mu_X)^2\right]
+ \amp = \Var(X),
+
+ so covariance generalizes the variance formula to two variables.
+ As with variance, there's an alternative formula more suited to doing computations:
+
+ \Var(X) \amp = \E\left[(X - \mu_X)(X - \mu_X)\right] = \E(X^2) - \mu_X^2
+ \Cov(X, Y) \amp = \E\left[(X - \mu_X)(Y - \mu_Y)\right] = \E(XY) - \mu_X\mu_Y
+
+
+
+
+
+
+ Consider the joint distribution table:
+
+
+
+ Joint Distribution
+
+
+
+ |
+ X = 0 |
+ X = 1 |
+
+
+
+ | Y = 0 |
+ 0.2 |
+ 0.1 |
+
+
+
+ | Y = 1 |
+ 0.05 |
+ 0.65 |
+
+
+
+
+
+ From the table, we can calculate the marginal distributions:
+
+ \Pr(X = 0) \amp = 0.25 \amp \Pr(Y = 0) \amp = 0.3
+ \Pr(X = 1) \amp = 0.75 \amp \Pr(Y = 1) \amp = 0.7
+
+ So \E(X) = \mu_X = 0.75 and \E(Y) = \mu_Y = 0.7.
+ Then:
+
+ \E(XY) \amp = (0)(0)(0.2) + (1)(0)(0.1) + (0)(1)(0.05) + (1)(1)(0.65)
+ \amp = 0.65
+ \Cov(X, Y) \amp = \E(XY) - \mu_X \mu_Y
+ \amp = 0.65 - (0.75)(0.7)
+ \amp = 0.125.
+
+
+
+
+
+
+ Question: the formula \Var(X) = \E\left[(X - \mu_X)^2\right] makes it clear that variance cannot be negative, since squares are nonnegative.
+ What about \Cov(X, Y)?
+
+
+
+
+
+ In the previous example, since X, Y were both indicator random variables, the variances for each were simply equal to the sum of the second row/column.
+ Similarly, \E(XY) was equal to the X = 1, Y = 1 entry in the table.
+ Using two indicator random variables, significantly simplifies the covariance calculation, so we can vary the table and recalculate covariance quickly.
+ Consider the following joint distribution:
+
+
+
+ Joint Distribution
+
+
+
+ |
+ X = 0 |
+ X = 1 |
+
+
+
+ | Y = 0 |
+ 0.1 |
+ 0.4 |
+
+
+
+ | Y = 1 |
+ 0.3 |
+ 0.2 |
+
+
+
+
+
+ Then:
+
+ \Cov(X, Y) = 0.2 - (0.6)(0.5) = -0.1 \lt 0.
+
+
+
+
+
+
+ Question: how do we interpret \Cov(X, Y)? X - \mu_X is positive when X \gt \mu_X and negative when X \lt \mu_X.
+ Y - \mu_Y is positive when Y \gt \mu_Y and negative when Y \lt \mu_Y.
+ So the product (X - \mu_X)(Y - \mu_Y) is positive when X, Y are both larger or both smaller than their expected values, and negative when one is larger and one is smaller.
+ That is, covariance tries to quantify the tendency of X, Y to get big/small at the same time.
+
+
+
\ No newline at end of file