Notes up through 2-10

This commit is contained in:
2026-02-12 07:09:15 -05:00
parent addfb172b3
commit e5d3538af5
5 changed files with 704 additions and 0 deletions
+1
View File
@@ -24,6 +24,7 @@
\DeclareMathOperator{\E}{E} \DeclareMathOperator{\E}{E}
\DeclareMathOperator{\Var}{Var} \DeclareMathOperator{\Var}{Var}
\DeclareMathOperator{\Cov}{Cov} \DeclareMathOperator{\Cov}{Cov}
\newcommand{\L}{\mathcal{L}}
</macros> </macros>
+2
View File
@@ -29,6 +29,8 @@
<xi:include href="./notes/1-27.ptx" /> <xi:include href="./notes/1-27.ptx" />
<xi:include href="./notes/1-29.ptx" /> <xi:include href="./notes/1-29.ptx" />
<xi:include href="./notes/2-3.ptx" /> <xi:include href="./notes/2-3.ptx" />
<xi:include href="./notes/2-5.ptx" />
<xi:include href="./notes/2-10.ptx" />
</chapter> </chapter>
<chapter xml:id="quizzes"> <chapter xml:id="quizzes">
+259
View File
@@ -0,0 +1,259 @@
<?xml version="1.0" encoding="UTF-8"?>
<section xml:id="notes-02-10">
<title>Tuesday, Feb 10</title>
<introduction>
<p>
This is an outline of the topics we covered in class.
These notes are <em>not</em> a substitute for your own note-taking.
I highly recommend that you take your own notes during class.
If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
</p>
</introduction>
<subsection>
<title>Likelihood</title>
<p>
In probability theory, we start with some probability model and its parameter values, and we try to find the probabilities of seeing certain types of data.
In statistics, we start with the collected data, and we try to find the most likely values of parameters for some underlying probability model.
</p>
<definition>
<statement>
<p>
An <term>estimator</term> is a way of estimating a parameter value based on data collected.
</p>
</statement>
</definition>
<example>
<statement>
<p>
Suppose we flip a coin 100 times and see 52 heads.
Let <m>p</m> be the bias of the coin.
Then we might estimate:
<md>
<mrow> \widehat{p} = \frac{52}{100} </mrow>
</md>
(The notation <m>\widehat{p}</m> is sometimes used to indicate an estimator for <m>p</m>.)
</p>
</statement>
</example>
<definition>
<statement>
<p>
If we see <m>k</m> heads in <m>n</m> flips, then the estimator <m>\widehat{p} = \frac{k}{n}</m> is called the <term>common sense estimator</term> for the binomial distribution parameter <m>p</m>.
</p>
</statement>
</definition>
<definition>
<statement>
<p>
An estimator <m>\widehat{p}</m> is <term>unbiased</term> if <m>\E(\widehat{p}) = p</m>.
</p>
</statement>
</definition>
<example>
<statement>
<p>
Flip a coin 100 times, and let <m>S</m> be the number of heads.
Then <m>S \sim \Bin(100, p)</m>.
Let <m>\widehat{p} = \frac{S}{100} = \frac{S}{n}</m>.
Then:
<md>
<mrow> \E(\widehat{p}) \amp = \E\left(\frac{S}{n}\right) </mrow>
<mrow> \amp = \frac{1}{n} \E\left(S\right) </mrow>
<mrow> \amp = \frac{1}{n} (np) </mrow>
<mrow> \amp = p. </mrow>
</md>
So <m>\E(\widehat{p} = p</m>, i.e., the common sense estimator is unbiased.
</p>
</statement>
</example>
<definition>
<statement>
<p>
We perform an experiment and collect data.
Let <m>p</m> be an unknown parameter value.
The <term>likelihood function</term> is:
<md>
<mrow> \L(p) = \Pr(\text{data} \mid \text{paramater value is } p). </mrow>
</md>
</p>
</statement>
</definition>
<example>
<statement>
<p>
Let <m>S \sim \Bin(100, p)</m>.
Suppose we see 52 heads.
Then:
<md>
<mrow> \Pr(S = k) \amp = b(k; 100, p) = {100 \choose k}p^k (1 - p)^{100 - k} </mrow>
<mrow> \L(p) \amp = {100 \choose 52}p^{52} (1 - p)^{48} </mrow>
</md>
In the first line, the variable <m>k</m> represents the data.
In the second line, the variable <m>p</m> represents the parameter value.
Our goal, given the collected data, is to find the <term>maximum likelihood estimation</term> (MLE) for the parameter value.
</p>
<p>
<m>\L(p)</m> is a continuous function over a closed interval <m>p \in [0, 1]</m>, so we use the Closed Interval Method.
<md>
<mrow> \L'(p) \amp = {100 \choose 52}\left[ 52 p^{51} (1 - p)^{48} + p^{52}48 (1-p)^{47}(-1)\right] </mrow>
<mrow> \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 (1 - p) - 48p \right] </mrow>
<mrow> \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 - 100p \right] </mrow>
</md>
Now, we look for critical numbers in the interior of the interval:
<md>
<mrow> \L'(p) = 0 \text{ when } p = \frac{52}{100} </mrow>
</md>
Finally, we test the critical numbers and the endpoints of the interval to find the max:
</p>
<table>
<title>Check Candidate Locations for Max</title>
<tabular halign="center">
<row bottom="minor">
<cell><m>p</m></cell>
<cell><m>\L(p)</m></cell>
</row>
<row>
<cell><m>0</m></cell>
<cell><m>0</m></cell>
</row>
<row>
<cell><m>52/100</m></cell>
<cell><m>\gt 0</m></cell>
</row>
<row>
<cell><m>1</m></cell>
<cell><m>0</m></cell>
</row>
</tabular>
</table>
<p>
So <m>\widehat{p} = \frac{52}{100}</m> is the MLE.
</p>
<p>
More generally, a similar calculation will show that, with <m>k</m> heads in <m>n</m> flips, the MLE will be <m>\widehat{p} = \frac{k}{n}</m>.
</p>
</statement>
</example>
<example>
<statement>
<p>
Suppose we observe a cell, measuring the time <m>T</m> until a toxin molecule leaves the cell.
Then <m>T \sim \Exp(\lambda)</m> for some <m>\lambda</m>, with pdf
<md>
<mrow> f(t) =\lambda e^{-\lambda t}, \lambda \gt 0 </mrow>
</md>
If we see a toxin molecule leave at 0.3 min, what's the MLE for <m>\lambda</m>?
<md>
<mrow> \L(\lambda) = \lambda e^{-0.3 \lambda} \quad \text{(density, not probability)} </mrow>
</md>
We want to maximize <m>\L(\lambda)</m> over <m>\lambda \in (0, \infty)</m>, an open interval.
So we'll use the "Open Interval Method".
<md>
<mrow> \L'(\lambda) \amp = e^{-0.3\lambda} + \lambda e^{-0.3\lambda}(-0.3) </mrow>
<mrow> \amp = e^{-0.3\lambda}\left( 1 - 0.3 e^{-0.3\lambda}\right) </mrow>
</md>
Then <m>\L'(\lambda) = 0</m> when <m>\lambda = \frac{1}{0.3}\approx 3.33</m>.
Checking <m>\lambda</m> values to the left and the right:
<md>
<mrow> \L'(1) \amp = (+)(+) = (+) </mrow>
<mrow> \L'(10) \amp = (+)(-) = (-) </mrow>
</md>
So <m>\L'(\lambda) \gt 0</m> (and therefore <m>\L(\lambda)</m> is increasing) on <m>(0, 1/0.3)</m>, and <m>\L'(\lambda) \lt 0</m> (and therefore <m>\L(\lambda)</m> is decreasing) on <m>(1/0.3, \infty)</m>.
Now we can conclude that <m>\widehat{\lambda} = 1/0.3 \approx 3.33</m> is the location of a global (and not just local) maximum value.
</p>
<p>
More generally, if the observed time is <m>t</m>, then the MLE will be <m>\widehat{\lambda} = \frac{1}{t}</m>.
</p>
<p>
What if we had more data points? For example, suppose two toxin molecules leave the cell at <m>t_1 = 0.3</m> min and <m>t_2 = 0.5</m> min?
</p>
<table>
<title>Waiting Times</title>
<tabular halign="center">
<row bottom="minor">
<cell>Molecule</cell>
<cell>Time</cell>
<cell>Rate Estimate</cell>
</row>
<row>
<cell><m>1</m></cell>
<cell><m>0.3</m></cell>
<cell><m>1/0.3 \approx 3.33</m></cell>
</row>
<row>
<cell><m>2</m></cell>
<cell><m>0.5</m></cell>
<cell><m>1/0.5 = 2</m></cell>
</row>
</tabular>
</table>
<p>
How do we combine these data points? We could take the average of the rate estimates:
<md>
<mrow> \frac{3.33 + 2}{2} \approx 2.67 </mrow>
</md>
Alternatively, we could average the times first, then create a new rate estimate from the average time:
<md>
<mrow> \frac{0.3 + 0.5}{2} \amp = 0.4 </mrow>
<mrow> \frac{1}{0.4} \amp = 2.5 </mrow>
</md>
Both of these make some sense, but let's do a careful computation to be certain which way is correct (if either of them is!).
<md>
<mrow> \L(\lambda) \amp = \left( \lambda e^{-0.3 \lambda} \right)\left( \lambda e^{-0.5 \lambda} \right) </mrow>
<mrow> \amp = \lambda^2 e^{-0.3 \lambda - 0.5 \lambda} </mrow>
<mrow> \amp = \lambda^2 e^{- 0.8 \lambda} </mrow>
</md>
Using the Open Interval Method:
<md>
<mrow> \L'(\lambda) \amp = 2\lambda e^{-0.8 \lambda} + \lambda^2 e^{-0.8 \lambda} (-0.8) </mrow>
<mrow> \amp = \lambda e^{-0.8 \lambda} \left[ 2 - 0.8\lambda \right] </mrow>
</md>
So <m>\L'(\lambda) = 0</m> when <m>\lambda = \frac{2}{0.8} = 2.5</m>.
Testing points to the left and right:
<md>
<mrow> \L'(1) \amp = (+)(+)(+) = (+) </mrow>
<mrow> \L'(5) \amp = (+)(+)(-) = (-) </mrow>
</md>
Now <m>\L'(\lambda) \gt 0</m> (<m>\L(\lambda)</m> is increasing) on <m>(0, 2.5)</m>, and <m>\L'(\lambda) \lt 0</m> (<m>\L(\lambda)</m> is decreasing) on <m>(2.5, \infty)</m>.
Therefore, the MLE is <m>\widehat{\lambda} = \frac{2}{0.8} = 2.5</m>.
</p>
<p>
Tracing the values <m>2</m> and <m>0.8</m> throughout the calculation, we can see that the value <m>2</m> will generally match the number of waiting times collected, and the value <m>0.8</m> will be the sum of the waiting times.
So, generally, with collected waiting times of <m>t_1, \dotsc, t_n</m>, the MLE will be:
<md>
<mrow> \widehat{\lambda} = \frac{n}{t_1 + \dotsb + t_n} = \frac{1}{\text{avg time}}. </mrow>
</md>
</p>
</statement>
</example>
</subsection>
</section>
+235
View File
@@ -37,5 +37,240 @@
</p> </p>
</statement> </statement>
</definition> </definition>
<p>
We can write an alternative formula here:
<md>
<mrow> \Var(X) \amp = \E\left[ (X - \mu)^2 \right] </mrow>
<mrow> \amp = \E\left[ X^2 - 2\mu X + \mu^2 \right] </mrow>
<mrow> \amp = \E(X^2) - 2\mu \E(X) + \E(\mu^2) </mrow>
<mrow> \amp = \E(X^2) - 2\mu^2 + \mu^2 </mrow>
<mrow> \amp = \E(X^2) - \mu^2. </mrow>
</md>
This formula is generally more useful for performing computations.
</p>
<example>
<statement>
<p>
Let <m>X</m> indicate event <m>A</m> with <m>\Pr(A) = p</m>.
</p>
<sidebyside widths="30% 30%">
<table>
<title>Distribution for <m>X</m></title>
<tabular halign="center">
<row bottom="minor">
<cell><m>k</m></cell>
<cell><m>\Pr(X = k)</m></cell>
</row>
<row>
<cell><m>0</m></cell>
<cell><m>1 - p</m></cell>
</row>
<row>
<cell><m>1</m></cell>
<cell><m>p</m></cell>
</row>
</tabular>
</table>
<table>
<title>Distribution for <m>X^2</m></title>
<tabular halign="center">
<row bottom="minor">
<cell><m>k</m></cell>
<cell><m>\Pr(X^2 = k)</m></cell>
</row>
<row>
<cell><m>0^2</m></cell>
<cell><m>1 - p</m></cell>
</row>
<row>
<cell><m>1^2</m></cell>
<cell><m>p</m></cell>
</row>
</tabular>
</table>
</sidebyside>
<p>
Then <m>\E(X) = p</m> and <m>\E(X^2) = p</m>, so:
<md>
<mrow> \Var(X) = \E(X^2) - \left(\E(X)\right)^2 = p - p^2 = p(1 - p). </mrow>
</md>
</p>
</statement>
</example>
<example>
<statement>
<p>
Let <m>R</m> be the roll of a fair D6.
</p>
<sidebyside widths="30% 30%">
<table>
<title>Distribution for <m>R</m></title>
<tabular halign="center">
<row bottom="minor">
<cell><m>k</m></cell>
<cell><m>\Pr(R = k)</m></cell>
</row>
<row>
<cell><m>1</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>2</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>3</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>4</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>5</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>6</m></cell>
<cell><m>1/6</m></cell>
</row>
</tabular>
</table>
<table>
<title>Distribution for <m>R^2</m></title>
<tabular halign="center">
<row bottom="minor">
<cell><m>k</m></cell>
<cell><m>\Pr(R^2 = k)</m></cell>
</row>
<row>
<cell><m>1^2</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>2^2</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>3^2</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>4^2</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>5^2</m></cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell><m>6^2</m></cell>
<cell><m>1/6</m></cell>
</row>
</tabular>
</table>
</sidebyside>
<p>
Then <m>\E(R) = \frac{7}{2}</m> and <m>\E(R^2) = \frac{91}{6}</m>, so:
<md>
<mrow> \Var(R) = \E(R^2) - \left(\E(R)\right)^2 = \frac{91}{6} - \frac{49}{4} = \frac{35}{12}. </mrow>
</md>
</p>
</statement>
</example>
<example>
<statement>
<p>
Let <m>X \in [0, 1]</m> with pdf <m>f(x) = 2x</m>.
<md>
<mrow> \E(X) \amp = \int_0^1 x \cdot 2x\ dx </mrow>
<mrow> \amp = \int_0^1 2x^2\ dx </mrow>
<mrow> \amp = \frac{2x^3}{3}\bigg|_0^1 </mrow>
<mrow> \amp = \frac{2}{3} - 0 </mrow>
<mrow> \amp = \frac{2}{3} </mrow>
<mrow> \E(X^2) \amp = \int_0^1 x^2\cdot 2x\ dx </mrow>
<mrow> \amp = \int_0^1 2x^3\ dx </mrow>
<mrow> \amp = \frac{2x^4}{4}\bigg|_0^1 </mrow>
<mrow> \amp = \frac{1}{2} - 0 </mrow>
<mrow> \amp = \frac{1}{2} </mrow>
<mrow> \Rightarrow \quad \Var(X) \amp = \E(X^2) - \left(\E(X)\right)^2 </mrow>
<mrow> \amp = \frac{1}{2} - \frac{4}{9} </mrow>
<mrow> \amp = \frac{1}{18}. </mrow>
</md>
</p>
</statement>
</example>
<p>
WARNING!!! In general, variance is <em>not</em> linear! That is,
<md>
<mrow> \Var(X + Y) \amp \neq \Var(X) + \Var(Y) </mrow>
<mrow>\Var(kX) \amp \neq k \Var(X)</mrow>
</md>
But there are still relevant properties we can state here:
</p>
<theorem>
<statement>
<p>
Let <m>X</m> be a random variable and <m>k\in \R</m>.
Then:
<md>
<mrow> \Var(X + k) \amp = \Var(X) \amp \Var(kX) \amp = k^2\Var(X). </mrow>
</md>
Moreover, if <m>Y</m> is another random variable and <m>X, Y</m> are independent, then:
<md>
<mrow> \Var(X + Y) = \Var(X) + \Var(Y). </mrow>
</md>
</p>
</statement>
</theorem>
<example>
<statement>
<p>
Let <m>H_1, H_2, \dotsc, H_n</m> be indicators with parameter <m>p</m>.
Let <m>S = H_1 + \dotsb + H_n</m>, so <m>S \sim \Bin(n, p)</m>.
We know <m>\Var(H_i) = p(1-p)</m>, and <m>H_1, \dotsc, H_n</m> are independent.
So:
<md>
<mrow> \Var(S) \amp = \Var(H_1 + \dotsb + H_n) </mrow>
<mrow> \amp = \Var(H_1) + \dotsb + \Var(H_n) </mrow>
<mrow> \amp = p(1-p) + \dotsb + p(1-p) </mrow>
<mrow> \amp = np(1-p) </mrow>
</md>
</p>
</statement>
</example>
</subsection> </subsection>
</section> </section>
+207
View File
@@ -0,0 +1,207 @@
<?xml version="1.0" encoding="UTF-8"?>
<section xml:id="notes-02-05">
<title>Thursday, Feb 5</title>
<introduction>
<p>
This is an outline of the topics we covered in class.
These notes are <em>not</em> a substitute for your own note-taking.
I highly recommend that you take your own notes during class.
If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
</p>
</introduction>
<subsection>
<title>Summary</title>
<p>
Here are the expected value and variance formulas for common distributions.
Some of these, we've shown justification for.
Others requires techniques beyond the scope of the class to justify.
</p>
<table>
<title>Expected Value and Variance Formulas</title>
<tabular halign="center">
<row bottom="minor">
<cell>Distribution</cell>
<cell>Parameters</cell>
<cell>Expected Value</cell>
<cell>Variance</cell>
</row>
<row>
<cell>Indicator</cell>
<cell><m>p</m></cell>
<cell><m>p</m></cell>
<cell><m>p(1-p)</m></cell>
</row>
<row>
<cell>Binomial</cell>
<cell><m>n, p</m></cell>
<cell><m>np</m></cell>
<cell><m>np(1-p)</m></cell>
</row>
<row>
<cell>Geometric</cell>
<cell><m>p</m></cell>
<cell><m>\frac{1}{p}</m></cell>
<cell><m>\frac{1-p}{p^2}</m></cell>
</row>
<row>
<cell>Poisson</cell>
<cell><m>\lambda</m></cell>
<cell><m>\lambda</m></cell>
<cell><m>\lambda</m></cell>
</row>
<row>
<cell>Exponential</cell>
<cell><m>\lambda</m></cell>
<cell><m>\frac{1}{\lambda}</m></cell>
<cell><m>\frac{1}{\lambda^2}</m></cell>
</row>
</tabular>
</table>
</subsection>
<subsection>
<title>Covariance</title>
<definition>
<statement>
<p>
If <m>X, Y</m> are random variables with expected values of <m>\mu_X, \mu_Y</m>, then the <term>covariance</term> of <m>X</m> and <m>Y</m> is:
<md>
<mrow> \Cov(X, Y) \amp = \E\left[ (X - \mu_X)(Y - \mu_Y)\right]. </mrow>
</md>
</p>
</statement>
</definition>
<p>
Observe that:
<md>
<mrow> \Cov(X, X) \amp = \E\left[(X - \mu_X)(X - \mu_X)\right] </mrow>
<mrow> \amp = \E\left[(X - \mu_X)^2\right] </mrow>
<mrow> \amp = \Var(X), </mrow>
</md>
so covariance generalizes the variance formula to two variables.
As with variance, there's an alternative formula more suited to doing computations:
<md>
<mrow> \Var(X) \amp = \E\left[(X - \mu_X)(X - \mu_X)\right] = \E(X^2) - \mu_X^2 </mrow>
<mrow> \Cov(X, Y) \amp = \E\left[(X - \mu_X)(Y - \mu_Y)\right] = \E(XY) - \mu_X\mu_Y </mrow>
</md>
</p>
<example>
<statement>
<p>
Consider the joint distribution table:
</p>
<table>
<title>Joint Distribution</title>
<tabular halign="center">
<row bottom="minor">
<cell right="minor"></cell>
<cell><m>X = 0</m></cell>
<cell><m>X = 1</m></cell>
</row>
<row>
<cell right="minor"><m>Y = 0</m></cell>
<cell>0.2</cell>
<cell>0.1</cell>
</row>
<row>
<cell right="minor"><m>Y = 1</m></cell>
<cell>0.05</cell>
<cell>0.65</cell>
</row>
</tabular>
</table>
<p>
From the table, we can calculate the marginal distributions:
<md>
<mrow> \Pr(X = 0) \amp = 0.25 \amp \Pr(Y = 0) \amp = 0.3 </mrow>
<mrow> \Pr(X = 1) \amp = 0.75 \amp \Pr(Y = 1) \amp = 0.7 </mrow>
</md>
So <m>\E(X) = \mu_X = 0.75</m> and <m>\E(Y) = \mu_Y = 0.7</m>.
Then:
<md>
<mrow> \E(XY) \amp = (0)(0)(0.2) + (1)(0)(0.1) + (0)(1)(0.05) + (1)(1)(0.65) </mrow>
<mrow> \amp = 0.65 </mrow>
<mrow> \Cov(X, Y) \amp = \E(XY) - \mu_X \mu_Y </mrow>
<mrow> \amp = 0.65 - (0.75)(0.7) </mrow>
<mrow> \amp = 0.125. </mrow>
</md>
</p>
</statement>
</example>
<p>
Question: the formula <m>\Var(X) = \E\left[(X - \mu_X)^2\right]</m> makes it clear that variance cannot be negative, since squares are nonnegative.
What about <m>\Cov(X, Y)</m>?
</p>
<example>
<statement>
<p>
In the previous example, since <m>X, Y</m> were both indicator random variables, the variances for each were simply equal to the sum of the second row/column.
Similarly, <m>\E(XY)</m> was equal to the <m>X = 1, Y = 1</m> entry in the table.
Using two indicator random variables, significantly simplifies the covariance calculation, so we can vary the table and recalculate covariance quickly.
Consider the following joint distribution:
</p>
<table>
<title>Joint Distribution</title>
<tabular halign="center">
<row bottom="minor">
<cell right="minor"></cell>
<cell><m>X = 0</m></cell>
<cell><m>X = 1</m></cell>
</row>
<row>
<cell right="minor"><m>Y = 0</m></cell>
<cell>0.1</cell>
<cell>0.4</cell>
</row>
<row>
<cell right="minor"><m>Y = 1</m></cell>
<cell>0.3</cell>
<cell>0.2</cell>
</row>
</tabular>
</table>
<p>
Then:
<md>
<mrow> \Cov(X, Y) = 0.2 - (0.6)(0.5) = -0.1 \lt 0. </mrow>
</md>
</p>
</statement>
</example>
<p>
Question: how do we interpret <m>\Cov(X, Y)</m>? <m>X - \mu_X</m> is positive when <m>X \gt \mu_X</m> and negative when <m>X \lt \mu_X</m>.
<m>Y - \mu_Y</m> is positive when <m>Y \gt \mu_Y</m> and negative when <m>Y \lt \mu_Y</m>.
So the product <m>(X - \mu_X)(Y - \mu_Y)</m> is positive when <m>X, Y</m> are both larger or both smaller than their expected values, and negative when one is larger and one is smaller.
That is, covariance tries to quantify the tendency of <m>X, Y</m> to get big/small at the same time.
</p>
</subsection>
</section>