Skip to main content

Section Tuesday, Feb 10

This is an outline of the topics we covered in class. These notes are not a substitute for your own note-taking. I highly recommend that you take your own notes during class. If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.

Subsection Likelihood

In probability theory, we start with some probability model and its parameter values, and we try to find the probabilities of seeing certain types of data. In statistics, we start with the collected data, and we try to find the most likely values of parameters for some underlying probability model.

Definition 106.

An estimator is a way of estimating a parameter value based on data collected.

Example 107.

Suppose we flip a coin 100 times and see 52 heads. Let \(p\) be the bias of the coin. Then we might estimate:
\begin{gather*} \widehat{p} = \frac{52}{100} \end{gather*}
(The notation \(\widehat{p}\) is sometimes used to indicate an estimator for \(p\text{.}\))

Definition 108.

If we see \(k\) heads in \(n\) flips, then the estimator \(\widehat{p} = \frac{k}{n}\) is called the common sense estimator for the binomial distribution parameter \(p\text{.}\)

Definition 109.

An estimator \(\widehat{p}\) is unbiased if \(\E(\widehat{p}) = p\text{.}\)

Example 110.

Flip a coin 100 times, and let \(S\) be the number of heads. Then \(S \sim \Bin(100, p)\text{.}\) Let \(\widehat{p} = \frac{S}{100} = \frac{S}{n}\text{.}\) Then:
\begin{align*} \E(\widehat{p}) \amp = \E\left(\frac{S}{n}\right) \\ \amp = \frac{1}{n} \E\left(S\right) \\ \amp = \frac{1}{n} (np) \\ \amp = p. \end{align*}
So \(\E(\widehat{p} = p\text{,}\) i.e., the common sense estimator is unbiased.

Definition 111.

We perform an experiment and collect data. Let \(p\) be an unknown parameter value. The likelihood function is:
\begin{gather*} \L(p) = \Pr(\text{data} \mid \text{paramater value is } p). \end{gather*}

Example 112.

Let \(S \sim \Bin(100, p)\text{.}\) Suppose we see 52 heads. Then:
\begin{align*} \Pr(S = k) \amp = b(k; 100, p) = {100 \choose k}p^k (1 - p)^{100 - k} \\ \L(p) \amp = {100 \choose 52}p^{52} (1 - p)^{48} \end{align*}
In the first line, the variable \(k\) represents the data. In the second line, the variable \(p\) represents the parameter value. Our goal, given the collected data, is to find the maximum likelihood estimation (MLE) for the parameter value.
\(\L(p)\) is a continuous function over a closed interval \(p \in [0, 1]\text{,}\) so we use the Closed Interval Method.
\begin{align*} \L'(p) \amp = {100 \choose 52}\left[ 52 p^{51} (1 - p)^{48} + p^{52}48 (1-p)^{47}(-1)\right] \\ \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 (1 - p) - 48p \right] \\ \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 - 100p \right] \end{align*}
Now, we look for critical numbers in the interior of the interval:
\begin{gather*} \L'(p) = 0 \text{ when } p = \frac{52}{100} \end{gather*}
Finally, we test the critical numbers and the endpoints of the interval to find the max:
Table 113. Check Candidate Locations for Max
\(p\) \(\L(p)\)
\(0\) \(0\)
\(52/100\) \(\gt 0\)
\(1\) \(0\)
So \(\widehat{p} = \frac{52}{100}\) is the MLE.
More generally, a similar calculation will show that, with \(k\) heads in \(n\) flips, the MLE will be \(\widehat{p} = \frac{k}{n}\text{.}\)

Example 114.

Suppose we observe a cell, measuring the time \(T\) until a toxin molecule leaves the cell. Then \(T \sim \Exp(\lambda)\) for some \(\lambda\text{,}\) with pdf
\begin{gather*} f(t) =\lambda e^{-\lambda t}, \lambda \gt 0 \end{gather*}
If we see a toxin molecule leave at 0.3 min, what’s the MLE for \(\lambda\text{?}\)
\begin{gather*} \L(\lambda) = \lambda e^{-0.3 \lambda} \quad \text{(density, not probability)} \end{gather*}
We want to maximize \(\L(\lambda)\) over \(\lambda \in (0, \infty)\text{,}\) an open interval. So we’ll use the "Open Interval Method".
\begin{align*} \L'(\lambda) \amp = e^{-0.3\lambda} + \lambda e^{-0.3\lambda}(-0.3) \\ \amp = e^{-0.3\lambda}\left( 1 - 0.3 e^{-0.3\lambda}\right) \end{align*}
Then \(\L'(\lambda) = 0\) when \(\lambda = \frac{1}{0.3}\approx 3.33\text{.}\) Checking \(\lambda\) values to the left and the right:
\begin{align*} \L'(1) \amp = (+)(+) = (+) \\ \L'(10) \amp = (+)(-) = (-) \end{align*}
So \(\L'(\lambda) \gt 0\) (and therefore \(\L(\lambda)\) is increasing) on \((0, 1/0.3)\text{,}\) and \(\L'(\lambda) \lt 0\) (and therefore \(\L(\lambda)\) is decreasing) on \((1/0.3, \infty)\text{.}\) Now we can conclude that \(\widehat{\lambda} = 1/0.3 \approx 3.33\) is the location of a global (and not just local) maximum value.
More generally, if the observed time is \(t\text{,}\) then the MLE will be \(\widehat{\lambda} = \frac{1}{t}\text{.}\)
What if we had more data points? For example, suppose two toxin molecules leave the cell at \(t_1 = 0.3\) min and \(t_2 = 0.5\) min?
Table 115. Waiting Times
Molecule Time Rate Estimate
\(1\) \(0.3\) \(1/0.3 \approx 3.33\)
\(2\) \(0.5\) \(1/0.5 = 2\)
How do we combine these data points? We could take the average of the rate estimates:
\begin{gather*} \frac{3.33 + 2}{2} \approx 2.67 \end{gather*}
Alternatively, we could average the times first, then create a new rate estimate from the average time:
\begin{align*} \frac{0.3 + 0.5}{2} \amp = 0.4 \\ \frac{1}{0.4} \amp = 2.5 \end{align*}
Both of these make some sense, but let’s do a careful computation to be certain which way is correct (if either of them is!).
\begin{align*} \L(\lambda) \amp = \left( \lambda e^{-0.3 \lambda} \right)\left( \lambda e^{-0.5 \lambda} \right) \\ \amp = \lambda^2 e^{-0.3 \lambda - 0.5 \lambda} \\ \amp = \lambda^2 e^{- 0.8 \lambda} \end{align*}
Using the Open Interval Method:
\begin{align*} \L'(\lambda) \amp = 2\lambda e^{-0.8 \lambda} + \lambda^2 e^{-0.8 \lambda} (-0.8) \\ \amp = \lambda e^{-0.8 \lambda} \left[ 2 - 0.8\lambda \right] \end{align*}
So \(\L'(\lambda) = 0\) when \(\lambda = \frac{2}{0.8} = 2.5\text{.}\) Testing points to the left and right:
\begin{align*} \L'(1) \amp = (+)(+)(+) = (+) \\ \L'(5) \amp = (+)(+)(-) = (-) \end{align*}
Now \(\L'(\lambda) \gt 0\) (\(\L(\lambda)\) is increasing) on \((0, 2.5)\text{,}\) and \(\L'(\lambda) \lt 0\) (\(\L(\lambda)\) is decreasing) on \((2.5, \infty)\text{.}\) Therefore, the MLE is \(\widehat{\lambda} = \frac{2}{0.8} = 2.5\text{.}\)
Tracing the values \(2\) and \(0.8\) throughout the calculation, we can see that the value \(2\) will generally match the number of waiting times collected, and the value \(0.8\) will be the sum of the waiting times. So, generally, with collected waiting times of \(t_1, \dotsc, t_n\text{,}\) the MLE will be:
\begin{gather*} \widehat{\lambda} = \frac{n}{t_1 + \dotsb + t_n} = \frac{1}{\text{avg time}}. \end{gather*}