Tuesday, Feb 10

This is an outline of the topics we covered in class. These notes are not a substitute for your own note-taking. I highly recommend that you take your own notes during class. If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.

Likelihood

In probability theory, we start with some probability model and its parameter values, and we try to find the probabilities of seeing certain types of data. In statistics, we start with the collected data, and we try to find the most likely values of parameters for some underlying probability model.

An estimator is a way of estimating a parameter value based on data collected.

Suppose we flip a coin 100 times and see 52 heads. Let p be the bias of the coin. Then we might estimate: \widehat{p} = \frac{52}{100} (The notation \widehat{p} is sometimes used to indicate an estimator for p.)

If we see k heads in n flips, then the estimator \widehat{p} = \frac{k}{n} is called the common sense estimator for the binomial distribution parameter p.

An estimator \widehat{p} is unbiased if \E(\widehat{p}) = p.

Flip a coin 100 times, and let S be the number of heads. Then S \sim \Bin(100, p). Let \widehat{p} = \frac{S}{100} = \frac{S}{n}. Then: \E(\widehat{p}) \amp = \E\left(\frac{S}{n}\right) \amp = \frac{1}{n} \E\left(S\right) \amp = \frac{1}{n} (np) \amp = p. So \E(\widehat{p} = p, i.e., the common sense estimator is unbiased.

We perform an experiment and collect data. Let p be an unknown parameter value. The likelihood function is: \L(p) = \Pr(\text{data} \mid \text{paramater value is } p).

Let S \sim \Bin(100, p). Suppose we see 52 heads. Then: \Pr(S = k) \amp = b(k; 100, p) = {100 \choose k}p^k (1 - p)^{100 - k} \L(p) \amp = {100 \choose 52}p^{52} (1 - p)^{48} In the first line, the variable k represents the data. In the second line, the variable p represents the parameter value. Our goal, given the collected data, is to find the maximum likelihood estimation (MLE) for the parameter value.

\L(p) is a continuous function over a closed interval p \in [0, 1], so we use the Closed Interval Method. \L'(p) \amp = {100 \choose 52}\left[ 52 p^{51} (1 - p)^{48} + p^{52}48 (1-p)^{47}(-1)\right] \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 (1 - p) - 48p \right] \amp = {100 \choose 52} p^{51} (1-p)^{47}\left[ 52 - 100p \right] Now, we look for critical numbers in the interior of the interval: \L'(p) = 0 \text{ when } p = \frac{52}{100} Finally, we test the critical numbers and the endpoints of the interval to find the max:

Check Candidate Locations for Max p \L(p) 0 0 52/100 \gt 0 1 0

So \widehat{p} = \frac{52}{100} is the MLE.

More generally, a similar calculation will show that, with k heads in n flips, the MLE will be \widehat{p} = \frac{k}{n}.

Suppose we observe a cell, measuring the time T until a toxin molecule leaves the cell. Then T \sim \Exp(\lambda) for some \lambda, with pdf f(t) =\lambda e^{-\lambda t}, \lambda \gt 0 If we see a toxin molecule leave at 0.3 min, what's the MLE for \lambda? \L(\lambda) = \lambda e^{-0.3 \lambda} \quad \text{(density, not probability)} We want to maximize \L(\lambda) over \lambda \in (0, \infty), an open interval. So we'll use the "Open Interval Method". \L'(\lambda) \amp = e^{-0.3\lambda} + \lambda e^{-0.3\lambda}(-0.3) \amp = e^{-0.3\lambda}\left( 1 - 0.3 e^{-0.3\lambda}\right) Then \L'(\lambda) = 0 when \lambda = \frac{1}{0.3}\approx 3.33. Checking \lambda values to the left and the right: \L'(1) \amp = (+)(+) = (+) \L'(10) \amp = (+)(-) = (-) So \L'(\lambda) \gt 0 (and therefore \L(\lambda) is increasing) on (0, 1/0.3), and \L'(\lambda) \lt 0 (and therefore \L(\lambda) is decreasing) on (1/0.3, \infty). Now we can conclude that \widehat{\lambda} = 1/0.3 \approx 3.33 is the location of a global (and not just local) maximum value.

More generally, if the observed time is t, then the MLE will be \widehat{\lambda} = \frac{1}{t}.

What if we had more data points? For example, suppose two toxin molecules leave the cell at t_1 = 0.3 min and t_2 = 0.5 min?

Waiting Times Molecule Time Rate Estimate 1 0.3 1/0.3 \approx 3.33 2 0.5 1/0.5 = 2

How do we combine these data points? We could take the average of the rate estimates: \frac{3.33 + 2}{2} \approx 2.67 Alternatively, we could average the times first, then create a new rate estimate from the average time: \frac{0.3 + 0.5}{2} \amp = 0.4 \frac{1}{0.4} \amp = 2.5 Both of these make some sense, but let's do a careful computation to be certain which way is correct (if either of them is!). \L(\lambda) \amp = \left( \lambda e^{-0.3 \lambda} \right)\left( \lambda e^{-0.5 \lambda} \right) \amp = \lambda^2 e^{-0.3 \lambda - 0.5 \lambda} \amp = \lambda^2 e^{- 0.8 \lambda} Using the Open Interval Method: \L'(\lambda) \amp = 2\lambda e^{-0.8 \lambda} + \lambda^2 e^{-0.8 \lambda} (-0.8) \amp = \lambda e^{-0.8 \lambda} \left[ 2 - 0.8\lambda \right] So \L'(\lambda) = 0 when \lambda = \frac{2}{0.8} = 2.5. Testing points to the left and right: \L'(1) \amp = (+)(+)(+) = (+) \L'(5) \amp = (+)(+)(-) = (-) Now \L'(\lambda) \gt 0 (\L(\lambda) is increasing) on (0, 2.5), and \L'(\lambda) \lt 0 (\L(\lambda) is decreasing) on (2.5, \infty). Therefore, the MLE is \widehat{\lambda} = \frac{2}{0.8} = 2.5.

Tracing the values 2 and 0.8 throughout the calculation, we can see that the value 2 will generally match the number of waiting times collected, and the value 0.8 will be the sum of the waiting times. So, generally, with collected waiting times of t_1, \dotsc, t_n, the MLE will be: \widehat{\lambda} = \frac{n}{t_1 + \dotsb + t_n} = \frac{1}{\text{avg time}}.