Likelihood answers

This commit is contained in:
2026-09-25 14:21:58 +00:00
parent 1211bb8e8c
commit ee9a651e95
+215 -104
View File
@@ -3,38 +3,52 @@
<p>
So far, we've been concerned with probability theory.
Starting with a probability distribution and some parameter values, we've tried to answer questions like: What's the probability of seeing certain experimental results? Statistics is concerned with going in the other direction: Upon seeing the experimental results, can we determine the type of underlying probability distribution? Can we determine its parameters?
Starting with a probability distribution and some parameter values, we've
tried to answer questions like: What's the probability of seeing certain
experimental results?
Statistics is concerned with going in the other direction: Upon seeing the
experimental results, can we determine the type of underlying probability
distribution?
Can we determine its parameters?
</p>
<definition xml:id="def-estimator">
<statement>
<p>
An <term>estimator</term> is a value of a parameter computed from a sample of data.
An <term>estimator</term> is a value of a parameter computed from a
sample of data.
</p>
</statement>
</definition>
<example>
<p>
Suppose we find a coin on the street and don't know whether or not it's fair.
Suppose we find a coin on the street and don't know whether or not it's
fair.
We want to know the probability <m>p</m> of the coin coming up heads.
We might, for example, flip the coin <m>n</m> times and count the number <m>k</m> of heads.
We might, for example, flip the coin <m>n</m> times and count the number
<m>k</m> of heads.
Then, we'll estimate <m>p = \frac{k}{n}</m>.
We'll refer to this as a <term>common sense</term> estimator.
(Other distributions and parameter types will have different notions of "common sense".)
(Other distributions and parameter types will have different notions of
"common sense".)
</p>
</example>
<p>
An estimator is, itself, a random variable: it produces a numerical value based on the results of an experiment.
We'll use notation like <m>\est{p}</m> for a random variable which is an estimator for a parameter <m>p</m>.
(Similarly, <m>\est{\lambda}</m> would denote an estimator for a parameter called <m>\lambda</m>.)
An estimator is, itself, a random variable: it produces a numerical value
based on the results of an experiment.
We'll use notation like <m>\est{p}</m> for a random variable which is an
estimator for a parameter <m>p</m>.
(Similarly, <m>\est{\lambda}</m> would denote an estimator for a parameter
called <m>\lambda</m>.)
</p>
<definition xml:id="def-unbiased">
<statement>
<p>
An estimator <m>\est{p}</m> is called <term>unbiased</term> if <m>\E(\est{p}) = p</m>.
An estimator <m>\est{p}</m> is called <term>unbiased</term> if
<m>\E(\est{p}) = p</m>.
</p>
</statement>
</definition>
@@ -42,23 +56,30 @@
<example>
<statement>
<p>
Suppose we have a coin with parameter <m>p</m>, which we'll flip <m>n</m> times and count the number <m>k</m> of heads.
Suppose we have a coin with parameter <m>p</m>, which we'll flip
<m>n</m> times and count the number <m>k</m> of heads.
We use the unbiased estimator <m>\est{p} = \frac{k}{n}</m>.
In this case, notice that <m>k \sim \Bin(n, p)</m>, so we know <m>\E(k) = np</m>, although we don't know the value of <m>p</m>.
(We probably do know the value of <m>n</m>; after all, we're flipping the coin!) Now:
In this case, notice that <m>k \sim \Bin(n, p)</m>, so we know
<m>\E(k) = np</m>, although we don't know the value of <m>p</m>.
(We probably do know the value of <m>n</m>; after all, we're flipping
the coin!) Now:
<md>
<mrow> \E(\est{p}) = \E\left(\frac{k}{n}\right) = \frac{1}{n} \cdot \E(k) = \frac{1}{n} \cdot np = p. </mrow>
</md>
It's worth pausing for a moment to be appropriately impressed with ourselves.
It's worth pausing for a moment to be appropriately impressed with
ourselves.
We still don't know the true value of <m>p</m>.
But we managed to show that our common sense method of estimating <m>p</m> gives, on average, the correct value.
But we managed to show that our common sense method of estimating
<m>p</m> gives, on average, the correct value.
</p>
</statement>
</example>
<p>
We'd like to be able to collect some data and use that data to estimate the values of whatever parameters our distribution has.
Perhaps as a starting point, it would be good to identify the single most likely value of a parameter:
We'd like to be able to collect some data and use that data to estimate the
values of whatever parameters our distribution has.
Perhaps as a starting point, it would be good to identify the single most
likely value of a parameter:
</p>
<definition xml:id="def-likelihood">
@@ -69,7 +90,9 @@
<md>
<mrow> \L(p) = \Pr(\text{data} \mid \text{parameter value is } p). </mrow>
</md>
The value of <m>p</m> which maximizes the function <m>\L(p)</m> is called the <term>maximum likelihood estimation</term>, or <term>MLE</term>, of <m>p</m>.
The value of <m>p</m> which maximizes the function <m>\L(p)</m> is
called the <term>maximum likelihood estimation</term>, or
<term>MLE</term>, of <m>p</m>.
</p>
</statement>
</definition>
@@ -83,19 +106,24 @@
<md>
<mrow> \Pr(N = k) \amp = b(k; 100, p) = {100 \choose k} p^k (1 - p)^{100 - k} </mrow>
</md>
In this case, <m>k</m> is the data that we collect, and <m>p</m> is the value of the parameter.
In this case, <m>k</m> is the data that we collect, and <m>p</m> is the
value of the parameter.
Suppose we see <m>52</m> heads.
Then:
<md>
<mrow> \L(p) = {100 \choose 52} p^{52} (1 - p)^{48}. </mrow>
</md>
If we want to know the most likely value of the parameter <m>p</m>, then we should maximize <m>\L(p)</m> over the interval <m>0 \leq p \leq 1</m>.
If we want to know the most likely value of the parameter <m>p</m>, then
we should maximize <m>\L(p)</m> over the interval
<m>0 \leq p \leq 1</m>.
<md>
<mrow> \L'(p) \amp = {100 \choose 52} \left[ 52 p^{51}(1 - p)^{48} + p^{52} 48 (1 - p)^{47}(-1)\right] </mrow>
<mrow> \amp = {100 \choose 52} p^{51} (1 - p)^{47} \left[ 52 (1 - p) - 48p \right] </mrow>
<mrow> \amp = {100 \choose 52} p^{51} (1 - p)^{47} \left[ 52 - 100 p \right]. </mrow>
</md>
We can see that <m>\L'(p) = 0</m> when <m>p = 0, 1, \frac{52}{100}</m>, and the endpoints of the interval we're maximizing over are <m>p = 0, 1</m>, so we can build a table:
We can see that <m>\L'(p) = 0</m> when <m>p = 0, 1, \frac{52}{100}</m>,
and the endpoints of the interval we're maximizing over are
<m>p = 0, 1</m>, so we can build a table:
</p>
<table>
@@ -125,15 +153,19 @@
</table>
<p>
Note that we don't need to know the exact value of <m>\L(52/100)</m> in order to see that it's strictly positive, and therefore the maximum value of <m>\L(p)</m>.
Note that we don't need to know the exact value of <m>\L(52/100)</m> in
order to see that it's strictly positive, and therefore the maximum
value of <m>\L(p)</m>.
So, the MLE of <m>p</m> is <m>52/100</m>.
</p>
</statement>
</example>
<p>
We can see in the previous example that the MLE of <m>p</m> is also the common sense estimation of <m>p</m>.
This will be the case in general for the binomial distribution, so we won't need to redo this work over and over:
We can see in the previous example that the MLE of <m>p</m> is also the
common sense estimation of <m>p</m>.
This will be the case in general for the binomial distribution, so we won't
need to redo this work over and over:
</p>
<fact xml:id="fact-MLE-binomial">
@@ -141,13 +173,16 @@
<statement>
<p>
If we see <m>k</m> heads in <m>n</m> coin flips, then the MLE of the bias <m>p</m> is <m>\frac{k}{n}</m>.
If we see <m>k</m> heads in <m>n</m> coin flips, then the MLE of the
bias <m>p</m> is <m>\frac{k}{n}</m>.
</p>
</statement>
</fact>
<p>
As we remember from Calculus 1, maximizing a continuous function works slightly differently over a closed interval (like the previous example) or an open interval (like the next example).
As we remember from Calculus 1, maximizing a continuous function works
slightly differently over a closed interval (like the previous example) or
an open interval (like the next example).
</p>
<example>
@@ -155,31 +190,41 @@
<p>
Suppose a radioactive material emits particles as it decays.
Let <m>N</m> count the particles emitted.
Then <m>N \sim \Poiss(\lambda)</m> for some unknown rate parameter <m>\lambda</m>:
Then <m>N \sim \Poiss(\lambda)</m> for some unknown rate parameter
<m>\lambda</m>:
<md>
<mrow> \Pr(N = k) = p(k; \lambda) = \frac{\lambda^k}{k!} e^{-\lambda} </mrow>
</md>
Suppose we observe a sample of material for 1 hour and count 8 particles emitted.
Suppose we observe a sample of material for 1 hour and count 8 particles
emitted.
Then:
<md>
<mrow> \L(\lambda) = \frac{\lambda^8}{8!} e^{-\lambda} </mrow>
</md>
To maximize <m>\L(\lambda)</m> over the interval <m>0 \lt \lambda \lt \infty</m>, we start by finding critical points.
To maximize <m>\L(\lambda)</m> over the interval
<m>0 \lt \lambda \lt \infty</m>, we start by finding critical points.
<md>
<mrow> \L'(\lambda) \amp = \frac{1}{8!} \left[ 8 \lambda^7 e^{-\lambda} + \lambda^8 e^{-\lambda} (-1)\right] </mrow>
<mrow> \amp = \frac{1}{8!} \lambda^7 e^{-\lambda} \left[ 8 - e^{-\lambda} \right] </mrow>
</md>
The only critical point is <m>\lambda = 8</m>, but we have not justified that this is the location of a global maximum.
A critical point is only a potential location of a local minimum or maximum.
But with a bit more justification: observe that <m>\L'</m> is positive on the interval <m>(0, 8)</m> and negative on the interval <m>(8, \infty)</m>.
Therefore the function <m>\L</m> increases on <m>(0, 8)</m> and decreases on <m>(8, \infty)</m>.
So it must reach its maximum at <m>\lambda = 8</m>, which is therefore the MLE.
The only critical point is <m>\lambda = 8</m>, but we have not justified
that this is the location of a global maximum.
A critical point is only a potential location of a local minimum or
maximum.
But with a bit more justification: observe that <m>\L'</m> is positive
on the interval <m>(0, 8)</m> and negative on the interval
<m>(8, \infty)</m>.
Therefore the function <m>\L</m> increases on <m>(0, 8)</m> and
decreases on <m>(8, \infty)</m>.
So it must reach its maximum at <m>\lambda = 8</m>, which is therefore
the MLE.
</p>
</statement>
</example>
<p>
As with the binomial distribution, the calculation will be essentially the same regardless of the specific number of particles observed, so:
As with the binomial distribution, the calculation will be essentially the
same regardless of the specific number of particles observed, so:
</p>
<fact xml:id="fact-MLE-Poisson">
@@ -187,22 +232,28 @@
<statement>
<p>
If we observe a Poisson process and see <m>k</m> events occur, then the MLE of the rate parameter <m>\lambda</m> is <m>k</m>.
If we observe a Poisson process and see <m>k</m> events occur, then the
MLE of the rate parameter <m>\lambda</m> is <m>k</m>.
</p>
</statement>
</fact>
<p>
The situation for continuous random variables is similar, but slightly different.
To build the likelihood function, we should use the pdf of the continuous random variable.
So the likelihood function will give the probabiliy density given the data collected, rather than the probability itself.
The situation for continuous random variables is similar, but slightly
different.
To build the likelihood function, we should use the pdf of the continuous
random variable.
So the likelihood function will give the probabiliy density given the data
collected, rather than the probability itself.
</p>
<example xml:id="example-exponential-MLE">
<statement>
<p>
Suppose we observe a cell and measure the time <m>T</m> until a toxin molecule leaves the cell.
Then <m>T \sim \Exp(\lambda)</m> for some unknown rate parameter <m>\lambda</m>:
Suppose we observe a cell and measure the time <m>T</m> until a toxin
molecule leaves the cell.
Then <m>T \sim \Exp(\lambda)</m> for some unknown rate parameter
<m>\lambda</m>:
<md>
<mrow> f(t) = \lambda e^{-\lambda t}. </mrow>
</md>
@@ -214,12 +265,16 @@
<mrow> \amp = e^{-0.3\lambda}\left[ 1 - 0.3\lambda\right] </mrow>
</md>
The critical point is <m>\lambda = \frac{1}{0.3} \approx 3.33</m>.
Since <m>\L' \gt 0</m> on <m>(0, 3.33)</m> and <m>\L' \lt 0</m> on <m>(3.33, \infty)</m>, there is a global maximum at <m>\lambda \approx 3.33</m>, which is therefore the MLE.
Since <m>\L' \gt 0</m> on <m>(0, 3.33)</m> and <m>\L' \lt 0</m> on
<m>(3.33, \infty)</m>, there is a global maximum at
<m>\lambda \approx 3.33</m>, which is therefore the MLE.
</p>
<p>
The specific time <m>0.3</m> minutes doesn't particularly matter in this calculation.
Whatever the time <m>t</m>, essentially the same calculation will result in a MLE of <m>\lambda = 1/t</m>.
The specific time <m>0.3</m> minutes doesn't particularly matter in this
calculation.
Whatever the time <m>t</m>, essentially the same calculation will result
in a MLE of <m>\lambda = 1/t</m>.
But what if we collect multiple pieces of data?
</p>
@@ -266,12 +321,18 @@
</table>
<p>
How should take all of this data into account in our maximum likelihood estimation? We might consider taking the average of all of the separate rate estimations, which would give <m>1.87</m>.
Is this the most likely? We need some mathematical justification.
How should take all of this data into account in our maximum likelihood
estimation?
We might consider taking the average of all of the separate rate
estimations, which would give <m>1.87</m>.
Is this the most likely?
We need some mathematical justification.
</p>
<p>
To account for multiple, independent data points, we should multiply the probability densities for each in the creation of our likelihood function:
To account for multiple, independent data points, we should multiply the
probability densities for each in the creation of our likelihood
function:
<md>
<mrow> \L(\lambda) \amp = \left(\lambda e^{-0.3\lambda}\right)\left(\lambda e^{-0.8\lambda}\right)\left(\lambda e^{-0.5\lambda}\right)\left(\lambda e^{-0.6\lambda}\right)\left(\lambda e^{-0.9\lambda}\right) </mrow>
<mrow> \amp = \lambda^5 e^{-0.3\lambda - 0.8\lambda - 0.5\lambda - 0.6\lambda - 0.9\lambda } </mrow>
@@ -284,14 +345,20 @@
<mrow> \amp = \lambda^4 e^{-3.1\lambda} \left[5 - 3.1\lambda \right] </mrow>
</md>
The only critical point is <m>5/3.1 \approx 1.61</m>.
Since <m>\L' \gt 0</m> on <m>(0, 1.61)</m> and <m>\L' \lt 0</m> on <m>(1.61, \infty)</m>, there is a global maximum at <m>\lambda = 1.61</m>, which is therefore the MLE.
Since <m>\L' \gt 0</m> on <m>(0, 1.61)</m> and <m>\L' \lt 0</m> on
<m>(1.61, \infty)</m>, there is a global maximum at
<m>\lambda = 1.61</m>, which is therefore the MLE.
</p>
<p>
It may seem less clear how to generalize this calculation for other tables of data.
Observe that the value <m>3.1</m> is the sum of the five times in the table, so <m>3.1/5</m> is the average time.
The MLE turned out to be the reciprocal of the average time (just as the MLE with only one data point was the reciprocal of that one time).
Notice that this does <em>not</em> match the guess we made previously of averaging the individual rate estimations for each data point.
It may seem less clear how to generalize this calculation for other
tables of data.
Observe that the value <m>3.1</m> is the sum of the five times in the
table, so <m>3.1/5</m> is the average time.
The MLE turned out to be the reciprocal of the average time (just as the
MLE with only one data point was the reciprocal of that one time).
Notice that this does <em>not</em> match the guess we made previously of
averaging the individual rate estimations for each data point.
</p>
</statement>
</example>
@@ -301,67 +368,111 @@
<statement>
<p>
If we observe a Poisson process and see events occur after waiting times <m>t_1, t_2, \dotsc, t_n</m>, then the MLE of the rate parameter <m>\lambda</m> is <m>\frac{n}{t_1 + t_2 + \dotsb + t_n}</m>.
If we observe a Poisson process and see events occur after waiting times
<m>t_1, t_2, \dotsc, t_n</m>, then the MLE of the rate parameter
<m>\lambda</m> is <m>\frac{n}{t_1 + t_2 + \dotsb + t_n}</m>.
</p>
</statement>
</fact>
<exercises>
<exercise>
<statement>
<p>
Suppose a coin has an unknown probability of coming up heads.
We perform the experiment in five independent trials, during which it takes 4, 5, 4, 3, and 6 flips to see our first heads in each trial.
What is the maximum likelihood estimation for the probability of the coin coming up heads on a flip?
</p>
</statement>
</exercise>
<exercise>
<statement>
<p>
Suppose a coin has an unknown probability of coming up heads.
We perform the experiment in five independent trials, during which it
takes 4, 5, 4, 3, and 6 flips to see our first heads in each trial.
What is the maximum likelihood estimation for the probability of the
coin coming up heads on a flip?
</p>
</statement>
<exercise>
<statement>
<p>
Suppose a coin has an unknown probability of coming up heads.
We perform the experiment in <m>n</m> independent trials, during which it takes <m>k_1, k_2, \dotsc, k_n</m> flips to see our first heads in each trial.
Find a "common sense" MLE formula for the geometric distribution.
</p>
</statement>
<answer>
<p>
<m>5/27 \approx 0.227</m>.
</p>
</answer>
</exercise>
<hint>
<p>
You <em>could</em> set up a calculation analogous to <xref ref="example-exponential-MLE"/>.
Or, you could consider <xref ref="fact-MLE-binomial"/>.
</p>
</hint>
</exercise>
<exercise>
<statement>
<p>
Suppose a coin has an unknown probability of coming up heads.
We perform the experiment in <m>n</m> independent trials, during which
it takes <m>k_1, k_2, \dotsc, k_n</m> flips to see our first heads in
each trial.
Find a "common sense" MLE formula for the geometric distribution.
</p>
</statement>
<exercise>
<statement>
<p>
A particular store owner wants to approximate the average hourly rate at which customers come into the store.
They observe 80 customers enter during a particular 4-hour shift.
What is the maximum likelihood estimation for the hourly customer rate?
</p>
</statement>
</exercise>
<hint>
<p>
You <em>could</em> set up a calculation analogous to
<xref ref="example-exponential-MLE"/>.
Or, you could consider <xref ref="fact-MLE-binomial"/>.
</p>
</hint>
<exercise>
<statement>
<p>
A radioactive material emits particles at an unknown probabilistic rate <m>\lambda</m> particles per minute.
We observe particles emitted at times 1.1, 1.7, 1.3, 2.2, 1.9, and 1.8 minutes.
Write the likelihood function <m>\mathcal{L}(\lambda)</m> based on this data.
What is the maximum likelihood estimation for <m>\lambda</m>?
</p>
</statement>
</exercise>
<answer>
<p>
<m>\frac{n}{k_1 + k_2 + \dotsb + k_n}</m>.
</p>
</answer>
</exercise>
<exercise>
<statement>
<p>
Suppose a parameter <m>\theta</m> takes values in <m>[0, 1]</m> with likelihood function <m>\mathcal{L}(\theta) = \sqrt{\theta} - \theta^2</m>.
Find the maximum likelihood estimation of <m>\theta</m>.
</p>
</statement>
</exercise>
<exercise>
<statement>
<p>
A particular store owner wants to approximate the average hourly rate
at which customers come into the store.
They observe 80 customers enter during a particular 4-hour shift.
What is the maximum likelihood estimation for the hourly customer
rate?
</p>
</statement>
<answer>
<p>
<m>20</m>.
</p>
</answer>
</exercise>
<exercise>
<statement>
<p>
A radioactive material emits particles at an unknown probabilistic
rate <m>\lambda</m> particles per minute.
We observe particles emitted at times 1.1, 1.7, 1.3, 2.2, 1.9, and 1.8
minutes.
Write the likelihood function <m>\mathcal{L}(\lambda)</m> based on
this data.
What is the maximum likelihood estimation for <m>\lambda</m>?
</p>
</statement>
<answer>
<p>
<m>\mathcal{L}(\lambda) = \left( \lambda e^{-1.1\lambda} \right) \left( \lambda e^{-1.7\lambda} \right) \left( \lambda e^{-1.3\lambda} \right) \left( \lambda e^{-2.2\lambda} \right) \left( \lambda e^{-1.9\lambda} \right) \left( \lambda e^{-1.8\lambda} \right)</m>.
The MLE is <m>0.6</m>.
</p>
</answer>
</exercise>
<exercise>
<statement>
<p>
Suppose a parameter <m>\theta</m> takes values in <m>[0, 1]</m> with
likelihood function
<m>\mathcal{L}(\theta) = \sqrt{\theta} - \theta^2</m>.
Find the maximum likelihood estimation of <m>\theta</m>.
</p>
</statement>
<answer>
<p>
<m>\left(\frac{1}{4}\right)^{2/3} \approx 0. 37</m>.
</p>
</answer>
</exercise>
</exercises>
</section>