Skip to main content

Section Thursday, Jan 29

This is an outline of the topics we covered in class. These notes are not a substitute for your own note-taking. I highly recommend that you take your own notes during class. If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.

Subsection Expected Value

Example 70.

Let \(R\) be the roll of a fair D6. What is the average value of \(R\text{?}\) The sample space is \(\Omega = \{1, 2, 3, 4, 5, 6\}\text{.}\) To find the average:
\begin{align*} \text{Avg} \amp = 1\left(\frac{1}{6}\right) + 2\left(\frac{1}{6}\right) + 3\left(\frac{1}{6}\right) + 4\left(\frac{1}{6}\right) + 5\left(\frac{1}{6}\right) + 6\left(\frac{1}{6}\right) \\ \amp = \frac{21}{6} = \frac{7}{2} = 3.5 \end{align*}
That is: take each value of \(R\text{,}\) multiply by the probability, and add all the results together.

Example 71.

Given a distribution table:
Table 72. Distribution
\(s\) \(\Pr(S = s)\)
1 0.1
2 0.1
3 0.1
4 0.1
5 0.1
6 0.5
Then the average is:
\begin{gather*} \text{Avg} = 1(0.1) + 2(0.1) + 3(0.1) + 4(0.1) + 5(0.1) + 6(0.5) = 4.5. \end{gather*}

Definition 73.

Let \(X\) be a discrete random variable taking values \(x_1, x_2, \dotsc, x_n\) with probabilities \(p_1, p_2, \dotsc, p_n\text{.}\) The expected value of \(X\) is:
\begin{gather*} \E(X) = \sum_{i=1}^n x_ip_i = x_1p_1 + x_2p_2 + \dotsb + x_np_n. \end{gather*}

Example 74.

Let \(X\) indicate \(A\text{,}\) with \(\Pr(A) = p\text{.}\) Then:
\begin{gather*} E(X) = 0 \cdot \underbrace{\Pr(X = 0)}_{1 - p} + 1 \cdot \underbrace{\Pr(X = 1)}_{p} = 0(1-p) + 1p = p. \end{gather*}
That is: "E(indicator random variable) = Pr(event that it indicates)".

Example 75.

Suppose we flip a fair coin 100 times. Let \(H_1, H_2, \dotsc, H_{100}\) indicate heads on each flip. Then \(\Pr(H_i = 1) = \frac{1}{2}\text{,}\) so \(E(H_i) = \frac{1}{2}\) for every \(i\text{.}\)
What if \(X\) indicates a run of 4 heads starting at flip 3? That is, flips 3, 4, 5, and 6 must come up heads, and all other flips can come up either heads or tails. Since these flip results are independent, the probabilities multiply:
\begin{gather*} \Pr(X = 1) = \left(\frac{1}{2}\right)\left(\frac{1}{2}\right)\left(\frac{1}{2}\right)\left(\frac{1}{2}\right) = \frac{1}{16} = \E(X). \end{gather*}

Definition 76.

Let \(X \in [a, b]\) be a continuous random variable with pdf \(f(x)\text{.}\) The expected value of \(X\) is:
\begin{gather*} \E(X) = \int_a^b x f(x)\ dx. \end{gather*}
It’s worth putting this side-by-side with DefinitionΒ 73 to compare the structure of each formula. These both say: "multiply each value of the random variable by the probability, then accumulate all of those products". Expected value is a weighted average of random variable values, with the probabilities as the weights.

Example 77.

Let \(X \in [0, 1]\) with pdf \(f(x) = \frac{3}{2}\sqrt{x}\text{.}\) Then:
\begin{align*} \E(X) \amp = \int_0^1 x f(x)\ dx \\ \amp = \int_0^1 x \frac{3}{2} x^{1/2}\ dx \\ \amp = \int_0^1 \frac{3}{2} x^{3/2}\ dx \\ \amp = \frac{3}{2} \frac{2x^{5/2}}{5}\bigg|_0^1 \\ \amp = \frac{3}{5} \cdot 1^{5/2} - \frac{3}{5} \cdot 0^{5/2} \\ \amp = \frac{3}{5}. \end{align*}
What would happen if you forgot the \(x\text{?}\) Then the calculation would become:
\begin{align*} \int_0^1 f(x)\ dx = 1 = \text{total probability} \amp \end{align*}
This mistake will often be easy to catch. For example, this random variable takes values between 0 and 1, so it doesn’t seem very likely that the average value is 1!

Example 78.

Find \(\E(X^2)\) given a distribution for \(X\text{.}\)
Table 79. Distribution for \(X\)
\(x\) \(\Pr(X = x)\)
-1 0.4
1 0.1
2 0.2
3 0.3
We can start by writing a distribution table for \(X^2\text{:}\)
Table 80. Distribution for \(X^2\)
\(x\) \(\Pr(X^2 = x)\)
1 0.5
4 0.2
9 0.3
Then:
\begin{gather*} \E(X^2) 1(0.5) + 4(0.2) + 9(0.3) = 4. \end{gather*}

Example 82.

Revisiting the previous example, the theorem says we don’t need to first create the distribution table for \(X^2\text{.}\) We can use the distribution table for \(X\text{,}\) and just apply the square to each value:
\begin{gather*} \E(X^2) = (-1)^2(0.4) + (1)^2(0.1) + (2)^2(0.2) + (3)^2(0.3) = 4. \end{gather*}

Example 83.

The situation with continuous random variables is similar. Let \(X\in [0, 1]\) with pdf \(f(x) = \frac{3}{2}\sqrt{x}\text{.}\) Then:
\begin{align*} \E(X^2) \amp = \int_0^1 x^2 f(x)\ dx \\ \amp = \int_0^1 x^2 \frac{3}{2} x^{1/2}\ dx \\ \amp = \int_0^1 \frac{3}{2} x^{5/2}\ dx \\ \amp = \frac{3}{2}\cdot \frac{2x^{7/2}}{7}\bigg|_0^1 \\ \amp = \frac{3}{7}\cdot 1^{7/2} - 0\cdot 0^{7/2} \\ \amp = \frac{3}{7}. \end{align*}
Notice, in particular, that \(\E(X^2) \neq \E(X).\)

Example 85.

Let \(N\) be the number of heads in \(n\) coin flips with bias \(p\text{.}\) Then \(N\sim \Bin(n, p)\text{.}\)
\begin{gather*} \E(N) = \sum_{i=0}^n i\cdot b(i; n, p) = \sum_{i = 0}^n i \cdot {n \choose i} p^i (1-p)^{n - i}. \end{gather*}
Gross.
Instead of calculating this directly, define \(H_1, H_2, \dotsc, H_n\) to indicate heads on each flip. Then:
\begin{align*} N \amp = H_1 + H_2 + \dotsb + H_n \\ \E(N) \amp = \E(H_1 + H_2 + \dotsb + H_n) \\ \E(N) \amp = \E(H_1) + \E(H_2) + \dotsb + \E(H_n) \\ \amp = p + p + \dotsb + p \\ \amp = np. \end{align*}

Example 86.

Flip a fair coin 100 times. Then, the expected number of heads is:
\begin{gather*} \underbrace{(100)}_{n}\underbrace{(0.5)}_{p} = 50. \end{gather*}

Example 87.

Flip a fair coin 100 times. What is the expected number of runs of 4 heads? As in the binomial EV calculation, define indicator random variables \(R_1, R_2, \dotsc, R_{97}\) each indicating a run of 4 heads starting at the specified flip. Then \(\E(R_i) = \frac{1}{16}\) for each \(i\text{.}\) Let \(T\) be the number of runs of 4 heads. Then:
\begin{align*} \E(T) \amp = \E(R_1 + \dotsb + R_{97}) \\ \amp = \E(R_1) + \dotsb + (R_{97}) \\ \amp = \frac{1}{16} + \dotsb + \frac{1}{16} \\ \amp = \frac{97}{16} \approx 6. \end{align*}

Example 88.

Consider a geometric distribution \(T \sim \Geom(p)\text{.}\)
Table 89. Geometric Distribution
\(k\) \(\Pr(T = k)\) with \(p = \frac{1}{2}\)
1 \(p\) \(\frac{1}{2}\)
2 \((1-p)p\) \(\frac{1}{4}\)
3 \((1-p)^2p\) \(\frac{1}{8}\)
4 \((1-p)^3p\) \(\frac{1}{16}\)
\(\vdots\) \(\vdots\) \(\vdots\)
Then \(\E(T) = (1)\left(\frac{1}{2}\right) + (2)\left(\frac{1}{4}\right) + (3)\left(\frac{1}{8}\right) + \dotsb\) is an infinite summation.
Instead, consider the following argument. If we flip a coin until we see heads, we either see heads on flip 1 or not. In the first case, \(T = 1\text{.}\) In the second case, what is the average value of \(T\text{?}\) Starting at flip 2, it will take on average \(\E(T)\) flips to see heads. Since we already flipped the coin once, the total number of flips will be \(1 + \E(T)\text{.}\) So, we can write:
\begin{align*} \E(T) \amp = 1 \left(\frac{1}{2}\right) + (1 + \E(T)) \left(1 - \frac{1}{2}\right) \\ \E(T) \amp = \frac{1}{2} + \frac{1}{2} + \frac{\E(T)}{2} \\ \E(T) - \frac{\E(T)}{2} \amp = 1 \\ \frac{\E(T)}{2} \amp = 1 \\ \E(T) \amp = 2. \end{align*}