Thursday, Jan 29

This is an outline of the topics we covered in class. These notes are not a substitute for your own note-taking. I highly recommend that you take your own notes during class. If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.

Expected Value

Let R be the roll of a fair D6. What is the average value of R? The sample space is \Omega = \{1, 2, 3, 4, 5, 6\}. To find the average: \text{Avg} \amp = 1\left(\frac{1}{6}\right) + 2\left(\frac{1}{6}\right) + 3\left(\frac{1}{6}\right) + 4\left(\frac{1}{6}\right) + 5\left(\frac{1}{6}\right) + 6\left(\frac{1}{6}\right) \amp = \frac{21}{6} = \frac{7}{2} = 3.5 That is: take each value of R, multiply by the probability, and add all the results together.

Given a distribution table:

Distribution s \Pr(S = s) 1 0.1 2 0.1 3 0.1 4 0.1 5 0.1 6 0.5

Then the average is: \text{Avg} = 1(0.1) + 2(0.1) + 3(0.1) + 4(0.1) + 5(0.1) + 6(0.5) = 4.5.

Let X be a discrete random variable taking values x_1, x_2, \dotsc, x_n with probabilities p_1, p_2, \dotsc, p_n. The expected value of X is: \E(X) = \sum_{i=1}^n x_ip_i = x_1p_1 + x_2p_2 + \dotsb + x_np_n.

Let X indicate A, with \Pr(A) = p. Then: E(X) = 0 \cdot \underbrace{\Pr(X = 0)}_{1 - p} + 1 \cdot \underbrace{\Pr(X = 1)}_{p} = 0(1-p) + 1p = p. That is: "E(indicator random variable) = Pr(event that it indicates)".

Suppose we flip a fair coin 100 times. Let H_1, H_2, \dotsc, H_{100} indicate heads on each flip. Then \Pr(H_i = 1) = \frac{1}{2}, so E(H_i) = \frac{1}{2} for every i.

What if X indicates a run of 4 heads starting at flip 3? That is, flips 3, 4, 5, and 6 must come up heads, and all other flips can come up either heads or tails. Since these flip results are independent, the probabilities multiply: \Pr(X = 1) = \left(\frac{1}{2}\right)\left(\frac{1}{2}\right)\left(\frac{1}{2}\right)\left(\frac{1}{2}\right) = \frac{1}{16} = \E(X).

Let X \in [a, b] be a continuous random variable with pdf f(x). The expected value of X is: \E(X) = \int_a^b x f(x)\ dx.

It's worth putting this side-by-side with to compare the structure of each formula. These both say: "multiply each value of the random variable by the probability, then accumulate all of those products". Expected value is a weighted average of random variable values, with the probabilities as the weights.

Let X \in [0, 1] with pdf f(x) = \frac{3}{2}\sqrt{x}. Then: \E(X) \amp = \int_0^1 x f(x)\ dx \amp = \int_0^1 x \frac{3}{2} x^{1/2}\ dx \amp = \int_0^1 \frac{3}{2} x^{3/2}\ dx \amp = \frac{3}{2} \frac{2x^{5/2}}{5}\bigg|_0^1 \amp = \frac{3}{5} \cdot 1^{5/2} - \frac{3}{5} \cdot 0^{5/2} \amp = \frac{3}{5}. What would happen if you forgot the x? Then the calculation would become: \int_0^1 f(x)\ dx = 1 = \text{total probability} \amp This mistake will often be easy to catch. For example, this random variable takes values between 0 and 1, so it doesn't seem very likely that the average value is 1!

Find \E(X^2) given a distribution for X.

Distribution for <m>X</m> x \Pr(X = x) -1 0.4 1 0.1 2 0.2 3 0.3

We can start by writing a distribution table for X^2:

Distribution for <m>X^2</m> x \Pr(X^2 = x) 1 0.5 4 0.2 9 0.3

Then: \E(X^2) 1(0.5) + 4(0.2) + 9(0.3) = 4.

If X takes on values x_1, \dotsc, x_n with probabilities p_1, \dotsc, p_n, and h is any function, then: \E(h(X)) = \sum_{i=1}^n h(x_i)p_i.

Revisiting the previous example, the theorem says we don't need to first create the distribution table for X^2. We can use the distribution table for X, and just apply the square to each value: \E(X^2) = (-1)^2(0.4) + (1)^2(0.1) + (2)^2(0.2) + (3)^2(0.3) = 4.

The situation with continuous random variables is similar. Let X\in [0, 1] with pdf f(x) = \frac{3}{2}\sqrt{x}. Then: \E(X^2) \amp = \int_0^1 x^2 f(x)\ dx \amp = \int_0^1 x^2 \frac{3}{2} x^{1/2}\ dx \amp = \int_0^1 \frac{3}{2} x^{5/2}\ dx \amp = \frac{3}{2}\cdot \frac{2x^{7/2}}{7}\bigg|_0^1 \amp = \frac{3}{7}\cdot 1^{7/2} - 0\cdot 0^{7/2} \amp = \frac{3}{7}. Notice, in particular, that \E(X^2) \neq \E(X).

If X, Y are random variables with finite expected value and k \in \R, then: \E(X+Y) \amp = E(X) + E(Y) \E(kX) \amp = kE(X)

Let N be the number of heads in n coin flips with bias p. Then N\sim \Bin(n, p). \E(N) = \sum_{i=0}^n i\cdot b(i; n, p) = \sum_{i = 0}^n i \cdot {n \choose i} p^i (1-p)^{n - i}. Gross.

Instead of calculating this directly, define H_1, H_2, \dotsc, H_n to indicate heads on each flip. Then: N \amp = H_1 + H_2 + \dotsb + H_n \E(N) \amp = \E(H_1 + H_2 + \dotsb + H_n) \E(N) \amp = \E(H_1) + \E(H_2) + \dotsb + \E(H_n) \amp = p + p + \dotsb + p \amp = np.

Flip a fair coin 100 times. Then, the expected number of heads is: \underbrace{(100)}_{n}\underbrace{(0.5)}_{p} = 50.

Flip a fair coin 100 times. What is the expected number of runs of 4 heads? As in the binomial EV calculation, define indicator random variables R_1, R_2, \dotsc, R_{97} each indicating a run of 4 heads starting at the specified flip. Then \E(R_i) = \frac{1}{16} for each i. Let T be the number of runs of 4 heads. Then: \E(T) \amp = \E(R_1 + \dotsb + R_{97}) \amp = \E(R_1) + \dotsb + (R_{97}) \amp = \frac{1}{16} + \dotsb + \frac{1}{16} \amp = \frac{97}{16} \approx 6.

Consider a geometric distribution T \sim \Geom(p).

Geometric Distribution k \Pr(T = k) with p = \frac{1}{2} 1 p \frac{1}{2} 2 (1-p)p \frac{1}{4} 3 (1-p)^2p \frac{1}{8} 4 (1-p)^3p \frac{1}{16} \vdots \vdots \vdots

Then \E(T) = (1)\left(\frac{1}{2}\right) + (2)\left(\frac{1}{4}\right) + (3)\left(\frac{1}{8}\right) + \dotsb is an infinite summation.

Instead, consider the following argument. If we flip a coin until we see heads, we either see heads on flip 1 or not. In the first case, T = 1. In the second case, what is the average value of T? Starting at flip 2, it will take on average \E(T) flips to see heads. Since we already flipped the coin once, the total number of flips will be 1 + \E(T). So, we can write: \E(T) \amp = 1 \left(\frac{1}{2}\right) + (1 + \E(T)) \left(1 - \frac{1}{2}\right) \E(T) \amp = \frac{1}{2} + \frac{1}{2} + \frac{\E(T)}{2} \E(T) - \frac{\E(T)}{2} \amp = 1 \frac{\E(T)}{2} \amp = 1 \E(T) \amp = 2.