Week 1

This is an outline of the topics we covered in the first week of class. These notes are not a substitute for your own note-taking. I highly recommend that you take your own notes during class. If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.

Tuesday 1/13 Sec 1.1: Sets

In prior math courses, you mostly asked deterministic questions. Now, we need new tools to model randomness.

A toxin molecule in a cell has a certain chance each minute to leave the cell.

A patient takes a diagnostic test for a disease and wants to know the chance that they have the disease based on the test result.

The sample space, often denoted \Omega, is the set of all possible results of an experiment. A single result is called an outcome, while a collection of results is called an event.

An experiment consists of rolling a standard 6-sided die (D6). The sample space is \Omega = \{1, 2, 3, 4, 5, 6\}. One possible event is A = \{2, 4, 6\}, i.e., the event that the result of the roll is even.

Note: we'll use notation like D6 to indicate a 6-sided die with faces 1, 2, 3, 4, 5, 6. Similarly, for example, D4 will indicate a 4-sided die with faces 1, 2, 3, 4.

The symbol \in means "is an element of", as in 4 \in A.

The symbol \subset means "is a subset of", as in A \subset \Omega. This means that every element of the set A is also an element of the set \Omega.

Consider sets A and B, each contained inside \Omega. We can combine sets in a variety of ways:

  • Union

    The union of A and B is the set A \cup B = \{x \mid x \in A \text{ or } x \in B\}.

  • Intersection

    The intersection of A and B is the set A \cap B = \{x \mid x \in A \text{ and } x \in B\}.

  • Difference

    The set difference A-B is the set A - B = \{x \mid x \in A \text{ and } x \notin B\}.

  • Complement

    The complement of A is the set A^c = \{x \in \Omega \mid x \notin A\}.

  • Empty Set

    The empty set, usually written \emptyset or \{\}, is the set which contains no elements.

  • It's useful sometimes to draw pictures called Venn diagrams representing set interactions.

    Union

    Venn diagram showing sets A, B with the region representing A \cup B shaded.

    \begin{tikzpicture} \def\firstcircle{(180:1.75cm) circle (2.5cm)} \def\secondcircle{(0:1.75cm) circle (2.5cm)} \fill[gray!30] \firstcircle; \fill[gray!30] \secondcircle; \draw \firstcircle node[text=black,left] {$A$}; \draw \secondcircle node [text=black,right] {$B$}; \end{tikzpicture}
    Intersection

    Venn diagram showing sets A, B with the region representing A \cap B shaded.

    \begin{tikzpicture} \def\firstcircle{(180:1.75cm) circle (2.5cm)} \def\secondcircle{(0:1.75cm) circle (2.5cm)} \fill[white] \firstcircle; \fill[white] \secondcircle; \begin{scope} \clip \firstcircle; \clip \secondcircle; \fill[gray!30] \firstcircle; \end{scope} \draw \firstcircle node[text=black,left] {$A$}; \draw \secondcircle node [text=black,right] {$B$}; \end{tikzpicture}
    Difference

    Venn diagram showing sets A, B with the region representing A - B shaded.

    \begin{tikzpicture} \def\firstcircle{(180:1.75cm) circle (2.5cm)} \def\secondcircle{(0:1.75cm) circle (2.5cm)} \fill[gray!30] \firstcircle; \fill[white] \secondcircle; \draw \firstcircle node[text=black,left] {$A$}; \draw \secondcircle node [text=black,right] {$B$}; \end{tikzpicture}
    Complement

    Venn diagram showing set A \subset \Omega with the region representing A^c shaded.

    \begin{tikzpicture} \def\firstcircle{(0, 0) circle (2.5cm)} \fill [gray!30] (-4, -3) rectangle (5, 3); \fill[white] \firstcircle; \draw \firstcircle node[text=black] {$A$}; \draw (-4, -3) rectangle (5, 3) node [text=black,right] {$\Omega$}; \end{tikzpicture}
    Sec 1.2: Probability

    Next, we want to start assigning probabilities to each individual outcome so we can then find the probabilities of events.

    An experiment consists of rolling a D6. The sample space is \Omega = \{1, 2, 3, 4, 5, 6\}. We might assign probabilities as follows:

    Distribution for a fair die x \Pr(x) 1 1/6 2 1/6 3 1/6 4 1/6 5 1/6 6 1/6

    Note that we don't have to assign the same probability to each outcome. If we do, we call this distribution uniform. If we have some event, such as A = \{2, 4, 6\}, then we calculate the probability of the event by adding together the probabilities of the outcomes that make up the event: \Pr(A) = \Pr(2) + \Pr(4) + \Pr(6) = \frac{1}{6} + \frac{1}{6} + \frac{1}{6} = \frac{3}{6} = \frac{1}{2} Note that this is the same result that we would get if we counted the total number of outcomes in A and divided by the total number of outcomes in \Omega.

    Given a finite set A, the cardinality of A, written |A|, is the number of elements in A.

    If \Omega is a finite probability space with the uniform distribution and A \subset \Omega is an event, then: \Pr(A) = \frac{|A|}{|\Omega|}

    Suppose we have a weighted die that's much more likely to come up 6 than any other outcome.

    Distribution for a fair die x \Pr(x) 1 0.1 2 0.1 3 0.1 4 0.2 5 0.1 6 0.4

    With this distribution, if A = \{2, 4, 6\}, then: \Pr(A) = 0.1 + 0.2 + 0.4 = 0.7 \neq \frac{3}{6}.

    Pick a number from 1 to 10. What is the probability of picking 3?

    Some considerations:

    1. Are we using the uniform distribution?

    2. What even is the sample space here? Do we need to pick only integers, or is \pi an outcome here?

    3. What would "uniform" mean if the sample space was infinite?

    A probability distribution on a sample space \Omega assigns probabilities to every event, satisfying the following conditions:

    1. \Pr(\Omega) = 1.

    2. 0 \leq \Pr(A) \leq 1 for any event A.

    3. If A \cap B = \emptyset, then \Pr(A\cup B) = \Pr(A) + \Pr(B).

    Suppose we flip a coin until we see heads. The sample space is \Omega = \{H, TH, TTH, TTTH, \dotsc\}. What would "uniform" mean here? Is there any way we could assign the same probability to each individual outcome here?

    Instead of "uniform", what if we want to treat the coin as a fair coin? Then what would the distribution be?

    Distribution assuming a fair coin x # flips \Pr(x) H 1 1/2 TH 2 1/4 TTH 3 1/8 TTTH 4 1/16 \vdots \vdots T\dotsm TH n 1/2^n \vdots \vdots

    One last (very useful!) observation. Recall: If A\cap B = \emptyset, then \Pr(A\cup B) = \Pr(A) + \Pr(B).

    For any event A, A\cap A^c = \emptyset and A \cup A^c = \Omega. So: 1 = \Pr(\underbrace{A \cup A^c}_{\Omega}) = \Pr(A) + \Pr(A^c). We can rewrite this in two useful ways: \Pr(A) \amp = 1 - \Pr(A^c) \Pr(A^c) \amp = 1 - \Pr(A)

    Thursday 1/15 Conditional Probability

    Question: How does evidence (e.g., knowledge of one event occurring) change our knowledge of probabilities for other events?

    Roll a fair D6 two times. Let A = \{\text{sum } \geq 10\} and B = \{\text{first roll is } 6\}. A feels more likely if we already know B has occurred.

    The conditional probability of A given B is: , \Pr(A \mid B) = \frac{\Pr(A\cap B)}{\Pr(B)}

    \Pr(A \mid B) tells the proportion of B which is overlapped by A.

    Two overlapping circles representing events A and B sit inside a rectangle representing the sample space \Omega. The circle labeled B is shaded. The portion of that circle which is overlapped by the A circle is also filled in with slanted lines.

    \begin{tikzpicture} \def\firstcircle{(180:1.75cm) circle (2.5cm)} \def\secondcircle{(0:1.75cm) circle (2.5cm)} \fill [gray!30] \secondcircle; \begin{scope} \clip \firstcircle; \clip \secondcircle; \fill [pattern=north east lines] \firstcircle; \end{scope} \draw \firstcircle node[text=black] {$A$}; \draw \secondcircle node[text=black] {$B$}; \draw (-5, -3) rectangle (5, 3) node [text=black,right] {$\Omega$}; \end{tikzpicture}

    Continuing from the previous example, |\Omega| = 36. A \amp = \{(4, 6), (5, 5), (5, 6), (6, 4), (6, 5), (6, 6)\} B \amp = \{(6, 1), (6, 2), (6, 3), (6, 4), (6, 5), (6, 6)\} A \cap B \amp = \{(6, 4), (6, 5), (6, 6)\} So \Pr(A) = \frac{6}{36}, \Pr(B) = \frac{6}{36}, \text{and } \Pr(A\cap B) = \frac{3}{36}. Then: \Pr(A \mid B) = \frac{\Pr(A\cap B)}{\Pr(B)} = \frac{3/36}{6/36} = \frac{3}{6} = \frac{1}{2}. Notice that \Pr(A\mid B) is significantly larger than \Pr(A).

    Diagnostic Testing

    Setup: A patient takes a diagnostic test. Let P be the event that they test positive. Let D be the event that they have the disease.

    The sensitivity of a diagnostic test is \Pr(P \mid D). The specificity of a diagnostic test is \Pr(P^c \mid D^c).

    But, what the patient really wants to know is \Pr(D \mid P).

    A disease has a prevalence of 1%. A test has sensitivity of 90% and specificity of 91%. For a patient who gets a positive test result, what is the probability that they have the disease?

    \text{A) } 9/10 \amp \amp \text{B) } 8/10 \amp \amp \text{C) } 1/10 \amp \amp \text{D) } 1/100

    C!

    Bayes' Theorem (v1)

    For events A, B with nonzero probability: \Pr(B \mid A) = \frac{\Pr(A \mid B)\Pr(B)}{\Pr(A)}

    For example, if a patient sees a positive diagnostic test result, they might try to calculate: \Pr(D \mid P) = \frac{\Pr(P \mid D)\Pr(D)}{\Pr(P)} \Pr(P \mid D) is the sensitivity. \Pr(D) could be the prevalence. We don't have direct access to \Pr(P).

    Observation: \Omega = D \cup D^c, so P = (P\cap D) \cup (P\cap D^c).

    \begin{tikzpicture} \def\firstcircle{(0, 0) circle (2)} \def\leftside{(-3, -3) rectangle (-0.5, 3)} \def\rightside{(-0.5, -3) rectangle (4, 3)} \begin{scope} \clip\leftside; \fill [gray!50] \firstcircle; \end{scope} \begin{scope} \clip\rightside; \fill [pattern=north east lines] \firstcircle; \end{scope} \draw (0, 0) circle (2); \node at (2.5, 0) {$P$}; \draw (-0.5, 3) to (-0.5, -3); \node at (-1.5, -3.5) {$D$}; \node at (1.5, -3.5) {$D^c$}; \draw (-3, -3) rectangle (4, 3) node [text=black,right] {$\Omega$}; \end{tikzpicture}

    Observation 2: \Pr(P \mid D) \amp \frac{\Pr(P\cap D)}{\Pr(D)} \amp \amp \Rightarrow \amp \Pr(P \cap D) \amp = \Pr(P\mid D)\Pr(D) \Pr(P \mid D^c) \amp \frac{\Pr(P\cap D^c)}{\Pr(D^c)} \amp \amp \Rightarrow \amp \Pr(P \cap D^c) \amp = \Pr(P\mid D^c)\Pr(D^c) So: \Pr(P) = \Pr(P\mid D)\Pr(D) + \Pr(P\mid D^c)\Pr(D^c)

    Bayes' Theorem (v2)

    \Pr(B \mid A) = \frac{\Pr(A \mid B)\Pr(B)}{\Pr(A \mid B)\Pr(B) + \Pr(A \mid B^c)\Pr(B^c)}

    Continuing from the previous example: \Pr(D\mid P) \amp = \frac{\Pr(P\mid D)\Pr(D)}{\Pr(P\mid D)\Pr(D) + \Pr(P\mid D^c)\Pr(D^c)} \amp = \frac{(0.9)(0.01)}{(0.9)(0.01) + (1 - 0.91)(1 - 0.01)} \amp \approx 0.092 What if the patient got a negative test result instead? In that case, what is the probaiblity they do not have the disease? \Pr(D^c\mid P^c) \amp = \frac{\Pr(P^c\mid D^c)\Pr(D^c)}{\Pr(P^c\mid D^c)\Pr(D^c) + \Pr(P^c\mid D)\Pr(D)} \amp = \frac{(0.91)(1 - 0.01)}{(0.91)(1 - 0.01) + (1 - 0.9)(0.01)} \amp \approx 0.999

    Independent Events

    Question: \Pr(A \mid B) is supposed to capture how information about B affects the probability of A. What if it doesn't?

    Events A, B are independent if \Pr(A \mid B) = \Pr(A).

    Observation: If A, B have nonzero probability and are independent, then: \Pr(A\mid B) \amp = \Pr(A) \frac{\Pr(A\cap B)}{\Pr(B)} \amp = \Pr(A) \Pr(A\cap B) \amp = \Pr(A)\Pr(B) We can take this last equation as a definition of independence.

    Continuing , recall A = \{\text{sum} \geq 10\} and B = \{\text{1st roll is } 6\}. We found that \Pr(A \mid B) \neq \Pr(A), so A and B are not independent.

    Now consider the event C = \{\text{sum} = 7\}. We have: C \amp = \{(1, 6), (2, 5), (3, 4), (4, 3), (5, 2), (6, 1)\} \Pr(C) \amp = \frac{6}{36} = \frac{1}{6} B\cap C \amp = \{(6, 1)\} \Pr(B\cap C) \amp = \frac{1}{36} \text{therefore: } \Pr(C\mid B) \amp = \frac{\Pr(C\cap B)}{\Pr(B)} = \frac{1/36}{1/6} = \frac{1}{6} = \Pr(C) Therefore events B, C are independent.