Files
Fall-2026-Math-1044/source/notes/week01.ptx
T
2026-01-15 20:00:48 -05:00

795 lines
26 KiB
XML

<?xml version="1.0" encoding="UTF-8"?>
<section xml:id="notes-week-01">
<title>Week 1</title>
<introduction>
<p>
This is an outline of the topics we covered in the first week of class.
These notes are <em>not</em> a substitute for your own note-taking.
I highly recommend that you take your own notes during class.
If you ever miss a class for any reason, reach out to another student in class to get a copy of their notes.
</p>
</introduction>
<subsection>
<title>Tuesday 1/13</title>
<subsubsection xml:id="subsubsec-Sets">
<title>Sec 1.1: Sets</title>
<p>
In prior math courses, you mostly asked <term>deterministic</term> questions.
Now, we need new tools to model <term>randomness</term>.
</p>
<example>
<statement>
<p>
A toxin molecule in a cell has a certain chance each minute to leave the cell.
</p>
</statement>
</example>
<example>
<statement>
<p>
A patient takes a diagnostic test for a disease and wants to know the chance that they have the disease based on the test result.
</p>
</statement>
</example>
<definition xml:id="def-sample-space">
<statement>
<p>
The <term>sample space</term>, often denoted <m>\Omega</m>, is the set of all possible results of an experiment.
A single result is called an <term>outcome</term>, while a collection of results is called an <term>event</term>.
</p>
</statement>
</definition>
<example xml:id="example-sample-space">
<statement>
<p>
An experiment consists of rolling a standard 6-sided die (D6).
The sample space is <m>\Omega = \{1, 2, 3, 4, 5, 6\}</m>.
One possible event is <m>A = \{2, 4, 6\}</m>, i.e., the event that the result of the roll is even.
</p>
</statement>
</example>
<p>
Note: we'll use notation like D6 to indicate a 6-sided die with faces 1, 2, 3, 4, 5, 6.
Similarly, for example, D4 will indicate a 4-sided die with faces 1, 2, 3, 4.
</p>
<definition xml:id="def-subset">
<statement>
<p>
The symbol <m>\in</m> means "is an <term>element</term> of", as in <m>4 \in A</m>.
</p>
<p>
The symbol <m>\subset</m> means "is a <term>subset</term> of", as in <m>A \subset \Omega</m>.
This means that every element of the set <m>A</m> is also an element of the set <m>\Omega</m>.
</p>
</statement>
</definition>
<definition xml:id="def-set-operations">
<statement>
<p>
Consider sets <m>A</m> and <m>B</m>, each contained inside <m>\Omega</m>.
We can combine sets in a variety of ways: <dl>
<li>
<title>Union</title>
<p>
The <term>union</term> of <m>A</m> and <m>B</m> is the set <m>A \cup B = \{x \mid x \in A \text{ or } x \in B\}</m>.
</p>
</li>
<li>
<title>Intersection</title>
<p>
The <term>intersection</term> of <m>A</m> and <m>B</m> is the set <m>A \cap B = \{x \mid x \in A \text{ and } x \in B\}</m>.
</p>
</li>
<li>
<title>Difference</title>
<p>
The <term>set difference</term> <m>A-B</m> is the set <m>A - B = \{x \mid x \in A \text{ and } x \notin B\}</m>.
</p>
</li>
<li>
<title>Complement</title>
<p>
The <term>complement</term> of <m>A</m> is the set <m>A^c = \{x \in \Omega \mid x \notin A\}</m>.
</p>
</li>
<li>
<title>Empty Set</title>
<p>
The <term>empty set</term>, usually written <m>\emptyset</m> or <m>\{\}</m>, is the set which contains no elements.
</p>
</li>
</dl>
</p>
</statement>
</definition>
<p>
It's useful sometimes to draw pictures called <term>Venn diagrams</term> representing set interactions.
</p>
<sbsgroup widths="40% 40%">
<sidebyside>
<figure xml:id="fig-Venn-diagram-union">
<caption>Union</caption>
<image>
<description>
<p>
Venn diagram showing sets <m>A, B</m> with the region representing <m>A \cup B</m> shaded.
</p>
</description>
<latex-image>
\begin{tikzpicture}
\def\firstcircle{(180:1.75cm) circle (2.5cm)}
\def\secondcircle{(0:1.75cm) circle (2.5cm)}
\fill[gray!30] \firstcircle;
\fill[gray!30] \secondcircle;
\draw \firstcircle node[text=black,left] {$A$};
\draw \secondcircle node [text=black,right] {$B$};
\end{tikzpicture}
</latex-image>
</image>
</figure>
<figure xml:id="fig-Venn-diagram-intersection">
<caption>Intersection</caption>
<image>
<description>
<p>
Venn diagram showing sets <m>A, B</m> with the region representing <m>A \cap B</m> shaded.
</p>
</description>
<latex-image>
\begin{tikzpicture}
\def\firstcircle{(180:1.75cm) circle (2.5cm)}
\def\secondcircle{(0:1.75cm) circle (2.5cm)}
\fill[white] \firstcircle;
\fill[white] \secondcircle;
\begin{scope}
\clip \firstcircle;
\clip \secondcircle;
\fill[gray!30] \firstcircle;
\end{scope}
\draw \firstcircle node[text=black,left] {$A$};
\draw \secondcircle node [text=black,right] {$B$};
\end{tikzpicture}
</latex-image>
</image>
</figure>
</sidebyside>
<sidebyside>
<figure xml:id="fig-Venn-diagram-difference">
<caption>Difference</caption>
<image>
<description>
<p>
Venn diagram showing sets <m>A, B</m> with the region representing <m>A - B</m> shaded.
</p>
</description>
<latex-image>
\begin{tikzpicture}
\def\firstcircle{(180:1.75cm) circle (2.5cm)}
\def\secondcircle{(0:1.75cm) circle (2.5cm)}
\fill[gray!30] \firstcircle;
\fill[white] \secondcircle;
\draw \firstcircle node[text=black,left] {$A$};
\draw \secondcircle node [text=black,right] {$B$};
\end{tikzpicture}
</latex-image>
</image>
</figure>
<figure xml:id="fig-Venn-diagram-complement">
<caption>Complement</caption>
<image>
<description>
<p>
Venn diagram showing set <m>A \subset \Omega</m> with the region representing <m>A^c</m> shaded.
</p>
</description>
<latex-image>
\begin{tikzpicture}
\def\firstcircle{(0, 0) circle (2.5cm)}
\fill [gray!30] (-4, -3) rectangle (5, 3);
\fill[white] \firstcircle;
\draw \firstcircle node[text=black] {$A$};
\draw (-4, -3) rectangle (5, 3) node [text=black,right] {$\Omega$};
\end{tikzpicture}
</latex-image>
</image>
</figure>
</sidebyside>
</sbsgroup>
</subsubsection>
<subsubsection xml:id="subsubsec-Probability">
<title>Sec 1.2: Probability</title>
<p>
Next, we want to start assigning probabilities to each individual outcome so we can then find the probabilities of events.
</p>
<example>
<statement>
<p>
An experiment consists of rolling a D6.
The sample space is <m>\Omega = \{1, 2, 3, 4, 5, 6\}</m>.
We might assign probabilities as follows:
</p>
<table>
<title>Distribution for a fair die</title>
<tabular halign="center">
<row bottom="minor">
<cell><m>x</m></cell>
<cell><m>\Pr(x)</m></cell>
</row>
<row>
<cell>1</cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell>2</cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell>3</cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell>4</cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell>5</cell>
<cell><m>1/6</m></cell>
</row>
<row>
<cell>6</cell>
<cell><m>1/6</m></cell>
</row>
</tabular>
</table>
<p>
Note that we don't have to assign the same probability to each outcome.
If we do, we call this distribution <term>uniform</term>.
If we have some event, such as <m>A = \{2, 4, 6\}</m>, then we calculate the probability of the event by adding together the probabilities of the outcomes that make up the event:
<md>
<mrow> \Pr(A) = \Pr(2) + \Pr(4) + \Pr(6) = \frac{1}{6} + \frac{1}{6} + \frac{1}{6} = \frac{3}{6} = \frac{1}{2} </mrow>
</md>
Note that this is the same result that we would get if we counted the total number of outcomes in <m>A</m> and divided by the total number of outcomes in <m>\Omega</m>.
</p>
</statement>
</example>
<definition xml:id="def-cardinality">
<statement>
<p>
Given a finite set <m>A</m>, the <term>cardinality</term> of <m>A</m>, written <m>|A|</m>, is the number of elements in <m>A</m>.
</p>
</statement>
</definition>
<fact>
<statement>
<p>
If <m>\Omega</m> is a finite probability space with the uniform distribution and <m>A \subset \Omega</m> is an event, then:
<md>
<mrow> \Pr(A) = \frac{|A|}{|\Omega|} </mrow>
</md>
</p>
</statement>
</fact>
<example>
<statement>
<p>
Suppose we have a weighted die that's much more likely to come up 6 than any other outcome.
</p>
<table>
<title>Distribution for a fair die</title>
<tabular halign="center">
<row bottom="minor">
<cell><m>x</m></cell>
<cell><m>\Pr(x)</m></cell>
</row>
<row>
<cell>1</cell>
<cell>0.1</cell>
</row>
<row>
<cell>2</cell>
<cell>0.1</cell>
</row>
<row>
<cell>3</cell>
<cell>0.1</cell>
</row>
<row>
<cell>4</cell>
<cell>0.2</cell>
</row>
<row>
<cell>5</cell>
<cell>0.1</cell>
</row>
<row>
<cell>6</cell>
<cell>0.4</cell>
</row>
</tabular>
</table>
<p>
With this distribution, if <m>A = \{2, 4, 6\}</m>, then:
<md>
<mrow> \Pr(A) = 0.1 + 0.2 + 0.4 = 0.7 \neq \frac{3}{6}. </mrow>
</md>
</p>
</statement>
</example>
<example>
<statement>
<p>
Pick a number from 1 to 10.
What is the probability of picking 3?
</p>
<p>
Some considerations:
<ol>
<li>
<p>
Are we using the uniform distribution?
</p>
</li>
<li>
<p>
What even is the sample space here? Do we need to pick only integers, or is <m>\pi</m> an outcome here?
</p>
</li>
<li>
<p>
What would "uniform" mean if the sample space was infinite?
</p>
</li>
</ol>
</p>
</statement>
</example>
<definition xml:id="def-probability-distribution">
<statement>
<p>
A <term>probability distribution</term> on a sample space <m>\Omega</m> assigns probabilities to every event, satisfying the following conditions:
<ol>
<li>
<p>
<m>\Pr(\Omega) = 1</m>.
</p>
</li>
<li>
<p>
<m>0 \leq \Pr(A) \leq 1</m> for any event <m>A</m>.
</p>
</li>
<li>
<p>
If <m>A \cap B = \emptyset</m>, then <m>\Pr(A\cup B) = \Pr(A) + \Pr(B)</m>.
</p>
</li>
</ol>
</p>
</statement>
</definition>
<example>
<statement>
<p>
Suppose we flip a coin until we see heads.
The sample space is <m>\Omega = \{H, TH, TTH, TTTH, \dotsc\}</m>.
What would "uniform" mean here? Is there any way we could assign the same probability to each individual outcome here?
</p>
<p>
Instead of "uniform", what if we want to treat the coin as a fair coin? Then what would the distribution be?
</p>
<table>
<title>Distribution assuming a fair coin</title>
<tabular halign="center">
<row bottom="minor">
<cell><m>x</m></cell>
<cell># flips</cell>
<cell><m>\Pr(x)</m></cell>
</row>
<row>
<cell><m>H</m></cell>
<cell>1</cell>
<cell><m>1/2</m></cell>
</row>
<row>
<cell><m>TH</m></cell>
<cell>2</cell>
<cell><m>1/4</m></cell>
</row>
<row>
<cell><m>TTH</m></cell>
<cell>3</cell>
<cell><m>1/8</m></cell>
</row>
<row>
<cell><m>TTTH</m></cell>
<cell>4</cell>
<cell><m>1/16</m></cell>
</row>
<row>
<cell><m>\vdots</m></cell>
<cell></cell>
<cell><m>\vdots</m></cell>
</row>
<row>
<cell><m>T\dotsm TH</m></cell>
<cell><m>n</m></cell>
<cell><m>1/2^n</m></cell>
</row>
<row>
<cell><m>\vdots</m></cell>
<cell></cell>
<cell><m>\vdots</m></cell>
</row>
</tabular>
</table>
</statement>
</example>
<example>
<statement>
<p>
One last (very useful!) observation.
Recall: If <m>A\cap B = \emptyset</m>, then <m>\Pr(A\cup B) = \Pr(A) + \Pr(B)</m>.
</p>
<p>
For any event <m>A</m>, <m>A\cap A^c = \emptyset</m> and <m>A \cup A^c = \Omega</m>.
So:
<md>
<mrow> 1 = \Pr(\underbrace{A \cup A^c}_{\Omega}) = \Pr(A) + \Pr(A^c). </mrow>
</md>
We can rewrite this in two useful ways:
<md>
<mrow> \Pr(A) \amp = 1 - \Pr(A^c) </mrow>
<mrow> \Pr(A^c) \amp = 1 - \Pr(A) </mrow>
</md>
</p>
</statement>
</example>
</subsubsection>
</subsection>
<subsection>
<title>Thursday 1/15</title>
<subsubsection xml:id="subsubsec-Conditional-Probability">
<title>Conditional Probability</title>
<p>
Question: How does evidence (e.g., knowledge of one event occurring) change our knowledge of probabilities for other events?
</p>
<example>
<statement>
<p>
Roll a fair D6 two times.
Let <m>A = \{\text{sum } \geq 10\}</m> and <m>B = \{\text{first roll is } 6\}</m>.
<m>A</m> feels more likely if we already know <m>B</m> has occurred.
</p>
</statement>
</example>
<definition xml:id="def-conditional-probability">
<statement>
<p>
The <term>conditional probability</term> of <m>A</m> given <m>B</m> is: ,
<md>
<mrow> \Pr(A \mid B) = \frac{\Pr(A\cap B)}{\Pr(B)} </mrow>
</md>
</p>
<figure xml:id="fig-conditional-probability">
<caption><m>\Pr(A \mid B)</m> tells the proportion of <m>B</m> which is overlapped by <m>A</m>.</caption>
<image width="50%">
<description>
<p>
Two overlapping circles representing events <m>A</m> and <m>B</m> sit inside a rectangle representing the sample space <m>\Omega</m>.
The circle labeled <m>B</m> is shaded.
The portion of that circle which is overlapped by the <m>A</m> circle is also filled in with slanted lines.
</p>
</description>
<latex-image>
\begin{tikzpicture}
\def\firstcircle{(180:1.75cm) circle (2.5cm)}
\def\secondcircle{(0:1.75cm) circle (2.5cm)}
\fill [gray!30] \secondcircle;
\begin{scope}
\clip \firstcircle;
\clip \secondcircle;
\fill [pattern=north east lines] \firstcircle;
\end{scope}
\draw \firstcircle node[text=black] {$A$};
\draw \secondcircle node[text=black] {$B$};
\draw (-5, -3) rectangle (5, 3) node [text=black,right] {$\Omega$};
\end{tikzpicture}
</latex-image>
</image>
</figure>
</statement>
</definition>
<example xml:id="example-rolls-conditional">
<statement>
<p>
Continuing from the previous example, <m>|\Omega| = 36</m>.
<md>
<mrow> A \amp = \{(4, 6), (5, 5), (5, 6), (6, 4), (6, 5), (6, 6)\} </mrow>
<mrow> B \amp = \{(6, 1), (6, 2), (6, 3), (6, 4), (6, 5), (6, 6)\} </mrow>
<mrow> A \cap B \amp = \{(6, 4), (6, 5), (6, 6)\} </mrow>
</md>
So <m>\Pr(A) = \frac{6}{36}, \Pr(B) = \frac{6}{36}, \text{and } \Pr(A\cap B) = \frac{3}{36}</m>.
Then:
<md>
<mrow> \Pr(A \mid B) = \frac{\Pr(A\cap B)}{\Pr(B)} = \frac{3/36}{6/36} = \frac{3}{6} = \frac{1}{2}. </mrow>
</md>
Notice that <m>\Pr(A\mid B)</m> is significantly larger than <m>\Pr(A)</m>.
</p>
</statement>
</example>
</subsubsection>
<subsubsection xml:id="subsubsec-Diagnostic-Testing">
<title>Diagnostic Testing</title>
<p>
Setup: A patient takes a diagnostic test.
Let <m>P</m> be the event that they test positive.
Let <m>D</m> be the event that they have the disease.
</p>
<definition xml:id="def-sensitivity-specificity">
<statement>
<p>
The <term>sensitivity</term> of a diagnostic test is <m>\Pr(P \mid D)</m>.
The <term>specificity</term> of a diagnostic test is <m>\Pr(P^c \mid D^c)</m>.
</p>
</statement>
</definition>
<p>
But, what the patient really wants to know is <m>\Pr(D \mid P)</m>.
</p>
<example>
<statement>
<p>
A disease has a prevalence of 1%.
A test has sensitivity of 90% and specificity of 91%.
For a patient who gets a positive test result, what is the probability that they have the disease?
</p>
<p>
<md>
<mrow> \text{A) } 9/10 \amp \amp \text{B) } 8/10 \amp \amp \text{C) } 1/10 \amp \amp \text{D) } 1/100 </mrow>
</md>
</p>
</statement>
<answer>
<p>
C!
</p>
</answer>
</example>
<theorem xml:id="thm-Bayes-v1">
<title>Bayes' Theorem (v1)</title>
<statement>
<p>
For events <m>A, B</m> with nonzero probability:
<md>
<mrow> \Pr(B \mid A) = \frac{\Pr(A \mid B)\Pr(B)}{\Pr(A)} </mrow>
</md>
</p>
</statement>
</theorem>
<p>
For example, if a patient sees a positive diagnostic test result, they might try to calculate:
<md>
<mrow> \Pr(D \mid P) = \frac{\Pr(P \mid D)\Pr(D)}{\Pr(P)} </mrow>
</md>
<m>\Pr(P \mid D)</m> is the sensitivity. <m>\Pr(D)</m> could be the prevalence. We don't have direct access to <m>\Pr(P)</m>.
</p>
<p>
Observation: <m>\Omega = D \cup D^c</m>, so <m>P = (P\cap D) \cup (P\cap D^c)</m>.
</p>
<figure xml:id="fig-P-breakdown">
<caption></caption>
<image width="50%">
<description>
<p>
</p>
</description>
<latex-image>
\begin{tikzpicture}
\def\firstcircle{(0, 0) circle (2)}
\def\leftside{(-3, -3) rectangle (-0.5, 3)}
\def\rightside{(-0.5, -3) rectangle (4, 3)}
\begin{scope}
\clip\leftside;
\fill [gray!50] \firstcircle;
\end{scope}
\begin{scope}
\clip\rightside;
\fill [pattern=north east lines] \firstcircle;
\end{scope}
\draw (0, 0) circle (2);
\node at (2.5, 0) {$P$};
\draw (-0.5, 3) to (-0.5, -3);
\node at (-1.5, -3.5) {$D$};
\node at (1.5, -3.5) {$D^c$};
\draw (-3, -3) rectangle (4, 3) node [text=black,right] {$\Omega$};
\end{tikzpicture}
</latex-image>
</image>
</figure>
<p>
Observation 2:
<md>
<mrow> \Pr(P \mid D) \amp \frac{\Pr(P\cap D)}{\Pr(D)} \amp \amp \Rightarrow \amp \Pr(P \cap D) \amp = \Pr(P\mid D)\Pr(D) </mrow>
<mrow> \Pr(P \mid D^c) \amp \frac{\Pr(P\cap D^c)}{\Pr(D^c)} \amp \amp \Rightarrow \amp \Pr(P \cap D^c) \amp = \Pr(P\mid D^c)\Pr(D^c) </mrow>
</md>
So:
<md>
<mrow> \Pr(P) = \Pr(P\mid D)\Pr(D) + \Pr(P\mid D^c)\Pr(D^c) </mrow>
</md>
</p>
<theorem xml:id="thm-Bayes-v2">
<title>Bayes' Theorem (v2)</title>
<statement>
<p>
<md>
<mrow> \Pr(B \mid A) = \frac{\Pr(A \mid B)\Pr(B)}{\Pr(A \mid B)\Pr(B) + \Pr(A \mid B^c)\Pr(B^c)} </mrow>
</md>
</p>
</statement>
</theorem>
<example>
<statement>
<p>
Continuing from the previous example:
<md>
<mrow> \Pr(D\mid P) \amp = \frac{\Pr(P\mid D)\Pr(D)}{\Pr(P\mid D)\Pr(D) + \Pr(P\mid D^c)\Pr(D^c)} </mrow>
<mrow> \amp = \frac{(0.9)(0.01)}{(0.9)(0.01) + (1 - 0.91)(1 - 0.01)} </mrow>
<mrow> \amp \approx 0.092 </mrow>
</md>
What if the patient got a negative test result instead? In that case, what is the probaiblity they do not have the disease?
<md>
<mrow> \Pr(D^c\mid P^c) \amp = \frac{\Pr(P^c\mid D^c)\Pr(D^c)}{\Pr(P^c\mid D^c)\Pr(D^c) + \Pr(P^c\mid D)\Pr(D)} </mrow>
<mrow> \amp = \frac{(0.91)(1 - 0.01)}{(0.91)(1 - 0.01) + (1 - 0.9)(0.01)} </mrow>
<mrow> \amp \approx 0.999 </mrow>
</md>
</p>
</statement>
</example>
</subsubsection>
<subsubsection xml:id="subsubsec-Independent-Events">
<title>Independent Events</title>
<p>
Question: <m>\Pr(A \mid B)</m> is supposed to capture how information about <m>B</m> affects the probability of <m>A</m>.
What if it doesn't?
</p>
<definition xml:id="def-independent-events">
<statement>
<p>
Events <m>A, B</m> are <term>independent</term> if <m>\Pr(A \mid B) = \Pr(A)</m>.
</p>
</statement>
</definition>
<p>
Observation: If <m>A, B</m> have nonzero probability and are independent, then:
<md>
<mrow> \Pr(A\mid B) \amp = \Pr(A) </mrow>
<mrow> \frac{\Pr(A\cap B)}{\Pr(B)} \amp = \Pr(A) </mrow>
<mrow> \Pr(A\cap B) \amp = \Pr(A)\Pr(B) </mrow>
</md>
We can take this last equation as a definition of independence.
</p>
<example xml:id="example-rolls-independent">
<statement>
<p>
Continuing <xref ref="example-rolls-conditional"/>, recall <m>A = \{\text{sum} \geq 10\}</m> and <m>B = \{\text{1st roll is } 6\}</m>.
We found that <m>\Pr(A \mid B) \neq \Pr(A)</m>, so <m>A</m> and <m>B</m> are not independent.
</p>
<p>
Now consider the event <m>C = \{\text{sum} = 7\}</m>.
We have:
<md>
<mrow> C \amp = \{(1, 6), (2, 5), (3, 4), (4, 3), (5, 2), (6, 1)\} </mrow>
<mrow> \Pr(C) \amp = \frac{6}{36} = \frac{1}{6} </mrow>
<mrow> B\cap C \amp = \{(6, 1)\} </mrow>
<mrow> \Pr(B\cap C) \amp = \frac{1}{36} </mrow>
<mrow> \text{therefore: } \Pr(C\mid B) \amp = \frac{\Pr(C\cap B)}{\Pr(B)} = \frac{1/36}{1/6} = \frac{1}{6} = \Pr(C) </mrow>
</md>
Therefore events <m>B, C</m> are independent.
</p>
</statement>
</example>
</subsubsection>
</subsection>
</section>