Probability
Events, sample spaces and Venn diagrams 2.03
- Sample space: the set of all possible outcomes of an experiment.
- Event: a subset of the sample space; is the probability that event occurs.
- Complement A': the event that does not occur.
- : both and occur. : at least one of and occurs.
- , and the probabilities of all distinct outcomes in a sample space sum to .
- P(A')=1-P(A) — often the fastest route to 'at least one' probabilities.
- A Venn diagram splits the sample space into the four disjoint regions , A\cap B', A'\cap B and A'\cap B', whose probabilities must total .
- Fill a Venn diagram from the intersection outwards: put in the overlap first, then subtract it to get the 'only ' and 'only ' regions, then use the total of for the outside region.
- P(A')=1-P(A) is especially useful for 'at least one' questions, where directly listing every favourable outcome would need many more cases than its single complementary event.
The addition rule and mutually exclusive events 2.03
- Mutually exclusive events: events that cannot occur together, so .
- Addition rule: , where subtracting the intersection corrects for the outcomes counted twice when and overlap.
- For mutually exclusive events this simplifies to .
- Mutually exclusive and independent are different, and in fact almost opposite, ideas: if two events with non-zero probability are mutually exclusive then knowing one occurred tells you the other did not, so they cannot be independent.
- Subtracting is only needed when and can overlap — for genuinely mutually exclusive events that term is already zero, so both forms of the rule agree.
Conditional probability and tree diagrams 2.03
- Conditional probability : the probability that occurs given that has occurred.
- , describing how the probability of changes once it is known that has occurred.
- Rearranging gives the multiplication rule , which is exactly what multiplying along a branch of a tree diagram does.
- Conditioning restricts attention to the part of the sample space where holds, which is why the denominator is rather than .
- Draw the tree with the first-stage events as the initial branches and the second-stage events branching from each.
- Write conditional probabilities on the second-stage branches.
- Multiply along a path to obtain the probability of that combination of outcomes.
- Add the probabilities of all paths that satisfy the required event.
- For a reversed conditional probability, divide the probability of the single relevant path by the total probability of the conditioning event.
In practice1% of people have a condition. A test is positive for 95% of those with it and 2% of those without. Find and .
- multiply along the path
- \displaystyle P(C'\cap+)=0.99\times0.02=0.0198
- add the paths
- most positives are false positives
- The probabilities on the branches leaving any one node must sum to ; this is the quickest check that a tree has been drawn correctly.
- and are different quantities and are only equal when — confusing them is the single most common error in conditional probability.
- 'Without replacement' problems are the classic use of tree diagrams with changing conditional probabilities — the second-stage branches must be updated to reflect what was removed at the first stage.
Independence 2.03
- Independent events: events for which the occurrence of one does not change the probability of the other.
- Two events are independent exactly when .
- Equivalently, and are independent when .
- Independence is a condition that must be tested, not assumed, since many pairs of events that sound unrelated are not actually independent in a strict probabilistic sense.
- In an exam, 'show that and are independent' means compute both and and state that they are equal — asserting independence from context earns nothing.
- Mutually exclusive events with non-zero probabilities are never independent — checking for independence when the events are already known to be mutually exclusive is a wasted step.
Worked examples
Worked example
Events and satisfy , and:
(a) Find .
(b) Determine whether and are independent.
(c) Find and .
Show worked solution
(a) Rearranging the addition rule,
(b) Test the condition:
but:
Since , the events are not independent. (c):
and:
(3 s.f.).
Note:
confirming the dependence found in (b), and note also that and are quite different numbers.
Worked example
A factory makes a component on two machines.
Machine A makes of output, of which is defective; machine B makes the other , of which is defective.
A component is chosen at random.
(a) Find the probability it is defective.
(b) Given that it is defective, find the probability it came from machine A.
Show worked solution
(a) Multiply along each branch of the tree and add:
and:
so:
(b) Use the conditional probability formula with the single relevant path on top:
(3 s.f.).
Although machine A makes of all components, it accounts for under half of the defective ones, because its defect rate is the lower of the two.
The binomial distribution
The binomial distribution: model and conditions 2.04
- : is the number of successes in trials, each independently a success with probability .
- The four conditions for a binomial model are: a fixed number of trials ; exactly two possible outcomes per trial; a constant probability of success ; and independence between trials.
- The binomial coefficient counts the number of orders in which successes can occur among trials, which is why it appears as a multiplier in the probability formula.
- The conditions should be checked explicitly before applying the model: a scenario that violates one of them — for example sampling without replacement from a small population, where changes after each selection — is not truly binomial.
- Sampling without replacement from a very large population is close enough to binomial to model as such, since barely moves; state this assumption when you rely on it.
- 'Two possible outcomes per trial' doesn't mean the underlying scenario has only two outcomes overall — a die roll has six, but 'success' can be defined as one specific outcome (or group of outcomes) versus everything else.
Binomial calculations, mean and variance 2.04
- for .
- Cumulative probabilities come from tables or a calculator: is tabulated directly, , and .
- Mean: . Variance: , and standard deviation .
- The mean and variance are derivable from the definition but are worth knowing directly for quick estimation and for sanity-checking a calculated probability against where the bulk of the distribution actually lies.
- Because is discrete, the inequalities matter: and differ by the whole of .
- Translate the wording carefully — 'at least ' is , 'more than ' is , 'at most ' is , and 'fewer than ' is .
- need not be a whole number even though itself only takes whole-number values — the mean of the distribution is an average over many repeats, not a value can actually take.
Worked examples
Worked example
Given , find .
Show worked solution
(3 s.f.).
Worked example
A machine produces items independently, each faulty with probability .
A sample of items is taken and is the number of faulty items.
(a) State the distribution of and find its mean and standard deviation.
(b) Find .
(c) Find .
Show worked solution
(a):
since there is a fixed number of independent trials with constant probability and two outcomes per trial.
Mean:
Variance:
so the standard deviation is:
(3 s.f.). (b):
Here:
Adding:
so:
(3 s.f.). (c):
(3 s.f.).
The normal distribution
The normal distribution 2.04
- : a continuous, symmetric, bell-shaped distribution fully described by its mean and variance .
- The curve is symmetric about , so , and the total area beneath it is .
- Roughly of the distribution lies within standard deviation of the mean, within , and within .
- Because the variable is continuous, for any single value , so and are identical.
- No simple closed-form antiderivative exists for the probability density function, so probabilities are found from standardised tables or a calculator's built-in normal-distribution function.
- The second parameter in is the variance, not the standard deviation: has standard deviation and variance .
- The 68/95/99.7 figures are worth keeping as an instant sanity check on a calculated probability — a computed answer wildly inconsistent with them usually signals a standardising error.
Standardising and using normal tables 2.04
- Standard normal variable: .
- : the tabulated cumulative probability .
- .
- By symmetry, , so tables need only list positive .
- , and for a symmetric interval .
- Standardising converts any normal-distribution question into a question about the standard normal distribution, which is why only one table (or one calculator function) is needed for every possible normal distribution, regardless of its particular mean and variance.
- Write down the distribution and the probability required, in symbols.
- Sketch the bell curve and shade the region asked for — this alone prevents most sign errors.
- Standardise each boundary value using .
- Read from tables and combine using the complement, symmetry or difference rules as the sketch dictates.
- For an inverse problem (given a probability, find the value), read from the percentage-points table and then solve .
In practice. Find , and the value with .
- standardise both ends
- percentage points table
- The commonest error is quoting when the shaded region is the upper tail: check the sketch before writing the final line.
- A negative -value is completely normal and expected whenever the boundary lies below the mean — there is no need to treat a negative standardised value as a sign of an error.
Approximating the binomial by the normal 2.04
- For large , can be approximated by — matching the mean and variance of the binomial exactly.
- The approximation is generally considered reasonable when and , which is also when the binomial is close to symmetric.
- A continuity correction of is applied to account for approximating a discrete distribution with a continuous one.
- Check that and .
- State the approximating distribution .
- Apply the continuity correction: becomes , becomes , and becomes .
- Standardise and read the probability from tables.
In practice. Use a normal approximation to find .
- both above 5
- variance
- continuity correction
- the exact binomial value is 0.0445
- The approximation exists because exact binomial probabilities become computationally awkward for large ; with a modern calculator it is often the continuity correction itself, rather than the arithmetic, that is being examined.
- Getting the direction of the shift wrong reverses the correction and typically changes the answer in the third significant figure or worse — always ask which whole numbers must remain inside the region.
- Sketching the discrete probabilities as a bar chart with the continuous normal curve overlaid makes the direction of each continuity correction visually obvious, rather than something to recall as a rule.
Worked examples
Worked example
Given , find .
Show worked solution
Standardise:
From the standard normal distribution,
(3 s.f.).
Worked example
A fair coin is tossed times.
Using a suitable approximation, estimate the probability of obtaining at least heads.
Show worked solution
Let be the number of heads, so:
Check the conditions:
and:
so the normal approximation is valid.
The approximating distribution is:
with .
Apply the continuity correction: becomes:
keeping the whole of inside the region.
Standardise:
Then:
(3 s.f.).
For comparison, the exact binomial value is , so the approximation is good to three decimal places.
Per disputationem veritatem quaerimus