2.01

Statistical sampling

Populations, censuses and samples 2.01

Definitions
  • Population: the complete set of items, people or measurements under investigation.
  • Census: a survey of every member of the population.
  • Sample: a subset of the population, surveyed in order to represent the whole.
  • Sampling frame: the list of population members from which the sample is actually drawn.
  • Sampling unit: one individual member of that frame.
Key results
  • A census is completely accurate about the population but is usually expensive, slow, and sometimes impossible (for example when testing is destructive).
  • A sample is cheaper and faster but can only ever estimate the population, so every sample carries sampling error.
  • A sample is only as good as its frame: conclusions apply to the population the frame actually describes, not the one intended.
  • Stratified sampling: the number taken from a stratum is ; for 120 office staff out of 200 employees and a sample of 25, .
  • Systematic sampling: from a frame of items take every th item, , starting at a random item among the first ; for and , .
Notes
  • Even a perfectly random method fails to represent the population if the sampling frame is itself incomplete or unrepresentative — for example a phone-number list that omits households without a landline.
  • Non-response is a form of bias in its own right — if the people who don't respond to a survey differ systematically from those who do, the sample skews even with a perfectly designed random selection.

Sampling techniques 2.01

Definitions
  • Simple random sampling: every member of the population has an equal chance of selection, typically by generating random numbers against a numbered list.
  • Systematic sampling: members are selected at a fixed interval through an ordered list, starting from a randomly chosen point.
  • Stratified sampling: the population is divided into groups (strata) and each stratum is sampled in proportion to its size.
  • Opportunity (convenience) sampling: the sample simply consists of whoever is easiest for the researcher to reach.
Method
  1. Number every member of the sampling frame from to .
  2. For a simple random sample of size , generate distinct random numbers in that range and select the corresponding members.
  3. For a systematic sample, compute the interval , choose a random start between and , then take every th member thereafter.
  4. For a stratified sample, compute each stratum's share as , round sensibly to whole people, then sample randomly within each stratum.

In practiceA school has 400 Year 12 and 600 Year 13 students. Take a stratified sample of 50, and a systematic sample of 50 from the full list.

  1. stratified: each year's share
  2. Then choose 20 and 30 students at random within each year, by random numbers.
  3. systematic: every 20th
  4. Choose a random start from 1 to 20, say 7, and take students 7, 27, 47, …, 987.
Notes
  • Each method carries its own risk of bias: opportunity sampling in particular tends to over- or under-represent groups that are more or less accessible to the researcher.
  • Systematic sampling becomes biased if the list has a periodic pattern whose period matches the sampling interval.
  • Stratified sampling is the method of choice when the population contains groups expected to behave differently, since it guarantees each group appears in proportion.
  • Within each stratum, members are still chosen by simple random sampling — stratified sampling controls which groups appear and in what proportion, not how individuals within a group are picked.

Worked example

Worked example

A college of students is divided into three faculties: Arts ( students), Science () and Humanities ().

Describe how to select a stratified sample of students, and explain one weakness of instead surveying the first students to arrive one morning.

Show worked solution

Each faculty contributes in proportion to its size.

Arts:

students.

Science:

students.

Humanities:

students.

These total:

as required.

Within each faculty, number the students from upwards and use random numbers to select the required count, so that selection inside each stratum is simple random.

Surveying the first arrivals is opportunity sampling: it over-represents students who arrive early — those living nearby or with first-period classes — so the sample is biased and the faculty proportions are left to chance rather than controlled.

2.02

Measures of location and spread

Measures of location 2.02

Definitions
  • Mean: , or for data given in a frequency table.
  • Median: the middle value once the data is placed in order.
  • Mode: the most frequently occurring value.
Key results
  • The mean uses every value in the data set, which makes it efficient but sensitive to outliers.
  • The median is resistant to extreme values, so it is preferred for strongly skewed data such as incomes or house prices.
  • The mode is the only measure of location available for purely categorical data, and is the least informative for continuous data.
  • For grouped data the mean can only be estimated, by treating every value in a class as sitting at that class's midpoint.
  • Example: for : , median , mode .
  • Grouped data: classes () and () have midpoints 5 and 20, so .
  • Coding: if then .
Notes
  • Choose the measure that suits the data, not the one that is easiest to compute: quoting a mean for a heavily skewed data set is technically correct but misleading.
  • In a symmetric distribution the mean and median coincide; a mean noticeably larger than the median signals positive skew (a long right tail).
  • For a frequency table, always divide by (the total frequency), not by the number of distinct rows in the table — a common error when the data is grouped rather than listed individually.

Measures of spread 2.02

Definitions
  • Range: largest value minus smallest value.
  • Interquartile range: , the spread of the middle of the data.
  • Variance: the mean squared deviation from the mean, .
  • Standard deviation: .
Key results
  • Expanding the bracket gives the computational form , which is far quicker when the data is supplied as the summary statistics and .
  • Standard deviation returns the spread to the original units of the data, making it directly comparable to the mean; variance, in squared units, is harder to interpret physically.
  • The IQR ignores the extreme quarter of the data at each end, so it is the natural companion to the median just as standard deviation is to the mean.
Notes
  • The range depends on only two values and is destroyed by a single outlier, so it is the weakest measure of spread.
  • A common error is to compute instead of — the mean must be squared before subtracting.
  • Variance can never be negative — a negative result from signals an arithmetic error, most often in computing .

Quartiles, the five-number summary and outliers 2.02

Cumulative frequency and quartiles (Data presentation and interpretation)
Definitions
  • Five-number summary: minimum, lower quartile , median , upper quartile , maximum.
  • Outlier: a value unusually far from the rest of the data.
Key results
  • A standard outlier rule flags any value more than beyond the nearest quartile: below or above .
  • An alternative rule sometimes used flags values more than two standard deviations from the mean; a question will always state which rule to apply.
Method
  1. Order the data and find the median.
  2. Find as the median of the values below the overall median, and as the median of the values above it.
  3. Compute and the two fences and .
  4. List any data values lying outside those fences.

In practiceFind the quartiles and any outliers in 3, 5, 7, 8, 9, 11, 12, 14, 30.

  1. ordered; the median is 9
  2. medians of each half
  3. the fences
  4. 30 lies above the upper fence, so it is an outlier.
Notes
  • Outliers should be investigated rather than automatically discarded, since a genuine outlier can be the most informative point in a data set.
  • Only discard a value when there is a substantive reason to believe it is an error, such as a mis-typed reading or a broken instrument — and say so explicitly.
  • For grouped data, quartiles are estimated by linear interpolation, or read off a cumulative frequency graph at , and .
  • State clearly which outlier rule was used before applying it — a value can be flagged by the rule but not by the two-standard-deviations rule, or vice versa.

Worked examples

Worked example

Find the mean, variance and standard deviation of the data set .

Show worked solution

Mean:

Deviations from the mean are , with squares , summing to .

Variance:

Standard deviation:

(3 s.f.).

Worked example

For the data:

find the median and quartiles, test for outliers using the rule, and compare the mean with the median.

Show worked solution

There are ordered values, so the median is the th: .

The five values below it are , whose median is ; the five above it are:

whose median is .

So:

and:

The fences are and .

Only lies outside, so is an outlier.

The total is , so:

(3 s.f.), noticeably above the median of : the single large value pulls the mean up but leaves the median untouched, which is exactly why the median is preferred here.

Worked example

A sample of values has and .

(a) Find the mean and standard deviation.

(b) A twenty-first value, , is then added.

Find the new mean and standard deviation.

Show worked solution

(a):

Using the computational form,

so:

(3 s.f.).

(b) The new totals are:

and:

with .

New mean:

(3 s.f.).

New variance:

so:

(3 s.f.).

The extra value sits above the old mean, so both the mean and the spread increase slightly.

2.02

Data presentation

Histograms 2.02

Histograms (Data presentation and interpretation)
Definitions
  • Histogram: a display for continuous grouped data in which the area of each bar is proportional to the frequency of its class.
  • Frequency density: , plotted on the vertical axis.
Key results
  • When class widths are unequal it is bar area, not bar height, that represents frequency — this is what distinguishes a histogram from a bar chart.
  • The number of values in part of a class is estimated by assuming the data is spread uniformly across that class, so a portion of the class contributes a proportional share of its frequency.
  • Example: the class with frequency 30 has width 15, so its frequency density is .
  • Part of a class: the number of values with is estimated as .
Notes
  • Bar charts are for categorical data and have gaps between bars; histograms are for continuous data and have none.
  • Label the vertical axis 'frequency density', not 'frequency' — marks are routinely lost for this.
  • To read off a frequency from a histogram bar, multiply frequency density by class width — never read the bar's height alone as if it were the frequency.

Box plots and comparing distributions 2.02

Box plots, outliers and comparing distributions (Data presentation and interpretation)
Definitions
  • Box plot: a diagram of the five-number summary, with the box spanning to , a line at the median, and whiskers to the extreme values that are not outliers.
Key results
  • Two box plots drawn on the same scale compare distributions at a glance: box position compares location, box width compares spread, and the position of the median line within the box indicates skew.
  • Outliers are conventionally plotted as separate crosses beyond the whiskers rather than being absorbed into them.
  • , the spread of the middle 50% of the data.
  • Outlier rule (when the question gives it): below or above . For and , IQR and the limits are and .
Notes
  • A comparison must be written in context and must address both location and spread — 'the medians differ' alone is not a comparison of two distributions.
  • A long whisker on one side combined with the median sitting off-centre in the box both point the same way — use them together as corroborating evidence of skew, not just one alone.

Worked example

Worked example

The time minutes taken by people to complete a task is grouped as: , frequency ; , frequency ; , frequency ; , frequency .

(a) Find the frequency densities.

(b) Estimate how many people took between and minutes.

(c) Estimate the mean time.

Show worked solution

(a) Frequency density is frequency divided by class width:

(b) The interval sits inside the class , covering of its minutes; assuming a uniform spread, the estimate is:

people.

(c) Using midpoints :

so the estimated mean is:

minutes (3 s.f.).

2.02

Correlation and regression

Scatter diagrams and correlation 2.02

Scatter diagrams and correlation (Data presentation and interpretation)
Definitions
  • Bivariate data: paired observations on the same individuals.
  • Product moment correlation coefficient : a measure of the strength and direction of the linear relationship between two variables.
Key results
  • ranges from (perfect negative linear correlation) through (no linear relationship) to (perfect positive linear correlation).
  • A scatter diagram is the correct first display for bivariate data, and shows at once whether a linear model is even plausible.
Notes
  • measures linear association only: a strong non-linear relationship, such as points lying neatly on a parabola, can still give a value of close to zero.
  • Always sketch or inspect the scatter diagram before quoting , since a single outlier can move substantially.
  • Correlation observed in a sample is not, by itself, evidence of correlation in the population — that requires the hypothesis test of the following topic.
  • Correlation, however strong, never proves causation — a third, unmeasured variable can drive both quantities without either directly causing the other.

Regression lines and prediction 2.02

Regression lines and prediction (Data presentation and interpretation)
Definitions
  • Least-squares regression line of on : the straight line that minimises the sum of the squared vertical distances between the line and the data points.
  • Explanatory (independent) variable ; response (dependent) variable .
Key results
  • The gradient is interpreted in context as the change in per unit increase in ; the intercept is the predicted value of when , which is often physically meaningless.
  • Interpolation — estimating for a value of inside the range of the original data — is reasonably reliable.
  • Extrapolation — estimating outside that range — is far less reliable, since there is no evidence the same linear relationship continues to hold there.
Notes
  • The regression line of on is designed to predict from ; using it backwards to predict from is not valid, because it minimises vertical distances only.
  • A regression line can always be calculated, even for data with no real linear relationship — check and the scatter diagram before trusting any prediction it gives.
  • Substitute a prediction back into context and sanity-check the units and the rough size of the answer — a numerically 'correct' prediction that is physically absurd (e.g. a negative height) signals the model has been pushed too far.

Worked example

Worked example

For students, is hours of revision (ranging from to ) and is a test score out of .

The regression line is:

and .

(a) Interpret the gradient.

(b) Estimate the score of a student who revises for hours.

(c) Comment on using the line to estimate the score after hours of revision.

(d) Does the value of show that revision causes higher scores?

Show worked solution

(a) The gradient means each additional hour of revision is associated with an increase of about marks, on average, across this sample.

(b) Substituting :

so about marks.

This is interpolation ( lies inside the data range to ), so it is reasonably reliable.

(c) Substituting gives:

which exceeds the maximum possible mark of .

This is extrapolation far outside the data range, where there is no evidence the linear relationship still holds — the estimate is meaningless.

(d) No.

indicates strong positive linear association only; the test cannot rule out a third factor, such as general motivation, driving both variables.