2.05

The language and logic of a test

Hypotheses, tails and the significance level 2.05

Tails and the significance level (Statistical hypothesis testing)
Definitions
  • Null hypothesis : a specific claimed value for a population parameter, assumed true for the purpose of the calculation.
  • Alternative hypothesis : the statement of how the true value might differ from that claim.
  • One-tailed test: specifies a direction, such as or .
  • Two-tailed test: specifies only that the parameter differs, such as .
  • Significance level: the threshold probability, fixed in advance, for how unlikely the observed data must be under before is rejected.
Key results
  • The significance level, commonly , is also the probability of rejecting when it is in fact true, so lowering it makes false alarms rarer but genuine effects harder to detect.
  • In a two-tailed test at the level the significance is split between the tails, giving in each.
Notes
  • The significance level must be chosen before looking at the data; selecting it afterwards to make a result 'significant' invalidates the test.
  • The direction of a one-tailed test must also come from the context stated before the data is seen, not from the direction the data happens to point.
  • A two-tailed test is the more cautious default when there is no prior reason to expect the effect in a particular direction — reserve a one-tailed test for when the context genuinely rules out the other direction in advance.

The logic of a test and how to state a conclusion 2.05

Definitions
  • Test statistic: the quantity computed from the sample whose distribution is known under .
  • Critical region: the set of values of the test statistic for which is rejected.
Key results
  • The two possible outcomes are: reject in favour of (the data would have been sufficiently unlikely under ), or conclude there is insufficient evidence to reject .
  • The two comparison routes are equivalent: comparing the probability of the observed result with the significance level, or comparing the test statistic with the critical value.
Method
  1. Define the parameter being tested in words, in context.
  2. State and , and the significance level.
  3. State the distribution of the test statistic assuming is true.
  4. Compute either the probability of an outcome at least as extreme as the one observed, or the value of the test statistic.
  5. Compare with the significance level or the critical value.
  6. Write the conclusion in context, referring back to the original claim.

In practiceA coin is suspected of bias towards heads. In 20 throws it shows 15 heads. Test at the 5% level.

  1. one-tailed, 5%
  2. at least as extreme as observed
  3. Reject : there is sufficient evidence at the 5% level that the coin is biased towards heads.
Notes
  • Never write that the test 'accepts' or 'proves' , since failing to find evidence against a claim is not the same as proving the claim true.
  • A conclusion stated only as 'reject ' is incomplete: it must be translated back into the language of the original context.
  • Statistical significance is not practical importance — a very large sample can make a tiny, meaningless difference significant.
  • A test that fails to reject has not shown is true — it has only shown the sample doesn't give enough evidence against it; a larger sample might yet reject it.
2.05

Testing a binomial proportion

Hypothesis testing for a binomial proportion 2.05

Key results
  • Under the number of successes follows , and the test finds the probability of an outcome at least as extreme as the one observed.
  • For an upper-tail test () compute ; for a lower-tail test () compute .
  • If that probability is less than the significance level, is rejected.
Method
  1. Define as the population proportion in context and state against the appropriate .
  2. State that under .
  3. Compute the tail probability for the observed value, using cumulative binomial tables or a calculator.
  4. Compare with the significance level (or with half of it, in each tail of a two-tailed test).
  5. Conclude in context.

In practice30% of voters are claimed to support a party. In a random sample of 25, 3 do. Test at the 5% level whether the claim is wrong.

  1. two-tailed
  2. compare with half the level, in the lower tail
  3. Do not reject : insufficient evidence at the 5% level that the proportion differs from 30%.
Notes
  • For small samples, exact binomial probabilities are used rather than a normal approximation, since the conditions for that approximation ( and ) may not hold — always check these conditions before deciding which method to use.
  • The tail probability must include the observed value itself, not just the outcomes beyond it: 'at least as extreme as' means , not .
  • State as a population proportion in the actual context ('the proportion of all such patients who improve'), not as an abstract symbol — a mark is routinely reserved for this contextual definition.

Critical regions and actual significance level 2.05

Critical regions and actual significance level (Statistical hypothesis testing)
Definitions
  • Actual significance level: the true probability that the test statistic falls in the critical region when is true.
Key results
  • Because a binomial variable is discrete, the critical region cannot usually be made to carry exactly of the probability, so its actual significance level is somewhat less than the nominal level.
  • The critical region is the largest tail whose total probability under still does not exceed the significance level.
Method
  1. Tabulate cumulative probabilities under near the relevant tail.
  2. For an upper tail, find the smallest with the significance level; the critical region is .
  3. Quote the actual significance level as that probability , expressed as a percentage.
  4. Check the neighbouring value to confirm that including it would push the probability above the significance level.

In practiceUnder , , with at the 5% level. Find the critical region.

  1. upper-tail probabilities
  2. the smallest with
  3. including 9 would push it to 11.33%
Notes
  • Once the critical region is known, testing any observed value is immediate — this is why critical regions are computed in advance for repeated quality-control checks.
  • Always check the value just outside the critical region as well as the boundary itself — this confirms the region really is the largest one still under the significance level, not merely a plausible-looking guess.

Worked examples

Worked example

A coin is tossed times and lands heads times.

Test, at the significance level, whether the coin is biased towards heads.

Show worked solution

Let be the probability that the coin lands heads.

, (one-tailed, since bias towards heads is specified).

Under , , where is the number of heads.

The probability of a result at least as extreme as the one observed is:

Since:

the result lies in the critical region, so is rejected.

There is significant evidence at the level that the coin is biased towards heads.

Worked example

A standard treatment is known to succeed for of patients.

In a trial of patients given a new treatment, only improve.

Test, at the significance level, whether the new treatment has a lower success rate.

Show worked solution

Let be the probability that a patient improves under the new treatment.

, (one-tailed, lower).

Under , .

The lower-tail probability is:

(from cumulative binomial tables).

Since:

is rejected.

There is significant evidence at the level that the new treatment has a lower success rate than the standard one.

Note that a normal approximation would be inappropriate here: although:

the sample is small and the exact binomial is available.

Worked example

It is claimed that of customers at a shop pay in cash.

A random sample of customers is taken to test against at the significance level.

Find the critical region and the actual significance level of the test.

Show worked solution

Under , .

Using cumulative binomial tables,

and:

For an upper-tail critical region : taking gives:

which exceeds , so is not in the critical region.

Taking gives:

which is less than .

So the critical region is , and the actual significance level is , or .

It falls short of the nominal because is discrete: no critical region of this form has probability exactly .

2.05

Testing for correlation

Hypothesis testing for correlation 2.05

Key results
  • The test asks whether an observed sample correlation coefficient provides evidence of a genuine linear correlation in the underlying population, rather than arising purely by chance in this particular sample.
  • The hypotheses are stated in terms of the population correlation coefficient : against , or .
  • The sample value is compared against a critical value read from tables, which depends on the sample size and the chosen significance level.
  • Reject when exceeds the critical value; larger samples have smaller critical values, so a modest can be significant if is large.
Notes
  • For a two-tailed test, use the critical value listed for half the significance level in each tail.
  • A significant result says only that a linear association exists in the population; it says nothing about how strong or how useful that association is.
  • Larger samples make smaller values of significant, so a 'significant' correlation from a large data set can still correspond to a fairly weak linear relationship in practical terms.

Correlation and causation 2.05

Key results
  • A statistically significant correlation is evidence of a linear association, not evidence of causation.
  • Two variables can be strongly correlated because one causes the other, because both are driven by a third factor, or by coincidence, and a hypothesis test for correlation cannot itself distinguish between these explanations.
Notes
  • Distinguishing causation requires a controlled experiment with randomised allocation, not a larger observational sample.
  • When asked to comment on a significant correlation, name a plausible third factor for the context given — that is what the mark is for.
  • 'Reverse causation' is worth considering too — for two correlated variables and , it's possible influences rather than the assumed direction influences .

Worked example

Worked example

For a random sample of pairs of observations, the product moment correlation coefficient is .

Test at the significance level whether there is positive correlation in the population, and state whether the conclusion would change at the level.

Show worked solution

Let be the population product moment correlation coefficient.

, (one-tailed).

For at the one-tailed level, the critical value from tables is .

Since:

is rejected: there is significant evidence at the level of positive linear correlation in the population.

At the one-tailed level the critical value for is , and:

so at that level there would be insufficient evidence to reject .

In either case the test shows association, not causation.

2.05

Testing a normal mean

The distribution of the sample mean 2.05

The distribution of the sample mean (Statistical hypothesis testing)
Key results
  • If and a random sample of size is taken, then the sample mean satisfies .
  • The standard deviation of the sample mean is therefore , which shrinks as grows — quadrupling the sample size halves the spread of the sample mean.
  • This is why larger samples give more precise estimates, and why the same observed difference from becomes more significant as increases.
Notes
  • Note the divisor: it is the variance that is divided by , so the standard deviation is divided by , not by .
  • Don't confuse (the spread of individual measurements) with (the spread of the sample mean) — they answer different questions and are never interchangeable in a formula.

Hypothesis testing for a normal mean 2.05

Hypothesis test for a normal mean (Statistical hypothesis testing)
Key results
  • Testing a claimed population mean with known population standard deviation uses under .
  • The test statistic is , compared against a critical value from the standard normal distribution.
  • A two-tailed test at the significance level uses critical values ; if the calculated falls beyond either of these, is rejected in favour of .
  • Other standard critical values: for a one-tailed test, for a one-tailed test, and for a two-tailed test.
Method
  1. State and the appropriate , with the significance level.
  2. State under .
  3. Compute .
  4. Compare with the critical value, or equivalently compare the tail probability with the significance level.
  5. Conclude in context.

In practiceMasses are . Test against at the 5% level, given and .

  1. two-tailed: 2.5% in each tail
  2. Reject : there is evidence at the 5% level that the mean is not 40.
Notes
  • The test requires the population standard deviation to be known and the underlying variable to be normally distributed (or the sample to be large); if is estimated from the sample, a different test is needed.
  • The two routes agree exactly, because the critical value is simply the whose tail probability equals the significance level — quoting both is a useful self-check.
  • For a two-tailed test, always compare (or the observed mean's absolute distance from ) with the positive critical value — comparing a negative directly against a positive critical value gives a false 'not significant' result.

Worked examples

Worked example

A machine is set to fill bottles with a mean volume of ml; the population standard deviation is known to be ml.

A random sample of bottles has a mean volume of ml.

Test, at the significance level, whether the mean volume has changed.

Show worked solution

, (two-tailed).

Test statistic:

The two-tailed critical values are .

Since:

the test statistic lies in the critical region, so is rejected: there is significant evidence at the level that the mean volume has changed.

Worked example

Packets are filled to a claimed mean of g with known population standard deviation g.

A sample of packets has mean g.

Test at the level whether the mean has decreased, and comment on what would happen at the level.

Show worked solution

, (one-tailed).

Under ,

so the standard deviation of the sample mean is:

Test statistic:

The one-tailed critical value is , and:

so is rejected: there is significant evidence at the level that the mean has decreased.

Equivalently, the tail probability is:

giving the same conclusion.

At the level the critical value is , and does not reach it (equivalently:

), so there would be insufficient evidence to reject — the same data can be significant at one level and not at another.