5.03

Probability density functions

Probability density functions 5.03

Probability density functions (Continuous random variables)
Definitions
  • Probability density function (PDF): a function with and , such that .
  • Mode: the value of at which is greatest.
Key results
  • and .
  • .
  • for any single value, so and give the same probabilities.
Method
  1. Find an unknown constant from total area 1, over the whole range where is non-zero.
  2. Sketch first. Symmetry often gives at once, and the sketch locates the mode, which may be at an end point.

In practice for . Find , and .

  1. total area 1
  2. the PDF is symmetric about
Notes
  • is a density, not a probability: it can exceed 1, as on the sheet where .
  • For a piecewise PDF, integrate each piece over its own interval and add.

Worked example

Worked example

for and 0 otherwise.

Find , , and .

Show worked solution

so:

so:

5.03

Cumulative distribution functions

Cumulative distribution functions, medians and quartiles 5.03

The cumulative distribution function (Continuous random variables)
Definitions
  • Cumulative distribution function (CDF): .
  • Median : . Lower and upper quartiles: and .
Key results
  • f(x)=F'(x) wherever is differentiable.
  • is continuous and non-decreasing, from 0 to 1. .
Method
  1. To find from a piecewise , integrate piece by piece, adding the total probability of the earlier pieces as the constant, so that is continuous at each join.
  2. Define for all real , including below the range and above it.

In practice for . Find , the median and the mean.

  1. define for every real
  2. reject : outside the range of
  3. mean below the median: negative skew
Notes
  • When solving produces more than one root, keep only the one inside the range of .
  • Comparing mean and median indicates skew. On the sheet, median mean : negative skew.

Worked example

Worked example

for , with below and above.

Find and the median.

Show worked solution
f(x)=F'(x)=\tfrac14(x+1)

for .

Median:

so and .

Only:

lies in .

5.03

The exponential and continuous uniform distributions

The exponential distribution 5.03

The exponential distribution (Continuous random variables)
Definitions
  • Exponential distribution : for .
Key results
  • , so .
  • and ; the median is .
  • Memoryless property: .
  • If events occur as a Poisson process at rate , the time between successive events is .
Notes
  • The memoryless property makes it a model for components that do not wear out: a bulb that has lasted 1000 hours is as good as new. Where wear matters, it is a poor model.
  • Watch the parameter: a mean lifetime of 200 hours means , not 200.
  • Why it is memoryless: .

The continuous uniform distribution 5.03

The continuous uniform distribution (Continuous random variables)
Definitions
  • Continuous uniform (rectangular) distribution : for .
Key results
  • and .
  • for .
Notes
  • Rounding errors are uniform: a length rounded to the nearest cm has error .
  • Probabilities are lengths of intervals as fractions of ; no integration is needed.
  • Where the variance comes from: , so .

Worked examples

Worked example

Bulbs have exponentially distributed lifetimes with mean 200 hours.

Find the probability that a bulb lasts more than 300 hours, and the probability that a bulb that has lasted 200 hours lasts at least another 300.

Show worked solution

so:

By the memoryless property the second probability is the same, 0.223: the bulb's age is irrelevant under this model.

Worked example

Lengths are rounded to the nearest centimetre.

Find the variance of the rounding error and the probability that the error is less than 0.2 cm in size.

Show worked solution

The error is , with variance:

5.07

The Wilcoxon signed-rank test

The Wilcoxon signed-rank test 5.07

The Wilcoxon signed-rank test (Non-parametric tests)
Definitions
  • Non-parametric test: a test that makes no assumption that the population follows a particular distribution, such as the normal; hypotheses are about the median.
  • Wilcoxon signed-rank test: a test for the median of a single population, or for paired data, using the ranks of the sizes of the differences from the hypothesised median.
Key results
  • = sum of the ranks of positive differences, = sum of the ranks of negative differences, with .
  • Test statistic : for a two-tailed test, the smaller of and ; for a one-tailed test, the sum that predicts will be small. Reject if is less than or equal to the critical value.
  • For large : , used with a continuity correction.
  • Paired samples: apply the test to the differences within each pair, with : median difference .
Method
  1. Subtract the hypothesised median from every value; discard zeros and reduce accordingly.
  2. Rank the absolute differences from smallest to largest, giving tied values the mean of the ranks they occupy, then attach the signs.
  3. Find and compare with the tabulated critical value for and the significance level.

In practiceData: 23, 18, 26, 21, 29, 25, 17, 24. Test : median against : median at the 5% level.

  1. no zeros, so
  2. tied sizes share the mean of their ranks
  3. check:
  4. predicts small negative ranks; the critical value for is 5
  5. Do not reject at the 5% level, although the result is close.
Notes
  • The test assumes the population is symmetric about its median; it uses both the signs and the sizes of the differences.
  • Unlike most tests met so far, a small statistic is the significant one.

Worked example

Worked example

Ten values are 53, 47, 58, 61, 49, 56, 52, 63, 45, 57.

Use a Wilcoxon signed-rank test at the 5% level to test whether the population median exceeds 50.

Show worked solution

Differences from 50: 3, −3, 8, 11, −1, 6, 2, 13, −5, 7.

Ranks of : 1→1, 2→2, the two 3s share 3.5, 5→5, 6→6, 7→7, 8→8, 11→9, 13→10.

Negative ranks sum to:

positive to:

so .

The one-tailed 5% critical value for is 10.

: reject .

There is evidence that the median exceeds 50.

5.07

The Wilcoxon rank-sum test

The Wilcoxon rank-sum test 5.07

The Wilcoxon rank-sum test (Non-parametric tests)
Definitions
  • Wilcoxon rank-sum test: a test of whether two independent samples come from populations with the same median (identical populations, against one shifted relative to the other).
Key results
  • Rank all values together, where is the size of the smaller sample. = sum of the ranks of the smaller sample.
  • Test statistic = the smaller of and . Reject if is less than or equal to the critical value for and .
  • For large samples: .
Method
  1. State hypotheses about the population medians, choosing one- or two-tailed from the context.
  2. Rank the combined data, find , then , and compare with tables.
  3. For a one-tailed test, check that the smaller sample's ranks are small (or large) in the direction predicts.

In practiceSamples A: 12, 15, 18, 20 and B: 14, 22, 25, 27, 30. Test at the 5% level (two-tailed) whether the population medians differ.

  1. rank all nine together
  2. A is the smaller sample, ,
  3. critical value for , , two-tailed 5%
  4. Do not reject : no significant evidence that the medians differ.
Notes
  • Use the rank-sum test for two independent samples and the signed-rank test for paired data. Applying the wrong one is a common error.
  • The non-parametric tests are less powerful than a or test when the population really is normal, but valid when it is not.

Worked example

Worked example

Times, in seconds, to complete a task: method A 12, 15, 19, 22; method B 17, 25, 28, 31, 34.

Test at the 5% level whether method A tends to give shorter times. (Critical value for , , one-tailed 5%: 12.)

Show worked solution

: the two populations have the same median; : A's median is lower.

Ranking all nine: A has ranks 1, 2, 4 and 5, so ;

so .

: reject .

There is evidence at the 5% level that method A gives shorter times.

5.08

Correlation

Testing correlation: Pearson and Spearman 5.08

Pearson and Spearman correlation (Correlation and regression)
Definitions
  • Product moment correlation coefficient : the strength of linear association.
  • Spearman's rank correlation coefficient : the PMCC of the ranks, when there are no ties.
Key results
  • Both lie between and . for any strictly monotonic relationship; only for an exact straight line.
  • Test with when the data come from a bivariate normal distribution (an elliptical scatter).
  • Use when that is doubtful, when the data are ranks, or when the relationship is monotonic but not linear.
Method
  1. Rank each variable separately. Find for each pair, then and .
  2. Compare or with the critical value for and the significance level, choosing one or two tails from the wording of .

In practiceTwo judges rank six entries: A gives 1, 2, 3, 4, 5, 6 and B gives 1, 3, 2, 4, 6, 5. Test for positive association at the 5% level.

  1. no ties, so the formula is exact
  2. critical value for , one-tailed 5%
  3. Reject : there is evidence of positive association between the judges' rankings.
Notes
  • Tied values receive the mean of the ranks they would occupy; the formula is then only approximate, so compute the PMCC of the ranks.
  • Significant correlation does not show causation, and near 0 does not mean there is no relationship, only no linear one.

Worked example

Worked example

Two judges rank 8 entries and .

Test at the 5% level whether there is positive agreement.

Show worked solution

, .

The one-tailed 5% critical value for is 0.6429.

Since:

reject : there is evidence of positive agreement between the judges.

5.08

Linear regression

Least-squares regression 5.08

Least-squares regression (Correlation and regression)
Definitions
  • Least-squares regression line of on : the line that minimises the sum of the squares of the residuals.
  • Residual: the observed value of minus the value predicted by the line, .
Key results
  • and , with and .
  • The line always passes through .
  • Use the line of on to estimate from a given , when is the independent (controlled) variable.
Method
  1. Calculate and from the summary statistics, then and ; give the equation in context with units.
  2. If the data have been coded, substitute the coding into the equation to recover the line in the original variables.

In practice, , , , . Find the regression line of on .

  1. it passes through , as it must
Notes
  • Predictions within the range of the data (interpolation) are reliable when the correlation is strong; extrapolation beyond it may not be.
  • Large residuals point to outliers or to a relationship that is not linear.

Worked example

Worked example

For six data points, , , and .

Find the regression line of on and estimate when .

Show worked solution

and:

and:

So:

and at , (interpolation, so reliable if the correlation is strong).