5.05

Unbiased estimates and the central limit theorem

Unbiased estimates and the distribution of the sample mean 5.05

The central limit theorem (Hypothesis tests and confidence intervals)
Definitions
  • Unbiased estimator: one whose expectation equals the parameter it estimates.
  • Sample mean , an unbiased estimate of .
  • Unbiased estimate of : .
Key results
  • For any population with mean and variance : and .
  • If , then exactly.
  • Central limit theorem: for any population, when is large.
Method
  1. Standardise with the standard error: .
  2. If is unknown and is large, use in its place.
Notes
  • Divide by , not , for an unbiased variance estimate: deviations from are on average smaller than deviations from .
  • The theorem is about the distribution of , not of . The population itself can stay as skewed as it likes.
  • Say when the CLT is being used: 'since is large, is approximately normal by the central limit theorem'.

Worked examples

Worked example

A sample of 10 values has and .

Find unbiased estimates of and .

Show worked solution

.

Worked example

has mean 4 and variance 9 but an unknown distribution.

Find approximately the probability that the mean of a random sample of 50 exceeds 4.5.

Show worked solution

By the CLT, since is large, .

5.05

Hypothesis tests for a population mean

Testing a population mean 5.05

Testing a population mean (Hypothesis tests and confidence intervals)
Key results
  • If , or is large (central limit theorem), then under the test statistic is (approximately) .
  • One-tailed tests at 5% and 1% use 1.645 and 2.326; two-tailed tests use 1.960 and 2.576.
  • If is unknown and is large, the unbiased estimate is used in its place.
Method
  1. State and in terms of the population mean , with the significance level and whether the test is one- or two-tailed.
  2. Calculate (or the -value, or the critical region for ) and compare.
  3. Conclude in context, without claiming certainty: 'there is sufficient evidence at the 5% level that…'.
Notes
  • Say why the distribution of is normal: either the population is normal, or is large enough for the central limit theorem.
  • A two-tailed test at the 5% level rejects exactly when lies outside the 95% confidence interval.

Worked example

Worked example

The mass of cereal in a box has standard deviation 8 g.

A random sample of 64 boxes has mean 52.1 g.

Test at the 5% level whether the mean mass exceeds 50 g.

Show worked solution

, .

Since is large,

under by the central limit theorem.

(equivalently:

).

Reject : there is evidence at the 5% level that the mean mass exceeds 50 g.

5.05

Confidence intervals

Confidence intervals for a mean 5.05

Confidence intervals (Hypothesis tests and confidence intervals)
Definitions
  • A 95% confidence interval: an interval constructed by a method which, over repeated samples, contains the true parameter 95% of the time.
Key results
  • Normal population with known , or large : , with replaced by when it is unknown and is large.
  • , and for 90%, 95% and 99% intervals.
  • The width is : four times the sample size halves the width.
Method
  1. Compute and the standard error, then .
  2. To judge a claim : if lies outside a 95% interval, a two-tailed test at the 5% level would reject it.
Notes
  • Never say 'there is a 95% probability that lies in this interval'. is fixed; the probability belongs to the method that produced the interval.
  • Higher confidence gives a wider interval. More data gives a narrower one.

Worked example

Worked example

A random sample of 80 bags has mean mass 32.4 g and standard deviation 5.1 g.

Find a 95% confidence interval for the population mean, and comment on a claim that g.

Show worked solution

is large, so use : SE:

The interval is:

that is:

33 lies inside the interval, so the data do not contradict the claim at the 5% level.

5.06

χ² goodness-of-fit tests

The χ² statistic and its distribution 5.06

The χ² distribution (χ² tests)
Definitions
  • Test statistic: , where are observed and expected frequencies under .
  • Degrees of freedom : the number of cells minus the number of independent constraints on the expected frequencies.
Key results
  • If every , is approximately under .
  • Large means observed and expected frequencies disagree, so the critical region is always the upper tail.
  • has mean and is skewed to the right, less so as grows.
Method
  1. Find each , combine adjacent cells if any , then count the cells that remain.
  2. Calculate , find , and compare with the critical value from tables or the calculator.
  3. Conclude in context: 'there is (or is not) sufficient evidence that…'.
Notes
  • Use frequencies, never proportions or percentages: scales with the sample size.
  • Not rejecting does not prove the model is right. It only shows the data are consistent with it.

Goodness-of-fit tests 5.06

Goodness of fit (χ² tests)
Key results
  • .
  • Models tested include the discrete uniform, binomial, Poisson and geometric distributions, and distributions given by a table of probabilities.
Method
  1. State : the model fits, with any given parameter values. State : it does not.
  2. If a parameter is not given, estimate it from the data (the sample mean for a Poisson ) and subtract an extra degree of freedom.
  3. Expected frequency = total × model probability. The last cell is usually 'this value or more', so the probabilities sum to 1.
Notes
  • Combining cells is done before counting , and it reduces .
  • Check the subtraction for estimated parameters. It is the most common error in this topic.

Worked example

Worked example

Goals in 100 matches: 0 goals 15 times, 1 goal 30, 2 goals 28, 3 goals 15, 4 or more 12.

The sample mean is 1.8.

Test at the 5% level whether a Poisson model fits.

Show worked solution

: goals follow a Poisson distribution.

Using Po(1.8), the probabilities are 0.1653, 0.2975, 0.2678, 0.1607 and 0.1087 (by subtraction), so:

all at least 5.

.

One parameter was estimated, so , with critical value 7.815.

do not reject .

The Poisson model is consistent with the data.

5.06

χ² tests for association

Contingency tables: tests for association 5.06

Contingency tables: testing for association (χ² tests)
Definitions
  • Contingency table: frequencies classified by two factors, with rows and columns.
Key results
  • Under (no association), .
  • , counted after any rows or columns are merged.
Method
  1. State : there is no association between the two factors. State : there is an association.
  2. Tabulate the contributions ; after a significant result, the largest contributions show where the association lies, and whether is above or below there.
Notes
  • Association is not causation: a significant result says the factors are related in this population, not why.
  • Merging must make sense in context: combine neighbouring age bands, not unrelated categories.

Worked example

Worked example

120 students are classified by whether they play an instrument and by grade.

Yes: A 25, B 15, C 10.

No: A 15, B 25, C 30.

Test at the 1% level for association.

Show worked solution

Column totals are 40 each and row totals 50 and 70, so in the 'Yes' row and in the 'No' row.

with:

The 1% critical value is 9.210 and:

so reject : there is evidence of association.

The largest contributions show that players gain more A grades and fewer C grades than expected.