Free resource

Hypothesis testing & power

Type I, Type II and power as pictures instead of definitions: two sampling distributions, a critical value between them, and two shaded areas that shrink and swell as you move the sliders. The trade-off examiners keep asking about becomes something you have seen.

μ₀ = 100μ = 104.0
X̄ under H₀X̄ under the true meanType I · α = 0.05Type II · β = 0.36Critical value

Power 0.639

Power against sample size, everything else fixed

0.250.500.751.00n = 20n = 40n = 60n = 80n = 100

For fixed n, shrinking α drags the critical value outward and β grows: the trade-off has no free lunch. The only thing that improves both errors at once is more data — or a bigger true effect. Slide μ towards μ₀ and watch power collapse towards α: tiny effects are nearly undetectable at any reasonable n.

The picture worth memorising

A test fixes α — the probability of rejecting a true null — by placing the critical value so that exactly that much of H₀’s sampling distribution lies beyond it. Everything else follows from the geometry: β is whatever slice of the TRUE distribution falls short of that same critical value, and power is the rest of it. Push the true mean away from μ₀ and the distributions separate, β melts, power climbs. Tighten α and the critical value slides outward, growing β — the see-saw that only more data can lift both ends of, because √n narrows both curves at once. The p-value fits the same picture: it is the tail area beyond your OBSERVED statistic rather than beyond the critical value, which is why p < α and “reject” are the same statement.

The machinery behind it. The sampling distributions come from the central limit theorem, and the critical values from the statistical tables. Memori is a flashcard app built by actuarial students, with a ready-made CS1 set in the shop. Join the beta.

For education only.