Statistics
Estimate, test and interpret honestly: sampling, likelihood, confidence and significance, regression, and the ways published numbers mislead.
From probability to inference
Probability runs from a known model to the data it produces, and every real question runs the other way, from one batch of data to an unknown model that will never be known exactly.
Describing a batch of numbers
Location, spread and shape reduce a sample to a handful of numbers, and every choice among them answers a different question and hides a different thing.
Where the data comes from
Random selection is what turns a batch of numbers into evidence about a population, and the two failures that break it, coverage and non-response, are not fixed by collecting more.
The sampling distribution
The sample mean is centred on the population mean and spreads as , and the central limit theorem makes its distribution normal whatever the population looked like.
What makes an estimator good
Bias, variance and mean squared error make "good" precise, explain the that everybody divides by, and show that an unbiased estimator is not automatically the one to use.
Maximum likelihood
One recipe turns any probability model into an estimator, reproduces the sample mean and sample proportion as special cases, and hands back a standard error from the curvature of the log likelihood.
Confidence intervals
A pivot turns an estimate and its standard error into a range with a stated long-run coverage, and the 95 per cent is a property of the procedure rather than of the interval in front of you.
Testing a hypothesis
A null model, a test statistic and a tail probability answer the question of whether an effect is there at all, and the -value means something much narrower than the use made of it.
Errors, power and sample size
A test can fail in two directions, only one of which is controlled by the threshold, and a study without enough observations to detect an effect will exaggerate it whenever it does.
Comparing two groups
Paired and independent designs, the pooled and Welch tests, the difference of two proportions on the Salk polio trial, and why effect size has to be reported next to significance.
Counts and categories
When the outcome is a category rather than a measurement, the chi-squared statistic compares observed counts with the counts a model predicts, for goodness of fit and for independence.
Fitting a line
Least squares derived by calculus, the slope in terms of covariance, correlation and read narrowly, inference on the slope, and residuals as the check that any of it was appropriate.
Confounding and randomisation
Berkeley's admissions figures reverse when the departments are separated, adjustment can only handle variables somebody measured, and randomising is the one step that turns an association into a cause.
How statistics mislead
Every method in this course is correct and still routinely produces false findings, because of choices made before the analysis, tests that were not counted, results that were never published, and numbers stripped of their base rate.