Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Statistics

Estimate, test and interpret honestly: sampling, likelihood, confidence and significance, regression, and the ways published numbers mislead.

01

From probability to inference

Probability runs from a known model to the data it produces, and every real question runs the other way, from one batch of data to an unknown model that will never be known exactly.

02

Describing a batch of numbers

Location, spread and shape reduce a sample to a handful of numbers, and every choice among them answers a different question and hides a different thing.

03

Where the data comes from

Random selection is what turns a batch of numbers into evidence about a population, and the two failures that break it, coverage and non-response, are not fixed by collecting more.

04

The sampling distribution

The sample mean is centred on the population mean and spreads as σ/n, and the central limit theorem makes its distribution normal whatever the population looked like.

05

What makes an estimator good

Bias, variance and mean squared error make "good" precise, explain the n-1 that everybody divides by, and show that an unbiased estimator is not automatically the one to use.

06

Maximum likelihood

One recipe turns any probability model into an estimator, reproduces the sample mean and sample proportion as special cases, and hands back a standard error from the curvature of the log likelihood.

07

Confidence intervals

A pivot turns an estimate and its standard error into a range with a stated long-run coverage, and the 95 per cent is a property of the procedure rather than of the interval in front of you.

08

Testing a hypothesis

A null model, a test statistic and a tail probability answer the question of whether an effect is there at all, and the p-value means something much narrower than the use made of it.

09

Errors, power and sample size

A test can fail in two directions, only one of which is controlled by the threshold, and a study without enough observations to detect an effect will exaggerate it whenever it does.

10

Comparing two groups

Paired and independent designs, the pooled and Welch t tests, the difference of two proportions on the Salk polio trial, and why effect size has to be reported next to significance.

11

Counts and categories

When the outcome is a category rather than a measurement, the chi-squared statistic compares observed counts with the counts a model predicts, for goodness of fit and for independence.

12

Fitting a line

Least squares derived by calculus, the slope in terms of covariance, correlation and r2 read narrowly, inference on the slope, and residuals as the check that any of it was appropriate.

13

Confounding and randomisation

Berkeley's admissions figures reverse when the departments are separated, adjustment can only handle variables somebody measured, and randomising is the one step that turns an association into a cause.

14

How statistics mislead

Every method in this course is correct and still routinely produces false findings, because of choices made before the analysis, tests that were not counted, results that were never published, and numbers stripped of their base rate.

Final Test