Northwood Metrics

Services

Regression and multilevel modelling

We estimate the relationship between an outcome and the things that might explain it, in a form that survives another analyst reading it.

Most quantitative questions in health services reduce to a regression of some kind, and most of the difficulty lies in choosing which kind and defending it. Continuous outcomes, binary outcomes, counts, ordered categories and proportions each imply a different model family, and the wrong family produces estimates that are precise and wrong rather than visibly broken.

Health data is almost never a flat table of independent observations. Patients sit within practices, participants within groups, repeated measurements within people, and areas within regions. Multilevel and mixed-effects models represent that structure directly and report how much of the variation belongs at each level, which is frequently the most useful output of the whole exercise. Where the structure is a nuisance rather than a question, fixed effects and cluster-robust variance estimation handle it.

Functional form is examined rather than assumed. A relationship with a floor, a ceiling or a threshold fitted as a straight line produces an average that describes nobody in the data, so non-parametric fits, splines and polynomial terms are used to establish the shape before a parametric form is committed to. Interactions are included where the question is about differences between groups, rather than inferred from separate models on separate subsamples.

Results are presented as predicted probabilities and marginal effects on the scale the reader thinks in, with the coefficient tables available underneath. A log-odds table answers a question almost nobody asked.

overallgroup Agroup Bgroup Cgroup Dno difference
Illustrative An estimate can look settled overall and not settled underneath.

Discuss a piece of work

Describe the programme, the data and the deadline, and we will say what is feasible.

Get in touch