Survival and time-to-event analysis
We analyse how long something takes rather than whether it happened, and keep the people it has not yet happened to inside the analysis.
Many of the questions a service asks are questions about timing. When do members stop attending, how long do people wait, how quickly do they return, and how long does a change persist once a programme ends. Reducing any of those to a binary outcome at a fixed follow-up point throws away most of the information and answers a question nobody asked.
The reason to use time-to-event methods rather than a logistic model is censoring. At the point the data is extracted, some people have not yet had the event and may never have it, and excluding them biases the estimate while treating them as event-free biases it the other way. Kaplan–Meier estimation and proportional-hazards models keep them in the analysis for exactly as long as they were observed.
Where the proportional-hazards assumption does not hold — and over a long follow-up it frequently does not — the models are extended rather than the assumption ignored: time-varying coefficients, stratification, and parametric survival models where the shape of the baseline hazard is itself of interest. Engagement that changes over time enters as a time-varying covariate rather than as a baseline characteristic.
Competing risks are separated from censoring. Somebody who completes a programme has not merely been lost to follow-up for drop-out, and treating the two as the same event overstates one and understates the other.
Discuss a piece of work
Describe the programme, the data and the deadline, and we will say what is feasible.