On this page:
Missing data are common in surveys and cohort studies. If not handled appropriately, they can introduce bias and reduce statistical power. Our goal is to explain how researchers working with CLS cohort data can use simple, robust methods that help them use all available data and produce more reliable findings.
Missing data in longitudinal studies happen for two reasons:
Study members do not participate in a sweep of data collection at all. This is known as “sweep non-response”. If they never return to the study, it is known as “attrition”. Reasons for this can include that a study member has died or emigrated, cannot be contacted, or chooses not to participate.
Some study members participate in a sweep of data collection, but do not answer every question. This is known as “item non-response”. For example, a person may feel unable to answer a question or choose not to answer it.
In short, sweep non-response means all data are missing at that sweep, while item non-response means only some answers are missing.
Missing data can affect research in two ways:
When data are missing, fewer observations are available. This leads to less precise estimates with wider confidence intervals, meaning that we are less likely to be able to draw definitive conclusions.
A perhaps more important issue is that missing data are often patterned by cohort members’ individual characteristics and circumstances. People who do not respond may differ systematically from those who do – for example, they may have lower income, worse health, or more disadvantaged backgrounds. This can lead to misleading results if not handled properly. For example, estimates may overstate or understate relationships relative to the true relationship in the target population
At CLS, we have developed an approach to deal with missing data and reduce bias. It builds on the rich data cohort members have provided to us over the years they have participated in our studies.
We suggest the use of well-known methods which aim to:
• use all available information
• account for differences between respondents and non-respondents
• reduce bias and improve precision.
These methods all rely on an important assumption called missing at random (MAR). This means that differences between respondents and non-respondents can be explained using the data we have observed.
Rather than filling in a single value for each missing observation, MI generates several plausible replacement values, producing multiple “completed” datasets. Each dataset is analysed separately, and the results are then pooled using specific rules (Rubin’s Rules) that account for the uncertainty introduced by the missing data.
The core idea of IPW is to reweight each observation by the inverse of its probability of it being included in the analysis, with these probabilities typically obtained by first fitting a model for inclusion in the analysis. Observations that were unlikely to be observed are given more weight, effectively creating a pseudo-population in which the weighted analysis sample is representative of the whole sample. IPW can also be used for confounder control.
FIML estimates model parameters using all available information, without imputing missing values at all. Rather than discarding incomplete cases or filling in gaps, FIML incorporates every observation’s available data directly into the likelihood function. The model finds the parameter values that would have been most likely to produce the observed data, even when observations are incomplete. It is particularly popular in structural equation modelling and other latent variable frameworks.
We can add extra information to our analyses. This is done using auxiliary variables – variables not included in the main analysis.
The most useful auxiliary variables are those that:
At CLS we have implemented a systematic data-driven approach to identify predictors of non-response in each cohort. These variables can be considered for use as auxiliary variables. A full list of the identified predictors of non-response at each sweep in each cohort is available in our Handling missing data in the CLS cohort studies user guide.
Missing at random assumption made more plausible: evidence from the 1958 British birth cohort
J Clin Epidemiol
Letter to the editor: Don’t forget survey data: ‘healthy cohorts’ are ‘real-world’ relevant if missing data are handled appropriately
Longitudinal and Life Course Studies
How to mitigate selection bias in COVID-19 surveys: evidence from five national cohorts
European Journal of Epidemiology
A data-driven approach to understanding non-response and restoring sample representativeness in the UK Next Steps cohort
Longitudinal and Life Course Studies
A data driven approach to address missing data in the 1970 British birth cohort
BMC Medical Research Methodology
A data driven approach to handling missing data in the UK Millennium Cohort Study
SocArXiv
NCDS response and missingness
The level of response for every major NCDS sweep, predictors of non-response and how to handle missing data.
BCS70 response and missingness
The level of response for every major BCS70 sweep, predictors of non-response and how to handle missing data.