Exposure measurement error: A short methodological primer and playlist

Contributors: Emma Nichols, Erik Meijer, Yao-Yi Chiang, Alden L. Gross, Eleanor Hayes-Larson, Sindana D. Ilango, Kayleigh Keller, Regina A. Shih, Adam A. Szpiro, Jennifer Weuve, Sara D. Adar

Studying the exposome, or the totality of exposures that individuals experience across the life course requires measuring features of the exposome. Although researchers seek to optimize measurement, some measurement error is unavoidable. Understanding, quantifying, and accounting for error in exposome measurement is an important methodological consideration in exposome research. While error in an outcome value is handled directly by many modeling approaches (with exceptions), most traditional methods assume no measurement error in the exposure. Error in measuring the exposome can have different impacts on the estimation of health effects depending on the type of measurement error, the relationship between measurement error and other variables in the model (e.g., if measurement error is associated with the outcome, referred to as differential measurement error), and the modeling framework used. This primer and publication playlist seeks to briefly explain key concepts in defining, quantifying, and accounting for measurement error, while embedding and summarizing key citations that cover these concepts in substantial depth.

Simple additive measurement error structures

We first focus on two most common structures of simple additive measurement error:

When the mean measurement error is 0 and measurement error is unrelated to the true value (for classical error) or observed value (for Berkson error) in addition to being unrelated to other variables in the model, bias from either classical measurement error or Berkson measurement error is well understood in linear regression models. In these settings with a single continuous exposure variable:

Classical measurement error:

There is a large body of research within psychology, epidemiology, and other fields that focuses on this form of bias, typically called “attenuation bias.” A number of procedures have been proposed to use estimates of reliability (a ratio that depends on the magnitude of exposure variation and measurement error variation) to correct for attenuation given a set of assumptions.

Muchinsky PM. The Correction for Attenuation. Educational and Psychological Measurement. 1996 Feb 1;56(1):63–75. doi:10.1177/0013164496056001004

Brief summary: This paper provides historical context and discusses controversies around the correction for attenuation within the psychology literature. The psychology literature focuses on corrections for the correlation between two variables (rather than adjustments for coefficients estimated within a regression framework, as is typically the focus in other literature). The paper can help readers understand the historical foundations of scholarly investigation of the impact of measurement error on measures of association between two variables.

Additional approaches in the econometrics literature to adjust for classical measurement error include:

Schennach SM. Mismeasured and unobserved variables. In: Handbook of Econometrics. Elsevier; 2020. p. 487–565.

Brief summary: A technical reference summary of econometric concepts and approaches to handling measurement error. This book chapter includes the two econometric approaches highlighted above along with a variety of approaches to deal with more complex measurement error structures, and includes details on identification of available estimators.

Berkson measurement error:

More complex measurement error structures

When models are nonlinear, or when measurement error structures are more complex (e.g., associated with the outcome or errors in other variables), other approaches are needed. Given the common occurrence of complex measurement error patterns, existing literature cautions researchers against the use of simple heuristics described above, which apply to non-differential, independent classical measurement error of continuous variables, but do not apply to more complex situations. We review a few of the most common complex forms of measurement error below, and use directed acyclic graphs (DAGs) to illustrate these structures. DAGs are adapted from Hernan & Robins, 2020, which covers a variety of topics in causal inference, including measurement error.

Measurement Bias and “Noncausal” Diagrams. In: Hernan M, Robins J. What If. Boca Raton: Chapman & Hall/CRC; 2020. p. 125-131.

Brief summary: This section of Hernan & Robins’ book on causal inference focuses on measurement error and the representation of measurement error in DAGs. They review different types of measurement error, map them to DAGs, and provide examples of how that measurement error could appear in practical examples.

In all following DAGs, we consider exposure (X) and outcome (Y). We use a * to denote a mismeasured variable and U to denote measurement error. In one instance, we also include a second exposure (W). For example, classical measurement error in the exposure would be encoded as:

Complex, alternative structures include:

Zeger SL, Thomas D, Dominici F, Samet JM, Schwartz J, Dockery D, et al. Exposure measurement error in time-series studies of air pollution: concepts and consequences. Environ Health Perspect. 2000 May;108(5):419–26. doi:10.1289/ehp.00108419 PubMed PMID: 10811568; PubMed Central PMCID: PMC1638034.

Brief Summary: This paper creates a framework for understanding measurement error in models focused on air pollution, including considering the case of multipollutant models with dependent measurement error between the exposures. The authors present a series of simulation results, with succinct summaries of findings that can be used as rules of thumb, and also implement analyses to quantify measurement error in a realistic analysis of real-world data.

Jurek AM, Greenland S, Maldonado G. Brief Report: How far from non-differential does exposure or disease misclassification have to be to bias measures of association away from the null? Int J Epidemiol. 2008 Apr 1;37(2):382–5. doi:10.1093/ije/dym291

Brief Summary: This brief paper presents multiple scenarios based on published data showing the sensitivity of measures of association to small changes in measurement error parameters related to misclassification that may be common in epidemiologic research. Findings illustrate the importance of considering complex measurement error patterns.

These forms of error can also occur in combination. For instance, it is possible to have differential, dependent measurement error.

Quantifying and adjusting for more complex measurement error structures

Oftentimes, validation data (i.e., external data or a subset of internal data containing both mismeasured and true measurements) or repeated measurements can be useful in estimating bias or adjusting findings in the presence of complex measurement error structures. We highlight key topic areas and papers related to various types of measurement error structures, and approaches to adjust for measurement error:

Regression calibration: Regression calibration uses validation data to predict the true value (i.e., the value without measurement error) based on the potentially mis-measured value and a set of covariates. Alternatively, regression calibration can leverage information about error variance from other sources, such as from repeated noisy measurements. Information from the calibration model is then used to adjust the biased regression coefficient. Standard errors can be estimated using a bootstrap or via an asymptotic variance estimator.

Multiple imputation for measurement error correction (MIME): MIME treats measurement error as a missing data problem. Internal validation data is needed, and these data are used to derive an imputation model leveraging covariate and outcome data as well as the mismeasured variable.

Simulation extrapolation (SIMEX): SIMEX relies on a good estimate of error variance, which, as in the case of regression calibration, can come from validation data or repeated noisy measurements. The SIMEX procedure involves repeatedly simulating additional error in the mismeasured variable and re-estimating the association of interest. These simulations enable the estimation of a smooth curve describing the relationship between the magnitude of error and estimated coefficient, which can then be extrapolated to estimate what the coefficient would be without measurement error (i.e., under the condition where the measurement error variance is 0). This procedure is a generalization of bootstrap-based measurement error bias correction.

Innes GK, Bhondoekhan F, Lau B, Gross AL, Ng DK, Abraham AG. The Measurement Error Elephant in the Room: Challenges and Solutions to Measurement Error in Epidemiology. Epidemiologic Reviews. 2021 Dec 30;43(1):94–105. doi:10.1093/epirev/mxab011

Brief Summary: This easy to digest guide reviews conceptual topics in measurement error and summarizes common approaches to address measurement error in epidemiologic studies, including regression calibration, MIME, and SIMEX. For each approach, the review summarizes the type of mismeasured data the approach can be applied to, requirements around validation data, and available software packages to facilitate use. Citations highlight both foundational papers and examples of applications from the literature.

Simulation analysis: Simulations are a flexible tool that can be used to ask questions about the impact of measurement error on findings in scenarios where the truth is known. Simulation analysis to probe the effects of measurement error is often discussed in the context of literature on quantitative bias analysis, though this umbrella term describes a broader suite of methods and approaches, including approaches to quantify the impacts of other forms of bias.

Fox MP, MacLehose RF, Lash TL. SAS and R code for probabilistic quantitative bias analysis for misclassified binary variables and binary unmeasured confounders. Int J Epidemiol. 2023 May 4;52(5):1624–33. doi:10.1093/ije/dyad053 PubMed PMID: 37141446; PubMed Central PMCID: PMC10555728.

Brief Summary: This tutorial describes how simulated-based quantitative bias can be used in the setting of a misclassified binary variable. The authors include example code in both SAS and R and walk readers through the process of conducting a simulation-based quantitative bias analysis and interpreting the results.

Joint modeling: Joint models for the exposure-outcome relationship and the measurement error process can also be used to directly account for exposure measurement error with or without validation data. These models typically simultaneously estimate a model for the relationship between the outcome and the observed (with error) exposure alongside a model for the relationship between the true exposure and the observed (with error) exposure, with the models linked by a latent variable representing true exposure. Estimation is possible by bringing in additional information about the magnitude of the measurement error (typically the error variance). Hierarchical Bayesian estimation approaches are well suited for estimating these joint models and are often used in this setting.

Richardson S, Gilks WR. A Bayesian Approach to Measurement Error Problems in Epidemiology Using Conditional Independence Models. American Journal of Epidemiology. 1993 Sep 15;138(6):430–42. doi:10.1093/oxfordjournals.aje.a116875

Brief Summary: Describes how hierarchical Bayesian models can be used to correct for measurement error in epidemiologic research. The paper describes how ancillary data from a validation studies or repeated measures could be used to estimate models and confirms performance in a simple simulation analysis.

Measurement error in geospatial exposures based on prediction models: Measurement error in spatial exposures resulting from prediction models (e.g., air pollution, heat) are complex, as the prediction process typically results in highly correlated measurement error with features of classical and Berkson errors. Existing approaches to adjust for this form of measurement error typically require re-estimation of the exposure model, which may not be feasible for researchers due to data available restrictions.

Sheppard L, Burnett RT, Szpiro AA, Kim SY, Jerrett M, Pope CA, et al. Confounding and exposure measurement error in air pollution epidemiology. Air Qual Atmos Health. 2012 Jun 1;5(2):203–16. doi:10.1007/s11869-011-0140-9

Brief Summary: This approachable review discusses both confounding and exposure measurement error in the context of air pollution research, though concepts are relevant to other spatially predicted exposures as well. The discussion of exposure measurement provides both a conceptual framework for why exposure measurement error matters and is complex in this setting, as well as a review of existing literature on available approaches to adjust for measurement error, with key references providing more technical information.

Latent variable modeling: Latent variable models can be used when the construct of interest is not directly observed but can be inferred from a larger number of variables hypothesized to be caused by the latent variable. In a concrete example, cognitive functioning can be conceptualized as a latent variable, and modeled using observed responses to a battery of cognitive tests. In addition to estimating the measurement structure of cognition or other latent variables, latent variable models can be embedded in a broader structural equation modeling framework. By simultaneous estimating a measurement model for the potentially mismeasured variable and the health effects model in the same modeling framework, this approach can explicitly consider and account for measurement error.

Conclusion

Measurement error in capturing features of the exposome is an underappreciated aspect of exposome research that can have important impacts on research findings. We note that measurement error in other components of health outcome models, including in covariates and in outcomes is beyond the scope of this primer, but can have important (and different) impacts. We also did not consider the impacts of measurement error on more complex model structures, including machine learning models or mixtures models that consider many potentially mismeasured exposures simultaneously, which are important areas for future research investment. Nonetheless, this primer can serve increase awareness and understanding of both basic concepts in measurement error and available approaches that can help researchers quantify the impacts of measurement error in research and correct findings for potential bias due to measurement error.

SIGN UP FOR OUR MAILING LIST


The GECC is funded by the National Institute on Aging (NIA) U24AG088894.