By comprising a massive amount of repeated data, epidemiological cohort studies can address increasingly sophisticated questions on mechanisms, risk factors and progression of diseases that account for the dynamic and multidimensional aspect of health processes. Current analytical methods, however, fail to simultaneously embrace the dynamic and continuous-time nature of the processes, their complex interrelationships, and the imperfectness of their measurements (noisy, sparse and irregular visits, missing data, heterogeneity, multiple types) thus hampering the epidemiological research. EDyLES envisions to develop, apply and promote flexible dynamic models along with epidemiologically meaningful estimands, exploiting the complex longitudinal data available in cohorts to address in-depth questions involving time-varying processes in cohort studies. EDyLES is structured into 5 complementary tasks: 1. We will extend the joint modeling technique to capture complex interrelationships among time-varying processes and/or clinical events (using nonlinear effects, interactions, weighted cumulative exposures, direct effects of intermediate processes) using scalable estimation procedures. 2. We will assess how current machine learning (ML) techniques adapt to the noisy, irregular and truncated observation of dynamic processes either as exposures or outcomes, and propose innovations for random forests and some artificial neural network architectures (reservoir computing, neural controlled differential equations) considering the fusion of methods with biostatistical techniques. 3. We will translate model outputs into epidemiologically meaningful estimands that embrace the dynamic nature of the processes, namely counterfactual mediation contrasts (e.g., path-specific indirect effects), epidemiological indicators (e.g., expected durations in health states defined from repeated data), and ML-derived measures of association (e.g., surfaces of counterfactual hazard ratios over time). 4. Leveraging large cohort data (3C, FMSA, CKD-REIN), we will apply the methods to three motivating contexts illustrating the diversity of questions to be addressed: • in brain health in aging: what are the pathways of time-varying modifiable exposures (e.g., cardio-metabolic health)? • in Multiple System Atrophy: what are the expected timings of major impairments and the dysautonomic and cerebral mechanisms underlying the clinical progression? • in chronic kidney disease: what are the roles of anemia and iron deficiency in disease progression including major adverse cardiovascular events and quality of life? 5. We will promote the methods developed by providing didactic introductions and comprehensive comparisons of methods using case studies or simulation studies, and develop open-source user-friendly software for dissemination to the Public Health community. EDyLES is a collaborative multi-disciplinary project gathering researchers in biostatistics, computer science, and epidemiology, along with clinicians in neurology and nephrology, coming from 12 research teams. The project's added value is manifold. From a biostatistical perspective, EDyLES will deliver novel analytical tools along with open-source software solutions for modeling the complex longitudinal data collected in cohort studies. These methodologies will leverage complementary assets of biostatistical models and ML techniques. The project will also address pivotal epidemiological questions developed directly by epidemiologists and clinicians from each domain. Beyond a better understanding of the disease progression and the mechanisms at play, they will help determine optimal therapeutic approaches in MSA and CKD, and preventive targets in cerebral aging. Although initially motivated by 3 pathologies, we anticipate that the techniques developed within EDyLES will benefit other areas of Public Health.
