To access the market, drugs must demonstrate a favorable benefit-risk profile, supported by solid evidence. Randomized controlled trials (RCTs) are the gold standard for producing such evidence, as they provide a rigorous framework. Yet, they have limitations, such as a lack of representativeness of several sub-groups, short follow-up periods, and high costs. Conversely, observational studies are more accessible due to the increasing availability of large databases, but are considered insufficiently reliable because they are exposed to certain biases. However, recent advances offer new perspectives. Among them, target trials emulation (TTE) aims to mimic the structure of RCTs to produce valid causal inference from observational data. Pioneering work from the RCT-DUPLICATE initiative has evaluated whether TTE yields results consistent with those of RCTs. Yet, these methods still provide insufficient agreement with RCTs for regulatory decision-making, and sources of variability in treatment effect estimation remain to be explored. Different healthcare systems data cover populations with different risk factor distributions, and contain different variables. Because almost all the research on TTE comes from the same team, using North American claims data, we hypothesize that differences in sources of data affect the external validity of these studies. In addition, emulating a target trial involves many methodological choices to emulate each component of the trial. To date, the impact of these methodological choices on the variability of estimations remains scarcely explored.To ensure that TTE generates reliable and actionable evidence, it is crucial to better understand the sources of variability in treatment effect estimates. The two main objectives of the present project are therefore: 1. To explore how we can target the RCT estimand using different sources of observational data (i.e., insurance claims and electronic health records, from different countries), and evaluate the reproducibility (inter-team) and replicability (inter-base) of emulation. 2. To assess the variability in treatment effect estimates according to methodological choices (e.g control arm, outcomes, etc.), and to decipher which of these choices influence, and to which extent, the difference in results between the RCT and the emulated trial. These objectives will be addressed through five workpackages. First, we will define the causal research questions of twenty selected RCTs previously emulated by the RCT-DUPLICATE initiative, and describe the emulation process for each essential component (e.g. eligibility criteria, treatment strategies, follow-up start and duration, outcomes, etc). We will assess the feasibility of accurately emulating these target trials using two distinct data sources: the French National Health Insurance Database (SNDS) and the UK's Clinical Practice Research Datalink (CPRD). Inferential analyses will subsequently be conducted by different teams, and we will evaluate the replicability (two different teams, using two different sources of data) and the reproducibility (two different teams, same source of data) of emulations. In the fourth workpackage, we will use the vibration of effect framework to quantify the impact of methodological choices on the stability of results and on the robustness of the research findings. Finally, we aim to develop an empirical distribution of the design effect, through simulations, to give an order of magnitude of the sample size required, compared to that of the target RCT. This project will produce original results on the reproducibility and replicability of emulated studies using nation-wide, European health databases. It will further provide tools to avoid methodological pitfalls when emulating target trials from observational data. Altogether, this will offer key information on the reliability of these approaches to complement RCTs for decision-making.
