knitr::opts_chunk$set( collapse = TRUE, comment = "#>", # pharmaversesdtm is suggested, only evaluate if it is installed eval = requireNamespace("pharmaversesdtm", quietly = TRUE) )
This R Markdown document illustrates example usage of the RESIDE package with multiple related tables, using the Demographics (DM) and Adverse Events (AE) SDTM domains from the pharmaversesdtm package. The tables are linked by a subject identifier (USUBJID), DM has one row per subject and AE has one row per adverse event.
Load the RESIDE package and set a seed for reproducibility and store the folder directory for export / import.
# Load the Library library(RESIDE) # Load dplyr for data manipulation library(dplyr) # Set the seed set.seed(1234) # Store the folder path used for import / export folder_path <- tempdir()
Store the tables in a named list, the names are used to identify the tables in the marginal distributions.
# Store the tables in a named list dfs <- list( dm = pharmaversesdtm::dm, ae = pharmaversesdtm::ae ) # Number of rows in each table sapply(dfs, nrow) # Number of subjects in each table sapply(dfs, function(df) length(unique(df$USUBJID))) # Treatment arms of the subjects table(dfs$dm$ARM) # Severity of the adverse events table(dfs$ae$AESEV)
Join the adverse events to the demographics of the subject and fit a logistic regression model for the odds of an adverse event being moderate or severe, by age, sex and treatment arm. The same function is used for the original and synthesised data.
# Join the adverse events to the demographics of each subject prepare_ae_data <- function(dm, ae) { ae |> select(USUBJID, AESEV) |> inner_join(select(dm, USUBJID, AGE, SEX, ARM), by = "USUBJID") |> # Remove screen failures and adverse events without a severity filter(ARM != "Screen Failure", AESEV != "") |> mutate( MOD_SEV = AESEV %in% c("MODERATE", "SEVERE"), ARM = relevel(factor(ARM), "Placebo") ) } ae_original <- prepare_ae_data(dfs$dm, dfs$ae) # Proportion of moderate or severe adverse events by treatment arm prop.table(table(ae_original$ARM, ae_original$MOD_SEV), 1) # Fit a logistic regression model glm.original <- glm( MOD_SEV ~ AGE + SEX + ARM, data = ae_original, family = binomial ) # Output the coefficients of the model summary(glm.original)$coefficients
NB Adverse events from the same subject are not independent, a mixed effects model would be more appropriate, however a logistic regression model is used here for simplicity.
Use the get_marginal_distributions() function to get the marginal distributions of the tables, specifying the column that identifies each subject using the subject_identifier parameter.
# Get the Marginal Distributions of both tables marginals <- get_marginal_distributions( dfs, subject_identifier = "USUBJID" ) # Summarise the Marginal Distributions summary(marginals)
Columns that are present in every table (other than the subject identifier), are identified as common columns, in this case the study identifier STUDYID.
Export the marginal distributions using the export_marginal_distributions() function, using the force parameter to override any existing files.
# Export the Marginal Distributions export_marginal_distributions(marginals, folder_path = folder_path, force = TRUE)
Import the exported marginal distributions using the import_marginal_distributions() function
# Import the Marginal Distributions imported_marginals <- import_marginal_distributions(folder_path = folder_path)
Synthesise data from the imported marginals using the synthesise_data function, for multiple tables a named list of data frames is returned.
# Synthesise the tables from the imported Marginal Distributions (without correlations) sim_dfs <- synthesise_data(imported_marginals) # Number of rows in each table sapply(sim_dfs, nrow) # Number of subjects in each table sapply(sim_dfs, function(df) length(unique(df$USUBJID)))
The number of rows and subjects of each table are maintained, synthesised subjects are identified by a number, which links the subjects between the tables.
Fit the same logistic regression model as earlier except this time on the synthesised data.
ae_sim <- prepare_ae_data(sim_dfs$dm, sim_dfs$ae) # Proportion of moderate or severe adverse events by treatment arm prop.table(table(ae_sim$ARM, ae_sim$MOD_SEV), 1) # Fit a logistic regression model on the synthesised data glm.sim <- glm( MOD_SEV ~ AGE + SEX + ARM, data = ae_sim, family = binomial ) # Output the coefficients of the model summary(glm.sim)$coefficients
Without correlations the variables are synthesised independently, so there is no relationship between the treatment arm and the severity of the adverse events.
Synthesise data from the imported marginals with assumed correlations, using the correlations parameter to specify a list of correlations created with the correlation() function. Categorical variables are correlated using a single category, specified with factor_name.x or factor_name.y, and the tables of the variables can be specified with df_name.x and df_name.y.
# Synthesise the tables specifying assumed correlations sim_dfs_cor <- synthesise_data( imported_marginals, correlations = list( # Subjects on placebo are more likely to have mild adverse events correlation( "ARM", "AESEV", 0.2, df_name.x = "dm", df_name.y = "ae", factor_name.x = "Placebo", factor_name.y = "MILD" ), # Male subjects are younger correlation("AGE", "SEX", -0.2, factor_name.y = "M") ) )
Correlated variables are synthesised together, one row per subject, and joined to each table by subject. Therefore correlations between tables are between subjects, and a correlated variable takes a single value for each subject within a table, in this case every adverse event of a subject has the same severity.
Using the synthesised data (with correlations) fit the same logistic regression model as earlier.
ae_sim_cor <- prepare_ae_data(sim_dfs_cor$dm, sim_dfs_cor$ae) # Proportion of moderate or severe adverse events by treatment arm prop.table(table(ae_sim_cor$ARM, ae_sim_cor$MOD_SEV), 1) # Fit a logistic regression model on the synthesised data (with correlations) glm.sim.cor <- glm( MOD_SEV ~ AGE + SEX + ARM, data = ae_sim_cor, family = binomial ) # Output the coefficients of the model summary(glm.sim.cor)$coefficients
With the assumed correlation, adverse events of subjects on placebo are less likely to be moderate or severe, allowing the analysis to be tested before access to the original data is granted.
NB The correlation specified is that of the underlying multivariate normal distribution (Copula), the correlation observed in the synthesised data will differ, particularly for categorical variables. It is not possible to entirely maintain all the marginal distributions when specifying correlations.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.