View source: R/predict_lucid.R
| predict_lucid | R Documentation |
Predict cluster assignment and outcome using new data on G, Z, and optional Y.
If g_computation = TRUE, prediction uses only the G-to-X path from the
fitted model and returns counterfactual-style predictions under modified G.
This function can also be used to extract latent cluster assignments when using
the training data as input.
predict_lucid(
model,
lucid_model = NULL,
G,
Z = NULL,
Y = NULL,
CoG = NULL,
CoY = NULL,
response = TRUE,
g_computation = FALSE,
verbose = FALSE
)
model |
A model fitted and returned by |
lucid_model |
Optional; "early", "parallel", or "serial". Auto-detected
from |
G |
Exposures, a numeric vector, matrix, or data frame. Categorical variable should be transformed into dummy variables. If a matrix or data frame, rows represent observations and columns correspond to variables. |
Z |
Omics data, and required for every model type unless
The requirement is not arbitrary: the E-step forms the posterior from the
omics likelihood, so with no |
Y |
Outcome, a numeric vector. Categorical variable is not allowed. Binary outcome should be coded as 0 and 1. |
CoG |
Optional, covariates to be adjusted for estimating the latent cluster. A numeric vector, matrix or data frame. Categorical variable should be transformed into dummy variables. |
CoY |
Optional, covariates to be adjusted for estimating the association between latent cluster and the outcome. A numeric vector, matrix or data frame. Categorical variable should be transformed into dummy variables. |
response |
If |
g_computation |
If |
verbose |
A flag indicates whether detailed information is printed in console. Default is FALSE. Applies consistently to all three model types (early, parallel, serial). |
A list containing:
inclusion.p |
Posterior inclusion probabilities for latent clusters (a
matrix for "early"; a list by layer for "parallel" and "serial"). Columns are
ordered by cluster, matching the row order of the model's |
pred.x |
Predicted latent-cluster labels (a numeric vector for "early";
a list by layer for "parallel" and "serial"), obtained as the maximum a
posteriori column of |
pred.y |
Predicted outcome values. For binary outcomes, this is class
labels when |
pred.z |
Predicted omics means under g-computation mode
( |
Supplying Y makes the cluster prediction supervised: the outcome
enters the posterior alongside G and Z, as it does during fitting. Omitting
it predicts clusters from G and Z alone, which is what is wanted when the
outcome is unavailable or must not inform the assignment.
# prepare data (a small subset keeps the example quick)
G <- sim_data$G[1:150, ]
Z <- sim_data$Z[1:150, ]
Y_normal <- sim_data$Y_normal[1:150, , drop = FALSE]
# fit lucid model
fit1 <- estimate_lucid(G = G, Z = Z, Y = Y_normal, lucid_model = "early", K = 2,
family = "normal", max_itr = 20, max_tot.itr = 50)
# prediction on training set (lucid_model is auto-detected from fit1's class)
pred1 <- predict_lucid(model = fit1, G = G, Z = Z, Y = Y_normal)
pred2 <- predict_lucid(model = fit1, G = G, Z = Z)
# g-computation style prediction using only G
pred_g <- predict_lucid(model = fit1, G = G, Z = NULL, g_computation = TRUE)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.