The package code is distributed under the package's MIT license. The bundled microbiome data are derived from the upstream resources described below. The package license does not replace or narrow applicable upstream attribution terms.
The valencia2k, valencia_linf_hypercube_1k,
valencia13k_dcst_depth2_merged, and valencia13k_dcst_depth3_merged objects
are derived from the VALENCIA training data:
The source-preparation scripts in data-raw/ record the deterministic
subsampling and dCST construction parameters. Original sample identifiers are
not retained in the merged 13k assignment objects.
The agp_gut object is derived from public American Gut Project 16S records
and deidentified metadata:
The full analysis first selected at most 5,000 samples with library size at
least 1,000 using random seed 42. data-raw/create_agp_gut_subset.py then
creates the 766-sample package object using an explicit deterministic rule:
Prevotella_7, Pasteurellaceae,
Akkermansia, or Staphylococcus;Eukaryota and Unassigned labels; andPhenotype fields are joined only after membership is fixed and cannot
influence selection. Exact run-ID membership, selection reasons, dCST labels,
and derived phenotype fields are retained in
inst/extdata/agp_gut_meta.csv. The script also reconstructs the bundled
count matrix from the full abundance table and retained 5,000-sample dCST
assignments. data-raw/build_agp_gut.R then constructs the package object.
Because inclusion probabilities differ by dCST, agp_gut is suitable for
demonstrating data alignment, filtering, normalization, and dCST construction
only. Its phenotype frequencies, effect sizes, and p-values must not be
interpreted as population estimates.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.