Nothing
#' Retrieve Ensembl Annotation Data
#'
#' Retrieves pre-generated and versioned Ensembl annotation resources for
#' use with GRIN2. Annotation files are downloaded from the GRIN2 annotation
#' repository, verified using MD5 checksums, and cached locally by default
#' for subsequent analyses.
#'
#' @param genome.assembly Character string specifying the genome assembly.
#' Currently supported: `"Human_GRCh38"`.
#' @param ensembl.version Integer specifying the Ensembl release.
#' Currently supported: `110`.
#' @param annotation.type Character string specifying the annotation resource
#' to retrieve. One of `"gene"`, `"exon"`, `"regulatory"`, or `"all"`.
#' Specifying `"all"` retrieves all three annotation resources.
#' @param cache Logical indicating whether downloaded annotation files should
#' be cached locally. Default is `TRUE`. Cached files are verified against
#' their expected MD5 checksums before use.
#' @param cache.dir Optional character string specifying the directory in
#' which annotation files should be cached. If `NULL`, the standard GRIN2
#' user cache directory is used.
#' @param force.download Logical indicating whether annotation files should
#' be downloaded again even when valid cached copies are available.
#' Default is `FALSE`.
#' @param quiet Logical indicating whether download and status messages should
#' be suppressed. Default is `FALSE`.
#'
#' @details
#' GRIN2 uses pre-generated, versioned annotation resources to provide
#' reproducible genomic analyses without requiring a live connection to
#' Ensembl BioMart during annotation retrieval.
#' This function downloads resources from the GRIN2 annotation repository when
#' a valid cached copy is unavailable, when \code{force.download = TRUE}, or
#' when a cached file fails checksum verification. Internet access is therefore
#' required for the initial download. Once a valid resource has been cached,
#' it can be reused without an internet connection.
#'
#' The available annotation resources include:
#'
#' \strong{Gene annotation}
#'
#' Gene-level annotation based on the specified Ensembl release. The resource
#' includes protein-coding genes as well as non-coding gene biotypes, including
#' long non-coding RNAs (lncRNAs), microRNAs (miRNAs), small nuclear RNAs
#' (snRNAs), small nucleolar RNAs (snoRNAs), pseudogenes, and other annotated
#' gene types. Each gene is represented by its Ensembl gene identifier and
#' genomic interval, together with gene name, biotype, strand, chromosome band,
#' and gene description where available. The columns \code{gene},
#' \code{chrom}, \code{loc.start}, and \code{loc.end} provide the genomic
#' interval required by GRIN2.
#'
#' \strong{Exon annotation}
#'
#' The exon annotation represents one selected transcript per gene. For
#' protein-coding genes, the MANE Select transcript is used when available.
#' MANE (Matched Annotation from NCBI and EMBL-EBI) Select identifies a
#' representative transcript with matching exon structure and sequence between
#' Ensembl/GENCODE and RefSeq. For genes without a selected MANE transcript,
#' including non-coding genes, the Ensembl Canonical transcript is used as the
#' representative transcript. The Ensembl Canonical transcript is the
#' representative transcript designated by Ensembl for a gene.
#'
#' This resource is used primarily for exon-level GRIN2 analyses, in which
#' selected lesion types, such as mutations and indels, can be evaluated using
#' exon-based rather than whole-gene target sizes when the analysis is restricted
#' to protein-altering alterations. It can be supplied
#' to \code{\link{grin.stats}} through the \code{exons.annotation},
#' \code{exon.chrom.size}, and \code{exon_level} arguments.
#'
#' This resource is used primarily for exon-level GRIN2 analyses, in which
#' selected lesion types are evaluated using exon-based rather than whole-gene
#' target sizes. It can be supplied together with chromosome-level exon target
#' sizes through the \code{exons.annotation}, \code{exon.chrom.size}, and
#' \code{exon_level} arguments used in the GRIN2 analysis workflow.
#'
#' The resource contains the individual exons belonging to the selected
#' transcript for each gene and includes exon genomic start and end coordinates,
#' exon number, Ensembl exon identifier, transcript identifier, transcript
#' biotype, transcript start and end coordinates, canonical-transcript status,
#' MANE Select status, and gene-level annotation including gene name, gene
#' boundaries, gene biotype, description, strand, and chromosome band. In this
#' resource, \code{loc.start} and \code{loc.end} specifically represent the
#' genomic start and end coordinates of each exon.
#'
#' \strong{Regulatory annotation}
#'
#' Regulatory annotation based on the Ensembl Regulatory Build. The current
#' GRCh38/Ensembl 110 resource contains predicted enhancer regions,
#' predicted promoter regions, CTCF-binding sites, and open chromatin
#' regions, together with gene intervals used to establish relationships
#' between regulatory elements and nearby genes.
#'
#' Regulatory features are additionally annotated with genes whose genomic
#' intervals overlap the regulatory feature, the nearest upstream and
#' downstream genes, their Ensembl gene identifiers, and the genomic
#' distance between each regulatory feature and the corresponding nearest
#' gene. The resource also reports genes located within 500 kb upstream
#' and within 500 kb downstream of each regulatory feature. These fields
#' facilitate interpretation of regulatory regions that may influence
#' nearby genes even when the regulatory element does not directly overlap
#' a gene.
#'
#' Annotation resources are downloaded only when required. When
#' \code{cache = TRUE}, a successfully downloaded and checksum-verified file
#' is stored in the GRIN2 user cache and reused in subsequent calls.
#' Annotation files are verified against their expected checksums before use.
#' If a cached file fails checksum verification, a fresh copy is downloaded
#' and verified. A newly downloaded file that does not match the expected
#' checksum is rejected and the function returns an error.
#'
#' @return If \code{annotation.type} is `"gene"`, `"exon"`, or
#' `"regulatory"`, the corresponding annotation data frame is returned.
#' If \code{annotation.type = "all"}, a named list with elements
#' \code{gene}, \code{exon}, and \code{regulatory} is returned.
#'
#' @examples
#' \dontrun{
#' # Requires internet access if a valid cached copy is unavailable.
#' gene.annotation <- get.ensembl.annotation(
#' genome.assembly = "Human_GRCh38",
#' ensembl.version = 110,
#' annotation.type = "gene"
#' )
#'
#' head(gene.annotation)
#' }
#' @export
get.ensembl.annotation <- function(
genome.assembly = "Human_GRCh38",
ensembl.version = 110L,
annotation.type = c("gene", "exon", "regulatory", "all"),
cache = TRUE,
cache.dir = NULL,
force.download = FALSE,
quiet = FALSE)
{
annotation.type <- match.arg(annotation.type)
####################################################################
# 1. Validate arguments
####################################################################
if (length(genome.assembly) != 1L || is.na(genome.assembly) ||
!is.character(genome.assembly) || !nzchar(genome.assembly))
stop("`genome.assembly` must be one non-empty character value.",
call. = FALSE)
if (length(ensembl.version) != 1L || is.na(ensembl.version) ||
!is.numeric(ensembl.version) || !is.finite(ensembl.version) ||
ensembl.version %% 1 != 0)
stop("`ensembl.version` must be one non-missing integer value.",
call. = FALSE)
ensembl.version <- as.integer(ensembl.version)
if (length(cache) != 1L || is.na(cache) || !is.logical(cache))
stop("`cache` must be TRUE or FALSE.", call. = FALSE)
if (length(force.download) != 1L || is.na(force.download) ||
!is.logical(force.download))
stop("`force.download` must be TRUE or FALSE.", call. = FALSE)
if (length(quiet) != 1L || is.na(quiet) || !is.logical(quiet))
stop("`quiet` must be TRUE or FALSE.", call. = FALSE)
####################################################################
# 2. Identify requested annotation resources
####################################################################
resources <- grin2.annotation.resources()
matching.resources <- resources[
resources$genome.assembly == genome.assembly &
resources$ensembl.version == ensembl.version, , drop = FALSE
]
if (nrow(matching.resources) == 0L)
stop(paste0("No pre-generated GRIN2 annotation bundle is available for ",
genome.assembly, " and Ensembl release ",
ensembl.version, "."),
call. = FALSE)
requested.types <- if (annotation.type == "all")
c("gene", "exon", "regulatory") else annotation.type
matching.resources <- matching.resources[
matching.resources$annotation.type %in% requested.types, , drop = FALSE
]
if (!all(requested.types %in% matching.resources$annotation.type))
{
missing.types <- setdiff(requested.types,
matching.resources$annotation.type)
stop(paste0("The requested annotation resource(s) are unavailable: ",
paste(missing.types, collapse = ", "), "."),
call. = FALSE)
}
bundle.versions <- unique(matching.resources$bundle.version)
if (length(bundle.versions) != 1L)
stop(paste0("Multiple annotation bundle versions were identified for ",
genome.assembly, " and Ensembl release ",
ensembl.version, "."),
call. = FALSE)
####################################################################
# 3. Determine cache directory
####################################################################
if (is.null(cache.dir))
{
cache.dir <- file.path(
tools::R_user_dir(package = "GRIN2", which = "cache"),
"annotations",
genome.assembly,
paste0("Ensembl_", ensembl.version),
paste0("bundle_", bundle.versions)
)
} else {
if (length(cache.dir) != 1L || is.na(cache.dir) ||
!is.character(cache.dir) || !nzchar(cache.dir))
stop("`cache.dir` must be NULL or one non-empty character value.",
call. = FALSE)
cache.dir <- path.expand(cache.dir)
}
if (isTRUE(cache))
{
dir.create(cache.dir, recursive = TRUE, showWarnings = FALSE)
if (!dir.exists(cache.dir))
stop(paste0("Could not create the annotation cache directory: ",
cache.dir),
call. = FALSE)
}
####################################################################
# 4. Retrieve requested annotation resources
####################################################################
annotation.objects <- lapply(requested.types, function(current.type)
{
current.resource <- matching.resources[
matching.resources$annotation.type == current.type, , drop = FALSE
]
grin2.retrieve.annotation.resource(
resource.row = current.resource[1L, , drop = FALSE],
genome.assembly = genome.assembly,
ensembl.version = ensembl.version,
cache = cache,
cache.dir = cache.dir,
force.download = force.download,
quiet = quiet
)
})
names(annotation.objects) <- requested.types
####################################################################
# 5. Return annotation
####################################################################
if (annotation.type == "all")
return(annotation.objects)
annotation.objects[[1L]]
}
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.