R/get.ensembl.annotation.R

Defines functions get.ensembl.annotation

Documented in get.ensembl.annotation

#' Retrieve Ensembl Annotation Data
#'
#' Retrieves pre-generated and versioned Ensembl annotation resources for
#' use with GRIN2. Annotation files are downloaded from the GRIN2 annotation
#' repository, verified using MD5 checksums, and cached locally by default
#' for subsequent analyses.
#'
#' @param genome.assembly Character string specifying the genome assembly.
#'   Currently supported: `"Human_GRCh38"`.
#' @param ensembl.version Integer specifying the Ensembl release.
#'   Currently supported: `110`.
#' @param annotation.type Character string specifying the annotation resource
#'   to retrieve. One of `"gene"`, `"exon"`, `"regulatory"`, or `"all"`.
#'   Specifying `"all"` retrieves all three annotation resources.
#' @param cache Logical indicating whether downloaded annotation files should
#'   be cached locally. Default is `TRUE`. Cached files are verified against
#'   their expected MD5 checksums before use.
#' @param cache.dir Optional character string specifying the directory in
#'   which annotation files should be cached. If `NULL`, the standard GRIN2
#'   user cache directory is used.
#' @param force.download Logical indicating whether annotation files should
#'   be downloaded again even when valid cached copies are available.
#'   Default is `FALSE`.
#' @param quiet Logical indicating whether download and status messages should
#'   be suppressed. Default is `FALSE`.
#'
#' @details
#' GRIN2 uses pre-generated, versioned annotation resources to provide
#' reproducible genomic analyses without requiring a live connection to
#' Ensembl BioMart during annotation retrieval.
#' This function downloads resources from the GRIN2 annotation repository when
#' a valid cached copy is unavailable, when \code{force.download = TRUE}, or
#' when a cached file fails checksum verification. Internet access is therefore
#' required for the initial download. Once a valid resource has been cached,
#' it can be reused without an internet connection.
#'
#' The available annotation resources include:
#'
#' \strong{Gene annotation}
#'
#' Gene-level annotation based on the specified Ensembl release. The resource
#' includes protein-coding genes as well as non-coding gene biotypes, including
#' long non-coding RNAs (lncRNAs), microRNAs (miRNAs), small nuclear RNAs
#' (snRNAs), small nucleolar RNAs (snoRNAs), pseudogenes, and other annotated
#' gene types. Each gene is represented by its Ensembl gene identifier and
#' genomic interval, together with gene name, biotype, strand, chromosome band,
#' and gene description where available. The columns \code{gene},
#' \code{chrom}, \code{loc.start}, and \code{loc.end} provide the genomic
#' interval required by GRIN2.
#'
#' \strong{Exon annotation}
#'
#' The exon annotation represents one selected transcript per gene. For
#' protein-coding genes, the MANE Select transcript is used when available.
#' MANE (Matched Annotation from NCBI and EMBL-EBI) Select identifies a
#' representative transcript with matching exon structure and sequence between
#' Ensembl/GENCODE and RefSeq. For genes without a selected MANE transcript,
#' including non-coding genes, the Ensembl Canonical transcript is used as the
#' representative transcript. The Ensembl Canonical transcript is the
#' representative transcript designated by Ensembl for a gene.
#'
#' This resource is used primarily for exon-level GRIN2 analyses, in which
#' selected lesion types, such as mutations and indels, can be evaluated using
#' exon-based rather than whole-gene target sizes when the analysis is restricted
#' to protein-altering alterations. It can be supplied
#' to \code{\link{grin.stats}} through the \code{exons.annotation},
#' \code{exon.chrom.size}, and \code{exon_level} arguments.
#'
#' This resource is used primarily for exon-level GRIN2 analyses, in which
#' selected lesion types are evaluated using exon-based rather than whole-gene
#' target sizes. It can be supplied together with chromosome-level exon target
#' sizes through the \code{exons.annotation}, \code{exon.chrom.size}, and
#' \code{exon_level} arguments used in the GRIN2 analysis workflow.
#'
#' The resource contains the individual exons belonging to the selected
#' transcript for each gene and includes exon genomic start and end coordinates,
#' exon number, Ensembl exon identifier, transcript identifier, transcript
#' biotype, transcript start and end coordinates, canonical-transcript status,
#' MANE Select status, and gene-level annotation including gene name, gene
#' boundaries, gene biotype, description, strand, and chromosome band. In this
#' resource, \code{loc.start} and \code{loc.end} specifically represent the
#' genomic start and end coordinates of each exon.
#'
#' \strong{Regulatory annotation}
#'
#' Regulatory annotation based on the Ensembl Regulatory Build. The current
#' GRCh38/Ensembl 110 resource contains predicted enhancer regions,
#' predicted promoter regions, CTCF-binding sites, and open chromatin
#' regions, together with gene intervals used to establish relationships
#' between regulatory elements and nearby genes.
#'
#' Regulatory features are additionally annotated with genes whose genomic
#' intervals overlap the regulatory feature, the nearest upstream and
#' downstream genes, their Ensembl gene identifiers, and the genomic
#' distance between each regulatory feature and the corresponding nearest
#' gene. The resource also reports genes located within 500 kb upstream
#' and within 500 kb downstream of each regulatory feature. These fields
#' facilitate interpretation of regulatory regions that may influence
#' nearby genes even when the regulatory element does not directly overlap
#' a gene.
#'
#' Annotation resources are downloaded only when required. When
#' \code{cache = TRUE}, a successfully downloaded and checksum-verified file
#' is stored in the GRIN2 user cache and reused in subsequent calls.
#' Annotation files are verified against their expected checksums before use.
#' If a cached file fails checksum verification, a fresh copy is downloaded
#' and verified. A newly downloaded file that does not match the expected
#' checksum is rejected and the function returns an error.
#'
#' @return If \code{annotation.type} is `"gene"`, `"exon"`, or
#'   `"regulatory"`, the corresponding annotation data frame is returned.
#'   If \code{annotation.type = "all"}, a named list with elements
#'   \code{gene}, \code{exon}, and \code{regulatory} is returned.
#'
#' @examples
#' \dontrun{
#' # Requires internet access if a valid cached copy is unavailable.
#' gene.annotation <- get.ensembl.annotation(
#'   genome.assembly = "Human_GRCh38",
#'   ensembl.version = 110,
#'   annotation.type = "gene"
#' )
#'
#' head(gene.annotation)
#' }
#' @export
get.ensembl.annotation <- function(
    genome.assembly = "Human_GRCh38",
    ensembl.version = 110L,
    annotation.type = c("gene", "exon", "regulatory", "all"),
    cache = TRUE,
    cache.dir = NULL,
    force.download = FALSE,
    quiet = FALSE)
{
  annotation.type <- match.arg(annotation.type)

  ####################################################################
  # 1. Validate arguments
  ####################################################################

  if (length(genome.assembly) != 1L || is.na(genome.assembly) ||
      !is.character(genome.assembly) || !nzchar(genome.assembly))
    stop("`genome.assembly` must be one non-empty character value.",
         call. = FALSE)

  if (length(ensembl.version) != 1L || is.na(ensembl.version) ||
      !is.numeric(ensembl.version) || !is.finite(ensembl.version) ||
      ensembl.version %% 1 != 0)
    stop("`ensembl.version` must be one non-missing integer value.",
         call. = FALSE)

  ensembl.version <- as.integer(ensembl.version)

  if (length(cache) != 1L || is.na(cache) || !is.logical(cache))
    stop("`cache` must be TRUE or FALSE.", call. = FALSE)

  if (length(force.download) != 1L || is.na(force.download) ||
      !is.logical(force.download))
    stop("`force.download` must be TRUE or FALSE.", call. = FALSE)

  if (length(quiet) != 1L || is.na(quiet) || !is.logical(quiet))
    stop("`quiet` must be TRUE or FALSE.", call. = FALSE)

  ####################################################################
  # 2. Identify requested annotation resources
  ####################################################################

  resources <- grin2.annotation.resources()

  matching.resources <- resources[
    resources$genome.assembly == genome.assembly &
      resources$ensembl.version == ensembl.version, , drop = FALSE
  ]

  if (nrow(matching.resources) == 0L)
    stop(paste0("No pre-generated GRIN2 annotation bundle is available for ",
                genome.assembly, " and Ensembl release ",
                ensembl.version, "."),
         call. = FALSE)

  requested.types <- if (annotation.type == "all")
    c("gene", "exon", "regulatory") else annotation.type

  matching.resources <- matching.resources[
    matching.resources$annotation.type %in% requested.types, , drop = FALSE
  ]

  if (!all(requested.types %in% matching.resources$annotation.type))
  {
    missing.types <- setdiff(requested.types,
                             matching.resources$annotation.type)

    stop(paste0("The requested annotation resource(s) are unavailable: ",
                paste(missing.types, collapse = ", "), "."),
         call. = FALSE)
  }

  bundle.versions <- unique(matching.resources$bundle.version)

  if (length(bundle.versions) != 1L)
    stop(paste0("Multiple annotation bundle versions were identified for ",
                genome.assembly, " and Ensembl release ",
                ensembl.version, "."),
         call. = FALSE)

  ####################################################################
  # 3. Determine cache directory
  ####################################################################

  if (is.null(cache.dir))
  {
    cache.dir <- file.path(
      tools::R_user_dir(package = "GRIN2", which = "cache"),
      "annotations",
      genome.assembly,
      paste0("Ensembl_", ensembl.version),
      paste0("bundle_", bundle.versions)
    )
  } else {
    if (length(cache.dir) != 1L || is.na(cache.dir) ||
        !is.character(cache.dir) || !nzchar(cache.dir))
      stop("`cache.dir` must be NULL or one non-empty character value.",
           call. = FALSE)

    cache.dir <- path.expand(cache.dir)
  }

  if (isTRUE(cache))
  {
    dir.create(cache.dir, recursive = TRUE, showWarnings = FALSE)

    if (!dir.exists(cache.dir))
      stop(paste0("Could not create the annotation cache directory: ",
                  cache.dir),
           call. = FALSE)
  }

  ####################################################################
  # 4. Retrieve requested annotation resources
  ####################################################################

  annotation.objects <- lapply(requested.types, function(current.type)
  {
    current.resource <- matching.resources[
      matching.resources$annotation.type == current.type, , drop = FALSE
    ]

    grin2.retrieve.annotation.resource(
      resource.row = current.resource[1L, , drop = FALSE],
      genome.assembly = genome.assembly,
      ensembl.version = ensembl.version,
      cache = cache,
      cache.dir = cache.dir,
      force.download = force.download,
      quiet = quiet
    )
  })

  names(annotation.objects) <- requested.types

  ####################################################################
  # 5. Return annotation
  ####################################################################

  if (annotation.type == "all")
    return(annotation.objects)

  annotation.objects[[1L]]
}

Try the GRIN2 package in your browser

Any scripts or data that you put into this service are public.

GRIN2 documentation built on Aug. 22, 2026, 5:09 p.m.