Nothing
#' GRIN Evaluation of Lesion Boundaries
#'
#' @description
#' Generates genome-wide genomic boundaries defined by the unique start and end
#' positions of interval-based lesions, such as copy-number gains or deletions,
#' for use as marker regions in GRIN analyses.
#'
#' @usage
#' grin.lsn.boundaries(lsn.data,
#' chrom.size)
#'
#' @param lsn.data A lesion data table containing interval-based lesions,
#' typically copy-number gains or deletions. The data may include related
#' subtypes analyzed together, such as gains and amplifications or heterozygous
#' and homozygous deletions.
#' @param chrom.size A chromosome size table with two required columns:
#' `"chrom"` containing chromosome identifiers and `"size"` containing
#' chromosome sizes in base pairs.
#'
#' @details
#' The function partitions each chromosome into non-overlapping genomic
#' boundaries using all unique lesion start and end positions observed in the
#' supplied lesion data. A new boundary begins at each lesion start position
#' and immediately after each lesion end position.
#'
#' As a result, large lesions may be divided into multiple smaller boundaries
#' when other lesions begin or end within the same genomic region. This allows
#' GRIN to evaluate recurrent lesion patterns at a finer resolution than the
#' original lesion intervals.
#'
#' Boundaries span the entire chromosome, including regions outside observed
#' lesions. The first boundary begins at position 1, and the final boundary
#' extends to the end of the chromosome.
#'
#' This approach is particularly useful for copy-number variation analyses in
#' which recurrent genomic regions are evaluated independently of existing gene
#' or feature annotations.
#'
#' @return
#' A `data.frame` with five columns:
#' \describe{
#' \item{gene}{Unique boundary identifier constructed from chromosome, start
#' position, and end position.}
#' \item{chrom}{Chromosome containing the boundary.}
#' \item{loc.start}{Start position of the boundary in base pairs.}
#' \item{loc.end}{End position of the boundary in base pairs.}
#' \item{diff}{Length of the boundary in base pairs.}
#' }
#'
#' @export
#'
#' @references
#' Cao, X., Elsayed, A. H., & Pounds, S. B. (2023). Statistical Methods
#' Inspired by Challenges in Pediatric Cancer Multi-omics.
#'
#' @author
#' Abdelrahman Elsayed \email{abdelrahman.elsayed@stjude.org} and
#' Stanley Pounds \email{stanley.pounds@stjude.org}
#'
#' @seealso \code{\link{grin.stats}},
#' \code{\link{genomewide.log10q.plot}}
#'
#' @examples
#' data(lesion_data)
#' data(hg38_chrom_size)
#'
#' # This analysis is lesion-type specific. For example, extract gains:
#' gain <- lesion_data[lesion_data$lsn.type == "gain", ]
#'
#' # Generate genome-wide lesion boundaries for gains:
#' lsn.bound.gain <- grin.lsn.boundaries(gain, hg38_chrom_size)
#'
#' # Run GRIN using lesion boundaries as markers instead of gene annotations:
#' GRIN.results.gain.bound <- grin.stats(gain,
#' lsn.bound.gain,
#' hg38_chrom_size)
#'
#' # The same approach can be applied to deletions or related
#' # copy-number subtypes analyzed together.
#'
grin.lsn.boundaries=function(lsn.data, # Lesion data file that should be limited to include gain OR deletions (If gains are splitted to gain and amplification based on the log2Ratio value of the CNV segmentation file, the two categories can be included, same for homozygous and heterozygous deletions)
chrom.size) # Chromosome size table
{
required.lsn.cols=c("chrom","loc.start","loc.end")
if (!all(required.lsn.cols %in% colnames(lsn.data)))
stop(
"'lsn.data' must contain the following columns: ",
paste(required.lsn.cols,collapse=", "),
"."
)
if (!all(c("chrom","size") %in% colnames(chrom.size)))
stop("'chrom.size' must contain columns named 'chrom' and 'size'.")
chr.size=chrom.size
chroms=chr.size$chrom
lsn.bound=data.frame()
for (i in chroms) {
chr=i
chr.lsns=lsn.data[lsn.data$chrom==chr,]
chrom.size=chr.size$size[chr.size$chrom==chr]
# specify all unique boundary positions for included CNVs
if (nrow(chr.lsns)>0)
{
breakpoints=unique(c(
1,
chr.lsns$loc.start,
chr.lsns$loc.end+1,
chrom.size+1
))
} else {
breakpoints=c(1,chrom.size+1)
}
breakpoints=sort(breakpoints)
breakpoints=breakpoints[
breakpoints>=1 &
breakpoints<=chrom.size+1
]
# boundaries are the regions between consecutive unique lesion breakpoints
# large lesions will be split into multiple boundaries when other lesions
# start or end within the same genomic region
all.start=breakpoints[-length(breakpoints)]
all.end=breakpoints[-1]-1
chr.lsn.loci=cbind.data.frame(
gene=paste0("chr",chr,"_",all.start,"_",all.end),
chrom=chr,
loc.start=all.start,
loc.end=all.end
)
lsn.bound=rbind(lsn.bound,chr.lsn.loci)
}
lsn.bound$diff=lsn.bound$loc.end-lsn.bound$loc.start+1
return(lsn.bound)
}
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.