| build_hypa | R Documentation |
Constructs a k-th order De Bruijn graph from sequential trajectory data and uses a hypergeometric null model to detect paths with anomalous frequencies. Paths occurring more or less often than expected under the null model are flagged as over- or under-represented.
build_hypa(
data,
order = 2L,
alpha = 0.05,
min_count = 5L,
p_adjust = "BH",
k = NULL
)
data |
A data.frame (rows = trajectories), list of character vectors,
|
order |
Integer scalar or integer vector. Order(s) of the De Bruijn
graph (default |
alpha |
Numeric. Significance threshold for anomaly classification (default 0.05). Paths with HYPA score < alpha are under-represented; paths with score > 1-alpha are over-represented. |
min_count |
Integer. Minimum observed count for a path to be
classified as anomalous (default 5). Paths with fewer observations
are always classified as |
p_adjust |
Character. Method for multiple testing correction of
p-values. Default |
k |
Deprecated. Former name of |
An object of class c("net_hypa", "cograph_network") with
components:
Data frame with path, from, to, observed, expected,
ratio, p_value, p_under, p_over, p_adjusted_under,
p_adjusted_over, anomaly, order
columns (one block of rows per requested order). The path
column shows the full state sequence
(e.g., "A -> B -> C"); from is the context (conditioning
states); to is the next state; ratio is
observed / expected; p_value is retained as an alias for
p_under, the raw lower-tail hypergeometric CDF value;
p_over is the inclusive upper-tail probability
P(X >= observed); p_adjusted_under and
p_adjusted_over are the corrected p-values for under- and
over-representation tests respectively.
Alias for scores (all orders, arrow notation).
Subset of scores classified as over-represented.
Subset of scores classified as under-represented.
Weighted adjacency matrix of the lowest-order De Bruijn graph.
cograph weight matrix of the lowest-order graph.
Fitted propensity matrix of the lowest-order graph.
cograph edge data.frame of the lowest-order graph.
Named list of per-order result lists.
Integer vector of orders actually built (sorted ascending).
Back-compatibility alias for order.
Significance threshold used.
Multiple testing correction method used.
Number of anomalous paths detected (all orders).
Number of over-represented paths (all orders).
Number of under-represented paths (all orders).
Total number of edges (all orders).
data.frame (id, label, name) of the
lowest-order De Bruijn graph nodes (arrow notation).
Logical. Always TRUE.
cograph meta list of the lowest-order graph.
Always NULL.
LaRock, T., Nanumyan, V., Scholtes, I., Casiraghi, G., Eliassi-Rad, T., & Schweitzer, F. (2020). HYPA: Efficient Detection of Path Anomalies in Time Series Data on Networks. SDM 2020, 460-468.
seqs <- list(c("A","B","C"), c("B","C","A"), c("A","C","B"), c("A","B","C"))
hyp <- build_hypa(seqs, order = 2)
trajs <- list(c("A","B","C"), c("A","B","C"), c("A","B","C"),
c("A","B","D"), c("C","B","D"), c("C","B","A"))
h <- build_hypa(trajs, order = 2)
print(h)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.