normalize: Normalize a co-occurrence matrix

View source: R/normalize.R

normalizeR Documentation

Normalize a co-occurrence matrix

Description

Applies a similarity normalization to a square co-occurrence matrix. The diagonal of the input matrix is used as the total occurrence count for each item. Operates entirely in sparse representation.

Usage

normalize(A, method = "none")

Arguments

A

A square symmetric matrix (dense or sparse) representing co-occurrence counts.

method

Character. Normalization method:

"none"

No normalization. Returns raw co-occurrence counts.

"association"

Association strength (probabilistic affinity index). s_{ij} = c_{ij} / (w_i \cdot w_j). Often recommended as the best normalization for co-occurrence data.

"cosine"

Salton's cosine. s_{ij} = c_{ij} / \sqrt{w_i \cdot w_j}.

"jaccard"

Jaccard index. s_{ij} = c_{ij} / (w_i + w_j - c_{ij}).

"inclusion"

Inclusion index (Simpson coefficient). s_{ij} = c_{ij} / \min(w_i, w_j).

"equivalence"

Equivalence index (Salton's cosine squared). s_{ij} = c_{ij}^2 / (w_i \cdot w_j).

Value

A normalized sparse matrix of the same dimensions.

Examples

# Create a small co-occurrence matrix
A <- matrix(c(10, 3, 1, 3, 8, 2, 1, 2, 5), nrow = 3,
            dimnames = list(c("a", "b", "c"), c("a", "b", "c")))
normalize(A, "association")
normalize(A, "cosine")
normalize(A, "jaccard")

bibnets documentation built on June 19, 2026, 1:06 a.m.