unsurv_compare: Compare cluster partitions against observed survival outcomes

View source: R/compare.R

unsurv_compareR Documentation

Compare cluster partitions against observed survival outcomes

Description

Summarizes and compares one or more cluster-label partitions of the same individuals against observed time-to-event outcomes. This is intended for comparing an unsurv curve-based partition against baseline partitions (e.g., PAM on a scalar risk summary, or PAM on covariate PCA scores), or for comparing the same partitioning rule applied to different patient sets (e.g., a partition-defining set and a held-out validation set) to check that survival separation generalizes.

Usage

unsurv_compare(labels, time, status, reference = 1)

Arguments

labels

A named list of cluster-label vectors (integer or factor), each of the same length as time/status. Names are used as method labels; unnamed elements are labeled "method1", "method2", etc.

time

Numeric vector of observed follow-up times.

status

Numeric/integer vector of event indicators (1 = event, 0 = censored).

reference

Name or integer index of the element of labels used as the reference partition for Adjusted Rand Index (ARI) agreement. Defaults to the first element.

Details

Requires the survival package for Kaplan-Meier medians.

For each partition, the Adjusted Rand Index quantifies agreement with the reference partition. A log-rank test is deliberately not reported: when a partition is itself fit to separate the curves (as unsurv and the baselines are), a log-rank test against those same labels is circular and close to guaranteed to be "significant," so it is not a fair basis for comparing methods. Instead, per-cluster Kaplan-Meier medians are reported, which is useful for checking whether the ordering of clusters by survival (e.g., "cluster 2 has better survival than clusters 1 and 3") is preserved across sets, such as a partition-defining set and an independent validation set.

Value

An object of class "unsurv_compare" with elements:

  • summary: one row per method with K, cluster-size range, and ARI against the reference partition.

  • cluster_summary: one row per method/cluster with size, event count, and Kaplan-Meier median survival.

  • labels, time, status, reference: the inputs, stored for plotting.

Examples

if (requireNamespace("survival", quietly = TRUE)) {
  set.seed(1)
  n <- 120
  time <- stats::rexp(n, 0.1)
  status <- sample(0:1, n, TRUE)
  labs <- list(
    unsurv_curve = sample(1:3, n, TRUE),
    scalar_risk = sample(1:3, n, TRUE)
  )
  cmp <- unsurv_compare(labs, time, status)
  print(cmp)
}

unsurv documentation built on Sept. 1, 2026, 1:06 a.m.