conetwork: Build a co-occurrence network from any field

View source: R/co-network.R

conetworkR Documentation

Build a co-occurrence network from any field

Description

With one field, entities are linked when they co-occur in the same document. With by, entities are linked when they share values of the by field across documents.

Usage

conetwork(
  data,
  field,
  by = NULL,
  sep = ";",
  counting = "full",
  similarity = "none",
  threshold = 0,
  min_occur = 1L,
  top_n = NULL,
  self_loops = FALSE,
  deduplicate = TRUE,
  format = "edgelist",
  strip_quotes = TRUE,
  id = NULL
)

Arguments

data

A data frame with column id and the specified field(s).

field

Character. The entity field — determines what the nodes are.

by

Character or NULL. What links the nodes. If NULL (default), entities are linked by co-occurring in the same document. If specified, entities are linked when they share values from the by field.

sep

Character or NULL. Delimiter for splitting character columns. Default ";". Set to NULL if columns are already list-columns.

counting

Character. Counting method. Default "full".

similarity

Character. Normalization method. Default "none".

threshold

Numeric. Minimum edge weight. Default 0.

min_occur

Integer. Minimum entity frequency. Default 1.

top_n

Integer or NULL. Return only the top n edges by weight. Default NULL (all edges).

self_loops

Logical. If TRUE, include self-loops (an entity linked to itself). Default FALSE.

deduplicate

Logical. If TRUE (default), each ⁠(paper, entity)⁠ pair is counted at most once — duplicate entries in the source data (e.g., the same author listed twice on a paper) are treated as one occurrence. Set to FALSE to count every raw occurrence.

format

Character. Output format:

"edgelist"

Default. A bibnets_network data frame with columns from, to, weight, count.

"gephi"

Gephi-ready data frame: Source, Target, Weight, Count, Type.

"igraph"

An igraph graph object (requires igraph).

"cograph"

A cograph_network object (requires cograph).

"matrix"

A sparse adjacency matrix.

strip_quotes

Logical. If TRUE (default), surrounding quote characters are removed from each entity, so a quoted CSV value such as "Alice" or ⁠""Alice""⁠ is treated as Alice. Set FALSE to keep quotes as part of the label.

id

Optional. Name of the column to use as the work identifier (the matrix-row dimension). If NULL (default), an existing id column is used when present, otherwise row numbers are used.

Details

Fields can be list-columns (already split) or character columns with delimiters (auto-split via sep).

Value

Depends on format: a bibnets_network data frame (default), a Gephi-ready data frame, an igraph graph, a cograph_network, or a sparse matrix.

Examples

data(biblio_data)

# Co-occurrence: keywords appearing in the same document
conetwork(biblio_data, "keywords")

# Authors linked by shared keywords
conetwork(biblio_data, "authors", by = "keywords")

# Keywords linked by shared authors
conetwork(biblio_data, "keywords", by = "authors")

# Journals linked by shared references (= journal coupling)
conetwork(biblio_data, "journal", by = "references", similarity = "cosine")

# Auto-splits semicolon-delimited string columns
d <- data.frame(id = 1:3, tags = c("ml; dl; nlp", "ml; cv", "dl; cv"))
conetwork(d, "tags")

bibnets documentation built on June 19, 2026, 1:06 a.m.