read_biblio: Read bibliometric data

View source: R/read-biblio.R

read_biblioR Documentation

Read bibliometric data

Description

Universal reader that handles files, folders, format detection, and generic CSV input. Accepts a single file, multiple files, or a directory.

Usage

read_biblio(
  path,
  format = "auto",
  id = NULL,
  authors = NULL,
  keywords = NULL,
  references = NULL,
  countries = NULL,
  affiliations = NULL,
  journal = NULL,
  sep = ";",
  list_cols = NULL,
  ...,
  actors = NULL
)

Arguments

path

Character. Path to a file, a vector of file paths, or a directory containing export files.

format

Character. File format:

"auto"

Default. Auto-detect from file content.

"scopus"

Scopus CSV.

"wos"

Web of Science plaintext.

"wos_tab"

Web of Science tab-delimited.

"bibtex"

BibTeX .bib file.

"ris"

RIS file.

"dimensions"

Dimensions CSV.

"lens"

Lens.org CSV.

"openalex_csv"

Flat OpenAlex CSV export (pipe-delimited fields).

"generic"

Any CSV. Map its columns with id, authors, keywords, references, countries, affiliations, journal. Inferred automatically when any of those arguments is supplied, so format = "generic" is optional in that case.

id

Character. Column name for document identifier. Only used when format = "generic". Default NULL (uses row numbers).

authors, keywords, references, countries, affiliations

Character. For format = "generic", the name of the source column to map onto that standard field. Its cells are split on sep into a list-column. For example authors = "Author Names" reads the ⁠Author Names⁠ column into the standard authors list-column.

journal

Character. For format = "generic", the name of the source column to use as the (scalar) journal field. Not split.

sep

Character. Delimiter for splitting the mapped multi-valued columns. Default ";".

list_cols

Character vector. For format = "generic", additional columns to split into list-columns in place (keeping their original names), for fields without a dedicated argument above.

...

Additional arguments passed to the format-specific reader.

actors

Deprecated. Use the entity arguments (authors, keywords, ...) or list_cols instead.

Value

A data frame.

Examples

# Auto-detect format from file content (here: a bundled OpenAlex CSV)
f <- system.file("extdata", "openalex_works.csv", package = "bibnets")
data <- read_biblio(f)
head(data[, c("id", "title", "year", "journal")])

# Read multiple files at once; auto-detects each format
f_scopus <- system.file("extdata", "scopus_sample.csv", package = "bibnets")
f_wos    <- system.file("extdata", "wos_sample.txt",  package = "bibnets")
combined <- read_biblio(c(f_scopus, f_wos))
head(combined[, c("id", "title", "year", "journal")])

# Read every supported export in a directory (here: the bundled extdata)
folder <- system.file("extdata", package = "bibnets")
all_data <- read_biblio(folder)
nrow(all_data)

# Custom CSV: map each source column onto a standard field by name.
# Naming columns implies format = "generic" (no need to pass it).
tmp <- tempfile(fileext = ".csv")
write.csv(data.frame(
  doc_id  = c("a", "b"),
  Authors = c("Smith J; Jones A", "Davis M"),
  Keywords = c("networks; bibliometrics", "analytics")
), tmp, row.names = FALSE)
generic <- read_biblio(tmp,
                       id = "doc_id",
                       authors = "Authors",
                       keywords = "Keywords",
                       sep = ";")
head(generic)

bibnets documentation built on June 19, 2026, 1:06 a.m.