deduplicate: Remove duplicate observations

View source: R/deduplicate.R

deduplicateR Documentation

Remove duplicate observations

Description

Identifies and removes duplicate rows. Supports exact duplication on all or selected columns, and fuzzy duplication for numeric columns using a rounding tolerance.

Usage

deduplicate(data, cols = NULL, method = "exact",
            key_cols = NULL, tol = 1e-8, max_dist = 1,
            verbose = FALSE)

Arguments

data

A data frame.

cols

The column indices or names used for duplication detection. If NULL and method = "exact", all columns are used. For method = "fuzzy", must be numeric columns.

method

Duplicate detection method. One of "exact" or "fuzzy".

key_cols

Deprecated. Use cols instead.

tol

Tolerance for fuzzy matching. Values are rounded to the number of decimal places indicated by tol before duplicate detection.

max_dist

Not used in current implementation.

verbose

Logical; if TRUE, prints number of duplicates removed.

Details

Fuzzy matching works by rounding numeric columns to a sufficient number of decimal places (derived from tol) and then applying exact duplicate detection.

Value

A data frame with duplicates removed.

Examples

# Exact duplicates on all columns
deduplicate(data[1:200, c(1, 4, 17:19)])
# Exact duplicates on selected columns
deduplicate(data[1:200, c(1, 4, 17:19)], cols = 3:5)
# Fuzzy duplicates with tolerance 0.01
deduplicate(data[1:200, c(1, 4, 17:19)], cols = 3:5, method = "fuzzy", tol = 0.01)

dataprep documentation built on Oct. 1, 2026, 5:07 p.m.