| deduplicate | R Documentation |
Identifies and removes duplicate rows. Supports exact duplication on all or selected columns, and fuzzy duplication for numeric columns using a rounding tolerance.
deduplicate(data, cols = NULL, method = "exact",
key_cols = NULL, tol = 1e-8, max_dist = 1,
verbose = FALSE)
data |
A data frame. |
cols |
The column indices or names used for duplication detection. If NULL and |
method |
Duplicate detection method. One of |
key_cols |
Deprecated. Use |
tol |
Tolerance for fuzzy matching. Values are rounded to the number of decimal places indicated by |
max_dist |
Not used in current implementation. |
verbose |
Logical; if |
Fuzzy matching works by rounding numeric columns to a sufficient number of decimal places (derived from tol) and then applying exact duplicate detection.
A data frame with duplicates removed.
# Exact duplicates on all columns
deduplicate(data[1:200, c(1, 4, 17:19)])
# Exact duplicates on selected columns
deduplicate(data[1:200, c(1, 4, 17:19)], cols = 3:5)
# Fuzzy duplicates with tolerance 0.01
deduplicate(data[1:200, c(1, 4, 17:19)], cols = 3:5, method = "fuzzy", tol = 0.01)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.