View source: R/clean_strings.R
| clean_strings | R Documentation |
Applies common string cleaning operations to character or factor columns: trimming whitespace, changing case, and applying regular expression substitutions. Numeric and other non-character columns are left untouched.
clean_strings(
data,
cols = NULL,
trim = FALSE,
tolower = FALSE,
toupper = FALSE,
pattern = NULL,
replacement = NULL,
verbose = FALSE
)
data |
A data frame, matrix, or character vector. |
cols |
Columns to clean. If |
trim |
Logical; if |
tolower |
Logical; if |
toupper |
Logical; if |
pattern |
Optional regular expression passed to
|
replacement |
Replacement string for |
verbose |
Logical; if |
The operations are applied in a fixed order:
trim (via trimws),
tolower then toupper,
pattern replacement (via gsub).
When both tolower and toupper are TRUE, the
uppercase conversion wins (it is applied last). In practice only
one of the two should be set.
For factor columns, the underlying integer codes are dropped and
the column becomes a character vector. If you need to keep the
factor type, convert the cleaned values back with
factor().
A data frame (or character vector, if the input was a vector) with the selected columns cleaned.
df <- data.frame(
id = 1:3,
name = c(" Alice ", "BOB", "Charlie "),
city = c("New York", "london", "Paris"),
stringsAsFactors = FALSE
)
# Only the character columns are touched by default
clean_strings(df, trim = TRUE, tolower = TRUE)
# Explicit column selection
clean_strings(df, cols = "name", trim = TRUE, toupper = TRUE)
# Regex replacement
clean_strings(df, cols = "city",
pattern = "\\s+", replacement = "_")
# Vector input
clean_strings(c(" A ", " b ", "C"), trim = TRUE, tolower = TRUE)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.