clean_url_columns: Clean URL Columns in a Data Frame

View source: R/clean_url_columns.R

clean_url_columnsR Documentation

Clean URL Columns in a Data Frame

Description

Applies 'rurl::get_clean_url' to specified columns of a data frame. URLs are cleaned under pagerankr's explicit canonicalization profile (see Details), with any arguments in '...' overriding individual knobs.

Usage

clean_url_columns(data_frame, columns = c("from", "to"), ...)

Arguments

data_frame

A data frame containing URL columns to be cleaned.

columns

A character vector specifying the names of the columns containing URLs. Defaults to 'c("from", "to")'.

...

'rurl::get_clean_url' arguments that override the canonicalization profile per key. Recognized knobs are the ones [canonical_profile()] pins: 'protocol_handling', 'case_handling', 'www_handling', 'trailing_slash_handling', 'index_page_handling', 'path_normalization', 'scheme_relative_handling', 'subdomain_levels_to_keep', 'host_encoding', 'path_encoding', 'url_standard', 'port_handling', 'query_handling', 'params_keep', 'params_drop', 'params_case_sensitive', 'sort_params', 'empty_param_handling', 'decode_plus'. Note 'url_standard' governs 'case_handling' and 'path_normalization', so 'rurl' rejects an override of either while the profile pins a standard selector.

Details

The canonicalization profile ([canonical_profile()]) pins every 'rurl' knob explicitly so node identities do not depend on ‘rurl'’s own (version-dependent) defaults, and keeps the cleaning and domain-filtering paths symmetrical. Most knobs equal ‘rurl'’s current defaults; six override them because they shape node identity – 'path_normalization', 'path_encoding', 'url_standard', 'port_handling', 'host_encoding' and 'query_handling'. In particular 'path_encoding = "keep"' holds that presentation dial at its only identity-preserving value, and 'url_standard = "whatwg"' is where the path-identity semantics actually live. See [canonical_profile()] for details.

NA values in the specified columns are preserved in the output. Downstream functions in the pagerankr workflow (such as get_unique_edges and pagerank) will automatically drop any edge where either from or to is NA.

Tokens that 'rurl' cannot parse as a URL (e.g. a dotless bare label such as '"A"', which newer 'rurl' normalizes to NA) are left as their raw input value rather than becoming NA. This keeps unparseable but non-missing node identities as opaque nodes instead of silently dropping them, so an odd URL in a crawl is scored as its own node rather than vanishing. Only genuinely missing (NA) inputs stay NA. This raw-fallback is a deliberate, accepted divergence from the sibling 'semantic' project (which drops such inputs); see [canonical_profile()] for why it does not desync the cross-repo node join.

Value

A data frame with the specified URL columns cleaned.

Examples

df <- data.frame(
  from = c(
    "http://example.com/path",
    "HTTPS://Example.com/PATH#frag", NA,
    "http://example.com/path"
  ),
  to = c(
    "www.another.com?q=1", "another.com/?q=1&b=2",
    "http://foo.bar", NA
  ),
  other_col = 1:4
)
cleaned_df <- clean_url_columns(df, columns = c("from", "to"))
print(cleaned_df)

# Pass extra arguments to rurl::get_clean_url via ...
cleaned_df_custom <- clean_url_columns(
  df,
  columns = c("from", "to"),
  protocol_handling = "http"
)
print(cleaned_df_custom)


pagerankr documentation built on Oct. 1, 2026, 5:09 p.m.