View source: R/clean_url_columns.R
| clean_url_columns | R Documentation |
Applies 'rurl::get_clean_url' to specified columns of a data frame. URLs are cleaned under pagerankr's explicit canonicalization profile (see Details), with any arguments in '...' overriding individual knobs.
clean_url_columns(data_frame, columns = c("from", "to"), ...)
data_frame |
A data frame containing URL columns to be cleaned. |
columns |
A character vector specifying the names of the columns containing URLs. Defaults to 'c("from", "to")'. |
... |
'rurl::get_clean_url' arguments that override the canonicalization profile per key. Recognized knobs are the ones [canonical_profile()] pins: 'protocol_handling', 'case_handling', 'www_handling', 'trailing_slash_handling', 'index_page_handling', 'path_normalization', 'scheme_relative_handling', 'subdomain_levels_to_keep', 'host_encoding', 'path_encoding', 'url_standard', 'port_handling', 'query_handling', 'params_keep', 'params_drop', 'params_case_sensitive', 'sort_params', 'empty_param_handling', 'decode_plus'. Note 'url_standard' governs 'case_handling' and 'path_normalization', so 'rurl' rejects an override of either while the profile pins a standard selector. |
The canonicalization profile ([canonical_profile()]) pins every 'rurl' knob explicitly so node identities do not depend on ‘rurl'’s own (version-dependent) defaults, and keeps the cleaning and domain-filtering paths symmetrical. Most knobs equal ‘rurl'’s current defaults; six override them because they shape node identity – 'path_normalization', 'path_encoding', 'url_standard', 'port_handling', 'host_encoding' and 'query_handling'. In particular 'path_encoding = "keep"' holds that presentation dial at its only identity-preserving value, and 'url_standard = "whatwg"' is where the path-identity semantics actually live. See [canonical_profile()] for details.
NA values in the specified columns are preserved in the output. Downstream functions in the pagerankr workflow (such as get_unique_edges and pagerank) will automatically drop any edge where either from or to is NA.
Tokens that 'rurl' cannot parse as a URL (e.g. a dotless bare label such as '"A"', which newer 'rurl' normalizes to NA) are left as their raw input value rather than becoming NA. This keeps unparseable but non-missing node identities as opaque nodes instead of silently dropping them, so an odd URL in a crawl is scored as its own node rather than vanishing. Only genuinely missing (NA) inputs stay NA. This raw-fallback is a deliberate, accepted divergence from the sibling 'semantic' project (which drops such inputs); see [canonical_profile()] for why it does not desync the cross-repo node join.
A data frame with the specified URL columns cleaned.
df <- data.frame(
from = c(
"http://example.com/path",
"HTTPS://Example.com/PATH#frag", NA,
"http://example.com/path"
),
to = c(
"www.another.com?q=1", "another.com/?q=1&b=2",
"http://foo.bar", NA
),
other_col = 1:4
)
cleaned_df <- clean_url_columns(df, columns = c("from", "to"))
print(cleaned_df)
# Pass extra arguments to rurl::get_clean_url via ...
cleaned_df_custom <- clean_url_columns(
df,
columns = c("from", "to"),
protocol_handling = "http"
)
print(cleaned_df_custom)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.