| html_url | R Documentation |
Reads a URL-valued attribute from each node and resolves it against the document's base URL with the reference-resolution algorithm of RFC 3986, section 5 (https://www.rfc-editor.org/rfc/rfc3986#section-5): dot segments are removed and relative, query-only, fragment-only and protocol-relative references are handled. Nothing is fetched.
html_url(x, attr = "href", base_url = NULL)
x |
A |
attr |
The attribute holding the URL. |
base_url |
The base URL to resolve against, overriding the
document's; |
The base URL is base_url if given; otherwise the document's first
<base href> resolved against the base_url given to html_parse();
otherwise that base_url alone.
This is RFC 3986, not the WHATWG URL Standard browsers implement: there is no IDNA processing, no percent-encoding of characters that need it, and no special handling of backslashes. Leading and trailing ASCII whitespace is removed from the attribute; a reference that still contains whitespace or a control character, or whose scheme is invalid, is malformed.
A character vector as long as x: the resolved URL of each
node. NA where the node lacks the attribute, the reference is
malformed, or a relative reference has no absolute base URL to resolve
against; and for missing nodes. The original attribute value stays
available through html_attr().
Other extraction:
html_forms(),
html_links(),
html_list(),
html_table(),
html_table_cells()
doc <- html_parse(
"<a href='../img/a.png'>A</a><a href='//cdn.example.org/b'>B</a>",
base_url = "https://example.org/docs/page.html"
)
html_url(html_elements(doc, "a"))
html_url(html_elements(doc, "a"), base_url = "http://other.test/x/")
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.