knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(zuhtml)
HTML from the web is untrusted input. zuhtml bounds the work and memory
any page can cost, per call, with html_limits():
html_limits()
max_input is checked before parsing. max_memory bounds the parser's
native memory and the document it builds; it is the real guard, since
some markup costs far more memory per byte than other markup.max_depth bounds nesting while parsing. Tree construction takes time
quadratic in nesting depth, so a check after parsing would come too
late: 100,000 nested elements would take 16 seconds. With the limit it
fails at once.max_nodes, max_errors (parse problems kept), max_table_cells and
max_selector_length bound the rest.Exceeding a limit is a zuhtml_limit_error, raised after every native
allocation has been released. It says which limit and by how much:
err <- tryCatch( html_parse(strrep("<div>", 1e5)), zuhtml_limit_error = function(e) e ) conditionMessage(err) err$limit
Tighter limits suit a service that parses pages from strangers:
strict <- html_limits(max_input = 2 * 1024^2, max_memory = 64 * 1024^2, max_depth = 128) doc <- html_parse("<p>Small page</p>", limits = strict)
A string is already text: it is used as UTF-8. Raw bytes are decoded, in
order of preference, with a byte-order mark, the encoding you give, the
page's own <meta> declaration, or UTF-8:
bytes <- as.raw(c(0x3c, 0x70, 0x3e, 0x63, 0x61, 0x66, 0xe9)) # "<p>caf\xe9" html_text_clean(html_parse(bytes, encoding = "latin1"))
Invalid input is an error, never silently replaced:
try(html_parse(bytes))
The declaration is found as a browser finds it, by scanning the first
1024 bytes for <meta charset> or its http-equiv form. Labels mean what
they mean to browsers, so iso-8859-1 is read as windows-1252, which
makes byte 0x93 a curly quote rather than a control character:
page <- c(charToRaw("<meta charset=iso-8859-1><p>"), as.raw(0x93), charToRaw("Quoted"), as.raw(0x94)) doc <- html_parse(page) html_text_clean(doc) html_info(doc)[c("encoding", "encoding_source")]
When you fetch a page, pass the charset from the HTTP Content-Type
header as encoding: it takes precedence over the page's declaration, as
it does in a browser. A byte-order mark takes precedence over both; one
that contradicts encoding is an error.
Handle errors by class, never by message text:
tryCatch( html_elements(html_parse("<p>"), "p:hover"), zuhtml_selector_error = function(e) paste("unsupported at", e$position) )
See ?zuhtml-conditions for the classes and their fields.
<script>
elements, event-handler attributes and javascript: URLs. Do not treat
html_serialize() output as safe to embed in another page.html_read() reads a URL with base R's url(),
without headers, cookies or retries; html_url() is string arithmetic
and fetches nothing.html_text_clean() follows fixed, documented rules.Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.