| html_parse | R Documentation |
html_parse() parses one string or raw vector of HTML the way a browser
does: omitted end tags, unquoted attributes, character references and
misnested elements are repaired by the HTML parsing algorithm, never
rejected. html_read() reads and parses one file, URL or connection.
html_parse(
x,
encoding = NULL,
base_url = NULL,
comments = TRUE,
limits = html_limits()
)
html_read(path, ...)
x |
One string, or a raw vector of encoded bytes. |
encoding |
The encoding of raw input, as a name |
base_url |
The document's URL, used to resolve relative links; |
comments |
Whether to keep comment nodes. |
limits |
Resource limits from |
path |
One file path, URL or connection; see "Reading files, URLs and connections". |
... |
Arguments passed on to |
html_parse() never treats a string as a file name or URL.
A zuhtml_document.
html_read() reads its input as raw bytes, then decodes and parses them
as html_parse() does:
a string with an http, https, ftp, ftps or file scheme (in
any case) is read with url(), and becomes the document's base_url
unless base_url is given. A redirect is not seen, so after one the
base URL is the address asked for. HTTP headers are not read either:
the encoding comes from a byte-order mark, encoding or the page's
<meta>. For anything more (headers, authentication, retries), fetch
with an HTTP client and pass the body to html_parse();
any other string is a file path;
a connection, such as gzfile() or rawConnection(), is read as
readBin() reads one: an unopened connection is opened in "rb" mode
and closed afterwards; an open one must be in binary mode, is read
from its current position, and is left open. A non-blocking pipe or
socket that has no data yet is an error rather than a short read.
Reading stops with a zuhtml_limit_error as soon as the input passes
four times max_input bytes, before a larger file is read at all. A
file that does not exist, an unreachable URL, and a connection that
cannot be read are zuhtml_input_errors. The input is always read
whole before parsing, because decoding needs all of it.
A character string is taken as text: it is converted to UTF-8 with
enc2utf8(), and encoding must be NULL or "UTF-8". A string marked
as "bytes" is rejected; pass a raw vector instead.
A raw vector is decoded with, in order of precedence, a byte-order mark
(UTF-8, UTF-16LE or UTF-16BE), encoding, a declaration in the page, or
UTF-8. A byte-order mark that contradicts encoding is an error, and so
is any byte sequence that is invalid in the chosen encoding: nothing is
replaced silently. A fetcher that knows the HTTP charset should pass it
as encoding.
The declaration is found as browsers find it, by the HTML standard's
prescan of the first 1024 bytes for <meta charset="..."> or
<meta http-equiv="Content-Type" content="...; charset=...">. The
prescan skips comments and the insides of tags, but not the text of
scripts. Labels are those of the Encoding Standard, which maps several
to a superset: "iso-8859-1", "latin1" and "us-ascii" mean
windows-1252, "gb2312" means GBK, and a UTF-16 label means UTF-8 (the
bytes read as ASCII, so they are not UTF-16). An unknown label is
ignored. html_info() reports the encoding used and its source.
A file saved in another encoding without updating its declaration, as
some tools do when they convert pages to UTF-8, decodes wrongly or fails
to decode, as it would in a browser. Pass its real encoding as
encoding.
A leading byte-order mark is removed. Input containing a NUL character after decoding is rejected.
html_problems() for the parse errors that were repaired;
html_limits(); zuhtml-conditions for the errors these functions
raise.
Other parsing:
html_fragment(),
html_info(),
html_limits(),
html_problems(),
zuhtml_info()
doc <- html_parse("<p>Hello <b>world</b>")
doc
# Raw bytes in a declared encoding:
html_parse(as.raw(c(0x3c, 0x70, 0x3e, 0xe9)), encoding = "latin1")
# From a file:
path <- tempfile(fileext = ".html")
writeLines("<title>Saved page</title><p>Text", path)
html_read(path)
# From a compressed file, through a connection:
gz <- tempfile(fileext = ".html.gz")
writeLines("<p>Compressed", gzfile(gz))
html_read(gzfile(gz))
# From a URL, when online; relative links resolve against it:
if (interactive()) {
doc <- html_read("https://cran.r-project.org/web/packages/")
head(html_links(doc, absolute = TRUE))
}
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.