First release.
Bundles the 'Gumbo' HTML5 parser 0.14.0 from the maintained fork at
https://codeberg.org/gumbo-parser/gumbo-parser, with local patches that
add a parse-time nesting-depth limit, remove the library's only
printf(), and fix two memory-safety bugs in its <selectedcontent>
support that fuzzing found (a use-after-free and a NULL dereference, both
reachable from untrusted HTML) and an uninitialized read in fragment
parsing. No system library is needed.
New html_parse() and html_read() parse HTML from a string, raw vector
or local file under explicit resource limits from the new
html_limits(). Nesting depth is bounded while parsing, and all parser
memory goes through an allocation ledger with a budget, so a failed
parse always releases everything. Raw input is decoded from a byte-order
mark, an explicit encoding, the page's <meta> declaration or UTF-8;
invalid input is an error, never silently replaced.
Parsed documents are converted into a compact immutable tree that holds no reference to the parser or the input. All 1,686 applicable html5lib tree-construction tests that the bundled parser ships produce the expected tree.
New html_fragment() parses markup in the context of a given element,
as innerHTML does.
Nodes are zuhtml_nodesets: vectors of nodes tied to their document,
with missing nodes where an aligned operation has no answer. Navigate
with html_children(), html_parent(), html_ancestors(),
html_next_sibling(), html_previous_sibling(), html_root(),
html_document() and html_template_content(); read values with
html_name(), html_namespace(), html_type(), html_attr(),
html_attrs(), html_classes() and html_text(). html_info()
describes a document. lapply() and friends over a nodeset pass one
node at a time; rep(), rev(), unique() and c() keep nodesets
nodesets.
New extraction functions. html_text_clean() gives text as a reader
wants it: scripts and styles skipped, whitespace collapsed outside
<pre>, line breaks at <br> and block elements. html_list() reads a
<ul>/<ol> as text or a tree, without nested items leaking into their
parents. html_table() and html_tables() read tables into data frames
of character columns, with row and column spans, rowspan="0", row
groups, header detection and an error rather than a silent overwrite
for overlapping cells. html_url() resolves URL attributes with RFC
3986 reference resolution, honouring <base href>, and html_links()
lists a page's links.
New html_elements(), html_element(), html_matches() and
html_filter() select elements with a documented subset of CSS
selectors: type, universal, ID, class and attribute selectors (with the
i and s flags), the four combinators, selector lists, :scope,
:root, :empty, :first-child, :last-child, :only-child,
:nth-child(), :nth-of-type() and :not(). Anything else is a
zuhtml_selector_error pointing at the offending position.
html_element() keeps one result per input node, so extracted columns
stay aligned.
New html_closest() finds each node's nearest ancestor matching a
selector, and html_strings() returns the text pieces html_text()
joins, keeping element boundaries. html_serialize(pretty = TRUE) lays
block-level elements out on indented lines for reading.
New html_title(), html_meta(), html_json_ld() and html_microdata()
read page metadata: the document title, every <meta> tag (OpenGraph,
Twitter cards and Dublin Core included), JSON-LD blocks (parsed with
jsonlite if asked) and microdata items per the HTML standard, itemref
included.
New html_table_cells() returns one row per table cell with its grid
position, spans, section, whether it is a header, and its links.
html_tables() gains match = to keep tables whose text matches a
regular expression. html_table() gains convert =, decimal = and
thousands = to convert columns that are entirely numbers or logicals,
keeping identifiers with leading zeros as text.
New html_markdown() converts nodes to CommonMark: headings,
paragraphs, emphasis, code, fenced code blocks, block quotes, nested
lists, links and images with resolved URLs, and GFM pipe tables for data
tables. Markdown-significant characters in text are escaped, and only
emphasis that CommonMark parses back is written.
The <meta> declaration of raw input is found as browsers find it, by
the HTML standard's prescan of the first 1024 bytes, and read with the
Encoding Standard's labels (so iso-8859-1 means windows-1252).
html_info() reports where the encoding came from in
encoding_source.
New html_forms() describes each form and the controls it owns, by the
HTML standard's form-owner rules (form= references included), with
DOM values for every control type and the options of each select. It
only inspects: nothing is submitted.
html_read() reads a URL (through base R's url(), which also sets the
default base_url) or any connection, such as gzfile(), as well as a
file path. Reading stops once the input passes four times max_input.
New html_serialize() (and as.character() on nodesets) writes nodes
as normalized HTML following the WHATWG serialization algorithm, outer
or inner.
New html_problems() lists the parse errors the parser repaired, with
package-owned codes and positions.
Errors are classed conditions under zuhtml_error; see
?zuhtml-conditions.
New zuhtml_info() reports the bundled parser version and patches, and
self-tests the compiled parser.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.