Parses real-world 'HTML' with a bundled copy of the 'Gumbo' parser (<https://codeberg.org/gumbo-parser/gumbo-parser>), which follows the 'WHATWG' parsing algorithm, so that no system library is required. Documents become immutable trees navigated with a documented subset of 'CSS' selectors. Attributes, text, lists, tables, links, forms and page metadata ('JSON-LD', microdata) are extracted into ordinary character vectors, lists and data frames, and nodes convert to 'Markdown'. Input is a string, raw bytes, a file, a URL or a connection, and raw input is decoded as browsers decode it, from a byte-order mark or a '<meta>' declaration. Parsing is bounded by limits on input size, native memory and nesting depth.
Package details |
|
|---|---|
| Author | Pedro Baltazar [aut, cre, cph], Google Inc. [cph] (Gumbo, bundled in src/vendor/gumbo), Bjoern Hoehrmann [cph] (UTF-8 decoder in src/vendor/gumbo/utf8.c) |
| Maintainer | Pedro Baltazar <pedrobtz@gmail.com> |
| License | MIT + file LICENSE |
| Version | 0.1.0 |
| URL | https://github.com/pedrobtz/zuhtml https://pedrobtz.github.io/zuhtml/ |
| Package repository | View on CRAN |
| Installation |
Install the latest version of this package by entering the following in R:
|
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.