| zuhtml-package | R Documentation |
zuhtml parses real-world HTML the way a browser does, with a bundled copy of the 'Gumbo' parser, and extracts ordinary R objects from it: nodes selected by a documented subset of CSS, attributes, text, lists, tables and links. It needs no system library and has no hard dependencies.
It reads files, URLs and connections, but has no HTTP client of its own, and does not run JavaScript, sanitize, edit documents or support XPath.
Parse with html_parse(), html_read() or html_fragment().
Select elements with html_elements() and html_element(), or move
around with html_children() and friends.
Read values with html_text_clean(), html_attr() and
html_serialize(); extract structures with html_table(),
html_list(), html_links() and html_url().
Every call runs under html_limits(), and every error is a classed
condition: see zuhtml-conditions.
vignette("zuhtml") walks through a complete extraction.
Maintainer: Pedro Baltazar pedrobtz@gmail.com [copyright holder]
Authors:
Pedro Baltazar pedrobtz@gmail.com [copyright holder]
Other contributors:
Google Inc. (Gumbo, bundled in src/vendor/gumbo) [copyright holder]
Bjoern Hoehrmann (UTF-8 decoder in src/vendor/gumbo/utf8.c) [copyright holder]
Useful links:
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.