zuhtml-package: zuhtml: parse HTML with a bundled Gumbo parser

zuhtml-packageR Documentation

zuhtml: parse HTML with a bundled Gumbo parser

Description

zuhtml parses real-world HTML the way a browser does, with a bundled copy of the 'Gumbo' parser, and extracts ordinary R objects from it: nodes selected by a documented subset of CSS, attributes, text, lists, tables and links. It needs no system library and has no hard dependencies.

Details

It reads files, URLs and connections, but has no HTTP client of its own, and does not run JavaScript, sanitize, edit documents or support XPath.

Getting started

  • Parse with html_parse(), html_read() or html_fragment().

  • Select elements with html_elements() and html_element(), or move around with html_children() and friends.

  • Read values with html_text_clean(), html_attr() and html_serialize(); extract structures with html_table(), html_list(), html_links() and html_url().

  • Every call runs under html_limits(), and every error is a classed condition: see zuhtml-conditions.

vignette("zuhtml") walks through a complete extraction.

Author(s)

Maintainer: Pedro Baltazar pedrobtz@gmail.com [copyright holder]

Authors:

Other contributors:

  • Google Inc. (Gumbo, bundled in src/vendor/gumbo) [copyright holder]

  • Bjoern Hoehrmann (UTF-8 decoder in src/vendor/gumbo/utf8.c) [copyright holder]

See Also

Useful links:


zuhtml documentation built on Oct. 6, 2026, 5:06 p.m.