html_list: Extract an HTML list

View source: R/list.R

html_listR Documentation

Extract an HTML list

Description

Reads one ⁠<ul>⁠ or ⁠<ol>⁠ element. In "text" mode, the result has one string per item: the item's cleaned text (see html_text_clean()), without the text of any list nested inside it, so child items do not leak into their parent. A nested list still separates the text around it with a line break. In "tree" mode, nested lists become children of the item that contains them, including lists inside wrapper elements such as ⁠<div>⁠.

Usage

html_list(x, mode = c("text", "tree"))

Arguments

x

A zuhtml_nodeset holding exactly one ⁠<ul>⁠ or ⁠<ol>⁠ element, as from html_element().

mode

"text" or "tree".

Details

Items are the ⁠<li>⁠ children of the list, in source order; empty items are "" and duplicates are kept.

Value

  • "text": a character vector, one string per item.

  • "tree": an object of class zuhtml_list, a list with type ("ul" or "ol") and items, a list with one element per item; each item is a list with text (as in text mode) and children (a list of zuhtml_list objects, one per list nested in the item).

See Also

Other extraction: html_forms(), html_links(), html_table(), html_table_cells(), html_url()

Examples

doc <- html_parse("<ul><li>Apples<li>Tools<ul><li>Hammer<li>Saw</ul></ul>")
items <- html_element(doc, "ul")
html_list(items)
html_list(items, mode = "tree")

zuhtml documentation built on Oct. 6, 2026, 5:06 p.m.