html_strings: Text pieces of nodes

View source: R/text.R

html_stringsR Documentation

Text pieces of nodes

Description

The text nodes of each node's subtree, in tree order, as separate strings: the pieces html_text() concatenates. Boundaries between elements are kept, which matters when the markup, not whitespace, separates values (⁠<td>1</td><td>2</td>⁠ is "1", "2", not "12"). Text inside ⁠<script>⁠, ⁠<style>⁠ and ⁠<template>⁠ is skipped, as are comments.

Usage

html_strings(x, trim = FALSE, drop_empty = FALSE)

Arguments

x

A zuhtml_document or zuhtml_nodeset.

trim

If TRUE, remove leading and trailing whitespace from each piece.

drop_empty

If TRUE, drop pieces that are empty (after trimming, when trim = TRUE).

Value

A list as long as x of character vectors; NA_character_ for a missing node. A text node is its own single piece.

See Also

Other node values: html_attr(), html_markdown(), html_name(), html_serialize(), html_text(), html_text_clean()

Examples

doc <- html_parse("<p>One <b>two</b>\n  <i> three </i></p>")
p <- html_element(doc, "p")
html_strings(p)
html_strings(p, trim = TRUE, drop_empty = TRUE)

zuhtml documentation built on Oct. 6, 2026, 5:06 p.m.