html_text: Text content of nodes

View source: R/text.R

html_textR Documentation

Text content of nodes

Description

html_text() is the structural text of each node: the text nodes in its subtree, concatenated in tree order, exactly as parsed. It inserts no separators, trims nothing and keeps ⁠<script>⁠ and ⁠<style>⁠ text; it does not include comments or the contents of ⁠<template>⁠ elements.

Usage

html_text(x, recursive = TRUE)

Arguments

x

A zuhtml_document or zuhtml_nodeset.

recursive

If FALSE, only the direct text children of each node.

Value

A character vector as long as x. An element with no text is ""; a text, comment or processing-instruction node is its own content; a doctype and a missing node are NA.

See Also

Other node values: html_attr(), html_markdown(), html_name(), html_serialize(), html_strings(), html_text_clean()

Examples

doc <- html_parse("<p>Hello <b>big</b> world</p>")
p <- html_children(html_children(html_root(doc))[2])
html_text(p)
html_text(p, recursive = FALSE)

zuhtml documentation built on Oct. 6, 2026, 5:06 p.m.