html_attr: Attributes of elements

View source: R/attributes.R

html_attrR Documentation

Attributes of elements

Description

html_attr() reads one attribute from every node; html_attrs() reads all of them; html_classes() splits the class attribute into tokens. Values are decoded (⁠&⁠ becomes &). An attribute that is present but empty is "", distinct from an absent one: test a boolean attribute such as disabled by presence, !is.na(html_attr(x, "disabled")).

Usage

html_attr(x, name, default = NA_character_)

html_attrs(x)

html_classes(x)

Arguments

x

A zuhtml_document or zuhtml_nodeset.

name

The attribute name: a single string.

default

The value for nodes that lack the attribute, including nodes that are not elements: a single string, possibly NA.

Details

Attribute names on HTML elements match regardless of ASCII case, as in a browser; on SVG and MathML elements they match exactly (viewBox). Namespaced attributes in foreign content are named with their prefix, as in "xlink:href". Where an element has the same attribute more than once, the first wins, as the HTML parser decides.

Value

  • html_attr(): a character vector as long as x; NA for missing nodes.

  • html_attrs(): a list as long as x of named character vectors, in source order; NA_character_ for missing nodes.

  • html_classes(): a list as long as x of character vectors of class tokens; NA_character_ for missing nodes.

See Also

Other node values: html_markdown(), html_name(), html_serialize(), html_strings(), html_text(), html_text_clean()

Examples

doc <- html_parse(
  "<a href='/x' class='btn  primary' data-id=7>Go</a><input disabled>"
)
body <- html_children(html_root(doc))[2]
nodes <- html_children(body)
html_attr(nodes, "href")
html_attr(nodes, "href", default = "")
html_attrs(nodes)
html_classes(nodes)
!is.na(html_attr(nodes, "disabled"))

zuhtml documentation built on Oct. 6, 2026, 5:06 p.m.