| html_text_clean | R Documentation |
Text as a reader would want it from a page, rather than exactly as
parsed. It is not a browser's innerText: there is no layout or CSS, and
hidden elements are included. The rules are fixed and documented:
text inside <script>, <style> and <template> is skipped, and so
are comments;
runs of whitespace (space, tab, newline, carriage return, form feed)
collapse to a single space, except inside <pre>, <textarea>,
<listing> and <plaintext>, where whitespace is kept;
<br> is a line break; each block element is a line break before and
after, and adjacent block boundaries give a single line break. The
block elements are address, article, aside, blockquote, body,
caption, center, dd, details, dialog, dir, div, dl,
dt, fieldset, figcaption, figure, footer, form, h1 to
h6, head, header, hgroup, hr, html, legend, li,
listing, main, menu, nav, ol, optgroup, option, p,
plaintext, pre, section, summary, table, tbody, tfoot,
thead, title, tr, ul and xmp;
table cells (<td>, <th>) are separated by a space.
html_text_clean(x, trim = TRUE, nbsp = TRUE)
x |
A |
trim |
If |
nbsp |
If |
A character vector as long as x: the cleaned text of each
element, document, fragment or text node; NA for other nodes and for
missing nodes.
html_text() for the text exactly as parsed.
Other node values:
html_attr(),
html_markdown(),
html_name(),
html_serialize(),
html_strings(),
html_text()
doc <- html_parse(paste0(
"<div><h1>Title</h1>\n <p>Some <b>bold</b> text.<br>New line",
"<script>ignored()</script></p><pre> kept\n as is</pre></div>"
))
div <- html_element(doc, "div")
cat(html_text_clean(div))
html_text(div)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.