zuhtml parses real-world HTML the way a browser does and turns it into ordinary R values: character vectors, lists and data frames. It bundles the Gumbo HTML5 parser, so it needs no system library, and it has no hard dependencies.
html_problems() lists what was repaired."0012" stays "0012".<meta> tags, JSON-LD, microdata) and forms,
and convert any node to Markdown with html_markdown().encoding, or the page's <meta charset>.html_read() reads a file, a URL or any R connection. zuhtml has no HTTP
client of its own: a URL goes through base R's url(), and a fetcher
that needs headers or authentication hands it the body. zuhtml does not
run JavaScript or sanitize HTML.
install.packages("zuhtml")
The development version, from GitHub:
# install.packages("pak")
pak::pak("pedrobtz/zuhtml")
library(zuhtml)
doc <- html_parse('
<div class=product><h2>Sencha</h2><span class=price>3.50</span>
<a href="sencha.html">details</a></div>
<div class=product><h2>Genmaicha</h2>
<a href="genmaicha.html">details</a></div>',
base_url = "https://example.org/shop/"
)
cards <- html_elements(doc, ".product")
data.frame(
name = html_text_clean(html_element(cards, "h2")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
#> name price url
#> 1 Sencha 3.50 https://example.org/shop/sencha.html
#> 2 Genmaicha <NA> https://example.org/shop/genmaicha.html
html_element() returns one result per card, with a missing value where a
card has no price, so the columns stay aligned.
html_read() reads a file, a connection or a URL. A URL becomes the
document's base URL, so relative links resolve against the page:
doc <- html_read("https://cran.r-project.org/web/views/")
html_title(doc)
#> [1] "CRAN Task Views"
rows <- html_elements(doc, "table tr")
views <- data.frame(
topic = html_text_clean(html_element(rows, "td:nth-child(2)")),
url = html_url(html_element(rows, "a"))
)
head(views, 3)
#> topic url
#> 1 Actuarial Science https://cran.r-project.org/web/views/ActuarialScience.html
#> 2 Agricultural Science https://cran.r-project.org/web/views/Agriculture.html
#> 3 Anomaly Detection https://cran.r-project.org/web/views/AnomalyDetection.html
The getting started guide walks through a complete extraction. There are also guides to selectors, tables and lists, and limits, encodings and safety.
zuhtml is MIT-licensed. The bundled Gumbo parser is Apache-2.0, and
LICENSE.note
explains how the two apply.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.