knitr::opts_chunk$set( collapse = TRUE, comment = "#>", fig.align = "center", fig.width = 6, fig.height = 5.5, out.width = "80%", fig.retina = 2 )
library(dataprep) library(ggplot2)
dataprep provides two families of plotting helpers, both built on
top of the C++ descriptive backends:
descplot() — descriptive statistics (n, na, mean, sd,
median, trimmed, min, max, IQR) computed in C++ and
displayed as line or bar charts.
percplot() — top and bottom percentile summaries, useful for
detecting heavy tails and percentile-based outlier cutoffs. The
geometry depends on the column names: lines when they are
numeric, grouped bars otherwise.
Both share the same interface style: a data frame, a numeric range,
and optional grouping. Use data1 (7,640 rows × 7 columns) for
quick demos, and data (7,640 × 65) for full-size examples.
Under the hood, descplot() calls descdata() (which calls
desc_stats_cpp()), and percplot() calls percdata() (which
calls quantile() from base R on each column). Both return a
ggplot object, so all usual ggplot2 layers apply.
When variable names are essentially numeric (e.g. particle
diameters such as 3.16, 3.55, ...), descplot() draws a line
plot with a log-scaled x axis.
descplot(data1, cols = 3:7) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
Pass a subset of statistics by index or by name to focus the plot.
descplot(data1, cols = 3:7, stats = c("na", "min", "max", "IQR")) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
The stats argument accepts both forms — numeric indices
(1:9) and character names ("na", "min", "max", "IQR")
— and can mix them.
When variable names are character (e.g. aerosol mode names
Nucleation, Aitken, Accumulation), descplot() falls back
to a bar chart.
descplot(data1, cols = 3:7) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
descplot(data1, cols = 3:7, stats = c("min", "max", "IQR")) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
descplot(data, cols = 5:65)
If you only need the numbers (not the plot), call descdata()
directly. It returns a data frame with one row per variable and
one column per statistic.
descdata(data1, cols = 3:7, stats = c(2, 3, 4, 7:9))
percplot() computes the extreme percentiles (0 to 0.5 and 99.5 to
100 by default) and draws them against the variable axis. This is
the visual companion to condextr() and percoutl().
percplot(data1, cols = 3:7, group = 2) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
percplot(data1, cols = 3:7, group = 2, part = "top") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
percplot(data1, cols = 3:7, group = 2, part = "bottom") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
When the column names are numeric, percplot() draws curves
instead of bars. The 61 size-bin columns of data (cols = 5:65)
have numeric names, so the default num_xaxis = "auto" selects a
log-scaled x axis:
percplot(data, cols = 5:65, group = 4) + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
For numeric variable names, the x axis can be forced to linear
scale with num_xaxis = "numeric".
percplot(data, cols = 5:65, group = 4, num_xaxis = "numeric") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
The num_xaxis argument controls how the x axis is treated when
column names are numeric:
"auto" (default) — uses a log scale if the column names are
evenly spaced on a log scale with a range ratio of at least
1000; otherwise keeps them as factor levels."log" / "numeric" — force the corresponding scale."character" / "factor" / "keep" / FALSE — keep the
column names as factor levels.percdata() returns the same table that percplot() draws.
percdata(data1, cols = 3:7, group = 2, part = "top")
ggplot2Both descplot() and percplot() return ggplot objects, so all
usual ggplot2 layers apply.
percplot(data1, cols = 3:7, group = 2) + ggplot2::theme_bw(base_size = 11) + ggplot2::labs(title = "Percentile plots by month", x = "Variable", y = "Value") + ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))
A typical diagnostic workflow combines data_report(),
na_diagnose(), and the two plot families:
# 1. Overview of the whole table data_report(data, cols = 5:65) # 2. Per-column NA run statistics na_diagnose(data, cols = 5:65) # 3. Descriptive statistics of the raw data descplot(data, cols = 5:65) # 4. Percentile curves of the raw data percplot(data, cols = 5:65, group = 4)
After running dataprep() you can compare the raw and cleaned
versions in the same plot by stacking them with a g column:
cleaned <- dataprep(data, cols = 5:65, group = 4) percplot( rbind( transform(data[names(cleaned)], g = "original"), transform(cleaned, g = "preprocessed") ), cols = 5:ncol(cleaned), group = ncol(cleaned) + 1 )
Design philosophy — vignette("dataprep-philosophy").
Why the cleaning pipeline has the shape it does.
Cleaning pipeline walkthrough — vignette("dataprep-cleaning").
Step-by-step execution of the four cleaning steps on data.
Performance and cross-engine consistency —
vignette("dataprep-performance"). Benchmark tables and
8-engine consistency checks.
Upgrading from 0.1.5 to 0.1.8 —
vignette("dataprep-migration"). Behaviour changes and the
migration checklist.
Fast reshaping — vignette("dataprep-melt-dcast").
sessionInfo()
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.