dataprep: descriptive statistics and diagnostic plots

knitr::opts_chunk$set(
  collapse  = TRUE,
  comment   = "#>",
  fig.align = "center",
  fig.width = 6,
  fig.height = 5.5,
  out.width = "80%",
  fig.retina = 2
)
library(dataprep)
library(ggplot2)

Overview

dataprep provides two families of plotting helpers, both built on top of the C++ descriptive backends:

Both share the same interface style: a data frame, a numeric range, and optional grouping. Use data1 (7,640 rows × 7 columns) for quick demos, and data (7,640 × 65) for full-size examples.

Under the hood, descplot() calls descdata() (which calls desc_stats_cpp()), and percplot() calls percdata() (which calls quantile() from base R on each column). Both return a ggplot object, so all usual ggplot2 layers apply.

Descriptive statistics

Line plot, numeric variable names

When variable names are essentially numeric (e.g. particle diameters such as 3.16, 3.55, ...), descplot() draws a line plot with a log-scaled x axis.

descplot(data1, cols = 3:7) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Selected statistics

Pass a subset of statistics by index or by name to focus the plot.

descplot(data1, cols = 3:7,
         stats = c("na", "min", "max", "IQR")) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

The stats argument accepts both forms — numeric indices (1:9) and character names ("na", "min", "max", "IQR") — and can mix them.

Bar chart, character variable names

When variable names are character (e.g. aerosol mode names Nucleation, Aitken, Accumulation), descplot() falls back to a bar chart.

descplot(data1, cols = 3:7) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Control facet layout

descplot(data1, cols = 3:7, stats = c("min", "max", "IQR")) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Full-size data

descplot(data, cols = 5:65)

The underlying table

If you only need the numbers (not the plot), call descdata() directly. It returns a data frame with one row per variable and one column per statistic.

descdata(data1, cols = 3:7, stats = c(2, 3, 4, 7:9))

Percentile plots

Full percentile range

percplot() computes the extreme percentiles (0 to 0.5 and 99.5 to 100 by default) and draws them against the variable axis. This is the visual companion to condextr() and percoutl().

percplot(data1, cols = 3:7, group = 2) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Top percentiles only

percplot(data1, cols = 3:7, group = 2, part = "top") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Bottom percentiles only

percplot(data1, cols = 3:7, group = 2, part = "bottom") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Percentile curves of the raw data

When the column names are numeric, percplot() draws curves instead of bars. The 61 size-bin columns of data (cols = 5:65) have numeric names, so the default num_xaxis = "auto" selects a log-scaled x axis:

percplot(data, cols = 5:65, group = 4) +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

Numeric axis control

For numeric variable names, the x axis can be forced to linear scale with num_xaxis = "numeric".

percplot(data, cols = 5:65, group = 4, num_xaxis = "numeric") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

The num_xaxis argument controls how the x axis is treated when column names are numeric:

The underlying table

percdata() returns the same table that percplot() draws.

percdata(data1, cols = 3:7, group = 2, part = "top")

Combining with ggplot2

Both descplot() and percplot() return ggplot objects, so all usual ggplot2 layers apply.

percplot(data1, cols = 3:7, group = 2) +
  ggplot2::theme_bw(base_size = 11) +
  ggplot2::labs(title = "Percentile plots by month",
                x = "Variable", y = "Value") +
  ggplot2::theme(axis.text.x = ggplot2::element_text(angle = 30, hjust = 1))

A diagnostic workflow

A typical diagnostic workflow combines data_report(), na_diagnose(), and the two plot families:

# 1. Overview of the whole table
data_report(data, cols = 5:65)

# 2. Per-column NA run statistics
na_diagnose(data, cols = 5:65)

# 3. Descriptive statistics of the raw data
descplot(data, cols = 5:65)

# 4. Percentile curves of the raw data
percplot(data, cols = 5:65, group = 4)

After running dataprep() you can compare the raw and cleaned versions in the same plot by stacking them with a g column:

cleaned <- dataprep(data, cols = 5:65, group = 4)

percplot(
  rbind(
    transform(data[names(cleaned)], g = "original"),
    transform(cleaned,              g = "preprocessed")
  ),
  cols  = 5:ncol(cleaned),
  group = ncol(cleaned) + 1
)

Where to go next

Session info

sessionInfo()


Try the dataprep package in your browser

Any scripts or data that you put into this service are public.

dataprep documentation built on Oct. 1, 2026, 5:07 p.m.