dataprep: upgrading from 0.1.5 to 0.1.8

knitr::opts_chunk$set(collapse = TRUE, comment = "#>",
                      fig.width = 6, fig.height = 5.5,
                      out.width = "75%", fig.retina = 2)
library(dataprep)

What changed

Three behaviour changes affect row and column counts. All three are bug fixes, but each one changes the output on real data. If your downstream analysis depends on exact row counts, read the quantified comparison in the last section before upgrading.

Interface changes

Besides the behaviour changes above, several functions changed their argument interface between 0.1.5 and 0.1.8. The old start / end pair was replaced by a single cols argument (character names, integer indices, or logical mask):

| Function | 0.1.5 arguments | 0.1.8 arguments | |----------------|------------------------|-----------------| | varidele | start, end | cols | | obsedele | start, end | cols | | condextr | start, end | cols | | percoutl | start, end | cols | | optisolu | start, end | cols | | dataprep | start, end | cols | | descdata | start, end | cols | | descplot | start, end | cols | | percdata | start, end | cols | | percplot | start, end | cols | | shorvalu | start, end | cols | | melt | cols (id columns) | id / measure.vars |

A call such as varidele(data, 3, 15) was interpreted as start = 3, end = 15 in 0.1.5, but in 0.1.8 it is parsed as cols = 3 (the 15 is dropped or becomes fraction). You must rewrite it as varidele(data, cols = 3:15). The same applies to every function in the table.

The full description is in news(package = "dataprep"). The design reasoning behind the pipeline is in vignette("dataprep-philosophy").

Minimal reproduction of the boundary change

Consider five observations sampled every 10 minutes, with a valid value only at the two ends:

df <- data.frame(
  date = as.POSIXct("2024-01-01 00:00:00", tz = "UTC") + 0:4 * 600,
  x    = c(1, NA, NA, NA, 5)
)

obsedele(df, cols = "x", half = 30)

The three interior rows are each 10, 20, or 30 minutes from the nearest valid anchor. Under 0.1.5 the third row was deleted because the delete condition was inclusive (dl >= half); under 0.1.8 it is retained because the delete condition is strict (dl > half && dr > half), so dl == half satisfies "within half minutes".

Minimal reproduction of the multi-column change

Consider two channels with mutually exclusive missing runs:

df <- data.frame(
  date = as.POSIXct("2024-01-01 00:00:00", tz = "UTC") + 0:9 * 600,
  x    = c(1, NA, NA, NA, NA, NA, NA, NA, NA, 5),
  y    = c(NA, NA, NA, NA, 2, NA, NA, NA, NA, NA)
)
nrow(obsedele(df, cols = c("x", "y"), half = 60))

Under 0.1.5, the x and y NA runs were merged before computing the run length. This changed the run boundaries and could either over-delete rows or retain rows that should have been deleted. Under 0.1.8 the two channels are checked independently: a row is deleted when any selected column has a run longer than half minutes on both sides.

Implementation evolution

The retention criterion has been stable across releases, but the implementation has improved steadily. The table below summarises the three generations:

| Release | Strategy | Complexity | |---|---|---| | 0.1.0 | Borrowed running mean: expand the series onto a regular grid with tidyr::complete(), compute a 59-minute centred moving average on a temporary column, and use its emptiness pattern to flag long runs. | O(grid length) per subset | | 0.1.5 | Run-length encoding: use data.table::rleid() and rowid() to collapse consecutive NAs into runs. The retention threshold was half * num grid rows, where num is the leading number in the by string: with by = "5 min", half = 30 this was a 150-minute window. | O(n) time, O(n) temporary storage | | 0.1.8 | Anchor scan: for each missing value, look up the nearest non-missing anchor on each side and compare the two time distances directly. No grid, no run-length state. | O(n) time, O(1) extra allocation per column |

Each generation produces the same deletion decision on the same input, but the constant factors shrink. The 0.1.8 anchor scan is the first version that is fast enough to run interactively on full-year data: on SMEAR I Varrio 2025 (49,422 rows × 61 numeric channels), obsedele() now runs in 0.05 s against 11.6 s in 0.1.5 on Ubuntu 25.10 (about 232×), and in 0.035 s against 22.5 s on Windows 11 Pro for Workstations (about 648×). See vignette("dataprep-performance") for the full benchmark.

Quantified effect on a full-year dataset

On SMEAR I Varrio 2025 (49,422 rows × 61 numeric channels, 10-minute sampling), running the same pipeline with the same parameters:

| Stage | 0.1.5 | 0.1.8 | Δ | |---|---:|---:|---:| | varidele | 25 columns deleted | 25 columns deleted | 0 | | obsedele | 1,494 rows deleted | 1,496 rows deleted | +2 | | condextr | 1,868 rows deleted | 1,863 rows deleted | −5 | | shorvalu | 50,376 NAs filled | 50,387 NAs filled | +11 | | dataprep final | 46,060 rows | 46,063 rows | +3 |

Net change: 0.006% of the input. The six rows that differ between versions all sit at run boundaries where the anchor distance is within one sampling interval of half minutes.

Where the difference comes from

The six differing rows fall into two groups:

The net result is 3 additional rows in the final output.

Migration checklist

result_015 <- dataprep::dataprep(data, cols = 5:65, group = 4)
result_017 <- dataprep(data, cols = 5:65, group = 4)

kept_only_by_017 <- dplyr::anti_join(result_017, result_015, by = "date")
kept_only_by_015 <- dplyr::anti_join(result_015, result_017, by = "date")

What did not change

For users who only call melt() and dcast(), or who only use dataprep for descriptive statistics (descdata, na_diagnose, percdata, percplot, descplot), there is no behaviour change. Those functions have been re-implemented in 0.1.8 for speed (melt and dcast) or reorganised internally (data_report, dry_run), but the output on every tested input is identical to 0.1.5 except where noted in the news file.

Performance reference

The 0.1.8 release rewrites every heavy cleaning routine in C++. The table below compares against 0.1.5 on three dataset sizes from the same source (SMEAR I Varrio forest). All numbers are speed-up ratios (0.1.5 time / 0.1.8 time); a value below 1.0× means 0.1.8 is slightly slower on that cell.

| Function | 500 rows | 7,640 rows | 49,422 rows (Ubuntu 25.10) | |---|---:|---:|---:| | varidele | 1.1× | 1.1× | 11.6× | | obsedele | 203× | 424× | 232× | | condextr | 196× | 217× | 1146× | | optisolu | 188× | 77× | 109× | | dataprep | 185× | 228× | 247× |

On Windows 11 Pro for Workstations, the same full-year pipeline gives obsedele ≈ 648×, condextr ≈ 839×, shorvalu ≈ 81×, optisolu ≈ 25× (at cores = 32), and the integrated dataprep call ≈ 173×. varidele is around 1.17× on this cell; this is expected, since varidele is a single colMeans(is.na(.)) in both versions and the new code path has little room for improvement.

Note on optisolu cores. The 0.1.5 implementation could crash when cores > 16. The benchmark above used cores = 16 for both versions to keep the comparison fair. 0.1.8 loads the package on each worker, exports the input data once per worker, and runs each (interval, times) case as a separate task, so cores = 64 is safe. The practical speed-up on a many-core host is larger than the table above. Also note that the optimal parameter values returned by optisolu() may differ slightly between versions; re-run optisolu() after upgrading if exact values matter.

Test environments

Benchmarks and checks were run on two reference hosts. Only the core configuration is listed here; full details are in vignette("dataprep-performance").

Where to go next

Session info

sessionInfo()


Try the dataprep package in your browser

Any scripts or data that you put into this service are public.

dataprep documentation built on Oct. 1, 2026, 5:07 p.m.