percoutl: Traditional percentile-based outlier removal

View source: R/percoutl.R

percoutlR Documentation

Traditional percentile-based outlier removal

Description

Removes values above a top percentile and below a bottom percentile. Thresholds are computed from quantiles. Optionally, observation deletion based on consecutive missing values can be performed after outlier removal.

Usage

percoutl(data, cols = NULL, group = NULL, top = 0.995,
         bottom = 0.0025, by = "min", half = 30,
         date_col = NULL, cores = NULL, verbose = FALSE)

Arguments

data

A data frame or matrix. If a numeric vector is supplied, only outlier marking (no observation deletion) is performed because there is no time axis.

cols

Column indices or names of numeric variables. If NULL, all numeric columns are used.

group

Optional grouping column for group-wise outlier removal.

top

Top percentile threshold. Values above this quantile are removed.

bottom

Bottom percentile threshold. Values below this quantile are removed.

by

Time unit for observation deletion (see obsedele).

half

Half window size, in minutes, for consecutive missing deletion.

date_col

Time column index or name. If NULL, automatically detected.

cores

Number of OpenMP threads. Passed to obsedele.

verbose

Logical; if TRUE, prints timing and deletion counts.

Details

This method is a one-size-fits-all approach and may remove non-outliers or fail to remove some outliers. It is provided for comparison with condextr, which uses a point-by-point weighted conditional extremum criterion.

Value

A data frame with outliers removed.

Author(s)

Chun-Sheng Liang <chun-shengliang@qq.com>

References

1. Example data is from https://smear.avaa.csc.fi/download. It includes particle number concentrations in SMEAR I Varrio forest.

Examples

percoutl(obsedele(data[1:500, c(1, 4, 17:19)], cols = 3:5, group = 2),
         cols = 3:5, group = 2)

dataprep documentation built on Oct. 1, 2026, 5:07 p.m.