Extract.rwl: Subset an rwl Object

Extract.rwlR Documentation

Subset an rwl Object

Description

Subset a rwl object by series (columns) and years (rows), keeping the class and the provenance record of the data.

Usage

## S3 method for class 'rwl'
x[i, j, drop]

## S3 method for class 'rwl'
subset(x, subset, select, drop = FALSE, ...)

Arguments

x

an object of class "rwl" with series as columns and years as rows, such as that produced by read.rwl.

i, j

indices for the rows (years) and columns (series), used exactly as in [.data.frame: numbers, negative numbers, logicals, or names. The year of a row is its name, so x["1900", ] selects a year and x[time(x) %in% 1800:1900, ] selects a window of them.

drop

logical. As in [.data.frame: if the result has one column and drop is TRUE, a numeric vector is returned instead of a one-series rwl object. Note that subset defaults to drop = FALSE, as subset.data.frame does.

subset, select

for the subset method, expressions naming the years (rows) to keep and the series (columns) to keep, used exactly as in subset.data.frame.

...

Not used.

Details

Subsetting an rwl object works as it does for a data.frame, with i indexing years and j indexing series. Four things are different.

Dropping series returns the years the remaining series cover. A collection is ragged, and the years at the ends of an rwl object usually belong to a few long series, so dropping some of the series ordinarily leaves years in which nothing that remains was measured. The result runs from the first year to the last year in which any of the selected series has a measurement. Empty years between those two are kept: an rwl object holds one row per year, and an interior year cannot be dropped without breaking the sequence.

This applies to calls that drop series without indexing years – x[, j], x[j], and subset given only select. Two kinds of call leave the years alone. A call that indexes years returns exactly the years it asks for, measured or not: x[i, ], x[i, j], head, tail, common.interval, and subset with its subset argument. A year window is how an rwl object is aligned with something else, such as a climate series or a second collection, and returning fewer years than were asked for would put that alignment out. And a call that keeps every series, such as the reordering x[, order(names(x))] or x[], has dropped nothing to trim for, whatever empty years the object already held.

Two results are not trimmed. A subset holding no measurement in any year is returned whole rather than emptied, and rwl.check reports it. A single series taken with drop = TRUE is a numeric vector, which carries no years at all, so it is returned at full length and stays aligned with the object it came from; use drop = FALSE for a one-series rwl object trimmed to its own years.

subset drops empty series; [ keeps them. Taking years can leave a series with no values at all: in ca533[time(ca533) %in% 1800:1899, ], the series CAM132, CAM152, CAM161 and CAM201 end before 1800 and are left as columns of nothing but NA. An empty column is not a series. chron, rwi.stats and the plots skip it, and summary.rwi, interseries.cor, corr.rwl.seg and detrend either drop it or fail on it.

subset and window.rwl therefore drop any series left with no values, and a message names the series dropped. If no series is left with values, that is an error. Use one of these to take a span of years for analysis.

[ keeps every series it is asked for, empty or not, on purpose. x[i, j] always returns the columns that j names, in that order, so the result stays lined up with anything indexed alongside it: a vector of series names, an ids table from read.ids, or a second object with the same series. If [ dropped columns, those would fall out of step without a word, and code that indexes with [ – dplR's own included – could no longer count on names(x[, j]) being the series it asked for. So row subsets made with [, and head and tail, can return empty series. Drop them with window or subset, or find them with colSums(!is.na(x)) == 0.

Years must stay consecutive. An rwl object holds one row per year in increasing order, and dplR reads the year of a measurement from its row name, so time.rwl, plot.rwl, detrend, chron, rwl.stats and the crossdating functions all take the next row to be the next year. A subset such as x[c(1, 5, 9), ] leaves that untrue. Where row subsetting leaves years that are not consecutive and increasing, because rows were taken from the middle, repeated, or reordered, the result is returned as a plain data.frame, with a warning, and without the provenance record. The measurements are unchanged; what is withdrawn is the claim that they are a set of dated series.

The provenance record is carried and cut to fit. The "dplR.provenance" attribute attached by read.tucson follows the subset. What describes the file – its name, the reader, the header lines – is kept. What describes series – the precision of each, the renames, the interior gaps, the problems found while parsing – is kept for the series that remain, and gaps are dropped where they fall outside the years that remain. mixed.precision is recomputed, since a file measured at two precisions can be subset down to series measured at one. A subset element is added, holding the series in the result, its first and last year, and all.series, which is TRUE when every series the file held is still present. A record that cannot be matched to a series in the result is dropped rather than kept, so an object whose series have been renamed since it was read carries a record that says nothing about them in preference to one that says something wrong.

common.interval cuts an object down to the years its series share, which is a different question from the years they cover between them.

Value

An object of class c("rwl", "data.frame"), or a numeric vector when one column is selected and drop is TRUE, or a data.frame when row subsetting has left the years no longer consecutive. When series are dropped and years are not indexed, the result runs only over the years the remaining series cover. [ returns every series asked for, including any with no values; subset drops those, with a message.

Author(s)

Andy Bunn

See Also

read.rwl, read.tucson, rwl.check, common.interval, time.rwl, subset.data.frame

Examples

library(utils)
data(ca533)

## Series are columns and years are rows. Selecting series drops the years
## that none of them covers: ca533 begins in 626, but these five series do
## not.
ca533.sub <- ca533[, 1:5]
class(ca533.sub)
range(time(ca533))
range(time(ca533.sub))

## Naming years returns those years, whether or not they hold data.
dim(ca533[time(ca533) %in% 1000:1099, 1:5])

## subset() works the same way.
range(time(subset(ca533, select = 1:5)))

## Years are subset by name or by time().
ca533.1800s <- ca533[time(ca533) %in% 1800:1899, ]
range(time(ca533.1800s))

## Four series end before 1800. `[` keeps them as empty columns, so the
## result still has the 34 series asked for; subset() and window() drop
## them and say which.
ncol(ca533.1800s)
names(ca533.1800s)[colSums(!is.na(ca533.1800s)) == 0]
ncol(subset(ca533, time(ca533) %in% 1800:1899))
ncol(window(ca533, 1800, 1899))

## A single series comes back as a vector unless drop is FALSE.
class(ca533[, 1])
class(ca533[, 1, drop = FALSE])

## The provenance record of the data follows the subset.
data(wa082)
gap.series <- attr(wa082, "dplR.provenance")$gaps$series
wa082.sub <- wa082[, gap.series, drop = FALSE]
attr(wa082.sub, "dplR.provenance")$gaps
attr(wa082.sub, "dplR.provenance")$subset$all.series

## Rows taken out of the middle break the one-row-per-year promise that
## dplR relies on. This warns, and gives back a data.frame.
not.rwl <- ca533[c(1, 5, 9), ]
class(not.rwl)

dplR documentation built on Oct. 1, 2026, 9:07 a.m.