| Extract.rwl | R Documentation |
Subset a rwl object by series (columns) and years (rows), keeping
the class and the provenance record of the data.
## S3 method for class 'rwl'
x[i, j, drop]
## S3 method for class 'rwl'
subset(x, subset, select, drop = FALSE, ...)
x |
an object of class |
i, j |
indices for the rows (years) and columns (series), used
exactly as in |
drop |
logical. As in |
subset, select |
for the |
... |
Not used. |
Subsetting an rwl object works as it does for a
data.frame, with i indexing years and j
indexing series. Four things are different.
Dropping series returns the years the remaining series cover. A
collection is ragged, and the years at the ends of an rwl object
usually belong to a few long series, so dropping some of the series
ordinarily leaves years in which nothing that remains was measured. The result runs
from the first year to the last year in which any of the selected series
has a measurement. Empty years between those two are kept: an
rwl object holds one row per year, and an interior year cannot be
dropped without breaking the sequence.
This applies to calls that drop series without indexing years –
x[, j], x[j], and subset given only select.
Two kinds of call leave the years alone. A call that indexes years
returns exactly the years it asks for, measured or not: x[i, ],
x[i, j], head, tail,
common.interval, and subset with its subset
argument. A year window is how an rwl object is aligned with
something else, such as a climate series or a second collection, and
returning fewer years than were asked for would put that alignment out.
And a call that keeps every series, such as the reordering
x[, order(names(x))] or x[], has dropped nothing to trim
for, whatever empty years the object already held.
Two results are not trimmed. A subset holding no measurement in any year
is returned whole rather than emptied, and rwl.check reports
it. A single series taken with drop = TRUE is a numeric
vector, which carries no years at all, so it is returned at full length
and stays aligned with the object it came from; use drop = FALSE
for a one-series rwl object trimmed to its own years.
subset drops empty series; [ keeps them. Taking
years can leave a series with no values at all: in
ca533[time(ca533) %in% 1800:1899, ], the series CAM132, CAM152,
CAM161 and CAM201 end before 1800 and are left as columns of nothing
but NA. An empty column is not a series. chron,
rwi.stats and the plots skip it, and
summary.rwi, interseries.cor,
corr.rwl.seg and detrend either drop it or
fail on it.
subset and window.rwl therefore drop any series
left with no values, and a message names the series dropped. If no
series is left with values, that is an error. Use one of these to
take a span of years for analysis.
[ keeps every series it is asked for, empty or not, on purpose.
x[i, j] always returns the columns that j names, in that
order, so the result stays lined up with anything indexed alongside it:
a vector of series names, an ids table from
read.ids, or a second object with the same series. If
[ dropped columns, those would fall out of step without a word,
and code that indexes with [ – dplR's own included – could no
longer count on names(x[, j]) being the series it asked for.
So row subsets made with [, and head and
tail, can return empty series. Drop them with
window or subset, or find them with
colSums(!is.na(x)) == 0.
Years must stay consecutive. An rwl object holds one row
per year in increasing order, and dplR reads the year of a measurement
from its row name, so time.rwl, plot.rwl,
detrend, chron, rwl.stats and
the crossdating functions all take the next row to be the next year. A
subset such as x[c(1, 5, 9), ] leaves that untrue. Where row
subsetting leaves years that are not consecutive and increasing, because
rows were taken from the middle, repeated, or reordered, the result is
returned as a plain data.frame, with a warning, and without the
provenance record. The measurements are unchanged; what is withdrawn is
the claim that they are a set of dated series.
The provenance record is carried and cut to fit. The
"dplR.provenance" attribute attached by read.tucson
follows the subset. What describes the file – its name, the reader, the
header lines – is kept. What describes series – the precision of each,
the renames, the interior gaps, the problems found while parsing – is
kept for the series that remain, and gaps are dropped where they fall
outside the years that remain. mixed.precision is recomputed,
since a file measured at two precisions can be subset down to series
measured at one. A subset element is added, holding the series in
the result, its first and last year, and all.series, which is
TRUE when every series the file held is still present. A record
that cannot be matched to a series in the result is dropped rather than
kept, so an object whose series have been renamed since it was read
carries a record that says nothing about them in preference to one that
says something wrong.
common.interval cuts an object down to the years its series
share, which is a different question from the years they cover
between them.
An object of class c("rwl", "data.frame"), or a numeric
vector when one column is selected and drop is TRUE, or a
data.frame when row subsetting has left the years no longer
consecutive. When series are dropped and years are not indexed, the
result runs only over the years the remaining series cover. [
returns every series asked for, including any with no values;
subset drops those, with a message.
Andy Bunn
read.rwl, read.tucson, rwl.check,
common.interval, time.rwl,
subset.data.frame
library(utils)
data(ca533)
## Series are columns and years are rows. Selecting series drops the years
## that none of them covers: ca533 begins in 626, but these five series do
## not.
ca533.sub <- ca533[, 1:5]
class(ca533.sub)
range(time(ca533))
range(time(ca533.sub))
## Naming years returns those years, whether or not they hold data.
dim(ca533[time(ca533) %in% 1000:1099, 1:5])
## subset() works the same way.
range(time(subset(ca533, select = 1:5)))
## Years are subset by name or by time().
ca533.1800s <- ca533[time(ca533) %in% 1800:1899, ]
range(time(ca533.1800s))
## Four series end before 1800. `[` keeps them as empty columns, so the
## result still has the 34 series asked for; subset() and window() drop
## them and say which.
ncol(ca533.1800s)
names(ca533.1800s)[colSums(!is.na(ca533.1800s)) == 0]
ncol(subset(ca533, time(ca533) %in% 1800:1899))
ncol(window(ca533, 1800, 1899))
## A single series comes back as a vector unless drop is FALSE.
class(ca533[, 1])
class(ca533[, 1, drop = FALSE])
## The provenance record of the data follows the subset.
data(wa082)
gap.series <- attr(wa082, "dplR.provenance")$gaps$series
wa082.sub <- wa082[, gap.series, drop = FALSE]
attr(wa082.sub, "dplR.provenance")$gaps
attr(wa082.sub, "dplR.provenance")$subset$all.series
## Rows taken out of the middle break the one-row-per-year promise that
## dplR relies on. This warns, and gives back a data.frame.
not.rwl <- ca533[c(1, 5, 9), ]
class(not.rwl)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.