pclm: Univariate Penalized Composite Link Model (PCLM)

View source: R/pclm_1D.R

pclmR Documentation

Univariate Penalized Composite Link Model (PCLM)

Description

Fit univariate penalized composite link model (PCLM) to ungroup binned count data, e.g. age-at-death distributions grouped in age classes.

Usage

pclm(
  x,
  y,
  nlast = NULL,
  offset = NULL,
  out.step = 1,
  ci.level = 95,
  verbose = FALSE,
  control = list(),
  omega = NULL,
  na.action = c("fail", "omit")
)

Arguments

x

Vector containing the starting values of the input intervals/bins. For example: if we have 3 bins [0,5), [5,10) and [10, 15), x will be defined by the vector: c(0, 5, 10).

y

Vector with counts to be ungrouped. It must have the same dimension as x.

nlast

Length of the last interval. In the example above nlast would be 5.

offset

Optional offset term to calculate smooth mortality rates. A vector of the same length as x and y, or one of the same length as the ungrouped output. See \insertCiterizzi2015;textualungroup for further details.

out.step

Length of estimated intervals in output. Values between 0.1 and 1 are accepted. Default: 1.

ci.level

Confidence level, as a percentage rather than a proportion, so 95 and not 0.95. Values in [50.1, 99.9] are accepted. It sets the width of both interval pairs in ci: the pointwise conf_lower and conf_upper, and the mass-preserving lower and upper scenarios. Default: 95.

verbose

Logical value. Indicates whether a progress bar should be shown or not. Default: FALSE.

control

List of fitting controls, given by name. A misspelled entry is an error rather than being silently ignored. An unnamed entry is matched positionally, so list(100) sets lambda and nothing else; naming every entry is strongly preferred. Any setting not supplied takes its default from control.pclm for this function, or control.pclm2D for pclm2D. See those pages for the meaning and default of each of lambda, kr, deg, int.lambda, diff, opt.method, max.iter and tol.

omega

Closing age of the distribution. An alternative to nlast: when it is given, the width of the last interval is taken as omega - max(x). Give one of the two, not both.

na.action

What to do with unobserved cells in y. "fail", the default, rejects them. "omit" drops the matching rows and lets the smoothing penalty bridge the gap, which is how a surface with interior gaps or a missing year is handled. Only NA counts as unobserved; infinite values are always an error.

Details

The PCLM method is based on the composite link model, which extends standard generalized linear models. It implements the idea that the observed counts, interpreted as realizations from Poisson distributions, are indirect observations of a finer (ungrouped) but latent sequence. This latent sequence represents the distribution of expected means on a fine resolution and has to be estimated from the aggregated data. Estimates are obtained by maximizing a penalized likelihood. This maximization is performed efficiently by a version of the iteratively reweighted least-squares algorithm. Optimal values of the smoothing parameter are chosen by minimizing Bayesian or Akaike's Information Criterion.

Value

The output is a list with the following components:

input

A list with arguments provided in input. Saved for convenience.

fitted

The fitted values of the PCLM model.

ci

A list with two kinds of interval and they are not interchangeable. lower and upper are the two mass-conserving scenarios: the distribution implied by a uniformly lower and a uniformly higher hazard, each rescaled so that it totals sum(fitted). Because the total is held fixed, the low scenario moves deaths towards older ages and the two curves cross fitted in the tail. They are therefore scenarios, not pointwise bounds, and they carry no coverage level. conf_lower and conf_upper are the pointwise marginal ci.level interval for the fitted values, computed as fitted * exp(-/+ qnorm * SE). These do satisfy conf_lower <= fitted <= conf_upper and they do not total sum(fitted). Use the first pair as low and high inputs to a life table, the second as a pointwise error bar on the estimate.

goodness.of.fit

A list containing goodness of fit measures: standard errors, AIC and BIC.

smoothPar

Estimated smoothing parameters. In the univariate model a named vector of lambda, kr and deg. In the two-dimensional model lambda splits into lambda.x for the age axis and lambda.y for the year axis, so the vector is lambda.x, lambda.y, kr, deg.

bin.definition

Additional values to identify the bins limits and location in input and output objects.

deep

A list of objects created in the fitting process. Useful in diagnosis of possible issues.

call

An unevaluated function call, that is, an unevaluated expression which consists of the named function applied to the given arguments.

References

\insertAllCited

See Also

control.pclm plot.pclm

Examples

# Data  
x <- c(0, 1, seq(5, 85, by = 5))
y <- c(294, 66, 32, 44, 170, 284, 287, 293, 361, 600, 998, 
       1572, 2529, 4637, 6161, 7369, 10481, 15293, 39016)
offset <- c(114, 440, 509, 492, 628, 618, 576, 580, 634, 657, 
            631, 584, 573, 619, 530, 384, 303, 245, 249) * 1000
nlast <- 26 # the size of the last interval

# Example 1 ----------------------
M1 <- pclm(x, y, nlast)
ls(M1)
summary(M1)
fitted(M1)
plot(M1)

# Example 2 ----------------------
# ungroup even in smaller intervals
M2 <- pclm(x, y, nlast, out.step = 0.5)
head(fitted(M2))
plot(M2, type = "s")
# Note, in example 1 we are estimating intervals of length 1. In example 2 
# we are estimating intervals of length 0.5 using the same aggregate data.

# Example 3 ----------------------
# Do not optimise smoothing parameters; choose your own. Faster.
M3 <- pclm(x, y, nlast, out.step = 0.5, 
           control = list(lambda = 100, kr = 10, deg = 10))
plot(M3)

summary(M2)
summary(M3) # not the smallest BIC here, but sometimes is not important.

# Example 4 -----------------------
# Grouped x & grouped offset (estimate death rates)
M4 <- pclm(x, y, nlast, offset)
plot(M4, type = "s")

# Example 5 -----------------------
# Grouped x & ungrouped offset (estimate death rates)

ungroupped_Ex <- pclm(x, y = offset, nlast, offset = NULL)$fitted # ungroupped offset data

M5 <- pclm(x, y, nlast, offset = ungroupped_Ex)

ungroup documentation built on Oct. 3, 2026, 5:06 p.m.