pip_collect: Collect step outputs

View source: R/pipeflow.R

pip_collectR Documentation

Collect step outputs

Description

Returns the outputs of the steps in a pipeline or view. With by = "step" (the default) the result is a named list of step outputs. With any other column, the outputs are grouped by the values of that column (typically "tags").

Usage

pip_collect(x, by = "step", as.table = FALSE, simplify = TRUE)

pip_collect_out(x, by = "step", as.table = FALSE, simplify = TRUE)

Arguments

x

A pipeflow pip or view.

by

Single step-table column name to group by.

as.table

If TRUE, return a data.table instead of a named list.

simplify

If TRUE (default), if the list of collected outputs contains exactly one step per group, the result is flattened by one level, otherwise it is returned as a grouped list.

Details

The result always takes one of two shapes:

  • Flat: a named list whose elements are the step outputs directly, e.g. list(s1 = 1, s2 = 2).

  • Grouped: a named list whose elements are themselves named lists of step outputs, one per group, e.g. list(io = list(s1 = 1, s2 = 2), model = list(s3 = 4)).

With simplify = TRUE (the default) the result is flat whenever every group contains exactly one step, which, for example, is always the case for by = "step" as step names are unique. If any group contains more than one step (or simplify = FALSE) the result is grouped. For list columns such as tags, a step with several entries contributes its output to every corresponding group, and steps without an entry (e.g. untagged steps) are omitted. You can use pip_view() to further narrow the selection before collecting.

Value

A named list of outputs

Lifecycle

Deprecated

pip_collect_out() is a legacy alias for pip_collect(). It raises a deprecation warning and will be removed in a future release.

Examples

p <- pip_new() |>
  pip_add("load", \(x = 1) x, tags = "io") |>
  pip_add("clean", \(x = ~load) x + 1, tags = "io") |>
  pip_add("model", \(x = ~clean) x * 2, tags = "model")
pip_run(p)

# By default, a flat named list with one entry per step
pip_collect(p)

# The same output as a data.table
pip_collect(p, as.table = TRUE)

# Group the outputs by tag ...
pip_collect(p, by = "tags")

# ... which is equivalent to
list(
  io = pip_view(p, tags = "io") |> pip_collect(),
  model = pip_view(p, tags = "model") |> pip_collect()
)

# Grouped table output
pip_collect(p, by = "tags", as.table = TRUE)

# Keep single-step groups nested
pip_collect(p, simplify = FALSE)

# Collect output from a view
v <- p[step %in% c("clean", "model"), ]
pip_collect(v)
pip_collect(v, as.table = TRUE)

pipeflow documentation built on Sept. 28, 2026, 1:06 a.m.