knitr::opts_chunk$set( comment = "#", prompt = FALSE, tidy = FALSE, cache = FALSE, collapse = TRUE ) old <- options(width = 100L)
Pipelines can get long, and often you want to focus on a subset of steps: a topic, a stage, or the steps that produce the outputs you care about. Views are {pipeflow}'s way of working on a subset of steps without copying anything. A view references the underlying pipeline, so every operation applied to a view (running it, updating parameters, tagging, locking, ...) writes through to the original pipeline, restricted to the steps covered by the view.
This vignette shows how to create and combine views, how to select steps
with the [ operator, and how to run only part of a pipeline.
Again, we use a very simplified example of a pipeline.
library(pipeflow) pip <- pip_new("my-pip") |> pip_add("load", \(n = 5) seq_len(n), tags = c("io", "daily")) |> pip_add("clean", \(x = ~load) x * 2, tags = c("io", "core")) |> pip_add("fit", \(x = ~clean) sum(x), tags = c("model", "core")) |> pip_add("report", \(x = ~fit) paste("result:", x), tags = "report")
As you see above, besides step and fun, pip_add() also allows to set
tags. We will use these tags as meta information to filter certain steps by
topic and/or output type. Let's do a first run before we move on.
(pip_run(pip, lgr = NULL))
pip_view() returns a view of the pipeline that contains only the
steps matching the given filters:
pip_view(pip, tags = "core")
Filters can be combined. By default, steps must match all filters (logical AND), while the values within a single filter are treated as alternatives (OR):
pip_view(pip, tags = "core", state = "done") pip_view(pip, step = c("clean", "fit"))
join = "union" keeps steps that match any filter:
pip_view(pip, tags = "report", step = "clean", join = "union")
With fixed = FALSE, filter values are interpreted as regular
expressions:
pip_view(pip, step = "^f", fixed = FALSE)
The available filters are step, params, state, exec, tags, and
depends. For example, to find all steps that depend on load and are
still new:
pip_reset(pip) # reset to initial state pip_view(pip, depends = "load", state = "new")
[The extract operator [ provides a data.table-like way of selecting
steps. It returns a view by default:
pip_run(pip, lgr = NULL) pip[c("load", "fit")] pip[1:3]
Boolean filters are evaluated in the context of the step table, so the
same columns as in pip_view() are available as variables:
pip[tags %like% "core"] pip[step %in% c("clean", "fit") & state == "done"]
Negative indices select all steps except the excluded ones:
pip[-2]
p[] returns a copy of the pipeline, and using two indices
(p[i, j]) extracts a the given rows and columns as a data.table:
pip2 <- pip[] pip[, c("step", "tags")] pip[c("load", "fit"), "out"]
While [ returns a view by default, view = FALSE builds a new,
self-contained pipeline containing the selected steps together with all
their upstream dependencies:
pip[c("fit", "report"), view = FALSE]
The printed message tells you how many steps were pulled in as upstream dependencies.
Views can be nested: applying pip_view() (or [) to a view narrows
the view further.
v1 <- pip_view(pip, tags = "io") # load, clean v1 v2 <- v1 |> pip_view(tags = "core") # clean only v2
view meta fieldUnder the hood, a view is defined by a vector of row indices covered by the
view (NULL for a "no view"). This vector is stored in the view meta
field of the pipeline object, so in principle you can also manipulate
the view directly by assigning to the meta field^[
Direct manipulation of the view usually is not needed and probably
mostly useful for debugging.
].
pip[["view"]] # NULL — not a view w <- pip w[["view"]] <- c("load", "clean") w w[["view"]] <- NULL # back to the full pipeline w
Running a view executes the covered steps together with any upstream
dependencies that are not up to date. The run log marks steps that belong
to the view as [view] and steps that were pulled in as dependencies as
[upstream]:
pip_reset(pip) pip_run(pip_view(pip, step = "report"))
Afterwards, the original pipeline is up to date for the covered steps.
Pipeline views are not only useful to inspect or run certain parts of the pipeline but also to filter and collect the final output of your analysis run. For more details on this see the next vignette Collect and group output.
options(old)
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.