ffm_batch: Run an FFmpeg Pipeline Over Many Files

View source: R/ffm_batch.R

ffm_batchR Documentation

Run an FFmpeg Pipeline Over Many Files

Description

Apply a pipeline-building function to every row of a jobs table and compile (and optionally run) the resulting FFmpeg command for each. This is the package's main batch function. It gives one reproducible compiled command per job, collected back into a tibble.

Usage

ffm_batch(
  jobs,
  .f,
  ...,
  run = TRUE,
  parallel = FALSE,
  verify = NULL,
  progress = FALSE,
  manifest = FALSE,
  checksums = FALSE
)

Arguments

jobs

A data frame with one row per job. Its column names are the arguments passed to .f.

.f

A function that takes a job's columns (by name) and returns an ffm pipeline object.

...

Additional arguments passed on to every call of .f.

run

A logical: run each compiled command through FFmpeg (TRUE, default) or only compile them for inspection (FALSE, a dry run).

parallel

A logical: map over jobs in parallel with furrr (TRUE) or sequentially (FALSE, default). Parallel runs follow the future plan that you set. With TRUE and the default sequential plan, jobs still run one at a time, and you get a warning. Set a plan first, for example future::plan(future::multisession).

verify

An optional output check applied to each job (only when run = TRUE). Give a named list of expected properties, or a function. A list, for example list(width = 1920), applies the same checks to every job. A function takes the job columns like .f (called pmap-style) and returns such a list for each job. Each job's output is passed to verify_media. Unlike ffm_run, a failed check is recorded, and does not stop the call. Adds a logical verified column (all checks passed), NA for jobs that did not run successfully.

progress

A logical: display a cli progress bar as the jobs run (TRUE) or run quietly (FALSE, default). Only applies when run = TRUE; safe (a no-op animation) in non-interactive sessions.

manifest

A logical. When TRUE (and run = TRUE), the batch records a provenance manifest and attaches it to the result. The manifest has each job's command, the FFmpeg and FFprobe versions, a timestamp and the output size. Read it with ffm_manifest. (default = FALSE)

checksums

A logical: when TRUE, the manifest also captures md5 checksums of each job's input(s) and output. Ignored unless manifest = TRUE. (default = FALSE)

Details

Each column of jobs is passed by name to .f (as purrr::pmap() does), so a job table with columns input, output and start calls .f(input = ..., output = ..., start = ...). .f must return a pipeline (see ffm_files). Give .f a ... argument if jobs carries columns it does not use.

Two jobs whose pipelines write to the same output path are refused before any job runs, under run = FALSE as well as run = TRUE. Paths are compared exactly as written. An output that writes no file may repeat. Such outputs are - (standard output), a pipe: URL, and an output whose last -f option is -f null, as ffm_output_options("-f null") gives.

with_timeout() explains how to limit how long R waits for each program in a job, and what happens when a program reaches the limit.

Value

jobs as a tibble with an added command column, which holds the compiled FFmpeg command for each job. When run = TRUE, it also has a logical success column. When verify is supplied, it also has a verified column. When manifest = TRUE, a provenance manifest is attached as an attribute; read it with ffm_manifest.

When the build lacks an encoder

A job can fail while its pipeline names a codec this FFmpeg build does not list as encodable. The section of the same name on ffm_run says how that is decided. Such a job still runs, and it still gets success = FALSE. After all the jobs run, ffm_batch() gives one warning of class tidymedia_encoder_unavailable for all such jobs. It names each encoder and the jobs that used it. Its fields hold one element for each job and codec: tm_encoder, tm_stream and tm_row, the job's row number in jobs. A task function's _batch form can pass ffm_batch() a table of its own, so these can differ from the rows of the table you gave: separate_audio_video_batch gives ffm_batch() two rows for each input, and normalize_audio_batch(two_pass = TRUE) leaves out silent inputs.

The same class is also an error, for example when a hardware encoder is missing and fallback = FALSE. A tryCatch() handler for the class therefore also catches this warning and loses the batch's result. Use withCallingHandlers() to act on the warning and keep the result.

If reading FFmpeg's encoder lists reaches the tidymedia.timeout limit, the batch gives one warning of class tidymedia_probe_timeout. The later jobs do not read the lists again. With parallel = TRUE, each worker reads them up to once. That warning comes before the tidymedia_encoder_unavailable one, so a tryCatch() handler for it also loses the result; use withCallingHandlers() for it too.

See Also

segment_video(), which is built on ffm_batch(); verify_media() for the verification spec and ffm_manifest() for the provenance manifest.

Other pipeline functions: ffm_codec(), ffm_compile(), ffm_concat(), ffm_copy(), ffm_crop(), ffm_drawbox(), ffm_drop(), ffm_files(), ffm_fps(), ffm_hstack(), ffm_jobs(), ffm_loudnorm(), ffm_map(), ffm_output_options(), ffm_overlay(), ffm_pixel_format(), ffm_run(), ffm_scale(), ffm_seek(), ffm_trim(), ffm_vstack(), print.tidymedia_ffm()

Examples

video <- system.file("extdata", "sample.mp4", package = "tidymedia")
jobs <- tibble::tibble(
  input  = c(video, video),
  output = c("a.mp3", "b.mp3")
)
# run = FALSE compiles one command per job without calling FFmpeg
ffm_batch(jobs, run = FALSE, .f = function(input, output, ...) {
  ffm_files(input, output) |>
    ffm_drop("video") |>
    ffm_codec(audio = "libmp3lame")
})

tidymedia documentation built on Oct. 11, 2026, 5:08 p.m.