dbc_to_parquet: Convert a DATASUS DBC file to Parquet format

View source: R/dbcturbo.R

dbc_to_parquetR Documentation

Convert a DATASUS DBC file to Parquet format

Description

A high-level convenience wrapper that first writes the .dbc file to a temporary CSV using the C engine, then converts that CSV into a .parquet file using the arrow package. It therefore requires temporary disk space for the CSV and is not an end-to-end streaming writer.

Usage

dbc_to_parquet(
  input_file,
  output_file,
  batch_size = 8192L,
  encoding = "CP850",
  verbose = FALSE,
  progress = NULL
)

Arguments

input_file

Character string. Path to the source .dbc file.

output_file

Character string. Path for the output .parquet file.

batch_size

Integer. Passed to dbc_to_csv. Default 8192L.

encoding

Character string. Source encoding of character fields. Default "CP850".

verbose

Logical. If TRUE, prints file metrics, a progress bar, and elapsed time upon completion. Default FALSE.

progress

Function or NULL. Optional callback for progress reporting.

Details

This is the recommended workflow for Big Data and epidemiological research, as Parquet files are heavily compressed, columnar, and preserve types.

Value

TRUE invisibly on success. Stops if the arrow package is not installed.

Examples


if (requireNamespace("arrow", quietly = TRUE)) {
  dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
  parquet <- tempfile(fileext = ".parquet")
  dbc_to_parquet(dbc, parquet)
  arrow::read_parquet(parquet)
  unlink(parquet)
}



dbcturbo documentation built on Oct. 10, 2026, 5:08 p.m.