read_dbc: Read a DATASUS DBC file into an R data frame

View source: R/dbcturbo.R

read_dbcR Documentation

Read a DATASUS DBC file into an R data frame

Description

High-level convenience wrapper. The native engine decodes small files directly into R columns without temporary files. Larger files use a temporary CSV and data.table::fread() (if available) or utils::read.csv().

Usage

read_dbc(
  file,
  batch_size = 4096L,
  encoding = "CP850",
  verbose = FALSE,
  cols = NULL,
  coerce_types = TRUE,
  engine = c("auto", "native", "data.table", "base"),
  native_threshold = 50 * 1024^2,
  ...
)

Arguments

file

Character string. Path to the .dbc file.

batch_size

Integer. Passed to dbc_to_csv. Default 4096L.

encoding

Character string. Source encoding of character fields. Default "CP850" (standard DATASUS legacy encoding).

verbose

Logical. Passed to dbc_to_csv. Default FALSE.

cols

Character vector or NULL. Names of the columns to keep in the returned object. NULL (default) returns all columns. Column names are validated against dbc_inspect metadata and selection is applied before the temporary CSV is written.

coerce_types

Logical. If TRUE (default), columns are automatically converted to their native R types based on the DBF field metadata:

  • D (date) → Date (DBF stores dates as YYYYMMDD strings; blank or "00000000" become NA).

  • N with decimals > 0 → numeric.

  • N with decimals == 0 → integer when values fit in R integers, otherwise numeric.

  • L (logical) → logical ("T"/"Y"/"S"/"1" → TRUE; "F"/"N"/"0" → FALSE; others → NA).

  • C (character) → unchanged.

engine

Character string selecting the reader. "auto" (default) uses the native engine for files no larger than native_threshold, then data.table when installed and R base otherwise. "native" always uses direct native decoding; "data.table" requires data.table; "base" always uses utils::read.csv().

native_threshold

Non-negative number of bytes. In "auto" mode, files at or below this size use the native engine. Defaults to 50 MiB. Set to 0 to always use the CSV route in "auto".

...

Additional arguments forwarded to the CSV reader.

Details

For files with more than one million records, use dbc_to_csv directly and load the resulting CSV with arrow::read_csv_arrow() or data.table::fread() for maximum performance.

Value

A data.frame or data.table (if data.table is installed).

Examples

dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
df <- read_dbc(dbc)
head(df)

# Read a subset of fields.
fields <- dbc_inspect(dbc)$fields$name[1:3]
selected <- read_dbc(dbc, cols = fields)
names(selected)


dbcturbo documentation built on Oct. 10, 2026, 5:08 p.m.