README.md

rbiogeme

rbiogeme is an R interface to Biogeme. Model expressions and data transformations are written in R, while Biogeme's native Python engine performs expression compilation, likelihood evaluation, differentiation, integration, optimization, simulation, and reporting.

The package is designed so that an R user can complete the full workflow from R. Python objects and Python callbacks are not part of the ordinary user API.

Development attribution

The whole package was implemented by ChatGPT 5.6 Luna under the supervision of Michel Bierlaire.

Installation

Install a source tarball downloaded from the release page or built from this repository:

install.packages(
  "/path/to/rbiogeme_0.1.2.tar.gz",
  repos = NULL,
  type = "source"
)

When working from a checkout, build the tarball first with R CMD build ..

The package requires R 4.3 or later and Python 3.12 or later. The default native requirement is currently biogeme==3.3.5. If no interpreter is selected, biogeme_setup() asks reticulate to provision an isolated environment for that requirement. If you already manage Python yourself, pass its executable explicitly to biogeme_setup().

If Biogeme is already installed in a Python environment, configure that interpreter before constructing a model or initializing the bridge:

library(rbiogeme)
biogeme_setup(python = "/absolute/path/to/python")

The configuration is session-wide and must be set before Python is initialized. Use biogeme_setup() for the first session, or biogeme_check() when the runtime is already configured. Use biogeme_diagnostics() to inspect the active interpreter and the native packages visible to it.

First session: from installation to results

For a first session, copy the following block into a fresh R session. biogeme_setup() provisions or verifies the configured native requirement and reports whether the environment is ready before any model is estimated.

library(rbiogeme)

check <- biogeme_setup()
if (!check$ready) {
  print(check)
  stop("The rbiogeme environment is not ready.")
}

database <- biogeme_database(
  "first_session",
  data.frame(
    choice = c(1, 2, 1, 2, 1, 2),
    time = c(10, 8, 12, 7, 11, 9),
    cost = c(5, 7, 6, 8, 5, 7)
  )
)

beta_time <- biogeme_beta("beta_time", start = 0)
beta_cost <- biogeme_beta("beta_cost", start = 0)
asc_2 <- biogeme_beta("asc_2", start = 0)

model <- logit_model(
  database = database,
  choice = "choice",
  utilities = list(
    `1` = beta_time * variable("time") + beta_cost * variable("cost"),
    `2` = asc_2 + beta_time * variable("time") +
      beta_cost * variable("cost")
  )
)

validation <- validate_model(model)
stopifnot(validation$valid)

output_directory <- tempfile("rbiogeme-first-session-")
dir.create(output_directory)
fit <- estimate(
  model,
  model_name = "rbiogeme_first_session",
  control = biogeme_control(
    output_directory = output_directory,
    generate_html = FALSE,
    generate_yaml = FALSE,
    save_iterations = FALSE
  )
)

summary(fit)
coef(fit)
predict(fit)

There are two supported environment paths:

| Situation | What to do | | --- | --- | | You do not already manage Python | Run biogeme_setup(). It asks reticulate to provision an isolated environment containing biogeme==3.3.5. | | You already have a Python environment | Install biogeme==3.3.5 into that environment, then run biogeme_setup(python = "/absolute/path/to/python"). |

Selecting an existing interpreter does not install Biogeme into it. For a manually managed environment, the installation step is performed outside R, for example with the environment's own package installer:

python3.12 -m venv /path/to/rbiogeme-venv
/path/to/rbiogeme-venv/bin/python -m pip install "biogeme==3.3.5"

Then configure that interpreter in R before any operation that initializes Python:

library(rbiogeme)
biogeme_setup(python = "/path/to/rbiogeme-venv/bin/python")

biogeme_setup() does not estimate a model. It provisions or selects the runtime, verifies the R and Python versions, confirms that Biogeme can be imported, and reports the corrective action when a check fails. Use biogeme_check() when the runtime is already configured and only a readiness check is needed. Use biogeme_diagnostics() when the detailed version list is needed.

A complete model in R

The following example creates a small two-alternative logit model. The model specification is entirely visible in R.

library(rbiogeme)

database <- biogeme_database(
  "demo",
  data.frame(
    choice = c(1, 2, 1, 2, 1, 2, 1, 2),
    income = c(1, 2, 1, 3, 2, 1, 3, 2)
  )
)

asc_2 <- biogeme_beta("asc_2", start = 0)
b_income <- biogeme_beta("b_income", start = 0)

model <- logit_model(
  database = database,
  choice = "choice",
  utilities = list(
    `1` = 0,
    `2` = asc_2 + b_income * variable("income")
  )
)

fit <- estimate(
  model,
  model_name = "rbiogeme_demo",
  control = biogeme_control(
    generate_html = FALSE,
    generate_yaml = FALSE,
    save_iterations = FALSE
  )
)

summary(fit)
coef(fit)
vcov(fit)
logLik(fit)

biogeme_beta() creates a named native parameter and variable() creates a data-variable node. Arithmetic and logical operators build a neutral expression tree; they do not evaluate the model in R. The complete tree is compiled once by the bridge before native Biogeme starts numerical work.

Main workflow

  1. Put numeric observations in a data.frame.
  2. Create a biogeme_database() and apply native database operations such as biogeme_database_define_variable() and biogeme_database_remove().
  3. Build expressions with variable(), biogeme_beta(), and the expression functions documented in the reference manual.
  4. Create a model with logit_model(), a specialized model constructor, or the generic biogeme_model().
  5. Estimate with estimate() or the relevant specialized operation.
  6. Run validate_model(model) before a long estimation when you want a native specification check without optimization.
  7. Use the native result through R methods such as summary(), coef(), vcov(), and logLik().
  8. Use predict(fit) for native choice probabilities or predict(fit, newdata = ...) for scenario data. Use simulate(), cross-validation, confidence intervals, diagnostics, and other post-estimation functions as separate R operations.

When an operation creates native HTML, YAML, NetCDF, pickle, iteration, or diagnostic files, supply output_directory = "/absolute/path/to/output" in biogeme_control() (or the explicit path argument documented by that operation). The package never silently uses the current working directory for these files. The distributed command-line examples therefore require an explicit --output=/path/to/output argument; tempdir() is appropriate for short-lived tests and demonstrations.

For models whose likelihood is not covered by a specialized constructor, biogeme_model() accepts a complete neutral likelihood expression, simulation expressions, optional weights, panel trajectory aggregation, draws, subsets, and parameter overrides.

Model families

The package includes R interfaces for:

Each operation delegates numerical work to native Biogeme. The R interface does not implement a second likelihood, optimizer, differentiation engine, or reporting engine.

Examples and guides

The inst/examples/ directory mirrors the corresponding native example groups and contains self-contained R scripts. Each script keeps its model specification visible and uses helpers only for explicit data preparation and command-line handling. After installation, the same tree is available through system.file("examples", package = "rbiogeme").

The package vignettes provide the user-oriented entry points:

Open the installed guides with:

vignette("getting-started", package = "rbiogeme")
vignette("modeling-workflows", package = "rbiogeme")
vignette("advanced-models", package = "rbiogeme")

Reproducibility and files

Use a fresh output directory for equivalence tests and examples. Set generate_yaml = FALSE, generate_html = FALSE, and save_iterations = FALSE when those files are not part of the operation being tested. Example scripts that generate sampled alternatives explicitly disable result recycling, so an old YAML or iteration file is not silently reused.

Random-draw and sampling examples reproduce native Biogeme's algorithm and parameter names. Their numerical results can differ across runs when the native operation uses random draws; deterministic models should agree up to floating-point precision.

Help and errors

The reference manual is available through R's normal help system. Start with ?rbiogeme, ?biogeme_check, and ?biogeme_model.

| Goal | Start with | Guide | | --- | --- | --- | | Estimate a standard choice model | logit_model() and estimate() | vignette("getting-started", package = "rbiogeme") | | Specify a custom likelihood | biogeme_model() | vignette("modeling-workflows", package = "rbiogeme") | | Work with panel observations | biogeme_panel_database() and panel_likelihood_trajectory() | vignette("modeling-workflows", package = "rbiogeme") | | Evaluate probabilities or scenarios | predict() and simulate() | vignette("modeling-workflows", package = "rbiogeme") | | Use Bayesian, MDCEV, Monte Carlo, catalog, hybrid-choice, or sampling features | The corresponding specialized constructor | vignette("advanced-models", package = "rbiogeme") |

Common first-session problems have direct remedies:

| Message or symptom | Remedy | | --- | --- | | Python was initialized before configuration | Restart R, call biogeme_config() first, then call biogeme_check(). | | Biogeme cannot be imported | Install biogeme==3.3.5 in the selected environment, restart R, and rerun the check. | | The Biogeme version is wrong | Select an environment containing exactly the configured requirement or change the requirement deliberately. | | A model reports a missing column | Compare biogeme_database_columns(database) with every name used by variable(). | | A model fails before estimation | Run validate_model(model); it checks the native specification without running the optimizer. | | Old result files affect an example | Use a fresh tempfile() output directory and disable YAML, HTML, and iteration files unless they are required. |

Errors preserve the operation being attempted and, when debug mode is enabled, can include the native traceback. Use biogeme_config(debug = TRUE) while diagnosing an environment or bridge problem.



Try the rbiogeme package in your browser

Any scripts or data that you put into this service are public.

rbiogeme documentation built on Sept. 29, 2026, 5:09 p.m.