knitr::opts_chunk$set( echo = TRUE, eval = FALSE, collapse = TRUE, comment = "#>" )
This guide collects the common estimation and post-estimation patterns. It
assumes that Python has been configured as described in getting-started.
Every model below is specified in R and compiled once into native Biogeme
expressions before numerical work begins.
The chunks are intentionally not evaluated during package documentation builds: they are complete, runnable examples, but estimation and simulation require a configured Python environment and may create native output files.
library(rbiogeme) data <- data.frame( choice = c(1, 2, 3, 1, 2, 3, 1, 2), time_1 = c(10, 8, 9, 12, 7, 10, 11, 8), time_2 = c(8, 10, 7, 9, 11, 8, 10, 12), time_3 = c(12, 9, 11, 8, 10, 12, 9, 11), cost_1 = c(5, 7, 6, 5, 8, 6, 5, 7), cost_2 = c(7, 5, 8, 6, 7, 5, 8, 6), cost_3 = c(9, 8, 7, 10, 8, 9, 7, 8), weight = c(1, 1, 2, 1, 1, 2, 1, 1), person = c(1, 1, 2, 2, 3, 3, 4, 4) ) database <- biogeme_database("workflows", data) b_time <- biogeme_beta("b_time", start = 0) b_cost <- biogeme_beta("b_cost", start = 0) asc_2 <- biogeme_beta("asc_2", start = 0) asc_3 <- biogeme_beta("asc_3", start = 0) utilities <- list( `1` = b_time * variable("time_1") + b_cost * variable("cost_1"), `2` = asc_2 + b_time * variable("time_2") + b_cost * variable("cost_2"), `3` = asc_3 + b_time * variable("time_3") + b_cost * variable("cost_3") ) output_directory <- tempfile("rbiogeme-workflows-") dir.create(output_directory) clean_control <- biogeme_control( output_directory = output_directory, generate_html = FALSE, generate_yaml = FALSE, save_iterations = FALSE )
Availability is an expression for each alternative. A value of 1 means
available and 0 means unavailable; a symbolic indicator can vary by
observation. If the utility names are labels instead of integer codes, supply
alternative_codes explicitly.
logit <- logit_model( database = database, choice = "choice", utilities = utilities, availability = list( `1` = 1, `2` = variable("cost_2") < 9, `3` = variable("time_3") < 12 ), weight = variable("weight") ) fit <- estimate(logit, model_name = "workflow_logit", control = clean_control) summary(fit) # Native probability columns for the fitted alternatives. predicted <- predict(fit) head(predicted)
The lower-level probability constructors are useful when the probability is part of a larger likelihood or simulation formula:
choice_probability <- logit_probability( utilities = utilities, availability = list(`1` = 1, `2` = 1, `3` = 1), alternative = variable("choice") ) choice_log_probability <- logit_log_probability( utilities = utilities, availability = list(`1` = 1, `2` = 1, `3` = 1), alternative = variable("choice") ) # The same expression can be used in a generic model. generic_logit <- biogeme_model( database = database, formula = choice_log_probability, probability = choice_probability, weight = variable("weight") )
Nested-logit structures are declarative. A non-trivial nest contains its alternative codes and a nest parameter; alternatives not listed in such a nest are treated as trivial nests.
mu_motor <- biogeme_beta("mu_motor", start = 1, lower = 1) nests <- nested_nests( choice_set = c(1, 2, 3), nests = list( nested_nest(mu_motor, alternatives = c(2, 3), name = "motor") ) ) nested <- nested_logit_model( database = database, choice = "choice", utilities = utilities, nests = nests ) nested_fit <- estimate(nested, model_name = "workflow_nested", control = clean_control) nested_logit_correlation(nested, beta_values = coef(nested_fit))
In a cross-nested model, an alternative can receive a symbolic or numeric
allocation in more than one nest. With sparse = TRUE, structurally omitted
allocations are allowed.
mu_a <- biogeme_beta("mu_a", start = 1, lower = 1) mu_b <- biogeme_beta("mu_b", start = 1, lower = 1) cnl_nests <- cross_nested_nests( choice_set = c(1, 2, 3), nests = list( cross_nested_nest(mu_a, c(`1` = 1, `2` = 0.5, `3` = 0)), cross_nested_nest(mu_b, c(`1` = 0, `2` = 0.5, `3` = 1)) ) ) cnl <- cross_nested_logit_model( database = database, choice = "choice", utilities = utilities, nests = cnl_nests ) cross_nested_sparsity_report(cnl_nests) cnl_fit <- estimate(cnl, model_name = "workflow_cnl", control = clean_control) cross_nested_logit_correlation(cnl, beta_values = coef(cnl_fit))
Declare the panel identifier once. Rows for one individual must be contiguous; the constructor checks this rather than silently changing the panel structure. For a random-coefficient panel model, place the trajectory and integration nodes in the generic expression and provide draw metadata.
panel_database <- biogeme_panel_database( "workflow_panel", data, panel_id = "person" ) panel_probability <- logit_probability( utilities = utilities, alternative = variable("choice") ) trajectory <- panel_likelihood_trajectory(panel_probability) b_time_random <- biogeme_beta("b_time_random", start = 0) random_utility <- list( `1` = (b_time_random + biogeme_beta("sd_time", start = 1, lower = 0) * draw("time_draw", "NORMAL")) * variable("time_1"), `2` = asc_2 + (b_time_random + biogeme_beta("sd_time", start = 1, lower = 0) * draw("time_draw", "NORMAL")) * variable("time_2"), `3` = asc_3 + (b_time_random + biogeme_beta("sd_time", start = 1, lower = 0) * draw("time_draw", "NORMAL")) * variable("time_3") ) random_probability <- logit_probability( utilities = random_utility, alternative = variable("choice") ) panel_model <- biogeme_model( database = panel_database, formula = log(monte_carlo(panel_likelihood_trajectory(random_probability))), draws = biogeme_draws( name = "time_draw", draw_type = "NORMAL_ANTI", number_of_draws = 128L, seed = 1223L ) ) panel_fit <- estimate(panel_model, model_name = "workflow_panel", control = clean_control)
panel_likelihood_trajectory() multiplies the observation-level
probabilities for each individual in native Biogeme. monte_carlo() then
averages over the declared draw design. Neither operation is evaluated by R.
For numerical quadrature over a standard normal variable, use
integrate_normal(expression, name, number_of_quadrature_points) instead.
Simulation is a separate operation from estimation. Give the model a named
simulations list, or pass one directly to simulate().
simulation_model <- biogeme_model( database = database, formula = choice_log_probability, simulations = list( probability = choice_probability, expected_time = variable("time_1") * choice_probability, cost_share = variable("cost_1") / (1 + variable("cost_1")) ) ) simulated <- simulate(simulation_model, beta = fit, control = clean_control) as.data.frame(simulated) # Expressions can also be supplied for one simulation call. simulate( simulation_model, expressions = list(probability = choice_probability), beta = fit, control = clean_control )
The package also exposes native derivative checks, confidence intervals, and cross-validation:
check_derivatives( model = generic_logit, model_name = "workflow_derivatives", control = clean_control, verbose = TRUE ) biogeme_confidence_intervals( model = simulation_model, beta_values = list(coef(fit), coef(fit)), expressions = list(probability = choice_probability), interval_size = 0.90, control = clean_control ) validate( model = logit, fit = fit, folds = 5L, groups = "person", seed = 1234, control = clean_control )
validate_model() checks the compiled native specification before estimation.
check_derivatives() compares native analytical derivatives with native finite
differences. biogeme_confidence_intervals() delegates repeated parameter
draws and quantile calculations to native Biogeme. validate() delegates fold
assignment, estimation, and scoring to native Biogeme.
There are two deliberately different result-file workflows:
yaml_file <- file.path(output_directory, "workflow.yaml") # Always estimate afresh and optionally write the named YAML result. fresh_fit <- estimate( logit, model_name = "workflow_file", yaml_file_name = yaml_file, control = clean_control ) # Loading is explicit. force = FALSE permits reading the named file; # force = TRUE ignores it and estimates again. loaded_fit <- estimate_or_load( logit, yaml_file_name = yaml_file, force = FALSE, model_name = "workflow_file", controls = clean_control ) # The ordinary R representation can also be saved and read explicitly. save_results(fresh_fit, yaml_file) read_results(yaml_file)
Use a fresh temporary directory in equivalence tests. This makes it clear which files belong to the current run and prevents a stale YAML or iteration file from changing the result.
biogeme_control() is a named R list. Named arguments such as seed,
number_of_draws, optimization_algorithm, variance_covariance_type, and
bootstrap_samples are mapped to their native controls. Additional named
arguments are passed through, which keeps the interface usable when a native
operation exposes a specialized control not yet given a dedicated R argument.
control <- biogeme_control( output_directory = output_directory, seed = 1234, numerically_safe = TRUE, optimization_algorithm = "automatic", number_of_draws = 128L, generate_html = FALSE, generate_yaml = FALSE, save_iterations = FALSE )
Use the model-specific help pages for the complete argument contract of each operation. The next guide covers Bayesian, Monte Carlo, MDCEV, catalog, assisted-specification, hybrid-choice, and sampled-alternative workflows.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.