View source: R/fedstat_get_data_ids.R
| fedstat_get_data_ids | R Documentation |
To query data from fedstat we need to POST some filters in form of filter numeric identificators. Most filters don't have some rule from which their ids can be generated based on filters titles and values. It seems like these ids are just indexes in the fedstat inner database. So in order to get the data, we first need to get the ids of the filter values by parsing specific part of java script source code on indicator web page.
fedstat_get_data_ids(
indicator_id,
...,
timeout_seconds = 180,
retry_max_times = 3,
httr_verbose = NULL
)
indicator_id |
character, indicator id/code from indicator URL. For example for indicator with URL https://www.fedstat.ru/indicator/37426 indicator id will be 37426 |
... |
other arguments passed to httr::GET |
timeout_seconds |
numeric, maximum time before a new GET request is tried |
retry_max_times |
numeric, maximum number of tries to GET |
httr_verbose |
|
It is known that the fedstat lags quite often. Sometimes site never responds at all. This is especially true for the most popular indicators web pages. In this regard, by default, a GET request is sent 3 times with a timeout of 180 seconds and with initially small, but growing exponentially, pauses between requests.
As a rule, requests to the indicator web page take much longer than requests
to get the data itself. A POST request for data is sent to
https://www.fedstat.ru/indicator/downloadData.do?format={format}
for all indicators and is often quite fast.
Note: CSRF tokens are single-use. Each call to fedstat_post_data_ids_filtered
consumes the token. For subsequent downloads, call fedstat_get_data_ids again.
The wrapper function fedstat_data_load_with_filters handles this automatically.
Correct filter_field_object_ids are needed to get data. For the sdmx format, these ids do not change anything, except for the standard data sorting, but their incorrect specification will lead either to incomplete data loading or to no data at all. For the excel format, these ids determine the form of data presentation, as in the data preview on the fedstat site. For now only default filter_field_object_ids are used, which are parsed from java script source code on indicator web page. Users can specify filter_field_object_ids for each filter_field in resulting data_ids table.
The returned data.frame also carries session and CSRF token metadata as attributes
(fedstat_handle, fedstat_base_url,
fedstat_csrf_token, fedstat_csrf_token_name, fedstat_indicator_id).
The fedstat_handle attribute holds an httr::handle object that preserves
session cookies for the downstream POST request.
These are used internally by fedstat_post_data_ids_filtered and must not be removed.
Note: CSRF tokens are single-use – each call to fedstat_post_data_ids_filtered
consumes the token. For subsequent downloads, call fedstat_get_data_ids again.
data.frame with all character type columns:
filter_field_id - id for filter field;
filter_field_title - filter field title string representation;
filter_value_id - id for filter field value;
filter_value_title - filter field value title string representation;
filter_field_object_ids - special strings that define the location of the filters fields. It can take the following values: lineObjectIds (filters in lines), columnObjectIds (filters in columns), filterObjectIds (hidden filters for all data);
fedstat_data_ids_filter,
fedstat_post_data_ids_filtered
## Not run:
# Get data filters identificators for CPI
data_ids <- fedstat_get_data_ids("31074")
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.