V3/V4 to V5 Migration Guide

library(epidatr)

The legacy Epidata APIs, including the V4 main endpoint (pub_covidcast()) and V3 other endpoints (pub_fluview(), pub_flusurv(), pub_wiki(), etc.), are transitioning to the V5 API. This transition is occurring source by source. All V3 and V4 sources will continue to operate until the migration is complete (tentatively scheduled for October 2026), and endpoints that are no longer updated will remain accessible on V3/V4. For new integrations, start directly on V5 and fall back to legacy endpoints only for sources that are not yet supported.

For the current list of sources and indicators available on the new API, see the V5 signals documentation.

This guide walks through the transition from pub_covidcast() and other legacy endpoints. While pub_covidcast() is the most widely used legacy endpoint, V3 endpoints differ in their function names and parameter conventions. The tables below compare both V4 (pub_covidcast()) and V3 (using pub_fluview() as an example) to their V5 equivalents.

Endpoint mapping

The legacy endpoints split into several purpose-built V5 routes determined by query type. The "V3 (Other Endpoints)" column highlights examples (pub_fluview(), pub_flusurv(), pub_wiki()) to illustrate differences across endpoints. Refer to each endpoint's documentation for specific behavior:

| Task | V4 (pub_covidcast) | V3 (Other Endpoints) | V5 Equivalent | |---|---|---|---| | Fetch latest data or snapshot as of a past date | pub_covidcast() (default or with as_of) | Endpoint-specific (pub_fluview() has no as_of) | epidata_snapshot() | | Fetch full revision history for a signal | pub_covidcast(issues = ...) | Supported by some (pub_fluview(), pub_flusurv() with issues) | epidata_archive() | | Discover sources, signals, geo types, and date ranges | pub_covidcast_meta(), covidcast_epidata() | Shared meta for some (pub_fluview_meta()) | epidata_meta() | | Access source-specific auxiliary tables | none | none | epidata_aux() | | Filter by publication lag | pub_covidcast(lag = ...) | Supported by some (pub_fluview(), pub_flusurv()) | none (compute report_time - reference_time) |

epidata() is a convenience wrapper that routes to epidata_archive() if you pass report_time, or to epidata_snapshot() if you pass snapshot_date (or neither).

Argument changes

Most pub_covidcast() arguments carry over to V5 with the same name, but some have been renamed, dropped, or added. Historical V3 endpoints do not share argument names with pub_covidcast(). Arguments for pub_fluview() are shown below as an example, but consult each endpoint's documentation for details:

| V4 argument (pub_covidcast) | V3 (pub_fluview) | V5 argument | Notes | |---|---|---|---| | source (data_source) | not exposed (identified by function name pub_fluview()) | source | Identifies the source dataset in V5 (replaces V4 source and V3 endpoint names). | | signals | none (implicit from endpoint) | signals | Identifies the specific signal name within the source. | | geo_type | not exposed (pub_fluview() supports only regions) | geo_type | Specifies geographic resolution (e.g., state, county, hhs, nation). | | geo_values | regions for pub_fluview() | geo_values | Removed from API query in V5 (queries return all locations for the requested geo_type). Filtered locally in R after the fetch. | | time_type | not exposed (pub_fluview() is always epiweeks) | none | Dropped. All V5 endpoints use standard calendar dates (Date). | | time_values | epiweeks for pub_fluview() | reference_time | Removed from API query in V5 (queries return all dates). Filtered locally in R after the fetch. | | as_of | none (pub_fluview() has no as_of) | snapshot_date | In V5, used only in epidata_snapshot() to fetch data known as of a past date. NULL returns the latest data. | | issues | issues (where supported) | report_time | In V5, used only in epidata_archive(). Accepts operators like "<2025-10-16>", or epirange(). (For a single date, use epidata_snapshot()). | | lag | lag (where supported) | none | Removed in V5. You can compute it yourself: fetch from epidata_archive() and filter by report_time - reference_time. See filtering by lag. | | none | none | fill_method | New in V5. Selects the imputation method when aggregating sub-geographies ("source", "fill_ave", or "fill_zero"). See below. | | none | none | ... | New in V5. Filters on source-specific dimensions (such as age_group or nwss_source). |

The new functions also add fill_method, which has no covidcast equivalent. Some sources publish several variants of the same signal that differ in how nulls were handled during geographic aggregation:

The default NULL returns all variants, so filter on this column (or pass a value to the argument) if you want exactly one time series per location.

Column changes

Response fields follow a similar pattern. In the table below, pub_fluview() serves as an example of an endpoint with custom fields. Column names vary across legacy endpoints (for example, pub_wiki() returns article, count, and hour):

| V4 column (pub_covidcast) | V3 (pub_fluview) | V5 column | Notes | |---|---|---|---| | source | not returned (implicit from endpoint) | dropped | Omitted in V5 responses because the source is already specified in the request. | | signal | none (implicit from endpoint) | signal | Identifies the signal name in V5. | | value | Endpoint-specific columns (e.g. num_ili, wili, ili) | value | Standardized metric value column across all V5 sources. | | not returned | not returned (implicit from endpoint) | geo_type | Explicitly included in V5 responses to identify geographic resolution. | | geo_value | region for pub_fluview() | geo_value | Standardized location identifier across all V5 responses. | | time_value | epiweek for pub_fluview() | reference_time | Standardized date in YYYY-MM-DD format representing the observation period. | | issue | issue (where returned) | report_time | Standardized date in YYYY-MM-DD format representing when the data point was published. Present in both snapshot and archive output. | | lag | lag (where returned) | dropped | Omitted in V5 responses. You can compute it yourself as report_time - reference_time. See calculating reporting lag. | | direction | none | dropped | Deprecated in V4 and dropped in V5. | | stderr, sample_size | none | ci_lower, ci_upper | Expresses uncertainty as explicit confidence interval bounds on value when provided by the data source. See Uncertainty columns below. | | missing_value, missing_stderr, missing_sample_size | none | dropped | Replaced in V5 by fill_method variants and plain NAs in value. | | none | none | fill_method | Indicates which null-handling imputation method was applied ("source", "fill_ave", or "fill_zero"). See above. |

Some sources also carry extra columns in the new API, for example age_group (pophive) and nwss_source, sample_index, pcr_target (nwss). For more information on whether the source you're interested in provides extra columns, please visit that source's documentation page.

Uncertainty columns

The covidcast columns stderr and sample_size have no fixed replacement. The shared schema carries only value; a source that quantifies uncertainty adds its own columns, such as ci_lower and ci_upper. Use the metadata function or the documentation to see which value columns a source returns:

meta_sleepcycle <- epidata_meta(source = "sleepcycle")
meta_sleepcycle$value_columns
#> [1] "ci_lower" "ci_upper" "value"

A query, before and after

V4 query example: NSSP COVIDcast

Fetching NSSP influenza ED visit percentages for two states, as the data looked on January 1, 2025:

old <- pub_covidcast(
  source = "nssp",
  signals = "pct_ed_visits_influenza",
  geo_type = "state",
  time_type = "week",
  geo_values = c("pa", "ca"),
  time_values = epirange(202440, 202501),
  as_of = 20250101
)
#> Warning: `pub_covidcast()` uses the V4 Epidata API.
#> ℹ Starting in October 2026, V4 is tentatively deprecated in favor of the V5 API.
#> ℹ See `vignette("migration-guide")` (or
#>   <https://cmu-delphi.github.io/epidatr/articles/migration-guide.html>) for the V5
#>   endpoints and how to move to them. Old data will remain available for at least a
#>   year, but new ingestion will end.
#> This warning is displayed once every 8 hours.
head(old)
#> # A tibble: 6 × 15
#>   geo_value signal     source geo_type time_type time_value direction issue     
#>   <chr>     <chr>      <chr>  <fct>    <fct>     <date>         <dbl> <date>    
#> 1 ca        pct_ed_vi… nssp   state    week      2024-09-29        NA 2026-09-13
#> 2 pa        pct_ed_vi… nssp   state    week      2024-09-29        NA 2026-09-13
#> 3 ca        pct_ed_vi… nssp   state    week      2024-10-06        NA 2026-09-13
#> 4 pa        pct_ed_vi… nssp   state    week      2024-10-06        NA 2026-09-13
#> # ℹ 2 more rows
#> # ℹ 7 more variables: lag <dbl>, missing_value <dbl>, missing_stderr <dbl>,
#> #   missing_sample_size <dbl>, value <dbl>, stderr <dbl>, sample_size <dbl>
new <- epidata_snapshot(
  source = "nssp",
  signals = "pct_ed_visits_influenza",
  geo_type = "state",
  geo_values = c("pa", "ca"),
  reference_time = epirange("2024-10-01", "2025-01-01"),
  snapshot_date = "2025-01-01"
)
head(new)
#> # A tibble: 6 × 7
#>   signal         report_time geo_type geo_value fill_method reference_time value
#>   <chr>          <date>      <chr>    <chr>     <chr>       <date>         <dbl>
#> 1 pct_ed_visits… 2024-12-27  state    ca        source      2024-10-05     0.140
#> 2 pct_ed_visits… 2024-12-27  state    ca        source      2024-10-12     0.140
#> 3 pct_ed_visits… 2024-12-27  state    ca        source      2024-10-19     0.160
#> 4 pct_ed_visits… 2024-12-27  state    ca        source      2024-10-26     0.200
#> # ℹ 2 more rows

Both queries return the same signal, just with renamed and reshaped columns:

names(old)
#> [1] "geo_value" "signal"    "source"    "geo_type" 
#>  [ reached 'max' / getOption("max.print") -- omitted 11 entries ]
names(new)
#> [1] "signal"      "report_time" "geo_type"    "geo_value"  
#>  [ reached 'max' / getOption("max.print") -- omitted 3 entries ]

V3 query example: FluView

For V3 endpoints like pub_fluview(), metric names that used to be separate columns (such as num_ili, ili, wili) become individual signal names queried via signals, and results are standardized into the single value column:

old_flu <- pub_fluview(
  regions = "nat",
  epiweeks = epirange(202440, 202445)
)
#> Warning: `pub_fluview()` uses the V4 Epidata API.
#> ℹ Starting in October 2026, V4 is tentatively deprecated in favor of the V5 API.
#> ℹ See `vignette("migration-guide")` (or
#>   <https://cmu-delphi.github.io/epidatr/articles/migration-guide.html>) for the V5
#>   endpoints and how to move to them. Old data will remain available for at least a
#>   year, but new ingestion will end.
#> This warning is displayed once every 8 hours.
head(old_flu[, c("release_date", "region", "epiweek", "wili", "ili")])
#> # A tibble: 6 × 5
#>   release_date region epiweek     wili   ili
#>   <date>       <chr>  <date>     <dbl> <dbl>
#> 1 2026-09-18   nat    2024-09-29  1.91  1.85
#> 2 2026-09-18   nat    2024-10-06  2.02  1.94
#> 3 2026-09-18   nat    2024-10-13  2.07  2.01
#> 4 2026-09-18   nat    2024-10-20  2.22  2.16
#> # ℹ 2 more rows
new_flu <- epidata_snapshot(
  source = "fluview_ilinet",
  signals = "wili",
  geo_type = "nation",
  geo_values = "us",
  reference_time = epirange("2024-10-01", "2024-11-15")
)
head(new_flu)
#> # A tibble: 6 × 7
#>   signal report_time geo_type geo_value fill_method reference_time value
#>   <chr>  <date>      <chr>    <chr>     <chr>       <date>         <dbl>
#> 1 wili   2025-09-12  nation   us        source      2024-10-12      2.02
#> 2 wili   2025-09-12  nation   us        source      2024-10-19      2.07
#> 3 wili   2025-09-12  nation   us        source      2024-10-26      2.22
#> 4 wili   2025-09-12  nation   us        source      2024-11-02      2.32
#> # ℹ 2 more rows

Revision history queries

Where you pass issues to pub_covidcast(), use epidata_archive() with report_time:

old_revisions <- pub_covidcast(
  source = "nssp",
  signals = "pct_ed_visits_influenza",
  geo_type = "state",
  time_type = "week",
  geo_values = "pa",
  time_values = epirange(202440, 202501),
  issues = epirange(202440, 202522)
)
head(old_revisions)
#> # A tibble: 6 × 15
#>   geo_value signal     source geo_type time_type time_value direction issue     
#>   <chr>     <chr>      <chr>  <fct>    <fct>     <date>         <dbl> <date>    
#> 1 pa        pct_ed_vi… nssp   state    week      2024-09-29        NA 2024-11-03
#> 2 pa        pct_ed_vi… nssp   state    week      2024-09-29        NA 2024-11-10
#> 3 pa        pct_ed_vi… nssp   state    week      2024-09-29        NA 2024-11-17
#> 4 pa        pct_ed_vi… nssp   state    week      2024-09-29        NA 2024-11-24
#> # ℹ 2 more rows
#> # ℹ 7 more variables: lag <dbl>, missing_value <dbl>, missing_stderr <dbl>,
#> #   missing_sample_size <dbl>, value <dbl>, stderr <dbl>, sample_size <dbl>
revisions <- epidata_archive(
  source = "nssp",
  signals = "pct_ed_visits_influenza",
  geo_type = "state",
  geo_values = "pa",
  reference_time = epirange("2024-10-01", "2025-01-01"),
  report_time = "<2025-06-01"
)
head(revisions)
#> # A tibble: 6 × 7
#>   signal        report_time geo_type geo_value fill_method reference_time  value
#>   <chr>         <date>      <chr>    <chr>     <chr>       <date>          <dbl>
#> 1 pct_ed_visit… 2024-11-08  state    pa        source      2024-10-05     0.0500
#> 2 pct_ed_visit… 2024-11-08  state    pa        source      2024-10-12     0.0700
#> 3 pct_ed_visit… 2024-11-08  state    pa        source      2024-10-19     0.0800
#> 4 pct_ed_visit… 2024-11-08  state    pa        source      2024-10-26     0.130 
#> # ℹ 2 more rows

If you filtered by lag, fetch the archive with epidata_archive() and filter afterwards:

# For an exact lag (e.g., 7 days):
revisions %>%
  filter(as.integer(report_time - reference_time) == 7)

# Or for maximum latency (e.g., at most 7 days of delay):
revisions %>%
  filter(as.integer(report_time - reference_time) <= 7)

Checking whether a source is available

Use epidata_meta() to see what a source offers in the new API. It returns signals, geo types, and the available reference_time and report_time ranges:

meta <- epidata_meta(source = "nssp")

# all the fields available for this source
names(meta)
#> [1] "report_time_range"    "reference_time_range" "signals"             
#> [4] "geo_types"           
#>  [ reached 'max' / getOption("max.print") -- omitted 4 entries ]

meta$signals # available signal names
#> [1] "pct_ed_visits_ari"       "pct_ed_visits_combined" 
#> [3] "pct_ed_visits_covid"     "pct_ed_visits_influenza"
#>  [ reached 'max' / getOption("max.print") -- omitted 5 entries ]
meta$geo_types # supported geography levels
#> [1] "census_division" "census_region"   "county"          "hhs"            
#>  [ reached 'max' / getOption("max.print") -- omitted 5 entries ]
meta$reference_time_range # earliest/latest reference_time available
#> $latest
#> [1] "2026-09-12"
#> 
#> $first
#> [1] "2022-10-01"
meta$report_time_range # earliest/latest report_time (publication date) available
#> $latest
#> [1] "2026-09-16T00:00:00"
#> 
#> $first
#> [1] "2024-04-18T00:00:00"

If epidata_meta() does not know the source yet, keep using pub_covidcast() (or the relevant {pub/pvt}_* function) for it and check back after package updates. The API mailing list announces sources as they move.

Endpoints kept for historical reference

Not every V4 endpoint is moving to V5. The functions below cover data sources whose collection has already ended (e.g. Google Flu Trends, the Twitter/HealthTweets signal, the various nowcasts). They are not part of the V4-to-V5 transition, so they are not deprecated and will keep working. The historical data they return is frozen and will remain available. They will just no longer receive new data.

| Function | Data source | |---|---| | pvt_cdc() | CDC total and by-topic webpage visits | | pub_covid_hosp_facility_lookup() | COVID hospitalization facility lookup | | pub_covid_hosp_facility() | COVID hospitalizations by facility | | pub_covid_hosp_state_timeseries() | COVID hospitalizations by state | | pub_delphi() | Delphi's ILINet outpatient doctor visits forecasts | | pub_dengue_nowcast() | Delphi's PAHO dengue nowcasts (Americas) | | pvt_dengue_sensors() | PAHO dengue digital surveillance sensors (Americas) | | pub_ecdc_ili() | ECDC ILI incidence (Europe) | | pub_gft() | Google Flu Trends flu search volume | | pvt_ght() | Google Health Trends health topics search volume | | pub_kcdc_ili() | KCDC ILI incidence (Korea) | | pvt_meta_norostat() | Metadata for the NoroSTAT endpoint | | pub_nidss_dengue() | NIDSS dengue cases (Taiwan) | | pub_nidss_flu() | NIDSS flu doctor visits (Taiwan) | | pvt_norostat() | CDC NoroSTAT norovirus outbreaks | | pub_nowcast() | Delphi's ILI Nearby nowcasts | | pub_paho_dengue() | PAHO dengue data (Americas) | | pvt_sensors() | Influenza and dengue digital surveillance sensors | | pvt_twitter() | HealthTweets total and influenza-related tweets | | pub_wiki() | Wikipedia webpage counts by article |



Try the epidatr package in your browser

Any scripts or data that you put into this service are public.

epidatr documentation built on Sept. 21, 2026, 5:08 p.m.