Skip to content

cast-API queries filter geo_values/reference_time only client-side — wide queries over-download, and limit under-counts silently because the cap applies before the local filter #403

Description

@docxology

Summary

While tracing the v5/cast query path, I found that epidata_snapshot()/epidata_archive() never place geo_values or reference_time in the request parameters: the full result set for the requested source/signal/geo_type is downloaded and only then filtered locally by .cast_filter(). Two consequences: (a) county-level queries download data for all geographies and discard most of it in memory (the over-download half of the OOM problem in #304 — PR #352's streaming mitigates the memory side but not the transfer); (b) when limit is set, the server-side cap applies before the local filter, so the merged output can have fewer rows than requested — or NA-padded columns in epidata_aux() — with no warning about the rows that were silently dropped. reference_time's local-only behavior is at least documented (R/endpoints.R:1465); geo_values' is not.

Evidence (dev tip 87b10e1)

  • Snapshot params contain no geo/reference filter — R/endpoints.R:1584-1592:
    params = list(
      source = source,
      signal = s,
      geo_type = g,
      fill_method = fill_method,
      snapshot_date = snapshot_date,
      extra_keys = extra_keys,
      limit = fetch_args$limit     # <- the only server-side row cap
    ),
    (archive params at R/endpoints.R:1696-1704, same shape)
  • The local drop happens afterwards — R/utils.R:174-179 (.cast_filter):
    if (!identical(geo_values, "*")) {
      actual_geo_values <- tolower(trimws(unlist(strsplit(geo_values, ","))))
      res <- res[res$geo_value %in% actual_geo_values, ]
    }
    if (!identical(reference_time, "*")) {
      res <- filter_by_timeset(res, "reference_time", parsed_reference_times)
    }
  • The documented asymmetry — R/endpoints.R:1465 documents reference_time as "applied locally ... after the API call"; no such note exists for geo_values.
  • Aux NA-padding — R/endpoints.R:1984-2000: base[value_cols] <- vctrs::vec_slice(aux[value_cols], idx) where idx contains NA for unmatched rows (e.g. when limit truncated the base before the merge).

Maintainer-runnable repro

epidata_snapshot(
  "nhsn", "pct_visits", "county", geo_values = "pa",
  fetch_args = fetch_args_list(dry_run = TRUE)
)$request$url
# URL carries no geo filter at all — the 'pa' restriction is applied locally after download

Impact

  • Resource/memory blow-up on wide pulls: the entire geo universe for the signal/geo_type is transferred and bound in one vctrs::vec_rbind() (the Provide way to avoid OOM with county-level data access #304 OOM reports; PR Stream large API responses to disk to prevent OOM crashes #352's streaming addresses the memory side, but the full transfer remains).
  • With limit, epidata_snapshot()/epidata_archive()/epidata_aux() can silently return fewer rows than the user believes were "all rows within the limit" — and aux merges can carry NA-padded columns — with no signal that the cap consumed rows the local filter would have kept.
  • The docs state the limit caveat ("does not guarantee the same rows (or even the same count)…"), but the interaction with the local filter is the surprising part.

Expected vs. actual

  • Expected: geo_values (like source/signal/geo_type) is applied server-side when the API supports it, or the over-download is documented as loudly as reference_time's local-only behavior; unmatched-needle cases warn.
  • Actual: silent full download + local filter; limit interacts with the local filter to under-count silently.

Suggested fix

  1. Send geo_values server-side if the v5 API accepts it (probe/coordinate with the API team); otherwise document the over-download in ?epidata_snapshot/?epidata_archive next to the reference_time note.
  2. Warn when any needle goes unmatched in the epidata_aux() merge, naming limit/the key filters as the likely cause (mirroring .check_cast_empty()'s partial-result warnings).

Related in-flight work

PR #352 (stream large responses to disk), #380 (comma-joined signals, one request per geo type), and #378 rework this area; the missing server-side geo filter and the limit/local-filter interaction appear to remain in all of them. Happy to consolidate there.

I checked for prior reports: #304 covers the OOM symptom; I found none covering the missing server-side filter or the silent under-count.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions