Provides a unified set of helper functions to access datasets from the Colorado Open Data platform < https://data.colorado.gov/>. Functions return results as tidy tibbles and support optional filtering, sorting, and row limits via the Socrata API. The package provides a consistent interface for discovering and downloading datasets from the Colorado Open Data Portal using human-readable dataset keys or official Socrata dataset identifiers.
coOpenData provides a lightweight R interface to the Colorado Open
Data Portal.
The package allows users to search, filter, and download datasets from the Colorado Open Data Portal directly into R without manually constructing API queries, handling JSON responses, or performing type conversion.
Designed for students, educators, researchers, journalists, civic
technologists, and analysts, coOpenData reduces the technical overhead
required to begin working with municipal Open Data while preserving
access to the underlying Socrata infrastructure.
coOpenData WorksThe package provides a streamlined interface to the Colorado Open Data Portal’s Socrata API.
Internally, coOpenData:
Most workflows begin with co_list_datasets(), which retrieves a live
catalog of datasets available through the Colorado Open Data Portal.
Datasets can then be downloaded using either:
key"e4mf-ggwf"The human-readable key is designed to improve readability and usability, while the U_ID is the stable identifier used by the Socrata platform.
The package provides three primary functions:
co_list_datasets() retrieves a live catalog of available Colorado
Open Data datasets, including human-readable keys, Socrata U_IDs,
names, and other available metadata.
co_pull_dataset() downloads cataloged datasets using either a
human-readable key or Socrata U_ID, with support for filtering,
ordering, date ranges, optional column-name cleaning, and optional
type coercion.
co_any_dataset() downloads data directly from a valid Socrata JSON
endpoint without requiring the dataset to appear in the package
catalog.
Datasets retrieved through co_pull_dataset() support arguments
including:
limitfiltersdatefromtodate_fieldwhereorderclean_namescoerce_typesAll functions return tibble outputs.
Advanced users may also provide raw SoQL conditions through the where
argument.
SoQL, or Socrata Query Language, is the query syntax used by Socrata-powered Open Data portals. Additional information is available from the Socrata developer documentation.
install.packages("coOpenData")
# install.packages("pak")
pak::pak("gomes-sh/coOpenData")
Alternatively:
# install.packages("remotes")
remotes::install_github("gomes-sh/coOpenData")
library(coOpenData)
library(dplyr)
# Browse available datasets
catalog <- co_list_datasets()
# Search for datasets containing a keyword
catalog |>
filter(grepl("KEYWORD", name, ignore.case = TRUE)) |>
select(key, uid, name)
# Pull a dataset using its U_ID
example_data <- co_pull_dataset(
dataset = "e4mf-ggwf",
limit = 100
)
# Pull the same dataset using its catalog key
example_data_by_key <- co_pull_dataset(
dataset = "data.colorado.gov",
limit = 100
)
# Pull filtered data
filtered_data <- co_pull_dataset(
dataset = "e4mf-ggwf",
limit = 100,
filters = list(
grade_levels = "ECE-5"
)
)
The filters argument accepts a named list and automatically constructs
the corresponding SoQL filtering conditions.
Multiple values may be supplied for one field:
filtered_data <- co_pull_dataset(
dataset = "e4mf-ggwf",
limit = 100,
filters = list(
grade_levels = c("ECE-5", "K-5")
)
)
Multiple fields may also be combined:
filtered_data <- co_pull_dataset(
dataset = "e4mf-ggwf",
limit = 100,
filters = list(
grade_levels = "ECE-5",
classification = "Charter"
)
)
Date filtering is available for datasets containing date or datetime fields:
date_filtered_data <- co_pull_dataset(
dataset = "e4mf-ggwf",
from = "2023-01-01",
to = "2024-01-01",
date_field = "last_verified",
limit = 100
)
When a dataset is not available through co_list_datasets(), it can be
downloaded directly using co_any_dataset().
endpoint_data <- co_any_dataset(
json_link = "https://data.colorado.gov/resource/e4mf-ggwf.json",
limit = 100
)
Use co_pull_dataset() for catalog-based workflows and
co_any_dataset() when working directly with a Socrata JSON endpoint.
A complete introductory workflow is available in the package vignette:
vignette("getting-started", package = "coOpenData")
The vignette demonstrates how to:
Complete documentation is available on the package website:
https://nyc-open-data-lab.github.io/coOpenData/
The website includes:
To run the package tests locally:
devtools::test()
To rebuild the documentation:
devtools::document()
To run a complete package check:
devtools::check()
To rebuild the pkgdown website:
pkgdown::build_site()
Contributions are welcome.
To report a bug, request a feature, or suggest an improvement, open an issue on GitHub:
https://github.com/nyc-open-data-lab/coOpenData/issues
Pull requests are also welcome. Before submitting a pull request, please ensure that:
devtools::check() completes successfullyShelby Lyn Gomes
Email: [email protected]
GitHub: @gomes-sh
Because the package retrieves metadata dynamically from the live Colorado Open Data catalog, newly published datasets may become available without requiring a package update.
Package updates may still be required when the portal changes its catalog structure, dataset metadata fields, or API behavior.
coOpenData is an independent project and is not affiliated with,
endorsed by, or maintained by Colorado or the organization responsible
for the Colorado Open Data Portal.