Reproducible Data Capsules with Provenance and Fallback

Tools for building brick-proof, reproducible, self-contained data capsules. Resolves open-data sources through the Comprehensive Knowledge Archive Network ('CKAN', < https://ckan.org/>) package_show and package_search endpoints, records and verifies provenance with Secure Hash Algorithm 256 ('SHA-256') digests and Internet Archive 'Wayback Machine' (< https://web.archive.org/>) snapshots, validates downloaded data against a pinned schema, and falls back to schema-driven synthetic data when the real source is unreachable. Run records are captured in a manifest plus a plain-language summary so any result can be traced back to its inputs. Distributional drift between a pinned capsule and a fresh fetch is tested with Kolmogorov-Smirnov, chi-square, population stability index, Jensen-Shannon divergence and 'Benford' first-digit screens, because a re-released extract can be statistically identical yet differ byte-for-byte, and a column can keep its name and type while having been silently rescaled. Manifests can be authenticated rather than only checksum-verified, with keyed digests ('HMAC-SHA-256', RFC 2104) or post-quantum hash-based signatures ('Winternitz' one-time signatures under a 'Merkle' tree, RFC 8391), and pinned chunk-wise through a 'Merkle' tree so a mismatch identifies which part of a capsule moved. Also ships a compiled C++ core (summary, robust and rank statistics, 'SHA-256', 'SHA-512' and 'CRC-32') that sibling packages in the 'rmorie' ecosystem reach through 'LinkingTo' for a single, shared numeric and provenance-hashing backend. For the published administrative tables these capsules usually hold, it computes period-over-period change matched on the period rather than the row, with the exact conditional-binomial interval for a ratio of counts and with a percentage-point reading kept distinct from a percent change, rendered to Hypertext Markup Language ('HTML'), Portable Document Format ('PDF'), delimited text, JavaScript Object Notation ('JSON') or Markdown. Interval categories such as "2 to 5" or "50+" are parsed to bounds and the dependence of any derived figure on the open top band is measured rather than assumed. Concentration is summarised by the 'Gini' coefficient, the Lorenz curve and tail-index estimation by exact discrete maximum likelihood; trend in a series of a few periods by the Mann-Kendall test with 'Theil-Sen' slopes, a permutation step-change scan and Poisson rate ratios; and region-coded counts by indirect standardisation, exact standardised incidence ratios, the empirical Bayes shrinkage of Clayton and 'Kaldor' (1987) , funnel-plot limits and Moran's I.


rmoriebricklayer

CRAN status R-CMD-check Codecov test coverage Project Status: Active Lifecycle: maturing r-universe License: AGPL v3

Brick-proof, reproducible data capsules for R.

rmoriebricklayer resolves open-data sources, records and verifies provenance, validates downloaded data against a pinned schema, and falls back to schema-driven synthetic data when the real source is unreachable — so any analysis result can be traced back to its exact inputs.

A checksum answers one question: are these the same bytes? The package exists because that is rarely the question that matters. A re-released extract can be statistically identical and differ byte-for-byte; a column can keep its name, type and row count while having been silently rescaled; and a digest anyone can recompute says nothing about who produced the data.

What it does

  • One call for a published table — analyse_table() takes a table of counts by period and group and returns what changed with exact intervals, p-values adjusted over the whole scan, the envelope that rounding and suppression in the release imply (published_bounds()), trend, rates if there is an exposure, and a drift screen against the prior capsule; report_analysis() writes it as Markdown or one HTML file and use_capsule_template() starts a capsule that runs as written. Start with vignette("getting-started").
  • CKAN resolution — resolve_via_ckan() / resolve_via_ckan_search() locate resources through a portal's package_show / package_search endpoints.
  • Provenance — load_provenance(), make_manifest(), record(), write_manifest_json(), and write_summary_txt() capture every run as a manifest plus a plain-language summary.
  • Integrity — sha256_file() / verify_sha256() hash and verify downloads; download_data() / friendly_download() fetch with a Wayback Machine fallback.
  • Categorical integrity — guard_recode(), decode_codes(), guard_levels(), audit_categories(), verify_recode(), verify_marginals(), odds_ratio_check(), relabel(), decode_labelled(), transfer_verify(), relabel_forensics() and a signed recode_manifest(): recodes that refuse anything unmapped or positional, an SPSS/Stata/SAS import checked against the source code book and frequency table, reported odds ratios recomputed under every relabelling, and the mechanical step behind a permutation named, so a swapped label is fixed on the day, not blamed on the software.
  • Schema validation — infer_schema() derives a pinnable schema from data you trust; validate_schema() checks names, types, ranges, value sets and missingness against it; rule() and the rule_*() library express the project-specific checks a generic schema cannot.
  • Drift detection — capsule_drift() asks whether the data moved, not just the bytes, with Kolmogorov-Smirnov, two-sample homogeneity, population stability index, Jensen-Shannon divergence and a Benford first-digit screen.
  • Signed provenance — capsule_sign() authenticates a manifest with a keyed digest or a post-quantum hash-based signature; merkle_root() pins a capsule chunk by chunk so a mismatch names which chunk moved; and chain_append() links manifests so the run history is tamper-evident, not only each run.
  • Description — profile_columns(), frequency_table(), correlation_table(), mahalanobis_outliers(), missingness_map() and mcar_test() (Little's test, with the EM estimator it requires) describe a capsule before you trust it.
  • Capsules larger than memory — exact block-wise moment accumulation, reservoir sampling, and HyperLogLog distinct counts, all in one pass.
  • Synthetic fallback — make_synthetic_column() / make_synthetic_csv() generate schema-driven stand-ins when the real source is down, so a pipeline still runs end-to-end.
  • Rates and shares — rate() gives events per population at any denominator (per = 1000, "100k", "1m") with the exact Poisson interval; share() gives percentage of a total with Wilson's interval. They are separate functions because a share of a total is not a rate per population, and labelling one as the other is the most common error in a published table. rate_change() gives the change in a rate between periods, conditioning on the two counts and correcting for the exposure ratio rather than treating two rates as measured numbers.
  • Change tables — yoy() computes period-over-period change matched on the period's own value rather than on row order, so a missing year is a gap instead of a quietly multi-year comparison. A percent off a small base is withheld with its reason; a column already in percent is reported in percentage points; a ratio of counts carries the exact conditional-binomial interval. yoy_write() renders to HTML, PDF, CSV, TSV, JSON or Markdown, format taken from the file name, with nothing outside base R.
  • Banded categories — parse_bands() reads the interval labels publishers actually use ("2 to 5", "50+", "under 18") and returns bounds; band_sensitivity() measures how far a result moves as the open top band's assumed cap varies, which is the dependence every figure computed from banded data carries.
  • Concentration and tails — gini(), lorenz(), top_share(), and hill_tail_index(), which maximises the exact discrete likelihood because the closed-form continuity correction is badly biased at the small thresholds administrative counts start from.
  • Short-series trend — trend_test() (Mann-Kendall with Theil-Sen), step_change() (permutation scan over splits, not the best split's own test) and count_trend() (Poisson rate ratio per period). Meaningful at the five-to-ten annual points an open-data extract actually has.
  • Region-coded counts — expected_counts() for indirect standardisation, sir() with the exact Poisson interval, eb_rates() for Clayton-Kaldor shrinkage, funnel_limits(), and morans_i().
  • Points, and the regions that contain them — region_map_integrity(), region_map_compare() and region_map_second_route() check a point-to-region assignment on its own terms, since an error in it reproduces perfectly in every table built on it; region_map_from_points() recomputes one by point in polygon when sf is available; and region_coverage() reports the population of the regions holding a unit while saying, each time it prints, why that share is not a rate denominator.
  • Stock and flow — adp(), alos() and stock_flow() read the same person-days two ways, per day and per person, after Lakner (1976). When stays lengthen the two move in opposite directions, so stock_flow() reports both and the exact decomposition between them.

Verification

Every hash, keyed hash, checksum, key derivation, base64 and JSON output is compared against an independent implementation — digest, openssl, jsonlite and base R's own inflater — over a length sweep crossing each construction's block boundaries, so the digest this package records for a set of bytes is the number anybody else would compute for them. Published vectors are checked too: SHA-512 (FIPS 180-4), HMAC-SHA-256 (RFC 4231), PBKDF2-HMAC-SHA256, BLAKE2b (RFC 7693) and CRC-32 (ITU V.42). The statistics are anchored on base R (stats::poisson.test, stats::glm, stats::cor.test, stats::qpois) or on closed forms recomputed by hand.

The worked example in examples/otis-mrp/ goes further than checking the package: it recomputes 147 published year-over-year tables across 29 datasets — 8,214 cells — from the source data and compares every one, alongside the descriptives, the matched sample and the causal estimates. It also checks what a cell-by-cell comparison cannot: three of those datasets reach the same population by different routes, so a wrong grain rule would move both sides of a cell comparison together and pass, while a01 distinct individuals against c01 and c04 totals fails. The datasets are not shipped — point OTIS_DATASETS_DIR at a copy you have, or set OTIS_YOY_DOWNLOAD=1 to fetch them from the province.

The compiled kernels are published for LinkingTo, and a consumer package is built and run against inst/include/rmoriebricklayer.h as part of the test suite — a signature mismatch is a compile error, while a misregistered name compiles cleanly and fails only when called.

The XMSS signature scheme is byte-compatible with the RFC 8391 reference implementation. The whole 2500-byte signature for XMSS-SHA2_10_256 -- index, randomiser, WOTS+ signature and authentication path -- matches it exactly, checked against embedded vectors in the test suite so the check needs no network.

The standardised schemes are byte-identical to OpenSSL. ML-DSA (FIPS 204) at all three parameter sets and SLH-DSA (FIPS 205) at all twelve -- six over SHAKE, six over SHA-2 -- are implemented here, with no system dependency. Every one of the fifteen is checked against OpenSSL 3.5: in deterministic mode the two implementations produce the SAME BYTES, over several message and context lengths, and each verifies the other's signatures. OpenSSL's keys and the digests of its signatures are embedded in the test suite, so the check needs no network and no system library.

ML-KEM (FIPS 203) is here too, at all three levels, along with the pre-hashed variants of both signature standards and ML-DSA's external-mu interface. ML-KEM keys generated from the same seed agree with OpenSSL's byte for byte, its ciphertexts decapsulate here to the secret it reports, and a corrupted ciphertext produces the same rejection secret in both -- which is the check that catches a wrong compression width, since compressing and decompressing with the same wrong width round-trips perfectly.

Signing is fast enough to be tested unconditionally: an SLH-DSA s parameter set signs in about a second, down from seven, after the Keccak round was made branch-free, the tweakable hash stopped heap-allocating a few million times per signature, and the SHA-2 sets learned to resume from a cached midstate.

That cross-check is the claim, not reference parity. This implementation matched the pq-crystals and sphincsplus reference code byte for byte while disagreeing with the standards in two places -- FIPS 204 and FIPS 205 both prepend a context domain separator that the reference code omits, and FIPS 205 reads the FORS indices most significant bit first where SPHINCS+ read them least significant bit first. A signature scheme that verifies only its own output passes every security-property test there is, so only an independent implementation can find that class of bug.

It is also verified against its security properties: a valid signature verifies, and every tampering of the message, signature, authentication path, index or key fails.

Installation

Released version from CRAN:

install.packages("rmoriebricklayer")

Latest build from r-universe (tracks main ahead of CRAN):

install.packages(
  "rmoriebricklayer",
  repos = c("https://rootcoder007.r-universe.dev",
            "https://cloud.r-project.org")
)

Development version from GitHub:

# install.packages("remotes")
remotes::install_github("rootcoder007/rmorie-bricklayer")

Quick example

The shortest path, on the OTIS table that ships with the package:

library(rmoriebricklayer)
otis <- read.csv(system.file("extdata", "otis_a01_individuals.csv",
                             package = "rmoriebricklayer"))
a <- analyse_table(otis, value = "individuals", period = "year",
                   by = c("table", "group"), rounding = 5)
a                                   # what changed, how sure, what was withheld
report_analysis(a, "otis.html")     # one self-contained file

The full capsule, step by step:

library(rmoriebricklayer)

prov <- load_provenance("provenance.json")     # pinned source + schema + hash
res  <- resolve_via_ckan(prov)                  # find the resource on the portal
path <- friendly_download(res$url, "data.csv")  # download (Wayback fallback)
verify_sha256(path, prov$sha256)                # integrity check
df   <- validate_schema(read.csv(path), prov)   # schema-validated data frame

man  <- make_manifest(project = "my-study")
record(man, "input", path)                      # trace the input
write_manifest_json(man, "manifest.json")

Then ask whether the data itself moved, and sign the answer:

# Did the distribution change, not just the bytes?
capsule_drift(reference_extract, fresh_fetch)

# Authenticate the manifest so a verifier knows who produced it.
key <- pqc_keygen()                              # post-quantum, hash-based
sig <- capsule_sign(core_sha256(readLines("manifest.json")), key)
capsule_verify(core_sha256(readLines("manifest.json")), sig,
               signing_public_key(key))

See vignette("drift") for the distributional checks and vignette("provenance") for signing, Merkle pinning and manifest chains.

Part of the MORIE family

rmoriebricklayer is the reproducibility / provenance layer of the MORIE ecosystem, alongside rmorie and rmoriedata.

Citation

If you use rmoriebricklayer in your research, please cite the software:

Ruhela, V. S. (2026). rmoriebricklayer: Reproducible Data Capsules with Provenance and Fallback. https://github.com/rootcoder007/rmorie-bricklayer

BibTeX (or run citation("rmoriebricklayer") after installation for the entry stamped with the exact installed version, sourced from inst/CITATION):

@Manual{ruhela_rmoriebricklayer_2026,
  title  = {rmoriebricklayer: Reproducible Data Capsules with Provenance and Fallback},
  author = {Ruhela, Vansh Singh},
  year   = {2026},
  url    = {https://github.com/rootcoder007/rmorie-bricklayer}
}

See CITATION.cff for the machine-readable metadata GitHub's "Cite this repository" button uses.

License

AGPL-3.0-or-later.

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct. By contributing, you agree to abide by its terms.

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("rmoriebricklayer")

0.5.1 by Vansh Singh Ruhela, 17 days ago


https://github.com/rootcoder007/rmorie-bricklayer, https://rootcoder007.github.io/rmorie-bricklayer/


Report a bug at https://github.com/rootcoder007/rmorie-bricklayer/issues


Browse source code at https://github.com/cran/rmoriebricklayer


Authors: Vansh Singh Ruhela [aut, cre] (ORCID:


Documentation:   PDF Manual  


AGPL-3 license


Imports methods, stats, utils

Suggests digest, jsonlite, knitr, openssl, pkgdown, rmarkdown, sf, stringi, testthat

System requirements: libcurl (deb: libcurl4-openssl-dev, rpm: libcurl-devel).


Imported by rmoriedata.


See at CRAN