Provides a disciplined precheck, execution, and diagnostics workflow
for survey weighting and raking. Weight construction requires design-only data,
a verified external target, outcome-blind planning, and human approval before
weight locking, with bilingual reports for decision and statistical audiences.
Converts calibrated and replicate weights into standard survey designs,
provides optional broom-style result projections, and records serializable
production pipeline provenance. Supports fixed, predeclared soft calibration
tolerances and categorical entropy balancing from verified margins, plus panel
attrition weighting, high-influence unit diagnostics, Fay's balanced repeated
replication, and opt-in parallel execution for long runs.
Calibration methods follow Deville and Saerndal (1992)
WFC is an R package for reviewable survey weighting. Version 2.0 makes one rule non-optional: weights may be built only from declared design variables and an independently sourced, verified target. Study outcomes enter only after the weights are locked.
This design serves two kinds of users:
The same result can be shown as a short decision view or a detailed statistical view. The two views come from one object and must agree.
Public weighting functions reject raw data frames, ordinary target objects, demo targets, changed identities, self-approval by an agent, and runtime changes to ID or base-weight roles. Manual targets, target shrinkage, inline target moments, manual pipeline modes, and runtime margin injection are not supported.
These controls reduce accidental and ordinary misuse. They cannot prove that a source document is truthful or stop a determined person from editing open source code. Accountable human review remains necessary.
Install the development version from GitHub:
remotes::install_github("weiandata/WFC")
A production run normally starts with four files:
.source.dcf evidence record for that target table.Use wf_target_template() if the target-table layout is unfamiliar:
library(WFC)
dims <- wf_dims(
age_group = c("18-34", "35-54", "55+"),
region = c("north", "south")
)
wf_target_template(
"population-margins-template.csv",
dims = dims
)
The template creates the data file and its companion DCF form. Complete the source fields, update the checksum after the data file is final, and obtain the file before inspecting study outcomes.
WFC also installs synthetic format examples named safe-target-example.csv,
safe-target-example.xlsx, and their separate .source.dcf files. They are
marked demo-only and cannot enter production planning.
library(WFC)
dims <- wf_dims(
age_group = c("18-34", "35-54", "55+"),
region = c("north", "south")
)
design_only <- read.csv("survey-design.csv")
analysis_data <- read.csv("survey-outcomes.csv")
design <- wf_prepare_design(
design_only,
id = "person_id",
calibration = c("age_group", "region"),
base_weight = "base_weight"
)
Every column in design_only must have a declared design role. Keep outcomes
such as satisfaction, vote choice, approval, score, or pass/fail in
analysis_data.
Population-count target:
target <- wf_import_target(
data_file = "population-margins.csv",
source_file = "population-margins.csv.source.dcf",
dims = dims,
key_map = c(age_group = "age_group", region = "region"),
count = "population_count",
production = TRUE
)
Independent reference-sample target:
reference_target <- wf_import_reference(
data_file = "reference-sample.csv",
source_file = "reference-sample.csv.source.dcf",
dims = dims,
feature = "reference_weight",
production = TRUE
)
Import verifies source completeness, declared selection timing, demo status, and the SHA-256 checksum. It does not decide whether the source is scientifically appropriate; that remains part of human review.
cell_plan <- wf_plan_cells(
design,
target,
dims,
min_cell = 5,
max_weight_ratio = 4
)
plan <- wf_plan_weights(
design,
target,
dims,
method = "raking",
bounds = c(0.3, 3),
min_cell = 5,
cell_plan = cell_plan
)
plan$ready
plan$issues
is.null(plan$weights)
Planning does not calculate weights. It records the exact inputs, checks, method, limits, and any deterministic category merge for review.
approval <- wf_approve_plan(
plan,
approver = "Qualified reviewer full name",
role = "Statistician",
note = "Reviewed source, support, method, limits, and intended use"
)
An AI agent may prepare the plan but may not create this attestation for itself. The name and role must identify the actual human reviewer.
locked <- wf_execute_plan(
plan,
approval,
design,
target
)
Changing the design, target, plan, or approval breaks the identity chain and stops execution.
analysis_ready <- wf_attach_weights(
analysis_data,
locked,
id = "person_id",
weight_name = ".weight"
)
impact <- wf_assess_impact(
locked,
analysis_data,
id = "person_id",
outcomes = c("satisfaction", "approved")
)
Impact assessment describes what the already locked weights change. It cannot re-plan, reapprove, or overwrite the weights.
decision_view <- wf_report(
locked,
audience = "decision"
)
statistical_view <- wf_report(
locked,
audience = "statistician"
)
impact_detail <- wf_report(
impact,
audience = "statistician"
)
wf_audit_export(locked, "weighting-audit.json")
The decision view emphasizes status, main risks, and next actions. The statistical view exposes full tables, convergence information, identities, and provenance for further analysis.
Advanced users may call wf_calibrate(), wf_rake(), wf_poststrat(),
wf_auto_trim(), or wf_autoweigh() directly, but their first inputs must
still be the unchanged design and target objects:
raked <- wf_rake(design, target, tol = 1e-8)
bounded <- wf_calibrate(
design,
target,
method = "logit",
bounds = c(0.3, 3)
)
trim_review <- wf_auto_trim(
design,
target,
caps = c(2, 4, 6, 8)
)
ID and base-weight columns come from wf_prepare_design() and cannot be
overridden at this stage.
An agent should preserve whole WFC objects and their identities. A minimal handoff object can be prepared like this:
handoff <- list(
plan = plan,
plan_identity = plan$identity,
design_identity = design$identity,
target_identity = target$identity,
required_human_action = "Review and approve the unchanged plan"
)
If an agent attempts approval, WFC returns wf_error_safety. Agents should read
the stable payload and stop:
refusal <- tryCatch(
wf_approve_plan(
plan,
approver = "Automated agent",
role = "assistant",
actor_type = "agent"
),
wf_error_safety = function(condition) condition$data
)
refusal[c("code", "severity", "field", "next_actions")]
Do not edit identities, manufacture a human name, widen limits, switch targets, or retry after viewing outcomes.
See Migrating from WFC 1.x to WFC 2.0. Behavior that selected desired results has no compatibility switch and no supported replacement.
WFC is licensed under GPL (>= 2). Copyright and contributor information are in
inst/COPYRIGHTS and DESCRIPTION.