Implements instrumental variables estimation and inference for
one endogenous regressor and one-way clustered errors. Includes the
cluster-jackknife IV estimator (CJIVE) of Frandsen, Leslie and McIntyre
(2025)
Clustered instrumental variables estimation and weak-instrument-robust inference for R.
clusterIV estimates CJIVE and implements cluster-jackknife Anderson–Rubin
and score inference for one endogenous regressor under one-way clustering.
CJIVE removes own-cluster first-stage terms that remain in 2SLS and
observation-level jackknife IV when observations are dependent within
clusters.
Setup | Syntax | Description | Options | Simulated example | Worked examples | Validation | FAQ | References
Install the released version from CRAN:
install.packages("clusterIV")
The weak-instrument-robust inference functions documented below
(cjar(), cjscore(), iv_infer()), fixed-effect absorption, and the
formula interfaces require clusterIV 0.2.0 or later; version 0.1.0
contains cjive() and iv_compare() only. Check the installed version
with packageVersion("clusterIV").
The vector interfaces expose all computational options directly:
cjive(y, x, z, cluster,
controls = NULL, weights = NULL, level = 0.95, intercept = TRUE,
method = c("auto", "dense", "leaveout_mean"),
fixed_effects = NULL, inference = c("asymptotic", "t"))
cjar(y, x, z, cluster,
controls = NULL, fixed_effects = NULL, weights = NULL,
intercept = TRUE, level = 0.95, beta0 = 0,
calibration = c("chisq", "normal"),
variance = c("plain", "crossfit"))
cjscore(y, x, z, cluster,
controls = NULL, fixed_effects = NULL, weights = NULL,
intercept = TRUE, level = 0.95, beta0 = 0,
variance = c("plain", "crossfit"))
iv_infer(y, x, z, cluster,
controls = NULL, fixed_effects = NULL, weights = NULL,
intercept = TRUE, level = 0.95, beta0 = 0,
calibration = c("chisq", "normal"),
inference = c("asymptotic", "t"),
tests = c("cjar", "cjscore"),
variance = c("plain", "crossfit"))
iv_compare(y, x, z, cluster,
controls = NULL, weights = NULL, level = 0.95,
intercept = TRUE, fixed_effects = NULL,
inference = c("asymptotic", "t"))
The same functions accept either restricted formula layout:
y ~ x | z # controls supplied through controls =
y ~ x | z | fe1 + fe2 # absorbed fixed effects
y ~ exog1 + exog2 | x ~ z # controls inside the formula
y ~ exog | fe1 + fe2 | x ~ z
Formula methods additionally accept data, subset, and
na.action = stats::na.omit. Formula sections are additive lists of bare
variable names. Transformed terms, multiway clustering, and multiple
endogenous regressors are not supported.
The cluster jackknife originates with Ligtenberg (2023, the first version of
the cited preprint). Frandsen, Leslie and McIntyre (2025) use it to construct
CJIVE. On the default dense path, the package residualises the outcome,
endogenous regressor, and instruments by Frisch–Waugh–Lovell before applying
the cluster jackknife. Consequently, cjive() and the CJIVE row of
iv_compare() agree for the same design. The explicitly requested
method = "leaveout_mean" shortcut is the documented exception.
iv_infer() is the main inference workflow. Its default panel contains:
beta0.The selected procedures share preprocessing, the instrument factorisation, leave-cluster-out work, and cluster sums.
The leave-cluster-out first stage works in whitened instrument coordinates and chooses the smaller of the cluster-size and instrument-dimension solves. It uses triangular solves rather than an explicit inverse. High-dimensional fixed effects are absorbed by a matrix-free joint projection and checked for weighted orthogonality before estimation; no dummy matrix is formed.
y and x are numeric vectors. Exactly one endogenous regressor is
supported.z is a numeric vector or matrix of excluded instruments, or a factor or
character grouping instrument for a judge/examiner design.cluster identifies one clustering dimension. A formula must name one bare
variable.weights supplies finite, strictly positive precision weights.controls supplies ordinary covariates. An intercept is included unless
intercept = FALSE.fixed_effects supplies one or more factors to absorb without dummy
expansion. In formulas, fixed effects can instead occupy the dedicated
formula section.method = "auto" and method = "dense" use the package-wide dense FWL
convention. method = "leaveout_mean" is an explicit special case for a
pure grouping-instrument design; it is never selected automatically.level sets the reported confidence level; beta0 sets the null tested by
CJAR and CJS.variance = "plain" uses Ligtenberg's plain variance estimator.
variance = "crossfit" uses the leave-cluster-out cross-fit construction
in Section 5.4 of Ligtenberg (2025). The cross-fit estimator solves one
leave-out system per cluster pair — G(G-1)/2 solves, inherent to its
definition — with each solve dispatched on the smaller of the stacked
cluster size and the instrument dimension.calibration = "chisq" uses the paper's shifted/scaled chi-square
approximation for CJAR; "normal" uses its asymptotic normal reference.
Neither is an exact finite-sample law.inference = "t" replaces the CJIVE normal critical value with a
t(G - 1) reference convention. It does not change the coefficient or
standard error and is not an exact finite-sample law under arbitrary
within-cluster dependence.tests selects any subset of "cjar" and "cjscore". Use NULL for a
CJIVE-only panel.Dense-path fits from cjive(), cjar(), cjscore(), and iv_infer()
report the maximum within-cluster leverage maxlev and the clustered
Montiel Olea–Pflueger effective first-stage statistic F_eff. CJAR and CJS
additionally report their polynomial tail diagnostics F_CJ and F_CJS.
iv_compare() reports estimates only, without these diagnostics. The
diagnostics describe different aspects of the design and are not automatic
model-selection rules.
library(clusterIV)
set.seed(42)
G <- 40L
ng <- 8L
n <- G * ng
dat <- data.frame(
cluster = rep(seq_len(G), each = ng), # dependence cluster
judge = factor(sample(1:8, n, replace = TRUE)), # grouping instrument
w = rnorm(n) # included covariate
)
cluster_shock <- rnorm(G)[dat$cluster]
dat$x <- 0.45 * as.numeric(dat$judge) + # endogenous regressor
0.30 * dat$w + cluster_shock + rnorm(n)
dat$y <- 1.20 * dat$x + 0.50 * dat$w + # outcome
cluster_shock + rnorm(n)
fit <- iv_infer(
y ~ w | x ~ judge, # outcome/control | endogenous ~ instrument
data = dat, # analysis data
cluster = cluster, # one-way clustering
beta0 = 1, # point null for CJAR and CJS
level = 0.95, # confidence level
variance = "plain" # Ligtenberg's plain variance estimator
)
fit
coef(fit$cjive) # CJIVE point estimate
confint(fit$cjar) # weak-ID-robust CJAR confidence set
fit$cjscore$p.value # CJS p-value for beta = 1
plot(fit) # p-value curves and reported sets
For an estimator comparison using the same residualised data and cluster-robust sandwich convention:
iv_compare(y ~ w | x ~ judge, data = dat, cluster = cluster)
This returns OLS, 2SLS, observation-level IJIVE (retaining the historical row
label "JIVE"), and CJIVE. It matches the four-estimator layout of FLM's
Table 1; it is not an empirical replication of that table.
Two precomputed vignettes run the package on real, named datasets:
vignette("queens-workflow") — the full workflow on Dube and Harish
(2020): 3,586 polity-years, 176 reign clusters, instruments that are not
strong, dense controls plus three absorbed fixed-effect dimensions. It
reproduces the published 2SLS anchors, walks through every diagnostic,
runs the published specification set, and shows the tidy()-to-LaTeX
export path.vignette("miami-bail") — the judge-leniency design of Frandsen, Leslie
and McIntyre (2025): 91,421 defendants, 146 judges, over one hundred
instrument columns, clustered by courtroom shift. It shows the workflow
at scale, including the instrument-hygiene step every judge-dummy design
needs.Neither dataset can be redistributed inside the package, so the vignettes ship with their outputs baked in; each states where to obtain the data and verifies it by hash.
The implementation is validated against external references and frozen brute-force oracles rather than against itself. In brief: CJIVE agrees with FLM's own released Stata implementation run on the deposited Miami-Dade sample (pointwise-identical constructed instrument; coefficient to the ado's single-precision limit); every CJAR/CJS statistic, variance polynomial, and confidence-set endpoint is gated at 1e-10 against literal dense transcriptions of the papers' formulas that were written and frozen before the fast paths existed; and at singleton clusters the tests reduce exactly to the Mikusheva–Sun and Matsushita–Otsu statistics.
iv_infer(). One call returns the CJIVE estimate, the CJAR confidence
set, and the CJS test from one shared preprocessing pass, and its printed
panel carries every diagnostic discussed below. cjive(), cjar(), and
cjscore() are the same computations as focused standalone calls;
iv_compare() adds the OLS/2SLS/IJIVE comparison row layout.
With clustered observations and many instruments, leaving out only the focal observation does not remove first-stage terms contributed by other observations in its cluster. CJIVE leaves out the entire cluster.
z be a factor rather than a numeric matrix?Use a factor or character vector for a grouping instrument such as judge or examiner identity. With an intercept or absorbed fixed effects, the package uses reference coding; it uses one column per level only when neither is present. Supply a numeric matrix when the excluded instruments are already constructed columns.
Report the CJAR confidence set as the headline weak-instrument-robust
result, and the CJS p-value when a specific point null is the question.
The two tests are both valid under weak and many instruments but have
different power profiles: the score test concentrates power near the null,
the AR test retains power against distant alternatives. Their confidence
sets therefore need not coincide, and neither disagreement between them
nor disagreement with the Wald interval is an error. When the CJAR and
Wald intervals disagree materially, the Wald interval is the one whose
validity is in question — check F_eff against its critical value.
The Wald interval assumes the estimator is approximately normal, which requires strong instruments; the CJAR set inverts a test that is valid without that assumption. With strong instruments the two nearly agree; with weak or many instruments the robust set is typically wider, shifted, or unbounded — that is information, not a bug. An unbounded set says the data cannot rule out arbitrarily large coefficients at this level.
F_CJ and F_CJS mean, and what is a "good" value?Each compares the tail behaviour of its test's inversion polynomials with
the critical value: F_CJ > crit holds exactly when the CJAR confidence
set is bounded, and likewise for F_CJS. They are boundedness criteria
in the spirit of a first-stage F, printed next to their own critical
values — compare each against the printed critical value, not against the
conventional rule of 10, and do not interchange them with F_eff: the
effective F and F_CJ measure different objects and can diverge as the
instrument count grows.
maxlev near one mean?maxlev is the largest spectral norm of a within-cluster projection block.
A value near one means that deleting one cluster leaves a nearly singular
instrument Gram matrix. Treat it as a conditioning warning and inspect the
instrument design and cluster sizes.
An unbounded set is a valid weak-identification outcome: the selected test does not exclude arbitrarily large coefficient values. An empty set means no coefficient value is accepted at the selected level. It is not, by itself, proof of invalid instruments or a misspecified model; inspect the design, assumptions, and numerical diagnostics.
method = "leaveout_mean"?Only when you deliberately want the printed leave-out group-mean formula for
a pure grouping-instrument design with intercept-only controls. The default
dense FWL path is the package-wide convention. The two forms differ through
the intercept direction by order n_g / n (about 1 / G in a balanced
design); the group-mean form is never chosen automatically.
Fast absorption makes a many-controls design easy to fit, but it does not make
the usual plug-in sandwich reliable in every such design. In one Table 1
simulation of Kolesar, Min, Wang and Zhang (2026), with n = 600, CJIVE's
nominal 5% test rejects 7.9% with 50 instruments and 50 controls, and 53.4%
with 150 of each. Their L2CO/L3CO corrections are not implemented here.
Ligtenberg's Section 5.3 result does cover cluster-specific controls, including
cluster fixed effects, because they introduce dependence only within clusters;
it should not be read as a general many-controls result.
tidy() and glance() methods are registered through the optional
generics package and return plain data frames — one row per inferential
procedure, with the complete confidence set retained in a conf.set
list-column and its topology in explicit shape/n.components columns.
They feed directly into knitr::kable(format = "latex") or
modelsummary; the queens vignette ends with a worked export. A disjoint
or unbounded set is never silently flattened into two numbers.
CJIVE reduces to improved JIVE after the package's FWL residualisation. Plain CJAR has the Mikusheva–Sun jackknife-AR numerator/statistic relationship, and plain CJS has the Matsushita–Otsu algebraic identity when there are no supplied controls or the data are already partialled. The cross-fit option uses Ligtenberg's leave-two-cluster-out variance and should not be given those plain-variance labels.
CJIVE is tested against dense leave-cluster-out calculations, a literal R translation of FLM's Stata/Mata algorithm, and FLM's own released Stata implementation run on the deposited Miami-Dade data. CJAR, CJS, and cross-fit variance are tested against direct implementations of their defining formulas, frozen before the fast paths were written. These checks establish numerical agreement with the stated algorithms; the cited papers supply the statistical assumptions and asymptotic results.
Ackerberg, D. A. and Devereux, P. J. (2009). Improved JIVE estimators for overidentified linear models with and without heteroskedasticity. Review of Economics and Statistics, 91(2), 351–362.
Frandsen, B., Leslie, E. and McIntyre, S. (2025). Cluster Jackknife Instrumental Variables Estimation. Review of Economics and Statistics. doi:10.1162/rest.a.263. See the estimator and empirical comparison.
Kolesar, M., Min, P., Wang, W. and Zhang, Y. (2026). Cluster-Robust Inference for Quadratic Forms. arXiv:2602.13537. See Table 1 and Sections 3–4.
Ligtenberg, J. W. (2025). Inference in clustered IV models with many and weak instruments. arXiv:2306.08559v3. See Sections 3–5 for CJAR, CJS, controls, and cross-fit variance.
Montiel Olea, J. L. and Pflueger, C. (2013). A robust test for weak instruments. Journal of Business & Economic Statistics, 31(3), 358–369.