15 Changelog
16 Changelog
The changelog is maintained in NEWS.md in the package source. The chapter below is auto-rendered from that file on every build of the docs site.
17 didgpu 0.1.2
17.1 Build
DIDGPU_NO_CUDA=1forces a CPU-only build even on a machine with the CUDA toolkit installed. Honoured bysrc/Makevars,src/Makevars.win,configureandconfigure.win.Without it there was no way to reproduce the configuration CRAN actually builds (their machines have no CUDA) on a developer box that does.
R CMD check --as-crannow reports:CUDA build Status: 1 NOTE (checking compiled code) CPU-only build Status: OKThe NOTE names only
cudart64_13.dllanddidgpu_cuda.dll– NVIDIA’s runtime and the package’s CUDA helper. Neither is an R extension, so neither callsR_registerRoutines, and neither exists in a CPU-only build. That the NOTE is CUDA-only is now demonstrable rather than asserted.
17.2 New arguments (Callaway-Sant’Anna)
didgpu_cs(base_period = c("varying", "universal")), defaulting to"varying"asdid::att_gtdoes. didgpu previously always used the universal base period, so under defaults every pre-treatment cell of the event study differed fromdid. Post-treatment cells and the overall ATT are the same under both. Behavior change: to keep the previous output, passbase_period = "universal".didgpu_cs(allow_unbalanced_panel = FALSE), as indid. On an unbalanced panel didgpu used, cell by cell, whatever units happened to be observed in both periods – which matched neither ofdid’s modes (on a real non-trade sales-growth panel, -0.060 againstdid’s -0.096). NowFALSE(the default) balances the panel first, dropping units not observed in every period, andTRUEuses the repeated-cross-section estimators of Sant’Anna and Zhao (2020): a port ofDRDID::reg_did_rc,std_ipw_did_rcanddrdid_rc, matching them to 2e-15 in the ATT and 3e-14 in the influence function. Behavior change on unbalanced panels.didgpu_cs(first_treat = NULL): the column holding each unit’s first-treatment period,did’sgname. Cohorts were always inferred from the treatment column, which puts a unit whose adoption-period row is missing into a later cohort; withallow_unbalanced_panel = TRUEthat is common, and didgpu now warns when it happens withoutfirst_treat.
Against did 2.5.1, didgpu_cs() now matches every ATT(g,t), the overall, event-study, group and calendar aggregates and all their SEs to <= 6e-16, for OR / IPW / DR, both base periods, both control groups, and balanced, balanced-by-dropping and repeated-cross-section panels. Also fixed along the way: with control_group = "notyet" and a universal base period, controls are now chosen at the later of the two periods compared, as in did; and an aggregated SE that is numerically zero (the universal base period’s own event time) is reported as NA, as aggte does.
did 2.3.0 has two bugs, both fixed in 2.5.0, that make it disagree with this: an integer gname silently removes the never-treated group (bcallaway11/did#264), and under allow_unbalanced_panel = TRUE its influence functions land on the wrong units, so every SE that combines cells changes when the units are merely renumbered (overall SE 0.071487 vs 0.072007 on one test panel; did 2.5.1 and didgpu both give 0.071407 either way). test-cs-unbalanced.R compares those combined SEs only against did >= 2.5.0.
17.3 Bug fixes
dCDH disagreed with
DIDmultiplegtDYNon panels with gaps or repeated firm-years. Every balanced-panel parity test passed, but on the annual tax panels of a real application – which drop loss-making years, leaving holes in 40-50% of firms’ histories – the estimates moved, most underonly_never_switchers = TRUE:tric_tax Effect_2 didgpu -0.0156 reference -0.0145 ntr_tax effects off by up to 5e-4, placebos by up to 7e-4.prep_panel()was missing four things the reference does before it balances the panel (did_multiplegt_main.R:119-300), and now does them in the reference’s order:- Collapsing un-aggregated data. A repeated (group, time) becomes one cell: weighted-mean outcome and treatment,
N_gt= the summed weight. didgpu kept both rows, double-counting them and – lags being taken by row – misaligning that group’s differences.ntr_taxhas 9 such firm-years. - The missing-treatment rules. When the last observation before a switch is not the period right before it, the switch date is unknown; the reference demotes the group to a control truncated at its last clean period, which under
only_never_switchersmakes it a never-switcher. Missing treatment inside a known span is imputed. - The sample restrictions, with
Gfixed between them. Cohorts with no variation in switch dates are dropped before the group count is fixed; (period, cohort) cells with no control, and switchers whose post-switch treatment averages back to baseline, are dropped after it and so still count inG. T_gper cohort, not per group – the last period in which the group’s baseline cohort still has a usable control.
ntr_taxandtric_taxnow agree to ~1e-17 in effects, placebos, the ATE and every SE, binary and multivalued, and theNandSwitcherscolumns agree too. The new cohort drop removes whole groups, which left holes in the internal group ids; the C++ and CUDA layouts assume ids 1..G, and briefly returned 0.07 where the R backend returned 0.47 before the ids were renumbered.test-gappy-panels.Rpins all of it, including backend agreement.- Collapsing un-aggregated data. A repeated (group, time) becomes one cell: weighted-mean outcome and treatment,
The CUDA SVD used by
didgpu_fect(method = "ife" / "mc")ignored GPU allocation failures.fect_svd_softthreshold_dev()checked none of itscudaMalloccalls (a TODO said so), and the truncated SVD missed two. If an allocation failed – plausible when several GPU jobs share one card – the next kernel wrote through a null pointer and R died with exit 139 and no error message. Every allocation is now checked; a failure returns an error code, and the R side, which already treated a failed GPU SVD as “fall back to the CPU”, does exactly that.This was found while investigating a segfault (exit 139) in a robustness suite run alongside other GPU jobs, but it does NOT explain that crash: the crash came during
method = "fe"steps, which never call this SVD code. That crash is still unexplained – it did not reproduce with the same call and data on an idle machine.Callaway-Sant’Anna aggregates were NA whenever a single ATT(g,t) cell could not be estimated. On an unbalanced panel a cohort can have no treated unit left in some period; that cell carries
att = NAwith weightn_treated = 0, and in RNA * 0is stillNA. On the trade subsamples of a real application the overall ATT was NA on all twelve. Unestimable cells are now dropped before weighting – exactly whatdid::aggte(na.rm = TRUE)does (diditself stops on them otherwise) – with a message saying how many; event times, calendar periods and cohorts left with no cell are dropped from the output, asdiddrops them. The CS placebo test excludes them the same way.Standard errors were a bootstrap approximation of the reference’s, not the reference’s.
DIDmultiplegtDYNcomputes SEs analytically from the estimator’s asymptotic linear representation. didgpu had no analytic path at all: withbootstrap_reps = 0theSEcolumn came back entirelyNA, and above zero it reported the bootstrap SD, which is a different estimator:didgpu (50 reps) 0.056174 0.051302 0.054590 0.059531 0.047412 reference 0.055504 0.054278 0.057839 0.052243 0.055922tests/testthat/test-reference-parity.Rhad said so in a comment – “SEs come from different estimators (bootstrap vs. analytic) and are not compared here” – so the gap was known and simply never closed, whileREADMEclaimed a bit-for-bit match on SEs.didgpu now computes the reference’s own influence functions. Effects, placebos and the ATE all match
DIDmultiplegtDYNto machine precision (0 to 2.1e-17), clustered and unclustered, on every backend, and acrossnormalized, multivalued and non-absorbing treatment, non-zero baseline dose, andswitchers = "in"/"out". The joint nullity tests come from the same influence functions via the covariance identity and match to 5e-16.bootstrap_repsnow defaults to0. Nothing in the default output needs a resample any more. This was also the package’s real performance problem: the reference does one pass, while didgpu’s old default of 100 reps meant 101 full fits. On a 400-unit x 104-period panel withonly_never_switchers = TRUE:DIDmultiplegtDYN 18.50s didgpu backend = r 3.13s 5.9x faster didgpu backend = auto 2.56s 7.2x fasterThe same specification previously ran ~5x SLOWER than the reference despite didgpu being 10-25x faster per fit. Setting
bootstrap_repsabove zero still works and no longer changes the reportedSE.One gap remains, and it is deliberate. With
controls(orcontinuous, which adds polynomial controls of its own) the reference subtracts a control-estimation correction from the influence function –part2_switchindid_multiplegt_dyn_core.R:502-535– that didgpu does not compute. Reporting the uncorrected number would have been wrong by about 9e-05, so the analytic path is switched off there and the bootstrap remains the SE source;didgpu()says so rather than returning a silentNAcolumn. Point estimates with controls are unaffected and still match the reference exactly (1.1e-16).vcov(), the joint tests anddidgpu_joint_placebo()now come from one analytic covariance matrix. When the SEs first became analytic,vcov()was still the bootstrap covariance, so its diagonal stopped matchingSE^2(0.0039 vs 0.0132) anddidgpu_joint_placebo(), which reads it, stopped reproducingp_jointplacebo(0.641 vs 0.678). The matrix is now built from the same influence vectors as the SEs –SE^2on the diagonal, polarisation off it – and all three read it.Joint tests under
normalized = TRUEwere wrong in the first analytic version. The reference divides each influence vector bydelta_kbefore polarising (did_multiplegt_main.R:1170-1175); the first version polarised the raw vectors against the normalised SEs. Now pinned againstDIDmultiplegtDYNon a multivalued design wheredelta_k != 1.backend = "reference"reportsDIDmultiplegtDYN’s own SEs and joint tests. It delegates to the reference for the point estimates but discarded the reference’s analytic SEs and reported the SD of didgpu’s outer bootstrap instead – a different estimator, and one that no longer matchedbackend = "r". Itsvcov()remains the bootstrap covariance, since the reference exposes no covariances to build one from.didgpu_compare()now fails on theSEand CI rows too. It scored only theEstimaterows, on the since-removed grounds that “reference SEs are analytical, ours are NA in that case”. Those columns now carry comparable numbers, so they are held to the same tolerance.Resuming a checkpoint written by a different version of didgpu is now refused. Cells carry whatever estimator produced them, and two have changed in this release (the ATE became
Av_tot_eff; the SEs became analytic), so resuming would silently mix old and new numbers in one result. Every checkpoint now records an estimator revision, anddidgpu()refuses to resume one written under a different revision – or one with no revision, which predates the stamp – and explains why. The package version alone could not do this: all of these changes shipped as 0.1.2.ATEdid not matchDIDmultiplegtDYN’sAv_tot_effunless treatment was binary and absorbing. didgpu reported the switcher-weighted mean of the per-event-time effects,ATE = sum_k (N_k * DID_k) / sum_k N_kbut the reference’s
Av_tot_effis an average total effect per unit of treatment and carries its own denominator (U_Gg_den_XX, built fromdelta_D_i_XXindid_multiplegt_dyn_core.R):ATE = sum_k (N_k * DID_k) / sum_k (N_k * delta_k)where
delta_kis the average treatment change among the event-time-k switchers. When treatment is binary and absorbing everydelta_kis exactly 1, the denominators coincide, and the two agree to the last bit – which is why the test suite and every worked example missed this.On one panel, holding the outcome fixed and changing only how the treatment is coded:
treatment coding didgpu (old) reference ratio binary, absorbing +0.35616730 +0.35616730 1.000 binary, non-absorbing +0.35616730 +0.47488973 0.750 multivalued dose (1, 3, 5) +0.35616730 +0.12020646 2.963 non-monotone (0 -> 2 -> 1) +0.35616730 +0.23744487 1.500The old column is constant: the reported ATE did not respond to the magnitude of the treatment at all, only to its timing. All four now agree with
DIDmultiplegtDYNto machine precision on every backend. Two consequences, both matching the reference:normalizeddoes not affect the ATE (it rescales only the per-event-time effects), and multiplying the treatment variable by a constantcnow divides the ATE bycinstead of leaving it unchanged.The old formula was duplicated in
.backend_cpu()and.backend_cuda()as well as the R core, andbackend = "auto"resolves to CUDA wherever a GPU is present, so this was the default path on a CUDA machine. All three now call one.ate_weighted().didgpu()’s help gained a@returnsection stating whatATEis; nothing in the docs had defined it before.Checkpoint cells written by an earlier version store the old ATE; resuming into such a directory is now refused outright (see below).
A user column named after one of the arguments broke several entry points.
[.data.tableevaluates itsiandjexpressions with the table’s COLUMNS in scope, and the affected functions filtered withd[!is.na(get(outcome)) & !is.na(get(group)) & ...]. A panel carrying a column literally calledoutcome,group,time,treatmentorweighttherefore shadowed the argument holding that column’s NAME, andget()received a vector instead of a string:numeric column -> "invalid first argument" character column -> "first argument has length > 1"Those are ordinary names in an analysis frame, so this failed on full panels while the same rows in a minimal four-column frame worked – which presents as “extra columns break it” rather than as a name collision. Fixed in
didgpu_bacon(),didgpu_fect(),didgpu_did_continuous(),didgpu_fhs()and theweightargument ofdidgpu(); column vectors are now resolved before subsetting. These were hard errors, never silent miscalculations, so no previously reported estimate is affected.didgpu_compare()required the caller to attach polars.DIDmultiplegtDYN(>= 2.x) calls polars through barepl$...but only Suggests it, so it never attaches polars itself. Withoutlibrary(polars)the reference fit died inside polars – “attempt to apply non-function”, or “Evaluation failed in$with_columns()” depending on version – anddidgpu_compare()surfaced that with nothing pointing at the cause. It now attaches polars when available and reports plainly what is missing when it is not.didgpu_fect(method = "mc")selected too little shrinkage. The lambda cross-validation held out RANDOM SCATTERED control cells. A low-rank model interpolates isolated holes far more easily than it extrapolates the contiguous block it actually has to predict, so the out-of-sample MSE was minimised at too small a lambda and the fit kept spurious factors.Validation now uses a rolling origin – the last
cv_nobsuntreated observations of each unit, shifted back one period per fold – and only on EVER-TREATED units, since those are the units whose counterfactuals must be predicted and whose task is extrapolation. Never-treated units have full histories, so holding out their final periods is an easier problem, and as the usual majority they dominated the MSE.Agreement with
fect::fect:panel before after long histories 6.911e-03 1.529e-05 mixed short histories 3.208e-03 4.009e-06Note also that
fect’s lambda is expressed inlambda / (T * N)units, so its reportedlambda.cvcorresponds tolambda * T * Nhere; passing its raw value todidgpu_fect()will not reproduce its fit.A caution when comparing against the reference: with neither
lambdanorCVsupplied,fect::fect(method = "mc")degenerates to pure two-way FE – its own message reads “No lambda is supplied. FEct is applied.” – and returns exactly itsmethod = "fe"answer. Comparisons against that run measure FE, not MC.All three fect methods now track the reference:
fe7.8e-07,ife6.9e-08,mc4.0e-06.didgpu_fect()kept units whose counterfactual is not identified, which mademethod = "ife"badly wrong – sometimes the wrong sign. A unit’s counterfactual comes only from its UNTREATED observations: the unit fixed effect needs at least one, an r-factor loading needs several. fect drops units below a minimum (fect.default):min.T0is1formethod = "fe"and5for"ife","mc","both","gsynth"and"cfe". didgpu kept every unit and extrapolated.With a weak factor structure that is a nuisance; with a strong one it is fatal. On a 100-unit panel with two latent factors and a TRUE ATT of +1.0:
factors truth fect didgpu before didgpu after L ~ 0.5 1.0 +1.00142 +0.97446 +1.00206 L ~ 1.5 1.0 +0.99880 +0.72146 +0.98872 L ~ 3.0 1.0 +1.00064 -0.17445 +0.94980The clue was that
fect’s returnedlambdawas 84 x 2 on a 100-unit panel: it had silently dropped 16 units, exactly those with fewer than 5 untreated periods.didgpu_fect()gains a documentedmin_T0argument.NULL(the default) follows fect’s rule;$n_units_droppedand$min_T0are returned and dropped units are warned about. This subsumes the earlier always-treated fix, since those units have zero untreated periods.Note a consequence: because the defaults differ by method,
method = "ife", r = 0does NOT generally equalmethod = "fe"– they fit different samples.fectbehaves the same way, and didgpu now reproduces both of its numbers on a panel where 22 of 60 units have short histories (fe+0.302820,ife(r = 0)+0.236833). Passmin_T0 = 1to make them agree.didgpu_fect(method = "fe")stopped before converging. The fit had a second stopping rule,abs(loss - prev_loss) < tol, wherelossis a SUM of squared residuals. Its absolute change falls below a tolerance meant for parameter units long before the parameters settle, so it fired first: on a 60x10 panel the fit exited after 6 iterations withdelta = 2.0e-04against the requested1e-05and reported itself finished. Convergence is now judged on the parameters alone.feconverges in 9 iterations and its agreement withfect::fectimproved from 2.81e-05 to 7.80e-07.
17.4 New features
didgpu_fect()reports convergence diagnostics and warns when a fit does not converge. The solver already producediter,deltaandlambdaper cell but nothing surfaced them, so a caller could not tell a converged fit from one that had exhaustedmax_iter– both returned a number and looked identical. Results now carry$diagnostics(iter,delta,converged,tol,max_iter,lambda,n_nonzero_singular) and a non-converged fit warns.This immediately exposed two problems that had been invisible:
fewas exiting after 6 iterations (fixed above), andmethod = "ife"exhaustsmax_iterwithout converging at default settings. The latter is a known outstanding defect –ifeis materially wrong as the factor structure strengthens – and is not yet fixed.The CUDA backend silently ignored most estimation options. Its compatibility guard tested only
controls,weightandtrends_nonparam, so every other option passed through to a kernel that does not implement it and a wrong number came back with no warning and no fallback.normalizedwas the damaging case. With a multivalued treatment,backend = "cuda"returned the UNnormalised effects, which do not vary with dose – so a dose-response analysis looked exactly as though the treatment had been binarised. AgainstDIDmultiplegtDYNwithnormalized = TRUEon a three-level dose:backend |diff| vs reference r 5.551e-17 cpu 5.551e-17 cuda 8.465e-01backend = "auto"resolves to"cuda"whenever a GPU is present, so this was the default path on a CUDA machine.dont_drop_larger_lowerwas separately dropped, because the CUDA path called.prep_panel()without forwarding it.The guard is now identical to the CPU backend’s and names the offending option when it falls back. All backends now agree with the reference to 5.551e-17 on a normalised multivalued treatment.
To be explicit, since this was reported as binarisation:
didgpu()does NOT binarise the treatment. Under the defaultnormalized = FALSEthe dCDH dynamic effect is the average outcome change among switchers, keyed on switching TIMES rather than dose magnitude, so rescaling a treatment that keeps the same switching pattern leaves the estimate unchanged – inDIDmultiplegtDYNtoo. Dose enters undernormalized = TRUE, where it moved the estimate by 0.56 and matched the reference exactly.didgpu_fect(method = "mc")ignored fixed effects and was not invariant to the level of the outcome. Athey et al. (2021) estimateY = L + unit FE + time FE, penalising the nuclear norm ofLalone. The implementation soft-thresholded the RAW outcome matrix, so the penalty shrank the level itself and the residualY - Y_hatabsorbed it.lambdacompounded this by being scaled to the singular values of the raw matrix, so the penalty also grew with the level.On a known-zero DGP the reported ATT moved with a pure location shift –
Y + 0gave+0.269,Y + 100gave+2.551– whilefe,ifeandfect::fectall returned-0.024at every level. On a positive, trending outcome it manufactured large, monotonically rising, significant effects wherefeandifeboth found a null.The fit now removes two-way fixed effects, soft-thresholds the SVD of the RESIDUAL, and adds the fixed effects back;
lambdais scaled to that residual. The estimator is now exactly level-invariant, recovers the known-zero null (-0.031), and tracksfect::fectto 6.9e-03 (was 2.9e-01). It is not yet bit-for-bit:lambdaselection still differs fromfect’s own cross-validation.didgpu_loo()reported a pre-treatment placebo instead of the ATT. The headline for adidgpu_cs_resultwasfit$aggregation$estimate[1], documented as working “for all four aggregations”. It does not: with the defaultaggregation = "event"row 1 is the MOST NEGATIVE event time, i.e. the longest pre-treatment horizon.It also concealed itself, because dropping a single entity seldom changes which cells populate the earliest lead – so nearly every entity returned an IDENTICAL estimate, which reads as “no entity is influential” rather than as a bug. On a 56-year panel with a 1985 cohort it returned the
e = -42cell (-0.006317) instead of the ATT (-0.032831), for bothby = "cohort"andby = "unit".didgpu_loo()now always reports the OVERALL ATT – then_treated-weighted mean over post-treatment cells – regardless of which aggregation the fit requested, so the answer no longer depends on an unrelated display choice. Both the fast re-aggregation path and the refit path were affected and both are fixed.didgpu_cs()influence functions were wrong; multiplier-bootstrap standard errors were two to eight times too narrow. The per-cell influence function was taken to be the treated units’ demeaned residual. It is not. It needs (a) the treated arm normalised byE[D], (b) the comparison arm normalised byE[p(X)(1-D)/(1-p(X))]– notE[D]– and (c) estimation-effect terms for the nuisance parameters, which load onto CONTROL units.ORset every control unit’s influence to zero outright;IPW/DRomitted the normalisers. Measured againstdid::att_gt()on a 200-unit panel,bootstrap_kind = "multiplier"returned SEs at ~0.13x (OR) and ~0.50x (IPW/DR) of the correct width. Point estimates were correct throughout, so nothing looked wrong. The defaultbootstrap_kind = "cluster"never touches these and was correct.The three per-cell estimators now mirror DRDID – the package
diditself calls – function for function (reg_did_panel,std_ipw_did_panel,drdid_panel). Verified againstdid::att_gt()to machine precision for all three methods, with and without covariates: max |diff| ~4e-16 on ATT(g, t) and ~1e-16 on its SE.Anyone who used
bootstrap_kind = "multiplier"should re-run: confidence intervals were far too narrow and p-values far too small.The batched CUDA inner kernel for
didgpu_cs()is disabled pending a matching rewrite. Its ATT agrees with the CPU path to ~4e-16, but it computes the OLD influence functions, and those now feed both the multiplier bootstrap and the aggregation SEs. After the CPU rewrite, per-cellmax |dSE|againstdid::att_gt()was 1.11e-16 for backends"r"and"cpu"but 1.96e-01 for"cuda". CS therefore routes through the validated CPU path, so all backends –"r","cpu"and"cuda"– now return identical numbers; verified at cellmax |dSE|= 1.11e-16 and aggregatemax |dSE|= 0 for every backend. Re-enable for kernel development withoptions(didgpu.cs_cuda_inner = TRUE).
17.5 New features
didgpu_cs()aggregations now carry standard errors and confidence intervals..cs_aggregate()previously propagated point estimates only, so$aggregationhad nosecolumn (while its empty-result stub declared one) anddidgpu_tidy()could only reportNA. Aggregate SEs are now derived from the influence functions the same waydid::aggte()derives them, including the correction for having ESTIMATED the aggregation weights (did:::wif) – without which SEs degrade badly at long event times, where few cohorts contribute (0.98x of correct at event 0, falling to 0.33x at event 9). Verified identical todid::aggte(type = "dynamic")for all three estimators, with and without covariates: max |diff| ~5e-16 on the aggregate estimate and ~3e-17 on its SE.Note that pre-treatment ATT(g, t) still differ from
didby construction: didgpu uses a universal base period,diddefaults to a varying one. Parity is asserted on post-treatment cells and non-negative event times.didgpu_fect(method = "ife", r = 0)now runs. An interactive-fixed-effects model with zero factors is the plain two-way FE model, and sweepingr = 0..kis the standard way to ask whether a result depends on the factor structure. It previously errored withrequires numeric/complex matrix/vector arguments:svd(M, nu = 0, nv = 0)omits theucomponent entirely (it is present only whennu > 0), so the truncated-SVD helper multipliedNULL.r = 0now returns an empty factor term and the fit reduces to two-way FE – verified identical tomethod = "fe".didgpu_tidy()accepts adidgpu_cs_result. The CS result class does not carry thedidgpu_resultparent, so tidy failed itsstopifnot()with the opaqueinherits(x, "didgpu_result") is not TRUE, even thoughdidgpu_loo()anddidgpu_honest_did()both document and accept that class. Tidy now emits one row peraggregationentry (kind = "aggregate") followed by one row per ATT(g, t) cell (kind = "att_gt")..cs_aggregate()propagates point estimates only, sostd.errorand the CI columns areNAon the aggregate rows rather than fabricated; the cell rows carry real SEs and CIs whenbootstrap_reps > 0. Passing anything else now names both accepted classes instead of printing a bareinherits()assertion.didgpu_cs()cluster bootstrap no longer crashes on panels whose columns collide with internal variable names..cs_bootstrap_se()held the panel as adata.tableand subset it withd[d[[args$group]] == u, ].[.data.tableevaluates itsiexpression with the table’s COLUMNS in scope, so a panel carrying a column literally nameddshadowed the localdwith the treatment VECTOR:d[[args$group]]becametreatment[["g"]]and failed withsubscript out of bounds. Sincedis an ordinary name for a treatment indicator, real panels hit this routinely, and because it fired only whenbootstrap_reps > 0it looked data-dependent rather than name-dependent – the same panel worked atreps = 0and crashed atreps > 0. didgpu’s own simulated panels never tripped it because their treatment column isD.The bootstrap now keeps the panel as a plain data.frame and resolves every column lookup before indexing, so no column name can shadow an internal. Point estimates and SEs are bit-identical on panels that previously worked. As a side effect the per-unit row lookup is precomputed once instead of rescanning the whole panel for every (replicate x pick), removing an O(B * n_units * nrow) cost.
didgpu_fect()no longer reports always-treated units’ levels as treatment effects. A unit treated in every observed period has no control cell, so its unit fixed effect (fe) / factor loading (ife,mc) is unidentified. The fitter set an unidentified unit effect to0, which made the imputed counterfactual the time effect alone – so the unit’s entire LEVEL landed in the residual and was reported as treatment effect. Always-treated units are selected on level (they are exactly the units already treated before the sample window opened), so the bias did not average out. On a known-zero DGP with 10 such units the reported ATE was +1.63 against a true effect of 0, and it survived every factor count andmethod = "fe". Units with no untreated period are now dropped before estimation with a warning, and the count is returned as$n_always_treated_dropped. This matches the referencefectpackage (“units whose number of untreated periods <1 are dropped automatically”) anddidgpu_bacon(), which already did this. On the same DGP all three methods now recover the null.Anyone who ran
didgpu_fect()on a panel containing always-treated units should re-run: the ATT was biased upward by their level.Note on
fectparity.didgpu_fect(method = "fe")agrees withfect::fect()definitionally but NOT bit-for-bit at the defaulttol = 1e-5: the alternating-projections fit stops early, leaving a gap of ~6e-5 on a 60-unit panel (~2e-6 attol = 1e-8, ~2e-8 attol = 1e-12). Pass a tightertolwhen exact agreement matters.fecthas been added toSuggestsso the comparison is now covered by the test suite.Reference parity restored on UNBALANCED panels. Baseline treatment
d_sqwas taken from the global first period rather than each group’s own first period with non-missing treatment. A group entering the panel late therefore hadd_sq = NA(the internal balancing merge creates the row but leaves treatment missing), which (a) madeF_gfall through toT_max + 1, reclassifying the group as a never-switcher, and (b) propagated the NA into the(time, d_sq)cohort-key encoding used by the fast backends, forcing a fallback branch that grouped cohorts differently from the reference. Backends"r","cpu"and CUDA consequently disagreed withDIDmultiplegtDYNon any panel with unequal group lengths – max |diff| 1.8e-01 on a 60-unit reprex – whilebackend = "reference"happened to agree. Baseline treatment now uses each group’s own first non-missing period, matching the reference definition exactly; the same reprex now agrees to 5.6e-17. Balanced-panel results are bit-identical to before.This went undetected by the randomized differential suite because
didgpu_simulate_panel()could only emit balanced panels, on which the two definitions coincide. The simulator gained alate_entry_fracargument (default0, RNG-stream preserving) andtests/testthat/test-unbalanced-parity.Rnow checks every CPU backend against the reference on an unbalanced panel.Users who ran didgpu on an unbalanced panel with any backend other than
"reference"should re-run: point estimates were affected.Bootstrap SEs survive partial-NA iterations. A kept bootstrap iteration can carry NA at some horizons (its resample has switchers overall but none reaching horizon j). Plain
sd()/cov()then returned NA — with few switchers this silently wiped out EVERY SE, CI and joint p-value while the point estimates looked fine. Per-horizon SEs are now computed from the finite draws with a minimum bootstrap support of 30 finite draws per horizon (below that the SE stays NA, honestly), and the stored$coef$vcovuses a pairwise-complete covariance.Joint tests: near-singular covariance now warns. With many horizons and few switchers, the bootstrap covariance behind the omnibus joint-effects/placebo chi-square can be near-singular; the statistic is then numerically unstable (tiny eigenvalues amplify arbitrary linear combinations).
.joint_pvalue()now excludes horizons with inadequate bootstrap support and emits a warning whenrcond(V) < 1e-10, advising a low-dimensional prespecified test (e.g. the leads nearest treatment) instead of the omnibus p.CUDA: unbalanced panels with late-entrant groups no longer crash with error 700. Groups unobserved at the global first period have NA baseline treatment (
d_sq); on the CPU path they fall out of every cohort mask and contribute zero, but on the CUDA path the NA flowed throughmatch()intocohort_keyasNA_integer_, which reaches the kernels asINT_MINand caused an illegal memory access (CUDA DID kernel failed with code 700) ink_finalize_dist_and_kernel— and error 700 poisons the CUDA context, so every subsequent didgpu call in the process failed too. Any real-world unbalanced panel (units entering the sample over time) hit this immediately. NA-key rows are now parked in a padding cohort whose kernel contribution is identically zero, matching CPU semantics bit-for-bit. Diagnosed withcompute-sanitizer(invalid 8-byte global read atbase + INT_MIN * 8); regression-tested intest-cuda-late-entrants.R, including a context-not-poisoned check.Bootstrap aggregation no longer crashes on degenerate resamples. Under sparse-switching treatments (few clean switchers), a bootstrap resample can contain no valid switcher cell at some horizon, yielding a zero-length effects vector.
.aggregate_to_result()’svapply()calls hard-required full-length vectors, so a single such iteration aborted the entire estimation withvalues must be length K ... result is length 0— and because the failure probability grows withbootstrap_reps, exactly the large-rep runs users want for final inference were the ones crashing. Degenerate iterations are now dropped with a warning that reports the count; SEs and the bootstrap covariance use the surviving iterations (standard failed-resample practice). Found while re-estimating thin subsamples at 2,000 reps; e.g. a 78-DAO/4-switcher spec had 103/2,000 degenerate resamples.
17.6 CRAN resubmission fixes
test-parallel.Rnow skips on CRAN. The previous submission’s pretest exceeded CRAN’s 2-core cap (_R_CHECK_LIMIT_CORES_) despite the test honouring the env var, producing the only test failure. Bothtest_thatblocks intest-parallel.Rnow callskip_on_cran(). The parallel path is fully exercised in our GitHub Actions CI matrix on every push.- README links rewritten to absolute GitHub URLs. Two references (
WINDOWS_BUILD_STATUS.md,BENCHMARKS.md) are intentionally excluded from the source tarball via.Rbuildignore; the README now points to their canonical GitHub URLs so they resolve from the rendered README on CRAN. One stale link (../NOTES_did_gpu_checkpointed.Rmd, an out-of-tree file that no longer exists) was removed. - DESCRIPTION typography. Single-quoted the software-name references
'Rcpp'and'CUDA'per CRAN convention.
18 didgpu 0.1.1
18.1 Bug fixes
- CUDA: support Blackwell (
sm_120) GPUs. The GPU build previously targeted only Turing–Hopper (sm_75–sm_90) with no PTX fallback, so on Blackwell parts (RTX PRO Blackwell, RTX 50xx) the kernels had no device image and effect estimates silently came back as all zeros while the CPU/R backends were correct. Addedcompute_120,sm_120plus acompute_120PTX target to bothsrc/Makevarsandsrc/Makevars.win. GPU effects are again bit-identical to the CPU/R backends on Blackwell (verified on an RTX PRO 5000 Blackwell, CUDA 13.2). - Windows build: ship
src/didgpu_cuda.def. The DLL export list was git-ignored (/src/*.def), so clean checkouts — and the r-universe /install_githubsource build — could not linkdidgpu_cuda.dll. It is now tracked. - Windows build: locate CUDA runtime DLLs under
bin/x64. CUDA 13 moved the redistributable DLLs frombin/tobin/x64/; the bundling step now searches both sodidgpu_cuda.dllloads at runtime.
18.2 Reference parity with DIDmultiplegtDYN 2.3.x
didgpu was originally validated bit-for-bit against an older DIDmultiplegtDYN. Two of its outputs were deliberately changed upstream; didgpu now tracks the current (fixed) behavior:
predict_hetstandard errors now use HC2. The reference switched the heterogeneity-regression variance from HC1 to HC2 (sandwich::vcovHC(type = "HC2")) in v2.3.1 (“explicit CI formulas” fix). didgpu now does the same (newsandwichdependency), so the predict_hetSE/t/LB/UB/pFcolumns match again.- Placebo
Ncounts each contributing cell once. For bidirectional panels the reported placebo sample size was double-counting controls shared between the switcher-in and switcher-out comparisons. It now uses the in-direction count, matching the reference’s per-rowcoalesce(in, out)combiner across bothin>outandout>inpanels. Point estimates were never affected.
18.3 Correctness fixes (found by randomized differential testing vs the reference)
same_switchers: the placebo now uses the same restricted switcher set as the effects. Undersame_switchers = TRUEthe placebo block was computed on the full switcher set rather than the consistent-switchers subset, producing a biased placebo estimate (and inflated placeboN) relative todid_multiplegt_dyn. The placebo distance now honours the effects-basedstill_switcherrestriction (the reference derives the placebo distribution from the same_switchers-gated effect distance), so placebo estimates and counts match again.trends_lin: no longer crashes on panels with zero estimable effects. On short panels where no group has theF_g-2pre-period thattrends_linrequires, result aggregation crashed with “length of ‘dimnames’ [1] not equal to array extent” (an unguarded row-name build on a 0-row table). It now returns an empty, no-estimable- effects result cleanly.- Weighted
N/Switcherscolumns now separate unweighted counts from weighted sums. On weighted panels (weight =) the four reported count columns were all populated from the same (weighted, truncated) switcher mass, soN/Switchersreported weighted sums instead of observation counts, andN.wlost the fractional weight (each cell’s weight was floored before summing). The estimator now reportsNandSwitchersas the true unweighted observation / switcher counts andN.w/Switchers.was the (unfloored) weighted sums, matchingdid_multiplegt_dynexactly. Point estimates were never affected — the weighted switcher mass that drives the Neyman pooling and ATE weights is unchanged; only the reported count columns moved. On unweighted panels all four columns coincide as before (bit-identical output). - Weighted estimates: Neyman direction-pooling no longer truncates the switcher mass. On weighted panels with switchers in both directions, the per-direction weighted switcher mass was floored to an integer before being used as the Neyman pooling weight (
w_in = N_in / (N_in + N_out)) and as the across-horizon ATE weight, biasing the pooled event-study estimates by ~1e-3. The exact (unfloored) mass is now used throughout the estimate path, matchingdid_multiplegt_dyn. Single-direction (switchers = "in"/"out") and unweighted panels were never affected (the weight is 0/1 or the mass is already integer). Found by randomized weighted×flag differential testing. - Reported effect/placebo count: drop trailing unestimable horizons. didgpu’s horizon clamp uses each group’s own data availability (
max L_g), which can be one step more permissive thandid_multiplegt_dyn’s cohort-levelT_gclamp. When that extra horizon has no switcher reaching it (NA estimate, zero switchers), the reference omits the row; didgpu now trims the trailing block of such unestimable effect/placebo rows so the reported horizon count matches. Estimable horizons, estimates, and all four count columns are unchanged. Verified across 359 boundary-stress panels (switchers in/out/both × effects 4–5 × weighted/unweighted): zero horizon count mismatches and, importantly, no case where didgpu reported an extra horizon with positive switchers. Found by overnight differential testing. - Placebos: drop the whole block when every horizon is unestimable. On weighted
trends_lin(and other short-panel cases) the placebo block can be entirely unestimable — no group has theF_g - q - 1pre-period any placebo horizon needs.did_multiplegt_dynreturnsNULLplacebos in that case; didgpu used to emit a placebo matrix of NAs. Aggregation now drops the placebo block when every placebo point estimate is NA, matching the reference. Found by extended weighted×flag differential testing (4 of 4 affected panels now match; all other scenarios untouched — verified 30 seeds × 14 scenarios = 420 weighted comparisons, 0 fails).
19 didgpu 0.1.0
First public release. Five estimator families plus a sensitivity layer, each with CUDA kernels for the hot paths.
19.1 Callaway-Sant’Anna (2021) — didgpu_cs()
est_method = c("OR", "IPW", "DR")— all three inner estimators. DR is Sant’Anna-Zhao (2020) doubly-robust.control_group = c("never", "notyet")— never-treated OR not-yet-treated.covariates =for OR/IPW/DR adjustment.- Pre-treatment placebos computed automatically; joint chi-square test on the placebo block via
fit$placebo. - Four aggregations (
event/group/calendar/overall) withdidgpu_cs_aggregate()for switching post-fit. bootstrap_kind = c("cluster", "multiplier")— cluster bootstrap on units, or multiplier wild bootstrap on per-unit influence functions (much faster for large B).- CUDA: all three inner regressions (OR / IPW / DR) run on the GPU (
src/cuda_cs_inner.cu) with per-row influence functions. OR uses an in-thread Cholesky per cell; IPW/DR add a per-cell IRLS logistic propensity model replicatingstats::glm.fit(ATT agrees with R to ~1e-8), and DR layers on the outcome-regression augmentation. With per-cell IFs, the cluster + multiplier bootstrap SEs all run on the GPU — the DR cluster bootstrap is ~192x faster than R. - Cross-validated against the reference
didpackage on simulated panels (max abs diff < 0.25 on event-study estimates).
19.2 TestMechs (Kwon & Roth 2026) — didgpu_test_sharp_null()
- All three test methods:
"CS"(Cox-Shi 2023),"ARP"(Andrews-Roth-Pakes 2023),"FSST"(Fang-Santos-Shaikh-Torgovitsky 2023). - Both binary mediator (K = 2) and multi-level (K >= 2) under no-defiers.
- Generic CS engine
.testmechs_cs_test(theta_hat, Sigma, A, A_eq, b_eq)is reusable for any moment-inequality test. - Nonparametric and Bayesian (Dirichlet) bootstrap of the partial-density vector beta.obs.
- CUDA bootstrap kernel
src/cuda_testmechs_bootstrap.cu(cuRAND multinomial; the main acceleration target). Live on Linux/WSL and wired through.testmechs_bootstrap_cuda; the “nonparametric” method runs on the GPU, “bayes” uses the R path. cuRAND vs R’s MT19937 differ per-replicate, so bootstrap moments match within Monte-Carlo error rather than bit-for-bit.
19.3 Leave-one-out robustness — didgpu_loo()
- Drops one entity at a time (cohort / unit / cluster / arbitrary column level) and re-fits the estimator.
- Works on all three estimator families:
didgpu_result,didgpu_cs_result,didgpu_fect_result. - Returns a
didgpu_loo_resultdata.frame withleave_out,estimate,delta, anddelta_pctcolumns, sorted byabs(delta)descending so the most-influential drop is on top. print()shows the top-N most influential rows with an interpretation hint;plot()draws a tornado plot of deltas.- Default
by = "cohort"(leave-one-cohort-out, the standard DiD diagnostic). Passby = "unit","cluster", or any column name to drop on a different key.
19.4 HonestDiD (Rambachan & Roth 2023) — didgpu_honest_did()
- Sensitivity analysis on event-study DiD estimates.
method = c("RM", "M")— relative-magnitudes OR smoothness bounds.- Reports the breakdown parameter (smallest Mbar at which the CI includes zero) so users can read off how robust their conclusion is.
- Works on both
didgpu_result(DIDmultiplegtDYN-style) anddidgpu_cs_result(Callaway-Sant’Anna) fits.
19.5 fect family (counterfactual-prediction estimators) — didgpu_fect()
19.6 Estimator (bit-identical to DIDmultiplegtDYN::did_multiplegt_dyn)
- Binary, multivalued, and continuous treatment.
effects,placebo,switchers = ""/"in"/"out", ATE.weight,controls,trends_nonparamcohort extension.only_never_switchers,same_switchers,dont_drop_larger_lower.normalized = TRUE(per-unit-of-treatment),trends_lin = TRUE(linear cohort trends with cumulative-recovery placebos).same_switchers_pl(placebo-side same-switchers gate; mirrors the reference’s constraint that it must be paired withsame_switchers).predict_het(heterogeneity regression with HC1 robust SEs and joint F-test).didgpu_by_path()for treatment-trajectory subgroup analysis (the equivalent of the reference’sby_pathargument).- Sample-size columns (
N,Switchers,N.w,Switchers.w) match the reference exactly. didgpu_equivalence(fit, delta)— pre-trends equivalence (TOST) test on the placebo estimates. Instead of “failed to reject a zero pre-trend” (weak, and worst exactly when underpowered), it tests H0: |pre-trend| >= delta and REJECTING is positive evidence the pre-trend is within +/- delta. Reports per-horizon and joint (intersection-union) verdicts plus the smallest defensible margin (breakdown_delta). Mirrorsdidgpu_fect_equivalence().didgpu_joint_placebo(fit, horizons)— the joint chi-square placebo test (p_jointplacebo) restricted to a chosen pre-treatment window, reusing the stored bootstrap covariance. Test parallel trends only over the leads you care about; the full-window call reproduces the headlinep_jointplaceboexactly.didgpu_bacon()— Goodman-Bacon (2021) decomposition of the static TWFE DiD into its 2x2 timing-group comparisons, with the total weight on “forbidden” already-treated-control comparisons as the bias diagnostic. Validated by the exact identity (weighted 2x2 sum == the TWFE coefficient fromdidgpu_twfe()). Balanced, binary, absorbing panels.didgpu_did_static()— de Chaisemartin & D’Haultfoeuille (2020) DID_M instantaneous estimator. Unlike the staggered-adoption methods it allows treatment to turn on AND off (non-absorbing): it compares each switcher’s period-over-period outcome change to same-baseline stayers and averages over all switch events, with a cluster bootstrap SE. Native reimplementation; cross-checked againstDIDmultiplegt::did_multiplegt.didgpu_freyaldenhoven()— Freyaldenhoven, Hansen & Shapiro (2019) pre-event panel event study.estimator = "OLS"is the two-way FE event study;estimator = "FHS"adds an auxiliary proxy covariate as an endogenous regressor and 2SLS-instruments it with a far policy lead to purge a confound that generates pre-trends. Native reimplementation of the first-difference parameterization; coefficients matcheventstudyr::EventStudy(OLS and FHS) to machine precision.didgpu_cs_continuous()— Callaway, Goodman-Bacon & Sant’Anna (2024) difference-in-differences with a CONTINUOUS treatment (dose). Estimates the dose-response curve: the level effect ATT(d) and the causal response ACRT(d) = ATT’(d), via a B-spline regression of the within-unit outcome change on the dose, vs a never-treated comparison; multiplier-bootstrap SEs. Native reimplementation (spline basis via splines2); ATT(d)/ACRT(d) matchcontdid::cont_didexactly.didgpu_did_continuous()— de Chaisemartin & D’Haultfoeuille (2024) continuous treatment with NO STAYERS. When the dose changes for (almost) every unit there is no pure control group, so identification is in first differences: with dY, dD the within-unit changes, the common trend E[dY|dD=0] is recovered from quasi-stayers (dD near 0), giving the level effect effect(d) = E[dY|dD=d] - E[dY|dD=0] and the average causal response ACR(d).estimator = "parametric"fits a polynomial in dD (sqrt(n));estimator = "nonparametric"is a local-linear (kernel) fit (n^2/5; flagged EXPERIMENTAL — no maintained R reference exists to bit-validate it). Multiplier-bootstrap SEs. Both estimators validated by simulation against a known dose-response.- All eight auxiliary estimators above were benchmarked to verify they belong on the CPU (none has a GPU-amenable hot path; see
BENCHMARKS.md). The audit also caught and fixed two quadratic bootstraps:didgpu_did_static’s cluster bootstrap was O(n_units^2) per replicate (pre-splitting by cluster makes it O(n); 12-23x faster, bit-identical SEs), anddidgpu_did_continuousno longer recomputes the O(n^2) overall-ACR on every bootstrap replicate (nonparametric bootstrap ~675x faster; reported effect(d)/ACR(d) unchanged).
19.7 Long-running workflow
- Per-cell checkpointing to disk with atomic writes (
saveRDStmp + rename) and append-onlymanifest.csv. Resumable on crash. didgpu_resume(checkpoint_dir, df, ...)— re-invokes with every stored arg restored frommeta.json.didgpu_bootstrap_more(checkpoint_dir, df, extra_reps)— extend a finished run with more bootstrap reps without rework.didgpu_by(df, by_var, ...)— per-subgroup fits, each with its own checkpoint subdirectory.n_workers > 1Lparallelises the bootstrap loop viaparallel::makeCluster; bit-identical to sequential at the same seed.
19.8 Backends
"r"— pure R viadata.table. 60× faster than the reference at 200 K rows."reference"— delegate toDIDmultiplegtDYN::did_multiplegt_dyn, used as the parity oracle."cuda"— live on Windows and Linux/WSL2 (built + verified end-to-end on an NVIDIA RTX 4000 Ada, CUDA 12.6; bit-identical results on both). On Windows it needs no admin rights: a user-local CUDA toolkit plus a two-DLL split (didgpu_cuda.dllbuilt by nvcc/MSVC, the R-facingdidgpu.dllbuilt by Rtools/MinGW, bridged by a pure-C ABI) sidesteps the MinGW↔︎MSVC link barrier — seeWINDOWS_BUILD_STATUS.md. Live GPU paths: the CS cluster bootstrap (179–228× faster than R via the influence-function shortcut), the CS multiplier bootstrap, the CS OR point estimate (bit-exact vs R), and the TestMechs nonparametric bootstrap. The fect SVD path is size-gated — it only engages for very large balanced panels, since cuSOLVER loses to CPU LAPACK on the small matrices typical of fect. Every GPU path falls back transparently to R when CUDA is unavailable or would be slower, sobackend = "cuda"is always safe. SeeBENCHMARKS.mdandtests/testthat/test-cuda-equivalence-grid.R(142 lock-step assertions). Tests skip GPU paths whennvcc/ a device is absent."cpu"— Rcpp+Eigen, scaffolded only.
19.9 R interface
- S3 methods:
print,summary,coef,confint,vcov,plot, plustidy,glance,augmentviabroom. - Plot is base-R (no
ggplot2dependency); event-study with stored CIs. - Diagnostic helpers:
didgpu_summarize_panel,didgpu_estimate_runtime,didgpu_compare(compare against the reference).
19.10 fect family (counterfactual-prediction estimators)
didgpu_fect(method = "fe")— two-way fixed effects, iterative demeaning of the controls-only outcome matrix.didgpu_fect(method = "ife")— Bai (2009) interactive fixed effects. Alternating fe-step + rank-r SVD of the residual matrix until convergence.didgpu_fect(method = "mc")— Athey et al. (2021) matrix completion. Iterative soft-thresholded SVD on the controls-only matrix.- All three reuse
didgpu()’s checkpoint / resume / parallel bootstrap infrastructure. Results are returned asdidgpu_fect_result(extendsdidgpu_result) so all the standard accessors (coef,confint,vcov,plot,tidy,glance) work the same way. - CUDA kernels for fect live in
src/cuda_fect_fe.cuandsrc/cuda_fect_svd.cu(the latter uses cuSOLVER’scusolverDnDgesvdjfor the SVD primitive shared byifeandmc) and are wired through R. However, they are size-gated: on the small, tall-skinny matrices typical of fect panels the per-iteration cuSOLVER SVD is 100–300× slower than R’s LAPACK (cuSOLVER handle + H2D/D2H overhead dwarfs the tiny SVD)..fect_cuda_svd_worthwhile()only routes to the GPU for very large balanced panels (n_units ≥ 2000andn_units·n_periods ≥ 2e5); below thatbackend = "cuda"transparently uses R’ssvd(). SeeBENCHMARKS.md.
19.11 Testing
- 300+ tests across 24 test files;
R CMD checkpasses with only pre-existing intentional WARN (CUDA.cufiles insrc/) and the declared GNU makeSystemRequirementsNOTE. - Adversarial fuzz harness (
tests/testthat/test-fuzz.R) covers 21 scenarios: vanilla / weight / controls / switchers / normalized / placebos / trends_lin / only_never + same_switchers / kitchen sink / trends_lin sink / multivalued / bootstrap-stability / parallel-equals-sequential / checkpoint round-trip / degenerate / very-small. DefaultDIDGPU_FUZZ_N = 8for fast CI; bump via env var for deep local runs (validated at N = 200, no failures).
19.12 Validation
- Monte Carlo CI coverage by switcher count (
tools/validation/): with the cluster bootstrap, 95% CIs achieve ~94% coverage at 8+ switchers, ~87% at 4, but only ~66% at 2 — the classic few-treated-clusters failure, not specific to didgpu. Point estimates are unbiased at every count. Practical rule: treat estimates identified off fewer than ~5 clean switchers as diagnostics; their bootstrap CIs materially undercover. The degenerate-resample drop does not distort coverage where support is adequate (8-13 switchers: nominal).