CausalSmith · seminar slides
Minimax Estimation of Optimal Treatment Values
With n observations and d discrete covariate cells, an explicit estimator attains the finite-sample minimax squared-risk order min{1,d/[nlog(ed)]} under fixed overlap.
slides for Minimax Estimation of Optimal Treatment Values with Discrete Covariates
Overview
For ϵ, the fixed overlap parameter, the minimax squared risk has order min{1,d/[nlog(ed)]}.
- The characterization covers binary treatments and outcomes, arbitrary cell masses, unequal propensities, and exact treatment-effect ties.
- A Jackson factorial estimator attains the upper bound.
- An embedded two-sample L1 problem supplies the matching converse.
- The same frontier applies to the identified causal optimal-treatment value.
Motivation
Consider a clinic choosing between two treatments within each of d diagnostic risk groups.
- The outcome is binary recovery.
- Each patient receives the treatment with the higher conditional recovery probability in that group.
- The target is the population recovery rate attained by these group-specific choices.
- A rich classification system creates many rare groups, even when the total sample is large.
- How accurately can we estimate the value of the optimal treatment choices?
Research Question
The value takes the larger of two estimated outcome means in every cell.
- At a treatment-effect tie, the maximum has an absolute-value cusp.
- Sampling noise then creates upward bias in the cellwise plug-in maximum.
- With many rare cells, these local errors accumulate across the alphabet.
- Fixed overlap supplies observations from both arms but leaves the tie-induced difficulty intact.
- What is the best possible worst-case squared error?
Setup
We observe independent triples Oi=(Xi,Ai,Yi).
- Xi is one of d covariate categories.
- Ai and Yi are binary treatment and outcome indicators.
- Cell x has mass px, treatment propensity πx, and arm means μ0x,μ1x.
- The target Ψ(P), the observed-law optimal value, averages maxaμax using the cell masses.
- Risk is worst-case mean-squared error over all observed laws satisfying fixed overlap.
Assumptions
- The n observed triples are independent and identically distributed.
- Consistency links the observed outcome to Y(a), the potential outcome under the received arm.
- Conditional exchangeability makes treatment independent of each potential outcome given the covariate cell.
- Fixed overlap keeps both treatment arms represented in every occupied cell.
Let [d]={1,…,d} be the finite covariate alphabet. For every observed law P under consideration and every x∈[d] with covariate-cell mass px=P(X=x)>0, the treatment propensity πx=P(A=1∣X=x) satisfies ϵ≤πx≤1−ϵ, where ϵ∈(0,1/2) is the fixed overlap constant.
Identification
informal · Theorem T-1 For d≥2 under fixed overlap, the observed optimal value is a sum of globally Lipschitz cell contributions, and every admissible observed law has a consistent, conditionally exchangeable causal completion.
informal · Theorem T-8 Under consistency, conditional exchangeability, and fixed overlap, the causal optimal-treatment value equals the observed-law optimal-regression value.
For the clinic, each cell contribution is its population share times the better arm’s recovery probability.
Related Literature
- Robins (1986) and Rosenbaum and Rubin (1983) provide the causal identification foundations.
- Manski (2004), Murphy (2003), Kitagawa and Tetenov (2018), and Athey and Wager (2021) study treatment rules, welfare, and policy learning.
- Hirano and Porter (2012) and Luedtke and van der Laan (2016) analyze nonregular optimal-value inference.
- Cai and Low (2011), Jiao et al. (2015), and Jiao et al. (2018) develop approximation and moment-matching methods for nonsmooth functionals.
- Jiao et al. (2018) settle the two-sample L1 distance with matching d/[nlog(en)] bounds; we transfer their converse through an equal-propensity embedding and build the estimator side here.
- Zeng et al. (2024) give d2/n2+1/n upper and d2/[n2log2n]+1/n lower rates for a linear treatment mean over a discrete alphabet; the cellwise maximum changes the frontier to min{1,d/[nlog(ed)]}.
- We connect these strands by characterizing the finite-sample minimax risk of the scalar optimal value over a growing categorical alphabet.
Key Idea
The direct plug-in estimator becomes unreliable when many cells are sparse and their treatment means are nearly tied.
- Pilot counts locate each cell’s four treatment–outcome probabilities at their own noise scale.
- A local Jackson polynomial replaces the maximum’s cusp by a degree proportional to log(ed).
- Centered factorial moments estimate the polynomial’s monomials without plug-in ratio bias.
- Clipping controls rare pilot failures and unstable high-degree terms.
- The logarithmic polynomial degree reduces the accumulated nonsmooth error to the d/[nlog(ed)] scale.
Estimator
- Small alphabets use the empirical-ratio estimator.
- Pilot counts determine a local rectangle for each cell’s four atom probabilities.
- Evaluation counts lift the local polynomial through centered factorial moments.
- Cell estimates are clipped, summed, and projected to the value range.
- Rao–Blackwellization averages over the auxiliary split and returns an estimator based on all n observations.
Upper Bound
informal · Theorem T-2 Under i.i.d. sampling and fixed overlap, the Jackson factorial estimator has squared risk at most Cϵmin{1,d/[nlog(ed)]} uniformly over every n≥1, d≥2, and admissible observed law.
The guarantee includes unknown unequal propensities, rare or null cells, boundary outcome means, and exact ties.
In the clinic example, the same estimator covers highly imbalanced risk-group frequencies as long as both treatments retain fixed overlap within occupied groups.
Lower Bound
informal · Theorem T-3 Under i.i.d. sampling and fixed overlap, an equal-propensity submodel makes the optimal value equal to one half plus one quarter of a two-sample L1 distance and transfers at least one sixteenth of its minimax squared risk.
informal · Theorem T-4 Under i.i.d. sampling and fixed overlap, every estimator has worst-case squared error at least cϵmin{1,d/[nlog(ed)]}.
How can two observational worlds have separated optimal values while remaining statistically difficult to distinguish?
Proof Sketch
- Construct equal-mass cells with propensity 1/2 and small treatment-effect contrasts around exact ties.
- Choose two symmetric contrast priors whose moments agree through logarithmic degree but whose average absolute contrasts differ.
- Poissonized cell counts have nearly identical mixtures because their likelihood expansions agree through the matched moments.
- The optimal values remain separated because the cellwise maximum depends on the absolute contrast.
- A testing argument converts this separation into the d/[nlog(ed)] lower bound, with a constant bound in the saturated regime.
Main Result
The explicit estimator’s upper bound and the all-estimator converse meet over the full fixed-overlap class.
There are universal tuning constants H0>0, 0<κ<1, and an integer cutoff D0>2 for the Jackson factorial estimator such that the following holds. For every fixed overlap constant ϵ∈(0,1/2), there are real constants cϵ and Cϵ with 0<cϵ≤Cϵ. For every n≥1 and d≥2, with Ld=log(ed) and with Rn,d,ϵ denoting the fixed-sample observed minimax squared risk in Definition P-12, cϵmin{1,nLdd}≤Rn,d,ϵ≤P∈Dd,ϵobssupEP⊗n[(Vn,dJF−Ψ(P))2]≤Cϵmin{1,nLdd}. Moreover, the fixed-sample minimax squared risk over the full-data causal completion class Pd,ϵ is exactly Rn,d,ϵ for the same n, d, and ϵ.
Thus the observed-law and causal-completion formulations have exactly the same finite-sample minimax order.
Implications
Let dn denote the covariate-alphabet size at sample size n.
informal · Theorem T-7 For fixed overlap, both observed and causal minimax risks converge to zero exactly when dn=o{nlog(en)}, and they have the parametric n−1 order exactly when dn=O(1).
- Bounded alphabets recover the usual parametric squared-risk scale.
- Growing alphabets remain uniformly learnable throughout the stated consistency region.
- In the clinic example, increasingly refined risk groups preserve uniform consistency when their number satisfies dn=o{nlog(en)}.
Additional Result
informal · Theorem T-6 With universal estimator tuning, the predecessor risk bracket remains valid, the Jackson factorial estimator equals the empirical-ratio estimator for 2≤d<D0, and bounded-alphabet sequences attain the n−1 minimax scale.
This comparison anchors the construction to the familiar cellwise estimator at small d.
Conclusion
- The optimal-treatment value is a nonsmooth large-alphabet functional generated by cellwise treatment-effect ties.
- Its finite-sample minimax squared risk is of order min{1,d/[nlog(ed)]} under fixed overlap.
- The Jackson factorial estimator attains this order using local approximation and unbiased polynomial lifting.
- Equal-propensity L1 embeddings and moment matching establish optimality over all estimators.
- The result yields exact consistency and parametric-rate boundaries for both observed and causal formulations.
- The rate characterization is uniform over all n≥1 and d≥2, with constants depending on fixed overlap.