CausalSmith · seminar slides

Minimax Estimation of Optimal Treatment Values

With nn observations and dd discrete covariate cells, an explicit estimator attains the finite-sample minimax squared-risk order min{1,d/[nlog(ed)]}\min\{1,d/[n\log(ed)]\} under fixed overlap.

Overview

For ϵ\epsilon, the fixed overlap parameter, the minimax squared risk has order min{1,d/[nlog(ed)]}\min\{1,d/[n\log(ed)]\}.

  • The characterization covers binary treatments and outcomes, arbitrary cell masses, unequal propensities, and exact treatment-effect ties.
  • A Jackson factorial estimator attains the upper bound.
  • An embedded two-sample L1L_1 problem supplies the matching converse.
  • The same frontier applies to the identified causal optimal-treatment value.

Motivation

Consider a clinic choosing between two treatments within each of dd diagnostic risk groups.

  • The outcome is binary recovery.
  • Each patient receives the treatment with the higher conditional recovery probability in that group.
  • The target is the population recovery rate attained by these group-specific choices.
  • A rich classification system creates many rare groups, even when the total sample is large.
  • How accurately can we estimate the value of the optimal treatment choices?

Research Question

The value takes the larger of two estimated outcome means in every cell.

  • At a treatment-effect tie, the maximum has an absolute-value cusp.
  • Sampling noise then creates upward bias in the cellwise plug-in maximum.
  • With many rare cells, these local errors accumulate across the alphabet.
  • Fixed overlap supplies observations from both arms but leaves the tie-induced difficulty intact.
  • What is the best possible worst-case squared error?

Setup

We observe independent triples Oi=(Xi,Ai,Yi)O_i=(X_i,A_i,Y_i).

  • XiX_i is one of dd covariate categories.
  • AiA_i and YiY_i are binary treatment and outcome indicators.
  • Cell xx has mass pxp_x, treatment propensity πx\pi_x, and arm means μ0x,μ1x\mu_{0x},\mu_{1x}.
  • The target Ψ(P)\Psi(\mathbb P), the observed-law optimal value, averages maxaμax\max_a\mu_{ax} using the cell masses.
  • Risk is worst-case mean-squared error over all observed laws satisfying fixed overlap.

Assumptions

  • The nn observed triples are independent and identically distributed.
  • Consistency links the observed outcome to Y(a)Y(a), the potential outcome under the received arm.
  • Conditional exchangeability makes treatment independent of each potential outcome given the covariate cell.
  • Fixed overlap keeps both treatment arms represented in every occupied cell.
Assumption A-4 (Fixed overlap)

Let [d]={1,,d}[d]=\{1,\ldots,d\} be the finite covariate alphabet. For every observed law P\mathbb P under consideration and every x[d]x\in[d] with covariate-cell mass px=P(X=x)>0p_x=\mathbb P(X=x)>0, the treatment propensity πx=P(A=1X=x)\pi_x=\mathbb P(A=1\mid X=x) satisfies ϵπx1ϵ, \epsilon\leq \pi_x \leq 1-\epsilon , where ϵ(0,1/2)\epsilon\in(0,1/2) is the fixed overlap constant.

Identification

informal · Theorem T-1 For d2d\ge2 under fixed overlap, the observed optimal value is a sum of globally Lipschitz cell contributions, and every admissible observed law has a consistent, conditionally exchangeable causal completion.

informal · Theorem T-8 Under consistency, conditional exchangeability, and fixed overlap, the causal optimal-treatment value equals the observed-law optimal-regression value.

For the clinic, each cell contribution is its population share times the better arm’s recovery probability.

Related Literature

  • Robins (1986) and Rosenbaum and Rubin (1983) provide the causal identification foundations.
  • Manski (2004), Murphy (2003), Kitagawa and Tetenov (2018), and Athey and Wager (2021) study treatment rules, welfare, and policy learning.
  • Hirano and Porter (2012) and Luedtke and van der Laan (2016) analyze nonregular optimal-value inference.
  • Cai and Low (2011), Jiao et al. (2015), and Jiao et al. (2018) develop approximation and moment-matching methods for nonsmooth functionals.
  • Jiao et al. (2018) settle the two-sample L1L_1 distance with matching d/[nlog(en)]d/[n\log(en)] bounds; we transfer their converse through an equal-propensity embedding and build the estimator side here.
  • Zeng et al. (2024) give d2/n2+1/nd^2/n^2+1/n upper and d2/[n2log2n]+1/nd^2/[n^2\log^2 n]+1/n lower rates for a linear treatment mean over a discrete alphabet; the cellwise maximum changes the frontier to min{1,d/[nlog(ed)]}\min\{1,d/[n\log(ed)]\}.
  • We connect these strands by characterizing the finite-sample minimax risk of the scalar optimal value over a growing categorical alphabet.

Key Idea

The direct plug-in estimator becomes unreliable when many cells are sparse and their treatment means are nearly tied.

  • Pilot counts locate each cell’s four treatment–outcome probabilities at their own noise scale.
  • A local Jackson polynomial replaces the maximum’s cusp by a degree proportional to log(ed)\log(ed).
  • Centered factorial moments estimate the polynomial’s monomials without plug-in ratio bias.
  • Clipping controls rare pilot failures and unstable high-degree terms.
  • The logarithmic polynomial degree reduces the accumulated nonsmooth error to the d/[nlog(ed)]d/[n\log(ed)] scale.

Estimator

Observed treatment–outcome counts Independent counts pilot and evaluation Pilot-local rectangles Jackson polynomial Factorial-moment cell estimates Clipped aggregate summed and projected All-data estimator
illustrative Boxes labeled “Observed treatment–outcome counts,” “Independent pilot and evaluation counts,” “Pilot-local rectangles,” “Jackson polynomial,” “Factorial-moment cell estimates,” “Clipped aggregate,” and “All-data estimator,” connected by left-to-right arrows.
  • Small alphabets use the empirical-ratio estimator.
  • Pilot counts determine a local rectangle for each cell’s four atom probabilities.
  • Evaluation counts lift the local polynomial through centered factorial moments.
  • Cell estimates are clipped, summed, and projected to the value range.
  • Rao–Blackwellization averages over the auxiliary split and returns an estimator based on all nn observations.

Upper Bound

informal · Theorem T-2 Under i.i.d. sampling and fixed overlap, the Jackson factorial estimator has squared risk at most Cϵmin{1,d/[nlog(ed)]}C_\epsilon\min\{1,d/[n\log(ed)]\} uniformly over every n1n\ge1, d2d\ge2, and admissible observed law.

The guarantee includes unknown unequal propensities, rare or null cells, boundary outcome means, and exact ties.

In the clinic example, the same estimator covers highly imbalanced risk-group frequencies as long as both treatments retain fixed overlap within occupied groups.

Lower Bound

informal · Theorem T-3 Under i.i.d. sampling and fixed overlap, an equal-propensity submodel makes the optimal value equal to one half plus one quarter of a two-sample L1L_1 distance and transfers at least one sixteenth of its minimax squared risk.

Two unknown categorical distributions Equal-propensity treatment table Optimal value containing L₁ distance Minimax lower bound
illustrative Boxes labeled “Two unknown categorical distributions,” “Equal-propensity treatment table,” “Optimal value containing L1L_1 distance,” and “Minimax lower bound,” connected by left-to-right arrows.

informal · Theorem T-4 Under i.i.d. sampling and fixed overlap, every estimator has worst-case squared error at least cϵmin{1,d/[nlog(ed)]}c_\epsilon\min\{1,d/[n\log(ed)]\}.

How can two observational worlds have separated optimal values while remaining statistically difficult to distinguish?

Proof Sketch

  • Construct equal-mass cells with propensity 1/21/2 and small treatment-effect contrasts around exact ties.
  • Choose two symmetric contrast priors whose moments agree through logarithmic degree but whose average absolute contrasts differ.
  • Poissonized cell counts have nearly identical mixtures because their likelihood expansions agree through the matched moments.
  • The optimal values remain separated because the cellwise maximum depends on the absolute contrast.
  • A testing argument converts this separation into the d/[nlog(ed)]d/[n\log(ed)] lower bound, with a constant bound in the saturated regime.

Main Result

The explicit estimator’s upper bound and the all-estimator converse meet over the full fixed-overlap class.

Theorem T-5 (Matched minimax frontier)

There are universal tuning constants H0>0H_0>0, 0<κ<10<\kappa<1, and an integer cutoff D0>2D_0>2 for the Jackson factorial estimator such that the following holds. For every fixed overlap constant ϵ(0,1/2)\epsilon\in(0,1/2), there are real constants cϵc_\epsilon and CϵC_\epsilon with 0<cϵCϵ. 0<c_\epsilon\leq C_\epsilon . For every n1n\geq 1 and d2d\geq 2, with Ld=log(ed)L_d=\log(ed) and with Rn,d,ϵ\mathfrak R_{n,d,\epsilon} denoting the fixed-sample observed minimax squared risk in Definition P-12, cϵmin{1,dnLd}Rn,d,ϵsupPDd,ϵobsEPn ⁣[(V^n,dJFΨ(P))2]Cϵmin{1,dnLd}. c_\epsilon \min\left\{1,\frac{d}{nL_d}\right\} \leq \mathfrak R_{n,d,\epsilon} \leq \sup_{\mathbb P\in\mathcal D_{d,\epsilon}^{\mathrm{obs}}} E_{\mathbb P^{\otimes n}}\!\left[ \left(\widehat V_{n,d}^{\mathrm{JF}}-\Psi(\mathbb P)\right)^2 \right] \leq C_\epsilon \min\left\{1,\frac{d}{nL_d}\right\}. Moreover, the fixed-sample minimax squared risk over the full-data causal completion class Pd,ϵ\mathcal P_{d,\epsilon} is exactly Rn,d,ϵ\mathfrak R_{n,d,\epsilon} for the same nn, dd, and ϵ\epsilon.

Thus the observed-law and causal-completion formulations have exactly the same finite-sample minimax order.

Implications

Let dnd_n denote the covariate-alphabet size at sample size nn.

informal · Theorem T-7 For fixed overlap, both observed and causal minimax risks converge to zero exactly when dn=o{nlog(en)}d_n=o\{n\log(en)\}, and they have the parametric n1n^{-1} order exactly when dn=O(1)d_n=O(1).

  • Bounded alphabets recover the usual parametric squared-risk scale.
  • Growing alphabets remain uniformly learnable throughout the stated consistency region.
  • In the clinic example, increasingly refined risk groups preserve uniform consistency when their number satisfies dn=o{nlog(en)}d_n=o\{n\log(en)\}.

Additional Result

informal · Theorem T-6 With universal estimator tuning, the predecessor risk bracket remains valid, the Jackson factorial estimator equals the empirical-ratio estimator for 2d<D02\le d<D_0, and bounded-alphabet sequences attain the n1n^{-1} minimax scale.

This comparison anchors the construction to the familiar cellwise estimator at small dd.

Conclusion

  • The optimal-treatment value is a nonsmooth large-alphabet functional generated by cellwise treatment-effect ties.
  • Its finite-sample minimax squared risk is of order min{1,d/[nlog(ed)]}\min\{1,d/[n\log(ed)]\} under fixed overlap.
  • The Jackson factorial estimator attains this order using local approximation and unbiased polynomial lifting.
  • Equal-propensity L1L_1 embeddings and moment matching establish optimality over all estimators.
  • The result yields exact consistency and parametric-rate boundaries for both observed and causal formulations.
  • The rate characterization is uniform over all n1n\ge1 and d2d\ge2, with constants depending on fixed overlap.