Minimax Estimation of Optimal Treatment Values with Discrete Covariates
Abstract
This paper characterizes the finite-sample minimax risk for estimating the scalar optimal-treatment value with a binary treatment, binary outcome, and a discrete covariate alphabet of size . Under consistency, conditional exchangeability, i.i.d. observed sampling, and a fixed overlap parameter , the causal value equals the observed-law functional that sums, over covariate cells, the larger of the two arm-specific outcome regressions. For sample size , the fixed-sample minimax squared risk satisfies The upper bound is attained by an explicit Jackson factorial estimator: pilot counts form local rectangles for each four-coordinate treatment–outcome cell table, a tensor Jackson polynomial approximates the cellwise maximum contribution, centered factorial moments lift the polynomial into unbiased count statistics under Poisson splitting, and Rao–Blackwellization returns an all-data fixed-sample estimator. The lower bound embeds normalized two-sample distance into an equal-propensity slice of the observed model and uses moment matching for the absolute-value cusp. The matched rate gives uniform consistency exactly when and gives the parametric minimax scale exactly along bounded-alphabet sequences.
Introduction
Optimal treatment choice assigns the arm with the larger conditional outcome mean to each covariate profile. In a discrete covariate model, the population value of this oracle rule is a scalar obtained by summing cell masses times the larger of two arm-specific regressions. This scalar is central in treatment-choice analysis because it reports the welfare level achieved by the unrestricted cellwise oracle assignment, the benchmark against which any particular learned rule is measured. The statistical difficulty comes from estimating many cellwise maxima when the covariate alphabet grows with the sample size.
This paper studies that difficulty under the fixed-overlap observed experiment. The data are i.i.d. triples , where , , and . The observed law belongs to , the fixed-overlap class defined in Definition 3, with overlap parameter . The target is the observed-law optimal-regression value from Definition 4. Under the full-data completion class in Definition 6, which imposes consistency, conditional exchangeability, and the same fixed-overlap observed margin, Proposition 2 identifies this observed value with the causal oracle value from Definition 9.
The main result is a matched all-regime finite-sample minimax law. With , Theorem 3 proves that, for every fixed , there are constants such that for every and , where , the fixed-sample observed minimax squared risk, is defined in Definition 5. The same theorem states that the corresponding minimax risk over the causal completion class equals this observed-law risk. Thus the scalar oracle value is uniformly estimable precisely up to the rate dictated by the alphabet-to-sample ratio .
The estimator achieving the upper bound is written out in closed form. The empirical-ratio estimator in Definition 10 serves as the bounded-alphabet branch. For growing alphabets, Algorithm 1 defines a Jackson factorial estimator that splits a Poissonized subsample into pilot and evaluation counts, builds a pilot-local rectangle for each covariate cell, approximates the globally defined cell contribution from Definition 8 by a tensor Jackson polynomial, replaces normalized monomials by centered factorial lifts, clips each cell estimate at its pilot scale, projects the aggregate to , and Rao–Blackwellizes over the auxiliary randomization. Theorem 1 gives the resulting uniform risk bound over . The construction is stated as a rate-achieving decision-theoretic procedure: every ingredient is defined by an explicit formula, yielding a complete statistical specification of the tensor-polynomial estimator and Rao–Blackwellized rule at the stated alphabet sizes and approximation orders.
In the proof, the logarithmic factor is tied to the absolute-value/ lower-bound mechanism for the cellwise maximum. On the overlap cone of Definition 7, the cell contribution equals the larger arm-specific regression contribution; after normalization, its hard local component has the geometry of an absolute value. The polynomial approximation and factorial-moment architecture follows the large-alphabet nonsmooth-functional literature, where approximation of , moment matching, and unbiased polynomial lifts generate effective logarithmic sample-size gains (Cai et al., 2011; Jiao et al., 2015; Wu et al., 2016; Han et al., 2018; Valiant et al., 2017). The tensor Jackson certificate in Lemma 4 supplies the local approximation and coefficient control needed for the four-coordinate treatment–outcome tables.
The converse embeds a classical large-alphabet testing problem into the treatment-value model. Proposition 3 maps two distributions into an equal-propensity observed law whose optimal-regression value is . This exact Markov reduction transfers the normalized two-sample minimax obstruction of Jiao et al. (2018) to the observed optimal-value problem. In the dense regime, Lemma 10 gives the corresponding moment-matched construction using the absolute-value approximation lower bound of Cai et al. (2011). Theorem 2 combines these ingredients into the all-estimator lower bound matching the Jackson factorial upper rate.
The rate has direct implications for growing categorical adjustment. Theorem 4 shows that, along alphabet sequences , the observed and causal minimax risks converge to zero exactly when . The same theorem characterizes the parametric minimax scale exactly by . These statements translate the matched finite-sample frontier into operational growth rules for discrete covariate tables under fixed overlap.
The analysis connects treatment-choice econometrics with nonregular and large-alphabet estimation. Treatment-choice and welfare-regret work studies optimal rules and policy performance (Manski, 2004; Murphy, 2003; Hirano et al., 2009; Kitagawa et al., 2018; Qian et al., 2011; Athey et al., 2021). Semiparametric analyses of optimal values describe identification, efficient influence functions, smoothing, and fixed-law inference under regularity or local conditions (Robins, 1986; Rosenbaum et al., 1983; van der Laan et al., 2015; Luedtke et al., 2016; Chakraborty et al., 2010; Hirano et al., 2012; Chen et al., 2023; Whitehouse et al., 2025; Wei, 2025). The present result gives a finite-sample minimax characterization for the scalar unrestricted cellwise oracle value over a growing discrete alphabet, with fixed overlap and arbitrary cell masses, null cells, boundary binary means, and treatment-effect ties inside the model class.
A verification appendix records the verification status and the cited external inputs for the anchored statements. The paper proceeds by reviewing related work, defining the observed and causal model classes, stating the estimator and minimax theorems, discussing implications and the sharp-constant question, and then giving the analytic and lower-bound details in the appendices.
Related work
This paper belongs to the treatment-choice literature that studies how observational or experimental data can support decisions that assign treatment by covariate profile. The classical welfare and empirical-welfare strands analyze treatment rules, regret, and policy choice when the object of interest is a decision rule or its welfare performance (Manski, 2004; Murphy, 2003; Hirano et al., 2009; Kitagawa et al., 2018; Qian et al., 2011; Athey et al., 2021). The present analysis focuses on the scalar value attained by choosing the better arm within each discrete covariate cell. That focus isolates the nonsmooth cellwise maximum that enters the optimal value and makes the alphabet size of the covariate distribution a primitive dimension of the estimation problem.
A second connection is to causal identification and semiparametric estimation of optimal values. Identification under consistency, conditional exchangeability, and overlap traces to the potential-outcome and inverse-probability traditions (Robins, 1986; Rosenbaum et al., 1983). Targeted-learning and efficient-influence-function analyses then study optimal-value parameters and data-adaptive treatment rules under regularity conditions that support asymptotic distribution theory (van der Laan et al., 2015; Luedtke et al., 2016). The fixed-overlap discrete model considered here keeps the same causal target once the observed law identifies it, and then asks for the finite-sample minimax risk of estimating the identified scalar optimal value.
The cellwise maximum also places the problem in the nonregular-inference literature. Optimal treatment values have kink points where two treatment arms have equal conditional means, and those ties create the familiar failure of ordinary differentiability for max-type functionals (Chakraborty et al., 2010; Hirano et al., 2012). Recent work develops smoothing, debiasing, and local-asymptotic analyses for related nonregular value and policy-learning targets (Chen et al., 2023; Whitehouse et al., 2025; Wei, 2025). The contribution here characterizes the minimax scale for the discrete-covariate scalar value under fixed overlap, with the dense-alphabet regime governed by approximation of the absolute-value singularity induced by the maximum.
That approximation structure links the analysis to large-alphabet estimation of nonsmooth distributional functionals. Polynomial approximation and moment matching have proved central for sharp rates in estimating entropy, support-related quantities, and distances between high-dimensional discrete distributions (Jiao et al., 2015; Wu et al., 2016; Han et al., 2018; Valiant et al., 2017), while Cai et al. (2011) supplies the sharp Gaussian absolute-value-cusp analysis used in the lower-bound comparison. The closest analogue is the normalized two-sample distance problem of Jiao et al. (2018), where the absolute-value functional yields logarithmic factors through best polynomial approximation. In the treatment-value problem, fixed overlap maps each covariate cell to a four-atom observed table and the maximum over arms creates the same absolute-value geometry after normalization.
Two large-alphabet papers are close enough to compare rate by rate. Zeng et al. (2024) study causal inference with unrestricted high-dimensional discrete covariates. Their target is a linear treatment-specific mean under an alphabet of size , estimated from i.i.d. observations; they show that regression, weighting, and doubly robust estimators attain mean-squared error of order , prove a minimax lower bound of order , and obtain the faster order under effect homogeneity or a known covariate distribution. Their aggregation is signed and linear, so the leading difficulty is bias accumulation across cells. The target of the present paper replaces that signed sum by the cellwise maximum, and the resulting frontier is : a different exponent in , a different consistency boundary, and a logarithmic gain whose source is the polynomial approximation of a cusp.
Jiao et al. (2018) give the second comparison. For two unknown -point distributions observed through independent Poissonized samples of size each, they construct minimax-rate estimators of and prove the matching squared-error obstruction, so their result is a full estimator-and-converse analysis of the two-sample distance. The present lower bound transfers that converse into the treatment-value model through the equal-propensity embedding of Proposition 3, and the upper bound is constructed here for the observational cell table. Three features of the observational problem require new work. The arm denominators are unknown and are estimated from the same sample, whereas the two-sample problem fixes the sample sizes by design; the propensities are unequal and constrained only by fixed overlap, so the two arms of a cell carry different effective sample sizes; and each covariate cell is a grouped four-atom table, where the two-sample problem presents a pair of scalar masses, which is why the upper bound needs a four-coordinate tensor Jackson construction with pilot-adaptive rectangles.
Standard minimax tools such as two-point and moment-matching lower bounds, together with approximation-theoretic upper bounds (Le Cam, 1986; Tsybakov, 2009; DeVore et al., 1993), provide the methodological basis for matching the estimator and lower bound in this setting.
Setup and assumptions
Throughout, denotes the covariate alphabet size, the fixed sample size, and the overlap constant. Sums over covariate cells run over , and treatment and outcome indices take values in . Generic positive finite constants may change from line to line; constants indexed by may depend on the fixed overlap constant. Vector norms are the usual finite-dimensional norms, with written explicitly when it is used.
We begin with the elementary projection notation used later by the estimators.
For real numbers , let be the interval projection .
The statistical experiment is a finite observational study. One unit is the observed triple =(,,), where the covariate lies in the finite alphabet and the treatment and outcome are binary. The observed law is a probability law on this finite sample space.
For each observed law , let denote the observed triple for unit . The observations are independent and identically distributed with common law .
⊢ LeanAssumption 1 is the standard independent and identically distributed sampling condition for the observed data (Zeng et al., 2024). It fixes the repeated-sampling experiment under each admissible observed law and makes the minimax risk below a fixed-sample decision problem.
For causal interpretation, we place the observed experiment inside a finite binary potential-outcome model. A full-data law governs on the same covariate alphabet.
The observed outcome satisfies almost surely.
⊢ LeanAssumption 2 is the standard consistency condition (Zeng et al., 2024). It ties the realized outcome to the potential outcome under the realized treatment arm, so the observed binary response carries the arm-specific potential-outcome information in treated and untreated cells.
For each , the potential outcome is conditionally independent of treatment given covariates:
⊢ LeanAssumption 3 is the standard armwise conditional exchangeability condition (Zeng et al., 2024). It encodes the observational-design requirement that, within each covariate cell, treatment assignment carries no further information about each potential outcome.
Let be the finite covariate alphabet. For every observed law under consideration and every with covariate-cell mass , the treatment propensity satisfies where is the fixed overlap constant.
⊢ LeanAssumption 4 is the standard fixed strong-overlap condition (Zeng et al., 2024). The propensity is bounded away from the edges on every occupied covariate cell , with denoting the corresponding cell mass; this keeps both treatment arms represented at the population level in each relevant cell.
We next define the observed-law model and the target functional. The observed margin records exactly the distribution of the observable triple induced by a full-data law.
For a full-data law , its observed margin is the law on defined by
⊢ LeanThe observed class parameterizes each covariate cell by its mass, propensity, and arm-specific binary outcome mean . Its atom probabilities are arranged into the cell vector , and the table across cells is denoted by .
Under Assumption 4, let The class consists of all laws on for which there exist , propensities , and means such that, for every and , where For each cell , write The support contains at most covariate cells, and on any null cell the parameters may be chosen arbitrarily within the displayed ranges.
⊢ LeanWithin this observed class, the estimand is the population value obtained by choosing the better arm in each covariate cell according to the conditional outcome means.
For as in Definition 3, with table , the observed-law optimal-regression functional is This value is independent of the permitted choice of on cells with .
⊢ LeanThe risk criterion evaluates estimators of uniformly over the fixed-overlap observed laws.
The fixed-sample minimax squared risk is where the infimum is over all measurable maps
⊢ LeanThe full-data causal class collects exactly the potential-outcome laws whose observed margins belong to the preceding observed-law model and whose potential outcomes satisfy the causal conditions already stated.
For every and , is the set of probability laws for such that:
(Support.) and -almost surely.
(Potential-outcome structure.) satisfies Assumptions 2 and 3.
(Observed fixed-overlap margin.) The observed margin from Definition 2 belongs to the fixed-overlap observed-law class under Assumption 4.
The analytic representation of the target works cell by cell. For a generic four-vector , the next definition records the overlap geometry in terms of arm masses and total mass .
The four-cell overlap cone is where, for each ,
⊢ LeanThe cell contribution is then extended from the overlap region to the whole nonnegative orthant by stabilizing the arm denominators.
For with , define the overlap-stabilized arm contribution and the globally defined cellwise maximum contribution At the origin,
⊢ LeanFinally, the causal optimal-treatment value is the full-data analogue of the observed optimal-regression value.
For a full-data law , the full-data causal optimal-treatment value is
⊢ LeanThe following result connects the observed-law representation, the analytic extension, and the causal completion class.
Let , let , and let in the sense of Definition 3. Then:
(Cell overlap.) For every , the four-vector belongs to the overlap cone of Definition 7.
(Observed value.) The observed-law optimal-regression value of Definition 4 satisfies
(Global extension bounds.) For every , the global extension of Definition 8 satisfies
(Global Lipschitz property.) For every ,
(Completion.) There exists a full-data law in the sense of Definition 6 whose observed margin satisfies . The completion is consistent and conditionally exchangeable, treatment is independent of the joint potential-outcome vector given , and the two potential outcomes are conditionally independent Bernoulli variables given .
(Observed margins.) For every full-data law , its observed margin of Definition 2 belongs to .
Proposition 1 establishes three facts used throughout the analysis. First, every observed cell lies in the overlap cone. Second, the observed value decomposes as the sum of the globally defined cell contributions . Third, the extension is bounded and Lipschitz on the nonnegative orthant, which supplies the analytic control used by the estimator construction in the main results.
The completion clauses give the causal interpretation of the observed model. Every observed law in admits a full-data law in with the same observed margin, and every law in returns an observed margin in . Thus the observed minimax problem is aligned with a potential-outcome class satisfying consistency, conditional exchangeability, and fixed overlap.
Let , let , and let be a full-data law satisfying in the sense of Definition 6. Then the full-data oracle value in Definition 9 agrees with the observed-law optimal-regression value in Definition 4 evaluated at the observed margin in Definition 2:
⊢ LeanProposition 2 identifies the full-data optimal-treatment value with the observed-law functional evaluated at the observed margin . The equality lets the minimax analysis target on while retaining its causal interpretation for laws in .
Main results
The setup in Definitions 3, 4, and 5 turns the value problem into uniform estimation of the observed-law optimal-regression value over the fixed-overlap observed model. This section gives estimators and finite-sample risk bounds for that problem. The final statements then transfer the same rate to the causal completion class through the identification result in Proposition 2.
For a fixed alphabet, the natural benchmark is the empirical cell-ratio estimator . It estimates each covariate-cell mass and each arm-specific success probability by its sample analogue, then maximizes over treatment arms cell by cell.
For and , define the empirical counts and Set with the convention . The empirical-ratio optimal-value estimator is
⊢ LeanThis plug-in form is useful at bounded alphabet sizes. The high-dimensional construction below instead estimates the cell contribution through a polynomial approximation to the Lipschitz extension from Definition 8. The construction follows the standard use of polynomial approximation for nonsmooth functionals (DeVore et al., 1993; Jiao et al., 2015).
The treatment-outcome coordinate set is A coordinate index is denoted by .
⊢ LeanThe set indexes the four treatment-outcome atom coordinates in each covariate cell. Given a pilot rectangle in these four coordinates, the next definition records the Jackson polynomial used to approximate the cellwise contribution.
For an integer , define the order-four Jackson kernel on by with removable values filled continuously and normalizing constant chosen so that For a rectangle define its coordinate centers and radii by The associated tensor polynomial is defined by the even product convolution
⊢ LeanThe order grows on the logarithmic alphabet scale, so the approximation error and coefficient growth can be balanced against sampling noise. In the algorithm, denotes projection onto the interval , and the constants , , and are fixed tuning constants chosen in the risk theorem.
The inputs are , the alphabet size , and tuning constants , , and .
Define the logarithmic scale, Poissonized scale, and polynomial degree by
If , output If and , output
In the remaining case, draw On the event , set
On the event , draw independent marks For each and coordinate , define pilot and evaluation counts and
Define pilot frequencies and rectangle half-widths by Set and Let and denote the vector of coordinate centers and coordinate radii of .
Expand the local polynomial as For , define the centered factorial lift
Define the clipping scale the unprojected cell estimate and the clipped cell estimate
Define the randomized projected aggregate by The all-data Rao–Blackwellized estimator is where the conditional expectation integrates only the auxiliary Poisson count and marks .
The all-data estimator uses the auxiliary Poisson count and random split only inside a Rao–Blackwellization. Conditional on the observed sample, the final object is a deterministic measurable estimator. The centered factorial lift in the construction gives unbiased polynomial terms under the Poissonized evaluation counts, while clipping and final projection keep the aggregate on the value scale.
For an estimator based on and an observed law , define where denotes the law of independent observations each distributed as .
The first risk statement is uniform over the observed model class with fixed overlap. It uses the i.i.d. sampling condition from Assumption 1 and the observed-law value from Definition 4.
Assume the sampling condition in Assumption 1: for every alphabet size , sample size , and observed finite law , the sample law is the canonical i.i.d. product law generated by .
Then there exist tuning constants , , and an integer cutoff such that the Jackson factorial estimator of Algorithm 1 satisfies the following guarantee whenever:
(Overlap.) The overlap constant satisfies .
(Sample size.) The sample size satisfies .
(Alphabet size.) The alphabet size satisfies .
(Observed model.) The observed law belongs to as in Definition 3.
For every such , there is a constant such that, for every such , and , Here is the observed-law optimal-regression value of Definition 4.
⊢ LeanTheorem 1 gives the estimator side of the minimax rate. For each fixed overlap constant, the constant controls the entire model class and all finite . The logarithmic denominator comes from approximating the cellwise maximum by a logarithmic-degree polynomial and then estimating the resulting polynomial coefficients through factorial moments.
The converse starts from a normalized two-sample problem. Let denote the -point probability simplex, and let and be two simplex points. The following embedding places this distance problem inside the equal-propensity slice of the observed model.
For , define the observed atom masses Complete the observed law by taking the potential outcomes to be conditionally independent Bernoulli variables, setting the treatment propensity equal to , and imposing consistency.
⊢ LeanLet be independent paired observations with , , and . The normalized two-sample minimax mean-squared error is where the infimum is over all measurable real-valued functions of the paired sample.
For , is the full-data law on with covariate mass , treatment propensity , outcome regressions and on cells with , conditionally independent Bernoulli potential outcomes with these means, treatment drawn independently of the potential outcomes given , and observed outcome . Equivalently, its observed margin has cell masses and , with a consistency-compatible completion.
Under this embedding, the observed optimal value is an affine transform of the distance. The reduction is stated as an exact product-kernel comparison, in the standard experiment-comparison sense of Le Cam (1986) and Tsybakov (2009).
Let , let , let be an observed law on , and let be a law on observed triples. Suppose that:
(Sampling.) The law is the -fold product law generated by , as in Assumption 1.
(Alphabet size.) .
(Overlap range.) .
Then, for every pair , let where is the observed margin from Definition 2 and is the embedding from Definition 16. The law belongs to in the sense of Definition 3, and, for every , Moreover, whenever , and the observed-law optimal-regression value from Definition 4 satisfies Finally, the explicit fixed- product Markov kernel obtained by applying the single-observation embedding kernel independently across the coordinates transforms the fixed paired-sample law generated by exactly into the product law . Consequently, for the fixed-sample minimax squared risk in Definition 5, where is the normalized two-sample minimax mean-squared risk.
⊢ LeanProposition 3 gives a concrete subexperiment inside the causal-value problem. On this subexperiment, estimating is exactly as hard as estimating a normalized distance up to the displayed factor. The Markov-kernel clause makes the comparison fixed-sample and preserves the product sampling structure.
The all-regime lower bound combines this reduction with the moment-matching lower-bound machinery for absolute-value and functionals in Cai et al. (2011) and Jiao et al. (2018).
Suppose the observed samples satisfy the independent product sampling law in Assumption 1. For every overlap constant , there exists a constant such that, for every and , the fixed-sample minimax squared risk in Definition 5 satisfies
⊢ LeanThe constant depends only on the overlap level. Theorem 2 supplies the information-theoretic scale matched by the Jackson factorial construction: the same expression , truncated at one, governs the worst-case squared risk over the observed model.
Putting the upper and lower inequalities together gives the finite-sample minimax rate. The statement also records the causal-completion equivalence supplied by the observed-margin bridge in Propositions 1 and 2.
There are universal tuning constants , , and an integer cutoff for the Jackson factorial estimator such that the following holds. For every fixed overlap constant , there are real constants and with For every and , with and with denoting the fixed-sample observed minimax squared risk in Definition 5, Moreover, the fixed-sample minimax squared risk over the full-data causal completion class is exactly for the same , , and .
⊢ LeanTheorem 3 characterizes the finite-sample minimax order for the observed-law and causal-completion formulations. The upper inequality is attained by the explicit estimator in Algorithm 1; the lower inequality inherits the approximation and moment-matching ingredients of Cai et al. (2011) and Jiao et al. (2018). The equality for the causal completion class means that the observed-data risk law is also the risk law for the causal optimal-treatment value.
For asymptotic interpretation, it is useful to name the causal minimax risk directly.
For , , and , write for the law of independent observations each distributed as the observed margin , and define the observed-sample causal minimax squared risk by where the infimum is over all measurable maps from to .
The final result translates the finite-sample rate into growth conditions for a sequence of alphabet sizes. Its lower side uses the same published lower-bound ingredients cited above, and its upper side uses the Jackson factorial estimator with the universal tuning constants from Theorem 3.
For every overlap constant with , and for every alphabet-size sequence satisfying for all , the following equivalences hold:
(Observed consistency.) The fixed-sample minimax squared risk of Definition 5 satisfies if and only if
(Causal consistency.) The causal minimax squared risk over satisfies if and only if
(Observed parametric boundary.) The observed minimax risk sequence has the parametric rate, meaning that there exist constants with such that, for all sufficiently large , if and only if the alphabet-size sequence is uniformly bounded asymptotically, meaning .
(Causal parametric boundary.) The causal minimax risk sequence has the parametric rate, meaning that there exist constants with such that, for all sufficiently large , if and only if .
Theorem 4 gives the operational phase diagram. Uniform consistency corresponds exactly to alphabet growth below . The parametric risk scale corresponds exactly to asymptotically bounded alphabet size, for both the observed-law target and the causal optimal-treatment value.
Discussion and limitations
The main results give a uniform mean-squared-error scale for the observed-law optimal value over the fixed-overlap model. This section translates that rate into several modeling consequences for categorical adjustment, records the comparison with the predecessor analysis on the same formal class, and isolates the sharp-constant question suggested by the matched upper and lower bounds.
There exist tuning constants , , and an integer cutoff , used by the Jackson factorial estimator in Algorithm 1, such that:
(Matched frontier.) For every overlap parameter , there are constants and with such that, for every and ,
(Bounded alphabets.) For every , every , every overlap parameter , and every observed sample, the Jackson factorial estimator in Algorithm 1 equals the empirical-ratio estimator in Definition 10:
(Bounded-alphabet sequences.) For every overlap parameter and every sequence with for all and for some finite , there are constants and with such that, for all sufficiently large ,
Proposition 4 places the matched-rate analysis on the same observed model class and target used by the earlier bracket. Its first clause is labelled Matched frontier because it is a consequence of the matched frontier, and not because the two displayed bounds meet: the lower expression and the upper expression are the predecessor’s nonmatching bracket, which Theorem 3 implies and strictly sharpens. The clause is recorded here so that the earlier bracket can be read off the present theorem on the same formal class; the sharp comparison is Theorem 3 alone. The second clause explains how the constructed estimator is anchored at small alphabets, and the third clause records the fixed-dimensional parametric consequence in sequence form. The external ingredients are the classical moment-matching approximation to in Cai et al. (2011) and the two-sample lower-bound mechanism in Jiao et al. (2018).
The statistical message is that, once fixed overlap is imposed, alphabet growth is what determines the nonparametric difficulty. Exact treatment-effect ties create the local nonsmoothness of the cellwise maximum, and the Jackson factorial construction in Algorithm 1 smooths that nonsmoothness at the logarithmic degree dictated by the alphabet scale. Unequal propensities enter through the overlap-stabilized representation in Definition 8: fixed overlap keeps the observed arm masses comparable enough for the same rate calculation, while the constants may depend on the overlap parameter.
Null cells and boundary outcome means are handled inside the model and estimator definitions themselves. The observed model class in Definition 3 permits zero covariate-cell mass, and the empirical-ratio estimator in Definition 10 fixes the zero-over-zero convention used by the bounded-alphabet branch. At the boundary of the binary outcome space, the global extension in Definition 8 keeps the cell contribution defined on the nonnegative orthant, so the same risk statement covers outcome means equal to zero or one.
For bounded alphabets, Proposition 4 gives the usual minimax scale after a finite burn-in along the sequence. This confirms that the logarithmic approximation gain is a growing-alphabet phenomenon: when the covariate alphabet is uniformly bounded, the optimal-value problem has the same squared-risk order as an ordinary finite-dimensional regular estimation problem, with constants depending on the alphabet bound and overlap.
Limitations and future work
The established domain is deliberately narrow. Every result in this paper is proved for a binary treatment, a binary outcome, a finite covariate alphabet, and a fixed overlap constant , with the comparison constants and depending on ; the target is a scalar and the loss is squared error; and the number of treatment arms is two and does not grow. Four directions therefore sit outside the present theorems. Overlap that vanishes with the sample size changes the effective per-cell information and is not covered by constants indexed by a fixed . A growing number of treatment arms replaces the two-arm maximum by a maximum over a growing set, which changes the approximation problem the Jackson degree is calibrated to. Continuous or unbounded outcomes, and continuous covariates, leave the finite four-atom cell representation on which the estimator and the embedding are built. Finally, the paper gives a point estimator and a squared-risk frontier; uniform confidence procedures for the optimal value under the same growing-alphabet asymptotics are a separate question, and the nonregularity at treatment-effect ties is what makes it a hard one.
The matched comparison constants leave open an exact efficiency calculation for nonsaturated growing alphabets. The natural normalization is determined by the term in Theorem 3; an exact constant would identify the efficiency price of estimating the cellwise maximum under unknown observational treatment imbalance.
The route rescales the pilot-local four-cell rectangles by their Poisson noise geometry, optimizes the tensor-polynomial bias–variance functional for , and dualizes the approximation constraint into moment-matched observed laws in for the target .
⊢ LeanDefinition 18 names the analytic route suggested by the upper-bound construction and the lower-bound duality. It combines the local Poisson geometry used by the Jackson factorial estimator with the approximation problem for , the globally defined cell contribution from Definition 8. This route is the candidate mechanism for moving from comparison constants to a sharp overlap-dependent constant.
A natural next question is whether, along each nonsaturated sequence with and the normalized risk converges to a finite positive limit, and whether an estimator attaining that limit can be derived through .
⊢ LeanRemark 1 states the resulting open question. The nonsaturated condition keeps the risk on the vanishing side of the consistency boundary in Theorem 4, and the normalization matches the leading growing-alphabet term in Theorem 3. A positive answer would refine the present minimax rate into an exact asymptotic constant for fixed overlap.
Appendices
Analytic ingredients for the upper bound
The upper bound in Theorem 1 rests on two local analytic devices. The first is a Jackson polynomial certificate for the globally defined cell contribution from Definition 8, localized to the pilot rectangle constructed in Algorithm 1. The second is a centered Poisson factorial lift that turns the local polynomial into a cell statistic whose bias and variance remain controlled at the cell mass scale.
We begin with the approximation component. The construction uses the four outcome-arm coordinate set and smooths a bounded continuous function on the normalized cube by the tensor product of the order-four Jackson kernel in Definition 12. The following coefficient envelope records the degree and one-norm control needed when the smoothed polynomial is later evaluated by factorial moments.
With the Jackson kernel notation of Definition 12, assume:
(Order.) is an integer.
(Function.) is continuous on the normalized cube , where is the four arm-outcome coordinate set.
(Envelope.) and for every .
Then there exists a polynomial in the four normalized coordinates such that, for every angle vector , Here cosine is applied coordinatewise. Moreover, every multi-index satisfies the total degree of is at most , and its coefficient one-norm satisfies
⊢ LeanWrite and define the tensor convolution For the positive order in the lemma, write , where with the removable values filled continuously, and set Then for every , and , so and . Thus and Fubini’s theorem for the nonnegative continuous tensor product then gives Continuity of on the normalized cube and continuity of the coordinatewise cosine map give the integrability needed for the displayed convolutions and for the linearity identities below.
First construct the polynomial representation. The Jackson frequency bound follows from the Chebyshev form of the raw kernel. If denotes the Chebyshev polynomial of the second kind, then Since , the fourth power is an even polynomial in . Hence there is a polynomial of degree at most such that Using , this gives Multiplication by the constant preserves the frequency bound, and a polynomial of degree at most in is a trigonometric polynomial with frequencies at most . Thus there are real coefficient sequences and , , such that
Fix all coordinates except . For an angle vector and , write Set, for , The period-box translation , with the box understood modulo the common -period in each coordinate, gives The translation is valid because the displayed integrand is continuous and is -periodic in each coordinate; the periodicity follows from the cosine term and the trigonometric-polynomial representation of . Define Substituting the finite expansion of , using the angle-subtraction identities for sine and cosine, and integrating the finite sum term by term gives Thus each one-coordinate section is a trigonometric polynomial of frequency at most . The same section is even in that coordinate: reflecting the integration variable preserves , the kernel is even, and cosine is even. Therefore the sine part cancels. Interpolating the resulting even trigonometric polynomial in one cosine coordinate at the nodes and repeating this coordinate by coordinate gives a polynomial in the four variables such that The interpolation basis has degree at most in the coordinate being added, while the induction hypothesis controls the other coordinates. Hence, for every , and
The convolution is bounded by the same envelope . Since on the normalized cube, Kernel nonnegativity gives By linearity of the integral and the unit-mass identity, Therefore
This angle-space bound transfers to the whole normalized cube. For any , take so that . The representation from the first step and the bound from the second step give
It remains to bound the coefficient one-norm. The coefficient estimate used here is the following. Let be a polynomial in normalized variables, each coordinate degree at most , and suppose Then For one variable, if has degree at most and on , then is an even trigonometric polynomial, Orthogonality on gives Since and the Chebyshev coefficient one-norm satisfies the polynomial identity yields For the multivariate estimate, write For each fixed , the preceding one-dimensional bound gives Applying the induction hypothesis to each coefficient polynomial and summing over gives which is exactly The case is the constant-polynomial case and follows directly from the displayed uniform bound.
Apply this coefficient estimate to , , and . With , The elementary bound and imply Consequently Since and , Because the base is at least and , Together with the representation, coordinate-support bound, and total-degree bound proved above, this is the asserted polynomial .
The envelope in Lemma 1 is deliberately expressed in normalized coordinates. After an affine change of variables from the pilot rectangle to , it gives a physical polynomial while preserving a centered expansion around , the rectangle center from Definition 12.
Using the overlap-parameter notation of Assumption 4, let , let be an integer, and let be a four-coordinate rectangle with Assume:
(Degree.) .
(Lower endpoints.) for every .
(Overlap parameter.) .
(Radii.) for every .
Then there exist a polynomial in the four physical coordinates and a polynomial in the four normalized coordinates such that:
(Physical degree.) Every coordinate degree of is at most .
(Normalized degree.) Every satisfies for every coordinate .
(Centered evaluation.) For every four-coordinate vector , Here is the globally defined cellwise maximum contribution from Definition 8.
(Jackson convolution.) For every angle vector , equals the order- tensor Jackson convolution from Definition 12 of precomposed with the affine parametrization of , evaluated at .
(Coefficient envelope.)
Define and If , then and , since each .
For nonnegative four-coordinate vectors , put For each arm , Definition 8 gives Moreover because when , and the zero-mass value is . If , nonnegativity gives , hence The case is symmetric.
Assume now that and . The elementary bounds and will be used throughout. If then so If then Since and , while we obtain If then The second summand is nonnegative, giving For the upper bound, and so The remaining mixed case, is analogous. Here The subtracted summand is nonnegative, hence Also and which yields Thus, for both arms, Since Definition 8 gives Consequently is continuous on , and for every , Set
Define the centered pullback on the normalized cube by The affine map sends into , so is continuous and Since , we have . Applying Lemma 1 to with envelope gives a polynomial in the normalized coordinates such that and
Substitute normalized coordinates into physical coordinates. Since every , there is a polynomial in the four physical coordinates satisfying and every coordinate degree of is at most . Define Adding a constant preserves the coordinate-degree bound, and therefore
For an angle vector , Thus Using the convolution identity for , The integrands are continuous on the compact period box, and the tensor Jackson kernel has unit mass, Therefore the constant centered term integrates exactly to the value added back: Thus For each coordinate, : both sine factors in its defining ratio change sign, and the ratio is raised to the fourth power. The coordinatewise substitution preserves and Lebesgue measure, and therefore changes to while leaving the product kernel unchanged. The last display is consequently the order- tensor Jackson convolution in Definition 12 of precomposed with the affine parametrization of , evaluated at .
Substituting the definition of into the coefficient estimate gives The constructed and have the asserted physical degree, normalized degree, centered-evaluation identity, Jackson-convolution identity, and coefficient envelope.
The extraction step supplies the polynomial used by the estimator, and its centered form is the one matched to the factorial lift below. The dependence on is the local scale that later becomes the Poisson noise scale after the pilot rectangle is chosen from the data.
For reference, we also isolate the tensor convolution operator used in these statements. This notation matches the product-kernel approximation scheme standard in Jackson-type polynomial approximation (DeVore et al., 1993).
For a positive integer , let on , with the removable values filled in continuously and with chosen so that . For a finite coordinate dimension and a function , the order- tensor Jackson convolution is the operator defined, for , by where cosine and addition are coordinatewise. Equivalently, the product kernel is normalized to have integral one over .
The approximation error is controlled through a weighted modulus along angular perturbations. The square-root boundary term in the main certificate comes from converting angular increments back to physical coordinates on a rectangle.
Let be a finite coordinate dimension, let be a positive integer, and let . Let , let and for each , and fix an angle vector . Suppose:
(Kernel.) The order- tensor Jackson convolution is formed from the order-four Jackson kernel in Definition 12.
(Continuity.) The function is continuous on the normalized cube .
(Weighted increment modulus.) For every shift , where cosine is applied coordinatewise.
Then
⊢ LeanLet For the positive integer in the statement, is the normalized order-four Jackson factor on , with removable values filled continuously. The raw fourth-power factor is nonnegative and has positive integral, so its normalized version satisfies for every . The tensor product kernel is normalized to have unit mass: The convolution is written in the shifted-variable form When , the product is the empty product and the sums below are empty sums, so these identities keep their displayed values. Continuity of on the normalized cube, coordinatewise continuity of cosine, and compactness of give the integrability used below.
The preceding unit-mass identity gives The subtraction is justified by the compact-domain integrability noted above.
Taking absolute values, using nonnegativity of , and applying the assumed weighted increment modulus pointwise yields
Distributing the finite sum through the integral gives For each selected coordinate , product integration and the one-dimensional unit mass give and
Write the raw kernel and its mass as so that . Let denote the Chebyshev polynomial of the second kind. The removable-value convention and the identity away from the removable points give The square expansion follows from the Chebyshev recurrence, with the displayed sum interpreted as an empty sum when . Using for positive integers , one obtains In particular,
The same Chebyshev-square expansion and the vanishing of all positive-frequency cosine integrals give For , If , then on this interval and If , put Since , After integration and normalization,
The elementary inequality is equivalent, since , to Multiplying by , integrating, and using the preceding second-moment estimate together with , gives
Combining the integral bound from the previous steps with the tensor moment reductions gives Here , , and justify multiplying the one-dimensional moment inequalities and summing them coordinatewise. This is the asserted bound.
Combining the coefficient envelope with the angular modulus gives the single polynomial certificate used in the upper-bound proof. It applies uniformly to nonnegative rectangles with positive radii, including the pilot-local rectangles that may meet coordinate faces or the vertex of the overlap cone.
There is a universal constant such that, for every overlap parameter with , there is a constant with the following property. Let:
(Degree.) .
(Rectangle.) have , , and positive radii for every .
Let be the tensor Jackson polynomial associated with , , and in Definition 12. Then the same polynomial satisfies all of the following conclusions: For every , There are a finite multi-index set and coefficients such that, for every , with and
⊢ LeanTake Fix , and set Let and let satisfy the displayed hypotheses. Applying Lemma 2 to this rectangle gives the tensor Jackson polynomial and a normalized centered polynomial . The same result gives, first,
Fix . Choose a bijection , and write Define the normalized point and angle vector by Since and every radius is positive, and On , set The lower-endpoint hypothesis gives , so the affine image of lies in . The global Lipschitz property in Proposition 1 therefore implies that is continuous on and that, for all , For every shift , the elementary inequality therefore gives Applying Lemma 3 with yields where the evaluation identity for is the Jackson-convolution identity supplied by Lemma 2. Finally, for each , because Reindexing the preceding sum over and using gives the asserted pointwise bound.
It remains to record the centered expansion and its coefficient envelope. By Lemma 2, for every four-coordinate vector , Let be the image of under the coordinate reindexing from to , and let be the corresponding reindexed coefficient. Then the preceding display becomes The normalized coordinate-degree conclusion of Lemma 2 gives and reindexing preserves the coefficient one-norm, so its coefficient envelope gives Since , Combining the last two displays proves the stated coefficient bound.
∎Lemma 4 is the analytic bridge between approximation and estimation. The pointwise bound controls the local bias of on , while the centered coefficient expansion controls the moments of the factorial estimator after replacing monomials by centered Poisson factorial polynomials.
We next record the pilot calculations that make the random rectangles usable. The good-pilot event is self-normalized: its radius is computed from the observed pilot count itself, so the event has the same scale in small and large cells.
Fix the four-coordinate set . Let have nonnegative coordinates, let , let , and let . Suppose the pilot counts are independent with With , define the normalized aggregate pilot score and define the self-normalized bad-pilot event Then, with ,
⊢ Lean1. For each , define and put Since , Thus and hence
2. We first record the scalar estimate used in the product step. Let , where , and define For every real , The Chernoff–Bennett bounds obtained from this identity give, for every , With , Indeed, on , so the bad-event inequality implies the displayed upper-tail event. On , the inequality implies and therefore , which is the displayed lower-tail event. Consequently
Next, and hence Since , . Set Then and . The preceding deterministic bound gives Using the moment-generating identity and for , The elementary exponential domination of powers then yields, for every integer with , For , define With the pointwise Young inequality gives Combining the moment bound with and the bad-event probability bound,
3. Now let be a finite coordinate set with , and let be independent with . Define If , the corresponding integral is zero. Otherwise, since each , for , and therefore When , the scalar bound from step 2 gives When , independence factors the expectation and the scalar moment and probability bounds give Thus, with Cauchy–Schwarz gives and hence Since , Also . Therefore
4. Apply the preceding product bound with The product-Poisson pilot law supplies the displayed coordinate laws and independence, and . Hence Combining this estimate with step 1 and using , Therefore Since , this is the asserted bound.
∎The exponential factor in Lemma 5 is paired with the logarithmic choice from Algorithm 1. This pairing makes the contribution from bad pilot rectangles summable across cells at the rate used in Theorem 1.
Fix a cell , four-cell mass vector , Poisson scale , alphabet logarithm , pilot-radius constant , and pilot counts . Put and . The coordinatewise good-pilot event used for the pilot rectangle is
On , the pilot rectangle radius is comparable to the local Poisson scale. The next lemma expresses that comparison in the aggregate form needed by the coefficient envelope.
Fix , an integer , pilot counts for in the four-coordinate set , and coordinate intensities . Let , set , and let be the canonical pilot rectangle of Algorithm 1, with radii Assume that the good-pilot event holds, where is the canonical pilot half-width in Algorithm 1. Then
⊢ LeanFix and set Since , , so . Hence . For each , write For the canonical tuning used in this lemma, the pilot half-width is calibrated with . Hence Also , and because both sides are nonnegative and The good-pilot event gives Substituting the displayed formula for and then using the preceding bound on , Thus
Define Since and , we have . From Equation 1, Moreover, Taking square roots gives
For the zero-truncated pilot interval, If , then . If , then . Therefore Combining this with the formula for and Equation 2,
Summing Equation 3 over the four coordinates, By Cauchy’s inequality on the four-point set , All terms are nonnegative, so Consequently, Since and , the desired bound follows:
∎The radius bound turns the abstract coefficient control into a stochastic bound indexed by , the cell mass. The following normalized envelope is the version applied directly to the centered polynomial expansion inside the estimator.
Let be the overlap parameter from Assumption 4, let be an integer, and let be a rectangle over the four outcome-arm coordinate set . Suppose that:
(Ordered endpoints.) for every .
(Nonnegative lower endpoints.) for every .
(Positive radii.) With one has for every .
Let , let , and let be the Jackson tensor polynomial attached to by Definition 12. Define where is the global cell extension from Definition 8 and . Then Moreover, every satisfies for every .
⊢ LeanThe ordered endpoints make a valid rectangle, and the remaining hypotheses give the lower-endpoint, positive-overlap, degree, and positive-radius conditions needed for Lemma 2. Apply that result to . Under the same positive-overlap and positive-radius conditions, the construction in Definition 12 identifies the extracted physical polynomial with .
Fix a bijection Write the normalized polynomial supplied by Lemma 2 as Define the coordinate transport and set Then, for every , The coefficient sum is unchanged by this relabeling:
Substitute in the preceding display. Since every , Thus, as a polynomial in ,
The last identity implies that every coefficient of outside is zero, and that for the coefficient equals . Hence
By the coefficient-envelope part of Lemma 2, Since , Also , because and every radius is positive. Combining these inequalities with the relabeling identity gives
It remains to record the coordinate-degree bound. Let . If , the polynomial identity in [proof:jackson-normalized-coefficient-envelope:normalized-expansion] gives , contradicting . Hence for some . The normalized-degree part of Lemma 2 gives Therefore, for each ,
The centered form in Lemma 7 aligns with the polynomial from Algorithm 1. This is the same unbiased polynomial-lifting principle used for nonsmooth functional estimation under Poisson sampling (Jiao et al., 2015).
Fix a cell , a four-cell mass vector , Poisson scale , and Jackson tuning. Let denote the product-Poisson law of the pilot and evaluation count vectors, and let be the coordinatewise good-pilot event used for the pilot rectangle in the Jackson–factorial construction. With denoting the clipped cell statistic built from the pilot and evaluation counts, define
The quantities in Definition 21 separate the contribution of pilot failures from the conditional bias and variance calculations on . The final lemma assembles the distributional identities, the conditional factorial moments, and the resulting cellwise controls.
Assume the product experiment satisfies the i.i.d. sampling law in Assumption 1. There is a universal Jackson tuning with self-normalized Poisson pilot-radius constant, polynomial-degree constant , and cutoff . For every overlap parameter , there is a constant such that the following assertions hold for every , every observed law on , and every positive sample size , whenever as in Definition 3. Set , let be the four-coordinate cell vector, and let be the marginal mass of cell .
In the uncapped marked experiment generated by together with i.i.d. observations from and independent fair pilot/evaluation marks, the count has law . The induced pilot/evaluation table has law equal to the product, over cells and coordinates , of two independent Poisson count arrays with coordinate means . For each fixed cell , the marginal law of the pilot and evaluation counts in that cell is therefore
For every cell , every fixed pilot table, every coordinate , every , and every , conditional expectation over the evaluation counts satisfies and
With the same tuning, the clipped Jackson-factorial cell statistic obeys, uniformly over all cells , and the pilot-failure contributions satisfy
⊢ LeanTuning and constants. Use the canonical Jackson tuning specified in the statement, whose numerical components are Fix and set Let be the constant supplied by Lemma 4, and let be the and constants in Lemma 5. Define All summands after the leading are nonnegative, so .
Cell notation. Fix , assume and as in Definition 3, and put Then . If is empty, the cellwise assertions are vacuous. Otherwise fix . Write The coordinates satisfy . Since , , and hence
Poisson splitting. The uncapped experiment draws , giving the asserted marginal law of . Conditional on inclusion, an observation in atom is independently marked pilot or evaluation with probability . Since , Poisson thinning gives pilot and evaluation intensities Thus Projecting this product law to the fixed cell gives
Centered factorial moments. By Algorithm 1, Equivalently, For , so coefficient comparison gives . With two formal variables, The coefficient of is therefore The evaluation counts are independent of the pilot counts, so with both identities hold conditionally on every fixed pilot table.
Pilot geometry and failure terms. For a fixed pilot table, define and Let and define the good-pilot event On , : if , then and ; if , then . Also and Define The interval definitions give Using the notation of Definitions 7 and 8, for any two nonnegative four-vectors , the global cell extension satisfies the four-coordinate Lipschitz bound For a fixed arm , write so when , and otherwise. If one of is zero, all coordinates of that vector vanish and gives the bound. On the positive-mass part, split according to whether or for . In the threshold-threshold case the arm value is . In the arm-active case use and . In the two mixed cases, the nonnegative gap between and is bounded by the corresponding changes in and , hence by ; the preceding estimates give the same -Lipschitz bound for each . Taking the maximum over the two arms preserves this bound. The clipping interval in Algorithm 1 therefore gives With and , set and
Good-pilot scales. By Lemma 6, on , Set Since , For , first Indeed, if , then gives . If , then and the same bound follows.
The logarithmic calibrations used below are and For , the calculus inequality for gives, with , which proves . Since , Consequently and With , the same calculus inequality gives so Multiplying by proves , and using once more gives
Good-pilot factorial bound. Fix a pilot table in , and write First, Indeed, the good-pilot inequality gives where the last inequality uses . Since , because . Combining the last two displays proves the ratio bound.
Also gives For any nonnegative integer , the centered-factorial second-moment identity from [item:pf-centered-factorial-moments] yields Let be the Jackson tensor polynomial on , and write the centered normalized expansion from Lemma 7 as with Define the factorial lift Independence of the evaluation coordinates and Minkowski’s inequality give The calibrated coefficient growth bound is Indeed, and the floor bound give Since , squaring the last display gives Also , and hence Because , Multiplying by the remaining exponential factor and using the preceding exponent comparison yields Therefore
Good-pilot bias. For the fixed good pilot, the factorial identities give Let Hence For each coordinate, To see this, if , then , and the quadratic inequality gives and , so If , then the good-pilot inequality gives and , whence Applying Lemma 4 at gives Since and , From the interval projection definition in Definition 1, a direct case split gives the following elementary consequence: for every real midpoint , every real input , and every radius , If lies inside the interval, the left side is zero. If , then the inequality follows after multiplying by from and the case is identical with the middle sign reversed. Applying this with midpoint , radius , and input , and using , gives Using the factorial bound from [item:pf-good-pilot-factorial-ltwo], Together with the good-pilot radius bound and , this implies
Good-pilot second moment. On , the four-coordinate Lipschitz bound from [item:pf-pilot-geometry-failure] and give A direct case split using the interval projection from Definition 1 gives the contraction Applying it with midpoint , radius , and input gives Therefore and hence After conditional expectation over the evaluation counts, because . Integrating over pilots on and using gives
Bad-pilot bounds. The pointwise domination from [item:pf-pilot-geometry-failure] and Lemma 5 give and Using and , Also and , so The displayed estimates are exactly the asserted pilot-failure bounds after enlarging by .
Bias. Disintegrating over the pilot table, the good-pilot conditional bound from [item:pf-good-pilot-bias] and the bad-pilot contribution give Using the estimate for and the definition of ,
Variance. Combining the good-pilot second-moment estimate with yields Using the bound on , the estimate for , and the definition of , Variance is bounded by the mean squared deviation about any fixed center, so Thus the stated bias, variance, and pilot-failure estimates hold uniformly in the cell , in , and in .
Lemma 8 supplies the cellwise estimates used to prove the upper bound. The first displayed identities identify the Poissonized marked experiment with independent pilot and evaluation arrays. The factorial identities then center the polynomial at the pilot rectangle center, and the final bounds control bias, variance, and pilot-failure terms at the same local scale.
Moment matching and lower-bound details
This appendix gives the dense lower-bound construction used in Theorem 2. The argument works inside an equal-mass, equal-propensity submodel and converts best polynomial approximation of into a separation in the optimal-value target. Its testing step follows the classical moment-matching method for nonsmooth functionals (Cai et al., 2011), with Poissonized mixtures used to make the likelihood calculation cellwise.
We first state the Poissonized risk that appears in the comparison. It is the same minimax squared loss criterion as Definition 5, evaluated under a Poisson observed-sample law and with the estimator projected to the target range.
For and , define the Poissonized minimax squared risk Here is the finite observed sample drawn from the Poisson observed-sample law with mean and atom law , the supremum ranges over the observed-law model class in Definition 3, the target is the observed optimal value from Definition 4, and the infimum ranges over all dense Poisson estimators on the -cell observed sample space.
⊢ LeanThe submodel is indexed by a dense contrast vector. Each cell has mass , treatment probability , and opposite perturbations of the two arm-specific outcome means.
For an integer and a contrast vector with for every , define as the full-data probability law on whose atom at is where if and if , and , .
⊢ LeanFor this law, the target in Definition 4 becomes an average of plus the common baseline. The lower bound therefore reduces estimation of the value to estimation of an absolute-value functional over many small coordinates.
The construction below fixes the moment-matching degree, scales the contrast amplitude, and defines the two product priors. The existence of the symmetric measures and with matching moments through degree and separation is the approximation-theoretic input from Cai et al. (2011); the same source gives constants with .
For , , the overlap parameter, a dense domain certificate and a dense prior family indexed by positive even degrees, define the dense moment-matching construction as follows. For every positive even , the family carries the prior-pair conditions that and are probability measures supported on , are symmetric, have matching moments through degree , and satisfy The construction has components:
(Dense law.) For each , is the potential-outcome law induced by the dense full-data masses, equivalently , , , and , with and conditionally independent Bernoulli variables given , independently of the potential outcomes, and .
(Approximation scale.) The approximation degree and best even-degree approximation error are The construction uses .
(Poisson intensity and amplitude.) The one-cell Poisson intensity and dense contrast amplitude are
(Coordinate priors.) The coordinate priors are
(Dense product priors.) The priors and are the pushforwards into the dense contrast class of the restrictions of and to , under coordinatewise scaling by .
(Observation mixtures.) The mixtures and are the observation-kernel mixtures induced by and , with independent Poissonized sample size .
(Risk and separation.) The risk component is the Poissonized minimax risk The prior target-mean separation is
The construction carries the supplied prior-pair certificate at degree and the identities
⊢ LeanThe amplitude condition keeps the constructed contrasts inside the dense law in Definition 23. In the range , the logarithmic degree is small enough relative to the one-cell Poisson intensity to ensure this admissibility.
Let satisfy and . For the lower-bound contrast amplitude defined in Definition 24,
⊢ LeanWrite as in Definition 24. Since , we have and hence, as real numbers, Moreover so Also , and therefore It follows that and, using ,
The assumption gives , so Consequently and hence
It remains to bound the square-root argument by . Since and , we have Dividing by the positive quantity gives Equivalently, because , Taking square roots yields Combining this with proves the claim.
∎The testing comparison is expressed in total variation. The next definitions isolate the distance, the likelihood tail left after moment matching, and the one-cell likelihood ratio used in the product calculation.
For probability laws and on a common measurable space, define , where the supremum ranges over measurable events .
For the dense lower-bound degree , Poissonized cell intensity , and dense amplitude , define the dense likelihood tail by
Fix a Poisson cell intensity . Let and be the two sign counts in one dense cell, and let denote expectation under the zero-contrast baseline law . For , the one-cell likelihood ratio relative to this baseline is In the dense lower bound, applies this definition with the cell contrast and .
Moment matching removes the first likelihood terms, so the remaining total-variation control is governed by the tail in Definition 26. The lower bound then applies the standard fuzzy-hypothesis testing reduction of Le Cam (1986) and Tsybakov (2009): two priors with separated target means, concentrated targets, and small induced total variation force squared risk at the square of the separation.
Suppose the product experiment satisfies the i.i.d. sampling law in Assumption 1. There is an integer cutoff such that, for every overlap constant , there is a constant with the following property. For all integers satisfying and :
(Degree and amplitude.) The lower-bound degree and amplitude obey
(Dense observed submodel.) For every , the dense equal-mass contrast submodel has observed margin in and satisfies, for every , with
(Moment-matched priors and mixtures.) The two dense product priors and , with induced Poissonized mixtures and , may be chosen so that their prior-mean separation satisfies and Under each of the two product priors, the target is concentrated around its prior mean: and
(One-cell likelihood identity.) For all contrasts ,
Consequently, and the fixed-sample and Poissonized risks satisfy In particular,
⊢ LeanLet By Lemma 9, the standing conditions and give The ceiling bounds also give If , then . Hence because is an integer above . Therefore . If also , then , and These relations, together with the amplitude range supplied above, prove the first group of claims.
For , the amplitude bound places every coordinate in . The dense law of Definition 24 has, in every cell, equivalently Since , the propensity belongs to , and the displayed means belong to . Thus the observed margin lies in the class of Definition 3. By Definition 4, By (Cai et al., 2011), for every positive even there are symmetric probability measures and on , with matching moments through degree , such that for universal constants . Choose these priors at , take their -fold products, and map each coordinate by . The preceding display for gives and hence The lower Cai–Low bound gives . The zero polynomial is an admissible degree- approximant to , and so . Together with , this yields The same reciprocal approximation bound supplies a universal cutoff such that, for every , Indeed, for , and . Thus, after increasing , and therefore For the one-cell likelihood identity, let and be independent counts under the zero-contrast baseline and define The Poisson probability-generating identity gives, for all , Let be the one-cell mixture density relative to the zero-contrast sign-count baseline : Expanding the exponential kernel above and using moment agreement through gives Define the supported one-cell sign kernel on by with independent coordinates in the first line. Its -cell product prior-predictive law is The Cauchy–Schwarz total-variation bound for one cell and the telescoping inequality for products yield Let map a Poisson observed sample to the array whose -th entry is the pair of counts with and in cell . Under the zero-contrast dense law, let be a regular conditional law. For every supported product prior above, this reconstruction kernel satisfies for every measurable observed-sample event . Since total variation weakly decreases under a common Markov kernel, Since Stirling’s lower bound implies, for every , Therefore Put . Since , , and , Also , and so Combining the last displays proves the total-variation chain in the statement.
Under either product prior, write , where are independent with common law . The support condition makes the restriction to immaterial, and Because , its variance is at most . Independence gives Chebyshev’s inequality, , and then yield For any estimator in the Poissonized experiment, set Use the test On the concentration events above, estimation error smaller than forces the test to select the correct prior index. Markov’s inequality gives Every test also satisfies Consequently The restriction to dense observed laws is legitimate because those laws lie in . In addition, for every , since , , and ; therefore projection of estimators onto weakly decreases squared loss.
The Cai–Low lower bound on , together with the amplitude identity, gives a universal constant such that Indeed, Hence It remains to compare the fixed-sample and Poissonized experiments. Given a fixed- estimator, project it onto , apply it to the first observations of a mean- Poisson sample on the event , and output on the complementary event. On the first event, the first observations have the -fold i.i.d. law from Assumption 1; on the second event, the projected squared loss is at most . Taking infima gives For the lower-tail absorption, fix any . Since , there is such that Increasing the cutoff to ensure , the dense condition gives . The mean- Poisson Chernoff bound gives Since and for , and therefore For and , because . Thus, with , there is a cutoff such that throughout the dense regime above that cutoff.
Choose Then and . The construction, observed-submodel identities, moment-matched priors, concentration bounds, total-variation chain, and one-cell likelihood identity have already been established for every , , and . Since , Since , Finally, , so Using the Poisson-tail comparison and the fixed-sample transfer, Because , this proves All displayed conclusions follow.
∎The lemma supplies the dense-regime lower bound used by Theorem 2. Its first clauses place the submodel inside the fixed-overlap class of Definition 3 and identify the target as an absolute-value average. The prior clauses then combine the approximation gap with a total-variation bound for the induced mixtures, matching the reduction for estimating nonsmooth -type functionals (Cai et al., 2011; Jiao et al., 2018). The final comparison transfers the Poissonized lower bound to the fixed-sample risk up to the displayed lower-tail term, yielding the rate contribution in the superquadratic-sample range.
Verification note
Scope of the formal layer
The formal layer of the paper consists of the fifty-one displayed assumptions, definitions, algorithm, propositions, lemmas, remarks, and theorems that appear in the main text and appendices. Forty of them are attached to a declaration in a machine-checked Lean development: four assumptions, sixteen definitions and one algorithm, ten lemmas, four propositions, one remark, and four theorems. Every displayed theorem, proposition, and lemma belongs to this group, so each such statement is the reader-facing rendering of a Lean declaration whose proof is checked by the Lean kernel, subject to the published inputs recorded below. The proofs printed in this paper are prose renderings of those Lean proofs; they are not themselves the checked objects.
The remaining eleven displayed definitions are presentation-level: they name objects used in the exposition and carry no separate Lean declaration. They are the interval projection, the fixed-law squared risk, the fixed-sample two-sample minimax risk, the equal-propensity normalized embedding, the causal minimax risk, the tensor Jackson convolution operator, the good-pilot event, the pilot-failure contributions, the total-variation distance, the dense likelihood tail, and the one-cell dense likelihood ratio. Each abbreviates a quantity that the checked statements express directly in their own terms; none carries a claim beyond the checked statements that use it.
Artifact, toolchain, and commit
The Lean development for this paper is the directory CausalSmith/Stat/STAT_DiscreteOptimalValueMinimaxMatched_Research of the CausalSmith repository, at commit 881a4d27aaf42f02b1f71128db9f0c6d2184c790. It is built with Lean 4 toolchain leanprover/lean4:v4.33.0 and Mathlib at revision db584cd6d46c92f209a44c0f1c829460d327499d, through the build command lake -d CausalSmith build. The declarations backing the displayed statements are named in the interactive companion to this paper. A source scan of the development finds no sorry and no added axiom, and the axiom audit of its main theorems reports only the standard Lean axioms propext, Classical.choice, and Quot.sound.
Objects anchoring each layer
The model and identification layer is anchored by Assumption 4, Assumption 1, Assumption 2, Assumption 3, Definition 3, Definition 6, Definition 2, Definition 4, Definition 5, Definition 9, Proposition 1, and Proposition 2. Together these statements specify the fixed-overlap observed experiment, the full-data causal completion class, and the equality between the observed optimal-regression value and the causal optimal-treatment value under the stated conditions.
The estimation layer is anchored by Definition 10, Definition 12, Algorithm 1, Definition 13, and Theorem 1. These statements define the empirical-ratio benchmark, the Jackson smoothing polynomial, the all-data Rao–Blackwellized estimator, the fixed-law squared risk, and the uniform upper bound for the estimator over the fixed-overlap observed-law class.
The lower-bound and comparison layer is anchored by Definition 14, Definition 15, Definition 16, Proposition 3, Definition 24, Lemma 9, Lemma 10, and Theorem 2.
The combined minimax conclusions are anchored by Theorem 3, Theorem 4, Proposition 4, Definition 18, and Remark 1. These statements give the matched fixed-sample rate, the sequence-level consistency and parametric-rate characterizations, the comparison with the predecessor bracket on the same formal class, and the stated route for the sharp-overlap-constant question.
Published inputs the certification is conditional on
Two analytic conclusions enter as published inputs rather than as checked steps: the polynomial-approximation and moment-matching construction for of Cai et al. (2011), and the Poissonized two-sample minimax lower bound of Jiao et al. (2018). The theorem-local footnotes identify the precise published conclusion each affected statement uses. The Jackson approximation material of DeVore et al. (1993) and the testing and minimax comparison tools of Le Cam (1986); Tsybakov (2009) supply the classical background for the constructions, whose uses in this paper are checked. Using those two published inputs, the Lean development certifies the lower-bound and matched-frontier results through the dependencies identified in the theorem-local footnotes and checks the remaining steps internally.
Proofs of the main results
Write . For a four-vector , set so that on , as in Definition 8.
Fix . The coordinates of are observed atom masses, hence for every . Moreover If , Definition 3 gives and therefore If , nonnegativity and force , so the same two inequalities hold. Thus by Definition 7.
For each arm , define the totalized cell mean On an occupied cell, the cone inequalities from the previous step imply and agrees with the outcome mean in Definition 3. Hence, for both arms, and therefore On a null cell all observed atom masses vanish, and both sides of this last identity are zero. Taking the maximum over gives with the displayed value unchanged by the permitted null-cell choice of . Summing over and using Definition 4 yields
Let . For every arm , . If , then . If , the denominator is positive and Taking the maximum over the two arms gives
It remains to prove the Lipschitz estimate for the global extension. For nonnegative , put For two quadruples, write When both totals are positive, split according to whether and . In the low–low case, In the low–high case, The second summand is nonnegative, and . Using gives and the lower bound follows from The high–low case is the same calculation with the two quadruples interchanged. In the high–high case, and because Together with and , this gives If one total is zero, nonnegativity forces that quadruple to vanish, and the bound follows from and . Applying this estimate to each arm, after merely permuting the four coordinates, yields Finally, which proves the asserted global Lipschitz property.
Define Construct a full-data mass function by where for and for . Since , every displayed mass is nonnegative. Also Thus these masses define a full-data law.
The construction gives consistency because atoms with have mass zero. Its observed margin is , since Using the observed-law factorization supplied by Definition 3, the potential-outcome atom obtained after summing over the observed outcome factors as Consequently, conditionally on , treatment is Bernoulli with parameter , independent of the joint potential-outcome vector, and are independent Bernoulli variables with parameters and . The observed margin is , so its overlap properties are exactly those of the given observed law; together with consistency and conditional exchangeability, Definition 6 places in .
Conversely, if , then Definition 6 includes the assertion that its observed margin satisfies the observed fixed-overlap model restrictions. Hence The preceding conclusions use the nonnegative-domain extension from Definition 8, so they are exactly the cell-overlap, value, global-bound, Lipschitz, completion, and observed-margin assertions stated above.
Let For and , write and define the totalized observed arm regression by Also write For and , put and define the totalized potential-outcome regression by
By Definition 2, for every , The support component of Definition 6 makes and binary, so the event is the disjoint union, over , of the events Finite additivity therefore gives Summing over and gives
Fix and , and suppose The consistency component of Definition 6 gives, for each , and hence The exchangeability component of Definition 6, applied to , , and outcome value , gives The double sum is . Dividing by the positive product yields
The support component of Definition 6 gives the finite partition . Expanding the outer expectation in Definition 9 over this partition gives By the defining convention for on null cells, this is equivalently Since , Definition 6 places in the observed-law class of Definition 3. Choose the representing parameters , , and from that representation. Summing its displayed atom factorization over and gives and summing over at fixed gives If , fixed overlap in Definition 3 gives , and hence If , the cell contribution to the value is zero for any permitted choice of . Therefore Definition 4 gives
Using , fix . If , then If , the overlap component of Definition 6 gives Since , Also so The identity from the previous step applies for and , giving Thus Summing over proves
Choose the universal tuning supplied by Lemma 8. Thus is the universal self-normalized Poisson pilot-radius constant, so , , and . Fix , let be the cell constant supplied by Lemma 8, and put Define Then . For , write Since and , Fix , , and . Let and let be the covariate-cell mass. By Proposition 1, The four-cell vectors partition the one-observation law, so Consequently
Consider the bounded-alphabet branch . Since and , this branch has . By Algorithm 1, Assume first that . For and , define with the convention from Definition 10, and put If , then and almost surely. If , the fixed-overlap representation in Definition 3 gives and the conditional success mean in that arm and cell is . Define The totalized-ratio identity is Conditional on , the arm count is binomial with success probability . For with , for , while the left-hand side is for . Since and , the centered residual has conditional second moment at most one; hence Also Let Expanding into diagonal and ordered-pair contributions gives The elementary inequality , applied with , and the condition give Using and , we obtain Moreover, The maximum map is one-Lipschitz in the sup norm, and Cauchy–Schwarz over the four pairs yields Therefore Since , , and ,
If , the empirical ratios lie in , their weights sum to one, and therefore Together with , this gives Also so Since , the desired bound follows throughout the bounded-alphabet branch.
Now suppose that and In this branch, Algorithm 1 outputs the constant estimator . Since , The scale condition gives and hence Since , the asserted inequality holds in this branch.
It remains to treat the Jackson branch Set Then . In the uncapped finite-Poisson marked experiment, draw then draw i.i.d. observations from , each with an independent fair pilot/evaluation mark. Let be the clipped cell statistic computed from the resulting cell- pilot and evaluation counts, and define By Lemma 8, the cell blocks are independent and, for every , Projection onto is nonexpansive for squared distance from , so Using independence, Since , Cauchy–Schwarz gives Summing the variance bounds gives The elementary logarithmic absorption bounds hold for every integer : the one-variable bound for gives the first inequality after multiplication by , and the second after squaring and multiplication by . With and , these imply and the scale inequality gives Consequently,
We pass from the uncapped comparison experiment to the fixed-sample estimator in Algorithm 1. Let be the product law of independent fair marks. For define Define the capped Poisson-prefix average In the present Jackson branch, the all-data estimator is the fair-mark average Jensen’s inequality for the mark average gives A second application of Jensen, now to the Poisson-prefix average, gives On , the prefix has the corresponding uncapped finite-Poisson marked law restricted to prefixes of length at most , so the nonoverflow contribution is bounded by On , the capped statistic equals , and , so the overflow contribution is at most Thus To bound the Poisson overflow term, let Since , both and are positive. For every nonnegative Chernoff tilt , the centered Poisson moment-generating function gives Consequently, for every , Markov’s inequality yields Taking gives the Bennett form The elementary inequality implies Now choose Then so For the present values, and therefore Thus Finally, follows from , and hence Using , Combining this overflow bound with the uncapped bound from the previous step, In the present branch , hence The three branches cover every , , and , giving the stated uniform bound.
Fix , and put The sampling condition in Assumption 1 supplies the product-sampling convention used in Definition 5.
For , define By Definitions 16 and 2, Hence The vector lies in , since and Set for every cell, and set On cells with , these quantities lie in , and the four identities in Equation 4 are exactly the factorization required in Definition 3. On cells with , nonnegativity gives , so all four atom masses vanish and the same factorization holds. Moreover, so whenever , Because , this propensity lies in . Thus .
Write By Definition 4, A null cell contributes zero. On a cell with , For each , Summing over and using gives
Define the single-observation kernel from a paired draw by Under one paired draw from , Therefore and similarly These are the atom masses in Equation 4, so the one-observation kernelized paired draw has law .
Define the fixed- product kernel on by For , this is the empty product kernel; for , it applies independently in each coordinate. Since the fixed paired-sample law is and the one-coordinate kernelized law is , finite-product independence gives This is the explicit fixed- product Markov kernel asserted in the proposition.
Let be any measurable observed-sample estimator. Since the observed sample space is finite, is bounded. Define the paired-sample pullback For each fixed paired sample , Equation 5 and Jensen’s inequality for give Integrating under and using Equation 6 yields Taking the supremum over , and using the admissibility proved above, gives The estimator is admissible for the paired problem in Definition 15; the zero estimator and the fixed pair chosen at the start give the required nonempty estimator and parameter classes. Hence Taking the infimum over all observed-sample estimators , as in Definition 5, gives
Combining the observed-model membership, the displayed formulas for , , and , the product-kernel identity, and the minimax transfer proves every asserted conclusion for the fixed pair . Since the pair was arbitrary, the proposition follows.
Fix . For an alphabet size , set where is used only for sample-size logarithms. Let be the fixed-sample two-sample risk of Definition 15. Also define the Poissonized two-sample risk where, under , the two histograms have independent coordinates and the infimum ranges over measurable histogram statistics whose squared loss is integrable for every and whose displayed worst-case risk is finite.
Lemma 10, using the moment-matching priors of (Cai et al., 2011), supplies an integer and a constant such that The Poissonized lower bound of (Jiao et al., 2018), with constants and , gives a universal such that For fixed-sample estimators this yields Indeed, from any fixed-sample estimator , construct a histogram statistic as follows. Let On , take the conditional expectation of applied to ordered -samples drawn from the two histograms; on , output . On , the conditional reconstruction has the fixed -sample law, so Jensen’s inequality gives the fixed-sample risk. Since , the squared loss on is at most . The Chernoff bound for gives and hence the displayed fixed-sample lower bound follows after taking the infimum over .
Choose so that and choose so that, for every , Set and
For every and , Put In every cell , take , , and . Under , set ; under , set . Since , the vector with coordinates lies in . The fixed range gives , so the displayed propensities satisfy Assumption 4. Finally places all arm-specific means in . Thus both laws belong to , and The one-observation chi-square divergence is so tensorization and give because . Therefore For two model points with target separation and product-law total variation at most , thresholding any estimator at the target midpoint and using gives a worst-case risk at least . With ,
If , then , and therefore
Assume . If , then . Since ,
Assume , , and . The dense lower bound from Lemma 10 gives Since and ,
It remains to treat , , and . Then Because , also , and hence
Suppose first that Define The preceding inequality implies . Since , this gives , hence , and therefore The ceiling definition also yields the lower bound Because and , this implies , and monotonicity of the logarithm gives The upper ceiling estimate is just as important: since , The denominator below is positive because . Consequently, and the two remaining gate estimates are the second following by dividing by the positive product . The fixed-sample bound at alphabet yields and the choice of gives Embedding into by padding the remaining coordinates with zero preserves the distance, so By Proposition 3, and consequently The saturated inequality gives , hence Since ,
Finally suppose that The fixed-sample bound at alphabet gives Since , The choice of therefore yields Also , so Using and Proposition 3,
The displayed alternatives cover every and . Thus the constant satisfies for all and .
∎Assumption 1 identifies the sampling law with the canonical product law. Apply Theorem 1 with this sampling law and choose its universal Jackson tuning constants , , and cutoff . Fix . The same result gives a constant such that, for all , , and , The lower-bound input is Theorem 2, using the moment-matching priors of (Cai et al., 2011) and the Poissonized minimax lower bound of (Jiao et al., 2018). It gives a constant such that, for all and , Set Then . For later comparison, when and , , and hence
The preceding lower bound gives For the middle inequality, define for each measurable observed-data estimator By Definition 5, where the infimum is over the same measurable sample maps. The estimator in Algorithm 1 is one admissible map, and squared-error risks are nonnegative: Therefore evaluating the infimum at this estimator gives
It remains to bound this particular worst-case risk. If the set over which the displayed supremum is taken is empty, the real-supremum convention gives and . Otherwise, for every , the upper bound from Theorem 1 gives where the second inequality uses . Taking the supremum over yields
Finally fix a measurable observed-data estimator and define its causal worst-case risk by The two attained-risk sets are equal. Indeed, if , then Proposition 1 gives and Proposition 2 gives Thus each causal risk value is an observed-law risk value: Conversely, if , then Proposition 1 supplies with and Proposition 2 again gives Hence the corresponding observed-law risk is attained in the causal class: Therefore for every measurable observed-data estimator . Taking the infimum over the common estimator class in Definitions 5 and 17 gives Together with the three inequalities above, this proves the result.
Fix . The published moment-matching conclusion of (Cai et al., 2011) and the published Poissonized lower bound of (Jiao et al., 2018) enter through Theorem 3 and Proposition 4. By Theorem 3, there are constants with such that, for every and , where The upper comparison uses the Jackson-factorial worst-case risk term displayed in Theorem 3.
Let satisfy for every . Define All displayed denominators are positive. We prove which is exactly . Suppose first that . Then eventually , hence , so . For all sufficiently large , Indeed, if , then If , set Then Since eventually and , so . Consequently, Because , this gives .
Conversely, suppose . For all , If , then , and therefore If , then , whence and therefore Since , we get , and then
The two-sided comparison with gives For the forward implication, the lower comparison gives for all . For the reverse implication, the upper comparison gives for all . Combining this equivalence with the preceding display proves observed consistency.
We next prove the observed parametric boundary. Suppose eventually for constants . The upper parametric bound and the lower frontier bound imply eventually. For all sufficiently large , the alternative would give and hence , a contradiction. Thus eventually , and so For , Therefore and hence is eventually bounded. This is .
Conversely, assume . Then there are a real number and an index such that for all . Enlarging over the finite set , choose an integer such that This finite enlargement is the passage from an eventual bound to a single global alphabet bound.
The bounded-alphabet sequence clause of Proposition 4 applied with this supplies constants with such that, for all sufficiently large , Together with the preceding implication, this proves the observed parametric equivalence.
It remains to transfer the two equivalences to the causal risk. The equality supplied by Theorem 3 gives, for every , Hence and observed consistency gives
The same equality transports the parametric-rate property itself: Indeed, the same eventual constants bound both sequences after replacing one risk by the other in the displayed equality. Composing this equivalence with the observed parametric equivalence proves the causal parametric boundary.
∎The published moment-matching construction and absolute-value approximation rate of (Cai et al., 2011), together with the published Poissonized two-sample lower-bound argument of (Jiao et al., 2018), are the literature inputs used by Theorem 3. Apply Theorem 3 and choose the tuning constants it supplies. Write them as , , and . They satisfy For each fixed , Theorem 3 gives constants and , with such that, for every and , with one has
1. Fix and . The logarithms and are positive. Since we have Define We prove If , the right-hand side is , while the left-hand side is at most . Suppose . If , then , so If , put Then , and gives . For every , where the last inequality is equivalent to . Hence so , and therefore Thus, throughout the case , Combining this with the square-root bound gives and hence the displayed comparison follows. Consequently, For the upper bound, gives , and therefore Taking and proves the matched-frontier assertion in the proposition.
2. Fix , , , and an observed sample with . Algorithm 1 selects its empirical-ratio branch whenever . Hence, for that sample,
3. Fix and a sequence such that for every and for some finite . Then . Define Since and , so For every , one has . With the bound gives , and therefore Thus the minimum in the matched frontier equals . Monotonicity of the logarithm gives and, since , Therefore, for every , Similarly, These inequalities hold for every , hence for all sufficiently large .
∎References
- Athey, Susan and Wager, Stefan (2021). Policy Learning with Observational Data. Econometrica. doi
- Cai, T. Tony and Low, Mark G. (2011). Testing Composite Hypotheses, Hermite Polynomials and Optimal Estimation of a Nonsmooth Functional. The Annals of Statistics. doi
- Chakraborty, Bibhas and Murphy, Susan A. and Strecher, Victor J. (2010). Inference for Non-Regular Parameters in Optimal Dynamic Treatment Regimes. Statistical Methods in Medical Research. doi
- Chen, Qizhao and Austern, Morgane and Syrgkanis, Vasilis (2023). Inference on Optimal Dynamic Policies via Softmax Approximation. . doi
- DeVore, Ronald A. and Lorentz, George G. (1993). Constructive Approximation. Springer-Verlag.
- Han, Yanjun and Jiao, Jiantao and Weissman, Tsachy (2018). Local Moment Matching: A Unified Methodology for Symmetric Functional Estimation and Distribution Estimation under Wasserstein Distance. Proceedings of the 31st Conference on Learning Theory. arXiv
- Hirano, Keisuke and Porter, Jack R. (2009). Asymptotics for Statistical Treatment Rules. Econometrica. doi
- Hirano, Keisuke and Porter, Jack R. (2012). Impossibility Results for Nondifferentiable Functionals. Econometrica. doi
- Jiao, Jiantao and Han, Yanjun and Weissman, Tsachy (2018). Minimax Estimation of the \(L_1\) Distance. IEEE Transactions on Information Theory. doi
- Jiao, Jiantao and Venkat, Kartik and Han, Yanjun and Weissman, Tsachy (2015). Minimax Estimation of Functionals of Discrete Distributions. IEEE Transactions on Information Theory. doi
- Kitagawa, Toru and Tetenov, Aleksey (2018). Who Should Be Treated? Empirical Welfare Maximization Methods for Treatment Choice. Econometrica. doi
- Le Cam, Lucien (1986). Asymptotic Methods in Statistical Decision Theory. Springer-Verlag. doi
- Luedtke, Alexander R. and van der Laan, Mark J. (2016). Statistical Inference for the Mean Outcome under a Possibly Non-Unique Optimal Treatment Strategy. The Annals of Statistics. doi
- Manski, Charles F. (2004). Statistical Treatment Rules for Heterogeneous Populations. Econometrica. doi
- Murphy, Susan A. (2003). Optimal Dynamic Treatment Regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology). doi
- Qian, Min and Murphy, Susan A. (2011). Performance Guarantees for Individualized Treatment Rules. The Annals of Statistics. doi
- Robins, James M. (1986). A New Approach to Causal Inference in Mortality Studies with a Sustained Exposure Period---Application to Control of the Healthy Worker Survivor Effect. Mathematical Modelling. doi
- Rosenbaum, Paul R. and Rubin, Donald B. (1983). The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika. doi
- Tsybakov, Alexandre B. (2009). Introduction to Nonparametric Estimation. Springer. doi
- Valiant, Gregory and Valiant, Paul (2017). Estimating the Unseen: Improved Estimators for Entropy and Other Properties. Journal of the ACM. doi
- van der Laan, Mark J. and Luedtke, Alexander R. (2015). Targeted Learning of the Mean Outcome under an Optimal Dynamic Treatment Rule. Journal of Causal Inference. doi
- Wei, Haoyu (2025). Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness. . doi
- Whitehouse, Justin and Chen, Qizhao and Austern, Morgane and Syrgkanis, Vasilis (2025). Inference on Optimal Policy Values and Other Irregular Functionals via Softmax Smoothing. . doi
- Wu, Yihong and Yang, Pengkun (2016). Minimax Rates of Entropy Estimation on Large Alphabets via Best Polynomial Approximation. IEEE Transactions on Information Theory. doi
- Zeng, Zhenghao and Balakrishnan, Sivaraman and Han, Yanjun and Kennedy, Edward H. (2024). Causal Inference with High-Dimensional Discrete Covariates. . doi
Comments on earlier versions
Anchored to: