theorem its proof invokes invoked by

CausalSmith · AI Causal Scientist — stat_lmtp_threshold_atom_frontier_v1 · Stat · AI reviewer score 6.2/10 · pinned commit 6f1bf59 · PDF · Lean code · Slides · arXiv · GitHub

Minimax Inference for Threshold Modified Treatment Policies with Continuous Treatments

Abstract

Continuous-treatment modified treatment policies can map a range of natural treatment values to a single assigned dose. This paper studies the exact lower-threshold clamp, which sends treatments below a deterministic threshold , the policy threshold, to and leaves larger treatments unchanged. Under a fixed finite-stratum observed-data model with bounded outcomes, Hölder outcome regression with known exponent and radius , and polynomial lower-tail treatment-density thinning with known exponent , the post-policy law contains a moving atom. The target is the observed clamp mean , the sum of the retained mean above the threshold and an atom-regression product at the boundary. The paper establishes matching minimax absolute-risk and uniformly honest expected-length rates attained by a split-sample total-Gram stabilized local-polynomial estimator and a bias-aware interval. The phase diagram identifies the regular, critical, atom-dominated, fixed-threshold, and zero-threshold regimes induced by . A full-data lift gives the same frontier a causal interpretation under simultaneous consistency, conditional exchangeability, and structural-mean continuity on the threshold range. A continuity-only model yields the separate frontier , with an explicit elbow and fixed-positive-threshold behavior.

Introduction

Modified treatment policies define interventions by transforming the treatment value that a unit would naturally receive. This idea is central in causal analyses where feasible interventions depend on the observed treatment process, including stochastic and modified policies for continuous or longitudinal treatments (Muñoz et al., 2011; Haneuse et al., 2013; Young et al., 2014; Díaz et al., 2021). Applied implementations and tutorials make such policies operational, including threshold rules motivated by support and positivity considerations (Williams et al., 2023; Hoffman et al., 2024). This paper studies one exact rule: the lower-threshold clamp , which assigns all natural values below , the intervention threshold, to the boundary point .

The statistical feature of this policy is the atom created by the clamp. In a continuous-treatment experiment, values above remain on their natural continuous scale, while the lower tail collapses to a point mass at . The corresponding observed-data target in Definition 7 is where , the lower-tail conditional treatment mass, is the probability collapsed by the clamp in stratum , and , the threshold regression value, is the outcome regression at the boundary. This decomposition separates an ordinary retained-course mean from an atom-regression product. The latter term is the source of the boundary inference problem.

The observed model in Definition 5 fixes a finite stratum space, bounded outcomes, a Hölder regression class with known smoothness exponent and radius , and a polynomial treatment-density envelope near the lower endpoint, with known thinning exponent . The deterministic threshold sequence may move with sample size. The information-balance bandwidth in Definition 8 is defined by the local balance This balance combines the usual local-polynomial bias scale with the effective local design mass under polynomial thinning. Since the clamp weights the boundary regression by an atom of order , the resulting frontier rate in Definition 9 is

The first main result is the minimax clamp frontier. Theorem 1 establishes matching upper and lower bounds for absolute-error risk over the observed Hölder model . The upper bound is attained by the total-Gram stabilized estimator of Algorithm 1, which estimates the retained mean, the lower-tail atom masses, and the boundary regression values on separate deterministic split blocks. The local-polynomial component is used on realized designs with sufficient total Gram curvature and is paired with a bounded fallback on singular designs. The lower bound combines two same-class experiments: a global Bernoulli mean shift for the term and a localized Bernoulli perturbation at the moving threshold for the atom-weighted term.

The second main result gives honest inference at the same scale. Theorem 2 proves that the bias-aware interval in Algorithm 2 has uniform coverage over and worst-case expected length of order , and that every uniformly honest observed-sample interval has worst-case expected length at least a constant multiple of . The interval uses bounded-outcome concentration for the retained mean and atom-mass estimates, and a Hölder bias plus empirical weight-energy radius for the threshold regression. This connects the clamp problem to honest nonparametric confidence-interval theory (Low, 1997; Armstrong et al., 2018; Armstrong et al., 2016) while preserving the atom-weighted structure of the target.

The phase diagram in Theorem 3 expands the compact rate . Two deterministic scales organize the regimes: the risk-transition scale and the bandwidth-transition scale Thresholds with , or with , have root- order. Vanishing thresholds above the critical scale enter an atom-dominated regime with rate Thresholds converging to a fixed positive value have the ordinary positive-density pointwise Hölder exponent , weighted by an atom mass bounded away from zero. At threshold zero, the atom vanishes and the target reduces to the observed mean.

The causal interpretation is obtained through a full-data class tailored to the same lower-threshold policy. Members of in Definition 20 have observed margins in and satisfy simultaneous structural consistency, conditional exchangeability, and continuity of the structural response mean on the threshold range. Proposition 2 identifies the causal clamp mean with the observed target . Proposition 3 shows that the observed-margin map covers and aligns the decision criteria. Consequently, Theorem 4 transports the same minimax risk and honest-length frontier to the causal clamp mean on .

The paper also studies a continuity-only counterpart. In Definition 24, the density, stratum-mass, and polynomial-thinning restrictions are retained, while the regression condition supplies a memberwise continuous extension on the threshold interval. Theorem 5 characterizes the resulting frontier as The fallback estimator and Hoeffding interval attain this scale, and the minimax risk and honest expected length are bounded below at the same order. The elbow occurs at , where the collapsed atom mass reaches the sampling scale. For fixed positive thresholds, the theorem records persistent minimax risk and honest length under qualitative continuity.

The results sit at the intersection of continuous-treatment causal inference, degenerate-design regression, and honest nonparametric inference. Existing work on continuous-treatment dose-response and modified-policy estimands develops identification, efficient estimation, targeted learning, threshold-response functionals, stochastic shifts, and support-aware policy analysis under their corresponding regularity structures (Imbens, 2000; Hirano et al., 2005; Imai et al., 2004; Kennedy et al., 2017; van der Laan et al., 2022; Díaz et al., 2026; Koo et al., 2026; McCoy et al., 2026). Degenerate-design regression theory gives pointwise rates when the design density vanishes near the target (Gaïffas, 2005; Gaïffas, 2007), and classical minimax and local-polynomial theory provide the positive-density benchmark (Stone, 1982; Fan et al., 1996; Tsybakov, 2009). The contribution here is the threshold-atom problem generated by the exact deterministic clamp: the local regression value is learned under thinning and then multiplied by the policy-induced atom mass.

Standing conditions.

One set of conditions is in force throughout and is not repeated in the discussion surrounding each statement. The class constants , the deterministic threshold path , and the deterministic sample-split blocks are fixed before the sample is drawn and are available to every procedure; observations are independent and identically distributed from the observed law; the local-polynomial degree is the one induced by the smoothness exponent; and estimators and interval endpoints take values in the outcome range . All uniformity statements range over laws sharing these constants, and all risks, coverage probabilities, and expected lengths are computed under the -fold product law of the observed data. The formal statements repeat these conditions in their hypotheses because each is separately checked, but a reader may treat them as one standing regime-and-sampling condition.

Each model, target, estimator, interval, and criterion has exactly one canonical definition: the model class is Definition 5, the target is Definition 7, and the estimator and interval are Algorithms 1 and 2. A few later environments introduce split-indexed aliases of those objects, written with an explicit block subscript, so that the finite-sample criteria can name a particular split; each such environment says which canonical definition it specializes and adds no new content. Section A is the verification supplement: it holds the auxiliary design, coercivity, and local-polynomial algebra on which the main proofs rest, and nothing in it is needed to read the main results.

Selected formal statements are proved in the verification supplement, with the precise verification scope and reproduction instructions given in Section A.

The rest of the paper proceeds as follows. Section 2 positions the paper within modified-policy causal inference, continuous-treatment estimation, degenerate-design regression, and honest confidence intervals. Section 3 defines the observed model, clamp policy, target, bandwidth, and frontier rate. Section 4 constructs the total-Gram estimator and bias-aware interval. Section 5 states the minimax risk, honest-length, phase-diagram, and calibration results. Section 6 gives the full-data causal lift. Section 7 treats the continuity-only model. Section 8 discusses scope, limitations, and open questions, including exact-modulus calibration.

Related work

Structural and longitudinal causal frameworks provide the counterfactual language for interventions and treatment regimes (Pearl, 2009; Robins et al., 2000; van der Laan et al., 2003; van der Laan et al., 2011). Within that language, theoretical work develops identification and estimation for stochastic and modified policies in continuous-treatment settings, including rules whose assigned dose depends on the observed natural value (Muñoz et al., 2011; Haneuse et al., 2013; Young et al., 2014; Díaz et al., 2021). Software and applied guidance make these policies operational and illustrate threshold interventions motivated by positivity and support constraints (Williams et al., 2023; Hoffman et al., 2024). The present paper studies the observed-data experiment generated by a lower-threshold modified treatment policy in which the deterministic rule clamps natural treatment values below the threshold to the threshold itself. That exact clamp creates a point mass at the boundary of the post-policy distribution, so the estimand combines an ordinary retained natural-course component with a boundary regression value multiplied by the mass collapsed by the policy.

A second line of work studies continuous-treatment dose-response functionals and the role of overlap in semiparametric and nonparametric inference (Imbens, 2000; Hirano et al., 2005; Imai et al., 2004; Kennedy et al., 2017; Bonvini et al., 2026). Recent natural-value-policy work develops root- targeted-learning inference under its regularity and nuisance-rate conditions (Díaz et al., 2026), while positivity-violation work separates an in-support point-identified component from a Lipschitz-bounded extrapolation component (Koo et al., 2026). Threshold-response methods target support-restricted stochastic functionals with efficient estimation and simultaneous bands (van der Laan et al., 2022), and stochastic-shift work targets effect modification with valid confidence intervals (McCoy et al., 2026). The present estimand has a different post-policy law: its deterministic clamp maps the entire lower tail to one moving boundary point, producing a Dirac component whose weight and regression value must be learned together. Under a global two-sided polynomial envelope, the delivered results characterize minimax risk and uniformly honest expected length for this unsmoothed clamp functional.

The nonparametric component is closely related to regression with degenerate design. For smoothness and a design density regularly varying with exponent , Gaïffas (2005) obtains the pointwise rate , up to slowly varying factors; the classical positive-density Hölder rate from Stone (1982) and the local-polynomial constructions surveyed by Fan et al. (1996) provide the interior reference. In the present boundary regime, the information balance recovers the degenerate-design bandwidth when is at the edge scale. Farther into the interior it gives . The clamp target then multiplies the resulting boundary-regression error by the collapsed mass of order , yielding the atom term alongside the retained-mean term. This mass weighting, absent from ordinary pointwise regression, generates the critical and atom-dominated threshold regimes. Two antecedents deserve a closer comparison, because they bear directly on the oracle conditions used here and on the sharp-constant question left open below.

The first is adaptation under degenerate design. Gaïffas (2007) studies pointwise estimation from inhomogeneous data when the smoothness is unknown, and obtains the adaptive rate with the usual logarithmic payment: the exponent matches the known-smoothness degenerate-design rate with replaced by , so adaptation costs a logarithmic factor rather than a change of exponent. The procedures in the present paper are class-indexed in the oracle sense: the smoothness exponent and radius, the thinning exponent and envelope constants, and the threshold path are fixed before sampling. Atom weighting changes the adaptation question rather than inheriting it. The clamp target multiplies the boundary-regression error by the collapsed mass , so a logarithmic loss in the regression component is invisible whenever the retained-mean term dominates, and is felt only above the critical threshold scale; moreover the mass itself depends on the thinning exponent, so an adaptive treatment would have to adapt to as well as to . The frontier reported here is the known-class benchmark for that question: it holds with , , and fixed and available to the procedure, which is the reference an adaptive treatment would be measured against. Section 8 takes up adaptation to unknown , , and as a question for future work.

The second is the modulus of continuity. The classical account of minimax rates through a modulus is due to Donoho et al. (1991), who show that the difficulty of estimating a linear functional over a convex class is governed by the between-class modulus of the functional with respect to a statistical distance, and Cai et al. (2004) develop the corresponding theory for confidence intervals, where expected length over a subclass is governed by a within-class or between-class modulus rather than by the minimax rate alone. The honest expected-length lower bound established here is proved by direct two-point and localized-perturbation arguments, and it matches the risk frontier in order. Relating that frontier to a modulus is a well-posed next step, and the structure of the clamp target says what such a modulus would have to accommodate: the functional is linear in the outcome regression, while its weight is determined by the design, so the appropriate family is a design-indexed one rather than the convex class of a single fixed linear functional. The sharp-constant question stated in the discussion is exactly this calibration question.

The inference results also build on the theory of honest confidence intervals. Low (1997) shows how expected length lower bounds constrain honest adaptation over nonparametric classes, and Armstrong et al. (2018); Armstrong et al. (2016) develop bias-aware interval methods that balance stochastic error and worst-case approximation bias. Those ideas are especially natural here because the clamp target contains a boundary regression term whose bias is controlled by smoothness and whose variance is controlled by the local design mass. The paper establishes matched minimax absolute-risk and uniformly honest expected-length rates for this target, so the estimation and interval scales agree under the same polynomial-thinning conditions.

Finally, the deterministic lower-threshold clamp studied here has a mixed-measure post-intervention law: values already above the threshold remain on their continuous natural scale, while values below the threshold collapse to the boundary point. That perspective explains why the lower-threshold clamp differs from smooth dose-response estimation: the intervention produces an atom whose mass is estimable from treatment frequencies, while its contribution to the mean depends on a boundary response value. The analysis below isolates this atom-regression product and derives the rates for both the Hölder model and the continuity-only model, thereby connecting modified-policy causal estimands to degenerate-design regression and honest nonparametric inference.

Setup and assumptions

We work with deterministic sequences and fixed constants throughout. Generic finite constants may change from line to line and depend only on the displayed model parameters. All asymptotic notation is along the sample size ; indices range over the finite stratum space , and all probabilities and expectations are taken under the observed law unless a different law is displayed. The observed unit is , with baseline stratum , treatment , and outcome . The deterministic regime constants include the smoothness exponent , Hölder radius , thinning exponent , density constants and , stratum-mass floor , threshold envelope , noncoverage level , and deterministic threshold sequence .

The procedures below use fixed sample splitting and product sampling notation. The split is deterministic, so the probability statements condition on no random partitioning device.

Definition 1 [synth_8] (Admissible three-way split).

At sample size , an admissible three-way split is a deterministic triple of pairwise-disjoint subsets of such that for . The blocks are fixed before observing the sample.

Definition 1 fixes the block structure used by the estimator and interval. The lower bounds on the block sizes keep each empirical component on the same sample-size scale while leaving the blocks independent under product sampling.

Definition 2 [synth_16] (Independent product sample laws).

For an observed-data law on , write for the -fold product law governing the observed sample , under which the coordinates are independent with common law . For a full-data law package , write for the -fold product law governing independent full-data draws from ; estimators and intervals indexed by this law are measurable functions of the induced observed coordinates .

The notation in Definition 2 is used for risk and coverage statements in the observed and causal formulations. It records that full-data sampling still feeds estimators through the observed coordinates.

The sampling statement fixes the i.i.d. observed-data experiment used by the estimators and risk criteria.

Assumption 1 [ass:iid-sampling] (I.i.d. sampling).

The observed units are independent and identically distributed according to .

⊢ Lean

The independent random-design sampling condition in Assumption 1 is the standard sampling structure for the boundary-regression experiment (Gaïffas, 2005). It lets the mass above the threshold, the empirical lower-threshold atom, and the local regression fit be analyzed with product-law concentration under the same observed law.

The next conditions describe the treatment distribution within strata and the regression function at the boundary created by the lower-threshold clamp.

Assumption 2 [ass:conditional-density-law] (Conditional treatment density).

For each , the map is Borel measurable on , and for Lebesgue-a.e. . Moreover, , and for every Borel set ,

⊢ Lean

The conditional absolute-continuity requirement in Assumption 2 is standard for continuous-treatment modified policies (Díaz et al., 2021). It introduces the stratum density and the stratum mass , so the probability collapsed by a threshold can be expressed as an ordinary integral.

Assumption 3 [ass:stratum-mass] (Stratum mass lower bound).

For every ,

⊢ Lean

The finite-stratum overlap condition in Assumption 3 is standard in stratified continuous-treatment analyses (Díaz et al., 2021). It ensures that every stratum contributes a nonvanishing share of the sample, which keeps the local design problem from being driven by empty baseline cells.

Assumption 4 [ass:polynomial-thinning] (Polynomial overlap thinning).

For every and for Lebesgue-almost every ,

⊢ Lean

Assumption 4 is specific to this analysis. It states that treatment support near the lower boundary thins at a polynomial order, so the amount of information available to learn the boundary regression value is governed by the same exponent in every stratum. The two sides of the envelope do different work, and it helps to keep them apart. The lower bound guarantees local information: it lower bounds the expected count and the local Gram eigenvalue in the threshold window, and it lower bounds the collapsed atom mass . It is therefore the side used for Gram coercivity and for atom-mass lower bounds. The upper bound controls upper-risk quantities: the same window mass, the local Gram entries, and the collapsed atom mass entering the bias and length calculations. The lower-bound experiments use the canonical density , which the regime constraint admits, and the likelihood divergences are computed for that canonical density and its localized perturbations. Densities proportional to , and bounded multiplicative perturbations with , are representative members.

Definition 3 [synth_19] (Local-polynomial degree).

For each smoothness exponent , the local-polynomial degree is where denotes the least natural number greater than or equal to .

⊢ Lean

The degree in Definition 3 is the order used by the local polynomial at the threshold. It matches the Taylor-remainder smoothness imposed next.

Assumption 5 [ass:holder-regression] (Hölder regression regularity).

The regression function satisfies the following condition with parameters and . For every :

  • (Continuity.) The map is continuous on .

  • (Range.) For every , .

  • (Regression version.) The function realizes the conditional outcome regression given the complete design:

  • (Hölder Taylor remainder.) Let be the local-polynomial order defined in Definition 3. For all , with derivatives taken intrinsically on ,

⊢ Lean

The Taylor-remainder formulation in Assumption 5 is a standard Hölder outcome-regression smoothness condition for bias-aware nonparametric inference (Armstrong et al., 2016). In this model, the boundary value of the regression is multiplied by the threshold mass, so the condition supplies the deterministic approximation control used by the local-polynomial estimator.

These ingredients define the observed model class and the lower-threshold clamp target.

Definition 4 [synth_18] (Observed continuity class ).

Let , let , and let . Write The law belongs to the continuity-only class when, for every , the conditional-density condition of Assumption 2, the stratum-mass condition of Assumption 3, and the polynomial-thinning condition of Assumption 4 hold, and an almost-everywhere version of has a unique continuous extension The Hölder class imposes the same three design conditions together with the Hölder regression condition of Assumption 5; its canonical definition is Definition 5, and the canonical continuity-only definition is Definition 24.

Definition 4 records the continuity-only class alongside the Hölder class. The later continuity-only section uses the selected continuous regression extension when smoothness is replaced by qualitative continuity on the threshold envelope.

Definition 5 [def:model-class] (Observed law class ).

For , let The observed-data law class is Membership means that satisfies:

  • (Conditional density.) The conditional-density condition in Assumption 2.

  • (Stratum mass.) The stratum-mass condition in Assumption 3.

  • (Polynomial thinning.) The polynomial-thinning condition in Assumption 4.

  • (Hölder regression.) The Hölder member condition in Assumption 5.

⊢ Lean

The observed class in Definition 5 collects the design and regression restrictions under one membership statement. This compact notation is the domain for the estimator, honest interval, and minimax criteria in the next two sections.

Definition 6 [def:clamp-policy] (Clamp policy and ).

For , the lower-threshold clamp map and the clamped treatment are

⊢ Lean

The map in Definition 6 leaves treatments above the threshold unchanged and moves lower treatments to the threshold. The resulting target therefore combines a natural-course component above with a boundary-regression component at .

Definition 7 [def:clamp-functional] (Clamp functionals , , and ).

For as in Definition 5, , and , define The unsmoothed observed-data clamp target is

⊢ Lean

In Definition 7, is the conditional probability mass collapsed by the clamp in stratum , is the retained mean above the threshold, and is the observed-data clamp target. The expression isolates the atom-regression product that drives the boundary inference problem.

At sample size the target is written : it is the functional of Definition 7 evaluated at the deterministic threshold .

The remaining definitions set the bandwidth, rate scale, and range projection used by the estimation and interval statements.

Definition 8 [def:bandwidth] (Information-balance bandwidth ).

The information-balance bandwidth for the threshold sequence is with when the displayed set is empty.

⊢ Lean

The bandwidth in Definition 8 is chosen by the local information available in a window to the right of the threshold. The factor carries the polynomial thinning from Assumption 4 into the window-mass calculation.

Definition 9 [def:frontier] (Frontier rate ).

The Hölder minimax rate is

⊢ Lean

The rate in Definition 9 combines the regular root- contribution from the retained mean and threshold mass with the atom-weighted Hölder approximation scale. This is the benchmark used for the minimax risk and honest expected-length results.

Definition 10 [synth_6] (Projection onto the outcome range).

For a real number , define . Applied to an estimator or interval endpoint, denotes this pointwise projection onto the outcome range .

Definition 11 [synth_2] (Continuity-only regression extension).

For and , denotes the unique continuous function on that agrees, almost everywhere on that interval under the conditional treatment law given , with a version of the conditional mean .

Definition 10 records the range projection used later for estimators and confidence endpoints. Because the target is built from an outcome regression taking values in , the projection keeps the reported numerical object on the same outcome scale.

Estimation and honest intervals

The estimator separates the three empirical tasks in : estimating the retained mean above the threshold, estimating the mass collapsed by the clamp, and estimating the boundary regression value that multiplies that mass. The deterministic split from Definition 1 assigns one block to the retained-course mean, a second to the lower-threshold atom masses, and a third to the local-polynomial fit at the boundary. This separation lets the interval treat the three components with direct concentration bounds for bounded variables. The local regression step follows the boundary-polynomial logic of Fan et al. (1996), with a stabilization rule tailored to the total Gram matrix in each stratum.

The realized-design notation is as follows. It records the rescaled local coordinate, the monomial basis, the stratum-specific local count, and the event on which the total Gram matrix has enough curvature to support the intercept fit.

Definition 12 [def:local-design-notation] (Local design notation).

For a realized sample , deterministic split blocks as in Definition 1, a stratum , local-polynomial degree , threshold , and bandwidth , put Define and For , define and set The good-design event is the event that and, for every , The exact local-polynomial intercept weight is

⊢ Lean

In Definition 12, the local count and the total localized Gram matrix are computed on the third split block . The normalized population moment matrix supplies the reference curvature induced by the polynomial thinning law in Assumption 4. The constant converts that reference curvature into a uniform empirical design threshold, and the good-design event is the realized condition under which the intercept weights are used.

With these quantities fixed, the point estimator is a plug-in estimator for the decomposition in Definition 7. It estimates each summand on its own split block and applies the range projection introduced in Definition 10.

Algorithm 1 [def:total-gram-estimator] (Total-Gram estimator ).

Given split blocks and the local-design quantities in Definition 12, define the estimator as follows.

  1. Compute the retained-course empirical mean

  2. For each , compute the empirical lower-threshold atom mass

  3. For each , compute the clipped local-polynomial threshold estimate

  4. Output the total-Gram clamp-target estimator

⊢ Lean

The estimator in Algorithm 1 uses the local-polynomial intercept only on . On the complementary realized designs, the fallback value keeps the atom-regression factor inside the bounded outcome range. The projection in the last line keeps the estimator on the same scale as .

For statements indexed by a full deterministic split , it is useful to record the same construction with the split made explicit.

Algorithm 2 [def:honest-interval] (Honest interval ).

Given the local-design quantities in Definition 12, define the interval as follows.

  1. Set the Hoeffding calibration multiplier and split-block stochastic-error radii

  2. For each , set the local regression bias-plus-noise radius

  3. For each , set the atom-weighted interval contribution

  4. Output the bias-aware honest interval

⊢ Lean

Definition 13 is the version used by the finite-sample risk and length criteria below. It makes explicit that the sampling split is fixed before the data are observed, while the good-design events and intercept weights are realized from the third block.

The confidence interval adds deterministic bias allowances to concentration radii. The stochastic terms use Hoeffding-type calibration for bounded outcomes and indicators (Hoeffding, 1963), while the local regression radius combines a Hölder bias bound with the empirical size of the intercept weights. This bias-aware construction is in the spirit of honest nonparametric inference with smoothness restrictions (Armstrong et al., 2018; Armstrong et al., 2016).

Definition 13 [synth_4] (Total-Gram stabilized estimator).

Let be a deterministic admissible three-way split, let be the information-balance bandwidth, and let be the local-polynomial order. Define the projection Set For each , let on , and let on . The split-indexed total-Gram estimator is

Here , the number of strata, enters only through the Hoeffding multiplier . The interval in Algorithm 2 treats the retained mean, atom masses, and local-regression intercepts as separate empirical components. On , carries both the uncertainty in and the bias-plus-noise radius for the threshold regression. On , the contribution covers the full feasible atom-regression range, which is the branch used to preserve the same bounded-outcome scale on singular realized designs.

The split-indexed interval records the same accounting with attached to each empirical quantity.

Definition 14 [synth_12] (Bias-aware interval for a fixed split ).

Let be an admissible deterministic three-way split at sample size . With and fixed, form the total-Gram weights and good-design events from the third block . Set , , and on , with on . Define . Let , , and . For each , set and on , while on . The bias-aware interval formed from is .

Definition 14 fixes the interval whose coverage and expected length are evaluated in the main statistical results. Its two branches mirror Algorithm 2: the regular branch combines atom-mass uncertainty with local-polynomial bias and noise, and the fallback branch uses the bounded outcome range when the realized design lacks sufficient Gram curvature.

The length criterion is pointwise in the observed sample and then averaged under the observed product law.

Definition 15 [synth_17] (Length of an observed-sample interval).

For an observed-sample measurable interval , define its length pointwise by . Thus expected lengths such as are taken with respect to the -fold observed-sample law.

Finally, the finite-sample criteria place the split-indexed estimator and interval on the common worst-case scale used in the minimax comparisons.

Definition 16 [synth_5] (Concrete stabilized risk and length).

For a deterministic admissible split , define the finite-sample worst-case risk of the total-Gram estimator by For the bias-aware interval formed from the same split, define its worst-case expected length by

Definition 16 evaluates the point estimator and interval over the observed-law class from Definition 5. The risk uses absolute error for the clamp target , and the length functional applies Definition 15 to the bias-aware interval from Definition 14.

Main statistical results

Section 4 constructed the split-sample total-Gram estimator and its bias-aware interval for the observed clamp target in Definition 7. We now state the corresponding performance guarantees. The product sampling law in Definition 2 and the interval-length convention in Definition 15 put point estimators, confidence intervals, and minimax benchmarks on the same finite-sample scale.

The scalar calibration at the end of the section uses a concrete Bernoulli pair. We record it before the main bounds so that the same notation can be used in the rate check.

Definition 17 [synth_10] (One-cell Bernoulli regressions).

In the one-stratum calibration with and , define the null Bernoulli regression by . For a threshold and bandwidth , define the localized alternative by , where .

The triangular alternative in Definition 17 is centered at the moving threshold and has height proportional to the bandwidth. It gives a transparent version of the same local perturbation that drives the lower-bound component of the general theory.

The minimax benchmarks optimize over all observed-sample procedures. The first criterion measures worst-case absolute accuracy for estimating the unsmoothed clamp target, and the second measures worst-case expected length among intervals with uniform coverage.

Definition 18 [synth_3] (Observed minimax risk and honest length).

For the observed clamp target over , the minimax absolute-error risk is where the infimum ranges over observed-sample measurable estimators taking values in . The minimax honest expected length is where ranges over observed-sample measurable interval procedures with endpoints in .

Thus is the best possible worst-case absolute accuracy for the observed clamp target, while is the best possible worst-case length under uniform honesty. These criteria match the honesty framework used in nonparametric confidence intervals (Low, 1997; Armstrong et al., 2018).

The main risk result gives matching upper and lower bounds at the rate in Definition 9. The upper bound is attained by the stabilized total-Gram estimator from Algorithm 1; the lower bound is witnessed by both a global mean shift and a localized threshold perturbation.

Theorem 1 [thm:minimax-risk] (Minimax clamp frontier).

Fix satisfying the regime conditions:

  • (Regime.) , , , , , , , , , and .

Then there exist constants and an amplitude , depending only on the displayed regime constants, such that, for every deterministic threshold sequence with for every and every deterministic sequence of three pairwise-disjoint sample blocks with for , for all sufficiently large , with the information-balance bandwidth from Definition 8 and as in Definition 9, the observed minimax absolute-error risk over the clamp model in Definition 5 and the worst-case i.i.d. risk over of the stabilized estimator using satisfy Moreover, for all sufficiently large , the lower bound is witnessed inside the same clamp model by:

  • (Global shift.) Laws satisfying Assumption 1, with the same -design distribution and the same and functions, having Bernoulli outcomes, and obeying with well-posed product chi-squared divergence, and

  • (Localized perturbation.) Laws satisfying Assumption 1, with the same -design distribution and the same and functions, having Bernoulli outcomes, and admitting a continuous bump satisfying , for every , for , and for , with regression shift for every and . The product chi-squared divergence is well posed, and

⊢ Lean

The two lower-bound constructions in Theorem 1 isolate the two terms in . A global Bernoulli shift produces the parametric contribution, while a Hölder-sized bump at the threshold produces the atom-weighted local contribution. The estimator in Algorithm 1 attains this combined scale, so the observed clamp problem has the same order from the attainable and unavoidable sides.

The corresponding interval result uses the same rate scale. Its statement combines finite-sample coverage for the constructed interval with a minimax lower bound for every honest observed-sample interval.

Theorem 2 [thm:honest-length] (Honest length frontier).

There are constants , with , for which the following holds.

  • (Regime constants.) The number of strata is a positive integer, , , , , , , , , and .

  • (Threshold path.) The deterministic thresholds satisfy for every .

  • (Sample splits.) For every , is a deterministic three-way split of the sample indices with pairwise disjoint blocks and for .

For all sufficiently large , set to be the information-balance bandwidth for from Definition 8, and set Then the stabilized interval has worst-case coverage and its worst-case expected length is comparable to : Moreover, the minimax honest expected length satisfies where is the infimum, over observed-sample confidence procedures with uniform coverage at least over , of their worst-case expected length. Equivalently, every observed-sample confidence procedure with has worst-case expected length at least :

⊢ Lean

The interval in Theorem 2 is calibrated to the same two sources of difficulty as the estimator. The retained-course and atom-mass terms are controlled by bounded-outcome concentration, while the threshold regression term carries the local Hölder bias and the realized Gram behavior encoded in Algorithm 2. The lower bound shows that uniform honesty over requires expected length of order , in line with modulus-based accounts of optimal confidence sets (Low, 1997; Armstrong et al., 2016).

The two theorems above leave the rate in the compact form . The next result expands that expression into the threshold regimes induced by the information-balance bandwidth. Two deterministic scales organize those regimes, the risk-transition scale and the bandwidth-transition scale , and they are restated inside Theorem 3. Table 1 collects the regimes, separating the geometry of the treatment support at the threshold from the resulting risk order.

Bandwidth and risk regimes for the lower-threshold clamp, with and . The third column records the geometry of the treatment support at the threshold; the fourth records which contribution determines the risk order.
Result type Condition Treatment-support geometry Conclusion
Bandwidth window reaches the thinned endpoint .
Bandwidth window sits away from the thinned endpoint .
Risk threshold near the thinned endpoint retained mean dominates: .
Risk threshold near the thinned endpoint retained mean and atom term balance: .
Risk and threshold vanishing but above the edge scale atom term dominates: .
Risk threshold interior to the treatment support atom-weighted pointwise Hölder rate: .
Risk no mass collapsed .
Theorem 3 [thm:phase-diagram] (Phase boundary regimes).

Let be a positive integer and let Define the phase and edge scales by Then Moreover, for every deterministic threshold sequence with for every , let be the information-balance bandwidth in Definition 8, and set as in Definition 9. The following conclusions hold:

  • Regular thresholds. If , then .

  • Critical thresholds. For every , if , then and .

  • Vanishing atom-dominated thresholds. If and , then

  • Fixed thresholds. For every , if , then

  • Zero threshold identities. For every and every satisfying Definition 5, where the last identity uses the bandwidth at threshold .

  • Zero threshold estimator. For every , every satisfying Definition 5, and every admissible three-way split as in Algorithm 1, the total-Gram estimator in Algorithm 1, formed at threshold with bandwidth , agrees under the product sampling law generated by with

  • Zero threshold stabilized guarantees. For every deterministic sequence of admissible three-way split blocks, the stabilized zero-threshold procedure has coverage at least for all sufficiently large . Its stabilized worst-case risk at threshold and its stabilized worst-case expected length at threshold with noncoverage level are both asymptotic to .

⊢ Lean

Theorem 3 identifies as the boundary at which the local atom contribution reaches the parametric scale, and as the smaller edge scale for the bandwidth algebra. Regular and critical threshold paths retain root- order. Vanishing thresholds above the critical scale inherit a degenerate-design nonparametric rate of the kind studied by Gaïffas (2005): the estimation point rides toward the thinned endpoint , so the local design mass degenerates at the polynomial rate implied by Assumption 4. A fixed positive threshold is an interior-design geometry. There the target point is interior to the treatment support and the density at that point is bounded away from zero, so the exponent is the ordinary positive-density pointwise Hölder rate (Stone, 1982; Donoho et al., 1991) multiplied by an atom weight bounded away from zero. The estimator still uses a one-sided window, because the clamp evaluates the regression at from above, but that is a property of the procedure rather than of the design geometry. At the exact zero threshold, the clamp target reduces to the observed mean and the total-Gram estimator reduces to the clipped split-sample outcome mean.

The section closes with a scalar calibration. It specializes the localized perturbation to one stratum, linear thinning, and Lipschitz smoothness, giving an explicit exponent check for the critical scale.

Proposition 1 [prop:one-cell-calibration] (One-cell calibration).

Let , and let be a deterministic threshold sequence satisfying Assume that , in the sense that In the one-stratum model with and , let be the information-balance bandwidth from Definition 8 evaluated at . For the concrete localized Bernoulli alternatives, define and Then Moreover,

⊢ Lean

In Proposition 1, is the product Kullback–Leibler scale for the localized Bernoulli alternative, while is the induced clamp-target separation. The balance keeps the localized alternative statistically contiguous at constant information scale, and the identity reproduces the critical exponent from Theorem 3 when .

Causal interpretation through full-data lifts

The observed-data results in Theorems 1 and 2 concern the clamp functional under the law of . This section supplies a full-data causal interpretation for the same target. One notational convention applies throughout this section and the continuity-only section. Every estimator and interval used in this paper is a measurable function of the observed sample, so its risk, coverage probability, and expected length are the same whether computed under the -fold product law of the observed margin or under the -fold product law of the full-data package. We therefore state all such quantities under the observed-margin product law , and use the full-data product law only where the statement genuinely concerns full-data objects. The structural full-data overlay augments the observed experiment with a full-data law , a member-specific latent-response space , a latent response variable , structural response maps , and a potential-outcome process . Within that overlay, the lower-threshold policy changes the realized treatment through the same clamp map introduced in Definition 6; the statistical analysis remains based on the observed sample.

The analysis rests on three full-data restrictions and one response-mean definition. The first two restrictions are standard causal conditions for a latent-response representation, and the third records the continuity required on the threshold interval used by the clamp policy.

Assumption 6 [ass:consistency] (Structural consistency).

Under , the following conditions hold:

  • (Latent response.) The latent response variable takes values in .

  • (Structural maps.) For every , the map is jointly Borel measurable.

  • (Potential and observed outcomes.) There is a single -null set outside which, for every ,

⊢ Lean

Assumption 6 is the structural consistency condition: a common latent response generates all potential outcomes and the observed outcome at the realized treatment, in the simultaneous NPSEM-style formulation used for modified treatment policies (Díaz et al., 2021). The common null set makes the equality simultaneous over doses, which is the form needed when the clamp map evaluates the response at .

Assumption 7 [ass:exchangeability] (Conditional exchangeability).

Under , the treatment and latent response variable are conditionally independent given .

⊢ Lean

Assumption 7 is the latent-response randomization condition (Díaz et al., 2021). Within each stratum, it lets the conditional treatment density weight the same latent-response distribution that determines the structural mean.

Definition 19 [synth_1] (Full-data response mean).

For a full-data law package with latent-response carrier , latent variable , structural maps , and observed stratum , the full-data response mean in stratum is When the law package is in the full-data continuity model, the map is continuous on for every .

The full-data response mean is the fiberwise structural analogue of the observed regression extension. It isolates the average structural response at dose among units in stratum , which is the quantity the bridge result compares with on the threshold interval.

Assumption 8 [ass:full-data-response-continuity] (Full-data response continuity).

Under the full-data law , for every , the full-data response mean is continuous on .

⊢ Lean

Assumption 8 is specific to this analysis. It places the continuity requirement directly on the structural response mean over the range where clamping can create a threshold evaluation, matching the interval on which the observed regression extension is used.

These conditions define the full-data causal class and the decision criteria used below.

Definition 20 [def:full-data-model-class] (Full-data causal class ).

is the collection of full-data law packages satisfying the following memberwise structure relative to Definition 5, Assumption 6, Assumption 7, and Assumption 8.

  • (Latent response space.) Each member carries its own standard Borel latent-response space and an -valued random variable .

  • (Structural response maps.) Each member carries jointly Borel maps and a jointly measurable potential-outcome process

  • (Full-data law.) Each member carries a probability law for

  • (Observed margin.) The -margin is some .

  • (Causal member conditions.) Within the member’s own carrier, the simultaneous latent-response consistency, latent randomization, and fiberwise structural-mean continuity on conditions hold.

The class is formed member by member, with membership determined on each member’s own standard Borel carrier.

⊢ Lean

The full-data class keeps the observed law inside while allowing each member to carry its own standard Borel latent-response space. This memberwise formulation is the natural setting for comparing causal and observed decision problems without imposing a common latent carrier across all laws.

Definition 21 [def:causal-clamp-mean] (Causal clamp mean ).

For a full-data law as in Definition 20, the clamp policy from Definition 6, and , define the causal clamp mean by

⊢ Lean

The causal clamp mean is the mean outcome under the same lower-threshold modified treatment policy used to define the observed clamp functional in Definition 7. Thus the causal estimand changes the interpretation of the response while preserving the policy rule and threshold range.

Definition 22 [def:causal-frontier-criteria] (Causal frontier criteria and ).

For the causal clamp mean in Definition 21, define the causal minimax absolute-error risk over estimators measurable with respect to the observed sample by For intervals measurable with respect to the same observed sample, define the causal minimax honest expected length by

⊢ Lean

The causal minimax risk and causal honest-length criterion compare observed-sample estimators and intervals under the full-data class through their induced observed-sample behavior. The procedures therefore use the same data as in the observed experiment while the target is the causal clamp mean.

The next proposition gives the identification step. It connects the structural mean under the full-data law to the observed regression extension and then identifies the causal clamp mean with the observed clamp target.

Proposition 2 [prop:causal-bridge] (Causal bridge identity).

Let satisfy Definition 20, let be its observed -margin, and let be the clamp map in Definition 6. Suppose the frontier constants satisfy

  • (Strata.) .

  • (Smoothness and thinning.) , , and .

  • (Density and mass constants.) , , , and .

  • (Threshold and coverage levels.) and .

For and , define the full-data response mean Then, for every and every Borel set , Moreover, for every and every , Consequently, for every , the following identities hold: and, with the causal clamp mean equals the observed clamp target

⊢ Lean

Proposition 2 is the fiberwise bridge. Conditional exchangeability converts observed outcomes within a stratum and treatment set into an integral of the structural response mean against the treatment density; continuity and the observed regression extension then identify that mean on the threshold interval. The final identity shows that the causal lower-threshold policy mean is exactly the clamp functional analyzed in the observed experiment.

A second structural fact records that the full-data class covers the observed model class exactly at the margin. The construction uses regular conditional distributions on standard Borel spaces, a standard measurable-disintegration tool (Kallenberg, 2002).

Proposition 3 [prop:observed-margin-surjectivity] (Observed margin surjectivity).

Fix and constants . Assume:

  • (Regime constants.) , , , , , , , , , , and .

  • (Observed model.) means that satisfies Definition 5 with constants .

  • (Full-data model.) means that satisfies Definition 20 with constants .

Then, for every , there exists whose observed -margin is exactly . Hence the observed-margin map from onto is surjective.

Moreover, for every and every , the causal frontier criteria in Definition 22 satisfy Here is the infimum, over observed-measurable estimators with values in , of the worst-case absolute-error risk over for the observed clamp target at threshold , and is the infimum, over uniformly honest observed-measurable confidence procedures at level , of the worst-case expected interval length over .

⊢ Lean

Proposition 3 establishes the decision-theoretic equivalence needed for causal transport. Every observed law in has a full-data representative in , and the risk and honest-length criteria agree at each fixed threshold. Thus lower bounds and attainable procedures for the observed clamp target can be read as statements about the causal clamp mean on the declared full-data class.

One technical convention remains for comparing expected lengths. The next definition fixes the order-preserving way real bounds are read inside the extended nonnegative reals.

Definition 23 [synth_11] (Real-to-extended embedding).

Under the notation and hypotheses of the objects introduced earlier in the paper, a real number is embedded into the extended nonnegative reals by The image is finite and equals the positive part of , so real expected-length bounds are compared with extended-real expected lengths through this embedding.

The embedding in Definition 23 lets the real constants in the rate bounds be compared with honest expected lengths when those lengths are represented in the extended nonnegative-real order. This is the scale used in the final causal length statement.

The main causal conclusion now follows by combining the bridge identity, observed-margin surjectivity, and the observed-data minimax and honest-interval results.

Theorem 4 [thm:causal-frontier-lift] (Causal frontier lift).

Let , and let . Suppose that

  • (Regime constants.) The constants satisfy

  • (Threshold sequence.) The deterministic thresholds satisfy for every .

  • (Split blocks.) For each , is a deterministic three-way split with pairwise disjoint blocks and

Let , the information-balance bandwidth for , be as in Definition 8, and let be the frontier rate from Definition 9. Then there exist constants , with , depending only on these regime constants, such that, for every threshold sequence and every deterministic split-block sequence satisfying the preceding conditions, all sufficiently large satisfy the following conclusions.

The total-Gram estimator in Algorithm 1, formed with split , order , bandwidth , and threshold , is observed-sample measurable. The corresponding bias-aware interval in Algorithm 2, formed with the same split, order, bandwidth, and threshold, is an observed-sample measurable interval.

For the causal minimax criteria and in Definition 22, Moreover, in the extended nonnegative-real order, after embedding the nonnegative real bounds into ,

⊢ Lean

Theorem 4 transports the observed minimax and honest-length frontier to the causal clamp mean on . The same total-Gram estimator and the same bias-aware interval remain observed-sample procedures, and their guarantees hold uniformly over the full-data class because Proposition 2 aligns the target while Proposition 3 aligns the induced decision criteria. The rate is therefore the same characterized in Definition 9 and Theorem 3: root- and boundary-driven regimes carry over with the causal interpretation supplied by the full-data lift.

Continuity-only clamp inference

Sections 3 and 5 impose a Hölder modulus for the regression at the lower threshold. We now record the enlarged continuity-only experiment, which keeps the same density, stratum-mass, and lower-tail thinning restrictions while replacing the Hölder regression member with qualitative continuity.

Definition 24 [def:continuity-model-class] (Continuity-only model class ).

is the set of laws satisfying Assumptions 2, 3, and 4 and the following continuity condition: for every , an almost-everywhere version of admits a continuous extension to . The continuity extension is required memberwise for each stratum.

⊢ Lean

The class in Definition 24 preserves the design restrictions that govern the amount of mass near the lower endpoint. Its regression condition supplies a selected point value at the clamp threshold through continuity, rather than through a quantitative Hölder modulus.

The corresponding target has the same retained-course plus threshold-atom decomposition as in Definition 7; the atom value is read from the continuous extension in the continuity-only class.

Definition 25 [def:continuity-clamp-functional] (Continuity clamp target ).

For and , with as in Definition 7 and as in Definition 24, define

⊢ Lean

The continuity-only scale is determined by the ordinary sampling fluctuation and the amount of mass moved by the lower-threshold clamp.

Definition 26 [def:continuity-frontier] (Continuity frontier ).

Define the continuity-only frontier rate by

⊢ Lean

The estimator and interval make the enlarged model explicit. They use the retained-course mean above the threshold and a fixed midpoint value for the lower-tail atom, with a bounded-sum allowance for the empirical lower-tail mass.

Algorithm 3 [def:continuity-fallback-estimator] (Continuity fallback estimator ).

The inputs are the deterministic split blocks , the threshold , and the observations .

  1. Define the retained-course empirical mean

  2. For each , define the empirical lower-threshold atom mass

  3. Output the fixed-regression-fallback estimator

⊢ Lean
Algorithm 4 [def:continuity-honest-interval] (Continuity honest interval ).

The inputs are the split blocks , the level , the estimator , and the empirical masses .

  1. For , define the split-specific Hoeffding radius

  2. Define the empirical total lower-threshold atom mass

  3. Output the Hoeffding interval

⊢ Lean

The concentration step for Algorithm 4 uses the standard bounded-sum inequality of Hoeffding (1963), applied separately to the retained-course and atom-mass split statistics.

The causal version of the continuity-only analysis uses the same structural ingredients as Definition 20, with the observed margin placed in .

Definition 27 [def:continuity-full-data-model-class] (Continuity full-data class ).

Define as the collection of full-data law packages with a memberwise standard Borel latent-response space , latent variable , jointly Borel stratum-specific response maps , and jointly measurable potential-outcome process , such that:

  • (Observed margin.) The observed -margin of belongs to in Definition 24.

  • (Consistency.) The package satisfies Assumption 6.

  • (Exchangeability.) The package satisfies Assumption 7.

  • (Response continuity.) The package satisfies Assumption 8.

The carrier spaces may vary with .

⊢ Lean

For later comparison, the next definition fixes the observed margin of a full-data law as the pushforward onto the observed coordinates.

Definition 28 [def:continuity-frontier-criteria] (Continuity frontier criteria and ).

For estimators and intervals measurable with respect to , define the continuity-only observed minimax absolute-error risk by Define the continuity-only observed minimax honest expected length by For the full-data continuity-only class, define the causal minimax absolute-error risk by Define the continuity-only causal minimax honest expected length by

⊢ Lean

The decision-theoretic criteria are recorded for both the observed continuity-only class and its full-data counterpart.

Proposition 4 [prop:continuity-causal-bridge] (Continuity causal bridge).

Let , let , and let be a full-data law package with observed margin . Suppose that

  • (Full-data continuity model.) with constants , as in Definition 27.

  • (Design constants.) , , , , , and .

For each and , define the full-data response mean by Then, for every and every , Moreover, if then, for every , where is the continuity-only clamp functional in Definition 25.

⊢ Lean

The first causal statement aligns the full-data response mean with the observed continuous extension and therefore identifies the causal clamp mean with the continuity-only observed target.

Proposition 5 [prop:continuity-observed-margin-surjectivity] (Observed-margin surjectivity).

Let and let . Suppose that the continuity-regime design constants used in Definitions 24, 27, and 28 hold for , and that .

Then the observed-margin map from onto is surjective: for every , there exists whose observed margin is . Moreover, for every and every , the four continuity-frontier criteria of Definition 28, evaluated at , satisfy

⊢ Lean

Proposition 4 shows that qualitative continuity turns almost-everywhere regression information into the pointwise threshold value needed after the clamp collapses the lower tail. Under the stated full-data continuity conditions, the structural response mean agrees with the observed continuous extension throughout the threshold envelope, and the causal clamp mean equals the observed continuity clamp target.

The next statement records that the continuity-only full-data class covers the observed continuity-only class exactly at the margin. Its construction uses regular conditional distributions on standard Borel spaces, as supplied by Kallenberg (2002).

Definition 29 [synth_13] (Observed margin of a full-data law).

For a full-data law package governing , define to be the observed-data marginal law of . Equivalently, if is the coordinate projection onto , then .

Proposition 5 puts the observed and causal continuity-only decision problems on the same footing. Every observed continuity-only law has a full-data representative satisfying the causal restrictions, and the minimax risk and honest expected-length criteria agree exactly for each sample size and threshold in the stated range.

The main continuity-only conclusion now combines the bridge identity, observed-margin surjectivity, and the fallback estimator and interval.

Theorem 5 [thm:continuity-only-frontier] (Continuity frontier characterization).

Fix and constants . Suppose Then there exist constants with such that the following statements hold.

  • (Uniform frontier bounds.) For every deterministic threshold sequence with for all , and for every deterministic sequence of admissible three-way split blocks, all sufficiently large satisfy the following bounds. With and with , , , and as in Definition 28, the interval has uniform coverage at least over , its worst-case expected length is at most , and Moreover, and the same estimator and interval satisfy with causal coverage at least uniformly over and causal worst-case expected length at most .

  • (Elbow and null-threshold rates.) For the continuity-only rate in Definition 26,

  • (Fixed positive thresholds.) For every deterministic threshold sequence with for all , every , and ,

⊢ Lean

Theorem 5 characterizes the continuity-only rate for both observed and causal formulations. The fallback estimator and Hoeffding interval attain the rate , while the minimax risk and honest expected length are bounded below at the same order. The honest-length comparison follows the nonparametric confidence-interval lower-bound phenomenon of Low (1997), specialized here to the threshold-atom decision problem.

The rate has a simple interpretation. When is at the elbow , the lower-tail mass term matches , so the continuity-only problem has a root- scale. At , the atom term vanishes and the bounded retained-course mean again gives the root- scale. For thresholds converging to a fixed positive value, the theorem records persistent minimax risk and honest length, reflecting the positive amount of treatment mass assigned a threshold regression value governed by qualitative continuity.

Discussion, limitations, and open questions

The results characterize how lower-tail overlap and smoothness jointly determine inference for threshold modified treatment policies with continuous treatments. In the Hölder model , Theorems 1 and 2 give the common rate , combining the ordinary sampling term with the local extrapolation term . The phase diagram in Theorem 3 explains how this expression changes with the threshold sequence: moving the clamp closer to the lower endpoint changes both the mass affected by the policy and the amount of information available near the threshold.

The continuity-only comparison in Theorem 5 separates quantitative smoothness from qualitative continuity. With a Hölder radius and exponent fixed in the model, the local polynomial construction uses observations in a shrinking neighborhood of and pays the bias term . With continuity alone, the delivered rate is , because the threshold regression contribution is controlled through boundedness and continuity rather than through a quantitative modulus. Taken together, the two frontiers show that atom mass and smoothness enter different parts of the problem: the mass below the threshold weights the regression value to be learned, while the smoothness condition determines how sharply that value can be inferred from nearby continuous-treatment observations.

The causal lifting results give the same statistical rates a full-data interpretation. Under the consistency, exchangeability, and response-continuity conditions in Assumptions 6, 7, and 8, Proposition 2 and Theorem 4 transfer the observed-data Hölder frontier to the causal clamp mean. The continuity-only bridge in Proposition 4 and Theorem 5 gives the analogous interpretation for the qualitative-continuity class. These statements make the role of the full-data class explicit: once the observed margin belongs to the relevant observed model and the stated causal restrictions hold, the observed clamp functional and the causal clamp mean have matching minimax criteria.

The exact-modulus perspective below records a more calibrated object behind the order results. The intervals in Algorithm 2 use analytic upper bounds on local bias and stochastic error. A sharper implementation can instead optimize directly over affine weights in the realized local design, in the spirit of modulus-based honest inference (Armstrong et al., 2018; Armstrong et al., 2016). The next three definitions name the realized affine-modulus handle, the integrated split-specific radius, and the resulting interval used to formulate the sharp calibration question.

Definition 30 [def:exact-modulus-handle] (Exact modulus radius).

With notation as in Definitions 12 and 5, fix under the displayed member conditions, a deterministic split with pairwise-disjoint blocks satisfying for , a stratum , a local-polynomial degree , a threshold , a bandwidth , and a multiplier . For a realized observed sample , with , set The affine weight set is For , let denote the exact Hölder bias term from the -Hölder regression class in Assumption 5. Define the realized affine-modulus radius by with the real-valued convention that this infimum is when the displayed feasible set is empty. The integrated exact-modulus radius is where is the canonical product law of the observed sample. Both and are radius, or half-length, quantities.

⊢ Lean
Definition 31 [synth_21] (Exact-modulus radius ).

Adopt the notation and hypotheses of the objects introduced earlier in the paper. Fix , a deterministic third split block , a stratum , and a multiplier . For a realized design , set Let Let be the class of continuous functions satisfying the declared order- Hölder Taylor-remainder bound with exponent and radius . For , define the realized Hölder bias functional The realized affine-modulus radius at multiplier is The integrated exact-modulus radius is where is the independent observed-design product law on the coordinates in .

If is an admissible deterministic three-way split at sample size , its worst integrated exact-modulus radius is Thus is a radius quantity computed from the third block of , and the associated untruncated symmetric full length is .

Definition 32 [synth_20] (Exact-modulus interval).

Assume the notation, model conditions, and preceding constructions in the objects introduced earlier in the paper. Fix an admissible deterministic three-way split at sample size . For a realized observed sample , let be the total-Gram estimator formed with the blocks in , and evaluate the realized affine-modulus radius on the -projection of using the third block , where Define The exact-modulus interval is Its worst-case expected length over the observed-data model is

Definitions 30, 31, and 32 record one concrete, fully specified object against which the sharp-constant question can be posed: an interval centered at the total-Gram estimator whose half-width is the realized affine-modulus radius, intersected with . It is important to be precise about what that object does and does not calibrate. Its radius solves the affine-modulus problem for the threshold-regression component alone: it is the supremum over strata of unweighted realized regression radii, it does not include the retained-mean deviation radius, the atom-mass deviation radii, or the atom weights by which those components enter the target, and it does not aggregate across strata in the way the target does. On the singular realized designs where the affine constraints are infeasible the realized radius is zero, so the object is not by itself a coverage-guaranteed procedure; the honest interval with a proved coverage guarantee is Algorithm 2. The zero-threshold regime makes the gap explicit: there the target is an ordinary bounded mean with frontier , while this radius continues to solve a local-regression problem whose natural scale is . The question below is therefore posed for a component-level modulus; extending it to a modulus for the complete clamp functional, with the regular components, atom weights, their uncertainty, stratum aggregation, and a singular-design fallback included, is the substantive part of the problem and is open. The integrated radius averages the design-dependent affine optimization under the observed-design product law for the third split block and is defined stratum by stratum, replacing the analytic threshold-regression radius by a realized optimization problem.

Limitations.

The scope of the results is the fixed finite-stratum design described in Definition 5, with outcomes bounded in , known smoothness radius and exponent in the Hölder model, and a known polynomial-thinning exponent governing lower-tail overlap. The causal interpretation is attached to the declared full-data classes in Definitions 20 and 27, and the continuity-only conclusions are attached to the continuity-only observed and full-data classes in Definitions 24 and 27. Within that scope, the theorems characterize the minimax risk and honest expected-length orders for the lower-threshold clamp target and its causal lift. A further limitation concerns evidence rather than scope. The case made here for the stabilized estimator and the bias-aware interval is an asymptotic-order case, supported at the level of a single analytic calibration cell in Proposition 1. Practical questions about total-Gram stabilization on singular realized designs and interval behavior at moderate sample sizes concern quantities whose finite-sample behavior need not resemble the eventual regime: the frequency of the Gram fallback, the realized coverage, and the realized expected length all depend on constants this analysis does not track. A systematic finite-sample evaluation across the regular, critical, atom-dominated, fixed-threshold, and near-zero regimes, reporting those quantities and comparing against unstabilized local polynomial and regular modified-policy benchmarks where the targets and conditions permit a fair comparison, is left to future work and is not attempted here.

Open questions.

Two questions remain open. The first is adaptation: the frontier established here is the known-class benchmark, with , , and fixed and available to the procedure, and whether the same frontier is attainable when those constants are unknown, and at what logarithmic cost, is not settled by the present analysis. The second concerns exact constants rather than rates. Modulus-based confidence-interval theory describes the difficulty of such problems through the local modulus of a functional over a smoothness class; Cai et al. (2004) gives a general adaptation framework for confidence intervals, while Armstrong et al. (2018); Armstrong et al. (2016) develop honest inference procedures that use modulus calculations for bias-aware intervals. The honest expected-length lower bound proved here is obtained by direct two-point and localized-perturbation arguments, and it matches the risk frontier in order. The clamp target is linear in the outcome regression but includes the design-dependent weight , so sharp modulus calibration is treated as a separate question. The exact-modulus constructions in Definitions 30, 31, and 32 identify one concrete candidate for studying sharp honest length: use the total-Gram center, compute the realized affine-modulus radius on the third split block, and compare its worst integrated radius to the minimax expected-length benchmark.

Remark 1 [oeq:sharp-constant] (Sharp length calibration).

The sharp constant question asks for the following property. For every satisfying there is a constant such that every deterministic threshold sequence with for all , and whose phase-boundary ratio satisfies either or has the following simultaneous calibration properties. First, where is the information-balance bandwidth from Definition 8 and is the frontier from Definition 9. Second, there exists a deterministic sequence of admissible three-way split blocks such that, for all sufficiently large , where is the exact-modulus interval obtained from the split using the exact-modulus construction of Definition 30. Its worst-case expected length satisfies Finally, the same sequence admits least-favorable mixed-experiment witnesses: there are a constant , an amplitude , and sequences such that, for all sufficiently large , the three laws satisfy the clamp-model member conditions; their -sample laws are the product laws in Assumption 1; the pair has the declared global Bernoulli shift at scale ; the pair has the declared localized Bernoulli perturbation at , , , and ; and With and with denoting the worst integrated exact-modulus radius from Definition 30 computed using the third block in , the matching limits are

⊢ Lean

A positive answer to Remark 1 would identify the sharp asymptotic full-length constant for the exact-modulus interval across the phase regimes covered by the displayed threshold condition. It would also connect the mixed regular-and-local perturbation comparison directly to the integrated realized radius, sharpening the order-optimal honest interval of Theorem 2 into an exact asymptotic calibration.

Appendices

Proofs and verification note

This appendix collects the auxiliary statements that support the main risk, honest-length, causal-bridge, and continuity-only conclusions. The statements are organized by mathematical role: the clamp pushforward and support facts, the bandwidth and local-design algebra, the stabilized estimator and interval bounds, the phase calculations, the population moment identities underlying total-Gram stabilization, and the uniqueness statement for the continuity-only regression extension.

The first two propositions record the measure-theoretic behavior of the exact lower-threshold clamp. They connect the deterministic policy in Definition 6 to the mixed law that appears in the observed clamp functional in Definition 7.

Proposition 6 [prop:pushforward-setup] (Clamp pushforward decomposition).

Let be a positive integer, let satisfy Definition 5, and let as in Definition 6. Suppose the constants satisfy:

  • (Smoothness and thinning.) , , , and .

  • (Stratum mass.) .

  • (Threshold and level.) and .

  • (Evaluation point.) and .

Define the lower-tail atom mass Then the conditional pushforward law of the clamped treatment satisfies and the atom mass obeys

⊢ Lean

Proposition 6 supplies the exact mixed-measure decomposition induced by the clamp. The continuous part retains the natural treatment values above the threshold, and the lower tail becomes a point mass at the threshold whose order is under Assumption 4. This atom-size calculation is the source of the weighting in Definition 9 and Definition 26.

Proposition 7 [prop:policy-support] (Policy support preservation).

Let , , , and satisfy:

  • (Model.) satisfies Definition 5 with constants .

  • (Regime constants.) , , , , , , , and .

  • (Stratum and threshold.) and .

Then, for conditional-treatment almost every in stratum , the clamped value from Definition 6 is a support point of the conditional natural-treatment law given : every open neighborhood of has positive conditional-treatment measure in stratum ,

⊢ Lean

Proposition 7 records the support property used by the causal bridge in Proposition 2. The polynomial lower envelope gives positive conditional treatment measure in every neighborhood of the clamp value, so the structural response at the clamped treatment is evaluated on the support generated by the observed design.

Lemmas 10, 11, 12, and 13 establish the bandwidth normalization and the realized local-polynomial algebra used by the total-Gram estimator. These statements translate the thinning envelope into local sample information and deterministic weight identities.

Lemma 1 [lem:shifted-moment-coercivity] (Uniform shifted coercivity).

Assume:

  • (Degree.) Fix a polynomial degree .

  • (Exponent.) Fix a thinning exponent .

  • (Nonnegativity.) The exponent satisfies , as in Assumption 4.

  • (Reference matrix.) Let denote the normalized population reference moment matrix in Definition 12.

Then there exists a constant such that, for every shift and every vector ,

⊢ Lean
Proof of Lemma 1.

For , set and Since , expansion of the square and the identity give

If , then is a nonzero polynomial. Its zero set in is finite, so choose with . At this point , while the integrand is continuous and nonnegative on . Hence

The map is continuous, because it is a finite sum of quadratic coordinate terms. Let The set is compact and nonempty, so attains its minimum on , say at . Since , strict positivity of at nonzero vectors gives . Define Then . For , the desired bound against is immediate. For , put Then , , and Therefore

It remains to compare this unshifted energy with the normalized shifted moment matrix. Fix , and define By the definition of ,

The denominator is positive. Indeed, is continuous and nonnegative on , and it is strictly positive near . With monotonicity of on gives

For , and hence Multiplying by and integrating yields

Since , the inequalities and imply

Combining the coercive bound for with the displayed expression for gives, for every ,

Lemma 10 turns the infimum definition of into an eventual equality at the information boundary. This equality is the normalization used in the risk upper bound and in the expected weight-energy calculation.

Lemma 2 [lem:lambda-star-positive] (Positive Gram constant).

In the local-design notation of Definition 12, suppose:

  • (Degree.) The polynomial degree is .

  • (Thinning exponent.) The exponent satisfies , as in Assumption 4.

  • (Envelope constants.) The envelope constants satisfy and .

Then the uniform population Gram lower-bound constant is strictly positive:

⊢ Lean
Proof of Lemma 2.

By Lemma 1, there is a constant such that, for every and every ,

Taking the infimum over all unit vectors gives

The indexing set is nonempty, so

Because and ,

Lemma 11 gives the population conditioning bound for the local monomial design. The lower bound is scaled by the local window mass, so it remains compatible with the global polynomial-thinning envelope in the lower-endpoint window.

Lemma 3 [lem:power-window-integral-lower] (Power window integral lower bound).

Using the thinning-exponent notation of Assumption 4, let be a threshold, a window width, and an exponent. Assume:

  • (Threshold.) .

  • (Window width.) .

  • (Exponent.) .

Then

⊢ Lean
Proof of Lemma 3.

Set Then

For every , the hypotheses and give and the bound recorded above, together with , gives Together with , monotonicity of real powers on nonnegative bases yields Thus

The left-hand side is the constant integral over an interval of length :

Since on and ,

Combining the window-length bound, the pointwise lower bound on , and the monotonicity of the window integral gives

Lemma 12 is the deterministic local-polynomial fact used on the good-design event. Moment reproduction gives the intercept interpretation at the threshold, while the absolute-weight and squared-weight bounds control the Hölder bias and stochastic radius in Algorithm 2.

Lemma 4 [lem:local-window-mass] (Local window mass identity).

Let , an observed-data law, belong to the model class of Definition 5, with conditional treatment densities, stratum masses, and polynomial thinning as in Assumptions 2, 3, and 4 for constants . Fix:

  • (Stratum.) A stratum .

  • (Threshold and width.) A threshold and a window width with .

  • (Window support.) .

Then the population mass of the local window in stratum is

⊢ Lean
Proof of Lemma 4.

Let Because , for every observation ,

Therefore

Integrating this identity gives

The window-support assumption gives . Applying the conditional-density law in stratum to the Borel set yields

The asserted identity follows from the two displayed evaluations of the local window weight.

Lemma 13 integrates the squared-weight scale over the random design. Under the balance equation, the average stochastic radius has the same order as the deterministic local-polynomial bias term.

The following two lemmas are the finite-sample ingredients for the stabilized procedures. They combine bounded-outcome concentration, deterministic fallback branches, and the local-design controls above.

Lemma 5 [lem:local-gram-entry-integral] (Local Gram integral identity).

Fix a law with declared constants , as in Definition 5, Assumption 2, Assumption 3, and Assumption 4. Let , , and . Assume:

  • (Constant signs.) , , and .

  • (Window.) and .

  • (Entries.) .

For , write as in Definition 12, and define Then

⊢ Lean
Proof of Lemma 5.

Write Since , for every real ,

Define the bounded measurable function For , the entries and are powers of , so their absolute values are at most one. Hence

The signs and , together with the density envelope on , give the pointwise domination used for integrability: for Lebesgue-a.e. , The last inequality follows from and , and it is preserved after multiplication by . Hence Consequently, if is measurable and , then so bounded measurable multiples are integrable for the conditional treatment measure in stratum . The hypothesis is carried with the model hypotheses used below for the stratum identity.

The stratum-restricted treatment law therefore has the following integral form: for every bounded measurable , To see this, apply the conditional-density condition in Assumption 2 first to indicators , where is Borel: Linearity gives the identity for nonnegative simple functions, monotone convergence gives it for nonnegative measurable functions, and decomposition into positive and negative parts gives it for bounded measurable ; the integrability just established makes both sides finite.

Applying this stratum identity to the bounded measurable function gives

By the definition of and ,

Finally, since and off , while on ,

Combining the displayed stratum decomposition with the preceding indicator rewriting yields

Lemma 14 gives the finite-sample coverage statement for the interval in Algorithm 2. The interval combines split-block concentration with a bias-aware local radius; on singular local designs, the bounded fallback radius covers the whole possible threshold-mean contribution in the affected stratum.

Lemma 6 [lem:scaled-window-coercivity] (Scaled window coercivity).

Let , and let be the uniform population Gram lower-bound constant from Definition 12, formed with the lower and upper thinning constants in Assumption 4. Suppose that

  • (Shape constants.) , , and .

  • (Window.) and .

  • (Coefficient vector.) .

Then

⊢ Lean
Proof of Lemma 6.

Put Since and , one has . The function is continuous on , is nonnegative there, and has value at . Hence it is positive on a neighborhood of , and the interval-integral positivity criterion gives .

By the definition of in Definition 12, expanding the quadratic form gives The uniform coercivity in Lemma 1 supplies the lower boundedness used to compare the defining infimum in with the -term. The usual rescaling from arbitrary vectors to the Euclidean unit sphere then gives Indeed, for this is immediate, while for applying the unit-sphere minimum to yields Multiplying by the nonnegative factor and using the displayed identity for gives the bound . Since , multiplying by gives

It remains to transport this inequality from to the window. With the change of variables , because and . The same substitution gives

Since , multiplying the unit-window inequality by this factor and using the two displayed identities gives exactly

Lemma 15 supplies the estimator side of the upper bound in Theorem 1. The retained-course and atom-mass empirical terms contribute the scale, while Lemma 13 aligns the threshold-regression stochastic scale with the Hölder bias scale after multiplication by the collapsed mass.

The phase analysis rests on a separate bandwidth comparison. The result below identifies the edge scale where the boundary form of the local design transitions to the interior form.

Lemma 7 [lem:intercept-weight-formula] (Intercept weight formula).

Fix the local-design notation of Definition 12 and the split-block construction of Algorithm 1. Let be a deterministic three-way split with pairwise disjoint blocks, each of cardinality at least . Fix a realized sample , a stratum , a degree , constants , a threshold , a window width , and an index .

Assume:

  • (Good design.) The good-design event holds for .

  • (Regression block.) The index satisfies .

  • (Stratum match.) The realized stratum satisfies .

  • (Local window.) With , one has .

Then the exact local-polynomial intercept weight satisfies

⊢ Lean
Proof of Lemma 7.

Let On , the totalized definition of the exact intercept weight uses the local-polynomial equivalent kernel:

The assumptions , , and give . Hence

Lemma 16 is the algebraic input behind the phase diagram in Theorem 3. It shows that the bandwidth is comparable to the boundary scale near the edge and to the usual thinned interior bandwidth when the threshold is asymptotically beyond that edge.

The total-Gram probability bound depends on population coercivity and elementary local-window identities. Lemmas 1, 2, 3, 4, 5, and 6 isolate these deterministic and integral facts before they are assembled in Lemma 17.

Lemma 8 [lem:good-gram-reproduction] (Good Gram reproduction).

Fix , a realized sample , a stratum , constants , and deterministic split blocks as in Algorithm 1. Use the local-design notation , , , , , and from Definition 12.

Assume that and that the good-design event holds for . Then the exact local-polynomial intercept weights reproduce every monomial coordinate at the threshold: for every ,

⊢ Lean
Proof of Lemma 8.

Write where the sum over all sample indices is the same as the sum over , since outside . Set On , , and the assumed gives . The good-design inequality is If , then the right side is zero, so and . Thus is invertible.

On , the weights are the equivalent-kernel weights, Consequently, for ,

Let be the -th coordinate vector. Since , Therefore

Lemma 1 gives a uniform lower eigenvalue bound for the shifted reference moment matrices. The bound is uniform in the shift parameter, which is what permits thresholds ranging from the boundary region to fixed positive interior points.

Lemma 9 [lem:intercept-weight-pointwise] (Pointwise intercept-weight bound).

Fix integers , a split , a realized sample , a stratum , constants , and an index . With , , , and as in Definition 12, assume:

  • (Split.) The split has three pairwise disjoint deterministic blocks , each of cardinality at least , as in Algorithm 1.

  • (Positive population Gram constant.) .

  • (Good design.) The good-design event holds at the realized sample .

Then the exact local-polynomial intercept weight at index satisfies

⊢ Lean
Proof of Lemma 9.

Let On , ; together with , this gives . The good-design inequality states that

First suppose that is active, meaning , , and . Put The coercivity inequality implies the inverse bound Indeed, , so and division by is harmless, with the case immediate.

By Lemma 7, Since , and hence

If is inactive, then the local kernel factor in the weight is zero on , so and the same bound follows because the right side is nonnegative.

Lemma 2 verifies the positivity of the Gram constant used in Definition 12. This makes the good-design event meaningful uniformly over the polynomial-thinning envelope.

Lemma 10 [lem:bandwidth-balance] (Eventual bandwidth balance).

Let denote the information-balance bandwidth from Definition 8 formed with . Suppose that:

  • (Hölder exponent.) The exponent appearing in Assumption 5 satisfies .

  • (Polynomial-thinning exponent.) The exponent appearing in Assumption 4 satisfies .

  • (Threshold cap.) .

  • (Threshold sequence.) is a deterministic real sequence satisfying for every .

Then, for all sufficiently large ,

⊢ Lean
Proof of Lemma 10.

Put and . The hypotheses give and . Since , choose so large that Because , this choice gives , and hence every satisfies . For , define The threshold condition gives , so because and is nondecreasing on . For continuity on , use the real-power convention . If , then on the whole interval and , which is continuous. If , then , and the map is continuous on because ; multiplying by the continuous factor gives continuity of . In both cases since . Continuity of on therefore gives a point with This point is strictly positive, since .

The same monotonicity gives strict increase of on : if , then , while Moreover by the preceding choice of , and , so Multiplying the strict inequality for and the weak inequality for the second power by this positive factor gives . Consequently the crossing set in Definition 8 is exactly Its infimum is , so . Thus, for every ,

Lemma 3 gives a simple lower bound on local design mass under the power envelope. It is used to compare the realized local count with .

Lemma 11 [lem:local-gram-coercivity] (Local Gram coercivity).

Fix an observed law as in Definition 5, with conditional-density, stratum-mass, and polynomial-thinning components as in Assumptions 2, 3, and 4. Fix constants , and , and a degree . Suppose that

  • (Design constants.) , , , and .

  • (Local window.) , , , and .

For an observation , set and, for , define the local monomial feature Then, with denoting the uniform population Gram lower-bound constant from Definition 12, every satisfies

⊢ Lean
Proof of Lemma 11.

Let The inequalities and give . Since , the condition is equivalent to . Hence Lemma 4 gives The left-hand side is .

For each pair , Lemma 5 gives the corresponding density-weighted local Gram entry. Summing those identities over and , and expanding , yields

The stratum-mass condition in Assumption 3 and give . Combining the displayed local mass identity for with the upper envelope in Assumption 4 gives The lower envelope in Assumption 4, together with , gives

By Lemma 2, the assumptions , , and imply ; in particular .

Applying Lemma 6 with this and coefficient vector gives

Combining the mass upper bound with and , then applying the scaled-window coercivity bound and the lower density-envelope bound, gives Using the Gram identity above for the final term proves

Lemma 4 expresses the local-window probability through the stratum mass and conditional density. It is the population counterpart of the count in Definition 12.

Lemma 12 [lem:good-gram-weights] (Good Gram weights).

Fix , a split with pairwise disjoint blocks satisfying for , an observed sample , and a stratum . Let . Suppose:

  • (Positive bandwidth.) .

  • (Positive population Gram constant.) The constant from Definition 12 satisfies .

  • (Good local design.) The good-design event from Definition 12 holds for .

Then the exact local-polynomial intercept weights and rescaled coordinates from Definition 12 satisfy, for every , They also obey the two deterministic weight bounds and, with denoting the local count in Definition 12,

⊢ Lean
Proof of Lemma 12.

Write On , , , and Thus is nonsingular: if , the coercivity bound with forces , hence . The resulting monomial identity is the good-Gram reproduction statement in Lemma 8.

For active observations with and , the weight is and all inactive terms in the sum over have weight . Therefore, for ,

Next let . The pointwise conclusion to be used is Lemma 9. For an active , set and . Coercivity and Cauchy–Schwarz give hence Since , and therefore for every active , while inactive weights vanish.

The number of active indices is . Summing the pointwise bound gives Squaring the same pointwise bound and summing gives

Lemma 5 identifies each population Gram entry with a one-dimensional integral over the local treatment window. Combined with the density envelope, it connects the empirical Gram matrix in Definition 12 to the shifted reference matrices.

Lemma 13 [lem:weight-energy-balance] (Balanced weight energy).

Let , let be an observed-data law on strata, let be a deterministic split, and fix a stratum . Let , and define the observed product sampling law by Suppose that:

  • (Model and sampling.) belongs to the observed model class in Definition 5 with fixed constants , including the polynomial-thinning, stratum-mass, and Hölder-regression components in Assumptions 4, 3, and 5; the observed sample is independently drawn from as in Assumption 1; and is a three-way split as in Algorithm 1.

  • (Numerical conditions.) , , , , , , and .

  • (Bandwidth and design.) , , and the uniform population Gram lower-bound constant from Definition 12 satisfies .

  • (Balance.) The bandwidth is information-balanced:

For the exact local-polynomial intercept weights from Definition 12,

⊢ Lean
Proof of Lemma 13.

Set Because , the event in the definition of is . Since and , this window is contained in . The conditional-density, stratum-mass, and polynomial-thinning components of Definition 5, equivalently Assumptions 2, 3, and 4, give For , the assumptions and imply . Hence Thus Since , and therefore The constants satisfy , , and , so .

Let . The admissible split-size condition in Definition 1 gives , and together with this yields Thus , and gives . Let be the local count in Definition 12, evaluated for the displayed . The count decomposition used for the weight energy is Indeed, on , Lemma 12 gives on , the weights in Definition 12 are zero. Consequently the energy is at most on the low-count event , while on its complement the good-design case gives and the case gives zero. Under Definition 2 and Assumption 1, the variables defining on the third split block are independent Bernoulli variables with mean . The Bernoulli lower-tail bound gives Integrating the pointwise decomposition gives Jensen’s inequality for the square root therefore gives

The balance condition rewrites the effective sample size as Together with and , this gives Also for , applied with . Therefore Taking square roots and using yields

Substituting the definitions of and gives exactly

Lemma 6 converts the normalized shifted-moment lower bound into the physical treatment scale. This is the population inequality used before applying concentration to the realized Gram matrix.

Lemmas 7, 8, and 9 record pointwise properties of the exact intercept weights on the good-design event. They are deterministic consequences of the total Gram matrix and the affine reproduction equations.

Lemma 14 [lem:stabilized-interval-honesty] (Stabilized interval honesty).

Fix integers and , deterministic split blocks , constants , and a threshold . Let be the information-balance bandwidth from Definition 8 evaluated at , and let be the stabilized interval from Algorithm 2 formed with split , bandwidth , and threshold . Suppose that:

  • (Split blocks.) consists of three deterministic, pairwise-disjoint subsets of with

  • (Regime constants.) The declared constants satisfy

  • (Sample size.) .

  • (Threshold range.) .

  • (Interior bandwidth.) and .

Then is an observed-sample measurable interval, and for every observed-data law in the clamp model of Definition 5, where is the unsmoothed observed-data clamp target from Definition 7.

⊢ Lean
Proof of Lemma 14.

Write , and let The regime assumptions imply , and the split-size and sample-size assumptions imply . The center in Algorithm 1 and every stratum contribution in Algorithm 2 are measurable functions of the observed sample. Hence the endpoint maps are measurable. The radius is nonnegative and , so these endpoints are ordered and define an observed-sample measurable interval.

Fix . Define By Definition 7 and the conditional-density component of Definition 5, Let and Hoeffding’s inequality on the first two split blocks gives The self-normalized weighted Hoeffding bound on the third block gives

Let Since ,

On , the retained-course and atom-mass deviations satisfy In the notation of Definition 12, set The regime assumptions give , , , and , so Lemma 2 yields

For a stratum with , the hypotheses , , and are therefore in force, and Lemma 12 gives the reproduction identities With these identities imply The inactive-window weights are zero, and on active summands and . The Hölder condition in Definition 5 therefore yields Since , projection onto is contractive around this target, and the decomposition gives On , the last absolute value is at most , so where is the local regression bias-plus-noise radius in Algorithm 2.

The two displayed deviation bounds imply the deterministic error bound To see the algebra, set and . The atom deviation gives Using Algorithm 1 and Definition 7 and contraction of the projection around the target , On , the product decomposition and the bounds , , and give On , both and lie in , so Together with , these are exactly the two branches of the interval contribution in Algorithm 2.

Because , the two displayed branches of the interval contribution place inside the clipped interval for every . Hence

Since was arbitrary, the same lower bound holds for every law in the stated model class, together with the measurability assertion proved above.

Lemma 7 gives the explicit intercept-weight representation for observations in the local stratum-window. It is the algebraic formula behind the threshold regression term in Algorithm 1.

Lemma 15 [lem:stabilized-risk-upper] (Stabilized risk upper bound).

Assume that the declared constants satisfy

  • (Strata.) .

  • (Smoothness and thinning.) and .

  • (Model radii.) , , , and with .

  • (Threshold and calibration.) and .

Then there is a constant , depending on these declared constants, such that the following holds. For every deterministic threshold sequence with for every , and for every deterministic split sequence in which each consists of pairwise disjoint subsets of with there is such that, for every , where is the model class in Definition 5, is the clamp target in Definition 7, is the total-Gram estimator in Algorithm 1, is the information-balance bandwidth for in Definition 8, and the rate is the frontier in Definition 9.

⊢ Lean
Proof of Lemma 15.

Let . From Lemma 17, choose constants such that the population Gram constant is positive and the following bound holds uniformly over , sample sizes, admissible splits, thresholds , and strata : if is the information-balance bandwidth from Definition 8 evaluated at , then

Define and The displayed constants are finite and nonnegative, and , by the regime assumptions and .

Fix satisfying for every , and fix an admissible split sequence . By Lemma 10, for all sufficiently large ,

After increasing , also . Then . Put Since and , we have and therefore

The balance identity gives Thus, using and , Indeed, with , the elementary bound gives , and .

For any , the projection is one-Lipschitz and the finite sum over strata gives the deterministic decomposition where all expectations on the right are under . The retained-course and atom-mass block averages are averages of -valued variables, hence

By Proposition 6 and ,

It remains to bound the local regression term. On , polynomial reproduction and the deterministic weight bound in Lemma 12 give the Hölder bias bound

The centered noise term satisfies because the conditional variance is bounded by and Lemma 13 controls the integrated square-root weight energy. On , both the fallback value and lie in , so the exceptional contribution is at most . Combining the good-event decomposition with the noise, bias, and tail bounds gives

Before the split-size and tail simplifications, the finite-sample bound is

Using the two split-block inequalities above and the local bound just obtained, uniformly in ,

The model class is nonempty. Take the centered Bernoulli law with , conditional treatment density , and outcome distribution Bernoulli with mean conditional on ; this law has regression and belongs to under the regime assumptions. Hence the same upper bound applies to the supremum over . Since ,

Finally, is exactly the frontier rate in Definition 9. This proves the asserted bound with the displayed .

Lemma 8 states the monomial reproduction property in the same notation used by the estimator. This property turns local-polynomial weighting into an estimate of the regression value at the threshold.

Lemma 16 [lem:bandwidth-phases] (Bandwidth phase bounds).

There exist constants with such that the following assertions hold. Let For every , every observed law , every , and every deterministic sequence , suppose that:

  • (Regime constants.) , , , , , , , , , and .

  • (Thinning.) The law satisfies the polynomial thinning condition in Assumption 4 with exponent and constants .

  • (Thresholds.) for every .

Let be the information-balance bandwidth from Definition 8 for . Then, for all sufficiently large , and If , then, for all sufficiently large ,

Moreover, for every , there exist constants , depending on , with such that, under the same regime-constant, thinning, and threshold conditions, if for all sufficiently large , then, for all sufficiently large ,

⊢ Lean
Proof of Lemma 16.

Set and choose At this initial choice of witnesses, the quotient is interpreted by the total real-division convention: it equals when and agrees with ordinary division when . Thus is a real constant for every value of . Since the base is positive, , and therefore Together with , this gives

Fix and a deterministic threshold sequence satisfying the hypotheses. The regime assumptions give , , and , hence . Let be the information-balance bandwidth from Definition 8; thus when the displayed set is nonempty, with fallback value when that set is empty. By Lemma 10, for all sufficiently large ,

The first balance assertion follows immediately:

Since , Dividing the balance identity by gives and therefore

Now assume On a sufficiently large eventual set, the balance identity holds, , and the convergence gives . Since on this set,

For those , , so the balance identity implies Taking -th roots gives

The comparison gives and hence

Consequently , and therefore Using the balance identity again, Thus

Under and , , so the chosen constant satisfies . With , the far-threshold comparison follows:

Finally fix . Set again using the same total real-division convention for the displayed quotient. Since , the positive-base property of real powers gives , and therefore Fix again and a deterministic threshold sequence satisfying the hypotheses, and let be the corresponding bandwidth from Definition 8. The regime assumptions again give and . Suppose that eventually On the eventual set where this bound and Lemma 10 both hold, , , and , so

The balance identity and give hence

Together with the threshold bound, this gives Therefore Since , the displayed inequality rearranges to

Under , , and , , so the chosen constant satisfies . Since , we obtain, for all sufficiently large ,

Lemma 9 gives the pointwise weight control implied by Gram coercivity. Together with Lemma 12, it bounds both individual leverage and aggregate variance.

Lemma 17 [lem:total-gram] (Total Gram stabilization).

Let be a positive integer and let the declared constants satisfy Then there are constants , depending only on these declared constants, such that the uniform population Gram lower-bound constant in Definition 12, evaluated at with , satisfies , and the following hold.

  • (Model and sampling.) as in Definition 5, and are sampled according to Assumption 1.

  • (Split and threshold.) are deterministic, pairwise-disjoint split blocks with for , and .

  • (Bandwidth and stratum.) is the information-balance bandwidth from Definition 8 at , and .

For the good-design event in Definition 12, On , the exact local-polynomial intercept weights in Definition 12 satisfy, for every , and also obey

⊢ Lean
Proof of Lemma 17.

Set The regime assumptions give , , and , so by Lemma 2.

Define and The constants are positive and, together with , depend only on the displayed regime constants.

Fix as in the statement, and write The regime assumptions give , , , and . The construction of in Definition 8 gives . Indeed, if the crossing set in Definition 8 is empty, then . If it is nonempty and is a member, then , , , and so . Hence the infimum is bounded below by this positive number and above by any member of the set, giving . Since , this also gives .

Let By Lemmas 4 and 3 and the thinning and stratum-mass assumptions,

For every , Lemma 11 gives Moreover and for every .

Let , and let be the elements of in increasing order. For a block vector , define and put The deterministic ordered-block map satisfies, for every measurable block event , because the selected coordinates are independent and each has law . Under this map, and, for every , where the last equality is the entrywise definition of the local Gram matrix in Definition 12.

Define the block good event The preceding identities give the exact event identity Indeed, is the indicator of the stratum-window event, is its local monomial coordinate, and , so and are exactly the count and quadratic form appearing in Definition 12.

If , the split condition gives , hence . The lower bound on gives . Applying the localized empirical Gram concentration bound to the independent block coordinates, with the population coercivity above and with and , gives

By the ordered-block product identity and , Combining with the lower bound on , Hence

If , then , , and imply Thus and the same probability bound follows from .

On , Lemma 12 gives, for every , which is the displayed reproduction identity. The same result also gives

Lemma 17 assembles the local-window identities, shifted coercivity, and concentration into the good-design probability bound used by the stabilized procedures. The exponential bound makes the singular-design fallback negligible on the scales used in Theorems 1 and 2, while the on-event identities preserve the usual local-polynomial intercept behavior.

The last auxiliary lemma for the appendix supports the continuity-only causal and observed targets. It establishes pointwise uniqueness of the continuous conditional-regression extension on the threshold envelope.

Lemma 18 [lem:continuity-regression-extension-unique] (Unique continuous regression extension).

Let be a positive integer, let , and let be a continuity-only observed clamp model with these constants, as in Definition 24. Suppose that

  • (Design constants.) The constants satisfy

  • (Continuous versions.) For , satisfies that is continuous on for every .

  • (Regression version property.) For ,

Then, for every ,

⊢ Lean
Proof of Lemma 18.

Fix . The two regression-version assumptions imply Define functions on the real line by The continuity hypotheses make these functions measurable. Set Then is measurable and The almost-sure equality of the two versions gives The conditional-density identity in Assumption 2 applies to this measurable subset of , so The stratum-mass condition in Assumption 3, together with , gives . Hence

It remains to justify the zero-integral argument on . The conditional-density measurability in Assumption 2 gives the needed measurability on the restricted set. Since , the set has finite Lebesgue measure. For Lebesgue-a.e. , the nonnegativity in Assumption 2, the upper envelope in Assumption 4, and the design inequalities and give Thus is integrable on . The nonnegativity part of Assumption 2 and the zero integral therefore yield The point has Lebesgue measure zero. For Lebesgue-a.e. with , the inclusion gives . The lower envelope in Assumption 4 and then give Combining this positivity with Lebesgue-a.e. on gives Consequently Lebesgue-a.e. on , and the definitions of and give

Finally, both functions and are continuous on . If they differed at some point of , continuity would make them differ on a nonempty relatively open interval in , which has positive Lebesgue measure, contradicting the almost-everywhere equality on . Hence Since was arbitrary, the conclusion holds in every stratum.

Lemma 18 justifies the pointwise notation for the continuity-only regression extension used in Definition 25. The positive treatment-density envelope gives enough support throughout for almost-everywhere equality of continuous versions to determine the extension at every threshold value.

Verification note.

The Lean kernel checks were run at repository commit 6f1bf59b1255336e29c22bdea39edd530cddc960 using the CausalSmith Lean toolchain leanprover/lean4:v4.33.0; the executable reports Lean 4.33.0, target x86_64-unknown-linux-gnu, Lean commit d8b18978322de05a8f3dba51ef03cf5461676c17. The kernel-checked assumption, class, definition, and algorithm anchors in this manuscript are the objects introduced earlier in the paper. The kernel-checked result and open-question anchors are the objects introduced earlier in the paper. Presentation-synthesized restatements such as Definitions 1, 2, 3, 4, 10, 11, 13, 14, 15, 16, 17, 18, 19, 23, 29, 31, and 32 are reader-facing anchors whose status is recorded separately in the verification contract. Remark 1 contributes a kernel-checked, well-formed encoding of the open sharp-constant question in the formal layer. Statement- and proof-faithfulness audits compare the displayed manuscript statements and proof narratives with those checked declarations; those audits are separate from Lean kernel checking. The theorem-local verification footnotes identify the published conditional-distribution input from Kallenberg (2002) and the bounded-difference concentration input from Hoeffding (1963) at the exact statements where those cited conclusions are invoked.

Reproducing the verification.

The material needed to reproduce the checks is organized as follows.

Archived source. The Lean development is released in the public CausalSmith repository at https://github.com/Jiyuan-Tan/CausalSmith; the state checked for this manuscript is the commit recorded above, and a snapshot of that commit is deposited in a persistent archive with a DOI at publication so that the checked state remains citable independently of the repository history. All declarations for this paper live in the single module tree CausalSmith/Stat/STAT_LmtpThresholdAtomFrontier_Research.

Dependencies. The build needs the Lean toolchain leanprover/lean4:v4.33.0, pinned by the repository’s lean-toolchain file and installed by elan, together with the Lean build tool lake, which is distributed with the toolchain. The only non-Lean prerequisites are git and curl. There are two Lean package dependencies, both pinned in the repository’s lockfile: the Causalean library, on which this development builds, and Mathlib, on which Causalean builds. No proprietary component, no external solver, and no network access beyond the initial fetch is required.

Build and check. From a clean environment, cloning the repository, checking out the recorded commit, fetching the pinned dependencies with lake exe cache get, and building the module tree above with lake build reproduces the kernel checks. A successful build is not by itself the claim made here: the reported status additionally requires that the sources contain no sorry, no admit, and no added axiom, and that #print axioms on each anchored declaration reports only the three standard Lean axioms. Both checks are scripted in the repository and are rerun against the archived commit.

Anchor-to-declaration manifest. The bundle accompanying this manuscript ships a machine-readable crosswalk that maps every numbered environment in the paper to the declaration that certifies it. Each record carries the manuscript anchor, the environment type and displayed number, the fully qualified Lean name, the declarations it uses, and a status field separating kernel-checked anchors from the presentation-synthesized restatements listed above, which carry no standalone declaration. A companion file records, for each anchor, the verbatim source text of the certifying declaration, so a reader can compare a displayed statement with the checked one without rebuilding the library. The interactive edition of this paper renders that manifest inline at each numbered environment.

Proofs of the main results

Proof of Theorem 1.

Fix the displayed regime constants. We use seven steps.

  1. Write for the observed minimax risk over the admissible estimator class used here, where the infimum ranges over observed-sample measurable estimators satisfying for every observed sample . For a split , put This is the finite-sample stabilized risk in Definition 16, with the product sample law interpreted as in Definition 2. Lemma 15 gives a constant such that, for every admissible deterministic and , for all sufficiently large , The estimator is observed-sample measurable and -valued. Hence it is an admissible competitor in the infimum defining , so

  2. We record the two-point device used for both lower bounds. Let , put , and suppose with well-posed product chi-squared divergence and For any admissible -valued observed-sample estimator , the testing inequality and give Multiplying by , the larger of the two risks under is at least The supremum over in the restricted minimax problem dominates these two risks for each ; taking the infimum over admissible yields

  3. For the parametric component, use the canonical design and Bernoulli regressions Both laws belong to . The stratum masses are , which meets the stratum-mass condition of Assumption 3 under the regime constraint ; the canonical density meets the conditional-density condition of Assumption 2 and, by the regime constraint , the polynomial-thinning envelope of Assumption 4; and each regression is constant, so every Taylor remainder in Assumption 5 vanishes and the Hölder condition holds for any . Finally for all , so both constant regressions take values in . The clamp target shifts by The product chi-squared divergence is well posed, and With , the two-point mixture bound gives a constant , depending only on the regime constants, such that for every admissible and all sufficiently large .

  4. For the localized lower-bound component, use a fixed dose bump centered at zero, with Let . Since , the number belongs to . Because is smooth and compactly supported, there is a finite constant such that Choose so small that For any and , define Then is continuous on , and there because For , differentiating the scaled bump times and using gives The one-dimensional Taylor remainder theorem on therefore yields Thus the localized Bernoulli regression satisfies the member condition in Assumption 5 uniformly over and . Set Then and For the contracted bandwidth , form Bernoulli laws on the same canonical design. For all sufficiently large , Lemma 10 gives Together with , this gives , , and . The displayed Taylor bound places both laws in . The target separation obeys The one-observation chi-squared increment satisfies Using the balance identity above, , and , Let be the one-observation chi-squared increment for this Bernoulli perturbation. The product law factorization gives Since and , the inequality yields Applying the two-point bound gives a constant such that for every admissible and all sufficiently large .

  5. The global witness in the theorem uses the same canonical design and For all sufficiently large , , so these are Bernoulli laws in , with shared -design, shared , and shared . They satisfy for every and , and The one-observation chi-squared increment is . Hence the product divergence is well posed and satisfies so

  6. For the localized witness appearing in the statement, use the same fixed bump . The amplitude construction above also permits a separate witness amplitude: choose such that For every and , the function is continuous on , takes values in , and the same scaled-derivative calculation gives Set For itself, define For all sufficiently large , Lemma 10 gives Together with , this gives , , and , so . The pair has Bernoulli outcomes, common design, common , common , and for every and . Since and the perturbation is nonnegative, The same integral calculation and the balance identity give Let be the one-observation chi-squared increment. The product law factorization gives The perturbation is bounded by , so the Bernoulli likelihood ratio is finite and square-integrable for one observation; absolute continuity and square integrability tensorize to the product laws, giving a well-posed product chi-squared divergence. Since and , the inequality gives

  7. Now set The amplitude asserted in the theorem is . Then , and , , , and . Increasing the eventual index so that all preceding bounds hold, put The two lower components imply Since , these inequalities and the choices of give Together with the admissibility and upper bounds above, By Definitions 16 and 2, is the finite-sample worst-case risk of the stabilized estimator computed under the observed i.i.d. product law. The global witness constructed above satisfies because and . The localized witness constructed above satisfies because and . These are the asserted eventual bounds and witness properties.

Proof of Theorem 2.
  1. Let and The expected-length upper bound used below is the following eventual statement: there is a constant , depending only on the declared regime constants, such that for every admissible threshold path and split path, for all sufficiently large . To see this, use Algorithm 2. Intersecting the symmetric interval with weakly shortens its length, and therefore where and is the stratum contribution in Algorithm 2. The split lower bounds in Definition 1 give once .

    For a fixed law and stratum , the integrated stratum contribution is bounded by Here is the finite -weight constant supplied on by Lemma 17. The displayed product bound is the sample-splitting step: the empirical atom estimate uses , the local weights use , and the mixed term factors as For the empirical atom envelope, define Then . The mean of one summand is The second equality is the conditional-density component of Definition 5, and the third is the definition of in Definition 7. Since , . Independence on the deterministic second block therefore yields Moreover, Proposition 6 gives , and , so The triangle inequality gives the integrated envelope Also, Lemma 13 gives for a finite constant , and Lemma 17 gives By Lemma 10, for all sufficiently large , The effective local sample size has the following lower bound. Put Since , , and , Because , this gives . Multiplying the balance identity by gives and therefore, since , Exponential decay along this polynomial scale is eventually bounded by . Combining the exponential-decay bound with the polynomial scale bound for , and using , yields uniformly in and , Since is fixed, summing over strata gives the displayed bound.

  2. Put . The lower-bound input is a two-point inequality for expected length. Because an arbitrary interval procedure need not have finite expected length, the worst-case length in this step is taken in , with denoting the lower integral of a nonnegative measurable function; for the bounded intervals of interest this is the ordinary expected length. If satisfy the likelihood-ratio square is -integrable, and , then every uniformly honest observed-sample interval procedure satisfies where the supremum and the lower integral are taken in and the nonnegative real number on the left is read as an element of as in Definition 23. The comparisons in this step are made in ; the final paragraph returns to the ordinary expected-length notation of Definitions 18 and 15 once the interval is known to be bounded. Indeed, coverage at and , transferred from to by gives . On that event the interval length is at least , and integrating this indicator gives the displayed extended-real bound.

    For the root component, set Let and have the common design and, conditionally on , let be Bernoulli with success probability For all sufficiently large , . The two laws therefore have outcomes in , the same density , the same masses , and constant Hölder regressions; the regime inequalities place both laws in . Their clamp-target gap is exactly For one observation the chi-square divergence is , hence independence gives Since , The two-point length inequality above, with , yields Thus, for some , eventually.

    For the local component, use the fixed compactly supported bump from the same two-point construction, normalized by The smoothness calibration for this particular bump gives an amplitude , depending only on , with the following uniform property. For every and , define Then is continuous on , takes values in , and satisfies the Hölder Taylor remainder for all . Indeed, applying the normalized Hölder gate of with radius and then shrinking the amplitude to be at most gives both the displayed Taylor bound and the range constraint.

    Set Let and have the same common design as the root pair, and let their conditional Bernoulli success probabilities be , where For all sufficiently large , Lemma 10 gives , hence , and the amplitude calibration places both laws in . The localized shift is nonnegative and equals at , so the collapsed atom contribution alone gives For one observation, write where is the common design law. Since and vanishes outside , Because , . With , , and the balance identity from Lemma 10, Independence gives The two-point length inequality therefore gives, for some , eventually. Taking half the smaller of and , and splitting into the two cases according to which component is larger, gives a constant such that eventually.

    The every-procedure version uses that a uniformly honest procedure is an admissible competitor in the minimax honest-length criterion. Hence there is a positive constant such that every uniformly honest observed-sample interval procedure satisfies in .

  3. Define Then , , , , and .

  4. Fix an admissible threshold path and split path . By Lemma 10, for all sufficiently large , Together with , this gives . Hence the hypotheses of Lemma 14 hold. That result supplies observed-sample measurability of and the modelwise coverage inequality for every , so the stabilized interval is uniformly honest. Taking the infimum over gives the displayed worst-case coverage bound.

  5. Put Since and , we have . The upper ingredient and yield The minimax lower ingredient and yield the extended-real comparison

    Since is itself uniformly honest by the coverage step, the procedure lower bound applies to this concrete interval. Combined with , it gives in . For the concrete stabilized interval the length is measurable, nonnegative, and bounded by one, so its worst-case length in agrees with its ordinary worst-case expected length, read in as in Definition 23. Since , this gives the real lower bound

    Finally, for every uniformly honest observed-sample interval procedure , the procedure lower bound gives the extended-real statement with the supremum taken in . This is the arbitrary-procedure worst-length criterion used in the comparison. For interval procedures whose pointwise lengths lie in as in Definition 15, the nonnegative lower integral equals the ordinary expected length under the positive-part embedding of Definition 23, and the display specializes to the corresponding real expected-length lower bound. These are the eventual length comparisons established by the argument.

Proof of Theorem 3.

Let The regime assumptions give , , and .

  1. The two deterministic scales satisfy for every . The exponent is positive because Hence , which proves

  2. Fix a deterministic threshold sequence with . Let be the information-balance bandwidth from Definition 8, define and let be the frontier rate in Definition 9, so that . By Lemma 10, for all sufficiently large , The phase comparisons below use this deterministic balance relation directly. In the regular regime , we also have, for all sufficiently large , If , then because . If , set Since and , the balance equation implies Therefore The exact power computation is Thus eventually, and eventually. Hence .

  3. Suppose with . By the scale separation , For this above-edge regime, the balance equation gives the closed-form bandwidth order. Indeed, for all sufficiently large , and . With , the balance equation gives . The inequality gives , and hence . Substituting this upper bound into the same balance equation yields Consequently Multiplying by the nonnegative factor and raising the bandwidth comparison to the positive power gives The normalized interior term satisfies Consequently . Since and both summands are nonnegative, also

  4. Suppose and . The scale separation again yields so the same direct balance comparison used at the critical scale gives Thus Moreover, Hence the root term is eventually bounded by a constant multiple of the displayed atom scale, and

  5. Suppose with . Since we have Together with the scale separation , this places the sequence above the edge scale. The same direct balance comparison used in the supercritical case gives For all sufficiently large , , and Because , it follows that

  6. Fix and . At threshold , Definition 7 gives, for every stratum , The conditional-density condition in Assumption 2 gives for every stratum . Summing over the finite stratum set yields , and the treatment support encoded by Definition 5 gives almost surely. Therefore the retained term in Definition 7 satisfies Since every is zero, Finally, , so Definition 9 gives

  7. Fix , , and an admissible split as in Definition 1. The conditional-density condition in Assumption 2 gives , so under almost surely. Since the product law in Definition 2 has each observed coordinate distributed as , every coordinate satisfies under almost surely, and the finite intersection of these coordinate events still has probability one. On that event, the finite-block average is interpreted as the totalized expression which is the displayed product with the real-inverse convention . Thus while Substituting these identities into Algorithm 1 gives, -almost surely, Thus the threshold-zero total-Gram estimator agrees under the product sampling law with the clamped block-zero outcome average.

  8. Apply Theorem 2 to the deterministic sequence and to a deterministic admissible split sequence . The identity at threshold zero gives eventual worst-case coverage and constants such that, eventually, the concrete stabilized length from Definition 16 satisfies Likewise, applying Theorem 1 to the same zero sequence and using gives constants such that, eventually, the concrete stabilized risk from Definition 16 satisfies Therefore the stabilized zero-threshold procedure has coverage at least for all sufficiently large , and its stabilized worst-case absolute-error risk and stabilized worst-case expected length at threshold are both .

Proof of Proposition 1.
  1. Put so that, on the eventual range where , Let For the infimum construction in Definition 8, set The bounds and make continuous and strictly increasing on , with . For all sufficiently large , , so the crossing set in Definition 8 is a nonempty interval whose left endpoint is the positive solution of . Hence, for all sufficiently large , Since , eventually , hence and The balance equation gives and therefore . The inequality is exactly Consequently , and then so eventually. Also , and the same balance equation forces eventually, since and would imply Thus This is the one-stratum far-edge instance of the bandwidth comparison in Lemma 16, specialized to the density and the condition .

  2. For all sufficiently large , the support of is contained in . On the left and right halves of this interval, and Hence

  3. Because and eventually, For every real with , the Bernoulli comparison used here is The lower bound follows from Pinsker applied to the one-point event , and the upper bound follows from the exact Bernoulli KL formula on the interval . Integrating this display against the nonnegative density and using the local perturbation-energy computation gives eventually. Therefore

  4. The collapsed atom mass is and . The retained-course contribution is supported on , so Thus Since eventually, eventually, and hence

  5. Let For all sufficiently large , , and direct exponent algebra gives Combining the bandwidth comparison from the first item with the target comparison from the fourth item gives constants such that eventually If , then for some , eventually. Since is increasing on , eventually, so .

    Conversely, assume . Then for some , eventually. With the two-sided display for , together with , yields eventually. Hence eventually, which is . This proves

Proof of Proposition 2.

Stratum product identity and the cell formula. Fix . From the regime constants and Definition 20, Let be the latent-coordinate marginal of the normalized full-data stratum law . The exchangeability member of Definition 20, together with Assumption 2 for the observed margin, gives the product formula Consequently, for every Borel , consistency on the common full-data event gives The measurability, boundedness, and restriction-to- bookkeeping in this calculation is supplied by the jointly measurable structural response, the simultaneous consistency event, and the bounded -valued response range in Definition 20.

Pointwise identification of the threshold-range regression. For every Borel , the stratum product formula for established for this stratum and equality of the observed margin give where the second equality is the conditional-expectation version property in Assumption 5 combined with Assumption 2. The products are integrable on : both regression curves are continuous there, and the polynomial envelope in Assumption 4 bounds by on . Hence Since and , the same thinning envelope gives Thus Lebesgue-a.e. on . Both functions are continuous on , so equality extends to every : This also covers the endpoint by right-continuity on the closed threshold interval.

Support preservation. Fix . The observed margin is in , and the regime constants satisfy the hypotheses of Proposition 7. Therefore, for each ,

Pathwise clamp decomposition. Because , the common consistency event applies both at and at . On that event, if , then and ; if , then and . Hence

The atom contribution. Define, for a full-data state , The bounded outcome and potential-outcome range make these integrable. The observed-margin identity gives For the lower-tail term, partition over the finite strata and use the stratum product formula for at the fixed dose :

Target equality. Integrating the pathwise clamp decomposition of and substituting the two integrated clamp contributions and yields Since , the identification on gives for every stratum, and therefore All endpoint cases in the closed range , including and , are covered by the same closed-interval hypotheses used above.

Proof of Proposition 3.
  1. The regime assumptions give , , and . Fix . Define the total regression extension The model assumptions in Definition 5 give that is Borel measurable, for every , and They also give continuity of on for every , since , and positivity of every stratum mass:

    Let , let By Kallenberg (2002), applied to the standard Borel design and outcome spaces, let be a conditional-distribution kernel of given , so that for all Borel and , Put Since -a.s., and the conditional-expectation identity gives equivalently -a.e. Hence the Borel set satisfies . Define the repaired kernel Then is a probability law supported on for every , -a.e., and

    Let be Lebesgue probability measure on . Realize the repaired kernel by its generalized inverse: for every and , let be The strict superlevel sets of this supremum are countable unions over rational , so is Borel. The right-continuity of gives the generalized-inverse identity, for every , Therefore

    On the product space define Because -a.e., the law of is exactly . The displayed structural identities and also give the simultaneous consistency condition in Definition 20. For bounded Borel and , the product construction yields, for every , which is the exchangeability condition in Definition 20. Finally, for , Thus is continuous on . For the remainder of the proof, write The constructed package belongs to , and its observed margin is . Hence This proves observed-margin surjectivity.

  2. Fix , including the empty-sample case , and fix . For each , Proposition 2 gives the target identity The causal criteria in Definition 22 evaluate every observed-sample estimator, coverage event, and length objective under the product law of the observed margin, in the notation of Definition 2. Hence, after the displayed target identity, each causal risk, coverage probability, and expected length depends on only through .

  3. Let be any observed-measurable estimator with values in . The target and sampling-law identity gives Taking suprema and using the surjectivity already proved, The admissible estimator class is the same on both sides, so taking the infimum over and using Definitions 22 and 18 gives

  4. Let be an observed-measurable interval procedure. The same target and product-law identities give Therefore, using surjectivity again, Thus the uniformly honest interval procedures are exactly the same for the causal and observed criteria. Their length objectives also agree: Taking the infimum over the common feasible class and using Definitions 22 and 18 gives

Proof of Theorem 4.
  1. For standard Borel spaces and every probability law on , the usual disintegration gives a probability kernel such that, for measurable and , This standard Borel disintegration supplies the regular conditional law needed for the observed-margin comparison in Proposition 3. Let Under the regime conditions, Proposition 3 gives surjectivity of from onto and, for every and every , For every and every such , Proposition 2 gives the target identity

  2. Let be the constants supplied by Theorem 1, and let be the constants supplied by Theorem 2. Set Then . Choose so that, for every , the eventual conclusions of Theorem 1, Theorem 2, and Lemma 10 hold, and also . Fix such an , and write Since , Lemma 10 gives and hence Moreover .

  3. The total-Gram estimator in Algorithm 1 is a measurable function of the observed sample, so is observed-sample measurable. The stabilized interval in Algorithm 2 has measurable endpoints built from the same finite sums and local-polynomial weights. The center is clipped to , and the displayed radius is nonnegative, so the clipped lower endpoint is no larger than the clipped upper endpoint. Thus is an observed-sample measurable interval.

  4. For the concrete estimator, the definitions of the full-data and observed risks, the identity , and the surjectivity of give Combining this equality with Theorem 1 and the choices , yields

  5. The same target transport gives the concrete coverage identity The right-hand side is at least by Theorem 2, so

  6. Expected length is transported by the same observed-margin argument. Since is a function of the observed sample, its length depends on only through , and the surjectivity of gives

    For every observed law , the clipping in Algorithm 2 gives Thus the real expected-length bound from Theorem 2 gives where the last inequality uses and . Combining this bound with the equality of the causal and observed expected-length criteria gives The same inequality may be viewed in through the order-preserving embedding of finite nonnegative real numbers; this is the extended order used below for the minimax length criterion.

  7. Finally, Lemma 14 makes an admissible honest observed-sample interval, since , , , and . Hence, by the definition of and the expected-length bound just proved, with the two sides read in by embedding the finite nonnegative quantities as finite extended values. The lower bound from Theorem 2, together with and , gives, in the same extended order, Substituting back , , and gives all displayed conclusions for every .

Proof of Proposition 4.

Write for the observed margin of , and write . The hypotheses in Definition 27 and the design constants give In particular, for every stratum, , and .

  1. Fix . Let be the continuity-only regression extension selected in Definition 11; the observed-margin component of Definition 27 supplies the membership of in . For every measurable , the full-data product identity from consistency, exchangeability, and the conditional-density law gives Indeed, under the normalized full-data law restricted to , the treatment coordinate has the conditional law with density on . Its latent-coordinate marginal is the pushforward equivalently, for every measurable , With this normalization, Assumptions 6, 7, and 2 give the product factorization of the treatment and latent coordinates in stratum . Applying that factorization to , and using Definition 19, identifies the stratum integral with The observed-margin identity gives The regression-version property of in Definition 11, together with the conditional-density law, gives Combining the three displays and using , The two products are integrable on : Definition 19, Assumption 8, and Definition 11 give continuity of the two mean functions on this compact interval, and Assumption 4 gives the envelope almost everywhere on . Equality of all measurable set integrals therefore gives For Lebesgue-a.e. , Assumption 4 also gives Since the endpoint has Lebesgue measure zero, the two continuous functions and agree Lebesgue-a.e. on . Applying the continuous-version uniqueness principle in Lemma 18,

  2. Fix . Since , also . On the full-data carrier define The common consistency null set in Assumption 6, the observed-treatment support supplied by the observed-margin component of Definition 27 and transferred to the full-data law, and Definition 6 give On the branch , treatment support puts in , the clamp policy returns , and consistency identifies with the observed outcome . On the branch , the clamp policy returns , so the displayed decomposition is the corresponding algebraic split. The retained summand is measurable because it is built from the observed outcome and the measurable threshold set . The atom summand is measurable because the potential-outcome process in Definition 27 is jointly measurable and is measurable. For integrability, the observed outcome support bounds by one almost surely, and Assumption 6 identifies with the -valued structural response on the consistency event; hence is also bounded by one almost surely.

  3. The retained summand depends only on the observed coordinates. The observed-margin identity therefore gives where is defined in Definition 7.

  4. For the atom summand, decompose over strata and use the conditional product law for , now with the dose argument fixed at and the treatment event : By Definition 7, so

  5. Combining the pathwise decomposition, integrability, and the two integral identities for the retained and atom summands, and using Definition 21, Applying the pointwise identity on at gives The right-hand side is exactly by Definition 25. Since was arbitrary, the asserted equality holds throughout the declared threshold range.

Proof of Proposition 5.
  1. The continuity-regime hypotheses give and also . The displayed inequalities are the continuity design constants needed in Proposition 4. The strict inequality , together with Assumption 3, gives positive mass to every stratum of any .

  2. Fix . Put and choose a measurable conditional-mean version whose restriction to is the continuous extension in Definition 11. Take a Borel regular conditional response kernel for given . Equivalently, for Borel and , Let Because almost surely, the raw conditional kernel and its clipped pushforward agree for -almost every , so their conditional means agree there. The conditional-distribution formula for conditional expectation, combined with the defining conditional-expectation version in Definition 24, gives directly Define the total measurable response mean Then for every , and is continuous on for every . Moreover, On the strip, the displayed almost-sure equality and show that clipping leaves the conditional mean unchanged. Off the strip, the equality is the definition of .

  3. Let and define The almost-sure agreement of with the outcome mean of gives , so replacing by preserves the joint observed law. Also for every . With denoting Lebesgue measure on , define The strict-superlevel sets of this supremum are countable unions over rational thresholds, so is jointly Borel measurable. The generalized-inverse identity on left rays gives Construct on the latent carrier by first sampling , taking the latent variable to be , and setting Define the structural response maps and potential outcomes by The simultaneous definition gives Assumption 6. To verify Assumption 7, it suffices in this finite-stratum setting to check the bounded-Borel product-moment factorization used by the construction. For every stratum , every bounded Borel , and every bounded Borel , let . The product law gives and . Substitution yields The observed margin is : the conditional law of given is , this kernel equals for -almost every , and agrees -almost surely with the original conditional response law because -almost surely. Finally, for every stratum and every , Definition 19 and the quantile identity give where follows from and Assumption 3. Since is continuous on , Assumption 8 holds. Therefore with the observed margin understood as in Definition 29. This proves surjectivity.

  4. Now fix and . For every , the design constants recorded above allow Proposition 4 to identify the causal and observed continuity targets: The observed product sample law in Definition 2 depends on only through the observed margin . Hence, for every observed-sample estimator satisfying the measurability and range requirements in Definition 28, The inclusion from left to right uses ; the inclusion from right to left uses surjectivity of the observed-margin map onto . Taking the supremum of equal value sets and then the same infimum over observed-sample estimators in Definition 28 yields The argument also covers , since it uses only equality of the defining value sets.

  5. The interval criteria are handled in the same way. For every observed-sample interval procedure satisfying the measurability requirement in Definition 28, the target identity and surjectivity give the equivalence if and only if Using the pointwise interval length from Definition 15, the same observed-margin argument gives equality of expected-length value sets: Taking the supremum over laws and then the infimum over honest interval procedures in Definition 28 yields Together with the risk equality, this proves the two displayed identities for every and every .

Proof of Theorem 5.

1. Lower experiments for the observed criteria. For and a threshold sequence write as in Definition 26. The lower-bound witnesses all use the same design: Since , , , and , this design satisfies the stratum-mass and thinning requirements in Definition 24. For the root term, set The two continuous regressions lie in , their targets differ by and the product chi-square divergence obeys For the atom term, let be the fixed smooth treatment bump with and whenever . Put and choose With define by regression and by regression . The bump properties give a continuous perturbation with , support contained in , and value . Therefore the atom contribution at the clamp point gives The same localization and boundedness give the design-integral estimate behind the Bernoulli chi-square calculation, under the displayed common design. Since and , the exponent is at most , and hence We use the following two-point risk bound. For probability laws on the sample space and real target values , every estimator satisfies To verify this, fix a probability measure dominating both laws and define the common-part measure Its total mass is , and integration of against gives the display. When , Cauchy–Schwarz applied to the likelihood ratio gives For the root pair, the product bound gives ; for the atom pair it gives . Hence the factors are positive in both two-point applications. The two target separations therefore yield constants such that, eventually, Taking gives For honest length, set For the root length witnesses, let and be the global-shift pair with . Then The second inequality follows from , at , and . For the atom length witnesses, let and be the localized-bump pair with the same amplitude and choose The localized-bump exponent satisfies so the same exponential-to-polynomial calculation gives For either length-witness pair, write the first law as , the second as , and the two target values as . If an interval covers under with probability at least for , then total variation and the union bound give Together with , the displayed chi-square bounds give On the event , the interval length is at least . Thus there are with in the extended nonnegative-real order, where nonnegative real lower bounds are embedded as finite extended values. With , the combination is by the case split or its complement, yielding

2. Upper risk for the fixed fallback estimator. Fix , a sample split , and . Define The estimator in Algorithm 3 is where . By Definition 25, the target atom contribution is Since , projection is contractive toward the target. Adding and subtracting the fallback atom term gives Here by the thinning envelope integrated over . The retained and atom indicators are -valued block averages, so Cauchy–Schwarz and the variance bound for give Integrating the pointwise bound therefore gives the explicit estimate For all , admissibility of gives for . Hence, uniformly over , for a constant depending only on the declared constants. The same estimator is an admissible competitor in Definition 28, so

3. Coverage and length of the continuity interval. Let The interval in Algorithm 4 is The coverage argument uses the following unit-range specialization of (Hoeffding, 1963): an average of independent -valued variables has two-sided deviation probability at radius at most . The total empirical atom mass is itself one such average, because the strata partition the sample space: Its population mean is Applied to the retained-course average and to this total lower-threshold atom average, with the radii in Algorithm 4, the Hoeffding bound gives On the complement of the two events, set The regression values lie in , so the unknown clamp atom contribution satisfies Since the concentration event also gives , the quantity lies between and . Hence Using Definition 25 and the contraction of projection onto toward , this gives so . The union bound yields uniform coverage at least . Also, Moreover, the displayed identity for gives Write for the canonical embedding of a nonnegative real into the extended nonnegative reals. Consequently, before the final uniform simplification one has the extended-real length bound Since , this immediately implies the deliberately weaker bound The first display follows from the pointwise inequality and the preceding expectation calculation; the second only enlarges the finite real upper bound by a nonnegative term. With and , this gives a constant such that and the same comparison holds in the extended nonnegative-real order after embedding the real bound as a finite extended value.

4. Transport to the causal continuity class. Let denote the observed margin in Definition 29. Since the estimator and interval are functions only of the observed coordinates, for every nonnegative or integrable observed-sample functional , and the analogous identity holds for observed-sample events. The causal minimax criteria in Definition 28 are defined with the full-data product law. For the observed-only estimator and interval used here, the preceding pushforward identity gives the corresponding fixed-procedure risk, coverage, and expected length as the same quantities computed under the observed-margin product law. By Proposition 4, every and every satisfy By Proposition 5, every is the observed margin of some member of , and the observed and causal criteria in Definition 28 agree exactly: The bridge identity and observed-margin surjectivity also identify the fixed-procedure observed and causal worst-case quantities. For the fallback estimator, For the interval, and the expected lengths satisfy

5. Choice of constants and eventual frontier bounds. Let be the constants obtained above and define Then . Intersecting the eventual sets from the lower-risk, lower-length, upper-risk, and upper-length bounds with the sample-size threshold gives, for every admissible threshold sequence and every admissible split sequence, all sufficiently large satisfying; at those sample sizes the positivity of and used in the coverage and upper-bound arguments follows from split admissibility. the asserted observed coverage and length upper bounds, and, in the extended nonnegative-real order, The equalities and causal bounds follow from the transport identities above with the same , , estimator, and interval.

6. Elbow, zero threshold, and fixed positive thresholds. Since , for , and therefore At , the identity gives If , then continuity of at the positive point gives The fixed-positive-threshold argument applies the atom-scale lower constructions directly. For the risk criterion, there is a constant such that eventually If , passing to limits in this eventual inequality would give , contradicting . For honest length, the corresponding lower construction gives a constant with eventually in the extended nonnegative-real order. Convergence of to zero would imply , again impossible. Hence neither fixed-positive-threshold criterion tends to zero.

Proof of Proposition 6.

Fix and . By Definition 5, the conditional treatment law in stratum is represented on by the density . By Definition 6, the clamp map is measurable and satisfies Splitting the conditional treatment law at and its complement gives Because this conditional treatment law is carried by , the retained part is The same density representation, together with , gives where the last equality is the definition of in Definition 7. These identities give the asserted pushforward decomposition.

It remains to bound . The conditional-density and polynomial-thinning conditions in Definition 5 give integrability of on and, for Lebesgue-a.e. , Since , the function is integrable on . Integrating the almost-everywhere inequalities yields The regime condition implies , and with , Substitution gives as required.

Proof of Proposition 7.

Fix the stratum , and let Thus is the conditional natural-treatment measure in stratum . The endpoints are -null, so By Assumption 4, for -almost every , because , , and for . Hence -almost everywhere. Since is obtained from by this density, , and therefore Fix such an , and put using Definition 6. Since , , and , we have Let be open with . Choose such that , and set Then and Consequently, It remains to transfer this positivity from to . For a measure with density, the zero-set identity gives, for every open , If , this identity gives Because -almost everywhere, the set in this display agrees with up to a -null set, contradicting . Thus . Since was arbitrary, is a support point of the conditional natural-treatment measure in stratum . The choice of was from a full -measure set, which gives the claimed almost-everywhere support preservation.

References

  • Díaz, Iván and Williams, Nicholas and Hoffman, Katherine L. and Schenck, Edward J. (2021). Nonparametric Causal Effects Based on Longitudinal Modified Treatment Policies. Journal of the American Statistical Association. doi
  • Hoffman, Katherine L. and Salazar-Barreto, Diego and Williams, Nicholas T. and Rudolph, Kara E. and D{\'i}az, Iv{\'a}n (2024). Studying Continuous, Time-Varying, and/or Complex Exposures Using Longitudinal Modified Treatment Policies. Epidemiology. doi
  • Williams, Nicholas T. and D{\'i}az, Iv{\'a}n (2023). {lmtp}: An {R} Package for Estimating the Causal Effects of Modified Treatment Policies. Observational Studies. doi
  • Ga{\"i}ffas, St{\'e}phane (2005). Convergence Rates for Pointwise Curve Estimation with a Degenerate Design. Mathematical Methods of Statistics. arXiv
  • Timothy B. Armstrong and Michal Kolesár (2016). Simple and Honest Confidence Intervals in Nonparametric Regression. Quantitative Economics. doi
  • Armstrong, Timothy B. and Koles{\'a}r, Michal (2018). Optimal Inference in a Class of Regression Models. Econometrica. doi
  • Kennedy, Edward H. and Ma, Zongming and McHugh, Matthew D. and Small, Dylan S. (2017). Nonparametric Methods for Doubly Robust Estimation of Continuous Treatment Effects. Journal of the Royal Statistical Society: Series B. doi
  • van der Laan, Lars and Zhang, Wenbo and Gilbert, Peter B. (2022). Nonparametric Estimation of the Causal Effect of a Stochastic Threshold-Based Intervention. Biometrics. doi
  • Bonvini, Matteo and Kennedy, Edward H. (2026). Fast Convergence Rates for Dose-Response Estimation. Electronic Journal of Statistics. doi
  • McCoy, David and Zhang, Wenxin and Hubbard, Alan and van der Laan, Mark J. and Schuler, Alejandro (2026). Data-Adaptive Identification of Effect Modifiers Through Stochastic Shift Interventions and Cross-Validated Targeted Learning. Biostatistics. doi
  • Kallenberg, Olav (2002). Foundations of Modern Probability. Springer. doi
  • Hoeffding, Wassily (1963). Probability Inequalities for Sums of Bounded Random Variables. Journal of the American Statistical Association. doi
  • Low, Mark G. (1997). On Nonparametric Confidence Intervals. The Annals of Statistics. doi
  • Robins, James M. and Hern{\'a}n, Miguel A. and Brumback, Babette (2000). Marginal Structural Models and Causal Inference in Epidemiology. Epidemiology. doi
  • Pearl, Judea (2009). Causality: Models, Reasoning, and Inference. Cambridge University Press.
  • van der Laan, Mark J. and Robins, James M. (2003). Unified Methods for Censored Longitudinal Data and Causality. Springer. doi
  • van der Laan, Mark J. and Rose, Sherri (2011). Targeted Learning. Springer. doi
  • Imbens, Guido W. (2000). The Role of the Propensity Score in Estimating Dose-Response Functions. Biometrika. doi
  • Hirano, Keisuke and Imbens, Guido W. (2005). The Propensity Score with Continuous Treatments. Applied Bayesian Modeling and Causal Inference from Incomplete-Data Perspectives. doi
  • Imai, Kosuke and van Dyk, David A (2004). Causal Inference With General Treatment Regimes. Journal of the American Statistical Association. doi
  • Muñoz, Iván Díaz and van der Laan, Mark (2011). Population Intervention Causal Effects Based on Stochastic Interventions. Biometrics. doi
  • Haneuse, Sebastian and Rotnitzky, Andrea (2013). Estimation of the Effect of Interventions that Modify the Received Treatment. Statistics in Medicine. doi
  • Young, Jessica G. and Hern{\'a}n, Miguel A. and Robins, James M. (2014). Identification, Estimation and Approximation of Risk Under Interventions that Depend on the Natural Value of Treatment Using Observational Data. Epidemiologic Methods. doi
  • Iván Díaz and Nicholas T. Williams and Paweł Morzywołek and Kara E. Rudolph (2026). Modified treatment policies that depend on the natural history of treatment. . arXiv
  • Koo, Taehyeon and Stuart, Elizabeth A. and Rudolph, Kara E. and Miles, Caleb H. (2026). Causal Effects of Modified Treatment Policies Under Positivity Violations: A Partial Identification Approach. . arXiv
  • Stone, Charles J. (1982). Optimal Global Rates of Convergence for Nonparametric Regression. The Annals of Statistics. doi
  • Fan, Jianqing and Gijbels, Irene (1996). Local Polynomial Modelling and Its Applications. Chapman and Hall.
  • Tsybakov, Alexandre B. (2009). Introduction to Nonparametric Estimation. Springer. doi
  • Donoho, David L. and Liu, Richard C. (1991). Geometrizing Rates of Convergence, II. The Annals of Statistics. doi
  • Cai, T. Tony and Low, Mark G. (2004). An Adaptation Theory for Nonparametric Confidence Intervals. The Annals of Statistics. doi
  • Donoho, David L. and Liu, Richard C. (1991). Geometrizing Rates of Convergence, II. The Annals of Statistics. doi
  • Cai, T. Tony and Low, Mark G. (2004). An adaptation theory for nonparametric confidence intervals. The Annals of Statistics. doi
  • Ga{\"i}ffas, St{\'e}phane (2007). On pointwise adaptive curve estimation based on inhomogeneous data. ESAIM: Probability and Statistics. doi