CausalSmith · seminar slides

Minimax Inference for Threshold Clamp Policies

For continuous treatments, we characterize the minimax risk and honest confidence-interval length for a lower-threshold clamp when overlap thins polynomially near the boundary.

Overview

  • The policy raises all treatments below the threshold δn\delta_n, the policy threshold, up to δn\delta_n.
  • This creates a moving atom at the boundary of the post-policy treatment law.
  • We estimate the clamp mean by combining a retained-course mean, a lower-tail atom mass, and a boundary regression value.
  • The frontier has two pieces: ordinary sampling noise and atom-weighted boundary learning.
  • The same frontier governs point estimation, honest interval length, and the causal lift.

informal · Theorem T-2 Under the finite-stratum Hölder model with the two-sided polynomial density envelope, the minimax absolute-error risk is at most and at least constant multiples of rnr_n.

Motivation

  • Modified treatment policies are a common way to define feasible effects for continuous treatments.
  • A threshold rule is natural when doses below a minimum practical level are reassigned to that minimum.
  • In applications, the threshold may move with sample size as investigators probe closer to the support boundary.
  • The inferential difficulty is concentrated at the boundary value created by the clamp.
  • The question is: how much precision is possible when the data become sparse near that boundary?
Natural A continuous exposure Threshold δ_n minimum level Split below δ_n above δ_n Lower tail below threshold Retained course observed dose kept Atom same dose δ_n Boundary response at δ_n Clamp mean averages both parts
illustrative Box-and-arrow schematic showing natural treatment A split by threshold δ_n into a retained course above the threshold and a lower-tail atom at δ_n, with the boundary response and retained course combining into the clamp mean.

Related Literature

  • Díaz et al. (2021), Williams and Díaz (2023), and Hoffman et al. (2024) make modified treatment policies operational for continuous treatments.
  • Kennedy et al. (2017) and Bonvini and Kennedy (2026) study continuous-treatment dose-response inference under regular support.
  • van der Laan et al. (2022) study stochastic threshold interventions with efficient inference and bands.
  • Gaïffas (2005) gives the closest degenerate-design pointwise regression benchmark.
  • Low (1997) and Armstrong and Kolesár (2018, 2016) supply the honest interval and modulus perspective.
  • Our contribution is the exact deterministic clamp frontier under polynomial thinning and moving thresholds.

Setup

  • Observed data are O=(X,A,Y)O=(X,A,Y): finite stratum, continuous treatment, bounded outcome.
  • The treatment density in stratum xx is πx(a)\pi_x(a).
  • The lower-threshold clamp is dδ(a)=max{a,δ}d_\delta(a)=\max\{a,\delta\}.
  • The observed clamp target is the retained mean above the threshold plus the lower-tail atom mass times the boundary regression.
  • The first-order object is the atom at δn\delta_n: its mass shrinks like δnκ+1\delta_n^{\kappa+1}, where κ\kappa is the thinning exponent.
Definition P-2a (Clamp functionals \(q_x(\delta)\), \(\nu_\delta(P)\), and \(\theta_\delta(P)\))

For PMP\in\mathcal M as in Definition P-1, xXx\in\mathcal X, and 0δδˉ0\leq\delta\leq\bar\delta, define qx(δ)=[0,δ]πx(a)da,νδ(P)=EP ⁣[Y1{A>δ}]. q_x(\delta)=\int_{[0,\delta]}\pi_x(a)\,\mathrm da, \qquad \nu_\delta(P)=\mathbb E_P\!\left[Y\mathbf 1\{A>\delta\}\right]. The unsmoothed observed-data clamp target is θδ(P)=νδ(P)+xXpxqx(δ)μxP(δ). \theta_\delta(P) = \nu_\delta(P) + \sum_{x\in\mathcal X}p_xq_x(\delta)\mu_x^P(\delta).

Assumptions

  • We work with fixed finite strata and i.i.d. sampling.
  • Every stratum has nonvanishing mass.
  • The treatment density obeys a two-sided envelope caκπx(a)c+aκc_-a^\kappa\le\pi_x(a)\le c_+a^\kappa on all of [0,1][0,1]; it is at the lower endpoint that this thins the design.
  • The outcome regression is Hölder smooth with exponent β\beta, the smoothness exponent, and radius LL.
  • These conditions describe the amount of information near the moving threshold.
Assumption A-6 (Polynomial overlap thinning)

For every xXx\in\mathcal X and for Lebesgue-almost every a[0,1]a\in[0,1], caκπx(a)c+aκ. c_- a^\kappa \leq \pi_x(a) \leq c_+ a^\kappa .

Assumption A-7 (Hölder regression regularity)

The regression function μxP\mu_x^P satisfies the following condition with parameters β\beta and LL. For every xXx\in\mathcal X:

  • (Continuity.) The map aμxP(a)a\mapsto \mu_x^P(a) is continuous on [0,1][0,1].
  • (Range.) For every a[0,1]a\in[0,1], μxP(a)[0,1]\mu_x^P(a)\in[0,1].
  • (Regression version.) The function μXP(A)\mu_X^P(A) realizes the conditional outcome regression given the complete design: EP[YX,A]=μXP(A)P-almost surely. \mathbb E_P[Y\mid X,A]=\mu_X^P(A) \quad P\text{-almost surely}.
  • (Hölder Taylor remainder.) Let =(β)\ell=\ell(\beta) be the local-polynomial order defined in Definition. For all s,t[0,1]s,t\in[0,1], with derivatives taken intrinsically on [0,1][0,1], μxP(t)j=0(μxP)(j)(s)j!(ts)jLtsβ. \left| \mu_x^P(t) - \sum_{j=0}^{\ell} \frac{(\mu_x^P)^{(j)}(s)}{j!}(t-s)^j \right| \le L|t-s|^\beta .

Information Balance

  • The threshold regression is learned locally around δn\delta_n.
  • The bandwidth hnh_n, the local window width, balances local-polynomial bias against local sampling information.
  • The resulting Hölder frontier rate combines sampling noise and atom-weighted boundary error.
Definition P-3 (Information-balance bandwidth \(h_n\))

The information-balance bandwidth hnh_n for the threshold sequence δn\delta_n is hn=inf{h(0,1δˉ]:nh2β+1(δn+h)κ1}, h_n = \inf\left\{ h\in(0,1-\bar\delta]: n h^{2\beta+1}(\delta_n+h)^\kappa\geq 1 \right\}, with hn=1δˉh_n=1-\bar\delta when the displayed set is empty.

Definition P-4 (Frontier rate \(r_n\))

The Hölder minimax rate rnr_n is rn=n1/2+δnκ+1hnβ. r_n=n^{-1/2}+\delta_n^{\kappa+1}h_n^\beta .

Estimator

  • Split the sample into three fixed blocks.
  • Use one block for the retained-course mean above δn\delta_n.
  • Use one block for the empirical lower-tail atom masses.
  • Use one block for the local-polynomial boundary regression.
  • Stabilize by checking the realized total Gram matrix and falling back to a bounded value on singular local designs.
Sample observed data Three blocks deterministic split Retained mean Block 1 above δ_n Atom masses Block 2 lower threshold Gram check realized local Gram bounded fallback Boundary regression Block 3 [δ_n, δ_n+h_n] Total-Gram estimator components combine Interval Hoeffding radii bias-aware radii
illustrative Box-and-arrow schematic showing an observed sample split into retained-mean, atom-mass, Gram-check, and boundary-regression components, then combined into the Total-Gram estimator and bias-aware interval.

Main Result

informal · Theorem T-2 The total-Gram estimator attains worst-case absolute error at most CrnC r_n, and every estimator has worst-case absolute error at least crnc r_n, under the stated Hölder clamp model.

Theorem T-2 (Minimax clamp frontier)

Fix J,β,κ,L,c,c+,pmin,δˉ,αJ,\beta,\kappa,L,c_-,c_+,p_{\min},\bar\delta,\alpha satisfying the regime conditions:

  • (Regime.) J1J\geq 1, β>0\beta>0, κ0\kappa\geq0, L>0L>0, c>0c_->0, cκ+1c+c_-\leq \kappa+1\leq c_+, pmin>0p_{\min}>0, pmin1/Jp_{\min}\leq 1/J, δˉ(0,1)\bar\delta\in(0,1), and α(0,1/2)\alpha\in(0,1/2).

Then there exist constants 0<c<C0<c<C and an amplitude a(0,1/4]a\in(0,1/4], depending only on the displayed regime constants, such that, for every deterministic threshold sequence (δn)n1(\delta_n)_{n\geq1} with δn[0,δˉ]\delta_n\in[0,\bar\delta] for every nn and every deterministic sequence Bn=(I0,n,I1,n,I2,n)B_n=(I_{0,n},I_{1,n},I_{2,n}) of three pairwise-disjoint sample blocks with Ij,nn/4|I_{j,n}|\geq \lfloor n/4\rfloor for j=0,1,2j=0,1,2, for all sufficiently large nn, with hnh_n the information-balance bandwidth from Definition P-3 and rn=n1/2+δnκ+1hnβ r_n=n^{-1/2}+\delta_n^{\kappa+1}h_n^\beta as in Definition P-4, the observed minimax absolute-error risk RnR_n^\star over the clamp model M\mathcal M in Definition P-1 and the worst-case i.i.d. risk over M\mathcal M of the stabilized estimator using BnB_n satisfy crnRnsupPMEP ⁣[θ^n,BnTGθδn(P)]Crn. c r_n \leq R_n^\star \leq \sup_{P\in\mathcal M} \mathbb E_P\!\left[ \left|\widehat\theta_{n,B_n}^{\mathrm{TG}}-\theta_{\delta_n}(P)\right| \right] \leq C r_n . Moreover, for all sufficiently large nn, the lower bound is witnessed inside the same clamp model by:

  • (Global shift.) Laws P0,P1MP_0,P_1\in\mathcal M satisfying Assumption A-1, with the same (X,A)(X,A)-design distribution and the same pxp_x and πx\pi_x functions, having Bernoulli outcomes, and obeying μxP1(a)=μxP0(a)+n1/2for every x and a[0,1], \mu_x^{P_1}(a)=\mu_x^{P_0}(a)+n^{-1/2} \quad\text{for every }x\text{ and }a\in[0,1], with well-posed product chi-squared divergence, P1nP0n,(dP1ndP0n1)2dP0n<, P_1^{\otimes n}\ll P_0^{\otimes n}, \qquad \int \left(\frac{dP_1^{\otimes n}}{dP_0^{\otimes n}}-1\right)^2\,dP_0^{\otimes n}<\infty, and cn1/2θδn(P1)θδn(P0),χ2 ⁣(P1n,P0n)C. c n^{-1/2} \leq \left|\theta_{\delta_n}(P_1)-\theta_{\delta_n}(P_0)\right|, \qquad \chi^2\!\left(P_1^{\otimes n},P_0^{\otimes n}\right)\leq C .
  • (Localized perturbation.) Laws Q0,Q1MQ_0,Q_1\in\mathcal M satisfying Assumption A-1, with the same (X,A)(X,A)-design distribution and the same pxp_x and πx\pi_x functions, having Bernoulli outcomes, and admitting a continuous bump b:RRb:\mathbb R\to\mathbb R satisfying b(0)=1b(0)=1, b(u)0b(u)\geq0 for every uu, b(u)=0b(u)=0 for u[1,1]u\notin[-1,1], and b(u)1|b(u)|\leq1 for u[1,1]u\in[-1,1], with regression shift μxQ1(t)μxQ0(t)=ahnβb ⁣(tδnhn) \mu_x^{Q_1}(t)-\mu_x^{Q_0}(t) = a\,h_n^\beta b\!\left(\frac{t-\delta_n}{h_n}\right) for every xx and t[0,1]t\in[0,1]. The product chi-squared divergence is well posed, Q1nQ0n,(dQ1ndQ0n1)2dQ0n<, Q_1^{\otimes n}\ll Q_0^{\otimes n}, \qquad \int \left(\frac{dQ_1^{\otimes n}}{dQ_0^{\otimes n}}-1\right)^2\,dQ_0^{\otimes n}<\infty, and cδnκ+1hnβθδn(Q1)θδn(Q0),χ2 ⁣(Q1n,Q0n)C. c\,\delta_n^{\kappa+1}h_n^\beta \leq \left|\theta_{\delta_n}(Q_1)-\theta_{\delta_n}(Q_0)\right|, \qquad \chi^2\!\left(Q_1^{\otimes n},Q_0^{\otimes n}\right)\leq C .

Honest Inference

informal · Theorem T-3 The bias-aware interval has uniform coverage at least 1α1-\alpha, expected length at most CrnC r_n, and every uniformly honest interval has expected length at least crnc r_n.

Theorem T-3 (Honest length frontier)

There are constants c,C(0,)c,C\in(0,\infty), with c<Cc<C, for which the following holds.

  • (Regime constants.) The number of strata JJ is a positive integer, β>0\beta>0, κ0\kappa\geq0, L>0L>0, c>0c_->0, cκ+1c+c_-\leq \kappa+1\leq c_+, pmin>0p_{\min}>0, pmin1/Jp_{\min}\leq 1/J, δˉ(0,1)\bar\delta\in(0,1), and α(0,1/2)\alpha\in(0,1/2).
  • (Threshold path.) The deterministic thresholds satisfy 0δnδˉ0\leq \delta_n\leq \bar\delta for every nn.
  • (Sample splits.) For every nn, Bn=(I0,n,I1,n,I2,n)B_n=(I_{0,n},I_{1,n},I_{2,n}) is a deterministic three-way split of the sample indices with pairwise disjoint blocks and Ij,nn/4|I_{j,n}|\geq \lfloor n/4\rfloor for j=0,1,2j=0,1,2.

For all sufficiently large nn, set hnh_n to be the information-balance bandwidth for δn\delta_n from Definition P-3, and set rn=n1/2+δnκ+1hnβ. r_n=n^{-1/2}+\delta_n^{\kappa+1}h_n^\beta . Then the stabilized interval CIn\mathrm{CI}_n has worst-case coverage inf{Pn{θδn(P)CIn(O1,,On)}:PM}1α, \inf\Bigl\{ P^{\otimes n}\{\theta_{\delta_n}(P)\in \mathrm{CI}_n(O_1,\ldots,O_n)\}: P\in\mathcal M \Bigr\}\geq 1-\alpha, and its worst-case expected length is comparable to rnr_n: crnsup{EPn[len(CIn)]:PM}Crn. c\,r_n \leq \sup\Bigl\{ \mathbb E_{P^{\otimes n}}\bigl[\operatorname{len}(\mathrm{CI}_n)\bigr]: P\in\mathcal M \Bigr\} \leq C\,r_n . Moreover, the minimax honest expected length satisfies crnLn, c\,r_n \leq L_n^\star, where LnL_n^\star is the infimum, over observed-sample confidence procedures with uniform coverage at least 1α1-\alpha over M\mathcal M, of their worst-case expected length. Equivalently, every observed-sample confidence procedure CnC_n with inf{Pn{θδn(P)Cn(O1,,On)}:PM}1α \inf\Bigl\{ P^{\otimes n}\{\theta_{\delta_n}(P)\in C_n(O_1,\ldots,O_n)\}: P\in\mathcal M \Bigr\}\geq 1-\alpha has worst-case expected length at least crnc\,r_n: crnsup{EPn[len(Cn)]:PM}. c\,r_n \leq \sup\Bigl\{ \mathbb E_{P^{\otimes n}}\bigl[\operatorname{len}(C_n)\bigr]: P\in\mathcal M \Bigr\}.

Phase Diagram

  • The threshold path determines which term in rnr_n is visible.
  • Near zero, the atom is small enough for root-nn behavior.
  • Past the critical scale, the atom-weighted boundary regression term governs the rate.
  • At a fixed positive threshold, the rate becomes the usual boundary nonparametric rate.
  • At threshold zero, the target is the ordinary observed mean.

informal · Theorem T-4 The phase diagram separates regular, critical, vanishing atom-dominated, fixed-threshold, and zero-threshold regimes for rnr_n.

Theorem T-4 (Phase boundary regimes)

Let JJ be a positive integer and let β>0,κ0,L>0,0<cκ+1c+,0<pmin1/J,0<δˉ<1,0<α<1/2. \beta>0,\qquad \kappa\geq 0,\qquad L>0,\qquad 0<c_-\leq \kappa+1\leq c_+,\qquad 0<p_{\min}\leq 1/J,\qquad 0<\bar\delta<1,\qquad 0<\alpha<1/2 . Define the phase and edge scales by δcrit,n=n1/{2(βκ+2β+κ+1)},δedge,n=n1/(2β+κ+1). \delta_{\mathrm{crit},n} = n^{-1/\{2(\beta\kappa+2\beta+\kappa+1)\}}, \qquad \delta_{\mathrm{edge},n} = n^{-1/(2\beta+\kappa+1)} . Then δcrit,nδedge,n. \frac{\delta_{\mathrm{crit},n}}{\delta_{\mathrm{edge},n}}\to\infty . Moreover, for every deterministic threshold sequence (δn)n1(\delta_n)_{n\geq 1} with δn[0,δˉ]\delta_n\in[0,\bar\delta] for every nn, let hnh_n be the information-balance bandwidth in Definition P-3, and set an=δnκ+1hnβ,rn=n1/2+δnκ+1hnβ a_n=\delta_n^{\kappa+1}h_n^\beta, \qquad r_n=n^{-1/2}+\delta_n^{\kappa+1}h_n^\beta as in Definition P-4. The following conclusions hold:

  • Regular thresholds. If δn/δcrit,n0\delta_n/\delta_{\mathrm{crit},n}\to 0, then rnn1/2r_n\asymp n^{-1/2}.
  • Critical thresholds. For every c0>0c_0>0, if δn/δcrit,nc0\delta_n/\delta_{\mathrm{crit},n}\to c_0, then ann1/2a_n\asymp n^{-1/2} and rnn1/2r_n\asymp n^{-1/2}.
  • Vanishing atom-dominated thresholds. If δn/δcrit,n\delta_n/\delta_{\mathrm{crit},n}\to\infty and δn0\delta_n\to0, then rnδnκ+1(nδnκ)β/(2β+1). r_n\asymp \delta_n^{\kappa+1}\bigl(n\delta_n^\kappa\bigr)^{-\beta/(2\beta+1)} .
  • Fixed thresholds. For every δ0(0,δˉ]\delta_0\in(0,\bar\delta], if δnδ0\delta_n\to\delta_0, then rnnβ/(2β+1). r_n\asymp n^{-\beta/(2\beta+1)} .
  • Zero threshold identities. For every nn and every PMP\in\mathcal M satisfying Definition P-1, qx(0)=0for every x,θ0(P)=YdP,rnδn=0=n1/2, q_x(0)=0\quad\text{for every }x,\qquad \theta_0(P)=\int Y\,dP,\qquad r_n\big|_{\delta_n=0}=n^{-1/2}, where the last identity uses the bandwidth hnh_n at threshold 00.
  • Zero threshold estimator. For every nn, every PMP\in\mathcal M satisfying Definition P-1, and every admissible three-way split B=(I0,I1,I2)B=(I_0,I_1,I_2) as in Definition P-5, the total-Gram estimator in Definition P-5, formed at threshold 00 with bandwidth hnh_n, agrees under the product sampling law generated by PP with zclamp[0,1] ⁣(1I0iI0Yi). z\mapsto \operatorname{clamp}_{[0,1]}\!\left(\frac{1}{|I_0|}\sum_{i\in I_0}Y_i\right).
  • Zero threshold stabilized guarantees. For every deterministic sequence B=(Bn)n1B_\bullet=(B_n)_{n\geq1} of admissible three-way split blocks, the stabilized zero-threshold procedure has coverage at least 1α1-\alpha for all sufficiently large nn. Its stabilized worst-case risk at threshold 00 and its stabilized worst-case expected length at threshold 00 with noncoverage level α\alpha are both asymptotic to n1/2n^{-1/2}.

Calibration

  • The one-stratum example makes the phase boundary concrete.
  • Take β=κ=1\beta=\kappa=1 and π(a)=2a\pi(a)=2a.
  • The localized Bernoulli alternatives keep likelihood distance bounded while moving the target by the atom-weighted boundary amount.
  • The critical threshold scale is where this movement matches n1/2n^{-1/2}.

informal · Theorem T-5 In the one-stratum linear-thinning case, hn(nδn)1/3h_n\asymp (n\delta_n)^{-1/3}, Δnδn2hn\Delta_n\asymp \delta_n^2h_n, and δnn1/10\delta_n\asymp n^{-1/10} is equivalent to Δnn1/2\Delta_n\asymp n^{-1/2}.

Key Idea

  • The clamp mean has two statistically different pieces.
  • The retained mean behaves like a bounded sample average.
  • The lower-tail atom multiplies the regression value at the threshold.
  • A naive boundary plug-in inherits degenerate-design instability near sparse support.
  • Total-Gram stabilization uses the realized local design only when it has enough curvature.
  • The atom weight shrinks the boundary-regression error from hnβh_n^\beta to δnκ+1hnβ\delta_n^{\kappa+1}h_n^\beta.

Proof Sketch

  • Upper bounds decompose the estimator into retained-course error, atom-mass error, and boundary-regression error.
  • Polynomial thinning gives the local sample size and Gram curvature scale in the threshold window.
  • Hölder smoothness gives a deterministic local-polynomial bias bound.
  • The lower bound uses two experiments: a global Bernoulli shift for n1/2n^{-1/2}, and a localized threshold bump for δnκ+1hnβ\delta_n^{\kappa+1}h_n^\beta.
  • Low-style honest-length lower bounds transfer these testing separations to interval length.

Causal Interpretation

  • The full-data class adds potential outcomes through a latent-response representation.
  • Consistency links observed outcomes to structural responses at the realized treatment.
  • Conditional exchangeability identifies the structural response mean within strata.
  • Response continuity aligns the structural mean with the observed regression value on the threshold range.

informal · Theorem T-1 Under the declared full-data causal conditions, the causal clamp mean equals the observed clamp target.

informal · Theorem T-7 Every observed law in the Hölder clamp model has a full-data lift, so the observed and causal minimax criteria coincide.

Causal Frontier

informal · Theorem T-6 The same total-Gram estimator, bias-aware interval, minimax risk rate, and honest-length rate rnr_n hold for the causal clamp mean over the full-data class.

Theorem T-6 (Causal frontier lift)

Let JNJ\in\mathbb N, and let β,κ,L,c,c+,pmin,δˉ,αR\beta,\kappa,L,c_-,c_+,p_{\min},\bar\delta,\alpha\in\mathbb R. Suppose that

  • (Regime constants.) The constants satisfy 0<J,β>0,κ0,L>0,c>0,cκ+1c+, 0<J,\qquad \beta>0,\qquad \kappa\geq0,\qquad L>0,\qquad c_->0,\qquad c_-\leq \kappa+1\leq c_+, pmin>0,pmin1J,0<δˉ<1,0<α<12. p_{\min}>0,\qquad p_{\min}\leq \frac1J,\qquad 0<\bar\delta<1,\qquad 0<\alpha<\frac12 .
  • (Threshold sequence.) The deterministic thresholds (δn)nN(\delta_n)_{n\in\mathbb N} satisfy 0δnδˉ 0\leq \delta_n\leq \bar\delta for every nNn\in\mathbb N.
  • (Split blocks.) For each nNn\in\mathbb N, Bn=(I0,n,I1,n,I2,n)B_n=(I_{0,n},I_{1,n},I_{2,n}) is a deterministic three-way split with pairwise disjoint blocks and n4I0,n,n4I1,n,n4I2,n. \left\lfloor \frac n4\right\rfloor\leq |I_{0,n}|,\qquad \left\lfloor \frac n4\right\rfloor\leq |I_{1,n}|,\qquad \left\lfloor \frac n4\right\rfloor\leq |I_{2,n}|.

Let hnh_n, the information-balance bandwidth for δn\delta_n, be as in Definition P-3, and let rn=n1/2+δnκ+1hnβ r_n=n^{-1/2}+\delta_n^{\kappa+1}h_n^\beta be the frontier rate from Definition P-4. Then there exist constants c,CRc,C\in\mathbb R, with 0<c<C0<c<C, depending only on these regime constants, such that, for every threshold sequence and every deterministic split-block sequence satisfying the preceding conditions, all sufficiently large nn satisfy the following conclusions. The total-Gram estimator θ^n\widehat\theta_n in Definition P-5, formed with split BnB_n, order =β1\ell=\lceil\beta\rceil-1, bandwidth hnh_n, and threshold δn\delta_n, is observed-sample measurable. The corresponding bias-aware interval CIn\mathrm{CI}_n in Definition P-6, formed with the same split, order, bandwidth, and threshold, is an observed-sample measurable interval. For the causal minimax criteria Rn,FR_{n,\mathrm F}^\star and Ln,FL_{n,\mathrm F}^\star in Definition P-9, crnRn,FsupPFMFE(PF)nθ^nψδn(PF)Crn, c r_n\leq R_{n,\mathrm F}^\star \leq \sup_{P^{\mathrm F}\in\mathcal M^{\mathrm F}} \mathbb E_{(P^{\mathrm F})^{\otimes n}} \left|\widehat\theta_n-\psi_{\delta_n}(P^{\mathrm F})\right| \leq C r_n, infPFMF(PF)n{ψδn(PF)CIn}1α, \inf_{P^{\mathrm F}\in\mathcal M^{\mathrm F}} (P^{\mathrm F})^{\otimes n}\{\psi_{\delta_n}(P^{\mathrm F})\in \mathrm{CI}_n\} \geq 1-\alpha, supPFMFE(PF)nlen(CIn)Crn. \sup_{P^{\mathrm F}\in\mathcal M^{\mathrm F}} \mathbb E_{(P^{\mathrm F})^{\otimes n}}\operatorname{len}(\mathrm{CI}_n) \leq C r_n. Moreover, in the extended nonnegative-real order, after embedding the nonnegative real bounds into R0{}\mathbb R_{\geq0}\cup\{\infty\}, crnLn,FCrn. c r_n\leq L_{n,\mathrm F}^\star\leq C r_n .

Continuity-only Comparison

  • The continuity-only model keeps the same design and thinning conditions.
  • It replaces quantitative Hölder smoothness with qualitative continuity of the regression extension.
  • The estimator uses the retained mean, atom masses, and a fixed bounded fallback for the atom regression.
  • The resulting frontier is driven by sampling noise plus the total lower-tail mass.

informal · Theorem T-8 In the continuity-only full-data class, the causal clamp mean equals the continuity-only observed clamp target.

informal · Theorem T-9 The continuity-only observed and causal minimax criteria coincide through observed-margin surjectivity.

informal · Theorem T-10 Over the continuity-only class, observed and causal minimax risk and honest-length rates are characterized by sn=n1/2+δnκ+1s_n=n^{-1/2}+\delta_n^{\kappa+1}, with the stated elbow and fixed-positive-threshold behavior.

Theorem T-10 (Continuity frontier characterization)

Fix JNJ\in\mathbb N and constants κ,c,c+,pmin,δˉ,αR\kappa,c_-,c_+,p_{\min},\bar\delta,\alpha\in\mathbb R. Suppose 0<J,0κ,0<c,cκ+1c+,0<pmin1J,0<δˉ<1,0<α<12. 0<J, \qquad 0\leq \kappa, \qquad 0<c_-, \qquad c_-\leq \kappa+1\leq c_+, \qquad 0<p_{\min}\leq \frac{1}{J}, \qquad 0<\bar\delta<1, \qquad 0<\alpha<\frac12. Then there exist constants c,CRc,C\in\mathbb R with 0<c<C0<c<C such that the following statements hold.

  • (Uniform frontier bounds.) For every deterministic threshold sequence (δn)n1(\delta_n)_{n\geq1} with δn[0,δˉ]\delta_n\in[0,\bar\delta] for all nn, and for every deterministic sequence B=(Bn)n1B_{\bullet}=(B_n)_{n\geq1} of admissible three-way split blocks, all sufficiently large nn satisfy the following bounds. With sn=n1/2+δnκ+1, s_n=n^{-1/2}+\delta_n^{\kappa+1}, and with Rn,contR_{n,\mathrm{cont}}^\star, Ln,contL_{n,\mathrm{cont}}^\star, Rn,cont,FR_{n,\mathrm{cont},\mathrm F}^\star, and Ln,cont,FL_{n,\mathrm{cont},\mathrm F}^\star as in Definition P-16, csnRn,contsupPMcontEPnθ^n,contθδn,cont(P)Csn, c s_n\leq R_{n,\mathrm{cont}}^\star \leq \sup_{P\in\mathcal M_{\mathrm{cont}}} \mathbb E_{P^{\otimes n}} \left|\widehat\theta_{n,\mathrm{cont}}-\theta_{\delta_n,\mathrm{cont}}(P)\right| \leq C s_n, the interval CIn,cont\mathrm{CI}_{n,\mathrm{cont}} has uniform coverage at least 1α1-\alpha over Mcont\mathcal M_{\mathrm{cont}}, its worst-case expected length is at most CsnC s_n, and csnLn,cont. c s_n\leq L_{n,\mathrm{cont}}^\star. Moreover, Rn,cont,F=Rn,cont,Ln,cont,F=Ln,cont, R_{n,\mathrm{cont},\mathrm F}^\star=R_{n,\mathrm{cont}}^\star, \qquad L_{n,\mathrm{cont},\mathrm F}^\star=L_{n,\mathrm{cont}}^\star, and the same estimator and interval satisfy Rn,cont,FsupPFMcontFE(PobsF)nθ^n,contψδn(PF)Csn, R_{n,\mathrm{cont},\mathrm F}^\star \leq \sup_{P^{\mathrm F}\in\mathcal M_{\mathrm{cont}}^{\mathrm F}} \mathbb E_{(P^{\mathrm F}_{\mathrm{obs}})^{\otimes n}} \left|\widehat\theta_{n,\mathrm{cont}}-\psi_{\delta_n}(P^{\mathrm F})\right| \leq C s_n, with causal coverage at least 1α1-\alpha uniformly over McontF\mathcal M_{\mathrm{cont}}^{\mathrm F} and causal worst-case expected length at most CsnC s_n.
  • (Elbow and null-threshold rates.) For the continuity-only rate in Definition P-12, sn(δn=n1/(2(κ+1)))n1/2,sn(0)n1/2. s_n\bigl(\delta_n=n^{-1/(2(\kappa+1))}\bigr) \asymp n^{-1/2}, \qquad s_n(0)\asymp n^{-1/2}.
  • (Fixed positive thresholds.) For every deterministic threshold sequence (δn)n1(\delta_n)_{n\geq1} with δn[0,δˉ]\delta_n\in[0,\bar\delta] for all nn, every δ0(0,δˉ]\delta_0\in(0,\bar\delta], and δnδ0\delta_n\to\delta_0, Rn,cont↛0,Ln,cont↛0. R_{n,\mathrm{cont}}^\star \not\to 0, \qquad L_{n,\mathrm{cont}}^\star \not\to 0.

Conclusion

  • We characterize estimation and honest inference for the exact lower-threshold clamp under a two-sided polynomial density envelope on [0,1][0,1].
  • The frontier is rn=n1/2+δnκ+1hnβr_n=n^{-1/2}+\delta_n^{\kappa+1}h_n^\beta.
  • The phase diagram explains how moving thresholds shift the problem between root-nn, atom-dominated, and fixed-threshold regimes.
  • The causal bridge gives the same rates a full-data interpretation under the stated consistency, exchangeability, and continuity conditions.
  • The continuity-only analysis separates the role of qualitative continuity from the quantitative Hölder modulus.