CausalSmith · seminar slides

Quotient-Law Inference for Latent-Class Mean Effects at Collisions

We estimate the population-weighted distribution of latent-class average treatment effects uniformly at the root-nn rate, including configurations where several classes share the same mean effect. This estimand is distinct from the distribution of individual contrasts Y(1)Y(0)Y(1)-Y(0).

Overview

  • Our target is the quotient law νP\nu_P, the population-weighted distribution of latent-class average treatment effects after aggregating classes with the same mean effect.
  • In the one-Wasserstein transport distance W1W_1, observable proxy moments determine this law through a gap-free Lipschitz map.
  • Law estimation and honest confidence reporting attain uniform root-nn accuracy.
  • Effect-ordered class masses incur the sharp inverse-gap cost on separated two-class configurations.

Motivation

  • Think of an observational treatment study where an unobserved response type affects treatment and outcomes.
  • Two proxy measurements carry complementary information about that latent type.
  • In our two-class benchmark, the class-average effects are 0.25ε0.25-\varepsilon and 0.25+ε0.25+\varepsilon, where ε\varepsilon is their displacement from a collision.
  • The population masses are 0.40.4 and 0.60.6.
  • At ε=0\varepsilon=0, both types contribute to one treatment-effect value; for ε>0\varepsilon>0, the law has two nearby atoms.
  • Can one infer the law of class-average effects continuously across this transition?

Target

  • The quotient law places each latent-class mass at its class-average treatment effect and adds the masses of coincident mean effects.
  • W1W_1 is the minimum distance-weighted mass required to transport one effect law into another.
  • In the benchmark, the quotient law moves continuously from two atoms to one aggregate atom as ε\varepsilon approaches zero.
  • It does not identify the distribution of individual contrasts Y(1)Y(0)Y(1)-Y(0); its atoms are conditional class means.
Latent class class mass Latent class class mass Class-average effect coincident mean Class-average effect coincident mean Quotient-law atom aggregate mass
illustrative Two latent-class boxes labeled by their masses point to class-average-effect boxes, and coincident mean-effect values merge into one quotient-law atom carrying the aggregate mass.

Model

  • We observe treatment, two proxy vectors, and the realized outcome for independent units.
  • The analyst supplies the fixed class cardinality kk and uniform bounds L,π0,σ0L,\pi_0,\sigma_0; consistency and latent ignorability give each class a causal mean treatment contrast.
  • Conditional separation makes one proxy a class measurement and the other an arm-specific class measurement.
  • Observable contributions are bounded.
  • The latent-arm positivity margin π0\pi_0 and proxy-rank margin σ0\sigma_0 remain fixed across the model class.
  • In our benchmark, the proxy matrices remain nonsingular at ε=0\varepsilon=0, and every latent class–treatment arm has positive mass.
Assumption A-9 (Latent-arm positivity)

The latent class and treatment arms satisfy min1uk, t{0,1}P(U=u,T=t)π0. \min_{1\le u\le k,\ t\in\{0,1\}} P(U=u,T=t)\ge \pi_0 .

Assumption A-10 (Proxy rank margin)

The reference- and target-proxy feature matrices satisfy min{σk(A0),σk(A1),σk(B)}σ0. \min\{\sigma_k(A_0),\sigma_k(A_1),\sigma_k(B)\}\ge \sigma_0 .

Observable Moments

  • The five-block summary S(P)S(P) consists of arm-specific proxy moments, outcome-weighted proxy moments, and the target-proxy mean.
  • Its empirical version S^n\widehat S_n is formed by sample averaging within treatment arms.
  • A compressed operator built from these moments has the latent treatment effects as eigenvalues.
  • Spectral anchors recover the masses attached to the corresponding eigenspaces.
  • The summary distance dSd_S adds the operator-norm errors of the four matrix blocks and the Euclidean error of the proxy mean.

informal · Theorem T-1 Under the fixed model restrictions, treatment arms have positive probability, the observable proxy moments have stable factorizations and singular-value margins, and latent effects remain within a fixed bounded interval.

informal · Theorem T-2 Under fixed dimensions and margins, the closure of feasible five-block summaries is compact, with nonemptiness equivalent to nonemptiness of the model class.

Related Literature

  • Miao et al. (2018) identify causal effects using conditionally independent proxies and rank conditions.
  • Mazaheri et al. (2025) and Virk et al. (2026) develop finite latent-effect identification and separated spectral recovery.
  • Heinrich and Kahn (2018) and Wu and Yang (2020) show how collision geometry shapes rates in ordinary finite mixtures.
  • Deo and Randrianarisoa (2023) develop honest Wasserstein confidence sets for mixture uncertainty.
  • Our contribution uses the proxy operator to obtain regular inference for the aggregated effect law across homogeneous, partially colliding, and separated configurations.

Key Idea

  • A direct plug-in eigendecomposition tracks individual eigenvectors and amplifies perturbations by the inverse effect gap.
  • Near a collision, individual eigendirections can rotate sharply while their combined spectral subspace remains stable.
  • We aggregate the anchor mass over each colliding eigenspace and transport that mass across the cluster diameter.
  • The aggregation absorbs simultaneous merges and splits into the W1W_1 geometry.
  • Uniform proxy conditioning controls the compressed operator; positivity controls the mass carried by each distinct atom.
  • Together these facts convert observable-moment error into law-level transport error without an effect-gap factor.

Main Result

informal · Theorem T-3 Under fixed dimensions, bounded contributions, latent-arm positivity, and proxy-rank margins, the W1W_1 distance between two quotient laws is at most a constant times the distance between their observable summaries, uniformly across collisions.

Theorem T-3 (Gap-free modulus extension)

Fix natural numbers k,dx,dzk,d_x,d_z and real constants L,π0,σ0L,\pi_0,\sigma_0. Assume:

  • (Latent dimension.) 2k2\le k, kdxk\le d_x, and kdzk\le d_z.
  • (Bound and margins.) 1L1\le L, 0<π0(2k)10<\pi_0\le (2k)^{-1}, and 0<σ010<\sigma_0\le 1.

Let Lτ=4Ldz/σ0L_\tau=4L\sqrt{d_z}/\sigma_0, and let K=S\mathcal K=\overline{\mathscr S} be the summary closure from Definition P-8. For two summaries s=(M0,M1,N0,N1,mX),q=(M0,M1,N0,N1,mX), s=(M_0,M_1,N_0,N_1,m_X), \qquad q=(M'_0,M'_1,N'_0,N'_1,m'_X), write dS(s,q)=M0M0+M1M1+N0N0+N1N1+(i(mX(i)mX(i))2)1/2. d_S(s,q) = \|M_0-M'_0\|+\|M_1-M'_1\| +\|N_0-N'_0\|+\|N_1-N'_1\| +\Bigl(\sum_i (m_X(i)-m'_X(i))^2\Bigr)^{1/2}. There exists a constant Cmod>0C_{\mathrm{mod}}>0 such that, for every pair of full-data laws P,QMP,Q\in\mathcal M, W1(νP,νQ)CmoddS(S(P),S(Q)). W_1(\nu_P,\nu_Q)\le C_{\mathrm{mod}}\,d_S(S(P),S(Q)). Moreover, there is a map F:KPk([Lτ,Lτ]) \overline F:\mathcal K\to \mathcal P_k([-L_\tau,L_\tau]) such that F\overline F is continuous and Borel measurable, W1(F(q),F(q))CmoddS(q,q)for all q,qK, W_1(\overline F(q),\overline F(q')) \le C_{\mathrm{mod}}\,d_S(q,q') \qquad\text{for all }q,q'\in\mathcal K, and F(S(Q))=νQfor every QM. \overline F(S(Q))=\nu_Q \qquad\text{for every }Q\in\mathcal M. This map is unique among continuous maps G:KPk([Lτ,Lτ])G:\mathcal K\to\mathcal P_k([-L_\tau,L_\tau]) that obey W1(G(q),G(q))CmoddS(q,q)for all q,qK W_1(G(q),G(q')) \le C_{\mathrm{mod}}\,d_S(q,q') \qquad\text{for all }q,q'\in\mathcal K and satisfy G(S(Q))=νQG(S(Q))=\nu_Q for every QMQ\in\mathcal M: every such GG satisfies G(q)=F(q)G(q)=\overline F(q) for all qKq\in\mathcal K.

The same Lipschitz law map covers the one-atom benchmark at ε=0\varepsilon=0 and its two-atom neighbors.

Estimation

  • Uniform concentration controls all five empirical-summary blocks at the usual sampling scale.
  • Our repair estimator selects the nearest feasible summary and applies the Lipschitz law map.
  • Our lattice estimator searches a finite constrained representation of discretized signal bases, conditioning matrices, masses, and effect locations for the best exact-real match to the empirical moments.
  • Both mechanisms return positive atomic effect laws.

informal · Lemma L-2 Uniformly over the model class, the empirical summary exceeds C0Llog(C0/η)/nC_0L\sqrt{\log(C_0/\eta)/n} error with probability at most the tail probability η\eta.

informal · Theorem T-5 With fixed dimensions and margins, both law estimators have W1W_1 error at most Clog(C/η)/nC\sqrt{\log(C/\eta)/n} simultaneously with probability at least 1η1-\eta, uniformly over the model class.

Theorem T-5 (Collision-uniform root-n rate)

Fix integers k,dx,dzk,d_x,d_z and constants L,π0,σ0L,\pi_0,\sigma_0. Suppose that

  • (Dimensions.) 2k2\le k, kdxk\le d_x, and kdzk\le d_z.
  • (Uniform bounds.) 1L1\le L, 0<π01/(2k)0<\pi_0\le 1/(2k), and 0<σ010<\sigma_0\le 1.

Then there is a constant C=C(k,dx,dz,L,π0,σ0)>0C=C(k,d_x,d_z,L,\pi_0,\sigma_0)>0 such that, for every n1n\ge 1, there exist a structured-lattice estimator λ^n\widehat\lambda_n taking values in Pk([Lτ,Lτ])\mathcal P_k([-L_\tau,L_\tau]), with Lτ=4Ldz/σ0, L_\tau=4L\sqrt{d_z}/\sigma_0, and a nearest-summary repair estimator ν^n=F(Π(S^n))\widehat\nu_n=\overline F(\Pi(\widehat S_n)) from Definition P-9, such that λ^n\widehat\lambda_n is the first lexicographic minimizer over a finite exhaustive structured lattice of well-formed candidates ϑ\vartheta for the criterion Jn(ϑ;S^n), J_n(\vartheta;\widehat S_n), with tie-breaking by the lattice order and output λϑ\lambda_\vartheta. Both estimators are Borel sample maps. For every η\eta satisfying 0<η<1/20<\eta<1/2, every PMP\in\mathcal M, and QP(n)Q_P^{(n)}-sample, QP(n) ⁣{max ⁣{W1(λ^n,νP),W1(ν^n,νP)}>Clog(C/η)n}η. Q_P^{(n)}\!\left\{ \max\!\left\{ W_1(\widehat\lambda_n,\nu_P), W_1(\widehat\nu_n,\nu_P) \right\} > C\sqrt{\frac{\log(C/\eta)}{n}} \right\}\le \eta . Moreover, for each treatment arm t{0,1}t\in\{0,1\}, QP(n){Nt,n=0}(1kπ0)n. Q_P^{(n)}\{N_{t,n}=0\}\le (1-k\pi_0)^n .

Estimator Pipeline

  • Thresholding extracts the kk-dimensional proxy signal space at the fixed scale π0σ02/2\pi_0\sigma_0^2/2.
  • The lattice compares candidate effect operators, proxy means, and anchor equations.
  • The first minimizing candidate supplies the estimated atomic law λ^n\widehat\lambda_n.
  • The mesh contributes 1/n1/\sqrt n, matching the sampling scale.

informal · Theorem T-15 For fixed dimensions and margins, the structured-lattice estimator has W1W_1 error at most Clat{dS(S^n,S(P))+1/n}C_{\mathrm{lat}}\{d_S(\widehat S_n,S(P))+1/\sqrt n\}, a polynomial-size candidate list, and the stated uniform root-nn tail bound.

informal · Theorem T-13 For fixed dimensions and supplied exact-real primitives, a finite-library estimator is Borel, uses at most C{n+n(4dzdx+dx)/2}C\{n+n^{(4d_zd_x+d_x)/2}\} operations, and attains the stated uniform root-nn tail bound.

Observed sample treatment, proxies, outcome Five-block summary empirical Ŝₙ Proxy signal threshold π₀σ₀²/2 k dimensions Lattice mesh 1/√n sampling scale Structured lattice operators, means, anchors polynomial candidate list Atomic quotient law first minimizing λ̂ₙ Wasserstein report C_lat{d_S+1/√n} uniform root-n tail
illustrative Observed treatment, proxies, and outcome flow to the empirical five-block summary, then to structured-lattice matching, and finally to an atomic quotient-law estimate and its Wasserstein confidence report.

Confidence Sets

  • For miscoverage level α\alpha, the summary radius rn,αr_{n,\alpha} describes simultaneous moment uncertainty.
  • The exact-real radius Rn,αR_{n,\alpha} adds the lattice approximation scale.
  • Our finite constrained exact-real set contains atomic laws with aggregate atom mass at least π0\pi_0 and transport distance at most Rn,αR_{n,\alpha} from λ^n\widehat\lambda_n.
  • A companion set maps nearby feasible summaries through the repaired law functional.

informal · Theorem T-6 Under fixed dimensions, boundedness, positivity, and rank margins, both nonempty confidence sets cover νP\nu_P jointly with probability at least 1α1-\alpha and have W1W_1 diameter at most their stated root-nn radii.

Theorem T-6 (Honest root-\(n\) confidence)

Fix integers k,dx,dzk,d_x,d_z and constants π0,σ0\pi_0,\sigma_0 such that 2k,kdx,kdz,0<π012k,0<σ01. 2\le k,\qquad k\le d_x,\qquad k\le d_z,\qquad 0<\pi_0\le {1\over 2k},\qquad 0<\sigma_0\le 1 . There exists a simultaneous-concentration constant C01C_0\ge 1 such that, for every boundedness radius L1L\ge 1, there is a constant Cmod>0C_{\mathrm{mod}}>0 for which the prescribed structured-lattice constant ClatC_{\mathrm{lat}} is positive and the following holds. For every sample size n1n\ge 1, there exist:

  • (Structured lattice estimator.) a measurable estimator λ^n\widehat\lambda_n taking values in Pk([Lτ,Lτ])\mathcal P_k([-L_\tau,L_\tau]), generated by a finite exhaustive list of well-formed structured-lattice candidates, ordered lexicographically, with λ^n\widehat\lambda_n equal to the effect law of the first lexicographic minimizer of the structured-lattice criterion;
  • (Summary repair data.) a continuous extension F\overline F on the summary closure K\mathcal K, a measurable nearest-summary selector Π\Pi with its empty-closure fallback, and the associated measurable repaired summary-to-law map.

For every miscoverage level α\alpha satisfying 0<α<12, 0<\alpha< {1\over 2}, let Cn,α\mathcal C_{n,\alpha}, Rn,αR_{n,\alpha}, and Cn,αalg\mathcal C^{\mathrm{alg}}_{n,\alpha} be the confidence-set objects of Definition P-10 formed from this repair data and this structured-lattice estimator, with m=π0m_\star=\pi_0 and center λ^n\widehat\lambda_n. Then, on every sample, both Cn,α\mathcal C_{n,\alpha} and Cn,αalg\mathcal C^{\mathrm{alg}}_{n,\alpha} are nonempty, and Cn,αalg\mathcal C^{\mathrm{alg}}_{n,\alpha} has the finite constrained transport representation νCn,αalgAtomFloor(π0,ν) and there exists a transport plan γ from ν to λ^n with cost at most Rn,α. \nu\in\mathcal C^{\mathrm{alg}}_{n,\alpha} \quad\Longleftrightarrow\quad \operatorname{AtomFloor}(\pi_0,\nu)\ \text{and there exists a transport plan }\gamma\text{ from }\nu\text{ to }\widehat\lambda_n \text{ with cost at most }R_{n,\alpha}. Moreover, for every full-data probability law PMP\in\mathcal M in Definition P-1, 1αQP(n){νPCn,αCn,αalg}. 1-\alpha \le Q_P^{(n)} \left\{ \nu_P\in\mathcal C_{n,\alpha}\cap\mathcal C^{\mathrm{alg}}_{n,\alpha} \right\}. Finally, on every sample, diamW1(Cn,α)4Cmodrn,α,diamW1(Cn,αalg)2Rn,α. \operatorname{diam}_{W_1}(\mathcal C_{n,\alpha}) \le 4C_{\mathrm{mod}}\, r_{n,\alpha}, \qquad \operatorname{diam}_{W_1}(\mathcal C^{\mathrm{alg}}_{n,\alpha}) \le 2R_{n,\alpha}.

informal · Theorem T-4 The repaired estimator is total and Borel, and its exact repaired-image confidence set is nonempty for every sample.

Cluster Report

  • The association radius ρn,α\rho_{n,\alpha} is the confidence radius divided by the atom-mass floor.
  • We connect estimated support points within 4ρn,α4\rho_{n,\alpha}; each connected component is an empirically unresolved effect cluster.
  • Each cluster receives a support interval and the range of compatible aggregate masses over the confidence set.
  • In the benchmark collision, the single effect cluster carries aggregate mass one.
  • External separation sharpens the mass interval for an isolated cluster.

informal · Theorem T-7 With probability at least 1α1-\alpha, the empirical clusters uniquely cover the true support, their intervals cover the associated atoms and aggregate masses, and their mass-interval widths are at most the stated inverse-external-gap bounds.

Theorem T-7 (Cluster-adaptive report validity)

Let k,dx,dzNk,d_x,d_z\in\mathbb N, π0R\pi_0\in\mathbb R, and σ0R\sigma_0\in\mathbb R. Suppose that

  • (Dimensions.) 2k2\le k, kdxk\le d_x, and kdzk\le d_z.
  • (Latent mass.) 0<π01/(2k)0<\pi_0\le 1/(2k).
  • (Rank margin.) 0<σ010<\sigma_0\le 1.

Then there is a simultaneous-concentration constant C01C_0\ge 1 such that, for every L1L\ge 1, if Lτ=4Ldzσ0, L_\tau=\frac{4L\sqrt{d_z}}{\sigma_0}, and s0=π0σ02,cV=4dxk,KD=4kLLτσ0,AD=83s0+32Ls02,Kf=16dxkL2σ02,cD=2KDcV+4kkLLτσ02+2kLσ0+kLτσ0,cm=2kLcV+k+kL,cb=k+2kLcV,cgrid=cD+cm+cb,Blat=2(AD+1)+cgrid, \begin{gathered} s_0=\pi_0\sigma_0^2,\qquad c_V=4\sqrt{d_xk},\qquad K_D=\frac{4\sqrt{k}LL_\tau}{\sigma_0},\\ A_D=\frac{8}{3s_0}+\frac{32L}{s_0^2},\qquad K_f=\frac{16\sqrt{d_x}kL^2}{\sigma_0^2},\\ c_D=2K_Dc_V+\frac{4k\sqrt{k}LL_\tau}{\sigma_0^2} +\frac{2\sqrt{k}L}{\sigma_0}+\frac{kL_\tau}{\sigma_0},\\ c_m=2\sqrt{k}Lc_V+k+kL,\qquad c_b=k+2\sqrt{k}Lc_V,\\ c_{\mathrm{grid}}=c_D+c_m+c_b,\qquad B_{\mathrm{lat}}=2(A_D+1)+c_{\mathrm{grid}}, \end{gathered} then the prescribed structured-lattice stability constant Clat=max{8Lτs0,(Lτ+KD+LKf)Blat} C_{\mathrm{lat}} = \max\left\{ \frac{8L_\tau}{s_0}, \, (L_\tau+K_D+LK_f)B_{\mathrm{lat}} \right\} is positive. For every n1n\ge 1, there exist a measurable structured-lattice estimator λ^n\widehat\lambda_n and repair data consisting of a continuous extension on K\mathcal K and a Borel nearest-summary selector, such that λ^n\widehat\lambda_n is obtained from an exhaustive lexicographically ordered finite list of well-formed structured-lattice candidates by choosing the first minimizer of Jn(ϑ;S^n)J_n(\vartheta;\widehat S_n). For every α\alpha with 0<α<1/20<\alpha<1/2, the estimator has atom floor π0\pi_0. Moreover, for every PMP\in\mathcal M in Definition P-1 and every sample, W1 ⁣(λ^n,νP)Clat{dS ⁣(S^n,S(P))+1n}. W_1\!\left(\widehat\lambda_n,\nu_P\right) \le C_{\mathrm{lat}} \left\{ d_S\!\left(\widehat S_n,S(P)\right)+\frac{1}{\sqrt n} \right\}. Define rn,α=C0Llog(C0/α)n,Rn,α=Clat(rn,α+1n),ρn,α=Rn,απ0, r_{n,\alpha} = C_0L\sqrt{\frac{\log(C_0/\alpha)}{n}}, \qquad R_{n,\alpha} = C_{\mathrm{lat}}\left(r_{n,\alpha}+\frac{1}{\sqrt n}\right), \qquad \rho_{n,\alpha}=\frac{R_{n,\alpha}}{\pi_0}, with S^n\widehat S_n as in Definition P-3, and set EP={ω:dS ⁣(S^n(ω),S(P))rn,α}. E_P=\left\{\omega: d_S\!\left(\widehat S_n(\omega),S(P)\right)\le r_{n,\alpha} \right\}. Then QP(n)(EP)1α. Q_P^{(n)}(E_P)\ge 1-\alpha. For every sample ωEP\omega\in E_P, the cluster report Rn,α\mathfrak R_{n,\alpha} centered at λ^n(ω)\widehat\lambda_n(\omega) satisfies the following properties for the representative atomic law νP\nu_P:

  • (Partition.) Each support point xsupp(νP)x\in\operatorname{supp}(\nu_P) belongs to a unique empirical component CK^n,αC\in\widehat{\mathscr K}_{n,\alpha} with xKC(P)x\in K_C(P).
  • (Association.) For every CK^n,αC\in\widehat{\mathscr K}_{n,\alpha}, the set KC(P)K_C(P) is nonempty and is contained in supp(νP)\operatorname{supp}(\nu_P).
  • (Support interval.) If CK^n,αC\in\widehat{\mathscr K}_{n,\alpha} and xKC(P)x\in K_C(P), then xx lies in the reported support interval for CC.
  • (Mass interval.) If CK^n,αC\in\widehat{\mathscr K}_{n,\alpha}, then the reported mass interval ICI_C contains νP(KC(P))\nu_P(K_C(P)).
  • (Singleton width.) If CK^n,αC\in\widehat{\mathscr K}_{n,\alpha} and card(KC(P))=1\operatorname{card}(K_C(P))=1, then the reported support interval for CC has length at most 4ρn,α4\rho_{n,\alpha}, and length(IC)8π0min{1,Rn,αδ(P)}. \operatorname{length}(I_C) \le \frac{8}{\pi_0} \min\left\{1,\frac{R_{n,\alpha}}{\delta(P)}\right\}.
  • (External-gap width.) If CK^n,αC\in\widehat{\mathscr K}_{n,\alpha} and ΔC(P)=+\Delta_C(P)=+\infty, then length(IC)=0\operatorname{length}(I_C)=0. For every CK^n,αC\in\widehat{\mathscr K}_{n,\alpha}, length(IC)8π0min{1,Rn,αΔC(P)}, \operatorname{length}(I_C) \le \frac{8}{\pi_0} \min\left\{1,\frac{R_{n,\alpha}}{\Delta_C(P)}\right\}, with the displayed right side read as zero when ΔC(P)=+\Delta_C(P)=+\infty.
  • (Constrained representation.) The computable class Cn,αalg\mathcal C^{\mathrm{alg}}_{n,\alpha} has the finite constrained-program representation ξCn,αalgAtomFloor(π0,ξ) and γ with transport cost at most Rn,α from ξ to λ^n(ω). \xi\in\mathcal C^{\mathrm{alg}}_{n,\alpha} \quad\Longleftrightarrow\quad \operatorname{AtomFloor}(\pi_0,\xi) \ \text{and}\ \exists\gamma\ \text{with transport cost at most } R_{n,\alpha} \text{ from }\xi\text{ to }\widehat\lambda_n(\omega).

Ordered Weights

  • The ordered target p(P)p^{\uparrow}(P) lists latent masses by increasing treatment effect when all kk effects are distinct.
  • The effect-gap scale gg places the smallest positive gap between g/2g/2 and 2g2g.
  • Converting W1W_1 error into individual mass error divides by the distance available to transport misallocated mass.
  • Wider gaps support precise labeling; shrinking gaps increase the cost of resolving effect order.

informal · Theorem T-8 On the gap-local stratum, our ordered-weight estimator has expected 1\ell_1 error at most Cmin{1,(ng)1}C\min\{1,(\sqrt n\,g)^{-1}\}.

Theorem T-8 (Ordered-weight upper bound)

Let k,dx,dzNk,d_x,d_z\in\mathbb N and L,π0,σ0RL,\pi_0,\sigma_0\in\mathbb R. Assume:

  • (Latent size.) 2k2\le k.
  • (Proxy dimensions.) kdxk\le d_x and kdzk\le d_z.
  • (Envelope.) 1L1\le L.
  • (Latent-arm margin.) 0<π0(2k)10<\pi_0\le (2k)^{-1}.
  • (Rank margin.) 0<σ010<\sigma_0\le 1.

Then there is a constant C=C(k,dx,dz,L,π0,σ0)>0C=C(k,d_x,d_z,L,\pi_0,\sigma_0)>0 such that, for every n1n\ge1, one can choose the simplex-valued measurable ordered-weight estimator p^n\widehat p_n^{\uparrow} from Definition P-12 so that, for every gap scale gg satisfying 0<g1/40<g\le1/4 and every full-data probability law PM(g)P\in\mathcal M(g), EQP(n) ⁣[p^np(P)1]Cmin{1,(ng)1}. E_{Q_P^{(n)}}\!\left[\left\|\widehat p_n^{\uparrow}-p^{\uparrow}(P)\right\|_1\right] \le C\min\left\{1,\left(\sqrt n\,g\right)^{-1}\right\}.

Lower Bounds

  • Our explicit two-class Bernoulli witness retains strict positivity and full proxy rank as its two effects collide.
  • For quotient laws, we compare the collision with an effect displacement of a/na/\sqrt n.
  • For ordered weights, a factorization-preserving path changes masses by a displacement hh while the observed Kullback–Leibler divergence scales as g2h2g^2h^2.
  • Le Cam (1986) and Polyanskiy and Wu (2019) then convert these close observed experiments into risk lower bounds.

informal · Theorem T-9 For 0ε1/80\le\varepsilon\le1/8, the two-class witness belongs to the uniformly conditioned model with the stated margins and has quotient law 0.4δ0.25ε+0.6δ0.25+ε0.4\,\delta_{0.25-\varepsilon}+0.6\,\delta_{0.25+\varepsilon}, including a nonsingular collision at ε=0\varepsilon=0.

informal · Theorem T-10 On the explicit two-class local experiments, every estimator has quotient-law risk at least c/nc/\sqrt n and ordered-weight risk at least cmin{1,1/(ng)}c\min\{1,1/(\sqrt n\,g)\}, with the stated observed-divergence certificates.

Theorem T-10 (Matching local lower bounds)

Fix cloc(0,1)c_{\mathrm{loc}}\in(0,1). There exist constants a,c,CRa,c,C\in\mathbb R such that 0<a18,c>0,C>0, 0<a\le \frac18,\qquad c>0,\qquad C>0, and, for every integer n1n\ge 1, the following statements hold for the explicit submodel k=dx=dz=2k=d_x=d_z=2, L=2L=2, π0=1/10\pi_0=1/10, and σ0=1/10\sigma_0=1/10.

  • (Quotient-law risk.) For every measurable quotient-law estimator ν~n\widetilde\nu_n with support radius Lτ=4Ldzσ0, L_\tau=\frac{4L\sqrt{d_z}}{\sigma_0}, there is a law PP in the local quotient-law experiment Lnν\mathcal L_n^\nu of Definition P-14 such that P=P0witorP=Pa/nwit, P=P_0^{\mathrm{wit}} \quad\text{or}\quad P=P_{a/\sqrt n}^{\mathrm{wit}}, with PεwitP_\varepsilon^{\mathrm{wit}} as in Definition P-13, and cnEQP(n) ⁣[W1(ν~n,νP)]. \frac{c}{\sqrt n} \le E_{Q_P^{(n)}}\!\left[W_1(\widetilde\nu_n,\nu_P)\right].
  • (Ordered-weight risk.) For every g(0,1/4]g\in(0,1/4] and every measurable ordered-weight estimator p~n\widetilde p_n taking values in the simplex, there is a law PP in the local labeled-weight experiment Ln,gp\mathcal L_{n,g}^{p} of Definition P-15 such that P=Pg,0orP=Pg,h(n,g),h(n,g)=amin{1,1ng}, P=P_{g,0} \quad\text{or}\quad P=P_{g,h(n,g)}, \qquad h(n,g)=a\min\left\{1,\frac{1}{\sqrt n\,g}\right\}, and cmin{1,1ng}EQP(n) ⁣[p~np(P)1]. c\min\left\{1,\frac{1}{\sqrt n\,g}\right\} \le E_{Q_P^{(n)}}\!\left[\|\widetilde p_n-p^{\uparrow}(P)\|_1\right].
  • (Quotient-law certificate.) The laws P0witP_0^{\mathrm{wit}} and Pa/nwitP_{a/\sqrt n}^{\mathrm{wit}} belong to the corresponding local quotient-law experiments, and their observed-data marginals and quotient laws satisfy DKL ⁣((Pa/nwit)O(P0wit)O)Cn,cnW1 ⁣(νP0wit,νPa/nwit). D_{\mathrm{KL}}\!\left((P_{a/\sqrt n}^{\mathrm{wit}})_O\Vert (P_0^{\mathrm{wit}})_O\right) \le \frac{C}{n}, \qquad \frac{c}{\sqrt n} \le W_1\!\left(\nu_{P_0^{\mathrm{wit}}},\nu_{P_{a/\sqrt n}^{\mathrm{wit}}}\right).
  • (Labeled-path certificate.) For every g(0,1/4]g\in(0,1/4], with h=h(n,g)=amin{1,1ng}, h=h(n,g)=a\min\left\{1,\frac{1}{\sqrt n\,g}\right\}, the displacement hh lies in the tangent-amplitude domain, the laws Pg,hP_{g,h} and Pg,0P_{g,0} belong to the corresponding local labeled-weight experiments, and DKL ⁣((Pg,h)O(Pg,0)O)Cg2h2,ipi(Pg,h)pi(Pg,0)=2h. D_{\mathrm{KL}}\!\left((P_{g,h})_O\Vert (P_{g,0})_O\right) \le C g^2h^2, \qquad \sum_i\left|p_i^{\uparrow}(P_{g,h})-p_i^{\uparrow}(P_{g,0})\right| = 2|h|.

informal · Theorem T-11 On the explicit two-class model, quotient-law minimax risk is bounded above and below by constant multiples of 1/n1/\sqrt n.

informal · Theorem T-12 On the explicit two-class gap stratum, ordered-weight minimax risk is bounded above and below by constant multiples of min{1,1/(ng)}\min\{1,1/(\sqrt n\,g)\}.

informal · Theorem T-14 A published-VMW comparator class inherits the applicable lower bound whenever it contains the displayed witness pair.

Scope of the Guarantees

  • The estimand is the population-weighted law of latent-class mean effects, rather than the law of individual treatment effects.
  • The procedures use supplied values of k,L,π0,σ0k,L,\pi_0,\sigma_0; the guarantees are uniform over the resulting fixed class.
  • The computational theorem gives a finite constrained representation and a fixed-dimensional exact-real operation bound. It does not claim polynomial bit complexity or a numerical implementation.
  • Matching minimax rates are proved on the explicit uniformly conditioned two-class specialization; the general model-class results are upper bounds.

Takeaways

  • Aggregating coincident class-average effects makes their population-weighted law a stable target for proxy-based causal inference.
  • Observable proxy moments determine this quotient law through a gap-free W1W_1 modulus.
  • Structured estimation and honest confidence reporting achieve uniform root-nn accuracy under fixed boundedness, positivity, and proxy-rank margins.
  • Cluster reports express uncertainty at the support resolution available in the sample.
  • Effect-ordered masses have the sharp clipped inverse-gap rate on the stated separated two-class configurations.