CausalSmith · seminar slides
Quotient-Law Inference for Latent-Class Mean Effects at Collisions
We estimate the population-weighted distribution of latent-class average treatment effects uniformly at the root-n rate, including configurations where several classes share the same mean effect. This estimand is distinct from the distribution of individual contrasts Y(1)−Y(0).
slides for Quotient-law Inference with Latent Treatment-effect Collisions
Overview
- Our target is the quotient law νP, the population-weighted distribution of latent-class average treatment effects after aggregating classes with the same mean effect.
- In the one-Wasserstein transport distance W1, observable proxy moments determine this law through a gap-free Lipschitz map.
- Law estimation and honest confidence reporting attain uniform root-n accuracy.
- Effect-ordered class masses incur the sharp inverse-gap cost on separated two-class configurations.
Motivation
- Think of an observational treatment study where an unobserved response type affects treatment and outcomes.
- Two proxy measurements carry complementary information about that latent type.
- In our two-class benchmark, the class-average effects are 0.25−ε and 0.25+ε, where ε is their displacement from a collision.
- The population masses are 0.4 and 0.6.
- At ε=0, both types contribute to one treatment-effect value; for ε>0, the law has two nearby atoms.
- Can one infer the law of class-average effects continuously across this transition?
Target
- The quotient law places each latent-class mass at its class-average treatment effect and adds the masses of coincident mean effects.
- W1 is the minimum distance-weighted mass required to transport one effect law into another.
- In the benchmark, the quotient law moves continuously from two atoms to one aggregate atom as ε approaches zero.
- It does not identify the distribution of individual contrasts Y(1)−Y(0); its atoms are conditional class means.
Model
- We observe treatment, two proxy vectors, and the realized outcome for independent units.
- The analyst supplies the fixed class cardinality k and uniform bounds L,π0,σ0; consistency and latent ignorability give each class a causal mean treatment contrast.
- Conditional separation makes one proxy a class measurement and the other an arm-specific class measurement.
- Observable contributions are bounded.
- The latent-arm positivity margin π0 and proxy-rank margin σ0 remain fixed across the model class.
- In our benchmark, the proxy matrices remain nonsingular at ε=0, and every latent class–treatment arm has positive mass.
The latent class and treatment arms satisfy 1≤u≤k, t∈{0,1}minP(U=u,T=t)≥π0.
The reference- and target-proxy feature matrices satisfy min{σk(A0),σk(A1),σk(B)}≥σ0.
Observable Moments
- The five-block summary S(P) consists of arm-specific proxy moments, outcome-weighted proxy moments, and the target-proxy mean.
- Its empirical version Sn is formed by sample averaging within treatment arms.
- A compressed operator built from these moments has the latent treatment effects as eigenvalues.
- Spectral anchors recover the masses attached to the corresponding eigenspaces.
- The summary distance dS adds the operator-norm errors of the four matrix blocks and the Euclidean error of the proxy mean.
informal · Theorem T-1 Under the fixed model restrictions, treatment arms have positive probability, the observable proxy moments have stable factorizations and singular-value margins, and latent effects remain within a fixed bounded interval.
informal · Theorem T-2 Under fixed dimensions and margins, the closure of feasible five-block summaries is compact, with nonemptiness equivalent to nonemptiness of the model class.
Related Literature
- Miao et al. (2018) identify causal effects using conditionally independent proxies and rank conditions.
- Mazaheri et al. (2025) and Virk et al. (2026) develop finite latent-effect identification and separated spectral recovery.
- Heinrich and Kahn (2018) and Wu and Yang (2020) show how collision geometry shapes rates in ordinary finite mixtures.
- Deo and Randrianarisoa (2023) develop honest Wasserstein confidence sets for mixture uncertainty.
- Our contribution uses the proxy operator to obtain regular inference for the aggregated effect law across homogeneous, partially colliding, and separated configurations.
Key Idea
- A direct plug-in eigendecomposition tracks individual eigenvectors and amplifies perturbations by the inverse effect gap.
- Near a collision, individual eigendirections can rotate sharply while their combined spectral subspace remains stable.
- We aggregate the anchor mass over each colliding eigenspace and transport that mass across the cluster diameter.
- The aggregation absorbs simultaneous merges and splits into the W1 geometry.
- Uniform proxy conditioning controls the compressed operator; positivity controls the mass carried by each distinct atom.
- Together these facts convert observable-moment error into law-level transport error without an effect-gap factor.
Main Result
informal · Theorem T-3 Under fixed dimensions, bounded contributions, latent-arm positivity, and proxy-rank margins, the W1 distance between two quotient laws is at most a constant times the distance between their observable summaries, uniformly across collisions.
Fix natural numbers k,dx,dz and real constants L,π0,σ0. Assume:
- (Latent dimension.) 2≤k, k≤dx, and k≤dz.
- (Bound and margins.) 1≤L, 0<π0≤(2k)−1, and 0<σ0≤1.
Let Lτ=4Ldz/σ0, and let K=S be the summary closure from Definition P-8. For two summaries s=(M0,M1,N0,N1,mX),q=(M0′,M1′,N0′,N1′,mX′), write dS(s,q)=∥M0−M0′∥+∥M1−M1′∥+∥N0−N0′∥+∥N1−N1′∥+(i∑(mX(i)−mX′(i))2)1/2. There exists a constant Cmod>0 such that, for every pair of full-data laws P,Q∈M, W1(νP,νQ)≤CmoddS(S(P),S(Q)). Moreover, there is a map F:K→Pk([−Lτ,Lτ]) such that F is continuous and Borel measurable, W1(F(q),F(q′))≤CmoddS(q,q′)for all q,q′∈K, and F(S(Q))=νQfor every Q∈M. This map is unique among continuous maps G:K→Pk([−Lτ,Lτ]) that obey W1(G(q),G(q′))≤CmoddS(q,q′)for all q,q′∈K and satisfy G(S(Q))=νQ for every Q∈M: every such G satisfies G(q)=F(q) for all q∈K.
The same Lipschitz law map covers the one-atom benchmark at ε=0 and its two-atom neighbors.
Estimation
- Uniform concentration controls all five empirical-summary blocks at the usual sampling scale.
- Our repair estimator selects the nearest feasible summary and applies the Lipschitz law map.
- Our lattice estimator searches a finite constrained representation of discretized signal bases, conditioning matrices, masses, and effect locations for the best exact-real match to the empirical moments.
- Both mechanisms return positive atomic effect laws.
informal · Lemma L-2 Uniformly over the model class, the empirical summary exceeds C0Llog(C0/η)/n error with probability at most the tail probability η.
informal · Theorem T-5 With fixed dimensions and margins, both law estimators have W1 error at most Clog(C/η)/n simultaneously with probability at least 1−η, uniformly over the model class.
Fix integers k,dx,dz and constants L,π0,σ0. Suppose that
- (Dimensions.) 2≤k, k≤dx, and k≤dz.
- (Uniform bounds.) 1≤L, 0<π0≤1/(2k), and 0<σ0≤1.
Then there is a constant C=C(k,dx,dz,L,π0,σ0)>0 such that, for every n≥1, there exist a structured-lattice estimator λn taking values in Pk([−Lτ,Lτ]), with Lτ=4Ldz/σ0, and a nearest-summary repair estimator νn=F(Π(Sn)) from Definition P-9, such that λn is the first lexicographic minimizer over a finite exhaustive structured lattice of well-formed candidates ϑ for the criterion Jn(ϑ;Sn), with tie-breaking by the lattice order and output λϑ. Both estimators are Borel sample maps. For every η satisfying 0<η<1/2, every P∈M, and QP(n)-sample, QP(n){max{W1(λn,νP),W1(νn,νP)}>Cnlog(C/η)}≤η. Moreover, for each treatment arm t∈{0,1}, QP(n){Nt,n=0}≤(1−kπ0)n.
Estimator Pipeline
- Thresholding extracts the k-dimensional proxy signal space at the fixed scale π0σ02/2.
- The lattice compares candidate effect operators, proxy means, and anchor equations.
- The first minimizing candidate supplies the estimated atomic law λn.
- The mesh contributes 1/n, matching the sampling scale.
informal · Theorem T-15 For fixed dimensions and margins, the structured-lattice estimator has W1 error at most Clat{dS(Sn,S(P))+1/n}, a polynomial-size candidate list, and the stated uniform root-n tail bound.
informal · Theorem T-13 For fixed dimensions and supplied exact-real primitives, a finite-library estimator is Borel, uses at most C{n+n(4dzdx+dx)/2} operations, and attains the stated uniform root-n tail bound.
Confidence Sets
- For miscoverage level α, the summary radius rn,α describes simultaneous moment uncertainty.
- The exact-real radius Rn,α adds the lattice approximation scale.
- Our finite constrained exact-real set contains atomic laws with aggregate atom mass at least π0 and transport distance at most Rn,α from λn.
- A companion set maps nearby feasible summaries through the repaired law functional.
informal · Theorem T-6 Under fixed dimensions, boundedness, positivity, and rank margins, both nonempty confidence sets cover νP jointly with probability at least 1−α and have W1 diameter at most their stated root-n radii.
Fix integers k,dx,dz and constants π0,σ0 such that 2≤k,k≤dx,k≤dz,0<π0≤2k1,0<σ0≤1. There exists a simultaneous-concentration constant C0≥1 such that, for every boundedness radius L≥1, there is a constant Cmod>0 for which the prescribed structured-lattice constant Clat is positive and the following holds. For every sample size n≥1, there exist:
- (Structured lattice estimator.) a measurable estimator λn taking values in Pk([−Lτ,Lτ]), generated by a finite exhaustive list of well-formed structured-lattice candidates, ordered lexicographically, with λn equal to the effect law of the first lexicographic minimizer of the structured-lattice criterion;
- (Summary repair data.) a continuous extension F on the summary closure K, a measurable nearest-summary selector Π with its empty-closure fallback, and the associated measurable repaired summary-to-law map.
For every miscoverage level α satisfying 0<α<21, let Cn,α, Rn,α, and Cn,αalg be the confidence-set objects of Definition P-10 formed from this repair data and this structured-lattice estimator, with m⋆=π0 and center λn. Then, on every sample, both Cn,α and Cn,αalg are nonempty, and Cn,αalg has the finite constrained transport representation ν∈Cn,αalg⟺AtomFloor(π0,ν) and there exists a transport plan γ from ν to λn with cost at most Rn,α. Moreover, for every full-data probability law P∈M in Definition P-1, 1−α≤QP(n){νP∈Cn,α∩Cn,αalg}. Finally, on every sample, diamW1(Cn,α)≤4Cmodrn,α,diamW1(Cn,αalg)≤2Rn,α.
informal · Theorem T-4 The repaired estimator is total and Borel, and its exact repaired-image confidence set is nonempty for every sample.
Cluster Report
- The association radius ρn,α is the confidence radius divided by the atom-mass floor.
- We connect estimated support points within 4ρn,α; each connected component is an empirically unresolved effect cluster.
- Each cluster receives a support interval and the range of compatible aggregate masses over the confidence set.
- In the benchmark collision, the single effect cluster carries aggregate mass one.
- External separation sharpens the mass interval for an isolated cluster.
informal · Theorem T-7 With probability at least 1−α, the empirical clusters uniquely cover the true support, their intervals cover the associated atoms and aggregate masses, and their mass-interval widths are at most the stated inverse-external-gap bounds.
Let k,dx,dz∈N, π0∈R, and σ0∈R. Suppose that
- (Dimensions.) 2≤k, k≤dx, and k≤dz.
- (Latent mass.) 0<π0≤1/(2k).
- (Rank margin.) 0<σ0≤1.
Then there is a simultaneous-concentration constant C0≥1 such that, for every L≥1, if Lτ=σ04Ldz, and s0=π0σ02,cV=4dxk,KD=σ04kLLτ,AD=3s08+s0232L,Kf=σ0216dxkL2,cD=2KDcV+σ024kkLLτ+σ02kL+σ0kLτ,cm=2kLcV+k+kL,cb=k+2kLcV,cgrid=cD+cm+cb,Blat=2(AD+1)+cgrid, then the prescribed structured-lattice stability constant Clat=max{s08Lτ,(Lτ+KD+LKf)Blat} is positive. For every n≥1, there exist a measurable structured-lattice estimator λn and repair data consisting of a continuous extension on K and a Borel nearest-summary selector, such that λn is obtained from an exhaustive lexicographically ordered finite list of well-formed structured-lattice candidates by choosing the first minimizer of Jn(ϑ;Sn). For every α with 0<α<1/2, the estimator has atom floor π0. Moreover, for every P∈M in Definition P-1 and every sample, W1(λn,νP)≤Clat{dS(Sn,S(P))+n1}. Define rn,α=C0Lnlog(C0/α),Rn,α=Clat(rn,α+n1),ρn,α=π0Rn,α, with Sn as in Definition P-3, and set EP={ω:dS(Sn(ω),S(P))≤rn,α}. Then QP(n)(EP)≥1−α. For every sample ω∈EP, the cluster report Rn,α centered at λn(ω) satisfies the following properties for the representative atomic law νP:
- (Partition.) Each support point x∈supp(νP) belongs to a unique empirical component C∈Kn,α with x∈KC(P).
- (Association.) For every C∈Kn,α, the set KC(P) is nonempty and is contained in supp(νP).
- (Support interval.) If C∈Kn,α and x∈KC(P), then x lies in the reported support interval for C.
- (Mass interval.) If C∈Kn,α, then the reported mass interval IC contains νP(KC(P)).
- (Singleton width.) If C∈Kn,α and card(KC(P))=1, then the reported support interval for C has length at most 4ρn,α, and length(IC)≤π08min{1,δ(P)Rn,α}.
- (External-gap width.) If C∈Kn,α and ΔC(P)=+∞, then length(IC)=0. For every C∈Kn,α, length(IC)≤π08min{1,ΔC(P)Rn,α}, with the displayed right side read as zero when ΔC(P)=+∞.
- (Constrained representation.) The computable class Cn,αalg has the finite constrained-program representation ξ∈Cn,αalg⟺AtomFloor(π0,ξ) and ∃γ with transport cost at most Rn,α from ξ to λn(ω).
Ordered Weights
- The ordered target p↑(P) lists latent masses by increasing treatment effect when all k effects are distinct.
- The effect-gap scale g places the smallest positive gap between g/2 and 2g.
- Converting W1 error into individual mass error divides by the distance available to transport misallocated mass.
- Wider gaps support precise labeling; shrinking gaps increase the cost of resolving effect order.
informal · Theorem T-8 On the gap-local stratum, our ordered-weight estimator has expected ℓ1 error at most Cmin{1,(ng)−1}.
Let k,dx,dz∈N and L,π0,σ0∈R. Assume:
- (Latent size.) 2≤k.
- (Proxy dimensions.) k≤dx and k≤dz.
- (Envelope.) 1≤L.
- (Latent-arm margin.) 0<π0≤(2k)−1.
- (Rank margin.) 0<σ0≤1.
Then there is a constant C=C(k,dx,dz,L,π0,σ0)>0 such that, for every n≥1, one can choose the simplex-valued measurable ordered-weight estimator pn↑ from Definition P-12 so that, for every gap scale g satisfying 0<g≤1/4 and every full-data probability law P∈M(g), EQP(n)[pn↑−p↑(P)1]≤Cmin{1,(ng)−1}.
Lower Bounds
- Our explicit two-class Bernoulli witness retains strict positivity and full proxy rank as its two effects collide.
- For quotient laws, we compare the collision with an effect displacement of a/n.
- For ordered weights, a factorization-preserving path changes masses by a displacement h while the observed Kullback–Leibler divergence scales as g2h2.
- Le Cam (1986) and Polyanskiy and Wu (2019) then convert these close observed experiments into risk lower bounds.
informal · Theorem T-9 For 0≤ε≤1/8, the two-class witness belongs to the uniformly conditioned model with the stated margins and has quotient law 0.4δ0.25−ε+0.6δ0.25+ε, including a nonsingular collision at ε=0.
informal · Theorem T-10 On the explicit two-class local experiments, every estimator has quotient-law risk at least c/n and ordered-weight risk at least cmin{1,1/(ng)}, with the stated observed-divergence certificates.
Fix cloc∈(0,1). There exist constants a,c,C∈R such that 0<a≤81,c>0,C>0, and, for every integer n≥1, the following statements hold for the explicit submodel k=dx=dz=2, L=2, π0=1/10, and σ0=1/10.
- (Quotient-law risk.) For every measurable quotient-law estimator νn with support radius Lτ=σ04Ldz, there is a law P in the local quotient-law experiment Lnν of Definition P-14 such that P=P0witorP=Pa/nwit, with Pεwit as in Definition P-13, and nc≤EQP(n)[W1(νn,νP)].
- (Ordered-weight risk.) For every g∈(0,1/4] and every measurable ordered-weight estimator pn taking values in the simplex, there is a law P in the local labeled-weight experiment Ln,gp of Definition P-15 such that P=Pg,0orP=Pg,h(n,g),h(n,g)=amin{1,ng1}, and cmin{1,ng1}≤EQP(n)[∥pn−p↑(P)∥1].
- (Quotient-law certificate.) The laws P0wit and Pa/nwit belong to the corresponding local quotient-law experiments, and their observed-data marginals and quotient laws satisfy DKL((Pa/nwit)O∥(P0wit)O)≤nC,nc≤W1(νP0wit,νPa/nwit).
- (Labeled-path certificate.) For every g∈(0,1/4], with h=h(n,g)=amin{1,ng1}, the displacement h lies in the tangent-amplitude domain, the laws Pg,h and Pg,0 belong to the corresponding local labeled-weight experiments, and DKL((Pg,h)O∥(Pg,0)O)≤Cg2h2,i∑pi↑(Pg,h)−pi↑(Pg,0)=2∣h∣.
informal · Theorem T-11 On the explicit two-class model, quotient-law minimax risk is bounded above and below by constant multiples of 1/n.
informal · Theorem T-12 On the explicit two-class gap stratum, ordered-weight minimax risk is bounded above and below by constant multiples of min{1,1/(ng)}.
informal · Theorem T-14 A published-VMW comparator class inherits the applicable lower bound whenever it contains the displayed witness pair.
Scope of the Guarantees
- The estimand is the population-weighted law of latent-class mean effects, rather than the law of individual treatment effects.
- The procedures use supplied values of k,L,π0,σ0; the guarantees are uniform over the resulting fixed class.
- The computational theorem gives a finite constrained representation and a fixed-dimensional exact-real operation bound. It does not claim polynomial bit complexity or a numerical implementation.
- Matching minimax rates are proved on the explicit uniformly conditioned two-class specialization; the general model-class results are upper bounds.
Takeaways
- Aggregating coincident class-average effects makes their population-weighted law a stable target for proxy-based causal inference.
- Observable proxy moments determine this quotient law through a gap-free W1 modulus.
- Structured estimation and honest confidence reporting achieve uniform root-n accuracy under fixed boundedness, positivity, and proxy-rank margins.
- Cluster reports express uncertainty at the support resolution available in the sample.
- Effect-ordered masses have the sharp clipped inverse-gap rate on the stated separated two-class configurations.