Quotient-law Inference with Latent Treatment-effect Collisions
Abstract
This paper studies proxy-based causal inference for latent-class mean treatment effects in a fixed finite-class model. Under the Virk–Mazaheri–Wu proxy moment structure with fixed latent cardinality (the number of latent classes), fixed proxy dimensions, bounded observable contributions, supplied latent-arm positivity and proxy-rank margins, the target is the quotient latent-effect law , the population-weighted distribution of latent-class average treatment effects with masses aggregated at coincident class means. The main population result establishes a gap-free Lipschitz modulus from the observable five-block summary , the vector of proxy moment summaries, to in , the one-Wasserstein distance. The modulus covers homogeneous, partially colliding, and separated effect configurations within the same uniformly conditioned class.
The statistical results turn this modulus into root- law estimation and confidence reporting. A nearest-summary repair estimator and a structured-lattice estimator, calibrated by the supplied constants , are Borel sample maps and achieve uniform root- tail bounds over the model class. In the fixed-dimensional unit-cost exact-real model, the lattice estimator has polynomial candidate growth in and centers an honest finitely represented exact-real Wasserstein outer confidence set. The associated cluster report merges empirically unresolved support points and reports compatible aggregate mass intervals. On separated gap-local strata, effect-ordered latent masses have clipped inverse-gap risk , where is the local effect-gap scale. Local and same-class lower bounds from explicit two-class proxy experiments establish root- sharpness for quotient-law estimation and the inverse-gap rate for ordered weights on the displayed two-class specialization.
Introduction
Proxy causal models use measurements related to latent confounding or latent classes to recover causal objects from observational data. In the potential-outcomes tradition (Splawa-Neyman et al., 1990; Rubin, 1974; Holland, 1986; Imbens et al., 2015), latent heterogeneity is typically organized through assumptions linking treatment assignment, outcomes, and observed covariates. Negative-control and proximal approaches enrich that structure by introducing proxy measurements with conditional-independence and rank properties (Miao et al., 2018; Shi et al., 2020; Tchetgen Tchetgen et al., 2024; Cui et al., 2024; Miao et al., 2023; Qi et al., 2024; Liu et al., 2024; Li et al., 2024; Ai et al., 2025; Saha et al., 2026). The finite-class proxy model of Virk et al. (2026) supplies a compressed observable operator whose spectral structure identifies latent-class mean treatment effects under separated recovery conditions. This paper develops the law-level inference theory for the same proxy moment geometry when equal or nearly equal class-average effects are represented through the probability law of effect values.
The central estimand is the quotient latent-effect law the distribution that places latent-class mass at the class treatment effect and aggregates the mass of classes with equal class means. The support points are latent-class average treatment effects; unit-level contrasts enter through these conditional means. The law records how much population mass has each class-average effect value, with collisions interpreted as shared support points. The observable input is the five-block summary formed from armwise proxy moments, outcome-weighted proxy moments, and the target-proxy mean.
The model class , the uniformly conditioned proxy causal class, fixes , imposes the VMW proxy conditional-independence and consistency restrictions, and uses supplied quantitative bounds: the observable envelope , a lower bound on every joint latent-arm probability, and a singular-value margin for the reference- and target-proxy feature matrices. These conditions yield the observed proxy factorizations and rank margins in Proposition 1, compactness of the feasible summary closure in Proposition 2, and a common compact support interval for latent-class mean effect laws. The fixed-margin formulation is the population domain for all upper bounds and confidence statements in the paper.
The first main result is the gap-free modulus in Theorem 1. It establishes a constant , depending only on the fixed dimensions and conditioning constants, such that for all . It also extends the summary-to-law map continuously to the compact summary closure. The statement is formulated in the one-Wasserstein metric , so perturbations move aggregate probability mass between effect values and treat simultaneous merges or splits of atoms through the same law-level metric.
This population result leads to two estimators. The nearest-summary repair estimator in Algorithm 1 projects an empirical summary to the compact feasible summary closure and evaluates the continuous extension. The structured-lattice estimator in Algorithm 2 searches over finite grids of signal bases, conditioning matrices, latent masses, and effect locations, using a criterion that matches the empirical compressed operator, target-proxy mean, and anchor equation. Theorem 2 establishes that both estimators are Borel sample maps and achieve uniform root- tail bounds over . The lattice construction in Proposition 4 has polynomial candidate growth in for fixed dimensions and constants.
The confidence results report law-level uncertainty in forms aligned with the quotient target. Algorithm 3 defines an exact repaired-image summary-inversion set through the repaired map and a lattice-centered Wasserstein outer set with a finite constrained exact-real representation. Theorem 3 gives simultaneous coverage and root- diameter bounds for these sets. Algorithm 4 and Theorem 4 then translate the Wasserstein confidence set into support-component intervals and aggregate mass intervals. The data-facing report uses the empirical lattice law, supplied constants, and the algorithmic confidence set; the true-law objects and name the validity quantities used to state coverage and width. Empirical atoms are connected at threshold , while support association and interval expansion use the radius ; the resulting mass intervals sharpen according to the external separation of the associated true component.
The paper also characterizes effect-ordered latent masses on separated strata. The ordered target , the vector of latent masses ordered by increasing effect, is studied on , the gap-local stratum with positive effect gaps at scale . Theorem 5 gives the upper risk bound This rate expresses the conversion from law-level transport error to ordered mass recovery when effect locations are separated at scale .
The lower bounds use explicit two-class proxy experiments. The witness law in Definition 55 verifies that a collision can occur inside the full-rank, strictly positive proxy model. Theorem 6 gives local two-point lower bounds for quotient laws and ordered weights, and Propositions 6 and 7 convert them into same-class minimax rates on the displayed two-class specialization. The quotient-law minimax rate is exactly root-, while the ordered-weight minimax rate is the clipped inverse-gap rate. Theorem 7 records the corresponding transfer to comparator classes stated in the published VMW vocabulary when those classes contain the relevant witness pair.
The paper proceeds as follows. The related-work section positions the argument within proxy causal identification, spectral finite-mixture recovery, and Wasserstein inference. The setup section defines the proxy model, observable summaries, quotient law, and compact summary closure. The main-results section proves the law-level modulus and develops the repair and lattice estimators. The confidence section constructs Wasserstein confidence sets, cluster reports, and ordered-weight inference on gap-local strata. The lower-bound section gives the local and same-class minimax converses and the VMW transfer statement. The discussion section collects extensions and open questions, and the appendices provide auxiliary probability, selection, spectral, finite-net, proof, and verification material.
Related work
This paper connects three literatures: proxy causal identification, finite-mixture and spectral recovery, and Wasserstein inference for singular discrete laws. In the Neyman–Rubin potential-outcome tradition (Splawa-Neyman et al., 1990; Rubin, 1974; Holland, 1986; Imbens et al., 2015), latent treatment-effect heterogeneity is usually interpreted through assumptions that relate treatment assignment, outcomes, and observed covariates. Proxy causal models enrich that structure by using measurements linked to latent confounding or latent classes, as in negative-control and proximal identification arguments (Miao et al., 2018; Tchetgen Tchetgen et al., 2024; Shi et al., 2020; Cui et al., 2024; Miao et al., 2023; Qi et al., 2024; Liu et al., 2024; Li et al., 2024; Ai et al., 2025; Saha et al., 2026). The present analysis studies the finite-class proxy setting in which the observed proxy moments identify the quotient latent-effect law of Definition 40: the probability law that aggregates latent classes sharing the same class-average treatment effect. This target is tailored to collisions in effect values, because the inferential object is the distribution of class-average effect values and their aggregate masses.
The closest point of contact is Virk et al. (2026). Their framework supplies the proxy moment structure, the compressed observable operator formalized here in Definition 39, and separated spectral recovery guarantees that make latent-class mean effects estimable when the relevant spectral quantities are well resolved. The present paper uses that operator as the population bridge from observable summaries to the quotient law and develops a modulus, with defined in Definition 19, for the law target. This perspective keeps the metric aligned with the identifiable probability measure: when several latent classes have the same class-average effect, their masses are combined at a common atom, and the stability statement is made for the quotient law. Label-sensitive lists enter through the separated ordered-weight results.
The identification argument also draws on classical and modern work on finite mixtures and tensor or spectral decompositions. Moment-based identification begins with Pearson (1894) and continues through finite-mixture identifiability results such as Teicher (1963), Lindsay (1995), and Allman et al. (2009). Econometric proxy and measurement-error models use related rank and completeness ideas (Hu, 2008; Hu et al., 2008; Kasahara et al., 2009; Bonhomme et al., 2016), while algorithmic spectral methods exploit low-rank decompositions and perturbation theory (Hsu et al., 2009; Anandkumar et al., 2014; Sidiropoulos et al., 2000; Bauer et al., 1960; Kato, 1995; Davis et al., 1970; Stewart et al., 1990; Mazaheri et al., 2025). Within this lineage, the contribution here is a stability and inference theory for the atomic effect law induced by the proxy causal model, with fixed known dimensions and uniform conditioning constants.
A second comparison is with Wasserstein rates for finite mixtures near singularities. Collision phenomena are central in the mixture literature: when component locations approach one another, labeled parameters and weak distributional metrics exhibit different local geometries (Nguyen, 2013; Ho et al., 2016; Heinrich et al., 2018; Ho et al., 2016; Manole et al., 2022; Nguyen et al., 2026). The rates in Heinrich et al. (2018) clarify how overfitting and merging components affect Wasserstein estimation. The present setting is causal and proxy-based, but the same mathematical distinction matters: law-level estimation remains regular for the quotient distribution of latent-class mean effects, while effect-ordered weights are governed by local separation. The paper makes that distinction explicit by pairing collision-uniform quotient-law inference with inverse-gap behavior for ordered weights on the stated gap-local strata.
A quantitative comparison makes the difference in local geometry explicit. The closest results differ from the present paper in three respects at once: the experiment that is observed, the target the loss is measured on, and the conditions under which the stated rate holds.
| result | experiment observed | target and loss | conditions and rate |
|---|---|---|---|
| Heinrich et al. (2018) | draws from an ordinary finite mixture | mixing measure, Wasserstein | strong identifiability, components fitted around an -component law: local minimax rate |
| Doss et al. (2023) | draws from a high-dimensional Gaussian location mixture | mixing measure, Wasserstein | components in dimension : rate |
| Wu et al. (2020) | draws from a Gaussian mixture | mixing measure, Wasserstein | moment projection onto the feasible moment space, attaining adaptive optimal rates |
| Virk et al. (2026) | proxy moments in the same causal model | labelled effects, features, and simplex-projected weights | rank, boundedness, positivity, and spectral separation: high-probability finite-sample first-order recovery for separated effects |
| this paper | proxy moments in the same causal model | quotient effect law , | fixed and uniform envelope, positivity, and rank margins: root-, uniformly over collision configurations |
| this paper | proxy moments in the same causal model | effect-ordered masses | additionally a positive effect gap at scale : clipped inverse-gap risk |
The table locates the source of the difference. The ordinary-mixture results observe draws from a mixture distribution, so the map from the sampling law back to the mixing measure degenerates exactly where components merge, and the exponents above deteriorate with the number of colliding components. The present experiment observes a uniformly conditioned proxy-moment operator instead, and the compressed observable operator of Virk et al. (2026) carries the latent effects in its spectrum. Aggregating the mass of coinciding effects into the quotient law removes the labelling that the collision destroys, and the surviving map from observable summaries to stays Lipschitz across merges and splits by Theorem 1. Root- quotient-law recovery and slower Wasserstein rates for ordinary mixing measures are therefore compatible: they measure different targets in different experiments. The labelled quantity that does retain collision sensitivity here is the effect-ordered mass vector, whose inverse-gap behaviour on gap-local strata is the direct analogue of the mixture exponents, and whose matching converse appears in Theorem 6.
Finally, the confidence-set construction is related to projection and inversion methods for weakly identified or singular models. Generalized method-of-moments traditions emphasize confidence regions obtained by inverting sample restrictions (Hansen, 1982). Recent Wasserstein approaches for finite mixtures, including Wu et al. (2020), Doss et al. (2023), Deo et al. (2023), Bing et al. (2026), and Niles-Weed et al. (2022), develop distributional metrics, singular-rate analyses, and confidence constructions for nonregular finite-mixture settings. Here the observable proxy summary is the primitive sample object, and the confidence reports are expressed directly in Wasserstein distance for the quotient law and as cluster-adaptive support and mass summaries. This yields an econometric reporting target that matches the causal estimand while retaining the finite-mixture insight that separated labels and quotient laws have different local statistical behavior.
Setup and assumptions
Throughout the paper, , , and are fixed finite dimensions, and treatment-arm and latent-class indices range over and . Vector norms are Euclidean, matrix norms are operator norms, and singular values are ordered decreasingly. Generic constants may depend on the fixed dimensions and conditioning constants, while asymptotic statements vary only with the sample size . Expectations and conditional laws are taken under the full-data probability law under discussion unless a subscript specifies otherwise.
The section is organized around one dependency chain, which the reader may find it useful to keep in view: the observed and full-data records fix the sampling object; the causal and proxy assumptions constrain it; the latent masses, means, and effects are the parameters those assumptions govern; the observable population moments and the quotient law are the two derived targets; the model class and the gap-local stratum collect the uniform conditions under which the theorems hold; and the empirical summary is the sample counterpart of the observable moments. The spectral, lattice, and confidence objects that appear alongside these are the machinery that later sections consume, and each is stated where the theorem citing it can be checked against it.
For reading, the logical dependencies are as follows. The observed record and full-data law supply the variables; the causal and proxy assumptions define the uniformly conditioned class ; the population moments , , and form the observable summary ; the quotient law is the population-weighted law of latent-class average treatment effects; and the empirical summary, spectral thresholding, lattice, confidence radii, and cluster objects are derived from those primitives. Some auxiliary objects are displayed where they are needed by later statements; the surrounding prose identifies their role in this dependency chain.
The statistical procedures are calibrated by supplied values of . A full-data law receives the stated guarantees when it satisfies the restrictions of with those supplied constants. A larger valid envelope , or smaller valid lower margins and , preserves membership after the constants and radii are recomputed, with the numerical bounds adjusted accordingly. The joint latent-arm margin encodes overlap within each latent class and arm, while the proxy-rank margin encodes stable variation in the reference- and target-proxy feature matrices.
The observed data consist of the treatment, two proxy vectors, and the realized outcome. The displayed objects below include both primitives and derived objects used later; the prose after each display identifies how it fits with the record spaces, observed margin, transport metric, and inferential target.
Let be an observable summary and put . For , write a real singular-value expansion , with . The thresholded Moore–Penrose inverse of the armwise proxy moment is
The observed record collects the binary treatment , target proxy , reference proxy , and observed outcome . The fixed target- and reference-proxy dimensions are and .
Let and be finite probability laws on . We write when is a finite transport plan with , for every , and for every . Its absolute-distance transport cost is .
The distance compares atomic laws by the minimum absolute-distance transport cost, so it measures the law of effect values. Label-specific comparisons enter later through the ordered-weight target on separated strata.
For a one-unit full-data law , VMW Assumption 4 holds when , , , and there exist positive constants such that , , and almost surely under , both treatment-arm probabilities satisfy for , and there is a population top- right singular basis of the stacked proxy moment for which
The generic full-data law carries the latent class , treatment, proxies, potential outcomes , and observed outcome. The observed map induces the observed-data marginal , which is the sampling law for the observable record.
Latent treatment-effect heterogeneity is described through class masses, class-specific potential-outcome means, and their differences.
For sample size , miscoverage level , concentration constant , and envelope constant , the simultaneous summary radius is This is the radius used in the event .
The mass records how much of the population belongs to latent class .
For confidence-set data with atom floor , outer radius , computed center , association radius , empirical component , and candidate endpoint value , say that holds when there exists a law and a transport plan from to such that and Thus the lower and upper endpoints of are the infimum and supremum of feasible values over the same ordered-support, atom-floor, simplex, and transport constraints.
The mean is the class-specific potential-outcome level for arm .
For a represented atomic record , define Repeated labelled locations contribute additively through the nonnegative masses .
The latent effect is the class-level treatment contrast.
The proxy structure is encoded by reference- and target-feature matrices. These matrices are the finite-dimensional objects through which the observable conditional moments factor.
For a prescribed structured-lattice estimator , let be the cardinality of its finite exhaustive candidate list. The listed candidates are indexed as and this enumeration is the duplicate-free lexicographically ordered list of all well-formed structured lattice points at the stated sample size, conditioning constants, and effect-support radius.
The matrix summarizes the reference proxy within latent classes and treatment arms.
For and any real matrix , the associated continuous linear operator from to is the ordinary finite-dimensional matrix action:
⊢ LeanThe matrix summarizes the target proxy within latent classes.
For a rectangular matrix , define the thresholded pseudoinverse The empirical compressed operator is
The radius is the common compact interval radius used for latent-class mean effect laws.
For a structured lattice point , define the structured comparison operator by For an empirical summary and threshold , define the empirical comparison operator by where is the singular-value-thresholded pseudoinverse retaining singular directions with singular value at least . In the lattice criterion, these operators are compared through
The gap measures separation among distinct positive-mass effect values and is used only for label-sensitive ordered-weight statements.
For natural numbers and and a real matrix , acts as the ordinary finite-dimensional linear map Its Euclidean operator norm is Thus comparisons of matrix-valued operators use the ordinary matrix operator norm, for example
The moments , , and are the population coordinates observed through the proxy experiment.
The analysis works with represented atomic laws while interpreting coincident atom locations as a single aggregate effect value.
For every and , consider the quotient of valid labelled at-most--atomic probability laws supported in the radius- interval by equality of their induced measures. For a quotient law in this space, is the noncomputably chosen valid labelled representative of the equivalence class .
The topology on this quotient space is the topology coinduced by the canonical quotient map from valid labelled laws to their induced-measure equivalence classes. Its measurable structure is the mapped measurable structure along the same quotient map. With this coinduced topology, the quotient space is compact.
⊢ LeanPut , , and . Let consist of the polar factors of all satisfying and . Let consist of all satisfying and . Define and let be the -fold product of the clipped -lattice in .
For , define and define the associated quotient law by
The vector is the all-ones vector, with coordinate for every latent-coordinate index .
The ordered target is used when the effects are separated enough to support label-sensitive mass recovery.
Put . For a summary , define the thresholded pseudoinverse by writing a rectangular matrix and setting . Set . For , with , invertible on , , and , define , , , and . The structured-lattice criterion at summary is For the empirical summary , write .
The model assumptions combine the proxy conditional-independence structure of Virk et al. (2026) with fixed quantitative margins. We first state the structural restrictions linking proxies, potential outcomes, treatment assignment, and the anchor coordinate.
For the structured-lattice estimator , let denote its lexicographically ordered finite candidate list. In the fixed-dimensional unit-cost exact-real model, defineThe term counts formation of the empirical summary , and counts exhaustive evaluation of the fixed-dimensional objective once for each , including the fixed-size arithmetic, comparison, singular-value, polar-decomposition, and root-isolation operations used by that candidate evaluation.
Reference-proxy separation is the VMW reference-proxy conditional independence condition (Virk et al., 2026). It makes the reference proxy vary with the latent class and treatment arm while separating it from the target proxy and outcome once those variables are fixed.
Write for the observed-record space of quadruples where , , , and .
⊢ LeanTarget-proxy separation is the VMW target-proxy conditional independence condition (Virk et al., 2026). It makes a latent-class measurement whose conditional mean is stable across treatment arms.
For represented atomic laws define by where the minimum is over all nonnegative couplings whose marginals are the represented masses:
⊢ LeanObserved-outcome consistency is the VMW causal consistency condition (Virk et al., 2026). It connects the realized outcome to the treatment-indexed potential outcome.
For a full-data probability law , let be the observed record. The observed-data marginal of is
⊢ LeanLatent ignorability is the VMW armwise latent ignorability condition (Virk et al., 2026). It places treatment assignment variation within latent classes in the role needed for the armwise proxy moments.
For each latent class , its latent-class mass is
⊢ LeanThe target-proxy anchor is the VMW anchor-coordinate normalization (Virk et al., 2026). It fixes the scale needed by the spectral anchors below.
The next conditions impose bounded contributions and uniform conditioning. The boundedness assumptions are standard VMW envelope conditions, while the positivity and rank margins are the fixed-margin structure specific to this analysis.
For treatment arm and latent class , the latent conditional potential-outcome mean is
⊢ LeanBounded target-proxy contributions follow the VMW bounded target-proxy condition (Virk et al., 2026). The radius supplies the common envelope used for summary concentration and compactness.
For each latent class , its latent-class treatment effect is
⊢ LeanThe proxy-product bound is the VMW bounded proxy cross-moment contribution condition (Virk et al., 2026). It controls the armwise moment coordinates.
For treatment arm , the reference-proxy feature matrix is defined entrywise by
⊢ LeanThe outcome-weighted proxy-product bound is the VMW bounded outcome-weighted proxy contribution condition (Virk et al., 2026). It controls the armwise moment coordinates.
The target-proxy feature matrix is defined entrywise by
⊢ LeanLatent-arm positivity is specific to this analysis. The supplied margin gives every latent class and treatment arm a fixed amount of population mass, which yields uniform arm probabilities and atom-mass control.
The derived effect-support radius is
⊢ LeanThe proxy-rank margin is specific to this analysis. The supplied singular-value margin gives the proxy feature matrices uniform conditioning, which is used to form stable compressed operators and fixed-radius effect supports.
For ordered weights, the paper restricts attention to local strata where all positive-mass latent effects are separated at a prescribed scale.
The effect gap of is with the convention when the quotient law has a single distinct support point.
⊢ LeanThe gap window is specific to this analysis. The local scale records the separation regime for ordered-weight rates.
For each , the observable armwise population moments and target-proxy population mean are
⊢ LeanDistinct latent effects are the VMW spectral separation condition (Virk et al., 2026). In this paper the condition is used for the gap-local ordered-weight target, while the quotient law below aggregates coincident effects.
Three representation levels are used for atomic laws. A raw labelled record consists of weight coordinates and location coordinates. A valid represented probability law is a raw record whose masses are nonnegative and sum to one, whose locations lie in the support interval, and whose repeated locations are interpreted by aggregating their masses. A quotient law is the equality class of valid represented laws that induce the same probability measure. The coordinate container in Definition 29 records the raw labelled representation; the , atom-floor, estimator, and confidence-set statements use the valid represented-law and quotient-law levels indicated in their hypotheses and conclusions.
We can now collect the full-data restrictions into the uniformly conditioned proxy model class.
For and radius parameter , the represented atomic-law class is the labelled coordinate space whose element consists of two real coordinate vectors, The abbreviation denotes the same represented space. The associated measure on is so coincident labelled locations contribute additively at that location. The topology on this represented space is induced by the coordinate map and its measurable structure is the pullback of the measurable structure on the two real coordinate vectors under the same map.
⊢ LeanThree distinct levels of representation are in play, and it is worth separating them once here because later definitions move between them. At the first level sits the labelled coordinate container just displayed: an element records real weights and real locations with no constraint tying them to a probability law. At the second level sit the valid labelled laws, those coordinate records whose weights are nonnegative, sum to one, and whose locations lie in the stated interval; the atom floor of Definition 31, the estimators, and the confidence sets are stated for laws of this kind. At the third level sits the quotient of Definition 12, formed from the valid labelled laws by identifying records that induce the same measure. The bridge between the levels is the induced measure displayed above, which sends coincident labelled locations to a single atom carrying their aggregate mass; compares laws through this measure, so it is insensitive to the labelling that a collision destroys. The estimand of Definition 40 is an object of the third level, which is what makes it the regular target under collisions, while the effect-ordered mass vector of Definition 30 retains second-level labelling information and is correspondingly governed by the effect gap.
The class fixes the latent cardinality , the boundedness radius, and the two conditioning margins. The qualitative proxy restrictions align with Virk et al. (2026); the fixed quantitative margins make the subsequent stability and concentration statements uniform over the class.
The observable summary is the finite-dimensional population object to which the statistical results are applied.
When the latent effects are distinct, is the vector of latent-class masses ordered by increasing effect. Otherwise,
⊢ LeanThe summary and metric put the four matrix moments and target-proxy mean in a single product space.
For a probability law , is the condition that every distinct support atom of has aggregate mass at least .
⊢ LeanThe arm count , empirical moments and , empirical mean , and empirical summary are total functions of the observed sample.
Under the full-data law ,
⊢ LeanUnder the full-data law ,
⊢ LeanThe vector implements the anchor coordinate from Assumption 5.
The compressed operator is the population spectral object inherited from the VMW moment structure.
Under the full-data law , almost surely.
⊢ LeanUnder the full-data law , for each ,
⊢ LeanThe signal basis , compressed operator , and anchors and provide the population bridge from observable summaries to latent-class mean effect moments.
Under the full-data law , almost surely.
⊢ LeanFor the boundedness radius , under the full-data law , almost surely.
⊢ LeanUnder , almost surely.
⊢ LeanThe quotient law is the primary inferential target: it records the distribution of latent-class mean effect values with the aggregate mass at each distinct value.
Under , almost surely.
⊢ LeanThe stratum is the domain for ordered-weight results, where the gap scale and distinctness conditions provide an effect ordering.
The latent class and treatment arms satisfy
⊢ LeanThe experiment contains the sample laws generated by all model laws in .
The next definition records the closure of the admissible summary image and the quotient-law functional on it.
The reference- and target-proxy feature matrices satisfy
⊢ LeanThe admissible image , closure , and functional define the population domain for the stability theorem in the next section. The moment identity is the cited spectral bridge from Virk et al. (2026); the homogeneous-effect clause records the corresponding single-atom specialization with common effect .
The smallest positive effect gap satisfies
⊢ LeanThe positive-mass latent effects take exactly distinct values.
⊢ LeanProposition 1 establishes the observable consequences of the full-data restrictions. It derives armwise positivity, the proxy moment factorization, compressed and stacked singular-value margins, observable envelopes, and the effect support radius. These conclusions place every inside the VMW moment structure under the fixed margins used here, while preserving the quantitative bounds needed for the quotient-law target.
The uniformly conditioned proxy causal model class is the class of full-data probability laws for records satisfying:
(Record space.) , , , , and .
(Parameter domain.) The dimensions and constants obey
(Model restrictions.) satisfies Assumptions 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10.
Proposition 2 supplies the compact population summary domain. Compactness supports nearest-summary selection in the next section, and the equivalence between nonempty and nonempty makes the domain statement exact for each fixed collection of supplied constants.
Several auxiliary definitions used by the empirical structured-lattice estimator and confidence reports are collected here because they share the same observable summaries, norms, and atomic-law coordinates.
The observable conditional moment summary is For summaries and , the summary metric is
⊢ LeanFor an observed sample of size , let be the number of observations in arm . Define the empirical armwise proxy moment and outcome-proxy moment using the total empty-arm convention: Let the empirical target-proxy mean be The empirical five-block observable summary is
⊢ LeanDefinition 34 records the coordinate formulas. The theorem-facing empirical-summary object below packages the same coordinates with the total empty-arm convention and Borel measurability used in sample-level statements.
For each , let Define and The empirical five-block observable summary is On the event , the corresponding armwise matrices and are the zero matrix, so is a total Borel function of the sample.
⊢ LeanLet be the first canonical basis vector.
⊢ LeanLet denote the finite ambient dimension of a Euclidean summary space . In the nearest-summary selection statement, is a compact feasible set and is the continuous summary distance used to choose nearest feasible summaries.
For matrices and with the same number of columns, define by vertical concatenation, Its row space is regarded as a subspace of .
Choose any matrix with orthonormal columns spanning the row space of the vertically stacked matrix . Define with spectral anchors
⊢ LeanFor as in Definition 32, the quotient latent-effect law is It is interpreted as a probability measure on the latent treatment effects, with coincident values of carrying their aggregate mass.
⊢ LeanHere and throughout, the latent treatment effects in are the class-average contrasts . Coincident class-average contrasts appear as a single support value with aggregate population mass.
For a local effect-gap scale satisfying , the gap-local stratum is the class of full-data laws in the uniformly conditioned proxy causal model class of Definition 32 that satisfy Assumptions 11 and 12.
⊢ LeanThe observed i.i.d. proxy experiment at sample size is
⊢ LeanFor and , the summary-closure data consists of the closed feasible-summary set together with a quotient-law functional For each admissible summary , is the quotient latent-effect law of a selected model representative satisfying . The data also records the published spectral moment identity and, in the homogeneous-effect specialization with common effect ,
⊢ LeanThe radius depends on the miscoverage level and concentration constant ; it is used for the summary event in the confidence results.
Fix and . Suppose
(Dimensions.) , , and .
(Margins.) , , and .
(Uniform model class.) The full-data law belongs to the uniformly conditioned proxy causal model class in Definition 32 with parameters .
Then the following quantitative observed-margin conclusions hold. For each treatment arm , and the observable armwise moment factors as Moreover, for every orthonormal signal basis spanning the signal space of , The stacked observed proxy moment satisfies The observable envelopes are Writing , for every latent class and treatment arm , and Finally, belongs to the published VMW parameter class represented by the supplied scope handle.
⊢ LeanLet and . Let be the class of full-data probability laws satisfying the uniformly conditioned proxy causal model restrictions with boundedness radius , latent-arm positivity margin , and proxy-rank margin . Define the admissible summary image and its closure by For summaries and , write Assume:
(Latent cardinality.) .
(Proxy dimensions.) and .
(Boundedness radius.) .
(Latent-arm margin.) .
(Proxy-rank margin.) .
Then there exists such that for every . Moreover, is compact, and
⊢ LeanTogether, the definitions and propositions in this section identify the population object, the observable summary space, and the compact domain on which the next section states the gap-free stability and estimation results.
Main results: law stability and estimation
The preceding section reduced the proxy causal model to the observable summary and the quotient law of latent-class average treatment effects. This section gives the population stability statement that makes the quotient law a regular law-level target, then turns the modulus into two estimators: a nearest-summary repair estimator defined on the compact summary closure and a finite structured-lattice estimator evaluated from the empirical summary with supplied . The constants below depend only on those fixed dimensions and conditioning constants.
Fix a simultaneous-concentration constant
⊢ LeanFor and , the simultaneous concentration radius is
⊢ LeanThe constant calibrates a uniform finite-sample concentration inequality for all five coordinates of . The radius is the corresponding -radius at confidence level , with the model envelope scaling the observable moment fluctuations.
For fixed proxy dimensions and , the ambient five-block observable-summary space is A point is written , with four matrix blocks and one target-proxy mean block, and the summary metric is The admissible image and its closure are subsets of , and the repair selector is a total Borel map from to .
The ambient space in Definition 46 lets the empirical summary be evaluated for every sample, while the feasible population summaries remain restricted to the compact closure from Proposition 2. The repair construction projects an empirical summary back to that population domain before applying the quotient-law functional.
Fix and . Let where is the admissible image in the five-block summary space. A summary-space repair datum consists of , , and maps with the following certified properties:
(Extension.) The map is continuous and, for every full-data law in the uniformly conditioned model class, sends its observable summary to the quotient latent-effect law .
(Nearest selector.) The map is Borel measurable and, whenever ,
(Fallback.) Whenever , the selector satisfies in the ambient five-block summary space for every .
(Repair measurability.) The sample-to-law map defined below is Borel measurable.
For an observed sample with empirical five-block summary , the repaired estimator is the total map
⊢ LeanThe estimator in Algorithm 1 is the repaired-summary route: it expresses the quotient-law estimator as a continuous extension evaluated at the nearest feasible summary. This directly links observable-summary uncertainty to uncertainty for the law.
The next construction replaces nearest-summary projection with a finite exhaustive lattice over the spectral factors that determine the quotient law.
The structured-lattice constant is the explicit prescribed constant given by the maximum of the threshold-stability constants and the grid-approximation constants displayed in the structured-lattice construction.
⊢ LeanFor integers , constants , and a structured-lattice estimator at radius is prescribed when, writing for the candidate count recorded in , there exist candidates and, for every sample of records in , a selected index such that:
(Well-formed exhaustive list.) Each listed candidate is well formed for the parameters . Conversely, every well-formed structured lattice point at radius appears at some listed index with
(No duplicates.) If two listed candidates agree in all four displayed coordinates, then .
(Lexicographic order.) The index order agrees exactly with the lexicographic order on structured lattice points:
(First minimizer.) For the empirical summary , the selected index minimizes the criterion computed with threshold : Ties are resolved by the first lexicographic index:
(Returned law.) On every sample , the estimator map recorded in returns the atomic effect law attached to the selected candidate:
The lattice program in Algorithm 2 searches over discretized signal bases, conditioning matrices, latent masses, and effect locations. Its criterion matches the empirical compressed operator, target-proxy mean, and anchor equation, so the returned law is tied to the same observable structure as the population quotient law.
Let , and let . Suppose that
(Dimensions.) , , and .
(Positivity calibration.) and
(Rank-margin calibration.) and .
Then there exists a simultaneous-concentration constant with such that the following holds. For every and with and , every full-data probability law on records in the uniformly conditioned proxy causal model class with radius , latent-arm positivity margin , and proxy-rank singular-value margin , and every tail probability satisfying , Moreover, for each treatment arm ,
⊢ LeanThe proof is deferred to Section D.
Lemma 1 is the sampling input for the law estimators. It gives a single high-probability radius for the empirical five-block summary and also bounds the probability of an empty treatment arm under the total empirical-summary convention.
Fix natural numbers and real constants . Assume:
(Latent dimension.) , , and .
(Bound and margins.) , , and .
Let , and let be the summary closure from Definition 43. For two summaries write There exists a constant such that, for every pair of full-data laws , Moreover, there is a map such that is continuous and Borel measurable, and This map is unique among continuous maps that obey and satisfy for every : every such satisfies for all .
⊢ LeanTheorem 1 is the population result behind the statistical theory. It establishes a Lipschitz modulus from observable summaries to quotient effect laws over the uniformly conditioned model class, and it extends the summary-to-law map continuously from the admissible image to the compact closure. The result is stated for the law that aggregates coincident effect values, so the same bound covers homogeneous, partially colliding, and separated latent-effect configurations.
The intuition is spectral but the conclusion is metric. The proxy moments determine a compressed effect operator through the VMW moment identity (Virk et al., 2026); perturbation control for finite-dimensional operators (Bauer et al., 1960; Kato, 1995) gives stable spectral moments; and the quotient law is compared in , a metric that transports aggregate probability mass between effect values. The continuous extension then makes feasible-summary projection a valid route from empirical summaries to law estimates.
Let and . Suppose that:
(Dimensions.) , , and .
(Model radii.) , , and .
(Sample size.) .
(Confidence parameters.) The miscoverage level satisfies , and the simultaneous-concentration constant satisfies .
(Empirical summary.) The empirical summary map of Definition 35 is measurable.
Then there exists repair data of the form specified in Algorithm 1 such that the estimator induced by is Borel measurable. Moreover, writing for the -closure of the admissible summary image, if , then for every sample , For every sample , the sharp summary-inversion confidence set associated with , with singleton value when and otherwise consisting of the laws with and , is nonempty.
⊢ LeanProposition 3 supplies the measurable total version of the nearest-summary construction. It also records nonemptiness of the exact repaired-image summary-inversion confidence set for every sample, preparing the confidence-set construction in the next section.
The structured-lattice estimator uses the same observable ingredients and replaces projection over the compact summary closure by a finite lattice search. The work count below fixes the computational scale certified here: a fixed-dimensional unit-cost exact-real benchmark, with arithmetic, comparison, singular-value, polar-decomposition, and fixed-degree root-isolation operations counted as unit primitives.
For the explicit structured-lattice estimator , let denote the fixed-dimensional exact-real operation count for forming , enumerating the lexicographically ordered candidate lattice, evaluating the criterion , and returning the first minimizer . In the fixed-dimensional unit-cost exact-real model, counts arithmetic, comparison, singular-value, and fixed-degree root-isolation operations and satisfies for a finite constant .
Fix integers and constants . Suppose that
(Dimension.) , , and .
(Radius and margins.) , , and .
Put and . Define Then , and there is a finite constant such that, for every , there is a total Borel summary rule whose empirical-summary estimator is the lexicographically first minimizer over the prescribed structured lattice with effect radius , atom floor , and threshold . For every sample, every distinct represented atom of has aggregate mass at least .
For every in Definition 32, the stacked observable proxy moment at , and at every summary with has exactly its first singular values at least and all remaining singular values below . The deterministic loss bound holds for every sample , where is the empirical summary in Definition 35. The candidate list and exhaustive-search operation count satisfy Finally, for every , every , and the observed product experiment ,
⊢ LeanProposition 4 establishes the empirical structured-lattice law estimator used for the algorithmic confidence set and cluster report. Its deterministic inequality converts summary error and lattice mesh width into error, while its high-probability conclusion follows by combining that inequality with the summary concentration in Lemma 1. The candidate bound records polynomial growth in for fixed in the exact-real operation model stated above.
The role of the lattice is to preserve the stabilized spectral structure at empirical summaries near the model image. The singular-value threshold keeps the compressed operator in the uniformly conditioned rank- regime; the grids over , , , and approximate the population spectral representation; and lexicographic tie-breaking turns the exhaustive search into a total Borel rule.
Fix integers and constants . Suppose that
(Dimensions.) , , and .
(Uniform bounds.) , , and .
Then there is a constant such that, for every , there exist a structured-lattice estimator taking values in , with and a nearest-summary repair estimator from Algorithm 1, such that is the first lexicographic minimizer over a finite exhaustive structured lattice of well-formed candidates for the criterion with tie-breaking by the lattice order and output . Both estimators are Borel sample maps. For every satisfying , every , and -sample, Moreover, for each treatment arm ,
⊢ LeanTheorem 2 is the law-estimation conclusion. Under the fixed dimension, envelope, positivity, and rank-margin conditions defining the uniformly conditioned class, both the nearest-summary repair estimator and the structured-lattice estimator achieve the same root- tail rate uniformly over . The statement also carries forward the armwise empty-sample bound, ensuring that the total empirical-summary convention is accompanied by an explicit probability control.
The theorem separates the regular law target from label-sensitive quantities. At the level of , coincident latent-class mean effects are represented by aggregate mass at a shared support point, and the modulus in Theorem 1 transfers root- summary concentration directly to law error. The next section uses the structured-lattice estimator as the center of exact-real finite-constraint confidence reporting and then treats ordered weights on the gap-local stratum, where effect separation determines the relevant local scale.
Confidence sets, cluster reports, and ordered weights
The root- law estimators from Theorem 2 lead directly to confidence reporting for the quotient law of latent-class average treatment effects. The first construction keeps two complementary objects in view: an exact repaired-image set of nearby feasible summaries through the repaired summary map, and a lattice-centered finitely represented exact-real outer set with an atom-floor restriction. The exact-real operation model is fixed-dimensional and unit-cost, with arithmetic, comparison, singular-value, polar-decomposition, and fixed-degree root-isolation primitives counted as the operations certified in Definition 48 and Proposition 4. This follows the econometric logic of confidence sets obtained by inverting sample restrictions (Robins et al., 2006), with the reporting metric adapted to discrete laws as in Wasserstein approaches to mixture uncertainty (Deo et al., 2023).
The inputs are fixed , constants , an observed sample with empirical summary , repair data on , and a prescribed structured-lattice estimator.
The fixed constants satisfy Define the effect-support radius and assume its nonnegativity:
Define the simultaneous concentration radius by
The repair data satisfy: is continuous on and extends the summary-to-quotient-law map on model summaries; is measurable; if , then if , then and the induced repaired estimator is measurable.
Let be the valid -atomic law returned at by the measurable prescribed structured-lattice estimator at support radius , generated by an exhaustive duplicate-free lexicographically ordered list of well-formed lattice candidates and the first lexicographic minimizer of with threshold
Define the original summary-inversion confidence set by
Define the atom-floor constant and computable radius by
Define the atom-floor subclass by
Output the computable Wasserstein confidence set
Algorithm 3 defines the exact repaired-image summary-inversion set and the lattice-centered outer set . The radius combines the concentration scale with the lattice mesh scale, while the atom floor carries the latent-arm positivity margin into the represented-law class. The lattice-centered set is finitely represented by the empirical lattice center, the transport metric, and fixed finite-dimensional constraints in the exact-real model.
The computational guarantee is an exact-real finite-representation statement. In the fixed-dimensional unit-cost model used in Definition 48 and Proposition 4, the lattice search has polynomial candidate growth and exact evaluation over the stated primitive operations. The constrained representation of gives finite atom, mass, and transport-plan constraints for membership and endpoint formulations. The exact repaired-image set is represented through on the compact repaired summary image, and Remark 1 records the polynomial-time exact-image problem for that original repaired set.
The cluster report is computed from , , , , , , and as follows.
Define the simultaneous concentration event and association radius by
Merge repeated atoms of . Join two distinct support points whenever their distance is at most , and define as the set of connected components of the resulting graph.
For each , define the associated true and candidate support sets by
Define the mass interval by In the ordered-support and transport-plan representation of , each endpoint is the value of a finite union of fixed-dimensional mixed-integer constrained programs. For each , introduce binary variables with and minimize or maximize subject to the ordered-support, atom-floor, simplex, and transport constraints defining .
Output the report For each component , define the external true-support gap by with value when .
The cluster report has a sample-facing layer and a theorem-side validation layer. The sample-facing output is formed from the observed , the supplied constants, the lattice center , and the constrained class : it merges empirical support points at threshold , expands each component by , and computes mass endpoints over the same constrained class. The true-law symbols , , , and name the concentration event, associated true support block, and external separation scale used to state coverage and width.
Theorem 3 establishes simultaneous coverage for both confidence constructions over the uniformly conditioned model class. The statement also gives samplewise nonemptiness and -diameter bounds, so the reported law-level uncertainty has root- scale under the same constants that control the estimator in Proposition 4.
The mechanism is the combination of summary concentration and the gap-free modulus. On the event that lies within of , the feasible-summary set includes a summary whose image is , and the structured-lattice center lies within the radius used to form . The transport representation in the theorem gives a finite-dimensional constrained problem over atom locations, masses, and couplings.
Law-level confidence sets can be difficult to read directly when nearby support points are statistically unresolved. The next construction converts the set into a component report: estimated support points are connected at threshold , and the mass attached to each empirical component is summarized by an interval over the same constrained confidence class.
The input is the positive estimated atomic law on .
Write uniquely in increasing support order as
Output the ordered-weight estimator
Ordering the finitely many support points and extracting their jump sizes are Borel operations, so is a total Borel map into .
⊢ LeanThe report replaces individual estimated atoms by empirical components formed at connection threshold . For each component , the support interval expands the component by , and the mass interval ranges over all laws in whose atoms can be transported to the lattice center within . The external gap is the separation scale relevant to the component’s aggregate mass.
Fix integers and constants such that There exists a simultaneous-concentration constant such that, for every boundedness radius , there is a constant for which the prescribed structured-lattice constant is positive and the following holds. For every sample size , there exist:
(Structured lattice estimator.) a measurable estimator taking values in , generated by a finite exhaustive list of well-formed structured-lattice candidates, ordered lexicographically, with equal to the effect law of the first lexicographic minimizer of the structured-lattice criterion;
(Summary repair data.) a continuous extension on the summary closure , a measurable nearest-summary selector with its empty-closure fallback, and the associated measurable repaired summary-to-law map.
For every miscoverage level satisfying let , , and be the confidence-set objects of Algorithm 3 formed from this repair data and this structured-lattice estimator, with and center . Then, on every sample, both and are nonempty, and has the finite constrained transport representation Moreover, for every full-data probability law in Definition 32, Finally, on every sample,
⊢ LeanTheorem 4 establishes the coverage and width properties of the component report on the concentration event. Every true support point is assigned to a unique empirical component, the reported support interval contains the associated true atom, and the reported mass interval contains the true aggregate component mass. When a component is separated from the remaining true support by , the mass-interval length contracts at the displayed inverse-gap scale.
The report is adaptive because transport cost controls how much mass can cross a gap. If estimated support points connected to nearby true atoms fall within the component threshold, the construction pools them and reports aggregate mass. When a component is externally separated, moving mass across that separation consumes budget, and the interval width reflects the ratio between the available radius and the external gap.
The final object in this section returns to effect-ordered latent weights. The quotient law treats coincident class-average effects as a single atom, while the ordered-weight target attaches masses to effect ranks on a gap-local stratum. The estimator below applies the repaired law construction and then orders the labelled representative when all labelled atoms are positive and distinct.
Let , , and . Suppose that
(Dimensions.) , , and .
(Latent mass.) .
(Rank margin.) .
Then there is a simultaneous-concentration constant such that, for every , if and then the prescribed structured-lattice stability constant is positive. For every , there exist a measurable structured-lattice estimator and repair data consisting of a continuous extension on and a Borel nearest-summary selector, such that is obtained from an exhaustive lexicographically ordered finite list of well-formed structured-lattice candidates by choosing the first minimizer of .
For every with , the estimator has atom floor . Moreover, for every in Definition 32 and every sample, Define with as in Definition 35, and set Then For every sample , the cluster report centered at satisfies the following properties for the representative atomic law :
(Partition.) Each support point belongs to a unique empirical component with .
(Association.) For every , the set is nonempty and is contained in .
(Support interval.) If and , then lies in the reported support interval for .
(Mass interval.) If , then the reported mass interval contains .
(Singleton width.) If and , then the reported support interval for has length at most , and
(External-gap width.) If and , then . For every , with the displayed right side read as zero when .
(Constrained representation.) The computable class has the finite constrained-program representation
Algorithm 5 defines as a simplex-valued ordered-weight rule derived from the repaired quotient-law estimator. Its target is from Definition 30, so the relevant population domain is the gap-local stratum in Definition 41.
Let and . Assume:
(Latent size.) .
(Proxy dimensions.) and .
(Envelope.) .
(Latent-arm margin.) .
(Rank margin.) .
Then there is a constant such that, for every , one can choose the simplex-valued measurable ordered-weight estimator from Algorithm 5 so that, for every gap scale satisfying and every full-data probability law ,
⊢ LeanTheorem 5 gives the ordered-weight risk bound on the separated gap-local class. The rate is clipped at one and improves as the product grows, capturing the scale at which effect ordering becomes informative for latent-stratum masses.
The contrast with the preceding law results is the statistical role of the gap. For the quotient law , Theorem 3 and Theorem 4 report uncertainty in and in aggregate component masses, with collisions represented by shared atoms. For ordered weights, Theorem 5 works on , where distinct effects and the gap window make the effect order a stable object at the scale shown in the bound.
Lower bounds and sharp rates
Rate matching is proved on the explicit uniformly conditioned two-class specialization , , , and . The construction uses two related two-class submodels. The first moves a collision witness away from a single class-average effect value at scale; the second changes ordered latent masses along a factorization-preserving path whose observed separation is proportional to the product of the effect gap and the weight displacement. The lower bounds combine the displayed local Kullback–Leibler neighborhoods with the standard two-point testing logic of Cam (1986) and Polyanskiy et al. (2019).
Fix a local Kullback–Leibler radius constant satisfying
⊢ LeanThe constant fixes the radius of the local experiments used below. Keeping this radius bounded away from one gives a common testing scale for the two-point comparisons.
For and , is the full-data law on with , , , , , latent masses and , and latent-arm weights Thus . Conditional on , the variables are Bernoulli with the displayed treatment probability, with , , , , and . Conditional on , is Bernoulli and independent of , with With the induced feature matrices and and latent-arm weight vector , the path preserves the armwise proxy factorization for , and .
For a gap and displacement , is the explicit two-class factorization-preserving path law obtained by perturbing the latent weights while preserving the stated proxy factorizations.
⊢ LeanThe divergence is used only for the observed one-record margins that generate the sample laws . Thus the lower bounds compare statistical experiments through the same observed records as the estimators and confidence reports.
The local neighborhoods center the two testing problems at the relevant collision or separated path endpoint.
The published VMW scope handle is the cited-publication interface for the proxy-model parameter class and recovery regime. For fixed latent cardinality , target-proxy dimension , reference-proxy dimension , and full-data law on , it records three law-level predicates: These predicates are publication-level gates whose concrete interpretations enter through separate cited gates. The handle also includes the published estimator which maps an observed sample of size from records to a triple consisting of treatment effects, an anchor-normalized target-proxy feature matrix, and real-valued weight coordinates.
⊢ LeanThe quotient-law lower bound applies the Le Cam local testing framework of Cam (1986) and Polyanskiy et al. (2019) to Assumption 13. It keeps the observed experiment within divergence of the collision witness, so an -sample comparison has bounded information distance.
For probability laws and on the same measurable space, define when , and set when . In the local experiments this divergence is applied to one-record observed-data margins such as and .
The ordered-weight lower bound applies the Le Cam local testing framework of Cam (1986) and Polyanskiy et al. (2019) to Assumption 14. The center has effect separation indexed by , matching the gap-local target used in Definition 41. The next display repeats the same divergence notation for the local-experiment statements that follow.
The witness laws are finite Bernoulli proxy models with explicit masses, treatment probabilities, proxy features, and potential-outcome means.
For probability laws and on the same measurable space, the Kullback–Leibler divergence from to is when , and when . In the local experiments this divergence is applied to one-record observed-data margins, including and .
The Kullback–Leibler divergence from , the observed-data marginal, to the observed margin of , the two-class collision witness at perturbation zero, satisfies
⊢ LeanProposition 5 places the collision witness inside the same uniformly conditioned class used for the upper bounds. The determinants and arm probabilities show that the testing pair is an interior proxy model under the displayed margins, while the quotient law moves from a single atom at to two nearby atoms for positive .
For labeled weights, the path changes latent masses while preserving the proxy factorizations that determine the observable moments.
The Kullback–Leibler divergence from , the observed-data marginal, to the observed margin of the two-class witness at perturbation satisfies
⊢ LeanFor , the two-class witness law is the law specified by Conditional on , the variables , , , and are mutually independent Bernoulli variables with the displayed conditional means. Conditional on and , the variable is independent of and has conditional mean given by . The observed outcome is
⊢ LeanThe family supplies the local perturbation for the ordered-weight converse. At it coincides with the separated witness at displacement , and changing changes the ordered masses while retaining the displayed armwise proxy factorizations.
The two-class witness also gives a compact picture of the reporting geometry. In this specialization the class masses are and , and the class-average effects are and . With association radius , the empirical graph in Algorithm 4 connects estimated atoms at threshold ; the support interval then expands each empirical component by .
The table serves as a reading guide for the existing witness. It illustrates how quotient aggregation and component reporting use the same masses while the merge threshold determines whether the report presents one aggregate interval or two separated component intervals.
The local quotient-law KL experiment is where is the model class in Definition 32.
⊢ LeanThe local labeled-weight KL experiment is where is the gap-local stratum in Definition 41.
⊢ LeanThe two local experiments distinguish the metrics under study. is centered at a collision and measures the law-level difficulty, whereas is centered on the gap-local stratum and measures the difficulty of recovering effect-ordered masses.
For every collision-witness effect displacement satisfying the explicit two-class Bernoulli proxy witness law belongs to with , , and . Moreover, and the quotient latent-effect law is At , the proxy feature and observable moment matrices above are nonsingular, the quotient latent-effect law is , and the two latent-class masses are distinct.
⊢ LeanIt is worth tracing the collision through the witness once, since every object in the paper is visible on it. The latent effects of are read off the displayed conditional means as and , carried by the class masses and , so the effect gap is and the quotient law is
| regime | latent effects | gap | quotient law | effect-ordered masses |
|---|---|---|---|---|
| one atom at of mass | labels carry no separation | |||
| at , at |
Transporting the mass and the mass each a distance onto the common location gives , so the estimand moves continuously into the collision and the modulus of Theorem 1 applies across the whole family at fixed conditioning constants. The two regimes are distinguished by which report they support. The quotient law and its confidence set of Algorithm 3 remain meaningful at , where the two classes have become one atom of mass . The effect-ordered masses are separated only while is positive, which is why Theorem 5 is stated on the gap-local strata and why its converse in Theorem 6 is driven by displacement along this same family. The cluster report of Algorithm 4 interpolates between the two pictures: while the estimated support points fall within the connection threshold they are returned as one component whose aggregate mass interval covers the combined mass, and once the separation exceeds that threshold they are reported as two components whose mass intervals sharpen with the external separation, in the manner quantified by Theorem 4.
Theorem 6 gives the two local converse rates on the displayed two-class specialization. For quotient laws, an displacement in the collision witness creates separation of order while keeping the observed Kullback–Leibler comparison at order . For ordered weights, the factorization-preserving path makes the observed divergence scale as , so the largest indistinguishable mass displacement is clipped at .
These local statements calibrate the upper bounds from the preceding sections on the same explicit two-class specialization. The quotient-law lower bound matches the root- law rate from Theorem 2; the ordered-weight lower bound matches the inverse-gap upper bound in Theorem 5.
Fix . There exist constants such that and, for every integer , the following statements hold for the explicit submodel , , , and .
(Quotient-law risk.) For every measurable quotient-law estimator with support radius there is a law in the local quotient-law experiment of Definition 56 such that with as in Definition 55, and
(Ordered-weight risk.) For every and every measurable ordered-weight estimator taking values in the simplex, there is a law in the local labeled-weight experiment of Definition 57 such that and
(Quotient-law certificate.) The laws and belong to the corresponding local quotient-law experiments, and their observed-data marginals and quotient laws satisfy
(Labeled-path certificate.) For every , with the displacement lies in the tangent-amplitude domain, the laws and belong to the corresponding local labeled-weight experiments, and
Proposition 6 converts the local quotient-law comparison into a same-class minimax statement on the displayed uniformly conditioned two-class model. The proposition states that the optimal uniform risk for the quotient law is of exact root- order on this specialization.
For the specialization , , , and , let be the uniformly conditioned proxy causal model class in Definition 32. The derived effect-support radius is There exist real constants and such that and, for every sample size , the following two statements hold:
(Lower bound.) For every measurable estimator mapping an -sample of observed records to a represented atomic law modulo the radius , there exists a full-data probability law such that
(Upper bound.) There exists a measurable estimator mapping an -sample of observed records to a represented atomic law modulo the radius such that, for every full-data probability law ,
Equivalently, on this same specialized class, for every , where the infimum ranges over measurable estimators with output radius .
⊢ LeanProposition 7 gives the corresponding same-class minimax statement for ordered weights on the same two-class specialization. The rate is governed by the product : when the effect gap is large relative to sampling noise, ordered masses are estimable at the inverse-gap scale, and the bound remains clipped by the diameter of the simplex-valued loss.
The final comparison records how the same witnesses transfer to classes formulated in the terminology of Virk et al. (2026). The following definitions separate the published structural gates, finite-regularity condition, and separated recovery regime used for that transfer.
There exist real constants and such that and, for every sample size and every gap scale with , the gap stratum from Definition 41 satisfies Here the infimum is over all measurable simplex-valued labeled weight estimators based on observed records in the explicit submodel , , , and , and is the effect-ordered latent-mass vector with the barycentric collision fallback.
⊢ LeanFor and , define
⊢ LeanA law satisfies the published VMW finite-regularity condition when , , , and there exist constants such that , , and almost surely, for , and there is a top- right singular basis for the stacked proxy moment satisfying and for each .
Fix with , , and . A one-unit full-data probability law belongs to the published VMW structural model when it satisfies the published qualitative proxy model conditions: reference-proxy separation , target-proxy separation , consistency almost surely, armwise latent ignorability for , full column rank of , , and , and strict latent-arm positivity for every and .
Definition 52, Definition 59, Definition 60, and Definition 61 provide the interface for comparing the present lower bounds with the published VMW recovery framework. The cited model conditions and separated recovery regime are used as external publication-level gates, while the converse below is stated in terms of exact containment of the two displayed witness pairs.
A full-data law lies in the separated recovery regime when it satisfies:
(Structural model.) The published VMW structural model conditions.
(Anchor normalization.) The first target-proxy coordinate satisfies almost surely.
(Finite regularity.) The published VMW finite-regularity condition.
(Spectral separation.) The latent-class treatment effects satisfy whenever .
In this regime, the separated-recovery estimator analyzed by Virk et al. (2026) is evaluated under its stated sample-size and radius conditions and returns simultaneous high-probability bounds for the treatment effects, the anchor-normalized feature matrix, and the simplex-projected mixture weights.
The calibrated displacement in Definition 58 is the path amplitude used in the labeled-weight two-point comparison. It is the same clipped scale that appears in the ordered-weight upper and lower bounds.
Fix , , and . There exist constants such that For every , the following hold.
(Quotient witness pair.) The two laws and from Definition 55 belong to the published VMW model and satisfy VMW Assumption 4, while lies outside the published separated recovery regime.
(Quotient lower bound.) For every class of full-data probability laws such that every belongs to the published VMW model and satisfies VMW Assumption 4, if contains laws equal to and , then every estimator of the quotient effect law with support radius has worst-case risk at least :
(Separated labeled path.) For every in the gap-scale domain, define Then is in the tangent-amplitude domain, and the endpoint laws and both belong to the published VMW model, satisfy VMW Assumption 4, and have spectrally separated latent effects.
(Labeled-weight lower bound.) For every in the gap-scale domain and every class of full-data probability laws such that every belongs to the published VMW model and satisfies VMW Assumption 4, if contains laws equal to and , then every estimator of the ordered latent weights satisfies
Theorem 7 states the converse comparison with Virk et al. (2026). Any comparator class satisfying the cited VMW model and finite-regularity gates inherits the displayed lower bound once it contains the two quotient witness laws or the two labeled-path endpoint laws. The transfer is therefore tied to concrete witness containment: the quotient comparison uses the collision pair, while the labeled-weight comparison uses separated endpoint laws at the calibrated displacement.
Discussion, extensions, and open questions
The results above identify the quotient law as the regular object for proxy-based latent-class mean treatment-effect heterogeneity. Theorem 1 characterizes a Lipschitz map from observable moment summaries to the aggregated latent-class mean effect law, and Theorem 2 transfers this population regularity to collision-uniform root- estimation in the observed i.i.d. experiment. The estimand in Definition 40 is therefore the natural law-level target when the empirical question concerns population mass at class-average effect values, with class labels handled through the separated ordered-weight results.
The confidence constructions in Algorithm 3 give two complementary ways to report this target. The summary-inversion set preserves the repaired feasible-image construction induced by on , while the algorithmic set is centered at the structured-lattice estimator and has an explicit finite transport-plan representation in the fixed-dimensional exact-real model. Theorem 3 establishes simultaneous coverage for both sets and root- diameter bounds in , so the two reports express the same law-level concentration through different computational representations.
The cluster report in Algorithm 4 translates this law-level uncertainty into component-level statements at the resolution delivered by the confidence radius. Estimated atoms are connected at threshold , while support association and interval expansion use ; the report then pools connected empirical components and attaches both a support interval and a compatible mass interval. Theorem 4 establishes that, on the simultaneous concentration event, these components partition the true support, contain the associated true atoms, and cover the corresponding aggregate masses. The interval-length bounds show how the report becomes sharper when a component is externally separated from the remaining support, with exact mass recovery for a component whose associated block exhausts the true support.
This reporting layer also clarifies the role of the ordered-weight results. Theorem 5, Theorem 6, and Proposition 7 characterize effect-ordered mass recovery on separated strata, where the positive gap scale converts law-level transport error into label-sensitive information. The quotient-law statements remain organized around on aggregated laws, while the ordered-weight statements quantify the additional inverse-gap cost of resolving labels by effect order.
Open questions
The computational distinction between the two confidence constructions leaves a sharp finite-dimensional problem. The structured-lattice estimator and the outer set have explicit constrained representations, and the original summary-inversion set retains the exact repaired-image target. The following remark records the corresponding exact-image question.
It remains open whether, in the fixed-dimensional exact-real model, one can compute exactly and in time polynomial in the sharp summary-inversion image together with its exact support-dependent cluster extrema, using neither compact nearest-summary optimization nor black-box evaluation of . The positive real law estimator , the honest outer set , and every interval in the report have explicit fixed-dimensional constrained representations, including cases in which the raw empirical operator has nonreal eigenpairs. The question concerns a sharp polynomial-time algorithm for the original summary-inversion image and its extrema.
⊢ LeanThe question in Remark 1 concerns exact computation of the sharp repaired-image set and its support-dependent extrema. A positive resolution would align the original summary-inversion report with the same kind of explicit fixed-dimensional representation already available for the lattice-centered outer report.
A second extension is adaptation to latent cardinality and unknown conditioning margins. The present model class in Definition 32 fixes , and the target space, structured lattice, confidence sets, and ordered-weight stratum are all calibrated to that fixed finite value together with supplied . An adaptive formulation would need a data-dependent dimension choice or margin choice together with a target convention for comparing quotient laws across candidate values of , while preserving the aggregation principle that makes stable at collisions.
Margin misspecification belongs to the same adaptation problem. The stated guarantees attach to the supplied model class: conservative supplied constants that still satisfy the defining inequalities yield valid conclusions after recomputing the displayed constants, whereas supplied constants violated by the data-generating law place that law outside the certified class. A diagnostic theory for selecting or stress-testing is a natural extension of the fixed-margin results.
The exact-real finite representations also leave ordinary computational questions for implementation work. The operation counts in Definition 48, Proposition 4, and Theorem 8 use fixed-dimensional unit-cost primitives, and the endpoint descriptions in Algorithm 4 are finite constrained programs over masses, support locations, and transport plans. Bit complexity for these endpoint programs, numerical conditioning of the singular-value and root-isolation primitives, and practical solvers for the reported intervals are open implementation questions beyond the exact-real certification.
A third direction concerns application-specific proxy mappings. Assumptions 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 give a finite proxy-causal environment in which the quotient-law and reporting results apply. In applied settings, the substantive task is to justify how observed proxy measurements instantiate these conditional-independence, normalization, boundedness, overlap, and rank conditions, following the model-building role that proxy variables play in the broader causal literature (Virk et al., 2026; van Amsterdam et al., 2022).
Appendices
Probability, selection, and concentration
This appendix collects the auxiliary statements used by the population, estimation, confidence, and lower-bound arguments. The first group supplies measurable selection and finite-sample probability control for the observable summary space introduced in Definition 33. The compact-selector statement is the measurable nearest-point device used by the repair map in Algorithm 1; it is the form of measurable extrema needed for a compact feasible summary set, in the spirit of Brown et al. (1973).
For every , every set , and every loss , suppose that
(Nonempty feasible set.) is nonempty.
(Compact feasible set.) is compact.
(Continuous loss.) The map is continuous.
Then there exists a Borel measurable map such that, for every ,
⊢ LeanIdentify with its standard finite-coordinate Euclidean realization. Define By the continuity hypothesis, is Borel measurable, and for each fixed the section is continuous on .
The compact measurable argmin selection theorem for a nonempty compact Euclidean action set gives a map such that is Borel measurable and, for every ,
Substituting the definition of gives which is the asserted selector property.
∎The selection lemma turns compactness of into a Borel rule, so the nearest-summary construction remains a sample-measurable estimator after applying the empirical summary map. The remaining probability statements in this subsection record the arm-mass and concentration primitives that support the total empty-arm convention.
Let be a full-data probability law in the setting of Definition 32. Suppose satisfies Assumption 9 with margin . Then, for each treatment arm ,
⊢ LeanFix an arm . The arm event is the disjoint union of its latent-arm cells: Therefore By Assumption 9, each summand is at least . Hence
∎The lower bound converts the latent-arm margin into an observed-arm margin. It is the elementary probability input behind the empty-arm probability bound in Lemma 1 and behind the observed VMW margin statement in Proposition 1.
The next statements collect the moment factorizations and envelope calculations that connect the full-data restrictions to the observable five-block summary.
Let be a full-data probability law carrying , with fixed target- and reference-proxy dimensions , in the uniformly conditioned proxy causal model class of Definition 32. Suppose that:
(Latent cardinality.) The number of latent classes satisfies .
(Target dimension.) The target-proxy dimension satisfies .
(Boundedness radius.) The uniform boundedness radius satisfies , with the bounded-target-proxy condition as in Assumption 6.
(Latent-arm margin.) The latent-arm positivity margin satisfies , with latent-arm positivity as in Assumption 9.
(Proxy separation.) The reference- and target-proxy separation conditions hold as in Assumptions 1 and 2.
Then, for each treatment arm , the observable armwise proxy moment from Definition 28 factors as where is the reference-proxy feature matrix of Definition 24, is the target-proxy feature matrix of Definition 25, and
⊢ LeanFix . Write For every latent class , latent-arm positivity gives Also, Lemma 3 gives
We prove the matrix identity entrywise. Fix coordinates and . By Lemma 5, Thus is integrable on each . The observed arm pulls back to the full-data arm, so
Since is the disjoint union of the cells , normalized additivity over this finite partition yields
On a fixed positive cell , put The coordinate bounds imply that and almost surely on the normalized restriction to . Applying reference-proxy separation to the bounded tests and , and then removing the clamps, gives
It remains to identify the target-proxy mean in the cell with the latent-class mean. Let Target-proxy separation on the positive class gives The two factors involving are ordinary normalized cell masses: and Because this ratio is positive, cancellation gives Using the same coordinate envelope to remove the clamps on the cell and on the class,
Combining the preceding displays and using Definitions 24 and 25, By the definition of , the right-hand side is exactly Since and were arbitrary, The arm was arbitrary, so the factorization holds for both treatment arms.
∎Let be a full-data probability law carrying , with as in Definition 32. Suppose that
(Latent cardinality.) .
(Target-proxy dimension.) .
(Anchor normalization.) satisfies Assumption 5.
Then the model envelope bounds the target proxy, the reference proxy, and the outcome–reference-proxy products -almost surely: and
⊢ LeanSince , the first target-proxy coordinate is available; write it as coordinate . We use the following coordinate bound: for any real matrix , Indeed, with the th Euclidean basis vector,
On the probability-one event from Assumption 6, Hence, for every target-proxy coordinate ,
On the probability-one event where Assumptions 5 and 7 both hold, For every reference-proxy coordinate , the entry of is , so the coordinate bound gives
On the probability-one event where Assumptions 5 and 8 both hold, For every reference-proxy coordinate , the entry of is . Therefore
The three displayed coordinatewise bounds hold on probability-one events. Since the coordinate sets are finite, these are exactly and
∎Let be a full-data probability law carrying , as in Definition 32. Suppose that
(Latent cardinality.) .
(Proxy dimensions.) and .
(Model constants.) , , and .
(Model components.) The model class includes the bounded target-proxy condition of Assumption 6, the latent-arm positivity condition of Assumption 9, and the proxy-rank margin condition of Assumption 10.
Then, for every latent class and every treatment state , the latent conditional potential-outcome mean of Definition 22 satisfies where is the derived effect-support radius from Definition 26.
⊢ LeanFix and . Define
The condition in Assumption 9, together with , gives In particular . Since we have Let .
We first bound the observed outcome on the latent treatment cell. Put Since gives , and since , we have . The rank condition in Assumption 10 gives Applying the least-singular-value lower bound to the th canonical vector gives Hence there exists such that Let Then Indeed, if this probability were zero, then contradicting the displayed coordinate lower bound.
Under , Assumption 1 makes independent of . Also, Lemma 5 supplies Define On , and therefore Independence gives Since the first factor is positive, Thus where the last equality uses Definition 26.
Now work under , and define The preceding positivity gives By Assumption 4, is independent of , hence independent of , under . On , Assumption 3 gives almost surely, and the cell bound from the previous step gives With we therefore have Independence gives and the positive arm probability yields Consequently
Finally, by Definition 22, The almost-sure bound just proved gives integrability and This is the desired bound.
Let , let , let be a full-data probability law carrying , and fix a latent class . Suppose that
(Dimensions.) , , and .
(Margins.) , , and .
(Model membership.) for the parameters , as in Definition 32, with the bounded-proxy, latent-arm-positivity, and proxy-rank-margin components recorded in Assumptions 6, 9, and 10.
Then the latent-class treatment effect , formed from the latent means in Definitions 22 and 23, satisfies where is the derived effect-support radius in Definition 26.
⊢ LeanFix . By Definition 23, Thus the triangle inequality gives Applying Lemma 6 to the two treatment states, Therefore
∎For every collision-witness effect displacement with , let be the two-class witness law of Definition 55. The latent-class masses from Definition 21 are Equivalently, these masses are constant in over this range and are distinct.
⊢ LeanFix with . For a Bernoulli success parameter and , put For the two witness classes in the paper labels, define and Thus Definition 55 assigns to the atom with the elementary mass The bounds on place every Bernoulli parameter above in , so these elementary masses are nonnegative. By Definition 21 and the finite witness-law expansion, for ,
The indicator restricts the finite sum to the slice, hence
For fixed and , summing over the nuisance Bernoulli coordinates gives because .
Therefore Evaluating this identity at the two paper-labeled classes yields
The displayed values are independent of , and they are distinct since
∎Together these factorization and envelope lemmas give the deterministic bridge from the conditional-independence restrictions to the matrices used in the compressed operator. The latent mean and effect envelopes place every quotient law in the common support interval fixed by Definition 26.
The following algebraic facts are used to pass from factored proxy moments to the compressed spectral representation.
For every collision-witness effect displacement with and every latent class , the latent-class treatment effect from Definition 23 under the two-class Bernoulli witness law from Definition 55 is
⊢ LeanFix . For a Boolean variable , write For latent class , put The range of puts all Bernoulli parameters used below in , so the finite weights of are nonnegative.
The elementary mass assigned to the point indexed by is where , , and is the displayed -mean from Definition 55. At this point,
For either , the latent conditional mean from Definition 22 is the finite conditional average
The denominator is because each Bernoulli factor sums to one over its binary coordinate.
For , so
For , and hence
Using the definition of the latent-class treatment effect in Definition 23,
Substituting the two values of gives
∎For any collision-witness effect displacement and any treatment arm , suppose that
(Lower bound.) .
(Upper bound.) .
For the two-class Bernoulli proxy witness law from Definition 55, its observed-data marginal from Definition 20 satisfies
⊢ LeanFix with and fix .
Define the observed arm set For a full-data record , let . Then The coordinate map is measurable and is measurable, so is measurable. By Definition 20,
For and , put The finite witness law in Definition 55 assigns to the atom indexed by the real weight where and The bounds on give and all the other displayed Bernoulli parameters also lie in . Hence every is nonnegative. Therefore the real mass of the full-data arm event is the ordinary finite sum
Selecting the treatment coordinate in the finite sum gives This is the two-case calculation and : exactly one value of has indicator one.
For each fixed and , Since for every , the four nuisance factors sum to one, and hence
Combining the preceding identities, If , then If , then Thus
For every collision-witness effect displacement satisfying , let be the two-class Bernoulli proxy witness law from Definition 55, and let be its target-proxy feature matrix from Definition 25. Then
⊢ LeanFix with . By Definitions 25 and 55, the entries of the target-proxy feature matrix are Since , with conditional means and , this gives, in the statement’s zero-based coordinate order, Therefore
∎For the two-class Bernoulli witness law and the reference-proxy feature matrix in Definitions 55 and 24, suppose
(Perturbation range.) The collision-witness effect displacement satisfies .
(Arm.) The treatment arm is binary, .
Then
⊢ LeanBy Definitions 55 and 24, the entry is the conditional mean of the th coordinate of in latent class and arm . The first coordinate of is identically one. Conditional on , the variable is Bernoulli with the displayed conditional means, while the remaining Bernoulli coordinates are nuisance factors whose conditional masses sum to one. Thus Equivalently, The assumptions ensure that these conditioning cells have positive mass, so the displayed conditional means are the ordinary finite conditional averages.
For a matrix , write There are two arm cases. If , then If , then Since the arm assumption is exactly , these two computations give .
∎Let be the factorization-preserving two-class path law of Definition 51. Suppose that
(Gap range.) .
(Tangent amplitude.) .
Then satisfies the proxy-rank margin in Assumption 10 with proxy-rank singular-value margin : the reference-proxy feature matrices and from Definition 24, and the target-proxy feature matrix from Definition 25, all formed from , have smallest singular value at least .
⊢ LeanDefinitions 50, 24, and 25 give the feature matrices of . Under , the denominators below are positive, and while Thus the matrices in Assumption 10, formed from , are , respectively.
The undisplaced reference features have singular-value slack . At , For every , and The variational characterization of the least singular value therefore gives
Define the reference-feature entry displacements Then Since , and the algebraic identities give For any real , because its action sends to . Weyl’s singular-value perturbation inequality therefore yields, for ,
The target feature is handled in the same perturbative form, starting from For every , so With we have The bound gives , hence Also for any real , since its action sends to . Another application of Weyl’s inequality gives
Combining the three estimates, This is the proxy-rank margin of Assumption 10 with .
These linear-algebra statements justify the rank thresholding and anchored diagonalization used by the lattice estimator. The condition-number bound provides an explicit finite-dimensional constant for perturbing between feature coordinates and the compressed spectral coordinates.
The Wasserstein comparisons use a one-dimensional dual potential and the coordinate relation between the Euclidean summary metric and .
Let . Suppose that
(Gap range.) .
(Path displacement.) .
Then the path law of Definition 51 belongs to the uniformly conditioned model class of Definition 32 with Equivalently, satisfies the core numerical domain for , together with Assumptions 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10, using latent-arm positivity margin and proxy-rank margin .
⊢ LeanWrite the two latent classes as , and write for the binary coordinates in the finite support of the path. Put and For a Bernoulli coordinate and success probability , set The path assigns the atom mass where and Since , A direct check of the two latent classes gives and follows from . Thus every displayed atom mass is nonnegative. Summing first over the Bernoulli success and failure values gives Hence is a probability law on the full-data record space with .
The numerical core domain in Definition 32 is
For Assumption 1, fix a latent-arm cell . For bounded measurable test functions of and of , the finite conditional sums over the atom weights factor as Indeed, conditional on , the factor depending on is , while the factors determining and are functions of and the fixed arm . The normalizing denominator is the same finite latent-arm mass in all three conditional means, so expanding the four binary nuisance sums gives the displayed product identity.
For Assumption 2, fix . For bounded measurable test functions of and of , After conditioning on , summing over uses for each arm . The remaining finite sum separates the -factor from the factors depending on , and hence on .
For Assumption 4, fix a treatment state and a latent class . For bounded measurable test functions of the potential outcome and of , This is the same finite-product calculation: the Bernoulli factor for depends on and , while the Bernoulli factor for is .
Every support point of records the observed outcome as , which gives Assumption 3. Every support point also has , which gives and hence Assumption 5.
The boundedness conditions hold supportwise. Since and with , For every vector , so . Since , Thus Assumptions 6, 7, and 8 hold with .
The latent-arm cell masses are precisely the four displayed values . From , and Therefore which is Assumption 9 with .
By Lemma 13, for the reference- and target-proxy feature matrices formed from . Hence Assumption 10 holds with .
Combining the core domain with the model restrictions just verified gives for , , , and , as required by Definition 32.
For real numbers and , suppose that
(Gap scale.) The local effect-gap scale satisfies .
(Path displacement.) The path displacement satisfies .
Then the path law of Definition 51 belongs to the gap-local stratum of Definition 41, with , , and . Equivalently, as specified by Definition 32, Assumption 6, Assumption 9, and Assumption 10, its smallest positive effect gap from Definition 27 satisfies the gap window of Assumption 11, and its latent effects from Definition 23 are distinct as in Assumption 12.
⊢ LeanWrite . The hypothesis gives . Also gives Thus the detailed path construction in Definition 50 applies throughout the tangent-amplitude domain and yields the full-data probability law .
Under the same bounds and , Lemma 14 gives for , , , and . This supplies the model restrictions in Definition 32, including Assumptions 6, 9, and 10.
It remains to check the gap-local conditions. By the explicit construction in Definition 50, the latent-class masses are The displacement bound gives These are the positive-mass classes entering Definition 27 and Assumption 12.
Again by Definition 50, the latent conditional potential-outcome means are Using Definition 23, the corresponding latent effects are
Since , these two effects are distinct, and their distance is The same value is obtained with the two classes interchanged, while equal-index pairs are excluded by the condition in Definition 27. Hence and therefore
Consequently, because and . Also, so the positive-mass latent effects take exactly distinct values. This is Assumption 12.
Combining , the gap domain , the displayed gap window, and distinct effects gives with , , and , as required by Definition 41.
∎Let , let be the local Kullback–Leibler radius constant, and let . Assume:
(Local radius.) , as in Definition 49.
(Gap scale.) The local effect-gap scale is in the gap window of Assumption 11, with .
(Path displacement.) The path displacement satisfies .
(Observed KL calibration.)
Then the path law from Definition 51 belongs to the local labeled-weight experiment of Definition 57 with , , and . Equivalently, as in Definition 41, and its observed-data marginal from Definition 20 obeys the neighborhood condition of Assumption 14:
⊢ LeanThe bound gives , hence also . On this range the latent masses belong to . The treatment Bernoulli probabilities are obtained by dividing by the corresponding latent masses, and the same bounds on make each numerator nonnegative and no larger than its denominator. The target-proxy probabilities also lie in . For the reference-proxy probabilities, gives positive denominators and while the remaining displayed values and already lie in . Finally and are Bernoulli probabilities because . Thus all finite product weights defining in Definition 50 are nonnegative, and summing those product weights over leaves the two latent masses, whose sum is one. Hence is a probability law. The same argument applies at , and Definition 50 gives
By Lemma 15, the hypotheses and imply with , , and . The local-radius hypothesis supplies
It remains to verify the local observed-neighborhood inequality. Let For , define Let be the atomic law on with these singleton masses. The map pushes to the observed margin: Therefore data processing gives
At , direct expansion of the sixteen visible cells gives Since , every entry in these two tables is at least . Hence and .
For each visible cell, the same finite expansion gives the exact displacement where All four coefficients are at most one, and , so for every visible cell.
For finite probability laws with positive base mass on every cell, using for the first inequality. Applying this with and , the preceding floor and displacement bound yield, cell by cell, There are visible cells, and therefore
Combining data processing with the last display and the assumed calibration gives Using , this is exactly the neighborhood condition in Assumption 14. Together with , it proves as defined in Definition 57.
∎The potential lemma is the dual device behind the modulus for represented atomic laws. The two summary-metric lemmas place the finite-dimensional summary topology and the selection loss on a common continuous footing.
The next group records the witness and path bookkeeping used in the lower-bound section. The statements give explicit masses, effects, observed arm probabilities, determinant margins, model membership, and local-experiment membership for the two-class laws.
For a labelled representative , define as follows. If every labelled weight is positive and the labelled atoms are distinct, then its th coordinate is the mass of the atom whose support rank is : If some labelled weight is zero or two labelled atoms coincide, set .
Fix and a real support radius. Let and be valid labelled atomic representatives with slots in the corresponding compact interval, as in Definition 29. If and represent the same aggregated probability measure, then their ordered-mass vectors coincide:
⊢ LeanWrite the two labelled representatives as Validity means with all labelled atoms lying in the prescribed compact interval. For any labelled representative , put If , there are no vector coordinates to check.
The equality of represented probability measures gives equality of aggregate mass at every real point: Indeed, by Definitions 29 and 6, the represented measure assigns the aggregate of the positive parts of the labelled weights at ; validity gives and . Evaluating the common represented measure on gives the displayed identity.
Consequently, If and , then the aggregate mass of at is positive. By [pf:ordered-masses:aggregate], the aggregate mass of at is positive, and nonnegativity of the labelled weights yields some with This proves one inclusion, and the reverse inclusion follows by interchanging and .
For a representative , define the rank of a labelled slot by The counted set excludes , so Moreover, if , then because the first counted set is contained in the second and belongs to the second but not the first. Hence, whenever is injective, the map is injective, and therefore bijective on the finite set .
The following full-support criterion holds: For the forward implication, is the image of under , so If the left side equals , both inequalities are equalities; thus every labelled weight is positive, and equality in the image-cardinality bound gives injectivity of the labelled locations. The reverse implication follows because all labelled slots are then positive and distinct.
Suppose first that By [pf:ordered-masses:support], Then [pf:ordered-masses:full-support] gives and both labelled-location maps are injective. For either such full-support representative , Definition 62 gives where the last equality uses the rank injectivity from [pf:ordered-masses:rank].
Fix a coordinate . Since the rank map for is bijective by [pf:ordered-masses:rank], choose such that The support equality in [pf:ordered-masses:support] gives with Using [pf:ordered-masses:aggregate] at , and using injectivity of both labelled-location maps to collapse the two sums, yields
For each , [pf:ordered-masses:support] supplies a slot such that The map is injective: if , then , and injectivity of the locations of gives . Thus is bijective. Since also , for every , The bijection therefore identifies the two sets counted by the ranks, so Combining this rank equality with [pf:ordered-masses:apply-rank,pf:ordered-masses:weight-match] gives Since was arbitrary, the ordered-mass vectors are equal in the full-support case.
It remains to consider By [pf:ordered-masses:support], Using [pf:ordered-masses:full-support], the positivity-and-injectivity condition in Definition 62 fails for both representatives. Hence both ordered-mass vectors take the barycentric fallback value: Thus the ordered-mass vectors are equal in the complementary case as well.
The two cases prove as required.
∎Let , a full-data probability law carrying , belong to the model class in Definition 32. Suppose that
(Latent classes.) .
(Positivity margin.) The latent-arm positivity margin satisfies .
(Model conditions.) The reference- and target-proxy separation, consistency, latent ignorability, latent-arm positivity, bounded outcome-proxy product, and bounded target-proxy conditions in Assumptions 1, 2, 3, 4, 9, 8, and 6 hold.
Then, for each treatment arm , Here is the observable armwise population moment in Definition 28, is the reference-proxy feature matrix in Definition 24, is the target-proxy feature matrix in Definition 25, and is the latent conditional potential-outcome mean in Definition 22.
⊢ LeanStep 1. Fix . For , put For a full-data event and a real-valued measurable function , write with the total convention . The latent-arm positivity assumption gives, for every , Also, Lemma 3 gives and since and , Thus all normalizations used below have positive denominators on each latent class, each latent-treatment cell, and the treatment arm.
Step 2. Let be the normalized restriction of to , equivalently the probability law with density with respect to . Latent ignorability in Assumption 4 gives independence of and under . Therefore Using the normalized-restriction formula on both sides, The positive denominators from Step 1 allow cancellation, giving where the last equality is Definition 22. On , consistency in Assumption 3 gives almost surely, hence
Step 3. Fix coordinates and . Target-proxy separation in Assumption 2, applied under , gives independence of and conditional on . Define The product identity under gives The three terms are Multiplying by and using Step 1 yields By Definition 25 and Step 2,
Step 4. Let be the normalized restriction of to . Reference-proxy separation in Assumption 1, applied under , gives independence of and on the latent-treatment cell. Applying the product identity to and , Equivalently, Combining this identity with Step 3 and Definition 24 gives the cellwise factorization
Step 5. Set The bounded outcome-proxy product condition in Assumption 8 implies almost surely, since every matrix entry is bounded by the operator norm. Thus is integrable on every latent-treatment cell. The treatment arm decomposes as the disjoint finite union so Dividing by the positive arm mass from Step 1 and rewriting each cell integral as a conditional mean gives Substituting Step 4,
Step 6. By Definition 20, is the pushforward of by . Hence the observed-arm conditional mean agrees with the corresponding full-data arm conditional mean: The armwise moment notation in Definition 28 identifies the entry of with this same full-data conditional moment. Therefore With the entry of is exactly the displayed sum. Since and were arbitrary,
∎Let be a full-data probability law carrying , and let . Suppose satisfies latent-arm positivity with margin , as in Assumption 9. Then, for each treatment arm , the diagonal matrix defines an injective linear map on .
⊢ LeanFix . For a latent coordinate , set By the definition of , this matrix is diagonal and its th diagonal entry is .
We first show that every diagonal entry is bounded below by the positivity margin. By Assumption 9, Moreover , so monotonicity gives Together with , this implies Since is a probability law, Hence, using , Dividing by the positive number yields In particular , so .
Now let and suppose . Taking the th coordinate and using diagonality gives Since , cancellation gives This holds for every latent coordinate , and therefore . Thus the linear map induced by is injective.
∎Let , , and . Suppose that the linear maps represented by and by , each with domain , are injective. Then
⊢ LeanFor any real rectangular matrix , let denote its Moore–Penrose inverse, characterized by If the matrix map is injective, then Indeed, the first Moore–Penrose identity gives Thus, for every vector , and injectivity of gives . Hence .
Apply this observation to and . Then and transposing the second identity gives Set We verify that is the Moore–Penrose inverse of . First, and For the symmetry equations, the same cancellations give so by the Moore–Penrose symmetry of . Also so by the Moore–Penrose symmetry of . Thus satisfies all four Moore–Penrose equations for , and uniqueness yields
Therefore, using associativity and , This is the claimed cancellation identity.
∎Let and be valid represented atomic laws with finite labelled slot sets: their weights are nonnegative and sum to one. Let be the cumulative-distribution-gap sign Kantorovich–Rubinstein potential selected for the pair . Then is one-Lipschitz, is normalized by and attains the Wasserstein distance of Definition 19: Moreover, for every , if every one-Lipschitz function satisfying obeys then
⊢ LeanWrite with , . Define and let The selected potential is Since , for all , so is one-Lipschitz. Also
We next record the finite one-dimensional transport identity used by the potential. For any finite transport plan from to , put Using the two marginal constraints of the plan gives Hence, with we have Integrating and using yields
Conversely, the finite transport polytope is compact and the quadratic cost therefore has a minimizer. If such a minimizer sent positive mass across some cut in both directions, shifting from the crossing pairs to would preserve the two marginals and decrease the quadratic cost by a contradiction. Thus the minimizer has no counterflow across any cut, and for that plan for every . Consequently Together with the preceding inequality and Definition 19, this gives
It remains to evaluate the potential. For an atom , set Then Interchanging the finite sums with the integral gives For , the expression in braces equals ; for , it equals . Therefore Taking absolute values and using the preceding display for gives
Finally, suppose and every normalized one-Lipschitz test function satisfies the displayed bound. Applying that bound to the just-constructed , which is one-Lipschitz and satisfies , gives
∎Let and be fixed target- and reference-proxy dimensions, and let be elements of the five-block observable summary space in Definition 33. Equip this space with its coordinate Euclidean distance . Then where and is the canonical coordinate identification.
⊢ LeanWrite For , put when , this set is empty and the finite sums below have value zero. Write where and are matrix blocks and , as in Definition 33.
Since is the Euclidean distance on the full coordinate vector, each coordinate-block projection is nonexpansive. Hence, for each , and In particular, and
We use the following finite-dimensional entrywise estimate. Let , let satisfy , and suppose For the canonical vector , Thus, for every with , Using the ordinary matrix action and operator-norm convention in Definitions 8 and 11, taking the supremum over the Euclidean unit ball gives
Apply the preceding estimate with to the four matrix differences , . The coordinate bounds from the first step yield
For the mean block, expand in the canonical basis of : The coordinate bounds and the triangle inequality give
Substituting the four matrix-block estimates and the mean-block estimate into the definition of in Definition 33 gives
For fixed target- and reference-proxy dimensions , let be the summary metric on observable moment summaries from Definition 33. The map from the product of the five-block summary space with itself to is continuous.
⊢ LeanBy Definition 46, for the fixed dimensions and the ambient summary space is Write The coordinate map is continuous for the product topology, and therefore each coordinate projection is continuous. The same finite-product argument covers the zero-dimensional cases, where the relevant products and sums have their usual singleton or empty interpretation.
For a pair , the four matrix-valued difference maps are continuous by continuity of the coordinate projections and of subtraction. If , the associated Euclidean linear operator depends continuously on , because this matrix-to-operator assignment is linear between finite-dimensional normed spaces. Composing with the continuous operator norm gives continuity of
The mean-coordinate difference is continuous as a map into . Its Euclidean norm is The summand coordinates are continuous, the index set is finite, and the square-root is continuous on the nonnegative range of the sum of squares. Hence is continuous.
Using Definition 33, This is a finite sum of continuous real-valued functions on . Therefore is continuous.
∎Let , a full-data probability law carrying , belong to the uniformly conditioned proxy causal model class in Definition 32. Suppose:
(Dimensions.) , where is the number of latent classes and is the target-proxy dimension.
(Radius.) , with the bounded target-proxy condition in Assumption 6.
(Margins.) and , with the latent-arm positivity and proxy-rank margin conditions in Assumptions 9 and 10.
Then the vertically stacked observable proxy moment satisfies where uses zero-based singular-value indexing and are the armwise population moments from Definition 28.
⊢ LeanSet and Let . The proxy-rank margin in Assumption 10 gives Since , the maps induced by and are injective. Hence
The latent-arm positivity condition gives for every . Since , each diagonal entry of is at least . Therefore, for every ,
Now fix . The least-singular-value bound for , restricted to its range, gives
Applying the preceding diagonal bound to gives
The least-singular-value bound for gives
Combining these three inequalities,
By the observable proxy moment factorization in Lemma 4, Thus
For vertical stacking gives the Pythagorean identity and hence
Consequently,
Since is a -dimensional subspace and , the variational subspace characterization of zero-indexed singular values yields
∎These witness calculations locate the explicit two-class comparisons inside the same model restrictions used for the upper bounds. The path-local statement calibrates the observed Kullback–Leibler radius to the product , which is the scale used by Theorem 6.
The final two auxiliary statements concern ordered masses as functions of represented atomic laws. They ensure that, when the support is separated, the ordered-weight target depends on the aggregate measure and is invariant across labelled coordinate representations.
Let , let , and let be a full-data probability law in the uniformly conditioned model class of Definition 32, with observable summary as in Definition 33. Assume that satisfies Assumptions 5, 10, and 9. Let be the -arm moment block from Definition 28, and set Then there exist a signal basis spanning the signal row space of and a real diagonalization of the compressed effect operator from Definition 39 such that:
for each arm , the compressed map is injective;
the stacked proxy moment has rank
for every index ,
the eigenvalue vector of the diagonalization is the latent-effect vector from Definition 23;
writing for the diagonalizing basis matrix, the left and right spectral anchors satisfy where is the latent-mass vector from Definition 21.
Write for the target-proxy feature matrix from Definition 25. The parameter domain in Definition 32 gives , , , , and . With zero-based singular-value indexing, Assumption 10 gives Hence the first singular values of are positive. Choose the thin singular-value factorization where has orthonormal columns and is invertible. Put Then
For each arm , define where is the reference-proxy feature matrix from Definition 24. At the model-generated summary , Lemmas 4 and 18 give and Multiplying by gives
Also, so the row space of each is contained in . For , latent-arm positivity in Assumption 9 gives for every , since and . Combining this with Assumption 10 and the standard product lower bound for the least signal singular value, Thus has row rank . Its row space is contained in the -dimensional space , hence Since the row space of is contained in the same space, the signal row space of equals Thus spans the signal row space of .
The same product-margin argument for each arm gives Because is an orthonormal basis for the signal row space, compression by preserves the nonzero signal singular values: Consequently and the linear map is injective for each .
Let By Lemma 24, so . On the other hand, if then the proxy factorizations give Hence , and therefore The threshold characterization follows immediately. If , monotonicity of singular values gives If , the rank identity gives , while Thus, for every singular-value index ,
Since is injective, its Moore–Penrose inverse is a left inverse on its range: Using the displayed factorizations, Therefore By the compressed-operator formula in Definition 39, where the last equality is Definition 23. Thus is a real diagonalizing basis for , its inverse is , and the eigenvalue vector is .
It remains to identify the spectral anchors. The target-proxy mean factors through the latent masses: Indeed, for each target-proxy coordinate , using Definitions 25 and 21. The anchor condition in Assumption 5, together with Definitions 36 and 25, gives because for each latent class , By Definition 39, Using , , and , and hence Similarly, These are the claimed left and right spectral-anchor identities.
Let and . Let , and let be a thin signal factorization of , written , with coordinate inverse . Let . Suppose that
(Feature bound.) For a constant as in Assumption 6, for every row and latent index .
(Coordinate-inverse bound.) For a constant as in Assumption 10, for every pair of coordinate indices .
Define and where is the entrywise-to-operator norm constant of Lemma 22. Then the real diagonalization associated with and has condition number at most .
⊢ LeanWrite For these constants we use the following finite-dimensional entrywise comparison. If , , and for all entries, then Indeed, writing for the standard coordinate vectors and for the coordinate identification in the definition of , each column satisfies Thus, for every , which gives the displayed operator-norm bound. In particular . Since the columns of are orthonormal, every coordinate of has absolute value at most one, and therefore The same comparison gives
Let and be the coordinate matrix and coordinate inverse from the thin factorization . Since , Thus, for coordinate indices , Consequently,
Define the forward and backward ambient maps Using submultiplicativity, the triangle inequality, and the bounds above, and similarly Both constants are nonnegative because , , and .
Let be the ambient orthonormal extension of the signal columns used in the associated real diagonalization, and let the diagonalizing basis vectors be The backward map is the inverse of the forward map. To see this, put . Then and the identities and give Hence the inverse-coordinate matrix of the basis is computed by first applying and then taking coordinates in the orthonormal basis .
Let be the matrix with columns , and let be its coordinate inverse. For every pair of ambient indices , because . Therefore Likewise, for the inverse-coordinate entries, so The condition number of the associated real diagonalization is the product of these two operator norms, hence
The definition matches the ordering convention in Definition 30 and Algorithm 5. The invariance lemma records the compatibility needed to read ordered masses from any valid representative of the same aggregated atomic law.
Spectral repair and finite-net benchmark
This appendix collects two auxiliary constructions used to connect the population modulus with sample-level procedures. The first is a spectral repair handle that turns empirical spectral information into a positive real atomic law even when eigenvalues are locally coalescent. The second is a class-advised finite-net benchmark that selects a stored admissible summary and then runs the exact spectral construction at that representative. Both constructions use finite-dimensional perturbation ideas familiar from matrix spectral analysis (Bauer et al., 1960; Kato, 1995; Davis et al., 1970; Stewart et al., 1990) and the moment-recovery perspective behind Wasserstein estimation from moments (Wu et al., 2020).
The repair handle records the finite moment target and the positive-law replacement used in the modulus-to-estimation route.
The constructive repair handle is defined for laws in Definition 32 as follows.
At the realized empirical summary, define the first spectral moments by
Among all valid positive -atomic laws , choose a minimizer of
Extract the repaired real probability measure from the minimizing positive -atomic law, transporting aggregate projector mass across each cluster diameter after replacing individual eigenvectors inside overlapping localization discs by Kato contour projectors.
For the labeled-weight converse, differentiate the Bernoulli factorization of and solve the linear score-cancellation equations for the proxy and treatment primitives so that a weight displacement preserves the factorization and the observed score has order .
Algorithm 6 isolates the spectral information needed for positive real-law repair. The moment vector contains the first empirical spectral moments, and the minimizing positive atomic law supplies a real probability measure compatible with those moments. The contour-projector clause expresses the standard perturbation treatment of clustered spectra: inside overlapping localization discs, aggregate eigenspaces carry the stable law-level mass.
The finite-net benchmark makes this repair logic an exact-real finite search after a class-dependent representative library has been fixed.
The inputs are a stored representative library and an observed sample .
For each representative , let be its stored five-block summary, and let be its position in the fixed lexicographic library order.
Define the spectral law attached to representative by the real atomic effect law returned by the thresholded exact-real spectral run for .
For the empirical five-block summary define the comparison criterion
If the library is nonempty, define the selected representative index by with the minimum taken in lexicographic order.
The finite-library spectral program returns the selected representative and its spectral law, when the library is nonempty, and returns the canonical zero law when the library is empty. Its returned effect-law component is
Algorithm 7 defines the library selector at the level of stored summaries. The criterion is the same five-block -type comparison used throughout the paper, and lexicographic tie-breaking makes the selected representative a total deterministic function of the sample and the library.
The estimator below specifies how the representative library is built from an -net of admissible summaries and how the selected representative is converted into an atomic law.
The input is the empirical summary ; the tuning constants are
Partition the coordinate box into a deterministic lexicographically ordered grid of half-open cubes whose side length is at most
If , set and stop.
For each grid cube that intersects , choose one representative Let be the smallest-index minimizer
At , form the vertically stacked proxy moment and retain the right singular directions whose singular values exceed . Denote the resulting orthonormal basis by Using this basis, form the compressed effect operator, left anchor, and right anchor from the compressed-operator formulas.
For each distinct eigenvalue of , define the spectral projector with the empty product equal to .
Output the atomic law
The estimator is class-advised through the choice of admissible representatives . Once the nearest representative is selected, the construction returns to the compressed-operator formula from Definition 39: thresholding fixes the rank- signal space, and the spectral projectors aggregate the anchor weights attached to each distinct eigenvalue.
Fix integers and real constants . Suppose
(Dimensions.) , , and .
(Bounds and margins.) , , and .
Then there are constants and such that, for every integer , there is a class-dependent finite library of representative summaries as in Algorithm 8 with the following properties. For every uniform provider of the fixed-parameter exact-real primitives on the admissible core parameter domain, the finite-net estimator is Borel measurable, and For every observed sample , the operation count of the same exact-real execution that returns is at most Its trace is exactly the empirical-summary trace, followed by the exhaustive finite-library search trace, and followed by the selected spectral-run trace when a representative is selected; if no representative is selected, the final spectral-run trace is empty.
If the program selects index on sample , then is a nearest library summary to the empirical summary of Definition 35: Among all tied nearest summaries, has minimal lexicographic rank: For the exact spectral run at , the estimator equals the run’s returned effect law. For each treatment value , the observed proxy moment at , multiplied by the run’s basis matrix, defines an injective Euclidean linear map. The stacked proxy moment at satisfies the exact threshold characterization for every singular-value index . Every eigenvalue of the compressed operator formed at with the run’s basis and spans appears among the eigenvalues returned by that run.
For every stored representative and every rectangular perturbation , implies that singular-value thresholding at recovers exactly the -dimensional matrix signal space for For every index , there is a model law from Definition 32 whose observable summary is , and the exact spectral run at returns the quotient law . Each returned representative effect law is a valid atomic law.
For every , every observed sample satisfies Finally, for every tail level with and every ,
⊢ LeanTheorem 8 establishes the finite-net benchmark. For fixed dimensions and fixed conditioning constants, the representative library has polynomial size in , the online exact-real operation count is polynomial in the same sense, and the selected summary is the lexicographically first nearest library representative. The deterministic inequality converts empirical summary error plus the mesh scale into quotient-law error, and the final tail bound applies the same concentration scale used for the main law estimator.
The intuition is that the admissible summary image is compact and finite-dimensional, so an -mesh supplies a representative close enough for the gap-free modulus to absorb the discretization error. At the chosen representative, the proxy-rank margin makes singular-value thresholding recover the correct -dimensional signal space under perturbations below . The exact spectral run then returns the quotient law for a model law with that stored summary, and the modulus transfers the distance from the empirical summary to distance between effect laws.
Proofs and verification note
This appendix collects the proof layer for the paper’s formal results. The arguments establish the population stability statement in Theorem 1, the measurable repair and estimation claims in Proposition 3, Proposition 4, and Theorem 2, the confidence and reporting guarantees in Theorems 3, 4, and 5, and the local converse statements in Theorem 6, Proposition 6, Proposition 7, and Theorem 7.
The proof of Theorem 1 uses the compact summary closure from Proposition 2 and the quotient moment representation in Definition 43. The finite-dimensional moment map gives a Lipschitz modulus from the observable summary coordinates to the law target on the compact domain. The extension clause follows by completing the admissible image under , with Algorithm 1 and Proposition 3 supplying the measurable nearest-summary selection needed to evaluate the extended map at empirical summaries.
The estimation proofs combine this deterministic modulus with concentration for the empirical five-block summary in Definition 35. For the repaired estimator, nearest-summary projection keeps the evaluated summary within the same concentration scale as the population summary. For the structured-lattice estimator, Algorithm 2 and Proposition 4 give a duplicate-free finite search whose deterministic error is bounded by the empirical summary error and the lattice scale. These two routes yield the joint root- tail statement in Theorem 2.
The confidence-set proofs use the same concentration event with the objects in Algorithm 3. The summary-inversion set inherits its diameter from the modulus on , while the algorithmic outer set inherits its radius from the deterministic lattice bound. The transport-plan formulation in Definition 2 gives the finite representation used in Theorem 3. The cluster-report argument then applies the atom floor and transport budget in Algorithm 4 to associate true atoms with empirical components, yielding the support containment, mass-coverage, and interval-length conclusions in Theorem 4.
For ordered weights, the proof of Theorem 5 works on the gap-local stratum in Definition 41. The positive effect gap converts law-level transport error into an error bound for the effect-ordered mass vector defined in Definition 30 and estimated by Algorithm 5. This gives the clipped inverse-gap upper rate stated in the theorem.
The lower-bound proofs use the explicit two-class witness in Definition 55, its validity certificate in Proposition 5, and the local experiments in Definitions 56 and 57. Standard two-point testing inequalities (Cam, 1986; Polyanskiy et al., 2019) convert the displayed Kullback–Leibler bounds and target separations into the local risk lower bounds in Theorem 6. The same comparisons give the same-class minimax rates in Propositions 6 and 7. The transfer statement in Theorem 7 expresses the same witness-pair logic for comparator classes stated in the published VMW vocabulary (Virk et al., 2026).
Verification note.
The Lean-checked scope consists of the matched theorem, proposition, and lemma declarations listed in the verification manifest for commit Definitions marked presentation-synthesized in that manifest are expository renderings of paper notation; definitions marked matched have the Lean declarations recorded by the manifest. The proof text in this appendix follows the matched declarations and preserves the % lean: audit markers that identify certified proof steps.
The Lean toolchain is leanprover/lean4:v4.33.0. The CausalSmith package is version 0.1.0; its manifest records the path dependency Causalean at the same repository commit, inherited path dependencies optlib at ../third_party/optlib and FoML at ../third_party/lean-rademacher, and the following external revisions: mathlib db584cd6d46c92f209a44c0f1c829460d327499d, plausible b7eb3304aeae834b12dda98993a37f6a41f6f0bb, LeanSearchClient 5f4d51b81cbd3f6b32b156bfad9056621a040404, importGraph 16f02aa7642864af59f1ff0e384a015994db9118, proofwidgets 4be2e3d5087eeb272cf5a8853b8f9dd025ef5957, aesop 3448c0bcc5ce01b2d1546e483ec3620e32df3d0e, Qq 92c15be17b7caf78c2ad767ec40f89052d908d81, batteries 4488d40d070b9700d4d5a6aa342f0d40c31b2a2d, checkdecls 3d425859e73fcfbef85b9638c2a91708ef4a22d4, and Cli 6130a47896ce867c6a4a55373441e59e565bad0f. The verification build command for the paper module is
Proposition 1 uses the VMW model-scope claim from (Virk et al., 2026). Theorem 7 uses that model-scope claim together with the VMW separated-recovery scope claim from (Virk et al., 2026). The remaining matched declarations listed below are derived from their displayed hypotheses and previously matched declarations.
| paper object(s) | Lean source file and declaration(s) |
|---|---|
| Proposition 1 | TObservedVMWMarginInclusion.lean: observed_vmw_margin_inclusion |
| Proposition 2 | TSummaryClosureCompact.lean: summary_closure_compact |
| Lemma 1 | Helpers/Concentration.lean: uniform_summary_concentration |
| Theorem 1 | TGapFreePositiveMeasureModulus.lean: gap_free_positive_measure_modulus |
| Proposition 3 | TSummaryRepairTotalBorel.lean: summary_repair_total_borel |
| Proposition 4 | TPolynomialLatticeEstimator.lean: polynomial_lattice_law_estimator |
| Theorem 2 | TCollisionUniformRootN.lean: collision_uniform_root_n |
| Theorem 3 | THonestRootNConfidence.lean: honest_root_n_confidence |
| Theorem 4 | TClusterAdaptiveReport.lean: cluster_adaptive_report |
| Theorem 5 | TLabeledWeightUpper.lean: labeled_weight_upper |
| Proposition 5 | TTwoClassWitnessValid.lean: two_class_witness_valid |
| Theorem 6 | TMatchingLocalLowerBounds.lean: matching_local_lower_bounds |
| Proposition 6 | TSameClassQuotientMinimax.lean: same_class_quotient_minimax |
| Proposition 7 | TSameClassLabeledMinimax.lean: same_class_labeled_minimax |
| Theorem 7 | TPublishedVMWConverseTransfer.lean: published_vmw_converse_transfer |
| Theorem 8 | TFiniteNetLawEstimator.lean: finite_net_law_estimator |
| appendix auxiliary lemmas | Helpers/ObservedLawAdapters.lean, Helpers/ObservedMarginAssembly.lean, Helpers/ConditionalMomentAdapters.lean, Helpers/OutcomeFactorization.lean, Helpers/AmbientOperatorBridge.lean |
| witness and path lemmas | Helpers/WitnessValidity.lean, Helpers/PathModelCertificates.lean, Helpers/PathLocalExperiments.lean, Helpers/WitnessKL.lean |
| metric, spectral, and ordered-mass lemmas | Basic.lean, Helpers/SummaryMetric.lean, Helpers/StructuredLatticeFunctionalCalculus.lean, Helpers/ModelSpectralConstruction.lean, Helpers/RepresentativeSpectralCertificate.lean, Helpers/OrderedMassStability.lean |
Reader-facing sections use one-based singular-value notation. Lean-facing auxiliary displays that expose array indices translate the th signal singular value into the corresponding zero-based index inside the verification layer.
Proofs of the main results
Fix satisfying the displayed hypotheses, and set All model restrictions used below are the components of in Definition 32.
For each treatment arm , Lemma 3 gives Since and , this arm probability is positive. Define The diagonal entries are so Lemma 4 yields
Latent-arm positivity gives, for every latent class , because . Hence, for every , and therefore The proxy-rank margin in Assumption 10 gives Using the product lower bound for full-column-rank factors in the factorization above,
Let and let have orthonormal columns spanning . For the fixed arm , the row space of is contained in : The lower bound from the previous step gives , so this row space has dimension at least . Since has dimension , the two spaces are equal. The map is an isometry from onto this row space, and composing with that isometry preserves its positive singular values. Thus
The stacked proxy signal margin is supplied by Lemma 24. Applied with the present model membership and numerical hypotheses, it gives with the displayed one-based singular-value indexing of the statement.
By Definition 28, The envelope components of give Since , Jensen’s inequality for the two conditional matrix means and for the unconditional vector mean yields
Fix and , and put Latent-arm positivity gives Let be the th canonical vector in . The proxy-rank margin implies Since , some coordinate satisfies Define If , then a contradiction. Hence By Lemma 5, The reference-proxy separation component in Assumption 1 gives . Therefore positive conditional probability for would give On this intersection, contradicting the product envelope. Thus On , the consistency component in Assumption 3 gives almost surely. Inside the conditional law given , the event has positive probability, and the latent-ignorability component in Assumption 4 gives . If had positive conditional probability given , independence would give a positive-probability violation on , contradicting the displayed cell bound. Hence Finally, Lemmas 6 and 7 give
It remains to identify membership in the published VMW parameter class. The scope in Definition 52, together with the cited VMW correspondence recorded in Definition 60 and (Virk et al., 2026), identifies with the qualitative structural conditions summarized there. Those conditions are supplied as follows. The proxy independences, consistency, and armwise latent ignorability are components of . The proxy-rank margin and give full column rank: Latent-arm positivity gives for every and , and hence the corresponding latent treatment probabilities are strictly positive. Together with , , and , these are exactly the qualitative conditions for .
First record the coordinate estimate used to pass from operator envelopes to coordinate boxes. For any , any , and any , where is the th canonical basis vector of .
Define the finite coordinate boxes and Each factor is a finite product of closed intervals, hence compact, and therefore is compact. For a five-block summary , define the coordinate map This map is a homeomorphism from the space of five-block summaries onto this finite product coordinate space. Thus is compact in the summary topology.
Let . For , put By Assumption 9, together with and , the arm has positive mass. Let be the probability law obtained by restricting to and normalizing by this positive mass. By Definition 28, the armwise blocks in are conditional means; on this positive arm, equivalently, and . By Assumption 7, the proxy-product envelope restricts to : -almost surely. Since , the matrix-valued integral is integrable and The same normalized-restriction argument with Assumption 8 gives Likewise, Assumption 6 gives Hence every coordinate of has absolute value at most . Combining these inequalities with the coordinate estimate from the first step and the summary tuple in Definition 33 gives Since was arbitrary,
The same block envelopes give the stated uniform -bound. For , For the final term, Therefore Taking the scalar bound and then instantiating the existential bound with proves the boundedness claim.
Since is compact and closed and contains , it also contains The set is closed by definition as a closure, so it is a closed subset of the compact set . Hence is compact.
Finally, because a closure has the same nonemptiness as the set it closes. By the definition the condition is equivalent to : a law gives , and any element of is for some . Thus is nonempty if and only if is nonempty.
Fix satisfying the displayed hypotheses. Let be the finite disjoint union of the following scalar summary coordinates: For from Definition 18, define Set The bounded target-proxy, bounded proxy-product, and bounded outcome-proxy components of Definition 32, namely Assumptions 6, 7, and 8, give, for every , where is the observed-data marginal of Definition 20.
Let be a finite coordinate-to-operator-norm constant such that, for every real matrix and every , Such a constant exists because, if each coordinate of a vector in has absolute value at most , then its Euclidean norm is at most ; applying this to each column and then taking the supremum over unit input vectors gives a finite constant depending only on . We shall also use the corresponding vector estimate
Put and choose Then . Since and , the product is positive, so all denominators in this definition are positive.
Now fix , , , and . Define For each , Hoeffding’s inequality under gives Indeed, -almost surely, the interval length is , and the choice reduces the exponential factor to .
Let By the union bound and , Also set The -almost-sure coordinate envelope transported to the product experiment gives
The population summary blocks obey For the armwise blocks, Lemma 3 gives positive arm probabilities, and Jensen’s inequality applied to the conditional moments in Definition 28 transfers the almost-sure bounds on and . For , Jensen’s inequality applies directly to the unconditional target-proxy mean.
It remains to prove a deterministic implication off the two bad events. Write By Lemma 3, transported through Definition 20, First suppose that On , For , put The coordinate integral identities are Because , the total empty-arm convention in Definition 34 gives and the same identity with and . Since each population entry is bounded by , the preceding displays imply, for both matrix types, The matrix norm conversion and the mean-coordinate bounds therefore yield
Now suppose that On , every empirical arm matrix has entries bounded by : if , the convention in Definition 34 gives the zero matrix, and if , the matrix is an average of entries whose absolute values are at most . Hence each empirical matrix block has operator norm at most , and Combining these empirical bounds with the population envelopes gives Since , and therefore
The two cases imply Consequently,
Finally fix . The event is the event that all observed records have . Since , The lower bound from Lemma 3, again through Definition 20, gives Thus and monotonicity of on gives This proves both asserted conclusions.
∎Put For matrix sizes , write with as in Lemma 22. Define and The assumptions and , together with nonnegativity of the entry-norm constants, give Set Since , , and , we have
Fix . Let be the target-proxy feature matrix of Definition 25. The proxy-rank margin gives Choose a thin singular-value factorization The singular-value lower bound implies For every coordinate and latent class , latent-arm positivity gives Under the normalized restriction of to , Lemma 5 gives Therefore Thus Lemma 26 applies to .
Define the ambient contrast By Lemmas 4 and 18, for each , where The proxy-rank margin and Lemma 19 make injective, and the proxy-rank margin makes injective. Hence Lemma 20 yields Subtracting the two arms gives
Let Then Extending the columns of to an orthonormal basis of , the columns obtained by applying form a real eigenbasis for . The signal eigenvalues are , and the remaining eigenvalues are . Consequently with The eigenvalue bound uses Lemma 7 on the signal coordinates and on the remaining coordinates.
It remains to record the two anchor identities used by the spectral representation. For each coordinate , where the latent classes form a finite measurable partition and is integrable by Lemma 5. Hence Also, since almost surely and , so the first canonical basis vector satisfies Therefore, for every one-Lipschitz satisfying , and
Fix . For each arm , Proposition 1 supplies Together with the proxy factorization, these inequalities give We first record the armwise perturbation estimate. Let have the same rank , and suppose Then because the operator norm of the Moore–Penrose inverse is the reciprocal of the smallest nonzero singular value. The Penrose equations, applied with the common rank condition, give the exact moving-space identity Here and are orthogonal projection residuals, so their operator norms are at most one. Taking norms term by term gives Hence For matrices , the algebraic split therefore implies, whenever , Applying this estimate with gives Summing over and using the definition of yields Moreover,
Let be the normalized attaining Kantorovich–Rubinstein potential for from Lemma 21. Then Use the diagonalizations in Equation 1 for and , written Let For any real number , define the aggregate projectors and The diagonal - matrices have operator norm at most one, so for and , The projectors extract their spectral values on the appropriate side: Hence In particular, when , Functional calculus with the two diagonalizations gives Since the aggregate projectors sum to the identity for each diagonalization, subtraction and expansion over yield where The equal-spectral-value branch is zero by the preceding vanishing identity. In the unequal branch, the one-Lipschitz property gives and therefore Because and , the triangle inequality gives Also, because , is one-Lipschitz, and every eigenvalue of lies in .
Using the representation identity for and , This proves the asserted modulus on model summaries:
Let For , choose a representative law with and define This definition is independent of the representative. Indeed, if another has , then the modulus just proved gives A zero finite transport cost between quotient atomic laws forces the same aggregated mass at each support point, so the two quotient laws are equal. Hence and, for all ,
Let be the ambient product metric on the five summary blocks: each matrix block has its Frobenius norm, the mean block has its Euclidean norm, and the finite product uses the maximum of the five block distances. Set By Lemma 22, For a matrix block , each entry obeys so summing the squared entries gives Since is the maximum of the four Frobenius block distances and the Euclidean mean-block distance, while is the sum of the four operator-norm distances and that same mean-block distance, (The assumptions cover the positive-dimensional case used here; in a zero-dimensional matrix block the unique empty matrix contributes zero.) Thus and induce the same topology and the same closures.
The first comparison and the -control on give Hence is Lipschitz on the summary image in the ambient metric.
Inside , put The equality of closures just proved makes dense in . The quotient atomic-law space is complete for : in labelled coordinates it is the metric quotient of the compact set Apply the dense-set extension theorem to the ambient-Lipschitz map The dense-set extension construction gives a Lipschitz, hence continuous, map that agrees with on , and the uniqueness property of this construction makes it the unique continuous extension with that agreement.
It remains to transfer the sharper -control. Fix . Choose sequences converging in the ambient metric to , respectively. Continuity of and Lemma 23 give For every , agreement on and the model-summary control in Equation 2 yield Passing to the limit proves Continuity of , together with the Borel structure on the quotient atomic-law space, makes Borel measurable.
For every , the summary belongs to , and hence Let be continuous, satisfy the stated -Lipschitz bound, and agree with on all model summaries. For any , choose with Continuity and agreement on imply Thus for all . Combining the model-summary modulus, the continuous extension, the Borel measurability, and this uniqueness proves the theorem.
Write By Proposition 2, is compact. Also gives , and Definition 26 with and gives . These two inequalities supply the positivity data needed for the atomic-law fallback . From Theorem 1 we take a continuous, hence Borel, extension satisfying, for every admissible model law ,
First suppose . Define the selector on the full five-block summary space by This map is Borel. The extension condition over model summaries is vacuous, because any admissible would give contradicting . The nonempty- nearest-point clause in Algorithm 1 likewise has a false premise. Thus these choices define admissible repair data . For every observed sample , Algorithm 1 gives so is the constant Borel map. In the same case, Algorithm 3 gives the retained set which is nonempty.
Now suppose . Let be the number of real coordinates in a five-block summary, and let be the coordinate homeomorphism from the summary space to Euclidean space. Set and, for , define the transported loss The set is nonempty and compact, and is continuous because is a homeomorphism and Lemma 23 gives continuity of . Applying Lemma 2 to , then transporting the selector back through , gives a Borel map such that, for every summary , Together with the from the first step, this defines repair data as in Algorithm 1; the fallback clause has premise , which is false in the present case.
For , the repaired estimator is The empirical summary map is Borel by hypothesis, is Borel by the preceding step, and is Borel because it is continuous. Hence is Borel measurable.
It remains in the nonempty case to verify nonemptiness of the retained summary-inversion set. By Definition 45, Since , , , and , one has , hence , and therefore For a fixed sample , put The selector property gives , and the defining formula for gives Moreover, Thus witnesses membership of in the displayed retained set, so that set is nonempty for every sample.
The two cases and cover all possibilities. In each case the constructed repair data has Borel selector , continuous extension , Borel repaired estimator , the asserted fallback when , and a nonempty retained summary-inversion set for every sample.
Set and order singular values nonincreasingly.
Let and be the structured lattice of Definition 13. Its four factors have cardinalities bounded by fixed-parameter constants times Since , the lattice height satisfies for a finite . Hence there is a constant , depending only on the displayed parameters, such that In the fixed-dimensional exact-real model of Definition 48, the work is the empirical-summary pass plus one fixed-size objective evaluation for each candidate, so
Choose from Lemma 1. Thus, for every , , and , The assumptions imply , , and . Therefore Define Then .
Fix . The lattice factors are finite and nonempty: canonical coordinate columns furnish an element of , a rounded scalar multiple of furnishes an element of , the inequality furnishes an element of , and . Enumerate in duplicate-free lexicographic order as For a summary , put where is the criterion in Definition 16. Thresholded singular-value pseudoinversion is Borel, each score is Borel, and a finite first-minimizer rule over Borel scores is Borel. Hence is a total Borel empirical-summary rule, with as in Definition 35, and it is the first minimizer prescribed by Algorithm 2. Every coordinate of every is at least , so aggregating coordinates at repeated effect locations gives for every sample .
Fix , and put Lemma 25 gives and the threshold characterization at . In addition, Lemma 24 gives the population signal margin; since the singular values are nonincreasing, In particular, has exactly its first singular values at least . Now let satisfy and set . The stacked-block comparison gives Weyl’s inequality yields Thus, for , while, for , This proves that also has exactly its first singular values at least and all remaining singular values below .
Fix a sample , and write Choose the population signal tuple By Definitions 25 and 21, and because the latent classes partition the record space with positive class masses, The chosen signal factorization therefore gives Using Assumption 5, the first target-proxy coordinate has conditional mean one in every latent class: Since , this yields The proxy and outcome moment factorizations from Lemmas 4 and 18 give, for each arm , The proxy-rank margin in Assumption 10 supplies injectivity of the linear maps represented by and . By Lemma 19, is injective, so is injective. Applying Lemma 20 with , , and yields Subtracting the two arms and using , together with , gives Moreover, where the effect envelope is supplied by Lemma 7. The associated quotient law is
Recall that . The definition of gives the four mesh inequalities For the signal basis , each entry has absolute value at most one because its columns are orthonormal. Coordinatewise clipped rounding therefore gives a grid matrix with Weyl’s inequality and the first mesh bound give Let be the polar factor . The polar-factor perturbation estimate at an orthonormal basis yields Next, since , coordinatewise clipped rounding in the radius gives with grid entries and Using and Weyl’s inequality, and using , For the weight vector, set . Since and , the quantities and satisfy and . Because , this implies , hence The floor vector satisfies and . Choose a subset of indices of size and add one to on that subset, obtaining integers with Thus belongs to , and summing the squared coordinate errors gives Finally, since , coordinatewise clipped rounding on the effect grid gives with Combining these four rounded components gives a well-formed lattice comparator satisfying
Assume first that For each treatment arm, The population armwise proxy block has rank , kth singular value at least , and . If denotes the rank- truncation retained at threshold , singular-value perturbation gives Define the thresholded empirical compressed operator, in the notation of Definition 16, by The equal-rank Moore–Penrose perturbation identity gives and therefore For the comparator , the mean residual decomposes as so the bounds , , , and give Similarly, and , , and yield For the operator residual, write The inverse identity and the singular-value bounds give Bounding the five displayed summands with , , and gives Consequently Let be the first minimizer. Minimality and the nonnegativity of the three terms in imply and Since all three residuals are bounded by
Let be the attaining Kantorovich–Rubinstein potential for supplied by Lemma 21. Thus and The selected structured operator and , extended by zero on the orthogonal complement of their signal spaces, have real diagonalizations with spectra in . Put For either signal factorization with coordinate matrix , the ambient eigenbasis uses the signal coordinates and an orthonormal basis on the orthogonal complement. Hence its condition number is bounded by For the selected lattice tuple, the defining lattice constraints give Therefore where the last two inequalities use , , and . Thus the selected diagonalization has condition number at most . For the population tuple, so the same calculation, with gives the condition-number bound for the population diagonalization as well. For the diagonalizations the divided-difference identity gives with when the two eigenvalues coincide. The coordinate perturbation in the last display is Since , each column of the Schur product has Euclidean norm bounded by the corresponding column norm of this coordinate perturbation. The fixed-dimensional column estimate therefore gives Multiplying by the exterior diagonalizer factors yields Because , this is For the selected law, put Since and , the functional calculus of the selected structured operator gives because on the orthogonal complement. Therefore The labelled integral of is and hence Because and Cauchy’s inequality yields For the population tuple, functional calculus gives since on the orthogonal complement. Using , , and , we obtain Together with and the residual bounds from the previous step give, in the small-error case, If instead , both laws are supported on , so the diameter bound from Definition 19 gives Finally . Combining the two cases with the definition of proves, for every sample,
The choice gives the displayed candidate-count and work bounds with the same . For the tail bound, fix . Since we have On the event the deterministic inequality gives Therefore Applying Lemma 1 under yields
Let in the finite-dimensional summary space equipped with . By Lemma 1, choose a simultaneous concentration constant . By Proposition 4, choose a positive lattice tail constant associated with the structured-lattice estimator. By Theorem 1, choose and a continuous extension such that and Set Then , and depends on .
Fix . The empirical summary is a Borel sample map: its coordinates are finite sums of Borel coordinate maps, treatment-arm indicators, products, and the total empty-arm normalizers in Definition 35. Applying Proposition 4 at this , choose the structured-lattice estimator . Its finite list contains each well-formed candidate once, in lexicographic order, and the returned law is where is the first lexicographic minimizer of over the structured lattice in Algorithm 2, with threshold . The same result gives Borel measurability of and, for every and every ,
Proposition 2 gives compactness of . If , take the selector for every ambient summary , and define the repaired estimator by the zero-law branch in Algorithm 1. If , identify the finite-dimensional summary space with Euclidean space. The loss is continuous by Lemma 23, so Lemma 2 gives a Borel map such that With this selector, set Continuity of , Borel measurability of , and Borel measurability of make a total Borel estimator into .
Fix . Then , so . For every sample , nearest-summary optimality gives By the triangle inequality for , Since , the Lipschitz extension bound gives
Fix , and define The definition of implies Hence Monotonicity of and of the square root gives and Therefore Indeed, on the complement of the right-hand union, the lattice error is at most , while the repaired-summary error is at most .
Apply the lattice tail bound with and Lemma 1 with . The displayed set inclusion and the union bound yield The arm-count conclusion is also supplied by Lemma 1: for each ,
By Lemma 1, choose such that, for every , , , and , Fix , and put Let be the summary closure from Definition 43. By Theorem 1, choose and a continuous extension such that and By Proposition 4, the prescribed structured-lattice constant is positive. For each , Proposition 4 supplies a measurable structured-lattice estimator , generated by the finite exhaustive duplicate-free lexicographic list of well-formed lattice candidates and by the first minimizer of , such that, for every sample and every , The same cited result gives, for every sample, For , define The parameter domain gives , and gives .
Construct the repair selector . By Proposition 2, is compact. If , set If , then Lemma 23 and Lemma 2 give a Borel map such that, for every ambient summary , Together with , this is the repair datum of Algorithm 1; the induced repaired summary-to-law map is measurable, and the displayed Lipschitz bound for is retained on .
Fix a sample . If , then , hence is nonempty. If , then , and Thus so is nonempty. The algorithmic set is nonempty because By Algorithm 3, and membership in is membership in the atom-floor subclass together with By Definition 19, this last inequality is equivalent to the existence of a finite nonnegative transport plan from to whose absolute-distance cost is at most . This gives the asserted finite constrained transport representation.
Fix . Since , the closure is nonempty. Define the bad event The concentration bound gives The empirical summary is Borel by Definition 35, and is continuous by Lemma 23; hence is measurable. Therefore For every , Nearest-summary optimality gives and the triangle inequality gives Since , the summary itself witnesses
On the same sample , Proposition 4 gives By symmetry of , It remains to verify the atom floor for . For each latent class , Assumption 9 as included in Definition 32 gives The two treatment cells partition the latent class, so where is the latent mass from Definition 21. For the raw latent law the aggregate mass at any distinct support value is because the displayed sum contains at least one latent class. Passing to the quotient law of Definition 40 only aggregates equal atom locations, so Definition 31 gives Thus Combining the two memberships on ,
Fix a sample . For the theoretical diameter, take any If , then both elements equal the zero-law fallback, so If , then by Algorithm 3 there are such that By symmetry and the triangle inequality, The Lipschitz bound for gives Taking the supremum over all pairs in the nonempty set yields
For the algorithmic diameter, take any By Algorithm 3, By symmetry and the triangle inequality, Taking the supremum over all pairs in the nonempty set gives
By Lemma 1, choose with such that, for every , , , and , Fix . With the displayed definitions of , Proposition 4 gives and, for every , a Borel structured-lattice estimator obtained from the exhaustive duplicate-free lexicographic list of well-formed candidates and the first minimizer of . It also gives, for every sample , and, for every , The extension on is supplied by Theorem 1. By Proposition 2, is compact, and by Lemma 23, is continuous. If , Lemma 2 supplies a Borel selector satisfying If , set for every ambient summary . Together with , these maps are repair data of the form specified in Algorithm 1; the induced sample-to-law map is Borel measurable by composition in the nonempty case and by the constant fallback in the empty case.
Fix , , and . Define Since , , , , and , we have The event is the complement of the deviation event in the first display with , and hence For every , the deterministic lattice bound gives
The representative law has atom floor . Indeed, for each latent class , by the latent-arm positivity component of . After coincident latent effects are aggregated as in Definition 40, every distinct support atom has aggregate mass at least . Since is symmetric on represented atomic laws, the preceding step gives Thus because Algorithm 3 defines
We shall use the following finite-transport implication. If two finite atomic probability laws and have atom floor and then For the first assertion, put If , then every coupling in Definition 19 moves the aggregate mass at , which is at least , by distance at least . Hence every transport cost is at least , contradicting the displayed Wasserstein bound. The second assertion follows by applying the same argument after transposing a coupling. Applying this implication with and using , gives mutual -closeness of the two supports.
The components are the connected components of under links of length at most . For such a component , The mutual support closeness from the previous step implies that every belongs to at least one such set. If belonged to two components, choose center atoms in the two components within of . These two center atoms are within of each other, so they lie in the same connected component. Hence For every component , reverse support closeness supplies a true atom within of any center atom in , so . The inclusion is part of the displayed definition of . If , then some satisfies , and therefore which is the reported support interval in Algorithm 4. Since , Consequently which proves the mass-interval containment.
Fix . Then and the triangle inequality gives The transport implication applied to and shows that their supports are mutually -close. For a component , write when , where is the finite real external gap.
If , then The partition just proved makes the unique empirical component associated with the true support. The same partition argument for gives Thus
Now suppose . The finite sets and are nonempty in this case, so Let be any transport plan from to . If then marginal balance gives Every positive-mass crossing pair in these two sums is separated by at least . For the first crossing case , , choose with Since is a center atom, support closeness from the center to gives with Then . Because , the definition of gives Also and therefore
For the second crossing case , , support closeness from to the center gives a center atom with Let be the empirical component containing . Support closeness from the center to gives with so . If , the uniqueness of the true association would force ; then , and contradicting . Hence Since , the external-gap definition yields Together with this gives The preceding crossing estimate and the displayed marginal identity imply Taking an optimal plan and using , if , then Since and , this is bounded by If , then both cluster masses lie in , so their difference has absolute value at most . When , this is at most . When , the inequality gives and hence Thus, in all finite-gap cases, with the top-gap case giving equality of the two masses.
For a fixed component , let The preceding step bounds every within of when , and within when . Since , the elementary extremal implication applies to the endpoints of . Therefore, if , then and, for every component, with the displayed right side read as zero when .
If , every center atom is within of a true atom . That true atom satisfies , and hence . Thus Applying this to the minimum and maximum points of the finite component gives The reported support interval therefore has length
In the same singleton case, the external gap dominates the global effect gap. If , then for every , because and are distinct positive-mass support points of . Taking the infimum over such gives If , the preceding zero-width conclusion applies. Since and , substituting the finite-gap comparison into the external-gap interval bound gives
Finally, Definition 19 defines as the minimum finite transport cost over couplings with the prescribed marginals. Therefore, for every represented atomic law , The forward implication takes an optimal finite transport plan attaining ; the reverse implication uses that any feasible plan upper-bounds the minimum. This is the finite constrained representation of .
Let be the constant supplied by Theorem 2. Define The hypotheses give , , , and , hence . Fix . By Theorem 2, choose the measurable nearest-summary quotient-law estimator from Algorithm 1. Set which is exactly the ordered-weight estimator of Algorithm 5. Its measurability follows from the measurability of and the Borel ordered-mass map; its values lie in the probability simplex by the ordered-mass construction, including its barycentric fallback branch.
Fix with , and fix . Write where is the quotient latent-effect law of Definition 40. The support radius gives for every , so is integrable under . For every , the tail bound in Theorem 2 and the inequality give Indeed, the displayed event is contained in the corresponding collision-uniform event because To integrate this tail, put For , define Then and Thus The layer-cake identity and the Gaussian tail bound yield
Let be the -slot representative of . By Lemma 17, because and represent the same aggregated atomic law. The gap-stratum assumptions give, for every latent class , and Definition 41 gives distinct latent effects and Consequently, if is any valid -slot atomic law and then For this last implication, take an optimal transport plan . Each true atom must have a positive-mass atom of within distance ; otherwise the mass transported from that atom alone would cost at least . The map is injective, hence bijective, because two -close images of distinct -separated true atoms would violate the triangle inequality. It also preserves the increasing order. Therefore With the row and column marginal identities give All off-matching mass moves at least , so and the displayed local stability bound follows.
Apply the preceding bound to the representative of . Since and represent the same law, If then If instead , the simplex diameter gives Thus, for every sample ,
Integrating the samplewise inequality and using the mean bound for , Independently, both and lie in the simplex, hence for every , and therefore
If , the preceding rate bound and give If , the simplex-diameter bound and give The constant depends only on , and the construction above supplies the required estimator for every .
Fix . For , write The law in Definition 55 is the finite law on points where . Its elementary mass at , for , is with The perturbation range gives so every displayed Bernoulli factor is a probability mass function. Summing successively over leaves , and . Thus is a probability law.
The parameter domain in Definition 32 is satisfied because The finite product form gives the conditional factorization requirements. For bounded measurable test functions , and These identities give Assumptions 1, 2, and 4. The construction gives and almost surely, which are Assumptions 3 and 5. Since and , so Assumptions 6, 7, and 8 hold with . The latent-arm masses are and each is at least , giving Assumption 9. The feature matrices are For every , Indeed, after multiplying the three differences by , respectively, the differences equal and Thus the smallest singular value of each of is at least , proving Assumption 10. Hence
Lemma 10 gives Also, Lemmas 11 and 12 give It remains to compute the observable proxy moment determinants. For and , put After summing out and , Definition 28 gives the conditional-normalized finite sum Since the - and -sums are Bernoulli means, this is For , the conditional latent weights are , and for they are . Therefore Taking determinants, and
Lemma 8 gives and Lemma 9 gives Since in Definition 40 places latent mass at latent effect ,
At , the singular-value inequalities in the second step give injectivity, hence nonsingularity, of . The determinant identities in the third step give nonsingularity of and . The quotient-law identity in the fourth step becomes Finally, Lemma 8 gives and these two latent-class masses are distinct.
Define Since , and Fix , and put Then , and
Let and let with as in Definition 18. For write For , define the finite visible witness cell mass by where The pushforward of this finite law by is , the observed-data marginal of Definition 20. At , every visible cell is bounded below by the first-class contribution: Only the treated outcome factor varies with , and hence For Kullback–Leibler divergence as in Definition 53, the elementary finite-carrier comparison is whenever and are probability vectors on a finite carrier and for every ; this follows from . Therefore where the first inequality also uses data processing under . Taking , Proposition 5 gives , and the displayed KL bound places both laws in from Definition 56.
We use the following two-point expected-risk inequality. Let be one-observation laws with integrable log-likelihood ratio and suppose Tensorization and Pinsker’s inequality give For an estimator in a metric space with targets , set The triangle inequality gives when , and the case is immediate. Thus so one of the two probabilities is at least . Since one of the two expected risks is at least . For simplex-valued estimators the same argument applied to one coordinate gives because . The loss is bounded on the stated compact support radius, and the simplex loss is bounded by , so the expectations used below are finite.
Apply the two-point inequality with The KL calculation gives By Proposition 5, Every coupling from the point mass at to the second law transports total mass one over distance . Hence, using Definition 19, Thus every measurable quotient-law estimator has one of the two local witness laws satisfying This proves the quotient-law risk assertion.
Fix , and define the calibrated displacement of Definition 58. Since and , Together with , this gives In particular , so lies in the tangent-amplitude domain. Moreover, and therefore Consequently The hypotheses of Lemma 16 are satisfied for , and also for . Hence
We prove the labelled-path KL and ordered-mass certificates. For , define The pushforward of this finite law by is . At , Definition 50 gives . For , the first latent class contributes at least For , because , both treated outcome probabilities lie in , and the first latent class contributes at least Therefore
For , all denominators appearing in Definition 50 are bounded away from zero; for example Substituting the displayed path probabilities of Definition 50, summing out the inactive potential outcome, and reducing the four cases gives and where , , and Since the treated outcome mass is the visible-cell displacement is Thus The same finite KL–chi-square bound, now applied with the KL divergence of Definition 53, and data processing under give
The latent effects on the path are so their order is fixed for . Also gives Using Definition 30, and hence
Apply the coordinate form of the two-point inequality to and to the ordered coordinate corresponding to the lower effect. The preceding bounds give and Therefore every measurable simplex-valued ordered-weight estimator has one of the two local laws satisfying The quotient-law certificate follows from the witness KL calculation with : and from The labelled-path membership, KL bound, tangent-amplitude statement, and ordered-mass identity were established above for every . All four asserted conclusions hold with the constants fixed at the start.
Apply Theorem 6 with local radius . It gives a constant such that, for every and every measurable quotient-law estimator , there is a probability law in the local quotient experiment satisfying By Definition 56, this law belongs to the fixed model class . Hence Taking the infimum over the same admissible estimators gives
For the upper bound, Theorem 2 applies at , , , and . Thus there is a constant such that, for every , there is a measurable structured-lattice estimator and a measurable repair estimator such that, for every and every , Set For , because , the logarithm is monotone on positive arguments, and . Therefore, with we have the uniform tail bound
The one-Wasserstein loss is nonnegative, so . Define Since , both and are positive. For every , put Then , and Substitution in the preceding tail bound gives Using the nonnegative-tail identity and the trivial bound for ,
Let The first term is positive, so , and . The estimator is admissible for the infimum, and the preceding display holds uniformly in . Hence, for every , Taking and gives and both displayed minimax inequalities.
For and , set Let be the class of measurable simplex-valued labelled weight estimators based on observed records in the specialized model and define
Apply Theorem 6 with . It gives a constant such that, for every , every , and every , there is a law satisfying By Definition 57, . Hence, for every admissible , and taking the infimum over yields
Apply Theorem 5 with It gives a constant such that, for every , there is an estimator satisfying for every and every . The target is the ordered aggregate mass vector of the quotient law from Definition 40: Definition 30 defines it as , and Lemma 17 makes this ordered vector depend only on the represented probability law. Therefore
Define Since and , so . Combining the bounds from the preceding two steps gives, for every and , This is the asserted minimax inequality.
Apply Theorem 6 with which lies in the domain of Definition 49. This gives constants , and an auxiliary positive Kullback–Leibler constant, such that and all conclusions of Theorem 6 hold for every . In the present specialization,
We record the published-scope implication used for the endpoint laws. Let for The restrictions in Definition 32 supply the qualitative proxy separations, causal consistency, armwise latent ignorability, anchor normalization, boundedness, latent-arm positivity, and proxy-rank margin. The proxy-rank margin gives so the three proxy feature matrices have full column rank. Latent-arm positivity gives Thus satisfies the qualitative structural conditions in Definition 60. The cited VMW model-scope correspondence, (Virk et al., 2026), identifies these qualitative conditions with membership in the published VMW model predicate attached to Definition 52.
It remains to verify VMW Assumption 4 in Definition 3. Put Choose a singular-value decomposition By Proposition 1, Let . Then the columns of are orthonormal and satisfy so is a population top-two right singular basis. Since , For , Hence each lies in the signal space generated by the two armwise row spaces. The factorization in Proposition 1 gives rank at most for this signal space, while gives rank at least . Therefore the span of is exactly the signal space. Applying the compressed-margin conclusion of Proposition 1 to this basis gives Together with these are the finite envelope, treatment-arm, top-right-basis, stacked-margin, and compressed-margin clauses of Definition 3, with Thus every such satisfies VMW Assumption 4.
Fix . The quotient-law certificate in Theorem 6, together with Definition 56, supplies probability laws belonging to with the fixed parameters above. By the implication from the preceding step, both endpoint laws belong to the published VMW model and satisfy VMW Assumption 4.
At the collision endpoint, Lemma 9 gives The cited recovery-regime correspondence, (Virk et al., 2026), identifies the separated recovery predicate attached to Definition 52, for , with the conjunction in Definition 61, including spectral separation of the two latent effects. The displayed equality violates that spectral-separation clause, so lies outside the published separated recovery regime.
Let satisfy the quotient-class hypotheses, and choose with For any estimator with support radius , the quotient-risk conclusion of Theorem 6 gives an endpoint such that If , use ; if , use . Equality as full-data laws identifies the observed product law and the quotient law, hence This is the quotient lower bound.
Fix in the gap-scale domain, so and set The labeled-path certificate in Theorem 6 supplies the tangent-amplitude condition and places in the local labeled-weight experiment . By Definition 57, both laws belong to the gap stratum , and therefore to with the fixed parameters. The implication from the first step gives published VMW model membership and VMW Assumption 4 for both endpoint laws.
Let be either or . Since , Definition 41 supplies the distinct-effects condition of Assumption 12. Latent-arm positivity gives Thus both latent classes have positive mass, and the distinct-effects condition yields This is the two-class spectral-separation condition in Definition 61.
Let satisfy the labeled-weight comparator-class hypotheses, and choose with For any ordered-weight estimator , the ordered-weight conclusion of Theorem 6 gives an endpoint such that If , use ; if , use . Equality as full-data laws preserves the observed product law and the ordered-mass target, hence This proves the labeled-weight lower bound.
Put By Theorem 1, choose such that, for any two model laws , Let and be the finite-library cardinality and work constants constructed below, and let be the constant supplied by Lemma 1. Define Then , and this single constant dominates the constants used for the library size, work bound, and probability bound.
Fix , and set The library is the deterministic half-open coordinate-grid library in the -coordinate summary box, with side length . Its coordinates are the entries of the four matrix blocks and the entries of the mean block. For each grid cube meeting the admissible image , choose one representative and order the representatives by the lexicographic order of the cube lower corners. The coordinate bounds following from Proposition 1 place every admissible summary coordinate in . If and lie in the same cube, then every scalar coordinate differs by at most , so for , and Using from Definition 33, Thus the stored representatives form an -net of ; when , the index set is empty and the program uses the fallback law in Algorithm 7.
Each coordinate lower corner has at most possible values. Since and , The lower-corner code is injective, hence After enlarging , the asserted cardinality bound follows.
Fix a uniform provider of the exact-real primitives. The estimator is the law component of in Algorithm 7. On a nonempty library, enumerate the finite index set as and update the current index by choosing the smaller pair in lexicographic order. For each fixed representative, is continuous by Lemma 23; therefore every update is Borel, being defined by finitely many weak and strict Borel inequalities. The empty selection is constant. Since is Borel by Definition 35, and the selected finite index is mapped to a stored spectral law, is Borel measurable.
The exact-real trace is assembled from the empirical-summary trace, the finite-library search trace, and the selected spectral run: Thus the returned law and the displayed trace come from the same execution.
The empirical-summary trace has arithmetic operations. Each library comparison uses scalar arithmetic operations, four fixed-dimensional singular-value or operator-norm operations, and one comparison. The selected spectral tail has three singular-value operations, one root-isolation operation, and one arithmetic operation. Therefore, for every sample , Combining this with the cardinality bound gives
If the program selects at sample , put The exhaustive fold preserves domination, where dominates at when and equality implies Induction over the scanned list yields, for every library index , and
Let be the exact spectral run at the selected representative . By the definition of the program output,
The same spectral execution supplies a basis , its signal spans, and an eigenvalue list . Its output fields give, for each , injective, the exact threshold characterization and completeness of the returned eigenvalue list: where is the compressed operator formed from , , and the run’s spans.
Fix any stored representative . Since , choose with Let , write , and set The latent-arm positivity condition in Assumption 9 gives for every . Let be the exact spectral run at , with signal basis , spans, anchors , returned eigenvalues , and compressed operator Set and Lemmas 4 and 18 give, for , The run’s armwise full-rank certificate makes injective for each . Taking , the first factorization shows that is injective and hence invertible. The same factorization also gives injectivity of : if for , then so injectivity of gives , and hence . The proxy factorization before postmultiplication by is Since is injective, its Moore–Penrose inverse is a left inverse on its range, and therefore Transposing gives By construction of the run’s span certificate, and consequently The second factorization can therefore be rewritten as Since the Moore–Penrose inverse is a left inverse on the range of an injective linear map, Subtracting the two treatment-arm identities gives Finally, the law of total expectation gives The column-space inclusion just proved implies that the orthogonal projector fixes . Since , Thus The target-proxy anchor condition gives , while and , so For the polynomial projector attached to a returned slot , the diagonalization gives Hence the returned slot weight satisfies for a spectral value and for a listed value outside the spectrum. The output validity of gives and . Completeness gives at least one returned slot for each latent-effect value, and the positivity of the corresponding cluster mass forces exactly one positive returned slot for each distinct latent-effect value; otherwise the total would exceed . Therefore, for every bounded test function , Thus the induced atomic measures agree: with coincident effects aggregated as in Definition 40. The exact spectral run at returns :
The validity certificate of the same result-bearing spectral execution states Consequently the returned weights form a probability vector and every returned atom lies in , so the represented finite atomic measure is a probability law supported on that interval.
For a stored representative , put Lemma 25 supplies the rank identity, and Lemma 24 supplies the singular-value margin: If satisfies Weyl’s singular-value inequality gives, for every , Thus while gives for , and hence Consequently the perturbed stacked moment satisfies the two threshold clauses Equivalently, exactly singular values of meet the threshold .
Fix and a sample . If , then is nonempty, so the library has a selected nearest index . By the covering property, choose such that Nearest-library optimality and the triangle inequality yield By Equations 3 and 4, Applying the modulus from Theorem 1 gives
Fix . Since , , and , Therefore On the event the deterministic inequality gives Thus the bad event in the theorem is contained in By Lemma 1, its -probability is at most .
References
- Brown, Lawrence D. and Purves, Roger (1973). Measurable Selections of Extrema. The Annals of Statistics. doi
- Virk, Hamza and Mazaheri, Bijan and Wu, Yihren (2026). The Spectral Structure of Latent Treatment Effects. . doi
- Mazaheri, Bijan and Squires, Chandler and Uhler, Caroline (2025). Synthetic Potential Outcomes and Causal Mixture Identifiability. Proceedings of The 28th International Conference on Artificial Intelligence and Statistics. doi
- Bauer, Friedrich L. and Fike, Charles T. (1960). Norms and Exclusion Theorems. Numerische Mathematik. doi
- Kato, Tosio (1995). Perturbation Theory for Linear Operators. Springer. doi
- Wu, Yihong and Yang, Pengkun (2020). Optimal Estimation of Gaussian Mixtures via Denoised Method of Moments. The Annals of Statistics. doi
- Heinrich, Philippe and Kahn, Jonas (2018). Strong Identifiability and Optimal Minimax Rates for Finite Mixture Estimation. The Annals of Statistics. doi
- Ho, Nhat and Nguyen, XuanLong (2016). Singularity Structures and Impacts on Parameter Estimation in Finite Mixtures of Distributions. . doi
- Polyanskiy, Yury and Wu, Yihong (2019). Dualizing {Le Cam}'s Method for Functional Estimation, with Applications to Estimating the Unseens. . doi
- Deo, Neil and Randrianarisoa, Thibault (2023). On Adaptive Confidence Sets for the Wasserstein Distances. Bernoulli. doi
- Robins, James M. and van der Vaart, Aad W. (2006). Adaptive Nonparametric Confidence Sets. The Annals of Statistics. doi
- Bing, Xin and Bunea, Florentina and Niles-Weed, Jonathan (2026). Estimation and Inference for the Wasserstein Distance Between Mixing Measures in Topic Models. Bernoulli. doi
- Miao, Wang and Geng, Zhi and Tchetgen Tchetgen, Eric J. (2018). Identifying Causal Effects with Proxy Variables of an Unmeasured Confounder. Biometrika. doi
- van Amsterdam, Wouter A. C. and Verhoeff, Joost J. C. and Harlianto, Netanja I. and Bartholomeus, Gijs A. and Puli, Anudeep M. and de Jong, Pim A. and Leiner, Tim and van Lindert, Alette S. R. and Eijkemans, Marinus J. C. and Ranganath, Rajesh (2022). Individual Treatment Effect Estimation in the Presence of Unobserved Confounding Using Proxies: A Cohort Study in Stage {III} Non-Small Cell Lung Cancer. Scientific Reports. doi
- Imbens, Guido W. and Rubin, Donald B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press.
- Splawa-Neyman, Jerzy and Dabrowska, D. M. and Speed, T. P. (1990). On the Application of Probability Theory to Agricultural Experiments: Essay on Principles. Section 9. Statistical Science. doi
- Rubin, Donald B. (1974). Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies. Journal of Educational Psychology. doi
- Holland, Paul W. (1986). Statistics and Causal Inference. Journal of the American Statistical Association. doi
- Hansen, Lars Peter (1982). Large Sample Properties of Generalized Method of Moments Estimators. Econometrica. doi
- Cam, Lucien Le (1986). Asymptotic Methods in Statistical Decision Theory. Springer. doi
- Pearson, Karl (1894). III. Contributions to the Mathematical Theory of Evolution. Philosophical Transactions of the Royal Society of London A. doi
- Teicher, Henry (1963). Identifiability of Finite Mixtures. The Annals of Mathematical Statistics. doi
- Lindsay, Bruce G. (1995). Mixture Models: Theory, Geometry, and Applications. Institute of Mathematical Statistics and American Statistical Association.
- Allman, Elizabeth S. and Matias, Catherine and Rhodes, John A. (2009). Identifiability of Parameters in Latent Structure Models with Many Observed Variables. The Annals of Statistics. doi
- Hu, Yingyao (2008). Identification and Estimation of Nonlinear Models with Misclassification Error Using Instrumental Variables: A General Solution. Journal of Econometrics. doi
- Hu, Yingyao and Schennach, Susanne M. (2008). Instrumental Variable Treatment of Nonclassical Measurement Error Models. Econometrica. doi
- Kasahara, Hiroyuki and Shimotsu, Katsumi (2009). Nonparametric Identification of Finite Mixture Models of Dynamic Discrete Choices. Econometrica. doi
- Bonhomme, St{\'e}phane and Jochmans, Koen and Robin, Jean-Marc (2016). Non-Parametric Estimation of Finite Mixtures from Repeated Measurements. Journal of the Royal Statistical Society: Series B. doi
- Anandkumar, Animashree and Ge, Rong and Hsu, Daniel and Kakade, Sham M. and Telgarsky, Matus (2014). Tensor Decompositions for Learning Latent Variable Models. Journal of Machine Learning Research. arXiv
- Hsu, Daniel and Kakade, Sham M. and Zhang, Tong (2009). A Spectral Algorithm for Learning Hidden Markov Models. Proceedings of the 22nd Annual Conference on Learning Theory. arXiv
- Sidiropoulos, Nicholas D. and Bro, Rasmus (2000). On the Uniqueness of Multilinear Decomposition of {N}-way Arrays. Journal of Chemometrics. doi
- Nguyen, XuanLong (2013). Convergence of Latent Mixing Measures in Finite and Infinite Mixture Models. The Annals of Statistics. doi
- Ho, Nhat and Nguyen, XuanLong (2016). Convergence Rates of Parameter Estimation for Some Weakly Identifiable Finite Mixtures. The Annals of Statistics. doi
- Manole, Tudor and Ho, Nhat (2022). Refined Convergence Rates for Maximum Likelihood Estimation under Finite Mixture Models. Proceedings of the 39th International Conference on Machine Learning. arXiv
- Doss, Natalie and Wu, Yihong and Yang, Pengkun and Zhou, Harrison H. (2023). Optimal Estimation of High-Dimensional Gaussian Location Mixtures. The Annals of Statistics. doi
- Niles-Weed, Jonathan and Rigollet, Philippe (2022). Estimation of Wasserstein Distances in the Spiked Transport Model. Bernoulli. doi
- Nguyen, Huy and Le, Dung and Rinaldo, Alessandro and Ho, Nhat (2026). On the Geometry of Separation in Finite Gaussian Mixtures. . doi
- Davis, Chandler and Kahan, William M. (1970). The Rotation of Eigenvectors by a Perturbation. {III}. SIAM Journal on Numerical Analysis. doi
- Stewart, G. W. and Sun, Ji-guang (1990). Matrix Perturbation Theory. Academic Press.
- Shi, Xu and Miao, Wang and Nelson, Jennifer C. and Tchetgen Tchetgen, Eric J. (2020). Multiply Robust Causal Inference with Double-Negative Control Adjustment for Categorical Unmeasured Confounding. Journal of the Royal Statistical Society: Series B. doi
- Tchetgen Tchetgen, Eric J. and Ying, Andrew and Cui, Yifan and Shi, Xu and Miao, Wang (2024). An Introduction to Proximal Causal Inference. Statistical Science. doi
- Cui, Yifan and Pu, Hongming and Shi, Xu and Miao, Wang and Tchetgen Tchetgen, Eric J. (2024). Semiparametric Proximal Causal Inference. Journal of the American Statistical Association. doi
- Miao, Wang and Hu, Wenjie and Ogburn, Elizabeth L. and Zhou, Xiaohua (2023). Identifying Effects of Multiple Treatments in the Presence of Unmeasured Confounding. Journal of the American Statistical Association. doi
- Qi, Zhengling and Miao, Rui and Zhang, Xiaoke (2024). Proximal Learning for Individualized Treatment Regimes Under Unmeasured Confounding. Journal of the American Statistical Association. doi
- Liu, Jiewen and Park, Chan and Li, Kendrick and Tchetgen Tchetgen, Eric J. (2024). Regression-Based Proximal Causal Inference. . doi
- Li, Kendrick and Linderman, George C. and Shi, Xu and Tchetgen Tchetgen, Eric J. (2024). Regression-Based Proximal Causal Inference for Right-Censored Time-to-Event Data. . doi
- Ai, Chunrong and Shan, Jiawei (2025). Efficient Estimation of Average Treatment Effects with Unmeasured Confounding and Proxies. . doi
- Saha, Aytijhya and Bates, Stephen and Shah, Devavrat (2026). Causal Inference with Categorical Unobserved Confounder via Mixture Learning. . doi
Comments on earlier versions
Anchored to: