theorem its proof invokes invoked by

CausalSmith · AI Causal Scientist — eid_crl_coverratio_mmd_genericity_v1 · Exact ID · AI reviewer score 6.6/10 · pinned commit f2e7e97 · PDF · Lean code · Slides · arXiv · GitHub

One Intervention per Latent Variable: Generic Identification of Nonlinear Causal Representations on Fixed-sign Compact Strata

Abstract

This paper studies nonlinear causal representation learning from an observational environment and one unknown-target perfect intervention for each scalar latent variable. In a positive compact-cube latent model with shared diffeomorphic mixing, causal minimality, and fixed own-coordinate derivative signs, likelihood ratios between interventional and observational laws define observable one-dimensional ratio coordinates. Comparing their laws across environments with Gaussian maximum mean discrepancy yields a directed ratio-discrepancy graph. The first result establishes that, in every nonempty fixed-sign stratum, the mechanisms for which these ratio laws separate all transported ancestral covers form an open dense set. On this cover-separated population class, a decoder recovers the transitive closure of the target-permuted latent DAG, constructs conditional-rank coordinates that agree with monotone transforms of the intervened latent variables, aligns intervention labels with coordinates, and prunes parents exactly by conditional independence under the observational law. Within the same structural class, equality of the observed environment-law family identifies the observed world up to componentwise coordinate changes and simultaneous relabeling of graph and targets. A sample-split procedure adds familywise-valid confidence edges for ancestral discoveries, conditional on a simultaneous first-stage likelihood-ratio error event. On the overlapping smooth compact-support faithful regime, the result gives an arbitrary-dimensional one-perfect-intervention-per-node counterpart to the bivariate law-separation route of von Kügelgen et al. (2023).

Introduction

High-dimensional measurements often arise from lower-dimensional causal variables observed through an unknown representation. Causal representation learning asks when distributional variation across environments identifies those variables, their causal graph, and the intervention labels that generated the variation (Schölkopf et al., 2021). In nonlinear settings with invertible mixing, this question joins three traditions: graphical causal discovery from conditional independences (Pearl, 2009; Spirtes et al., 2000; Peters et al., 2017), identifiable nonlinear independent component analysis from distributional changes (Hyvärinen et al., 2016; Hyvärinen et al., 2019; Khemakhem et al., 2020; Hyvärinen et al., 2023), and interventional discovery when targets are known, unknown, or partially specified (Hauser et al., 2012; Yang et al., 2018; Jaber et al., 2020; Dai et al., 2025; Zhang et al., 2026).

The setting here is an observational law and interventional laws generated by one perfect intervention for each of scalar latent coordinates. The target labels are permuted, so the analyst observes environment laws before recovering the alignment between environments and latent variables. The latent variables live on the compact cube, the observation map is shared and diffeomorphic across environments, the local mechanisms are positive and smooth, and each intervention likelihood ratio has a fixed own-coordinate derivative sign. Under these conditions, the likelihood ratio , the density ratio for environment against the observational law, is an observable scalar function whose latent form depends on the intervened coordinate and its parents.

The paper’s central object is the law of one ratio coordinate under another environment. For each ordered pair , the Gaussian maximum mean discrepancy , the kernel discrepancy between the laws of under environments and , measures whether intervention changes the distribution of ratio coordinate . Theorem 2 shows that, within every nonempty positive smooth causal-minimal fixed-sign stratum, the class on which these discrepancies are positive for all transported ancestral covers is open and dense. The proof uses explicit regular witnesses from Theorem 1 and analytic perturbations along normalized affine paths to move edge-specific second-moment contrasts away from cancellation.

On that generic population class, Theorem 3 gives an exact decoder. Starting from the environment laws, , the population decoder, forms the ratio-discrepancy graph, chooses a topological order, constructs conditional distribution transforms of log-ratios, aligns each environment with the coordinate recovered by its own rank, and prunes candidate predecessors using conditional independence under the observational law. The ratio graph recovers the transitive closure of , the target-permuted DAG; the conditional ranks equal , the intervention distribution transform of the target latent coordinate, or its sign-reversed version; and causal minimality makes the true parent set the unique inclusion-minimal admissible set. Thus the output graph is exactly , the latent DAG expressed on environment labels.

The result contributes to the unknown-intervention literature by using likelihood-ratio laws in place of scores, support geometry, quadratic precision structure, or supplied target metadata. The closest invertible-mixing comparator is von Kügelgen et al. (2023), who establish a bivariate one-intervention result under a continuous law-separation condition and an arbitrary-dimensional result with two paired perfect interventions per node. On the overlapping positive smooth compact-support faithful regime, Theorem 3 supplies an arbitrary-dimensional one-perfect-intervention-per-node population recovery statement and gives the Gaussian-MMD law-separation witness in the two-variable case. Relative to target-aligned approaches such as Wendong et al. (2023) and Yao et al. (2025), the ratio construction recovers the graph, order, and target alignment from the same observed law family.

The statistical layer translates population discrepancies into finite-sample edge statements. Theorem 4 studies a sample-split design: independent training folds produce fitted positive likelihood ratios, and independent evaluation folds compute empirical Gaussian-MMD discrepancies. Given a simultaneous first-stage ratio-error radius , the theorem provides a familywise event on which every selected edge in the confidence graph has an ancestral interpretation in the latent DAG. When every transported ancestral-cover discrepancy exceeds twice the simultaneous radius, the selected graph has the same transitive closure as . This connects the ratio decoder to kernel two-sample concentration (Gretton et al., 2012; Schrab et al., 2023) while keeping the first-stage ratio-estimation requirement explicit.

The final section formulates the next coordinate-error problem in a bounded Hölder regime. Under smoothness, lower-envelope, own-derivative margin, shared-mixing geometry, observed log-ratio regularity, and conditional-design assumptions, Remark 2 states a generated-rank rate target for conditional-CDF estimation with fitted log-ratio responses and fitted conditioning covariates. The target combines second-stage smoothing bias, local empirical variation, and first-stage error propagated through bandwidth neighborhoods, drawing on conditional-CDF and generated-regressor theory (Xie, 2023; Mammen et al., 2012; Lee, 2018).

The whole argument rests on two observations that the rest of the paper makes exact. Under the stated positivity, smoothness, causal-minimality, shared-diffeomorphic-mixing, one-perfect-intervention, fixed-sign, and cover-separation conditions, the ratios , the density ratios for interventional environments against the observational law, support ancestry recovery through population discrepancies, and the conditional-rank construction returns a monotone transform of the corresponding latent coordinate. Together with causal minimality, shared diffeomorphic mixing, and the one-perfect-intervention-per-node setup, positivity, smoothness, the sign condition, and cover separation make those two steps exact; the bookkeeping objects introduced below state these conditions precisely. A reader who wants the argument before its apparatus can read Algorithm 1 and Theorem 3 first and consult Section 3 for the objects they use. The formal layer accompanying the paper is machine-checked, with precise scope recorded in Section A.

Section 2 situates the paper in graphical causal inference, nonlinear ICA, causal representation learning, unknown-target interventions, and kernel testing. Section 3 introduces the positive compact-cube model, observed worlds, likelihood ratios, and ratio-law discrepancies. Section 4 proves generic ratio separation. Section 5 gives the population decoder and its exact recovery theorem. Section 6 establishes sample-split confidence edges, and Section 7 records the generated-rank frontier and further directions.

Related work

Graphical causal inference supplies the language for the paper’s target object: a latent causal model on scalar variables, indexed by the labeled vertex set and a finite DAG . Classical constraint-based discovery identifies Markov equivalence classes from conditional independence information under faithfulness-type restrictions (Meek, 1995). Recent work sharpens the role of typical or generic faithfulness in graphical models (Boeken et al., 2024). The present analysis uses that tradition in a different observational object: ratios across environments generated by one intervention per latent coordinate. The resulting separation statements characterize when the ratio laws distinguish the transported graph, rather than when conditional independences alone determine a Markov equivalence class.

Identifiable nonlinear ICA provides a second point of contact. The modern line from temporal, auxiliary-variable, and exponential-family restrictions establishes identifiability of latent coordinates through variation in observed distributions (Hyvärinen et al., 2001; Hyvärinen et al., 2016; Hyvärinen et al., 2019; Khemakhem et al., 2020; Hyvärinen et al., 2023). Causal representation learning extends this idea by connecting identifiable coordinates to causal variables and interventions (Brehmer et al., 2022; Lippe et al., 2022; Squires et al., 2022; Zhang et al., 2023). Here the source of variation is an observed environment family generated by a shared mixing map , with observed laws obtained from single-target interventions whose labels are permuted by . The shared invertible observation map keeps the ratio law invariantly comparable across environments, while the unknown target alignment makes the graph recovery problem genuinely interventional.

The closest comparator in the invertible-mixing and unknown-target direction is von Kügelgen et al. (2023). On the overlapping smooth, compact-support, faithful regime, both analyses exploit multiple environments whose targets are initially unlabeled. von Kügelgen et al. (2023) settles the bivariate case from one perfect intervention per node under a continuous law-separation condition, and von Kügelgen et al. (2023) reaches arbitrary dimension from two paired perfect interventions per node, with the one-intervention extension beyond two latent variables stated there as a conjecture. The present paper closes that count in a smaller model class rather than in theirs: on a positive compact latent cube with mixing, fixed own-derivative signs, and causal minimality, one perfect intervention per node suffices in every dimension, and the separation condition that makes this work is proved generic rather than assumed. The two results overlap rather than nest: their model class is the full Euclidean latent space at regularity, while the class used here is the compact cube at , and the comparison therefore holds on the intersection. Inside that intersection, positive Gaussian-kernel discrepancy supplies the bivariate law-separation witness that von Kügelgen et al. (2023) requires. The present paper characterizes a complementary ratio-law route: likelihood ratios and log-ratios produce observable discrepancies that separate transported ancestral covers and then support population decoding of the intervention labels and parent-pruned graph.

Several recent papers study interventional causal discovery with unknown or partially known targets. Hauser et al. (2012) analyzes interventional Markov equivalence under known targets, Yang et al. (2018) studies general interventional equivalence classes, and Jaber et al. (2020), Dai et al. (2025), and Zhang et al. (2026) develop graphical and algorithmic theory for target uncertainty. Related latent-variable and representation-learning approaches include Kivva et al. (2021), Xie et al. (2024), Jiang et al. (2023), Buchholz et al. (2023), Ahuja et al. (2023), Varıcı et al. (2024), Varıcı et al. (2025), Bing et al. (2024), and Ng et al. (2025). Score-based analyses of the same design carry the closest published intervention counts. Varıcı et al. (2025) identify latent variables from score differences across interventional environments and show that one stochastic hard intervention per node suffices under linear transformations, while two per node suffice under general transformations; Varıcı et al. (2024) develop the general-identifiability side of that program. The present result keeps one perfect intervention per node under a nonlinear diffeomorphic transformation, at the price of the compact positive fixed-sign class and a genericity condition, and it works with likelihood-ratio laws and their kernel embeddings rather than with score functions. Within this literature, Wendong et al. (2023) and Yao et al. (2025) are especially close in their use of interventional or invariant structure for causal representations. The comparison turns on inputs: those approaches use graph, order, or target-alignment information appropriate to their identification strategies, whereas the ratio construction starts from the environment laws and recovers an observable directed discrepancy graph from pairwise ratio variation.

The statistical layer draws on kernel mean embeddings and maximum mean discrepancy. The Gaussian-kernel discrepancy construction uses the bounded kernel to compare one-dimensional ratio laws, following the embedding and two-sample testing literature (Sriperumbudur et al., 2010; Sriperumbudur et al., 2011; Gretton et al., 2007; Smola et al., 2007; Gretton et al., 2012; Schrab et al., 2023). For the finite-sample confidence graph, sample splitting and independent environment samples align the empirical discrepancy analysis with standard concentration arguments for kernel averages. The generated-rank discussion uses tools for conditional distribution estimation and generated regressors, especially local conditional CDF theory and partial-mean arguments with first-stage inputs (Xie, 2023; Mammen et al., 2012; Lee, 2018). Recent sample-complexity work for causal representation learning and few-environment recovery, including Acartürk et al. (2024) and Lee et al. (2026), provides the surrounding statistical context for translating population separation into finite-sample procedures. The confidence-edge guarantee proved here is a familywise statement about which edges are selected, valid on an event whose first-stage component is an assumed uniform ratio-error contract. Remark 2 states, as a target for later work, the rate a conditional-rank coordinate estimator would have to meet, expressed through an abstract first-stage rate rather than through any particular estimator.

Setup and ratio observables

Throughout the paper, is the ambient number of scalar latent variables, is the labeled vertex set, and is a finite labeled DAG. Vectors are written componentwise, with latent coordinates taking values in the compact cube . Generic constants may change from line to line, norms are taken in the space indicated by their subscripts, and all index ranges are over unless stated otherwise. The sign vector records the prescribed own-coordinate derivative directions.

The structural model is specified by observational conditional densities , parent-independent intervention densities , a shared observation map , and observed environment laws on the observed support . The first group of conditions fixes the smooth positive latent model and the common invertible observation system.

Assumption 1 [ass:positive-normalized-smooth-mechanisms] (Positive smooth mechanisms).

For every , the observational mechanism and the intervention density are strictly positive, normalized densities on their respective closed cubes.

⊢ Lean

The smooth full-support compact-cube density condition in Assumption 1 places the observational and intervention mechanisms on a common regular support for the ratio construction in this paper. Positivity makes likelihood ratios well-defined on the latent cube, while normalization preserves the interpretation of each carrier as a conditional or marginal density.

Assumption 2 [ass:shared-diffeomorphic-mixing] (Shared diffeomorphic mixing).

The same map is a diffeomorphism from onto in every environment.

⊢ Lean

The shared diffeomorphic mixing condition in Assumption 2 is the standard invertible-observation restriction for this setting (von Kügelgen et al., 2023). It ensures that every environment is observed through the same smooth coordinate change, so differences across come from the latent intervention laws rather than from environment-specific measurement maps.

We next make the graph notation and latent product law explicit.

Definition 1 [synth_9] (Parent set ).

For a finite labeled directed acyclic graph on and a vertex , denotes the set of parents of in : For , write .

The notation in Definition 1 is used throughout to express local Markov factorization through parent coordinates. In particular, depends on its own coordinate and the parent set .

Definition 2 [synth_2] (Observational product density).

For a finite dimension , a labeled DAG on , a mechanism with observational conditional-density carriers , and a latent state , define

⊢ Lean

The product density in Definition 2 is the observational latent law associated with and the local mechanisms. The target- intervention law replaces the th observational mechanism by , and this replacement is formalized in the observed laws below.

Assumption 3 [ass:one-perfect-intervention-per-node] (Perfect intervention laws).

For the mechanism , the shared mixing map , the observed laws , the target permutation , and the supplied likelihood ratios , the following conditions hold:

  • (Observational pushforward.) The observational law satisfies with .

  • (Single-target pushforwards.) For every environment ,

  • (Radon–Nikodym ratios.) For every ,

  • (Latent ratio formula.) For every and every ,

⊢ Lean

The unknown-target stochastic perfect-intervention coverage condition in Assumption 3 is standard in this literature (von Kügelgen et al., 2023). It ties each nonobservational environment to exactly one latent target through the permutation , and it supplies the likelihood-ratio representation that turns observed laws into one-dimensional ratio coordinates. When , the ratio for an intervention targeting depends on and its parents through the displayed quotient.

Assumption 4 [ass:fixed-own-derivative-sign] (Fixed derivative sign).

For every , throughout the closed cube.

⊢ Lean

Assumption 4 orients each ratio in its own coordinate through a compact-support fixed-sign derivative condition. It records that the log-ratio changes monotonically in the intervened coordinate after applying the prescribed sign , which later supports the rank-based decoding step.

Assumption 5 [ass:causal-minimality] (Causal minimality).

For every edge of , the observational law satisfies

⊢ Lean

The local edge form of causal minimality in Assumption 5 is the standard positivity-compatible requirement that each displayed parent contributes to the observational conditional law (Spirtes et al., 2000). It anchors the graph to conditional distributions under , so the formal stratum below consists of mechanisms whose graph edges are statistically active.

Definition 3 [def:model-stratum] (Model stratum ).

The set is equipped with the relative product topology.

⊢ Lean

The model stratum collects the latent mechanisms satisfying the positivity, minimality, and signed-ratio restrictions for the fixed graph and sign vector. The relative product topology is the topology used later for generic separation statements.

The remaining definitions describe the observed data object and the two population separation classes used by the ratio-decoding argument.

Definition 4 [synth_19] (Observed world data).

For a mechanism on the labeled DAG with ambient vertex set , an observed world is the following data: Here and are maps from the latent state space to itself, is a permutation of , is a measure on for each environment , and is a real-valued ratio carrier for each target .

⊢ Lean

Definition 4 packages the objects that are available to population procedures: common coordinates, environment laws, target labels, and ratio carriers. In the canonical observed world used below, the mixing and unmixing maps are identities, so the same definitions can be evaluated directly on latent laws.

Definition 5 [def:edge-separated-set] (Edge separation set).

For a target-label permutation , define Here is the second-moment contrast in the canonical observed world generated by and :

⊢ Lean

The set records mechanisms for which every direct edge has a nonzero canonical second-moment contrast. The contrast compares the squared ratio coordinate under the observational law and under another intervention environment after the target relabeling .

Definition 6 [synth_13] (Population MMD discrepancy).

For a complete real Hilbert space , let be a feature map with and For an observed world with environment laws and ratio coordinates , and for , define the population MMD discrepancy by When the observed world is fixed by the context we abbreviate to .

⊢ Lean

The discrepancy in Definition 6 is the maximum mean discrepancy between the one-dimensional laws of ratio coordinate under environments and , using the bounded Gaussian kernel and its Hilbert-space feature map. This embeds the ratio-law comparison in the standard kernel two-sample framework (Gretton et al., 2012; Sriperumbudur et al., 2010; Sriperumbudur et al., 2011).

Definition 7 [def:cover-separated-set] (Cover-separated mechanisms).

Let be a complete real Hilbert space, and let be a feature map satisfying for all . For a target permutation , let be the canonical observed world generated by and : the mixing and unmixing maps are the identity, the target permutation is , the observational environment has law induced by ’s observational law, intervention environment has law induced by the intervention targeting , and its ratio carrier is The cover-separated set is where is the law of ratio under the observational environment of , and is the law of ratio under intervention environment of .

⊢ Lean

The cover relation selects ancestral covers in the latent DAG, and Definition 7 requires a positive ratio-law discrepancy on every such transported cover after applying . Thus is the population class on which the observable discrepancy graph contains the cover information needed by the decoder developed in the next section.

Generic ratio separation

The separation sets in Definitions 5 and 7 turn observable ratio laws into graph information. This section supplies the two ingredients behind that step. First, explicit three-node mechanisms show both a strict direct-edge discrepancy and an exact cancellation. Second, the genericity result shows that, within any nonempty fixed-sign stratum, the separating behavior is open and dense.

We begin with two concrete mechanisms on the graph , with node isolated. The first mechanism is the separating endpoint used to certify that the relevant analytic contrast can be nonzero.

Definition 8 [def:sparse-witness] (Sparse witness ).

On the graph with node isolated, define by and For a sign vector with a negative prescribed sign in coordinate , the corresponding coordinate is reflected by the map .

⊢ Lean

The affine dependence in Definition 8 creates a direct response of the ratio coordinate at node to interventions at node . Reflections preserve the fixed-sign convention from Assumption 4, so the same construction serves every prescribed sign pattern.

The next mechanism has the same graph and the same intervention tilts, but it places the parent dependence in a form that exactly balances the ratio law across the relevant environments.

Definition 9 [def:cancellation-witness] (Cancellation witness ).

On the graph with node isolated, define by For any other sign vector, is obtained from this mechanism by coordinate reflections.

⊢ Lean

The contrast between Definitions 8 and 9 is useful because the ratio-law condition can hold strictly at one regular point and vanish at another regular point. Thus the generic result below is a topological statement about typical mechanisms in the stratum, with the cancellation example marking the boundary behavior that analytic perturbations must move away from.

For the witness calculation and the generic statement, it is helpful to name the observed law system induced by a mechanism and a shared coordinate map.

Definition 10 [synth_28] (Coherence of observed laws with an observed world).

Let be an observed world generated by a mechanism on , with shared mixing map , inverse , and target permutation . We say that observed laws are when and, for each interventional environment label , , where and Equivalently, under and , the unmixed variable has densities and , respectively, on .

Definition 10 records the population object available to the ratio analysis: the same observation map pushes forward the observational law and each single-target intervention law. In the explicit witness world, the shared map and target permutation are identities, so the latent and observed densities coincide.

Definition 11 [synth_27] (Sparse-witness environment laws).

Let be the observed world generated by the sparse witness mechanism with identity mixing and identity target permutation on the three-node DAG with node isolated. For , denotes the environment- law in this world. Thus is the observational law with density and, for , is the intervention law at target , with density Because the mixing map is the identity, these latent densities are also the corresponding observed laws.

The graph-theoretic and probabilistic conventions used by the generic result are standard. We state them here to make the separation theorem self-contained.

Definition 12 [synth_24] (d-separation in a DAG).

Let be a DAG on , and let be pairwise disjoint. A path in the underlying undirected graph of is -active when every non-collider on the path is outside , and every collider on the path belongs to or has a directed path in to some vertex of . The sets and are d-separated by in when there is no -active path in the underlying undirected graph of with one endpoint in and the other endpoint in .

Definition 13 [synth_10] (Faithfulness to ).

An observational law on is faithful to the directed acyclic graph when, for all disjoint subsets , Equivalently, the conditional independence relations of are exactly the d-separation relations encoded by .

Definition 14 [synth_23] (Ancestral order in a DAG).

For a DAG on the finite labeled vertex set and vertices , write when is a strict ancestor of in , equivalently when there exists a directed path in of positive length from to .

Definition 15 [synth_4] (Ancestral cover relation).

For a DAG on the finite labeled vertex set , define to be the cover relation in the ancestor order of :

⊢ Lean

The cover relation in Definition 15 is the transitive-reduction analogue for the ancestor order. It selects the ancestral comparisons whose ratio-law discrepancies form the population input for the decoder in the next section.

We also record the ratio-law notation used in the witness certificate and in the open-dense theorem.

Definition 16 [synth_1] (Latent state space).

For , the latent state ranges over the real coordinate space

⊢ Lean
Definition 17 [synth_20] (Observational, intervention, and ratio laws).

For a mechanism , let denote the observational law on the observed support generated by and the shared mixing map . For , let denote the perfect-intervention law obtained from by replacing the -th observational conditional mechanism with the parent-independent intervention density , so its latent density is again pushed forward by to . Given a target permutation , write and, for , . For , let be the observed density ratio for environment label . For , define the probability law on obtained by pushing the canonical environment- observed law forward through the ratio coordinate .

Definition 18 [synth_3] (Intervention distribution function).

For a mechanism on a DAG over , with parent-independent intervention density at coordinate , define for and .

⊢ Lean

One further condition is recorded here for use by the explicit witness and by the comparison with prior work: the usual faithfulness condition, linking conditional independences in the observational density to d-separation in the graph. It is a strictly stronger requirement than the causal minimality used by Theorem 2 and Theorem 3, which quantify over the causal-minimal stratum; faithfulness enters only where it is cited by name.

Assumption 6 [ass:faithfulness] (Faithful observational law).

The observational latent density is faithful to the DAG .

⊢ Lean

Assumption 6 is the standard causal faithfulness condition (Spirtes et al., 2000). It rules the observational conditional independences to be exactly those encoded by d-separation, so graphical ancestry and the statistical independences used by causal minimality are aligned.

The explicit witnesses now certify the two endpoint behaviors needed for the perturbation argument.

Theorem 1 [prop:sparse-witness-certificate] (Sparse witness certificate).

For every sign vector , let be the construction in Definition 8 and let be the construction in Definition 9, both on the three-node DAG with node isolated. Let and be the corresponding observed worlds with identity mixing and identity target permutation, so that environment has the observational law, environment has the intervention law at target , and Then:

  • (Sparse witness regularity.) The mechanism satisfies Assumption 1, Assumption 4 with sign vector , and Assumption 6. Moreover, every observational conditional density is on , and every intervention density is on .

  • (Cancellation witness regularity.) The mechanism satisfies Assumption 1, Assumption 4 with sign vector , and Assumption 6. Moreover, every observational conditional density is on , and every intervention density is on .

  • (Sparse witness separation.) The sparse witness has and its Gaussian-kernel population discrepancy satisfies

  • (Cancellation witness equality.) For the cancellation witness, the law of under the observational environment equals the law of under environment , and

⊢ Lean

Theorem 1 gives a regular separating point and a regular cancelling point inside the same elementary graph geometry. The strict lower bounds for the sparse witness show that the direct-edge second-moment contrast and the Gaussian-kernel discrepancy can both detect the edge . The cancellation equality shows that equality of ratio laws can occur at structured mechanisms satisfying the same smoothness, sign, and faithfulness requirements, which is exactly the situation addressed by a generic rather than universal separation theorem.

The main population result converts that endpoint calculation into a statement for every nonempty fixed-sign stratum. Its topology is the relative product topology built into Definition 3.

Theorem 2 [thm:generic-cover-separation] (Generic cover separation).

Let be the ambient number of scalar latent variables. Assume:

  • (Finite labeled DAG.) is a finite labeled DAG on .

  • (Sign stratum.) is a sign vector with coordinates in , and the stratum of Definition 3 is nonempty.

  • (Target permutation.) is a permutation of .

Let be the direct-edge separation set of Definition 5, and let be the ancestral-cover separation set of Definition 7 formed with the Gaussian feature map for . Then is open and dense in , and Moreover, is open and dense in , and its complement in is meagre, closed, and nowhere dense. If has no directed edges, then

⊢ Lean

Theorem 2 establishes that ratio-law separation is typical in every nonempty stratum with fixed own-coordinate derivative signs. Openness says that a separated mechanism remains separated under sufficiently small perturbations of the normalized mechanisms. Density says that every admissible mechanism can be approximated by mechanisms whose direct-edge contrasts are nonzero and whose ancestral-cover discrepancies are positive. The meagreness and nowhere-density clauses give the corresponding topological description of the exceptional set.

The intuition is that the witness in Theorem 1 provides a local direction in which an edge-specific ratio contrast becomes nonzero. Along normalized analytic perturbation paths, a contrast that is nonzero at one endpoint can vanish only on a thin set of perturbation values unless it is identically zero. Applying this reasoning across the finitely many relevant edges and covers yields the open-dense separation class used by the population decoder in the next section.

Population decoding

The generic separation result in Theorem 2 supplies the observable input for a population construction. The construction is an oracle on the environment laws: its steps evaluate Radon–Nikodym derivatives, conditional-law versions, and conditional independences exactly, so it is a map from a law family to a graph and a set of coordinates rather than an estimator or an algorithm on data. Its order and pruning steps are well posed precisely on the cover-separated set, where Theorem 3 shows the ratio graph is acyclic and the admissible parent set is unique; on an arbitrary law family the ratio graph may contain cycles and the inclusion-minimal admissible set need not be unique, and the conclusions below are asserted only under the theorem’s hypotheses. Starting from the environment laws, the decoder forms the directed ratio-discrepancy graph, uses its order information to compute conditional distribution transforms, aligns each intervention label with the coordinate recovered by its own rank, and then prunes the candidate predecessors through conditional independence under the observational law. The output is an environment-label graph, so we first record the relabeling convention that translates the latent DAG into intervention labels.

Definition 19 [synth_17] (Permuted environment graph).

For an observed world with target permutation , the environment-label DAG is the directed graph on whose edge relation is

⊢ Lean

This graph is the structural object visible after intervention labels have been matched to latent coordinates. Because the target permutation is part of the observed-world description rather than known to the analyst, the decoder is formulated entirely in terms of laws and ratios.

The rank step requires a version of the conditional distribution function selected from observed laws. The next definition fixes that selection in a way that is compatible with regular conditional distributions and continuous support behavior.

Definition 20 [synth_18] (Selected conditional CDF).

For an observed probability-law family , an ordering, and a vertex , let be the predecessors of in that ordering. The selected conditional distribution function is defined as follows. If there exists a map such that

  • (Joint measurability.) is measurable as a function of ;

  • (Fixed-threshold version.) for every threshold , the function agrees almost surely, under the observed target- law of , with the raw regular conditional CDF of given ;

  • (Joint-law version.) the function agrees almost surely, under the observed target- joint law of , with ;

  • (Support continuity.) is continuous on times the support of the observed target- law of ,

then is one such map selected from the observed laws alone. In the complementary case, where is the raw regular conditional CDF of given under the observed target- law.

⊢ Lean

The continuity clause gives a population version of the probability integral transform along the conditional-ratio support. With that convention in place, the decoder can be stated as an observable map from environment laws to a pruned environment-label DAG.

Algorithm 1 [def:population-decoder] (Population decoder ).

Given observed laws , the population decoder is defined by the following steps.

  1. Compute the likelihood ratios and log-ratios and the Gaussian-MMD discrepancies .

  2. Form the observable directed ratio-discrepancy graph

  3. Choose a topological ordering of . For each , let be the predecessors of in that ordering and define

  4. Align environment with coordinate .

  5. For each , range over and identify the unique inclusion-minimal subset satisfying The output is the environment-label DAG

⊢ Lean

The first two steps use the law discrepancies studied in Theorem 2; the third step converts log-ratios into coordinates through conditional ranks; the final step applies the Markov factorization logic under , in the spirit of graphical-model recovery from conditional independences (Pearl, 2009; Spirtes et al., 2000). The rank transform has a structural target: conditionally on the recovered predecessors, the log-ratio distribution is the intervention-target distribution induced by the latent mechanism.

Definition 21 [synth_26] (Structural conditional-ratio CDF).

For a mechanism , an environment label , and target , write The structural conditional-ratio CDF at target is Equivalently, is the conditional distribution function of the target log-ratio in the target- intervention environment, conditional on the latent parent vector .

The structural CDF describes the rank that would be computed with latent parents in hand. The theorem below states that the observed conditional CDF selected in Definition 20 agrees with this structural object on the relevant support, so the rank coordinates produced by Algorithm 1 coincide with monotone transforms of the intervention-target latent coordinates.

Representation recovery is stated up to componentwise coordinate changes and a simultaneous relabeling of the DAG and targets. The equivalence relation records exactly those transformations.

Definition 22 [synth_8] (Componentwise equivalence).

For , let be an observed world over with mixing map , inverse-coordinate map , target permutation , and observed laws . With the same ambient number of scalar latent variables, write The worlds and are componentwise -equivalent when the following conditions hold:

  • (Observed support.) Their observed supports agree:

  • (DAG relabeling.) There is a permutation of such that, for every ,

  • (Target alignment.) For every ,

  • (Coordinate changes.) There are real functions and , one pair for each , whose restrictions to are and map into , such that, for every and every ,

  • (Unmixing relation.) For every and every ,

⊢ Lean

This is the usual observational equivalence scale for nonlinear component recovery: the coordinate system is identified up to separate invertible transformations of each latent coordinate, while the graph and target labels move together under the same permutation.

The ratio graph need not equal the parent graph before pruning, because a positive discrepancy can follow an ancestral path. The next definition fixes the graph operation used to state the exact order information recovered by the first stage.

Definition 23 [synth_16] (Transitive closure of a directed graph).

For a directed graph on a vertex set , denotes its transitive closure: the directed graph on with an arrow exactly when there is a directed path of positive length from to in .

The headline result combines the generic ratio separation condition with conditional ranks and conditional-independence pruning. It gives the population decoding guarantee for the one-perfect-intervention-per-node design.

Theorem 3 [thm:exact-ratio-decoder] (Exact ratio decoding).

Fix , a finite labelled DAG on , and a sign vector . Let be a mechanism on . Let assign an observed world to every mechanism on , and write , with shared mixing map , inverse map , target permutation , observed laws , likelihood ratios , and log-ratios . Assume:

  • (Smooth positive mechanisms.) satisfies Assumption 1.

  • (Causal minimality.) satisfies Assumption 5.

  • (Shared mixing.) The maps and in satisfy Assumption 2.

  • (One intervention per node.) The laws , ratios , and target permutation in satisfy Assumption 3.

  • (Own-derivative signs.) satisfies Assumption 4 with sign vector .

  • (Gaussian cover separation.) The stratum point belongs to the cover-separated set of Definition 7 with the Gaussian unit-norm feature map associated with and target permutation ; equivalently, every transported ancestral cover has strictly positive population Gaussian-MMD discrepancy.

Let be the population decoder of Algorithm 1, let , and let be the environment-label DAG. Then The ordering selected by is a topological ordering of . For every topological ordering of , with predecessor set for node , and every , the selected conditional CDF is a continuous version of the conditional ratio CDF. For each such , there exists a continuous version such that, for every latent point in the latent cube, Every continuous version agrees with the selected on the observed conditional-ratio support, and the selected version satisfies the same structural identity at every mixed latent point . Consequently, for every in the latent cube, For every topological ordering of and every , Under , for every such ordering, every , and every , Thus, for every such ordering and every , the set is admissible and is the unique inclusion-minimal admissible parent set. The observed laws are coherent with , and the parent-pruned graph returned by is exactly .

For any DAG on , mechanism , and observed world satisfying Assumption 1, Assumption 5, Assumption 4 with the same sign vector , Assumption 2, and Assumption 3, equality between the observed environment-law family of and the observed environment-law family of implies that is componentwise -equivalent to up to the simultaneous relabeling encoded by the target permutations. When , every edge of has strictly positive observed Gaussian-MMD discrepancy, Within the positive compact-cube, shared -mixing, causal-minimal regime, identifies the latent representation and structure from one perfect intervention per node, supplies the continuous law-separation witness used by von Kügelgen et al. (2023) when , and recovers the graph, topological order, and intervention alignment relative to Wendong et al. (2023) and Yao et al. (2025).

⊢ Lean

The theorem has three linked components. First, Gaussian cover separation makes the ratio-discrepancy graph rich enough to recover the transitive closure of the environment-label DAG. Second, every topological order of that graph supplies conditioning sets large enough for the selected conditional CDFs to agree with their structural counterparts, so the ranks are the coordinatewise transforms or their sign-reversed versions. Third, causal minimality turns the observational conditional independences among these ranks into exact parent sets, which converts the order information into the graph .

The final clauses state the representation comparison at the population level. Equality of the observed environment-law family pins down the recovered world up to the componentwise equivalence of Definition 22. In the two-variable case, the same conditions also yield strictly positive observed Gaussian-MMD discrepancy on every true edge, giving a compact-cube Gaussian-MMD counterpart to the continuous ratio-law witness used by von Kügelgen et al. (2023). For the multi-node recovery task, the decoder delivers graph, order, and intervention alignment under the one-intervention-per-node design through ratio observables and conditional ranks.

Sample-split confidence edges

The population decoder in Algorithm 1 and Theorem 3 turns positive ratio-law discrepancies into ancestral order information. This section adds the finite-sample graph layer. The sampling scheme uses independent training folds to construct the fitted ratios and independent evaluation folds, with environment sample sizes , to test the resulting one-dimensional ratio laws.

Definition 24 [synth_5] (Sample-split experiment world).

A sample-split experiment for a fixed observed environment family , whose observations take values in , consists of a measurable sample space together with:

  • (Experiment law.) A measure on .

  • (Evaluation fold.) Environment sample sizes , for , and evaluation observations, each taking values in the observed state space,

  • (Training fold.) Training sample sizes , for , and training observations, again valued in the observed state space,

  • (Fitted ratios.) For each full training fold , each , and each , a fitted ratio value For every , the map is jointly measurable; for every and , the map is measurable; and for every , , and .

  • (First-stage radius.) A first-stage radius , indexed by the first-stage failure level , satisfying for every .

  • (Error levels.) A familywise evaluation error level and a first-stage failure level .

⊢ Lean

Definition 24 isolates the objects needed for a sample-split analysis: a training sample, an evaluation sample, fitted positive ratios, a first-stage radius , a familywise level , and a first-stage failure level . The radius is an input to the analysis: the section takes as given a first-stage estimator delivering uniform ratio accuracy with probability at least , and proves what follows from it. What the section delivers is therefore an inference interface for the ratio graph, valid for any first-stage estimator meeting that contract. The second-stage half of the radius, which comes from evaluation-fold kernel concentration, is explicit in the sample sizes and alone. The next condition supplies the probabilistic independence that lets the evaluation fold be analyzed conditionally on the trained ratios.

Assumption 7 [ass:independent-environment-sampling] (Independent environment sampling).

The sample-split experiment is carried by a probability law, and each environment law is a probability measure. It satisfies:

  • (Training fold.) For each environment , the training observations are measurable, independent within that environment, and each training observation in environment has law .

  • (Evaluation fold.) For each environment , the evaluation observations are measurable, independent within that environment, and each evaluation observation in environment has law .

  • (Fold independence.) The complete training fold is independent of the complete array of evaluation observations across all environments.

  • (First-stage measurability.) For each latent coordinate and latent state , the ratio estimate is measurable with respect to the training fold, and the same estimate is jointly measurable in the training fold and . The training-fold event is measurable with respect to the training fold.

⊢ Lean

This is the standard independent multi-sample sampling with sample splitting condition used in kernel two-sample analysis (Gretton et al., 2012). The within-environment independence clauses identify each empirical law with its environment distribution, while fold independence makes the evaluation concentration conditional on the realized training output. The measurability clause makes the simultaneous first-stage ratio-error event a genuine training-fold event.

The evaluation statistic compares the fitted-ratio distributions in an interventional environment and the observational environment through a Gaussian maximum mean discrepancy. The following definitions fix the evaluation notation, the Hilbert-space embedding, the empirical discrepancy , and the simultaneous radius .

Definition 25 [synth_25] (Evaluation-fold observations).

In the sample-split experiment, for each environment let denote the independent evaluation-fold observations drawn from the observed law , independently of the training folds used to construct the ratio estimators . Thus is the -th evaluation observation from interventional environment , and is the -th evaluation observation from the observational environment.

Definition 26 [synth_21] (Feature Hilbert space and kernel mean embedding).

Let be the complete real Hilbert space in which the feature map takes its values. The feature map is normalized for the Gaussian kernel when For any probability law on , define the kernel mean embedding with the integral understood as the Hilbert-valued Bochner integral.

Definition 27 [synth_14] (Empirical MMD discrepancy).

For , outcome , and a real Hilbert-space feature map of unit norm satisfying , define

⊢ Lean
Definition 28 [synth_11] (Simultaneous MMD radius).

In the sample-split experiment of Definition 24, assume:

  • (First-stage radius.) is the deterministic radius in the simultaneous first-stage evaluation-law error event at failure level .

  • (Evaluation sizes.) is the evaluation sample size in environment .

  • (Nondegeneracy.) and .

  • (Levels.) and .

The simultaneous MMD confidence radius is This radius is the threshold used to convert the empirical discrepancies into simultaneous lower confidence bounds for the population discrepancies .

The first term in Definition 28 carries the training-fold ratio error into the MMD scale. The second term is the simultaneous evaluation fluctuation over all ordered pairs, using the unit-norm Gaussian feature representation in Definition 26; this is the same empirical-process layer underlying kernel two-sample testing and its aggregated variants (Gretton et al., 2012; Schrab et al., 2023).

The graph estimator thresholds empirical lower confidence bounds for the population discrepancies. Its output is an edge set on the environment labels.

Algorithm 2 [def:sample-split-confidence-graph] (Sample-split confidence graph ).

Under Assumption 7, with training and evaluation folds and tuning levels , the sample-split confidence graph is defined by the following steps.

  1. On the training folds, fit likelihood-ratio estimators

  2. On the independent evaluation folds, compute the empirical Gaussian-MMD discrepancies using .

  3. Output the directed graph

⊢ Lean

The rule in Algorithm 2 selects an arrow exactly when the lower confidence bound for the corresponding population discrepancy is positive. The theorem below states the resulting simultaneous coverage and its implication for ancestral edge discoveries.

Theorem 4 [thm:simultaneous-confidence-edges] (Simultaneous confidence edges).

Fix , a labeled DAG on , a sign vector , a mechanism , an observed world with target permutation , and a sample-split experiment . Suppose:

  • (Structural model.) The mechanism, shared mixing representation, and one-intervention-per-node design satisfy Assumptions 1, 2, and 3.

  • (Sample splitting.) The training and evaluation folds satisfy Assumption 7, and each environment has evaluation sample size .

  • (Levels.) The familywise evaluation level and first-stage failure level satisfy .

  • (First-stage event.) With where the displayed integrals are finite, the training-fold event obeys

Let be the simultaneous evaluation event Then, conditionally on the training fold and on , this event has probability at least : for every training-fold measurable event , Consequently, Moreover, for every , every selected arrow of the sample-split confidence graph from Algorithm 2 satisfies that is an ancestor of in . If, in addition, then for every ,

⊢ Lean

The intuition is a conditioning argument. On the first-stage event, the fitted ratios are close enough to the population ratios in every environment law to perturb the population Gaussian-MMD discrepancies by at most the first part of . Given the training fold, the evaluation observations are independent across the multi-sample experiment, so the empirical embeddings concentrate simultaneously over ordered pairs. The selected arrows therefore have familywise-valid ancestral interpretation, and when every transported ancestral cover has population discrepancy larger than twice the radius, the selected graph has the same transitive closure as .

Limitations and future work

The preceding sections establish population recovery under generic ratio separation and give sample-split confidence edges for ancestral discoveries. This section records the regularity regime in which a next statistical step can be formulated: coordinatewise error bars for the conditional-rank coordinates delivered by Algorithm 1. We fix a smoothness index , lower-envelope constant , radius bound , compact evaluation set , and latent boundary margin ; these parameters define the bounded subclass used to pose a generated-ratio, generated-rank problem.

Assumption 8 [ass:bounded-holder-radius] (Bounded Hölder radius).

Let . The parameters satisfy and . For every , the three functions satisfy the common Hölder bound of order and radius on their respective domains: and on , and on . That is, for each such function with domain , is on , for every and every , and for every .

⊢ Lean

Assumption 8 is the standard bounded Hölder ball condition used to obtain uniform local-polynomial control (Xie, 2023). It places the observational mechanisms, intervention densities, and latent-coordinate log-ratios on a common smoothness scale.

Assumption 9 [ass:density-lower-envelope] (Density lower envelope).

For the bounded regime under consideration, the lower-envelope constant satisfies:

  • (Range.) .

  • (Observational densities.) For every and every latent state ,

  • (Intervention densities.) For every and every ,

⊢ Lean

The lower-envelope restriction in Assumption 9 is the standard uniform full-support condition (Xie, 2023). In the present model it gives a common lower bound for both observational conditionals and intervention marginals, keeping likelihood ratios well behaved on the closed latent cube.

Assumption 10 [ass:own-derivative-margin] (Own-derivative margin).

For every and every ,

⊢ Lean

Assumption 10 is specific to this analysis. It strengthens the fixed-sign condition from Assumption 4 into a uniform margin, so the own-coordinate monotonicity used by the rank construction remains quantitatively separated from zero.

Assumption 11 [ass:common-interior-domain] (Common interior domain).

The common interior-domain condition for and radius is the conjunction:

  • (Radius.) .

  • (Compactness.) is compact.

  • (Observed interior.) , where .

  • (Latent boundary margin.) For every ,

⊢ Lean

The compact interior domain in Assumption 11 is the standard evaluation-region condition for uniform conditional-CDF estimation away from boundary effects (Xie, 2023). Here it is imposed simultaneously in observed space and after pullback through the shared mixing map.

Assumption 12 [ass:bounded-shared-mixing-geometry] (Shared mixing geometry).

The shared mixing map satisfies

⊢ Lean

Assumption 12 is specific to this analysis. It gives a common geometric envelope for transporting smoothness, volume, and interior-domain statements between latent and observed coordinates.

Assumption 13 [ass:observed-log-ratio-regularity] (Observed log-ratio regularity).

The observed log-ratios satisfy

⊢ Lean

The generated regressors used in the second stage are the observed log-ratios, and Assumption 13 is the standard uniformly bounded smooth-generated-regressor condition on an interior compactum (Mammen et al., 2012). It supplies the regularity needed to compare oracle and generated conditioning arguments.

The next objects describe the conditioning variables selected by the population decoder and the corresponding local design envelope.

Definition 29 [synth_12] (Conditioning-coordinate domain ).

Fix the predecessor set selected by the population decoder and write . For a compact evaluation domain , the conditioning-coordinate domain for node is the image of under these log-ratio coordinates: When , is the singleton subset of .

Definition 29 turns the selected predecessor set into the conditioning-coordinate support . When no predecessor is selected, the conditioning problem collapses to the unconditional one.

Assumption 14 [ass:conditional-design-envelope] (Conditional design envelope).

For every , the density of under satisfies for all .

⊢ Lean

Assumption 14 is the standard interior local-polynomial design-density envelope condition (Xie, 2023). The density controls how much data are available in the conditioning neighborhoods indexed by Definition 29.

Definition 30 [def:bounded-holder-subclass] (Bounded Hölder subclass ).

The bounded Hölder subclass is

⊢ Lean

Definition 30 collects the smoothness, support, margin, geometry, observed-regularity, and local-design requirements into the bounded subclass . The definition gives a single parameter space over which uniform generated-rank questions can be posed.

The statistical program also requires a regime for sample sizes, first-stage ratio accuracy, and second-stage smoothing.

Definition 31 [synth_6] (Bounded regime parameters).

For an observed world , a bounded generated-rank regime consists of:

  • (Order.) An order map .

  • (Environment sample sizes.) A sample-size sequence for each and each environment .

  • (First-stage ratio rate.) A uniform first-stage rate satisfying

  • (First-stage failure rate.) A first-stage failure sequence satisfying

  • (Second-stage bandwidth.) A second-stage bandwidth satisfying

⊢ Lean

In Definition 31, the rate measures uniform first-stage log-ratio error, records first-stage failure probability, and is the second-stage bandwidth. These sequences separate the two statistical layers: estimating likelihood ratios and estimating conditional ranks from generated responses and covariates.

Definition 32 [def:first-stage-ratio-contract] (First-stage ratio contract).

For a generated observed-world assignment , a decoder-selected bounded-regime assignment , sample-split experiments , a common environment sample-size sequence , rate sequences and , and subclass parameters , define The first-stage ratio contract holds when the following conditions are satisfied:

  • (Vanishing positive rates.) For every , and , and both sequences satisfy and as .

  • (Common selected-regime schedules.) For every and every as in Definition 30, the selected bounded regime has environment sample-size sequence , ratio-error rate , and first-stage failure rate .

  • (Sampling design.) For every and every , the experiment satisfies the independent-environment sample-splitting conditions in Assumption 7, and its environment sample sizes are exactly .

  • (Measurable high-probability event.) For every and every , the event is measurable under , and

⊢ Lean

Definition 32 defines the high-probability event on which all fitted log-ratios and their first derivatives are uniformly close to their population counterparts. Its form is deliberately matched to the generated-regressor literature, where first-stage perturbations enter through both fitted responses and fitted conditioning arguments (Pagan, 1984; Mammen et al., 2012; Lee, 2018).

Remark 1 [def:generated-rank-handle] (Generated-rank analysis handle).

, the generated-rank proof-strategy handle, records a nonassertive construction for , the proposed second-stage cross-fitted rank estimator. The estimator is formed by adapting Xie’s local-linear conditional-CDF estimator to cross-fitted fitted log-ratio responses and cross-fitted conditioning covariates. The handle decomposes the second-stage rank error into the oracle conditional-CDF process, indicator boundary crossings from the generated threshold, and local-design perturbations from generated conditioning arguments; it then linearizes the local-design perturbation component using the generated-covariate expansion of Mammen, Rothe, and Schienle. The second-stage weight convention, bandwidth selector, influence representation, attainment claim, and sharpness claim are treated as open components outside the supplied handle.

⊢ Lean

The handle organizes the proposed second-stage estimator around local-linear conditional-CDF estimation (Fan et al., 1996; Fan et al., 1996; Xie, 2023). Its decomposition isolates oracle smoothing variation from the additional terms created by generated thresholds and generated conditioning coordinates.

Definition 33 [synth_7] (Maximum conditioning dimension).

The maximum conditioning dimension of the selected ordering is where is the set of labels preceding in that ordering.

⊢ Lean

Definition 33 defines the effective nonparametric dimension . It is the dimension that appears in the local-neighborhood sample size for the conditional-CDF step.

Remark 2 [oeq:generated-rank-frontier] (Generated-rank frontier).

A natural next question is whether the handle yields uniform generated-rank estimators under the following conditions:

  • (First-stage contract.)

  • (Bandwidth scale.) The bandwidth satisfies , and

The target is a construction of cross-fitted estimators such that where When the first-stage log-ratio estimators additionally admit a uniform asymptotically linear representation on , future work could derive the corresponding expansion for , separating threshold crossings, generated-conditioning perturbations, oracle conditional-CDF variance, and smoothing bias, and could characterize the weakest sufficient first-stage remainder bound.

⊢ Lean

Remark 2 records the rate target in terms of the minimum environment sample size , that is , and the proposed rate . The three terms correspond to second-order smoothing bias, local empirical variation, and first-stage error propagated through bandwidth-scale neighborhoods. Meeting that target would supply a coordinate error analysis for the nonlinear representation delivered by Theorem 3, and would attach rank-coordinate uncertainty to the confidence-edge construction of Theorem 4. Whether a deployed pipeline such as ROPES (Kulkarni et al., 2025) satisfies the model assumed here is a separate question that this paper does not settle; the reference is to the kind of implementation such an error analysis would serve.

Several additional directions remain open. Exact-parent inference calls for conditional-dependence margins that distinguish direct parents from more distant ancestors after the population decoder has recovered the transitive structure. Broader support conditions would extend the compact-cube analysis to domains with boundary or tail behavior closer to applied continuous measurements. Vector-valued latent nodes would replace scalar monotone coordinates by multivariate intervention scores, changing both the separation argument and the conditional-rank step. Imperfect interventions would allow the intervention law to retain controlled parent dependence, connecting the one-intervention design here to general-environment formulations (Ng et al., 2025; Jin et al., 2024; Zhang et al., 2024; Ahuja et al., 2024) and to recent benchmark and application settings for interventional representation learning (Chen et al., 2025; Wang et al., 2024; Sun et al., 2024). Noninvertible observation maps would move beyond the shared diffeomorphic mixing geometry in Assumption 2 and require an identification target defined on observational equivalence classes rather than on recovered latent coordinates.

Appendices

Proofs and verification

This appendix collects the auxiliary statements that support the generic separation, population decoding, and sample-split confidence results. The first group records the explicit coordinate reflection and affine perturbation used to make a direct-edge ratio contrast nonzero inside a fixed-sign stratum.

Definition 34 [synth_22] (Coordinate reflection map).

For , define

⊢ Lean
Definition 35 [def:affine-path] (Edge affine path).

For a finite labeled DAG on , a sign vector , a stratum point as in Definition 3, an edge in , and , define the reflected edge-specific sparse endpoint in by where . The affine path is the mechanism on with

⊢ Lean
Definition 36 [synth_15] (Second-moment contrast).

For an observed world with environment laws and likelihood-ratio carriers , define, for ,

⊢ Lean
Lemma 1 [lem:analytic-edge-perturbation] (Analytic edge perturbation).

After the harmless target relabeling , fix the ambient number of scalar latent variables, a finite labeled DAG on , a sign vector , a mechanism as in Definition 3, and an edge of . Assume that the identity-target canonical observed world generated by satisfies Assumption 3. Let be the edge-specific affine path of Definition 35, and write for the direct-edge second-moment contrast computed in the identity-target canonical observed world generated by a mechanism . Along this path, There is an open set with such that:

  • (Integral identity.) For every , the displayed integral equals .

  • (Analyticity.) The map is real analytic on a neighborhood of every point of .

  • (Isolated zeros.) If and , then there is such that every with satisfies .

Moreover, . Finally, for every relative neighborhood of , there exists such that , , and .

⊢ Lean
Proof of Lemma 1.

Work after the stated relabeling, so the target permutation is the identity. Let denote the edge-specific sparse endpoint used in Definition 35, and write the affine extension for real as

Step 1: the positive analytic parameter set. What the rest of the argument uses is an open convex set on which every path mechanism is positive and normalized, together with real analyticity on of the parameter integral displayed in Step 2. We exhibit one such set. Strict positivity of , strict positivity of the sparse endpoint, continuity on the compact cubes, and finiteness of the node set give a number such that every is at least on its domain; any positive lower bound serves equally well below. Define For fixed and cube point, the two displayed inequalities define open half-lines in the parameter . Taking the infimum over the compact sets and gives continuous lower-envelope functions, so and are open; as intersections of affine strict-superlevel conditions, they are convex. If , then for every admissible argument and hence . On , all and are positive. Normalization and -smoothness are affine identities: for every real , every , and every , and the derivatives up to order three are the same affine combinations of the endpoint derivatives. Moreover, for every and ,

We now make the analytic-under-the-integral step explicit. Put and with the empty product equal to . Since the product has finitely many factors and every factor is affine in , there are measurable coefficient functions , continuous on the cube after choosing the standard continuous representatives on the ambient cube, such that Fix . The denominator satisfies on the compact latent cube, while By compactness, choose so that for , Then converges uniformly in . Multiplying by the polynomial , the coefficient bounds and the finite volume of give a summable dominating majorant on a smaller neighborhood of . Termwise integration therefore gives a power series for near . Thus the parameter integral is real analytic on a neighborhood of every point of .

Step 2: the integral identity. Since is an edge in a DAG, . For each , the construction of supplies positive normalized smooth mechanisms. In the identity-target canonical observed world generated by , the latent observational density, target- density, and target- ratio are and The lower bound on makes the ratio continuous and integrable on the compact latent cube. Therefore, with the empty product interpreted as , This proves the displayed identity for every . The right-hand side is the analytic parameter integral constructed above, so the same is true of .

Step 3: the endpoint is nonzero. At , the path is the embedded sparse endpoint. Write where is the coordinate reflection appearing in Definition 35. At the sparse endpoint, and for . Since the edge gives , the two selected coordinates form a two-element set. The integrand in the identity from Step 2 is therefore which depends only on ; the finite product over is , with the empty-product convention covering the case of no unused coordinates. Product integration on then gives Each is either the identity map or , hence preserves Lebesgue measure on . Define, for , Changing variables first in and then in gives

It remains to establish the scalar gap for this last integral. Introduce the auxiliary primitive and, for , set The identities and imply, by integration by parts, For , differentiating the defining integral for under the integral sign gives The differentiation is justified because and , so the denominator is bounded away from zero uniformly on the square.

For every and , the same bounds give . Together with , this yields Since we have the uniform derivative bound Consequently, for , Applying this secant estimate on the intervals from to and from to gives, for every , Because and , integration gives Hence, for , The exponential distribution function in lies in , and convexity of the exponential on gives ; therefore on . Multiplying the last derivative bound by the nonpositive function and integrating yields A direct integration of the exponential density gives so integration by parts gives Combining the rational factors, and hence Thus and in particular

Step 4: isolated zeros. Let and suppose . The set is convex, hence preconnected, and with . For a real analytic function on a preconnected set, the local dichotomy at is: either the function vanishes on a neighborhood of , or it is nonzero on a punctured neighborhood of . The first alternative would force the function to vanish throughout , contradicting the nonzero value at . Thus there is such that

Step 5: small stratum-preserving nonzero perturbations. Let be any neighborhood of in the topology of Definition 3; in particular this covers every relative neighborhood in the statement. The affine path satisfies so in the topology of Definition 3. Hence there is such that implies .

The conditions defining persist on a sufficiently short initial segment. Positivity, normalization, and smoothness are handled first. For every , convexity of the positive endpoint densities gives positivity of every and , and the affine identities show normalization; -smoothness follows by applying the same affine combination to derivatives up to order three.

For the fixed-sign part, define for , , and , where is the one-dimensional derivative on and is the own-coordinate derivative on the latent cube. Here abbreviates regarded as a function on the whole latent cube, and likewise . At , this is exactly the fixed own-coordinate log-ratio derivative required in Assumption 4. The numerator and denominator terms in are continuous in , and the denominators are bounded away from zero at . Since is compact and there are finitely many , the strict inequalities persist uniformly for all and all sufficiently small .

It remains to record the local witness used for causal minimality. For a positive normalized smooth mechanism , let be the compact product factorization on with local factor For an edge , a base point , child values , and parent values , define the local determinant where is obtained from by replacing the -coordinate by and the -coordinate by . In a positive compact product factorization, for all such . Indeed, after conditioning on , positivity lets one cancel the remaining factors, and conditional independence is exactly the rank-one condition that all two-by-two determinants of the child factor as a function of vanish. The compact-coordinate statement is the same as the coordinate conditional-independence statement in Assumption 5, because both use the same cube-supported observational density.

Since , for every edge there is a witness with For each fixed witness, the map is continuous, so the witness remains nonzero when the corresponding local factors are uniformly close to those of . Taking the minimum over the finite edge set gives such that every positive compact factorization satisfying has the same nonzero witnesses and hence satisfies causal minimality. If the edge set is empty, the condition is vacuous.

Along the affine path, so the factorization neighborhood above contains for all sufficiently small . Combining the closed-segment positive normalized smoothness with the local fixed-sign and causal-minimality neighborhoods, there is such that and imply .

It remains to choose avoiding the zero set. If , analyticity gives continuity at , so for all sufficiently small . If , the isolated-zero conclusion at gives the same conclusion for all sufficiently small nonzero . Hence there is such that Set Then , , , and .

The perturbation statement gives the local ingredient behind the density argument in Theorem 2: the affine path stays in the same fixed-sign stratum near the starting mechanism and produces a nonzero direct-edge second-moment contrast. The companion zero statement records the population implication for labels whose transported targets have no ancestral relation, using the Gaussian mean-embedding discrepancy that defines the ratio graph in Algorithm 1.

Lemma 2 [lem:ratio-nonancestor-zero] (Nonancestor ratio discrepancy vanishes).

Fix an ambient number of scalar latent variables, a finite labeled DAG on , a mechanism , and an observed world with environment laws , likelihood ratios , and target permutation . Assume Assumptions 1, 2, and 3. Let satisfy , and suppose that is not an ancestor of in . Then the Gaussian-kernel discrepancy of ratio between the observational environment and environment vanishes: where is the Gaussian kernel and is its feature map.

⊢ Lean
Proof of Lemma 2.
  1. Let be the latent observational law with density on , and let be the latent law for the single intervention at , with density on . On the cube define The denominator is positive on by Assumption 1. By Assumption 3, the observed laws are the -pushforwards of the latent laws, and the likelihood-ratio formula stated there identifies the two ratio laws as

  2. Since and is injective, . Together with the hypothesis that is not an ancestor of , this gives The set is parent-closed: if and , then , because a parent of is an ancestor of , and a parent of an ancestor of is again an ancestor of . Let be the coordinate projection. For and , write for the point of whose -coordinates are and whose complementary coordinates are . For such , set and where the coordinate is defined because . Parent-closedness implies that, for every , the factor evaluated at depends only on . Hence and The complement factors have unit iterated integral. Indeed, if a nonempty collection of complementary vertices remains to be integrated, choose one that is maximal in a topological ordering of the induced remaining set. Parent-closedness prevents this vertex from being a parent of any retained factor, and maximality prevents it from being a parent of any other remaining factor. Its coordinate therefore appears in the current product only through its own factor; integrating that factor gives one by the normalization of , and, when the vertex is , by the normalization of . Iterating this deletion, with the empty remaining collection as the base case, gives Consequently the two complement-marginal densities on are the same: Thus the projected latent laws agree: Since , the ratio map factors through this projection. Explicitly, the map defined by satisfies . Pushing the projected-law identity forward by gives and hence, by the ratio-law identifications in the preceding step,

  3. The population discrepancy is the norm of the difference of the two Gaussian-kernel mean embeddings of these ratio laws. Using the equality just proved,

The next statements isolate the triangular change of variables used by the conditional-rank decoder. They express predecessor log-ratios as coordinates for the already ordered predecessors and then factor the target intervention law conditional on those scores.

Definition 37 [synth_29] (Triangular predecessor-score range and inverse).

Fix an observed world , a mechanism , and an injective topological order for the observed ratio graph. For , write for the predecessor labels of in this order. The predecessor-score space is , with coordinates indexed by , and the predecessor latent cube is .

For , let be the latent vector that places in coordinate for and fills all remaining coordinates by . The triangular predecessor-score map is Equivalently, when and for ,

The realized predecessor-score range is the image Under the positive smooth mechanism, shared mixing, one-intervention-per-node, and fixed own-derivative-sign conditions used in Lemma 6, is injective and hence is a homeomorphism from onto . Its inverse is written so that for , is the unique predecessor latent vector with , and The range is the continuous image of a compact cube, hence a nonempty compact Borel subset of . Formulas indexed by an arbitrary score vector are totalized through a fixed Borel retraction which sends every to one fixed point of . Writing for this retraction and for the inverse keeps the two maps distinct: the composition is defined on all of , and off the convention carries no content, since every continuity and support claim is asserted on .

Lemma 3 [lem:predecessor-score-homeomorphism] (Predecessor score homeomorphism).

Let be a DAG on , let be a sign vector on , let be a mechanism on , and let be an observed world with observed laws , likelihood ratios , shared observation map , and target permutation . Assume:

  • (Regular mechanism and observation model.) The hypotheses of Assumptions 1, 2, 3, and 4 hold for , with prescribed sign vector .

  • (Observed ratio ordering.) The map is injective and is a topological ordering of the observable ratio-discrepancy graph constructed from the observed laws with the Gaussian feature map in Algorithm 1: whenever and , .

  • (Transported latent ordering.) The same order also respects the transported latent DAG: whenever is an edge of , .

Fix , set , and let For , let be the zero-filled latent point with for and all other coordinates equal to zero. Define by and write . Then is a homeomorphism from onto . In particular, the inverse map is continuous.

⊢ Lean
Proof of Lemma 3.

Write , , and . For , set , so This point lies in , since each displayed predecessor coordinate lies in and every filled coordinate is .

  1. The map is continuous. Indeed, is continuous coordinatewise: each coordinate is either one of the product-coordinate projections or the constant . For each , positivity gives and the regularity in the hypotheses gives continuity of both factors on the relevant closed cubes. Hence the quotient is a continuous positive function of , and applying preserves continuity. Since this holds for every coordinate , the product map is continuous.

  2. The map is injective. Let satisfy and put , . With , the latent ratio identity in the perfect-intervention hypotheses gives, for every , Thus equality of the score vectors is exactly equality of the predecessor log-ratio projections:

    We prove by strong induction on that Fix with , and assume the claim for all smaller order values. If , let . Then is an edge of , so the transported latent ordering gives Since , also , hence and . The induction hypothesis gives Therefore

    For , let be the point obtained from by replacing the coordinate by : Define, on , the one-coordinate score Strict positivity and regularity make continuous on and differentiable on . For , the point lies in the latent cube, so the fixed own-derivative sign hypothesis gives Since , the mean-value theorem yields, for any , for some , with positive right-hand side when and negative right-hand side when . Thus is strictly monotone on .

    Moreover, and parent locality, together with for all , gives so The equality therefore gives By strict monotonicity of , The induction is complete. Finally, for each , so in the product cube. Thus is injective.

  3. Since is finite, is compact, and is Hausdorff. Every closed subset of is compact, so its image under the continuous map is compact and hence closed in . Together with injectivity, this makes a topological embedding with closed image. Therefore the induced map is a homeomorphism onto its range. Consequently is continuous. The same argument covers the case , where both products are singleton spaces.

The homeomorphism clause supplies the support-level continuity needed for conditional distribution transforms: predecessor scores identify predecessor latent coordinates on the realized range. The following two marginal identities then remove coordinates outside the retained ancestral set while preserving the factors that enter the target log-ratio.

Lemma 4 [lem:equation-ten-intervention-marginal] (Intervention marginal identity).

Fix . Suppose that:

  • (Regular mechanisms.) The mechanism satisfies Assumption 1.

  • (Ratio-graph order.) The map is injective and is a topological ordering of the observable ratio-discrepancy graph in Algorithm 1: whenever , .

  • (Transported-DAG order.) The same order respects the transported latent DAG: whenever is an edge of , .

Let For and , let be the point with Then, for every ,

⊢ Lean
Proof of Lemma 4.

Put and .

  1. First, is parent-closed: Indeed, if , then in the transported environment-label graph, so the transported-DAG order gives Thus , hence . If instead for some , then in the transported graph, and the same order gives Thus again , hence .

  2. Define, for and , The intervention density at target is therefore For and , write for the point satisfying Since , the preceding display is the ordinary product density after replacing the -factor by .

    For a parent-closed subset , an index , and a normalized replacement density , define and let The replacement construction gives that is, the observational product factorization with the single -factor replaced. For this replaced factorization, parent-closed marginalization gives where the integrated point agrees with on . This is the usual finite product-density marginal identity for a parent-closed coordinate set: after the single-factor replacement, the factors again form a normalized directed factorization, and parent-closedness ensures that retaining leaves exactly the product of the retained factors. The sigma-finiteness and finite-coordinate integration bookkeeping are supplied by the unit-interval reference measure, while Assumption 1 supplies the normalized measurable factors and the normalized replacement density.

    Apply this identity with and . Unfolding the retained product gives because the erased product contains no -factor. Therefore

  3. Substituting and in the preceding identity gives, for every , which is the asserted identity.

Lemma 5 [lem:equation-ten-fiber-product] (Fiber product factorization).

In the setting of Assumption 1, fix an observed world with target permutation . Let be a numerical order on , and fix . Assume:

  • (Ratio-graph order.) The order is injective and topological for the observable ratio graph of Algorithm 1: whenever , one has .

  • (Transported-DAG order.) The same order respects every transported latent edge: whenever is an edge of , one has .

Write For and , let be the latent point whose coordinates in agree with and whose remaining coordinates are filled by . For every and every , if denotes with its -coordinate replaced by , then

⊢ Lean
Proof of Lemma 5.

Let , and put . For , write

  1. The intervention marginal factorization supplied by Lemma 4 is used here in its ambient density-carrier form: for the arbitrary point , The density carriers and are evaluated as their ambient functions on and , respectively, so the coordinate replacement is admissible for every and . Since , the scalar factor on the right is .

  2. It remains to identify the retained observational product after the coordinate replacement. We claim that Fix . By the definition there is with . The own coordinate satisfies because lies in . Now let . If , then is an edge of , so the transported-DAG order gives But gives , a contradiction. Hence Thus and agree on , and consequently Multiplying this equality over proves the claimed product identity; when the index set is empty, both products are the empty product .

  3. Substituting the product identity from the previous step into the marginal identity from the first step yields which is the asserted fiber product factorization.

Lemma 6 [lem:equation-eleven-kernel-factorization] (Kernel factorization of scores).

Let be fixed, let be a finite labeled DAG on , let be a sign vector with entries in , and let be an observed world generated by a mechanism . Assume:

  • (Structural regularity.) The mechanism and observed world satisfy Assumptions 1, 2, 3, and 4 with sign vector .

  • (Observed ratio order.) The map is injective and is a topological ordering of the observed Gaussian-MMD ratio graph of Algorithm 1: whenever contains , .

  • (Transported graph order.) The same order respects the transported latent DAG: whenever is an edge of , .

Fix , and put For each predecessor log-ratio vector , let be the law on of where is the triangular predecessor-score inverse composed with the retraction onto the realized score range , and is the latent point whose -coordinate is for , whose -coordinate is , and whose remaining coordinates are zero. Then, under the target- observed environment law , the joint law of the predecessor log-ratios and the target log-ratio factorizes as

⊢ Lean
Proof of Lemma 6.
  1. Let be the latent target- intervention law on , and let By Assumption 1, is a probability law. Define the total clamped predecessor-coordinate map Thus, for , For , let be the latent point with The triangular predecessor-score map is The assumptions in the statement make one-to-one on the compact predecessor cube, with realized range , as recorded in Definition 37. For and , put The clamping makes a Borel map on the full domain , while on it is exactly the target score in the statement. Since the retraction onto fixes every realized value and the inverse of sends back to , the kernel satisfies the fiber identity When , the predecessor cube is the singleton product and the same formula is the corresponding singleton-index identity.

  2. The latent predecessor coordinates and the target coordinate factor under : Indeed, Lemma 5 integrates the target- intervention density over all coordinates outside and leaves as the target-coordinate factor, multiplying a factor depending only on the retained predecessor coordinates. The normalization in Assumption 1 identifies this target factor with .

  3. Define the latent joint score map For a measure on and a Markov kernel from to , write for the measure on defined by for measurable . This is the product denoted by in the displayed statement. Mapping the product identity in the previous step through and using the fiber identity for gives the measure-kernel product factorization Indeed, for every measurable , which is exactly the displayed measure-kernel product.

  4. For -almost every , one has , because the support of is the closed latent cube. On this event, clamping is inactive. Fix . The point agrees with in coordinate . It also agrees on each parent of : if is an edge of , write . Then the transported edge is , and the transported graph order gives . Since , , so and hence . Thus the retained predecessor vector supplies the coordinate . By the parent-locality of and the latent ratio formula in Assumption 3, Thus

  5. The same comparison gives the target coordinate. If is a parent edge in , write . Then , and the transported graph order applied to gives , so . Therefore agrees with on and on every parent coordinate of . Hence, again by Assumption 3, for -almost every . Combining this with the preceding step,

  6. By Assumption 3, the observed target- law satisfies Therefore the almost-sure identity in the previous step yields For the marginal predecessor pushforward, the almost-sure identity from step 4 gives and the same pushforward calculation gives Substituting these two pushforward identities into the latent factorization from step 3 gives which is the asserted factorization, with the statement’s product understood as this measure-kernel product.

Together, these identities turn the target intervention environment into a product between predecessor scores and a one-dimensional kernel indexed by those scores. The continuous-rank clauses below state the corresponding conditional-CDF representation and the resulting rank coordinate used by the population decoder.

Lemma 7 [lem:exact-ratio-decoder-continuous-rank-clauses] (Continuous rank decoder clauses).

Let be the ambient number of scalar latent variables, let be the ambient finite labeled DAG on , let be the mechanism, let be the observed world with shared observation map , observed laws , likelihood ratios , log ratios , and target permutation , and let be a sign vector with for every . Assume:

  • (Mechanism regularity.) The mechanism satisfies Assumption 1.

  • (Shared mixing.) The observed world satisfies Assumption 2.

  • (Single-target interventions.) The observed laws and likelihood ratios satisfy Assumption 3.

  • (Own-derivative signs.) The mechanism satisfies Assumption 4 with sign vector .

  • (Ratio-order recovery.) For the observable directed ratio-discrepancy graph of Algorithm 1, built from the Gaussian kernel , the transitive closures agree:

For any injective topological ordering of , any , and , call a function a continuous conditional-CDF version of given when is jointly Borel measurable, for every fixed it represents the regular conditional probability for -almost every , the same identity holds for -almost every , and is continuous on the observed conditional-ratio support Then, for every such and every , with denoting the law-selected conditional CDF used by Algorithm 1, the following hold:

  • is a continuous conditional-CDF version of given .

  • There exists a continuous conditional-CDF version such that, for all and all , where is the latent point obtained from by replacing coordinate by .

  • Every continuous conditional-CDF version agrees with on :

  • For all and all ,

  • The rank coordinate returned by Algorithm 1 satisfies, for every , where .

⊢ Lean
Proof of Lemma 7.

Fix an injective topological ordering of , and fix . Write

  1. First record the order-recovery consequence that is passed to the ordered decoder argument. If contains , then the one-edge path is an element of . The assumed equality therefore gives a directed -path from to . Along each edge of this path, topologicality of for gives ; induction over the path and transitivity of yield Thus is also topological for the transported graph , and in particular every -parent of lies in .

  2. The supplied observed laws are probability laws. The observational law is the pushforward by of the normalized observational latent density, and each target-intervention law is the pushforward by the same of the corresponding normalized intervention density. Hence, for every , Consequently the decoder is computed from these laws themselves: writing for the law family it uses, In each subsequent conditional-CDF statement, the selected family is rewritten by this identity to the stated observed laws.

  3. Let be the probability kernel from Lemma 6, extended to all by first retracting onto the realized predecessor-score range. Thus, for every , is the law of where the retraction in supplies the off-range values. Define the total candidate The same formula applies when , with the one-point product space. Put By Lemma 6, The candidate has the four properties required of a continuous conditional-CDF version. First, is Borel measurable because it is the lower-half-line evaluation of a measurable probability kernel. Second, uniqueness of regular conditional distributions in the displayed product disintegration gives, for every fixed , The same disintegration, pulled back along the argument map , gives the joint almost-everywhere version For continuity, put On , the retraction in the definition of is the identity, and the predecessor-score homeomorphism rewrites as the lower-level integral of the one-dimensional score The positivity and smoothness in Assumption 1, together with the fixed own-derivative sign in Assumption 4, make this score continuous in and strictly monotone in on ; hence its lower-level integral varies continuously in on . The observed conditional-ratio support satisfies because under the predecessor projection is supported on realized predecessor scores. Thus is continuous on , and so it is a continuous conditional-CDF version of given .

  4. Since the class of continuous conditional-CDF versions is nonempty, the law-selected conditional CDF used by Algorithm 1 is itself one of these versions. This proves the first asserted clause: Taking proves the asserted existence of a continuous version once its pointwise latent formula is verified in the next step.

  5. Fix , and set The realized score belongs to , so the retraction used to define fixes it. The score-image kernel then evaluates at the predecessor coordinates reconstructed from . For each , let be this reconstructed point with target coordinate set to . The order-recovery step places every transported parent of in , and the inverse predecessor-score map reconstructs those predecessor coordinates from . Hence where is obtained from by replacing coordinate by . Parent locality of the mechanism gives Substituting this identity into the lower-half-line evaluation of yields, for every , This is the second asserted clause, with .

  6. Any two continuous conditional-CDF versions agree on the observed conditional-ratio support. To see this, fix another continuous version . For each fixed threshold , the defining conditional-CDF property gives Both sides are continuous in on . Therefore, for every and every , that is, for every , Equivalently, It remains to place realized predecessor scores in this support. For every , Indeed, is the pushforward by of the target- intervention law; by strict positivity of the latent density on the cube and the shared diffeomorphic mixing, lies in the support of . The predecessor-log-ratio projection is continuous on the observed support, so every neighborhood of has preimage with positive -mass. Applying the support uniqueness just proved to , and then using the previous step, gives This proves the third and fourth asserted clauses.

  7. It remains to evaluate the rank at the realized target log-ratio. Fix . The latent ratio formula in Assumption 3, together with the selected log-ratio version used by the decoder, gives the pointwise log-ratio identity By definition of the rank coordinate and by the fourth clause already proved, Define, for this fixed , For every , the point remains in . Positivity and smoothness from Assumption 1 make continuous on and differentiable in the interior, and its derivative is the own-coordinate log-ratio derivative at : The fixed sign condition in Assumption 4 gives Thus is strictly increasing on when , and strictly decreasing on when . Since we have If , strict increase makes the lower level set , so If , strict decrease makes the lower level set , and normalization of gives Thus, for every , This proves the rank-coordinate clause.

  8. The argument was for an arbitrary injective topological ordering of and an arbitrary . The five asserted clauses therefore hold for every such and every .

The last auxiliary bound supports the simultaneous confidence graph in Theorem 4. It is a Hilbert-space empirical-mean concentration statement for bounded Gaussian-kernel feature maps, matching the kernel mean-embedding discrepancy used in maximum mean discrepancy testing (Gretton et al., 2012; Schrab et al., 2023).

Lemma 8 [lem:bounded-rkhs-empirical-mean] (Conditional RKHS mean bound).

Fix , a labeled DAG on , a mechanism , an observed world with environment laws , and a sample-split experiment . Let be a complete real Hilbert space equipped with a measurable unit-norm feature map for the Gaussian kernel .

Assume:

  • (Conditional evaluation sampling.) The experiment satisfies the conditional evaluation-sampling conditions: the ambient and environment laws are probability measures; the complete training fold is measurable; the evaluation observations are measurable, independent within each environment, distributed according to their corresponding , and independent of the training fold; and the fitted ratio map is jointly measurable in the training fold and the evaluation point.

  • (Evaluation sizes.) Each environment has at least one evaluation observation, for every . Write

  • (Familywise level.) The error level satisfies .

For each training realization, let denote the fitted ratio obtained from the complete training fold, and define Then, conditionally on the complete training fold, the simultaneous event has probability at least . Equivalently, for every event measurable with respect to the complete training fold,

⊢ Lean
Proof of Lemma 8.

Fix an event measurable with respect to the complete training fold. It is enough to prove the displayed lower bound for this .

  1. Set Since , . Hence , and For every environment , and . Therefore where the equality uses , and the inequality uses monotonicity of the square root together with the nonnegativity of the numerator.

  2. Fix , an environment , and a training realization . Define for . The unit-norm feature-map assumption gives Under the product law , Indeed, the squared centered empirical mean expands into diagonal and off-diagonal Hilbert inner products; the off-diagonal terms vanish by product independence after centering, while the diagonal terms are bounded by the unit second moment. If is obtained from by replacing only coordinate by , then McDiarmid’s bounded-difference inequality therefore yields, for the bound

  3. Let The preceding step gives The conditional evaluation-sampling assumptions give measurability of these sections, independence of the complete training fold from the evaluation vector and the law . Hence, for every training-fold event , By the definitions of the empirical and fitted mean embeddings, together with the change-of-variables identity for the fitted ratio map, Thus

  4. Define the pairwise bad events and set The union bound and the preceding pairwise estimate give

  5. Let If , then each pair avoids , so the pairwise inequality from Step 1 implies . Consequently, Since all measures involved are finite, Because was an arbitrary complete-training-fold measurable event, this is the asserted conditional probability lower bound.

Verification record

The formal layer accompanying this paper is a Lean 4 development, built with toolchain leanprover/lean4:v4.33.0 against the Mathlib revision pinned in that repository. Its sources are the module tree CausalSmith/ExactID/EID_CrlCoverratioMmdGenericity_Research/. The theorem statements and proofs corresponding to the declaration-naming formal objects in this paper are machine-checked there, including Theorems 2, 3, 4, and 1 and the auxiliary statements displayed in this appendix. The certificate covers the frozen assumptions, definitions, auxiliary lemmas, and theorems recorded in the paper crosswalk; external citations and Remark 2 enter through the stated theorem-local comparisons, assumptions, and labelled future-work objects.

Proofs of the main results

Proof of Theorem 1.

Fix , and write The map is a unit-Jacobian bijection of . Put

  1. First consider the structural assertions. Since the intervention densities and the displayed child conditionals in Definitions 8 and 9 are normalized. Convexity gives so . Hence, for both and , The reflected mechanisms therefore have strictly positive normalized factors. Their factors are restrictions of smooth elementary functions, so every is on and every is on .

    For the unreflected child factor, and at nodes and the same own-coordinate derivative equals . Applying multiplies the -th own derivative by , so throughout the closed cube. This proves Assumption 1 and Assumption 4 for both witnesses.

    It remains to check faithfulness for the two observational laws. Define, for , Reflection preserves Lebesgue measure and sends to or , so For the sparse law, the child density is . Hence, for each fixed , and After integrating over and over the independent third coordinate, these identities give and If the first two coordinates were independent under , the last integral would be the product of the preceding two means, equal to zero; this contradicts . Thus the edge endpoints are dependent in the sparse observational law.

    For the cancellation law, put The function is continuous and nonconstant: indeed , while Therefore The cancellation child density is , and the same centered-coordinate calculation gives Independence of the first two coordinates under would force contradicting . Hence the edge endpoints are also dependent in the cancellation observational law. Finally, the factorization over with node isolated gives the local-Markov independence of nodes and . In this three-node graph, any failed d-separation between disjoint coordinate blocks places the endpoints of the unique edge in opposite blocks; if the conditioning block contains node , contraction with the independence of nodes and reduces it to the unconditional edge independence just ruled out. Both observational laws are therefore faithful to the three-node DAG, giving Assumption 6.

  2. We next prove the sparse second-moment separation. In the unreflected sparse construction define The product-density calculation gives and reflection changes variables in the same integral, so the identity holds for every sign vector.

    For , set The elementary exponential bounds imply Moreover so Integrating this derivative bound from to , with the direction reversed when , gives After integration over , using and we obtain Since and with , integration by parts yields The last integral is Therefore

  3. Let and be the laws of under and . The bounds above also give so both laws are supported on . The preceding item gives

    We now derive the quantitative Gaussian recovery estimate used on this common support. Put The power series for the exponential gives so is the Gaussian feature map. For probability laws supported on , write Reading the -th coordinate and bounding it by the Hilbert norm yields Define the degree- truncation For , the tail of the exponential series satisfies because after the first omitted term the successive ratios are bounded by . Since and , multiplication by gives the uniform approximation bound Consequently, for every such pair , Since the central-binomial estimate implies Thus The remaining rational tail comparison is Combining the last three displays gives, for laws supported on , Applying this estimate to , where , and using the strict moment gap above, we obtain

  4. It remains to identify the cancellation ratio law. Let be bounded and continuous. After expanding the observational and parent-interventional densities for , the difference of the two test-function expectations equals with the same value after the sign reflections by the changes of variables and . For fixed , define Since , is continuous and has an antiderivative . Using and , Fubini’s theorem applies on the compact square, hence the two expectations agree for every bounded continuous . Thus the law of under the observational environment equals its law under environment for .

    Since is the Hilbert norm of the difference of the Gaussian mean embeddings of these two ratio laws, equality of the laws gives

Combining the four numbered assertions gives exactly the four clauses of the theorem.

Proof of Theorem 2.

1. First, . If , then the cover relation in the ancestral order is a direct edge: Indeed, an ancestral relation in a DAG is either a child edge or factors through an intermediate ancestor; the cover property excludes the latter. Hence, for , Definition 5 gives Let and be the two ratio laws appearing in Definition 7: By Definition 3, the ratio is continuous and bounded on the compact latent cube. Hence there is such that, with the two ratio laws satisfy We use the following bounded-support Gaussian recovery fact. If finite real laws and are supported on the same interval , with , then To prove the fact, use the explicit Gaussian feature coordinates Equality of the two mean embeddings makes the integrals of equal for every , since the coefficient multiplying this weighted monomial is strictly positive. Hence the integrals of agree for every polynomial . For any continuous test function on and any , choose so that Then has equal integrals under and , and the common support gives Thus all continuous test integrals agree, so , and their second moments agree. The displayed implication is the contrapositive.

For the canonical observed world generated by and , the observational and interventional law formulas and the ratio formula of Assumption 3 hold directly from the product observational density, the target- intervention density, and the identity mixing map. Therefore the two second moments of the ratio laws are and the canonical second-moment contrast is Since this contrast is nonzero, the two second moments differ. The bounded-support recovery fact applied to and gives Thus for every transported ancestral cover, and .

2. The direct-edge separated set is open and dense, and the cover-separated set is open. For a directed edge , put For this paragraph write for the identity-target contrast obtained by intervening on latent coordinate and measuring the ratio coordinate . The canonical relabeling identity is This identity-target contrast is the cube integral The relative product topology of Definition 3, strict positivity of the denominators, and compactness of the cube make this contrast continuous on ; hence is open. For a fixed edge , density is obtained from the edge-specific affine path. Let be a relative neighborhood of . Since , the positivity and smoothness conditions in Definition 3 give strictly positive normalized densities on the latent cube. The identity-target canonical observed world is formed by taking the mixing map to be the identity, taking to be the law with density , and taking environment to be the law with density . With these canonical choices, the observational pushforward, the single-target pushforwards, the Radon–Nikodym identities, and the latent ratio formula listed in Assumption 3 are verified directly; in particular Thus the hypotheses of Lemma 1 are met for . The lemma gives a parameter value on the affine path, arbitrarily close to in the relative topology, for which the identity-target contrast is nonzero. Relabeling the canonical intervention environments back by , every contains a stratum point satisfying Thus each is dense. Since the edge set of is finite and is open and dense. Similarly, for each ordered pair , the map is continuous. Define, for , and set The observational and interventional mean embeddings in Definition 7 have the fixed-cube representations Consequently It remains to justify the displayed continuity. Fix . By strict positivity on the compact cube, there is a relative neighborhood and a number such that, for every , every , and all admissible arguments, The coordinate maps and are continuous in the relative product topology of Definition 3, hence uniformly continuous into on this local neighborhood after shrinking it. Therefore, as , and likewise The ratios above stay in a common compact interval. For the Gaussian feature map, and so is uniformly continuous on that interval. Combining the preceding uniform convergences gives The Bochner integral over the unit cube is continuous under this uniform convergence because and the same estimate holds for . Taking the difference of the two mean embeddings and then the Hilbert norm proves continuity of at . Since was arbitrary, the map is continuous on . A finite intersection of the open sets , indexed by the ancestral covers of Definition 7, gives openness of .

3. Since is dense and , every nonempty relatively open subset of meets . Hence is dense in .

4. The complement of has the asserted topology. Since is open, its complement is closed. Since is dense, the complement has empty interior; being closed with empty interior, it is nowhere dense. Every nowhere dense set is meagre, so is meagre, closed, and nowhere dense.

5. Finally assume has no directed edges. The defining condition for in Definition 5 is then vacuous. Also, if and , then the cover-to-edge implication gives the directed edge , which the no-edge hypothesis rules out. Thus The defining condition for in Definition 7 is vacuous. Thus Combining the preceding conclusions gives all conclusions.

Proof of Theorem 3.

Put , let , and write Throughout the proof the superscript is reserved for the observational environment, while for the notation denotes the single-target intervention environment whose target is ; hence always compares intervention environment with the observational law. The assumptions in Assumptions 1, 2, and 3 make each a probability law. They also give the coherence of with the observed world , in the following concrete sense. If denotes the likelihood-ratio function specified by the world for environment , and if denotes the observed-law likelihood ratio selected from , then The same coherence statement records that selected rank coordinates are functions of the common law family whenever the two worlds have the same observed laws. More precisely, if is compatible with , so that both worlds have the same environment-law family , then for every order , every , and every point in the observed support of , where each side is the selected rank coordinate computed from the same observed law family. Separately, the pointwise ratio formula in Assumption 3, together with the continuous observed-law representative of the likelihood ratio, gives, for every and every , If , all indexed conclusions below are over empty finite sets and the same displayed identities are vacuous; the argument is otherwise unchanged.

The cover-separation premise is stated for the canonical identity-mixing world. For an ancestral cover , Definition 7 gives The transfer from the canonical world to the supplied world is needed only on ancestral covers. Applied to the displayed cover, it rewrites the positive canonical population discrepancy as the corresponding supplied-world observed-law discrepancy, using Assumption 3 for the environment pushforwards, Assumption 2 for the inverse on observed support, and the almost-sure ratio coherence established above. Hence Thus every transported ancestral cover appears as an edge of .

Conversely, Lemma 2 applies under Assumptions 1, 2, and 3. Hence, whenever and is not an ancestor of in , the Gaussian-kernel discrepancy of ratio between the observational environment and intervention environment satisfies The lemma gives this conclusion by marginalizing over the complement of the ancestral set : the observational and intervention- marginal laws agree on the coordinates on which depends, so the two ratio pushforward laws, and therefore their Gaussian mean embeddings, coincide. We now pass from these discrepancy statements to the graph identity. Let denote the environment-label graph If is an edge of , then . The zero-discrepancy conclusion above shows that must be an ancestor of in ; hence every edge of is contained in the transitive closure of . Therefore For the reverse inclusion, take any ancestral relation . In the finite ancestral order, choose a saturated chain where each step is an ancestral cover. By Equation 2, each transported pair is an edge of . Thus reaches in , and Combining Equations 3 and 4 gives The graph reconstruction statement also supplies the decoder’s selected order as a topological ordering of . If is any topological ordering of and , then reaches in by Equation 5; applying the topological-order inequalities along the path gives

Fix such an order . By Equation 5, also respects every transported latent edge. For , set Let with the usual one-point product when . Embed into the full latent cube by the zero-filled map Define the predecessor score map and its range by If , then gives Indeed, every parent of has environment label preceding , hence belongs to ; on the own coordinate and on those parent coordinates the zero-filled point agrees with , and the local dependence of removes all other coordinates. By Lemma 3, applied to the present order , In particular, is compact and the inverse is continuous.

Now work under the target- environment. For , let be the filled-in latent point defined by The marginalization identity in Lemma 4 gives, for , The fiber version in Lemma 5 gives the same factorization after replacing by an arbitrary scalar . These identities are the density input for Lemma 7. Applying that result to the order , the law-selected conditional CDF used by is a continuous conditional-CDF version on the observed conditional-ratio support, and there is a continuous version such that, for all and , where Here is the value of the mechanism at the point obtained from by replacing coordinate with ; the two agree because a mechanism depends only on its own coordinate and its parents. The same cited result states that every continuous conditional-CDF version agrees with the selected version on the observed conditional-ratio support and that the selected version itself satisfies the structural identity at every mixed latent point:

Evaluating Equation 11 at and using the strict monotonicity from Assumption 4 gives The right-hand side is a componentwise continuous, strictly monotone transform of the latent coordinate.

We next prove the pruning characterization used by the decoder. Let For every with , the target assertion is The pruning characterization is supplied by the rank conditional-independence characterization applied at the present order. Its hypotheses are exactly the positivity, causal minimality, shared mixing, one-intervention, sign, and rank-identity clauses already established, together with the subset bookkeeping . It yields

Equation 14 says exactly that the admissible sets are the supersets of . Indeed, by Algorithm 1, a set is admissible precisely when the left side of Equation 14 holds. Thus is admissible by Equation 6, and if is any admissible set then . Conversely every superset of contained in is admissible. Hence is the unique inclusion-minimal admissible parent set for , for every topological ordering of . Applying this extraction to the selected order in Algorithm 1, the parent-pruned graph has edge relation which is exactly

It remains to prove representation uniqueness. Let , , and satisfy the same regularity, causal-minimality, sign, mixing, and one-intervention assumptions, and suppose has the same observed law family as . Write and for the mixing map and target permutation of ; these are in general different from and . Let be the canonical topological ordering of , and define the environment-label order The common-order lemma for equal observed law families gives that is a topological ordering of the common observable ratio graph. Since Equation 5 holds for , the same also respects the transported graph of ; by construction through , it respects the transported graph of .

For any coordinate label , define the signed intervention distribution functions Here is the intervention distribution function for coordinate of , and both signed functions are indexed by the same ambient sign vector . Using the continuous-rank construction at this common order gives, for , Applying Equation 14 to and to at the same order shows that both representations have the same unique minimal admissible parent sets. Therefore, with where and are the target permutations of and , respectively, the map applies the inverse target permutation of first and then the target permutation of , and carries -labels to -labels, one has, for every , Equality of the observational laws gives equality of observed supports, It remains to assemble the common ranks, aligned edges, and common support into componentwise equivalence. We use the following implication, proved in the formal development: if two worlds have the common rank identities in Equation 16, equal observed supports, and the edge alignment in Equation 17, then there are interval maps , one pair for each latent coordinate , such that and, for every observed point , The smoothness of these charts follows by composing the one-coordinate face embedding , the mixing map, the inverse mixing map on the common observed support, and the coordinate projection; the inverse charts are obtained by interchanging the two worlds. Together with Equation 17, this is exactly the componentwise equivalence of and up to the simultaneous relabeling .

When , every edge is an ancestral cover. Therefore Equation 2 gives, for each edge of ,

The established graph equality Equation 15, the order and predecessor containment Equations 5 and 6, the rank identity Equation 12, the uniqueness statement above, and the bivariate positivity Equation 20 supply the packaged comparison clauses: within the positive compact-cube mechanism class with shared mixing, the decoder identifies the latent representation and structure from one perfect intervention per node, supplies the continuous law-separation witness used by von Kügelgen et al. (2023) when , and recovers the graph, topological order, and intervention alignment relative to Wendong et al. (2023) and Yao et al. (2025). Collecting Equations 5, 11, 12, 6, 14, and 15, the coherence of with , the componentwise equivalence statement, Equation 20, and these comparison clauses gives every assertion of Theorem 3.

Proof of Theorem 4.

Let be the Hilbert space of the Gaussian unit-norm feature map . Write so that For , , and an outcome , set Define the evaluation concentration event The independent-environment sampling assumption supplies the conditional evaluation-sampling hypotheses in Lemma 8. Hence, conditionally on the training fold, has probability at least : for every training-fold measurable event ,

For the bridge from fitted to population embeddings, first note that the Gaussian normalization gives, for all , and therefore On , the defining bound gives, for every and , The measurability and integrability requirements here are exactly the training-fold and finite-integral clauses in Assumption 7 and in the definition of .

Now fix and . The empirical and population discrepancies are By the reverse triangle inequality, Thus Applying the conditional bound for to , which is training-fold measurable by Assumption 7, yields for every training-fold measurable . This is the stated conditional probability bound.

Taking gives Since and ,

It remains to identify the graph-theoretic consequences on . Fix , and suppose is selected by . Then and If were not an ancestor of in , the structural Assumptions 1, 2, and 3 imply that the ratio coordinate has the same law under and : Indeed, in the latent factorization, depends only on and the parents of ; an intervention at a non-ancestor of leaves the joint law of these ancestral coordinates unchanged. Therefore The event then gives contradicting the selection inequality. Hence every selected arrow satisfies that is an ancestor of in .

Assume now the cover margin The soundness just proved gives the forward inclusion a selected edge yields an ancestral chain from to in , and applying to each edge of that chain gives a directed path from to in .

For the reverse inclusion, let be an ancestral cover and put The cover relation implies , hence . By the margin and the event , so is selected by . Thus every transported ancestral cover of is an edge of .

Finally, every strict ancestral relation in the finite ancestor order of factors into covers: This follows by induction on the finite interval : if there is nothing to prove, while otherwise one chooses an intermediate with and concatenates the two shorter saturated chains. Hence each edge of , and therefore each path in , maps to a path in . We obtain Together with the forward inclusion,

References

  • von K{\"u}gelgen, Julius and Besserve, Michel and Wendong, Liang and Gresele, Luigi and Keki{\'c}, Armin and Bareinboim, Elias and Blei, David M. and Sch{\"o}lkopf, Bernhard (2023). Nonparametric Identifiability of Causal Representations from Unknown Interventions. Advances in Neural Information Processing Systems 36. doi
  • Wendong, Liang and Keki{\'c}, Armin and von K{\"u}gelgen, Julius and Buchholz, Simon and Besserve, Michel and Gresele, Luigi and Sch{\"o}lkopf, Bernhard (2023). Causal Component Analysis. Advances in Neural Information Processing Systems 36. arXiv
  • Jiang, Yibo and Aragam, Bryon (2023). Learning Nonparametric Latent Causal Graphs with Unknown Interventions. Advances in Neural Information Processing Systems 36. arXiv
  • Buchholz, Simon and Rajendran, Goutham and Rosenfeld, Elan and Aragam, Bryon and Sch{\"o}lkopf, Bernhard and Ravikumar, Pradeep (2023). Learning Linear Causal Representations from Interventions under General Nonlinear Mixing. Advances in Neural Information Processing Systems 36. arXiv
  • Ahuja, Kartik and Mahajan, Divyat and Wang, Yixin and Bengio, Yoshua (2023). Interventional Causal Representation Learning. Proceedings of the 40th International Conference on Machine Learning. arXiv
  • Var{\i}c{\i}, Burak and Acart{\"u}rk, Emre and Shanmugam, Karthikeyan and Kumar, Abhishek and Tajer, Ali (2025). Score-Based Causal Representation Learning: Linear and General Transformations. Journal of Machine Learning Research. arXiv
  • Ng, Ignavier and Xie, Shaoan and Dong, Xinshuai and Spirtes, Peter and Zhang, Kun (2025). Causal Representation Learning from General Environments under Nonparametric Mixing. Proceedings of the 28th International Conference on Artificial Intelligence and Statistics. arXiv
  • Boeken, Philip and Forr{\'e}, Patrick and Mooij, Joris M. (2024). Are Bayesian Networks Typically Faithful?. . arXiv
  • Yao, Dingling and Rancati, Dario and Cadei, Riccardo and Fumero, Marco and Locatello, Francesco (2025). Unifying Causal Representation Learning with the Invariance Principle. International Conference on Learning Representations. arXiv
  • Gretton, Arthur and Borgwardt, Karsten M. and Rasch, Malte J. and Sch{\"o}lkopf, Bernhard and Smola, Alexander J. (2012). A Kernel Two-Sample Test. Journal of Machine Learning Research.
  • Schrab, Antonin and Kim, Ilmun and Albert, M{\'e}lisande and Laurent, B{\'e}atrice and Guedj, Benjamin and Gretton, Arthur (2023). {MMD} Aggregated Two-Sample Test. Journal of Machine Learning Research. arXiv
  • Xie, Haitian (2023). Uniform Convergence Results for the Local Linear Regression Estimation of the Conditional Distribution. Statistics and Probability Letters. doi
  • Mammen, Enno and Rothe, Christoph and Schienle, Melanie (2012). Nonparametric Regression with Nonparametrically Generated Covariates. The Annals of Statistics. doi
  • Lee, Ying-Ying (2018). Partial Mean Processes with Generated Regressors: Continuous Treatment Effects and Nonseparable Models. . arXiv
  • Jaber, Amin and Kocaoglu, Murat and Shanmugam, Karthikeyan and Bareinboim, Elias (2020). Causal Discovery from Soft Interventions with Unknown Targets: Characterization and Learning. Advances in Neural Information Processing Systems 33.
  • Acart{\"u}rk, Emre and Var{\i}c{\i}, Burak and Shanmugam, Karthikeyan and Tajer, Ali (2024). Sample Complexity of Interventional Causal Representation Learning. Advances in Neural Information Processing Systems 37. doi
  • Lee, Inbeom and Jin, Tongtong and Aragam, Bryon (2026). Beyond Identifiability: Learning Causal Representations with Few Environments and Finite Samples. . arXiv
  • Spirtes, Peter and Glymour, Clark and Scheines, Richard (2000). Causation, Prediction, and Search. MIT Press. doi
  • Kulkarni, Pranamya and Datta, Puranjay and Var{\i}c{\i}, Burak and Acart{\"u}rk, Emre and Shanmugam, Karthikeyan and Tajer, Ali (2025). {ROPES}: Robotic Pose Estimation via Score-Based Causal Representation Learning. . arXiv
  • Pearl, Judea (2009). Causality: Models, Reasoning, and Inference. Cambridge University Press.
  • Peters, Jonas and Janzing, Dominik and Sch{\"o}lkopf, Bernhard (2017). Elements of Causal Inference: Foundations and Learning Algorithms. MIT Press.
  • Sch{\"o}lkopf, Bernhard and Locatello, Francesco and Bauer, Stefan and Ke, Nan Rosemary and Kalchbrenner, Nal and Goyal, Anirudh and Bengio, Yoshua (2021). Toward Causal Representation Learning. Proceedings of the IEEE. doi
  • Hyv{\"a}rinen, Aapo and Karhunen, Juha and Oja, Erkki (2001). Independent Component Analysis. Wiley. doi
  • Hyv{\"a}rinen, Aapo and Morioka, Hiroshi (2016). Unsupervised Feature Extraction by Time-Contrastive Learning and Nonlinear {ICA}. Advances in Neural Information Processing Systems 29. arXiv
  • Hyv{\"a}rinen, Aapo and Sasaki, Hiroaki and Turner, Richard E. (2019). Nonlinear {ICA} Using Auxiliary Variables and Generalized Contrastive Learning. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics. arXiv
  • Khemakhem, Ilyes and Kingma, Diederik P. and Monti, Ricardo Pio and Hyv{\"a}rinen, Aapo (2020). Variational Autoencoders and Nonlinear {ICA}: A Unifying Framework. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics. arXiv
  • Hyv{\"a}rinen, Aapo and Khemakhem, Ilyes and Morioka, Hiroshi (2023). Nonlinear Independent Component Analysis for Principled Disentanglement in Unsupervised Deep Learning. Patterns. doi
  • Brehmer, Johann and de Haan, Pim and Lippe, Phillip and Cohen, Taco (2022). Weakly Supervised Causal Representation Learning. Advances in Neural Information Processing Systems 35. arXiv
  • Lippe, Phillip and Magliacane, Sara and L{\"o}we, Sindy and Asano, Yuki M. and Cohen, Taco and Gavves, Efstratios (2022). {CITRIS}: Causal Identifiability from Temporal Intervened Sequences. Proceedings of the 39th International Conference on Machine Learning. arXiv
  • Squires, Chandler and Seigal, Anna and Bhate, Salil and Uhler, Caroline (2022). Linear Causal Disentanglement via Interventions. . arXiv
  • Zhang, Jiaqi and Squires, Chandler and Greenewald, Kristjan and Srivastava, Akash and Shanmugam, Karthikeyan and Uhler, Caroline (2023). Identifiability Guarantees for Causal Disentanglement from Soft Interventions. Advances in Neural Information Processing Systems 36. arXiv
  • Kivva, Bohdan and Rajendran, Goutham and Ravikumar, Pradeep and Aragam, Bryon (2021). Learning Latent Causal Graphs via Mixture Oracles. Advances in Neural Information Processing Systems 34. arXiv
  • Xie, Feng and Cai, Ruichu and Huang, Biwei and Glymour, Clark and Hao, Zhifeng and Zhang, Kun (2024). Generalized Independent Noise Condition for Estimating Latent Variable Causal Graphs. Journal of Machine Learning Research. arXiv
  • Meek, Christopher (1995). Strong Completeness and Faithfulness in Bayesian Networks. Proceedings of the Eleventh Conference on Uncertainty in Artificial Intelligence. arXiv
  • Hauser, Alain and B{\"u}hlmann, Peter (2012). Characterization and Greedy Learning of Interventional Markov Equivalence Classes of Directed Acyclic Graphs. Journal of Machine Learning Research. arXiv
  • Yang, Karren D. and Katcoff, Abigail and Uhler, Caroline (2018). Characterizing and Learning Equivalence Classes of Causal {DAG}s under Interventions. Proceedings of the 35th International Conference on Machine Learning. arXiv
  • Sriperumbudur, Bharath K. and Gretton, Arthur and Fukumizu, Kenji and Sch{\"o}lkopf, Bernhard and Lanckriet, Gert R. G. (2010). Hilbert Space Embeddings and Metrics on Probability Measures. Journal of Machine Learning Research. arXiv
  • Sriperumbudur, Bharath K. and Fukumizu, Kenji and Lanckriet, Gert R. G. (2011). Universality, Characteristic Kernels and {RKHS} Embedding of Measures. Journal of Machine Learning Research. arXiv
  • Gretton, Arthur and Borgwardt, Karsten M. and Rasch, Malte J. and Sch{\"o}lkopf, Bernhard and Smola, Alexander J. (2007). A Kernel Method for the Two-Sample-Problem. Advances in Neural Information Processing Systems 19. arXiv
  • Smola, Alexander J. and Gretton, Arthur and Song, Le and Sch{\"o}lkopf, Bernhard (2007). A Hilbert Space Embedding for Distributions. Algorithmic Learning Theory. doi
  • Fan, Jianqing and Gijbels, Irene (1996). Local Polynomial Modelling and Its Applications. Chapman and Hall/CRC.
  • Fan, Jianqing and Yao, Qiwei and Tong, Howell (1996). Estimation of Conditional Densities and Sensitivity Measures in Nonlinear Dynamical Systems. Biometrika. doi
  • Pagan, Adrian (1984). Econometric Issues in the Analysis of Regressions with Generated Regressors. International Economic Review. doi
  • Var{\i}c{\i}, Burak and Acart{\"u}rk, Emre and Shanmugam, Karthikeyan and Tajer, Ali (2024). General Identifiability and Achievability for Causal Representation Learning. Proceedings of the 27th International Conference on Artificial Intelligence and Statistics. arXiv
  • Bing, Simon and Ninad, Urmi and Wahl, Jonas and Runge, Jakob (2024). Identifying Linearly-Mixed Causal Representations from Multi-Node Interventions. Proceedings of the Third Conference on Causal Learning and Reasoning. arXiv
  • Jin, Jikai and Syrgkanis, Vasilis (2024). Learning Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity. Advances in Neural Information Processing Systems 37. arXiv
  • Zhang, Kun and Xie, Shaoan and Ng, Ignavier and Zheng, Yujia (2024). Causal Representation Learning from Multiple Distributions: A General Setting. Proceedings of the 41st International Conference on Machine Learning. arXiv
  • Ahuja, Kartik and Mansouri, Amin and Wang, Yixin (2024). Multi-Domain Causal Representation Learning via Weak Distributional Invariances. Proceedings of the 27th International Conference on Artificial Intelligence and Statistics. arXiv
  • Dai, Haoyue and Ng, Ignavier and Sun, Jianle and Tang, Zeyu and Luo, Gongxu and Dong, Xinshuai and Spirtes, Peter and Zhang, Kun (2025). When Selection Meets Intervention: Additional Complexities in Causal Discovery. . arXiv
  • Zhang, Mingxuan and Desai, Khushi and Kevlishvili, Sopho and Azizi, Elham (2026). Scalable Contrastive Causal Discovery under Unknown Soft Interventions. . arXiv
  • Chen, Guangyi and Deng, Yunlong and Zhu, Peiyuan and Li, Yan and Sheng, Yifan and Li, Zijian and Zhang, Kun (2025). {CausalVerse}: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations. . arXiv
  • Wang, Kun and Varambally, Sumanth and Watson-Parris, Duncan and Ma, Yi-An and Yu, Rose (2024). Discovering Latent Causal Graphs from Spatiotemporal Data. . arXiv
  • Sun, Yuewen and Kong, Lingjing and Chen, Guangyi and Li, Loka and Luo, Gongxu and Li, Zijian and Zhang, Yixuan and Zheng, Yujia and Yang, Mengyue and Stojanov, Petar and Segal, Eran and Xing, Eric P. and Zhang, Kun (2024). Causal Representation Learning from Multimodal Biomedical Observations. . arXiv