CausalSmith · seminar slides
One Intervention per Latent Variable
Under smooth invertible observations, compact overlap, causal minimality, and fixed intervention-sign strata, one perfect intervention per scalar latent variable recovers the latent DAG, target alignment, and representation up to relabeling and componentwise smooth coordinate changes.
Overview
- We observe an observational environment and n interventional environments.
- Each intervention perfectly replaces one scalar latent mechanism.
- The intervention targets are unknown.
- The observations are nonlinear mixtures of the latent variables through one shared diffeomorphism.
- We use likelihood-ratio laws across environments to recover ancestry, ranks, target labels, and parents.
Motivation
- Econometric environments often shift one latent structural component at a time.
- The analyst sees high-dimensional observables, not the structural variables.
- The labels of the shifted components may be unknown.
- The central question is whether one shifted environment per latent variable carries enough information for nonlinear causal representation learning.
- Our answer is affirmative on a smooth compact-support regime with generic ratio-law separation.
Setup
- V=(V1,…,Vn) is the scalar latent causal vector on a DAG G.
- X=f(V) is observed through the same smooth invertible map in every environment.
- P0 is observational; P1,…,Pn are single-target intervention laws.
- π, the unknown target permutation, maps environment labels to latent targets.
- Ri=dPi/dP0, the likelihood ratio for environment i, is the observable one-dimensional coordinate we compare.
The same map f is a C2 diffeomorphism from [0,1]n onto X in every environment.
Intervention Design
- A perfect intervention replaces the target conditional density by a parent-independent density.
- In latent coordinates, the ratio for environment e depends on the target π(e) and its parents.
- The ratio therefore carries local graph information through its distribution across environments.
For the mechanism θ, the shared mixing map f, the observed laws P0,…,Pn, the target permutation π, and the supplied likelihood ratios R1,…,Rn, the following conditions hold:
- (Observational pushforward.) The observational law satisfies P0=f#(p0(v)dv),p0(v)=i=1∏npi(vi∣vpaG(i)), with v∈[0,1]n.
- (Single-target pushforwards.) For every environment e∈[n], Pe=f#(pπ(e)(v)dv),pπ(e)(v)=qπ(e)(vπ(e))ℓ=π(e)∏pℓ(vℓ∣vpaG(ℓ)).
- (Radon--Nikodym ratios.) For every e∈[n], Re(X)=dP0dPe(X)P0-almost surely.
- (Latent ratio formula.) For every e∈[n] and every v∈[0,1]n, Re(f(v))=pπ(e)(vπ(e)∣vpaG(π(e)))qπ(e)(vπ(e)).
Assumptions
- Smooth positive mechanisms give common compact support and well-defined ratios.
- Fixed own-coordinate derivative signs make each target ratio monotone in its own latent coordinate.
- Causal minimality makes every graph parent statistically active under the observational law.
- These conditions define the fixed-sign model stratum ΘG,s.
For every i∈[n], si∂vilog{qi(vi)/pi(vi∣vpaG(i))}>0 throughout the closed cube.
Key Idea
- Compare the law of Ri under P0 with its law under each Pj.
- If intervening on j changes the law of Ri, draw j→i in the ratio-discrepancy graph.
- Gaussian maximum mean discrepancy makes this comparison observable from one-dimensional ratio samples.
- A positive discrepancy reveals ancestral information after the target permutation.
Generic Separation
informal · Theorem T-1 A smooth three-node separating witness has strictly positive direct-edge ratio discrepancy, while a smooth cancellation witness has zero discrepancy on the same edge.
- The witness pair shows both strict separation and exact cancellation inside regular mechanisms.
- Analytic perturbations move mechanisms away from the cancellation boundary.
- This supports a topological genericity statement in each fixed-sign stratum.
informal · Theorem T-2 In every nonempty fixed-sign causal-minimal stratum, Gaussian ratio-law separation of transported ancestral covers holds on an open dense set.
Decoder
- Build HD={j→i:j=i and Dji>0}, the observable ratio-discrepancy graph.
- Choose a topological order of HD.
- For each node, condition its log-ratio on earlier log-ratios.
- Convert the conditional log-ratio into a rank coordinate.
- Prune parents by conditional independence under the observational law.
Given observed laws (P0,…,Pn), the population decoder D is defined by the following steps.
- Compute the likelihood ratios and log-ratios Ri=dP0dPi,Li=logRi,i∈[n], and the Gaussian-MMD discrepancies Dji.
- Form the observable directed ratio-discrepancy graph HD={j→i:j=i and Dji>0}.
- Choose a topological ordering of HD. For each i, let Bi be the predecessors of i in that ordering and define Ui=Ci(Li∣LBi).
- Align environment i with coordinate Ui.
- For each i, range over A⊆Bi and identify the unique inclusion-minimal subset satisfying Ui⊥⊥P0UBi∖A∣UA. The output is the environment-label DAG GDπ={a→i: a∈Aimin},Aimin=⊆min{A⊆Bi:Ui⊥⊥P0UBi∖A∣UA}.
Main Result
informal · Theorem T-3 Under the stated smoothness and separation conditions, the decoder recovers the transported DAG, target alignment, and latent coordinates up to allowed componentwise changes.
- The ratio graph gives the ancestral order.
- Conditional ranks recover monotone transforms of the intervened latent variables.
- Conditional-independence pruning turns ancestry into exact parents.
- The recovered graph is the environment-label DAG Gπ.
Comparison
- von Kügelgen et al. (2023) identify the bivariate unknown-target case from one perfect intervention per node under a law-separation condition.
- In arbitrary dimension, their paired-intervention theorem uses two perfect interventions per node.
- On the overlapping positive C3 compact-cube faithful regime, our ratio-law route gives one perfect intervention per node in every dimension.
- Wendong et al. (2023) and Yao et al. (2025) use supplied graph, order, or target-alignment inputs; here the ratio laws recover those objects.
Why Ranks Work
- The fixed-sign condition makes the target log-ratio monotone in its own latent coordinate.
- Once earlier ratio coordinates encode predecessor information, conditioning removes parent variation.
- The conditional distribution transform maps the target log-ratio to the intervention distribution rank.
- Thus Ui agrees with a monotone transform of Vπ(i).
- This converts an observed likelihood-ratio object into a latent coordinate.
Why Parents Prune Exactly
- The ratio graph may contain ancestral arrows rather than parent arrows.
- After rank recovery, the Ui's behave like componentwise transforms of the latent variables.
- Under causal minimality and positivity, conditioning on the true parents screens off earlier nonparents.
- Any missing parent leaves a conditional dependence.
- The unique minimal admissible predecessor set is therefore the transported parent set.
Confidence Edges
- The finite-sample layer splits data into training and evaluation folds.
- Training folds estimate ratios Ri.
- Independent evaluation folds compute empirical Gaussian-MMD discrepancies Dji.
- Thresholding Dji by εN(α,η) gives confidence edges for ancestry.
Under Assumption A-6, with training and evaluation folds and tuning levels α,η, the sample-split confidence graph H is defined by the following steps.
- On the training folds, fit likelihood-ratio estimators Ri,i∈[n].
- On the independent evaluation folds, compute the empirical Gaussian-MMD discrepancies Dji,j,i∈[n], j=i, using Ri.
- Output the directed graph H={j→i:Dji−εN(α,η)>0}.
Statistical Guarantee
informal · Theorem T-4 Conditional on a simultaneous first-stage L1 ratio-error event, the sample-split MMD event holds with probability at least 1−α, selected arrows are true transported ancestral discoveries, and separated ancestral covers recover the transported transitive closure.
- The guarantee is simultaneous over all ordered pairs.
- Marginally, the MMD event has probability at least 1−α−η.
- When every transported ancestral cover clears the stated radius, the confidence graph recovers the same transitive closure as the population decoder.
Future Work
- The next statistical target is coordinate uncertainty for the recovered rank variables.
- The bounded Hölder subclass records smoothness, support, margin, geometry, regularity, and design envelopes.
- The proposed handle adapts local-linear conditional-CDF estimation to generated log-ratio responses and generated conditioning covariates.
- The frontier rate is stated as a uniform generated-rank target on compact interior sets.
informal · Definition P-12 Under the stated first-stage contract and bandwidth scale, the target is a cross-fitted generated-rank estimator with sup-norm error at most rN uniformly over the bounded Hölder subclass.
Conclusion
- One unknown-target perfect intervention per scalar latent variable can identify nonlinear causal representations in the fixed-sign compact-support regime.
- Generic Gaussian ratio-law separation supplies the observable ancestral order.
- Conditional ranks align intervention labels with latent coordinates.
- Observational conditional independence prunes ancestry to the exact transported DAG.
- Sample splitting adds simultaneous confidence edges for ancestral discoveries under a first-stage ratio-error contract.