CausalSmith · seminar slides

One Intervention per Latent Variable

Under smooth invertible observations, compact overlap, causal minimality, and fixed intervention-sign strata, one perfect intervention per scalar latent variable recovers the latent DAG, target alignment, and representation up to relabeling and componentwise smooth coordinate changes.

Overview

  • We observe an observational environment and nn interventional environments.
  • Each intervention perfectly replaces one scalar latent mechanism.
  • The intervention targets are unknown.
  • The observations are nonlinear mixtures of the latent variables through one shared diffeomorphism.
  • We use likelihood-ratio laws across environments to recover ancestry, ranks, target labels, and parents.

Motivation

  • Econometric environments often shift one latent structural component at a time.
  • The analyst sees high-dimensional observables, not the structural variables.
  • The labels of the shifted components may be unknown.
  • The central question is whether one shifted environment per latent variable carries enough information for nonlinear causal representation learning.
  • Our answer is affirmative on a smooth compact-support regime with generic ratio-law separation.

Setup

  • V=(V1,,Vn)V=(V_1,\ldots,V_n) is the scalar latent causal vector on a DAG GG.
  • X=f(V)X=f(V) is observed through the same smooth invertible map in every environment.
  • P0P^0 is observational; P1,,PnP^1,\ldots,P^n are single-target intervention laws.
  • π\pi, the unknown target permutation, maps environment labels to latent targets.
  • Ri=dPi/dP0R_i=dP^i/dP^0, the likelihood ratio for environment ii, is the observable one-dimensional coordinate we compare.
Assumption A-3 (Shared diffeomorphic mixing)

The same map ff is a C2C^2 diffeomorphism from [0,1]n[0,1]^n onto X\mathcal X in every environment.

Intervention Design

  • A perfect intervention replaces the target conditional density by a parent-independent density.
  • In latent coordinates, the ratio for environment ee depends on the target π(e)\pi(e) and its parents.
  • The ratio therefore carries local graph information through its distribution across environments.
Assumption A-4 (Perfect intervention laws)

For the mechanism θ\theta, the shared mixing map ff, the observed laws P0,,PnP^0,\ldots,P^n, the target permutation π\pi, and the supplied likelihood ratios R1,,RnR_1,\ldots,R_n, the following conditions hold:

  • (Observational pushforward.) The observational law satisfies P0=f#(p0(v)dv),p0(v)=i=1npi(vivpaG(i)), P^0 = f_\#\bigl(p^0(v)\,dv\bigr), \qquad p^0(v)=\prod_{i=1}^n p_i(v_i\mid v_{\operatorname{pa}_G(i)}), with v[0,1]nv\in[0,1]^n.
  • (Single-target pushforwards.) For every environment e[n]e\in[n], Pe=f#(pπ(e)(v)dv),pπ(e)(v)=qπ(e)(vπ(e))π(e)p(vvpaG()). P^e = f_\#\bigl(p^{\pi(e)}(v)\,dv\bigr), \qquad p^{\pi(e)}(v) = q_{\pi(e)}(v_{\pi(e)}) \prod_{\ell\ne \pi(e)} p_\ell(v_\ell\mid v_{\operatorname{pa}_G(\ell)}).
  • (Radon--Nikodym ratios.) For every e[n]e\in[n], Re(X)=dPedP0(X)P0-almost surely. R_e(X)=\frac{dP^e}{dP^0}(X) \quad\text{\(P^0\)-almost surely.}
  • (Latent ratio formula.) For every e[n]e\in[n] and every v[0,1]nv\in[0,1]^n, Re(f(v))=qπ(e)(vπ(e))pπ(e)(vπ(e)vpaG(π(e))). R_e(f(v)) = \frac{q_{\pi(e)}(v_{\pi(e)})} {p_{\pi(e)}(v_{\pi(e)}\mid v_{\operatorname{pa}_G(\pi(e))})}.

Assumptions

  • Smooth positive mechanisms give common compact support and well-defined ratios.
  • Fixed own-coordinate derivative signs make each target ratio monotone in its own latent coordinate.
  • Causal minimality makes every graph parent statistically active under the observational law.
  • These conditions define the fixed-sign model stratum ΘG,s\Theta_{G,s}.
Assumption A-5 (Fixed derivative sign)

For every i[n]i\in[n], sivilog{qi(vi)/pi(vivpaG(i))}>0 s_i\,\partial_{v_i}\log\{q_i(v_i)/p_i(v_i\mid v_{\operatorname{pa}_G(i)})\}>0 throughout the closed cube.

Key Idea

  • Compare the law of RiR_i under P0P^0 with its law under each PjP^j.
  • If intervening on jj changes the law of RiR_i, draw jij\to i in the ratio-discrepancy graph.
  • Gaussian maximum mean discrepancy makes this comparison observable from one-dimensional ratio samples.
  • A positive discrepancy reveals ancestral information after the target permutation.
Environment laws P⁰ and Pʲ law of Rᵢ Likelihood ratios one-dimensional ratio samples Gaussian MMD compare ratio laws observable discrepancy Discrepancy graph draw j→i changed law Rank decoding ancestral information target permutation
illustrative Box-and-arrow schematic showing observed environment laws feeding likelihood ratios, likelihood ratios feeding Gaussian MMD comparisons, comparisons feeding a ratio-discrepancy graph, and the graph feeding rank decoding.

Generic Separation

informal · Theorem T-1 A smooth three-node separating witness has strictly positive direct-edge ratio discrepancy, while a smooth cancellation witness has zero discrepancy on the same edge.

  • The witness pair shows both strict separation and exact cancellation inside regular mechanisms.
  • Analytic perturbations move mechanisms away from the cancellation boundary.
  • This supports a topological genericity statement in each fixed-sign stratum.

informal · Theorem T-2 In every nonempty fixed-sign causal-minimal stratum, Gaussian ratio-law separation of transported ancestral covers holds on an open dense set.

Decoder

  • Build HD={ji:ji and Dji>0}H_D=\{j\to i:j\ne i\text{ and }D_{ji}>0\}, the observable ratio-discrepancy graph.
  • Choose a topological order of HDH_D.
  • For each node, condition its log-ratio on earlier log-ratios.
  • Convert the conditional log-ratio into a rank coordinate.
  • Prune parents by conditional independence under the observational law.
Definition P-7 (Population decoder $\mathscr D$)

Given observed laws (P0,,Pn)(P^0,\ldots,P^n), the population decoder D\mathscr D is defined by the following steps.

  1. Compute the likelihood ratios and log-ratios Ri=dPidP0,Li=logRi,i[n], R_i=\frac{dP^i}{dP^0}, \qquad L_i=\log R_i, \qquad i\in[n], and the Gaussian-MMD discrepancies DjiD_{ji}.
  2. Form the observable directed ratio-discrepancy graph HD={ji:ji and Dji>0}. H_D=\{j\to i:j\ne i\text{ and }D_{ji}>0\}.
  3. Choose a topological ordering of HDH_D. For each ii, let BiB_i be the predecessors of ii in that ordering and define Ui=Ci(LiLBi). U_i=C_i(L_i\mid L_{B_i}).
  4. Align environment ii with coordinate UiU_i.
  5. For each ii, range over ABiA\subseteq B_i and identify the unique inclusion-minimal subset satisfying Ui ⁣ ⁣ ⁣P0UBiAUA. U_i\mathbin{\perp\!\!\!\perp}_{P^0}U_{B_i\setminus A}\mid U_A. The output is the environment-label DAG GDπ={ai: aAimin},Aimin=min{ABi:Ui ⁣ ⁣ ⁣P0UBiAUA}. G^\pi_{\mathscr D} = \{a\to i:\ a\in A_i^{\min}\}, \qquad A_i^{\min} = \min_{\subseteq} \left\{ A\subseteq B_i: U_i\mathbin{\perp\!\!\!\perp}_{P^0}U_{B_i\setminus A}\mid U_A \right\}.

Main Result

informal · Theorem T-3 Under the stated smoothness and separation conditions, the decoder recovers the transported DAG, target alignment, and latent coordinates up to allowed componentwise changes.

  • The ratio graph gives the ancestral order.
  • Conditional ranks recover monotone transforms of the intervened latent variables.
  • Conditional-independence pruning turns ancestry into exact parents.
  • The recovered graph is the environment-label DAG GπG^\pi.

Comparison

  • von Kügelgen et al. (2023) identify the bivariate unknown-target case from one perfect intervention per node under a law-separation condition.
  • In arbitrary dimension, their paired-intervention theorem uses two perfect interventions per node.
  • On the overlapping positive C3C^3 compact-cube faithful regime, our ratio-law route gives one perfect intervention per node in every dimension.
  • Wendong et al. (2023) and Yao et al. (2025) use supplied graph, order, or target-alignment inputs; here the ratio laws recover those objects.

Why Ranks Work

  • The fixed-sign condition makes the target log-ratio monotone in its own latent coordinate.
  • Once earlier ratio coordinates encode predecessor information, conditioning removes parent variation.
  • The conditional distribution transform maps the target log-ratio to the intervention distribution rank.
  • Thus UiU_i agrees with a monotone transform of Vπ(i)V_{\pi(i)}.
  • This converts an observed likelihood-ratio object into a latent coordinate.

Why Parents Prune Exactly

  • The ratio graph may contain ancestral arrows rather than parent arrows.
  • After rank recovery, the UiU_i's behave like componentwise transforms of the latent variables.
  • Under causal minimality and positivity, conditioning on the true parents screens off earlier nonparents.
  • Any missing parent leaves a conditional dependence.
  • The unique minimal admissible predecessor set is therefore the transported parent set.

Confidence Edges

  • The finite-sample layer splits data into training and evaluation folds.
  • Training folds estimate ratios R^i\widehat R_i.
  • Independent evaluation folds compute empirical Gaussian-MMD discrepancies D^ji\widehat D_{ji}.
  • Thresholding D^ji\widehat D_{ji} by εN(α,η)\varepsilon_N(\alpha,\eta) gives confidence edges for ancestry.
Definition P-8 (Sample-split confidence graph $\widehat H$)

Under Assumption A-6, with training and evaluation folds and tuning levels α,η\alpha,\eta, the sample-split confidence graph H^\widehat H is defined by the following steps.

  1. On the training folds, fit likelihood-ratio estimators R^i,i[n]. \widehat R_i, \qquad i\in[n].
  2. On the independent evaluation folds, compute the empirical Gaussian-MMD discrepancies D^ji,j,i[n], ji, \widehat D_{ji}, \qquad j,i\in[n],\ j\ne i, using R^i\widehat R_i.
  3. Output the directed graph H^={ji:D^jiεN(α,η)>0}. \widehat H = \left\{ j\to i: \widehat D_{ji}-\varepsilon_N(\alpha,\eta)>0 \right\}.

Statistical Guarantee

informal · Theorem T-4 Conditional on a simultaneous first-stage L1L^1 ratio-error event, the sample-split MMD event holds with probability at least 1α1-\alpha, selected arrows are true transported ancestral discoveries, and separated ancestral covers recover the transported transitive closure.

  • The guarantee is simultaneous over all ordered pairs.
  • Marginally, the MMD event has probability at least 1αη1-\alpha-\eta.
  • When every transported ancestral cover clears the stated radius, the confidence graph recovers the same transitive closure as the population decoder.

Future Work

  • The next statistical target is coordinate uncertainty for the recovered rank variables.
  • The bounded Hölder subclass records smoothness, support, margin, geometry, regularity, and design envelopes.
  • The proposed handle adapts local-linear conditional-CDF estimation to generated log-ratio responses and generated conditioning covariates.
  • The frontier rate is stated as a uniform generated-rank target on compact interior sets.

informal · Definition P-12 Under the stated first-stage contract and bandwidth scale, the target is a cross-fitted generated-rank estimator with sup-norm error at most rNr_N uniformly over the bounded Hölder subclass.

Conclusion

  • One unknown-target perfect intervention per scalar latent variable can identify nonlinear causal representations in the fixed-sign compact-support regime.
  • Generic Gaussian ratio-law separation supplies the observable ancestral order.
  • Conditional ranks align intervention labels with latent coordinates.
  • Observational conditional independence prunes ancestry to the exact transported DAG.
  • Sample splitting adds simultaneous confidence edges for ancestral discoveries under a first-stage ratio-error contract.