CausalSmith · seminar slides
Hidden-State Overlap in POMDP Policy Evaluation
We characterize the minimax risk for stationary off-policy evaluation from one finite POMDP trajectory, and a radius-calibrated partial-history estimator attains the full overlap frontier.
slides for Latent Stationary Overlap and Minimax Off-policy Evaluation in Finite POMDPs
Overview
- We observe one stationary behavior trajectory with observed state, action, and reward.
- The Markov state has an unobserved component.
- Behavior and target policies are known functions of the observed state.
- We study the stationary target-policy mean reward.
- The key population control is latent stationary overlap: target occupancy is bounded by behavior occupancy on the full observed-latent state.
- The risk surface depends on the overlap-distance radius qC, mixing scale t0, and log-overlap scale ζ.
informal · Theorem T-5 Uniformly over finite observed and latent alphabets, the minimax risk is comparable to the all-radius frontier, and a radius-calibrated PHIW estimator attains the upper bound.
Motivation
- In mobile health and treatment-policy evaluation, we often observe a coarse context but miss clinically relevant latent state.
- A target policy may change the long-run mix of latent states, even when actions are randomized with known probabilities.
- Current-action weighting corrects the action choice at the current observed state.
- Longer observable histories can partially recover the latent occupancy shift created by the target policy.
- The central question is: how much history is statistically worth using?
Running Example
- Think of an insulin policy that treats when observed glucose exceeds a threshold.
- Diet history is latent, but it affects future glucose and treatment needs.
- The behavior policy randomizes treatment, while the target policy follows the glucose threshold.
- One-step policy overlap controls action probabilities at observed glucose levels.
- Latent stationary overlap controls how far the target policy can move the stationary diet-glucose distribution.
- The theory chooses history length from that stationary-overlap radius.
Setup
- The joint state is St=(Xt,Ht): observed state Xt and latent state Ht.
- The behavior policy b generates the data; the target policy e defines the counterfactual value.
- We observe (Xt,At,Yt) for t=0,…,T−1.
- Rewards are bounded, and the trajectory starts in the behavior stationary law.
- The target estimand is the stationary mean reward under the target policy.
For the joint state St and the target-policy reward regression ge, θ(K,e):=s=(x,h)∈S∑de(s)ge(s)=s=(x,h)∈S∑de(s)a∈A∑e(a∣x)EK[Yt∣St=s,At=a].
Assumptions
- The process is a finite POMDP with a time-homogeneous reward-transition kernel.
- Sequential ignorability holds at the observed state: actions are drawn from b(⋅∣Xt).
- One-step policy overlap bounds e(a∣x)/b(a∣x) by L=exp(ζ).
- Uniform contraction gives geometric forgetting under both behavior and target policies.
- Latent stationary overlap bounds the target stationary law de by Cdb on the full joint state.
For the overlap constant C and experiment M, the latent stationary overlap condition holds when C≥1 and, writing de and db for the stationary laws of the target-policy and behavior-policy joint-state transition kernels, respectively, de(s)≤Cdb(s),s∈S.
Estimator
- Partial-history importance weighting uses only the last k+1 observable policy ratios.
- Larger k reduces latent-history bias through mixing.
- Larger k increases variance through products of one-step ratios.
- Clipping keeps the estimate on the reward scale.
For 0≤k<T, define Zt(k):=Ytj=t−k∏tρj, where ρt is the observable one-step policy ratio with the zero-cell convention. The clipped partial-history importance-weighted estimator is θk:=clip[−1,1]{T−k1t=k∑T−1Zt(k)}.
Radius Calibration
- The normalized overlap-distance radius is qC=(C−1)/C.
- At qC=0, target and behavior stationary occupancies coincide.
- As qC grows, longer observable histories become useful.
- The supplied radius C determines the depth used by the estimator.
For C≥1, define qC:=CC−1.
For T∈N and C≥1, define κTrad(C):=⎩⎨⎧0,min{⌊2T⌋,⌊2log(1/α)+logLlog(TqC2)⌋},TqC2≤1,TqC2>1.
Main Result
- The frontier has a parametric term and a hidden-state term.
- The elbow occurs when the overlap-distance radius is on the T−1/2 scale.
- The constants depend on t0 and ζ, while the statement is uniform over finite alphabets.
Fix t0>0 and ζ>0. Let RT(t0,ζ,C) be the minimax squared-error risk in Definition P-3, let qC=(C−1)/C be the overlap-distance radius in Definition P-10, and set β=2+t0ζ2,rT(t0,ζ,C)=T−1+T−βqC2(1−β). There exist constants clo,chi>0, with clo≤chi, and an integer T⋆≥1, all depending only on t0 and ζ, such that for every C≥1 and every T≥T⋆, clorT(t0,ζ,C)≤RT(t0,ζ,C)≤chirT(t0,ζ,C), and the radius-adaptive partial-history importance-weighted estimator with depth κTrad(C) from Definition P-11 satisfies m∈MT(t0,ζ,C)supEm[(θκTrad(C)−θm)2]≤chirT(t0,ζ,C). Moreover, for every sequence CT≥1, if TqCT2 is eventually bounded, then there is a constant cimm>0, depending on the sequence and on t0,ζ, such that eventually m∈MT(t0,ζ,CT)supEm[(θ0−θm)2]≤Tcimm.
Fixed-radius Slice
informal · Theorem T-1 For every fixed C>1, the minimax risk is between constants times T−β, and clipped PHIW at the balanced depth attains the upper bound.
Fix t0,ζ,C∈R. Assume:
- (Positive memory exponent.) t0>0.
- (Positive overlap exponent.) ζ>0.
- (Fixed latent-overlap radius.) C>1.
Let the exponent β be β=2+t0ζ2. There exist constants c⋆,c⋆∈R and an integer T⋆≥1 such that 0<c⋆≤c⋆ and, for every integer T≥T⋆, c⋆T−β≤RT(t0,ζ,C)≤c⋆T−β, where RT(t0,ζ,C) is the minimax squared-error risk in Definition P-3. Moreover, the clipped partial-history importance-weighted estimator θkT, with kT as in Definition P-5, satisfies M∈MT(t0,ζ,C)supEM[{θkT−θ(M)}2]≤c⋆T−β, with θ(M) the stationary target-policy mean reward in Definition P-2. For each integer T≥T⋆, there is an integer Q≥1 such that the lower bound is witnessed by the two signed-depth alternatives MQ,vT, v∈{−1,+1}, generated by Definition P-6 with one observed state and 2(Q+1) latent states: for every measurable observable-data estimator θ, under the usual squared-error integrability conditions for both alternatives, c⋆T−β≤v∈{−1,+1}maxEMQ,vT[{θ−θ(MQ,vT)}2].
Boundary Slice
informal · Theorem T-2 At unit latent overlap, the behavior and target stationary laws coincide, immediate weighting has at most order T−1 risk, and the minimax risk is on the T−1 scale.
Let t0>0 and ζ>0. At unit latent overlap, C=1, the following assertions hold. For every trajectory length T, every finite observed and hidden state cardinalities, and every experiment M∈MT(t0,ζ,1) in the model class of Definition P-1, the target stationary law de and the behavior stationary law db of the corresponding policy-induced joint-state transition kernels coincide: de=db. Moreover, there is a constant cimm>0, depending only on t0 and ζ, such that for every integer T≥1, the worst-case squared risk of the immediate observable estimator θ0 satisfies M∈MT(t0,ζ,1)supEM[(θ0−θ(K,e))2]≤Tcimm. Finally, for the minimax squared-error risk RT(t0,ζ,1) of Definition P-3, there are constants clo,chi>0 with clo≤chi and an integer T⋆≥1 such that, for every T≥T⋆, Tclo≤RT(t0,ζ,1)≤Tchi. The hidden-state exponent at these positive scales satisfies 2+t0ζ2<1.
Local Slice
informal · Theorem T-4 Along shrinking radii CT=1+δT, the minimax risk is comparable to the local frontier, and immediate weighting attains a parametric bound when TδT2 stays bounded.
The following conditions hold:
- (Fixed smoothness.) The constants satisfy t0>0 and ζ>0.
- (Shrinking radius.) The deterministic radius excess sequence (δT)T≥1 satisfies δT≥0 for every T.
- (Local limit.) The sequence satisfies δT→0.
- (Rate exponent.) Let β=2/(2+t0ζ).
- (Local rate and depth.) Define rT(δ)=T−1+T−βδT2(1−β) and κT=⎩⎨⎧0,min{⌊T/2⌋,⌊2log(1/α)+logLlog(TδT2)⌋},TδT2≤1,TδT2>1.
Then there exist constants cloc,Cloc with 0<cloc≤Cloc, depending only on t0 and ζ, such that, for all sufficiently large T, clocrT(δ)≤RT(t0,ζ,1+δT)≤ClocrT(δ), where RT is the minimax squared-error risk in Definition P-3. Moreover, the observable partial-history importance-weighted estimator θκT of Definition P-4 satisfies M∈MT(t0,ζ,1+δT)supEM[(θκT−θ(M))2]≤ClocrT(δ) for all sufficiently large T. Finally, if there is a constant D≥0 such that TδT2≤D for all sufficiently large T, then there is a constant cimm>0 such that M∈MT(t0,ζ,1+δT)supEM[(θ0−θ(M))2]≤Tcimm for all sufficiently large T.
Related Literature
- Hu and Wager (2023) give the closest POMDP baseline: partial-history weighting with the hidden-state exponent β=2/(2+t0ζ).
- We keep the same finite POMDP, known-policy, bounded-reward, mixing, and one-step-overlap setting.
- Our additional population control is latent stationary overlap on the full observed-latent state.
- Mehrabi and Wager (2025) provide the closest fully observed overlap comparator through stationary density-ratio control.
- Proxy and bridge approaches impose additional observability structure; our results quantify the partial-history route under latent stationary overlap.
Key Idea
- The estimator balances two forces.
- Bias shrinks because contraction makes remote latent history less relevant after a length-k observed window.
- Variance grows because each extra step multiplies another observable policy ratio.
- Latent stationary overlap scales the bias by qC.
- The calibrated depth sets the window length from the effective radius TqC2.
Lower-bound Idea
- The hard pair has one observed state, so the econometrician sees actions and rewards but not latent depth.
- Rare action histories move the latent state toward a terminal depth.
- The target value differs through the terminal sign.
- The observed laws remain close because the terminal state is hard to detect from the observed path.
- The reset-sign perturbation keeps latent stationary overlap within the prescribed radius.
Insulin Demonstration
- The finite grid has 360 joint states with latent diet coordinates.
- The behavior policy treats with probability 0.3.
- The target policy treats when glucose is at least 125.
- The refresh component gives one-half contraction.
- The certified stationary overlap constant is C=720.
- Immediate weighting has a strictly negative certified stationary bias interval.
informal · Theorem T-3 In the refreshed insulin grid, the structural constants satisfy the theorem’s contraction, action-overlap, and stationary-overlap conditions, while immediate weighting has certified negative bias.
For every horizon T, the refreshed insulin experiment Igrid of Definition P-9 satisfies:
- (State count.) Its observed--latent joint-state alphabet has cardinality 90⋅4=360.
- (Uniform contraction.) It satisfies Assumption A-6 with contraction parameter α=1/2.
- (Policy overlap.) It satisfies Assumption A-5 with one-step overlap constant L=10/3.
- (Stationary overlap.) It satisfies Assumption A-7 with stationary-overlap constant C=720.
Define the immediate-weighting bias diagnostic by Δimm(T)=s∑db(s)a∈{0,1}∑b(a∣sX)b(a∣sX)e(a∣sX)∫rK(ds′,dr∣s,a)−θ(K,e), where K,b,e are the transition kernel, behavior policy, and target policy of Igrid, sX is the observed coordinate of s, and db is the stationary law induced by b. This diagnostic is interval-certified as strictly negative: −0.006<Δimm(T)<−0.0058.
Takeaways
- Latent stationary overlap gives a quantitative radius for stationary POMDP off-policy evaluation.
- The all-radius frontier is T−1 plus the hidden-state partial-history term scaled by qC.
- Fixed positive radius recovers the Hu-Wager partial-history exponent.
- Unit overlap and sufficiently small local radii yield the parametric scale.
- Radius-calibrated PHIW is the constructive estimator matched to this frontier.